Score-Based Generative Modeling with Critically-Damped Langevin Diffusion
Tim Dockhorn, Arash Vahdat, Karsten Kreis
Introduction
Score-based generative models (SGMs) and denoising diffusion probabilistic models have emerged as a promising class of generative models (Sohl-Dickstein et al., 2015; Song et al., 2021c; b; Vahdat et al., 2021; Kingma et al., 2021). SGMs offer high quality synthesis and sample diversity, do not require adversarial objectives, and have found applications in image (Ho et al., 2020; Nichol & Dhariwal, 2021; Dhariwal & Nichol, 2021; Ho et al., 2021), speech (Chen et al., 2021; Kong et al., 2021; Jeong et al., 2021), and music synthesis (Mittal et al., 2021), image editing (Meng et al., 2021; Sinha et al., 2021; Furusawa et al., 2021), super-resolution (Saharia et al., 2021; Li et al., 2021), image-to-image translation (Sasaki et al., 2021), and 3D shape generation (Luo & Hu, 2021; Zhou et al., 2021). SGMs use a diffusion process to gradually add noise to the data, transforming a complex data distribution into an analytically tractable prior distribution. A neural network is then utilized to learn the score function—the gradient of the log probability density—of the perturbed data. The learnt scores can be used to solve a stochastic differential equation (SDE) to synthesize new samples. This corresponds to an iterative denoising process, inverting the forward diffusion.
In the seminal work by Song et al. (2021c), it has been shown that the score function that needs to be learnt by the neural network is uniquely determined by the forward diffusion process. Consequently, the complexity of the learning problem depends, other than on the data itself, only on the diffusion. Hence, the diffusion process is the key component of SGMs that needs to be revisited to further improve SGMs, for example, in terms of synthesis quality or sampling speed.
Inspired by statistical mechanics (Tuckerman, 2010), we propose a novel forward diffusion process, the critically-damped Langevin diffusion (CLD). In CLD, the data variable, (time along the diffusion), is augmented with an additional “velocity” variable and a diffusion process is run in the joint data-velocity space. Data and velocity are coupled to each other as in Hamiltonian dynamics, and noise is injected only into the velocity variable. As in Hamiltonian Monte Carlo (Duane et al., 1987; Neal, 2011), the Hamiltonian component helps to efficiently traverse the joint data-velocity space and to transform the data distribution into the prior distribution more smoothly. We derive the corresponding score matching objective and show that for CLD the neural network is tasked with learning only the score of the conditional distribution of velocity given data , which is arguably easier than learning the score of diffused data directly. Using techniques from molecular dynamics (Bussi & Parrinello, 2007; Tuckerman, 2010; Leimkuhler & Matthews, 2013), we also derive a new SDE integrator tailored to CLD’s reverse-time synthesis SDE.
We extensively validate CLD and the novel SDE solver: (i) We show that the neural networks learnt in CLD-based SGMs are smoother than those of previous SGMs. (ii) On the CIFAR-10 image modeling benchmark, we demonstrate that CLD-based models outperform previous diffusion models in synthesis quality for similar network architectures and sampling compute budgets. We attribute these positive results to the Hamiltonian component in the diffusion and to CLD’s easier score function target, the score of the velocity-data conditional distribution . (iii) We show that our novel sampling scheme for CLD significantly outperforms the popular Euler–Maruyama method. (iv) We perform ablations on various aspects of CLD and find that CLD does not have difficult-to-tune hyperparameters.
In summary, we make the following technical contributions: (i) We propose CLD, a novel diffusion process for SGMs. (ii) We derive a score matching objective for CLD, which requires only the conditional distribution of velocity given data. (iii) We propose a new type of denoising score matching ideally suited for scalable training of CLD-based SGMs. (iv) We derive a tailored SDE integrator that enables efficient sampling from CLD-based models. (v) Overall, we provide novel insights into SGMs and point out important new connections to statistical mechanics.
Background
where is the score function of the marginal distribution over at time .
If and are affine, the conditional distribution is Normal and available analytically (Särkkä & Solin, 2019). Different result in different trade-offs between synthesis quality and likelihood in the generative model defined by (Song et al., 2021b; Vahdat et al., 2021).
Critically-Damped Langevin Diffusion
where denotes the Kronecker product. The coupled SDE that describes the diffusion process is
There is a crucial balance between and (McCall, 2010): For (underdamped Langevin dynamics) the Hamiltonian component dominates, which implies oscillatory dynamics of and that slow down convergence to equilibrium. For (overdamped Langevin dynamics) the -term dominates which also slows down convergence, since the accelerating effect by the Hamiltonian component is suppressed due to the strong noise injection. For (critical damping), an ideal balance is achieved and convergence to occurs as fast as possible in a smooth manner without oscillations (also see discussion in App. A.1) (McCall, 2010). Hence, we propose to set and call the resulting diffusion critically-damped Langevin diffusion (CLD) (see Fig. 1).
Diffusions such as the VPSDE (Song et al., 2021c) correspond to overdamped Langevin dynamics with high friction coefficients (see App. A.2). Furthermore, in previous works noise is injected directly into the data variables (pixels, for images). In CLD, only the velocity variables are subject to direct noise and the data is perturbed only indirectly due to the coupling between and .
Notice that this objective requires only the velocity gradient of the log-density of the joint distribution, i.e., . This is a direct consequence of injecting noise into the velocity variables only. Without loss of generality, . Hence,
This means that in CLD the neural network-defined score model only needs to learn the score of the conditional distribution , an arguably easier task than learning the score of , as in previous works, or of the joint . This is the case, because our velocity distribution is initialized from a simple Normal distribution, such that is closer to a Normal distribution for all (and for any ) than itself. This is most evident at : The data and velocity distributions are independent at and the score of simply corresponds to the score of the Normal distribution from which the velocities are initialized, whereas the score of the data distribution is highly complex and can even be unbounded (Kim et al., 2021). We empirically verify the reduced complexity of the score of in Fig. 2. We find that the score that needs to be learnt by the model is more similar to a score corresponding to a Normal distribution for CLD than for the VPSDE. We also measure the complexity of the neural networks that were learnt to model this score via the squared Frobenius norm of their Jacobians. We find that the CLD-based SGMs have significantly simpler and smoother neural networks than VPSDE-based SGMs for most , in particular when leveraging a mixed score formulation (see next section).
2 Scalable Training
In HSM, the expectation over is essentially solved analytically, while DSM would use a sample-based estimate. Hence, HSM reduces the variance of training objective gradients compared to pure DSM, which we validate in App. C.1. Furthermore, when drawing a sample to diffuse in DSM, we are essentially placing an infinitely sharp Normal with unbounded score (Kim et al., 2021) at , which requires undesirable modifications or truncation tricks for stable training (Song et al., 2021c; Vahdat et al., 2021). Hence, with DSM we could lose some benefits of the CLD framework discussed in Sec. 3.1, whereas HSM is tailored to CLD and fundamentally avoids such unbounded scores. Closed form expressions for the perturbation kernel are provided in App. B.1.
(ii) Vahdat et al. (2021) showed that it can be beneficial to assume that the diffused marginal distribution is Normal at all times and parametrize the model with a Normal score and a residual “correction”. For CLD, the score is indeed Normal at (due to the independently initialized and at ). Similarly, the target score is close to Normal for large , as we approach the equilibrium.
which corresponds to training the model to predict the noise only injected into the velocity during reparametrized sampling of , similar to noise prediction in Ho et al. (2020); Song et al. (2021c).
3 Sampling from CLD-based SGMs
To sample from the CLD-based SGM we can either directly simulate the reverse-time diffusion process (Eq. (2)) or, alternatively, solve the corresponding probability flow ODE (Song et al., 2021c; b) (see App. B.5). To simulate the SDE of the reverse-time diffusion process, previous works often relied on Euler-Maruyama (EM) (Kloeden & Platen, 1992) and related methods (Ho et al., 2020; Song et al., 2021c; Jolicoeur-Martineau et al., 2021a). We derive a new solver, tailored to CLD-based models. Here, we provide the high-level ideas and derivations (see App. D for details).
Our generative SDE can be written as (with , , ):
It consists of a Hamiltonian component , an Ornstein-Uhlenbeck process , and the score model term . We could use EM to integrate this SDE; however, standard Euler methods are not well-suited for Hamiltonian dynamics (Leimkuhler & Reich, 2005; Neal, 2011). Furthermore, if was , we could solve the SDE in closed form. This suggests the construction of a novel integrator.
We use the Fokker-Planck operatorThe Fokker-Planck operator is also known as Kolmogorov operator. If the underlying dynamics is fully Hamiltonian, it corresponds to the Liouville operator (Leimkuhler & Matthews, 2015; Tuckerman, 2010). formalism (Tuckerman, 2010; Leimkuhler & Matthews, 2013; 2015). Using a similar notation as Leimkuhler & Matthews (2013), the Fokker-Planck equation corresponding to the generative SDE is , where and are the non-commuting Fokker-Planck operators corresponding to the and terms, respectively. Expressions for and can be found in App. D. We can construct a formal, but intractable solution of the generative SDE as , where the operator (known as the classical propagator in statistical physics) propagates states for time according to the dynamics defined by the combined operators . Although this operation is not analytically tractable, it can serve as starting point to derive a practical integrator. Using the symmetric Trotter theorem or Strang splitting formula as well as the Baker–Campbell–Hausdorff formula (Trotter, 1959; Strang, 1968; Tuckerman, 2010), it can be shown that:
Related Work
Relations to Statistical Mechanics and Molecular Dynamics. Learning a mapping between a simple, tractable and a complex distribution as in SGMs is inspired by annealed importance sampling (Neal, 2001) and the Jarzynski equality from non-equilibrium statistical mechanics (Jarzynski, 1997a; b; 2011; Bahri et al., 2020). However, after Sohl-Dickstein et al. (2015), little attention has been given to the origins of SGMs in statistical mechanics. Intuitively, in SGMs the diffusion process is initialized in a non-equilibrium state and we would like to bring the system to equilibrium, i.e., the tractable prior distribution, as quickly and as smoothly as possible to enable efficient denoising. This “equilibration problem” is a much-studied problem in statistical mechanics, particularly in molecular dynamics, where a molecular system is often simulated in thermodynamic equilibrium. Algorithms to quickly and smoothly bring a system to and maintain at equilibrium are known as thermostats. In fact, CLD is inspired by the Langevin thermostat (Bussi & Parrinello, 2007). In molecular dynamics, advanced thermostats are required in particular for “multiscale” systems that show complex behaviors over multiple time- and length-scales. Similar challenges also arise when modeling complex data, such as natural images. Hence, the vast literature on thermostats (Andersen, 1980; Nosé, 1984; Hoover, 1985; Martyna et al., 1992; Hünenberger, 2005; Bussi et al., 2007; Ceriotti et al., 2009; 2010; Tuckerman, 2010) may be valuable for the development of future SGMs. Also the framework for developing SSCS is borrowed from statistical mechanics. The same techniques have been used to derive molecular dynamics algorithms (Tuckerman et al., 1992; Bussi & Parrinello, 2007; Ceriotti et al., 2010; Leimkuhler & Matthews, 2013; 2015; Kreis et al., 2017).
Further Related Work. Generative modeling by learning stochastic processes has a long history (Movellan, 2008; Lyu, 2009; Sohl-Dickstein et al., 2011; Bengio et al., 2014; Alain et al., 2016; Goyal et al., 2017; Bordes et al., 2017; Song & Ermon, 2019; Ho et al., 2020). We build on Song et al. (2021c), which introduced the SDE framework for modern SGMs. Nachmani et al. (2021) recently introduced non-Gaussian diffusion processes with different noise distributions. However, the noise is still injected directly into the data, and no improved sampling schemes or training objectives are introduced. Vahdat et al. (2021) proposed LSGM, which is complementary to CLD: we improve the diffusion process itself, whereas LSGM “simplifies the data” by first embedding it into a smooth latent space. LSGM is an overall more complicated framework, as it is trained in two stages and relies on additional encoder and decoder networks. Recently, techniques to accelerate sampling from pre-trained SGMs have been proposed (San-Roman et al., 2021; Watson et al., 2021; Kong & Ping, 2021; Song et al., 2021a). Importantly, these methods usually do not permit straightforward log-likelihood estimation. Furthermore, they are originally not based on the continuous time framework, which we use, and have been developed primarily for discrete-step diffusion models.
A complementary work to CLD is “Gotta Go Fast” (GGF) (Jolicoeur-Martineau et al., 2021a), which introduces an adaptive SDE solver for SGMs, tuned towards image synthesis. GGF uses standard Euler-based methods under the hood (Kloeden & Platen, 1992; Roberts, 2012), in contrast to our SSCS that is derived from first principles. Furthermore, our SDE integrator for CLD does not make any data-specific assumptions and performs extremely well even without adaptive step sizes.
Some works study SGMs for maximum likelihood training (Song et al., 2021b; Kingma et al., 2021; Huang et al., 2021). Note that we did not focus on training our models towards high likelihood. Furthermore, Chen et al. (2020) and Huang et al. (2020) recently trained augmented Normalizing Flows, which have conceptual similarities with our velocity augmentation. Methods leveraging auxiliary variables similar to our velocities are also used in statistics—such as Hamiltonian Monte Carlo (Neal, 2011)—and have found applications, for instance, in Bayesian machine learning (Chen et al., 2014; Ding et al., 2014; Shang et al., 2015). As shown in Ma et al. (2019), our velocity is equivalent to momentum in gradient descent and related methods (Polyak, 1964; Kingma & Ba, 2015). Momentum accelerates optimization; the velocity in CLD accelerates mixing in the diffusion process. Lastly, our CLD method can be considered as a second-order Langevin algorithm, but even higher-order schemes are possible (Mou et al., 2021) and could potentially further improve SGMs.
Experiments
Architectures. We focus on image synthesis and implement CLD-based SGMs using NCSN++ and DDPM++ (Song et al., 2021c) with 6 input channels (for velocity and data) instead of 3.
Sampling. We generate model samples via: (i) Probability flow using a Runge–Kutta 4(5) method; reverse-time generative SDE sampling using either (ii) EM or (iii) our SSCS. For methods without adaptive stepsize (EM and SSCS), we use evaluation times chosen according to a quadratic function, like previous work (Song et al., 2021a; Kong & Ping, 2021; Watson et al., 2021) (indicated by QS).
Following Song et al. (2021c), we focus on the widely used CIFAR-10 unconditional image generation benchmark. Our CLD-based SGM achieves an FID of 2.25 based on the probability flow ODE and an FID of 2.23 via simulating the generative SDE (Tab. 1). The only models marginally outperforming CLD are LSGM (Vahdat et al., 2021) and NSCN++/VESDE with 2,000 step predictor-corrector (PC) sampling (Song et al., 2021c). However, LSGM uses a model with parameters to achieve its high performance, while we obtain our numbers with a model of parameters. For a fairer comparison, we trained a smaller LSGM also with parameters, which is reported as “LSGM-100M” in Tab. 1 (details in App. E.2.7). Our model has a significantly better FID score than “LSGM-100M”. In contrast to NSCN++/VESDE, we achieve extremely strong results with much fewer NFEs (for example, see in Tab. 3 and also Tab. 3)—the VESDE performs poorly in this regime. We conclude that when comparing models with similar network capacity and under NFE budgets , our CLD-SGM outperforms all published results in terms of FID. We attribute these positive results to our easier score matching task. Furthermore, our model reaches an NLL bound of , which is on par with recent works such as Nichol & Dhariwal (2021); Austin et al. (2021); Vahdat et al. (2021) and indicates that our model is not dropping modes. However, our bound is potentially quite loose (see discussion in App. B.5) and the true NLL might be significantly lower. We did not focus on training our models towards high likelihood.
To demonstrate that CLD is also suitable for high-resolution image synthesis, we additionally trained a CLD-SGM on CelebA-HQ-256, but without careful hyperparameter tuning due to limited compute resources. Model samples in Fig. 4 appear diverse and high-quality (additional samples in App. F).
2 Sampling Speed and Synthesis Quality Trade-Offs
We analyze the sampling speed vs. synthesis quality trade-off for CLD-SGMs and study SSCS’s performance under different NFE budgets (Tabs. 3 and 3). We compare to Song et al. (2021c) and use EM to solve the generative SDE for their VPSDE and PC (reverse-diffusion + Langevin sampler) for the VESDE model. We also compare to the GGF (Jolicoeur-Martineau et al., 2021a) solver for the generative SDE as well as probability flow ODE sampling with a higher-order adaptive step size solver. Further, we compare to LSGM (Vahdat et al., 2021) (using our LSGM-100M), which also uses probability flow sampling. With one exception (VESDE with 2,000 NFE) our CLD-SGM outperforms all baselines, both for adaptive and fixed-step size methods. More results in App. F.2.
Several observations stand out: (i) As expected (Sec. 3.3), for CLD, SSCS significantly outperforms EM under limited NFE budgets. When using a fine discretization of the SDE (high NFE), the two perform similarly, which is also expected, as the errors of both methods will become negligible. (ii) In the adaptive solver setting, using a simpler ODE solver, we even outperform GGF, which is tuned towards image synthesis. (iii) Our CLD-SGM also outperforms the LSGM-100M model in terms of FID. It is worth noting, however, that LSGM was designed primarily for faster synthesis, which it achieves by modeling a smooth distribution in latent space instead of the more complex data distribution directly. This suggests that it would be promising to combine LSGM with CLD and train a CLD-based LSGM, combining the strengths of the two approaches. It would also be interesting to develop a more advanced, adaptive SDE solver that leverages SSCS as the backbone and, for example, potentially test our method within a framework like GGF. Our current SSCS only allows for fixed step sizes—nevertheless, it achieves excellent performance.
3 Ablation Studies
We perform ablation studies to study CLD’s new hyperparameters (run with a smaller version of our CLD-SGM used above; App. E for details).
Mass Parameter: Tab. 5 shows results for a CLD-SGM trained with different (also recall that and are tied together via ; we are always in the critical-damping regime). Different mass values perform mostly similarly. Intuitively, training with smaller means that noise flows from the velocity variables into the data more slowly, which necessitates a larger time rescaling . We found that simply tying and together via works well and did not further fine-tune.
Initial Velocity Distribution: Tab. 5 shows results for a CLD-SGM trained with different initial velocity variance scalings . Varying similarly has only a small effect, but small seems slightly beneficial for FID, while the NLL bound suffers a bit. Due to our focus on synthesis quality as measued by FID, we opted for small . Intuitively, this means that the data at is “at rest”, and noise flows from the velocity into the data variables only slowly.
Mixed Score: Similar to previous work (Vahdat et al., 2021), we find training with the mixed score (MS) parametrization (Sec. 3.2) beneficial. With MS, we achieve an FID of 3.14, without only 3.56.
Hybrid Score Matching: We also tried training with regular DSM, instead of HSM. However, training often became unstable. As discussed in Sec. 3.2, this is likely because when using standard DSM our CLD would suffer from unbounded scores close to , similar to previous SDEs (Kim et al., 2021). Consequently, we consider our novel HSM a crucial element for training CLD-SGMs.
We conclude that CLD does not come with difficult-to-tune hyperparameters. We expect our chosen hyperparameters to immediately translate to new tasks and models. In fact, we used the same , , MS and HSM settings for CIFAR-10 and CelebA-HQ-256 experiments without fine-tuning.
Conclusions
We presented critically-damped Langevin diffusion, a novel diffusion process for training SGMs. CLD diffuses the data in a smoother, easier-to-denoise manner compared to previous SGMs, which results in smoother neural network-parametrized score functions, fast synthesis, and improved expressivity. Our experiments show that CLD outperforms previous SGMs on image synthesis for similar-capacity models and sampling compute budgets, while our novel SSCS is superior to EM in CLD-based SGMs. From a technical perspective, in addition to proposing CLD, we derive CLD’s score matching objective termed as HSM, a variant of denoising score matching suited for CLD, and we derive a tailored SDE integrator for CLD. Inspired by methods used in statistical mechanics, our work provides new insights into SGMs and implies promising directions for future research.
We believe that CLD can potentially serve as the backbone diffusion process of next generation SGMs. Future work includes using CLD-based SGMs for generative modeling tasks beyond images, combining CLD with techniques for accelerated sampling from SGMs, adapting CLD-based SGMs towards maximum likelihood, and utilizing other thermostating methods from statistical mechanics.
Ethics and Reproducibility
Our paper focuses on fundamental algorithmic advances to improve the generative modeling performance of SGMs. As such, the proposed CLD does not imply immediate ethical concerns. However, we validate CLD on image synthesis benchmarks. Generative modeling of images has promising applications, for example for digital content creation and artistic expression (Bailey, 2020), but can also be used for nefarious purposes (Vaccari & Chadwick, 2020; Mirsky & Lee, 2021; Nguyen et al., 2021). It is worth mentioning that compared to generative adversarial networks (Goodfellow et al., 2014), a very popular class of generative models, SGMs have the promise to model the data more faithfully, without dropping modes and introducing problematic biases. Generally, the ethical impact of our work depends on its application domain and the task at hand.
To aid reproducibility of the results and methods presented in our paper, we made source code to reproduce the main results of the paper publicly available, including detailed instructions; see our project page https://nv-tlabs.github.io/CLD-SGM and the code repository https://github.com/nv-tlabs/CLD-SGM. Furthermore, all training details and hyperparameters are already in detail described in the Appendix, in particular in App. E.
References
Appendix A Langevin Dynamics
Here, we discuss different aspects of Langevin dynamics. Recall the Langevin dynamics, Eq. (5), from the main paper:
As discussed in Sec. 3, Langevin dynamics can be run with different ratios between mass and squared friction . To recap from the main paper:
(i) For (underdamped Langevin dynamics), the Hamiltonian component dominates, which implies oscillatory dynamics of and that slow down convergence to equilibrium.
(ii) For (overdamped Langevin dynamics), the -term dominates which also slows down convergence, since the accelerating effect by the Hamiltonian component is suppressed due to the strong noise injection.
(iii) For (critical-damping), an ideal balance is achieved and convergence to occurs quickly in a smooth manner without oscillations.
In Fig. 5, we visualize diffusion trajectories according to Langevin dynamics run in the different damping regimes. We observe that underdamped Langevin dynamics show undesired oscillatory behavior, while overdamped Langevin dynamics perform very inefficiently, too. Critical-damping achieves a good balance between the two and mixes and converges quickly. In fact, it can be shown to be optimal in terms of convergence; see, for example, McCall (2010).
Consequently, we propose to set in CLD.
A.2 Very High Friction Limit and Connections to previous SDEs in SGMs
Let us re-write the above Langevin dynamics and consider the more general case with time-dependent :
To solve this SDE, let us assume a simple Euler-based integration scheme, with the update equation for a single step at time (this integration scheme would not be optimal, as discussed in Sec. 3.3., however, it would be accurate for sufficiently small time steps and we just need this to make the connection to previous works like the VPSDE):
Now, let us assume a friction coefficient . Since the time step is usually very small, this correspond to a very high friction. In fact, it can be considered the maximum friction limit, at which the friction is so large that the current step velocity (i) is completely cancelled out by the friction term (iii). We obtain:
Now the velocity update, Eq. (17), does not depend on the current step velocity on the right-hand-side anymore. Hence, we can insert Eq. (17) directly into Eq. (16) and obtain:
Re-defining and , we obtain
which corresponds to the high-friction overdamped Langevin dynamics that are frequently run, for example, to train energy-based generative models (Du & Mordatch, 2019; Xiao et al., 2021). Let’s further absorb the mass and the time step into the time rescaling, defining . We obtain:
where the last approximation is true for sufficiently small . However, this expression corresponds to
which is exactly the transition kernel of the VPSDE’s Markov chain (Ho et al., 2020; Song et al., 2021c). We see that the VPSDE corresponds to the high-friction limit of a more general Langevin dynamics-based diffusion process of the form of Eq. (11).
If we assume a diffusion as above but with the potential term (ii) set to , we can similarly derive the VESDE Song et al. (2021c) as a high-friction limit of the corresponding diffusion. Generally, all previously used diffusions that inject noise directly into the data variables correspond to such high-friction diffusions.
In conclusion, we see that previous high-friction diffusions require an excessive amount of noise to be injected to bring the dynamics to the prior, which intuitively makes denoising harder. For our CLD in the critical damping regime we can run the diffusion for a much shorter time or, equivalently, can inject less noise to converge to the equilibrium, i.e., the prior.
Appendix B Critically-Damped Langevin Diffusion
Here, we present further details about our proposed critically-damped Langevin diffusion (CLD). We provide the derivations and formulas that were not presented in the main paper in the interest of brevity.
Following Särkkä & Solin (2019) (Section 6.1), the mean and covariance matrix of obey the following respective ordinary differential equations (ODEs)
Notating , the solutions to the above ODEs are
where . For constant (as is used in all our experiments), we simply have . The correctness of the proposed mean and covariance matrix can be verified by simply plugging them back into their respective ODEs; see App. G.1.
With the above derivations, we can find analytical expressions for the perturbation kernel . For example, when conditioning on initial data and velocity samples and (as in denoising score matching (DSM)), the mean and covariance matrix of the perturbation kernel can be obtained by setting , , and .
In our experiments, the initial velocity distribution is set to . Conditioning only on initial data samples and marginalizing over the full initial velocity distribution (as in our hybrid score matching (HSM), see Sec. C), the mean and covariance matrix of the perturbation kernel can be obtained by setting , , and .
B.2 Convergence and Equilibrium
Our CLD-based training of SGMs—as well as denoising diffusion models more generally—relies on the fact that the diffusion converges towards an analytically tractable equilibrium distribution for sufficiently large . In fact, from the above equations we can easily see that,
which establishes .
Notice that our CLD is an instantiation of the more general Langevin dynamics defined by
which has the equilibrium distribution (Leimkuhler & Matthews, 2015; Tuckerman, 2010). However, the perturbation kernel of this Langevin dynamics is not available analytically anymore for arbitrary . In our case, however, we have the analytically tractable . Note that this corresponds to the classical “harmonic oscillator” problem from physics.
B.3 CLD Objective
To derive the objective for training CLD-based SGMs, we start with a derivation that targets maximum likelihood training in a similar fashion to Song et al. (2021b). Let and be two densities, then
where and are the marginal densities of and , respectively, diffused by our critically-damped Langevin diffusion. As has been shown in Song et al. (2021b), Eq. (39) can be written as a mixture (over ) of score matching losses. To this end, let us consider the Fokker–Planck equation associated with the critically-damped Langevin diffusion:
Similarly, we have . Assuming and are smooth functions with at most polynomial growth at infinity, we have
Using the above fact, we can compute the time-derivative of the Kullback–Leibler divergence between and as
Notice that due to the form of , we now have only gradients with respect to the velocity component . Combining the above with Eq. (39), we have
Note that the approximation holds if is sufficiently “close” to . We obtain a more general objective function by replacing with an arbitrary function , i.e.,
As shown in App. C, the above can be rewritten, up to irrelevant constant terms, as either of the following two objectives:
For both HSM and DSM, we have shown in App. B.1 that the perturbation kernels and are Normal distributions with the following structure of the covariance matrix:
We can use this fact to compute the gradient
where and is the Cholesky factorization of the covariance matrix . Note that the structure of implies that , where is the Cholesky factorization of , i.e,
and denotes those (latter) components of that actually affect .
where is sampled via reparameterization:
Note again that is different for HSM and DSM.
Analogously to prior work (Ho et al., 2020; Vahdat et al., 2021; Song et al., 2021b) an objective better suited for high quality image synthesis can be obtained by “dropping the variance prefactor”:
B.4 CLD-specific Implementation Details
B.5 Lower Bounds and Probability Flow ODE
Given the score model , we can synthesize novel samples via simulating the reverse-time diffusion SDE, Eq. (2) in the main text. This can be achieved, for example, via our novel SSCS, Euler-Maruyama, or methods such as GGF (Jolicoeur-Martineau et al., 2021a). However, Song et al. (2021b; c) have shown that a corresponding ordinary differential equation can be defined that generates samples from the same distribution, in case models the ground truth scores perfectly. This ODE is:
This ODE is often referred to as the probability flow ODE. We can use it to generate novel data by sampling the prior and solving this ODE, like previous works (Song et al., 2021c). Note that in practice won’t be a perfect model, though, such that the generative models defined by simulating the reverse-time SDE and the probability flow ODE are not exactly equivalent (Song et al., 2021b). Nevertheless, they are very closely connected and it has been shown that their performance is usually very similar or almost the same, when we have learnt a good . In addition to sampling the generative SDE in our paper, we also sample from our CLD-based SGMs via this probability flow approach.
With the definition of our CLD, the ODE becomes:
Notice the interesting form of this probability flow ODE for CLD: It corresponds to Hamiltonian dynamics () plus the score function term . Compared to the generative SDE (Sec. 3.3), the Ornstein-Uhlenbeck term disappears. Generally, symplectic integrators are best suited for integrating Hamiltonian systems (Neal, 2011; Tuckerman, 2010; Leimkuhler & Reich, 2005). However, our ODE is not perfectly Hamiltonian, due to the score term, and modern non-symplectic methods, such as the higher-order adaptive-step size Runge-Kutta 4(5) ODE integrator (Dormand & Prince, 1980), which we use in practice to solve the probability flow ODE, can also accurately simulate Hamiltonian systems over limited time horizons.
Importantly, the ODE formulation also allows us to estimate the log-likelihood of given test data, as it essentially defines a continuous Normalizing flow (Chen et al., 2018; Grathwohl et al., 2019), that we can easily run in either direction. However, in CLD the input into this ODE is not just the data , but also the velocity variable . In this case, we can still calculate a lower bound on the log-likelihood:
where denotes the entropy of (we have ). We can obtain a stochastic, but unbiased estimate of via solving the probability flow ODE with initial conditions and calculating a stochastic estimate of the log-determinant of the Jacobian via Hutchinson’s trace estimator (and also calculating the probability of the output under the prior), as done in Normalizing flows (Chen et al., 2018; Grathwohl et al., 2019) and previous works on SGMs (Song et al., 2021c; b). In the main paper, we report the negative of Eq. (60) as our upper bound on the negative log-likelihood (NLL).
Note that this bound can be potentially quite loose. In principle, it would be desirable to perform an importance-weighted estimation of the log-likelihood, as in importance-weighted autoencoders (Burda et al., 2015), using multiple samples from the velocity distribution. However, this isn’t possible, as we only have access to a stochastic estimate . The problems arising from this are discussed in detail in Appendix F of Vahdat et al. (2021). We could consider training a velocity encoder network, somewhat similar to Chen et al. (2020), to improve our bound, but we leave this for future research.
B.6 On Introducing a Hamiltonian Component into the Diffusion
Here, we provide additional high-level intuitions and motivations about adding the Hamiltonian component to the diffusion process, as is done in our CLD.
Let us recall how the data distribution evolves in the forward diffusion process of SGMs: The role of the diffusion is to bring the initial non-equilibrium state quickly towards the equilibrium or prior distribution. Suppose for a moment, we could do so with “pure” Hamiltonian dynamics (no noise injection). In that case, we could generate data from the backward model without learning a score or neural network at all, because Hamiltonian dynamics is analytically invertible (flipping the sign of the velocity, we can just integrate backwards in reverse time direction). Now, this is not possible in practice, since Hamiltonian dynamics alone usually cannot convert the non-equilibrium distribution to the prior distribution. Nevertheless, Hamiltonian dynamics essentially achieves a certain amount of mixing on its own; moreover, since it is deterministic and analytically invertible, this mixing comes at no cost in the sense that we do not have to learn a complex score function to invert the Hamiltonian dynamics. Our thought experiment shows that we should strive for a diffusion process that behaves as deterministically (meaning that deterministic implies easily invertible) as possible with as little noise injection as possible. And this is exactly what is achieved by adding the Hamiltonian component in the overall diffusion process. In fact, recall that it is the diffusion coefficient of the forward SDE that ultimately scales the score function term of the backward generative SDE (and it is the score function that is hard to approximate with complex neural nets). Therefore, in other words, relying more on a deterministic Hamiltonian component for enhanced mixing (mixing just like in MCMC in that it brings us quickly towards the target distribution, in our case the prior) and less on pure noise injection will lead to a nicer generative SDE that relies less on a score function that requires complex and approximate neural network-based modeling, but more on a simple and analytical Hamiltonian component. Such an SDE could then be solved easier with an appropriate integrator (like our SSCS). In the end, we believe that this is the reason why our networks are “smoother” and why given the same network capacity and limited compute budgets we essentially outperform all previous results in the literature (on CIFAR-10).
We would also like to offer a second perspective, inspired by the Markov chain Monte Carlo (MCMC) literature. In MCMC, “mixing” helps to quickly traverse the high probability parts of the target distribution and, if an MCMC chain is initialized far from the high probability manifold, to quickly converge to this manifold. However, this is precisely the situation we are in with the forward diffusion process of SGMs: The system is initialized in a far-from-equilibrium state (the data distribution) and we need to traverse the space as efficiently as possible to converge to the equilibrium distribution, this is, the prior. Without efficient mixing, it takes longer to converge to the prior, which also implies a longer generation path in the reverse direction—which intuitively corresponds to a harder problem. Therefore, we believe that ideas from the MCMC literature that accelerate mixing and traversal of state space may be beneficial also for the diffusions in SGMs. In fact, leveraging Hamiltonian dynamics to accelerate sampling is popular in the MCMC field (Neal, 2011). Note that this line of reasoning extends to thermostating techniques from statistical mechanics and molecular dynamics, which essentially tackle similar problems like MCMC methods from the statistics literature (see discussion in Sec. 4).
Appendix C HSM: Hybrid Score Matching
We begin by recalling our objective function from App. B.3 (Eq. (44)):
where is our score model. In the following, we dissect the “score matching” part of the above objective:
Instead, we can “denoise” only with the data distribution and marginalize over the entire initial velocity distribution , which results in
Using this relation, we can also find a connection between our CLD objective functions from App. B.3. In particular, we have
where we applied the chain-rule. Note that in Eq. (71) and Eq. (72), is sampled from the same distribution. Hence, acts as a common scaling factor, with the variance difference between HSM and DSM originating from the squared norm term. Hence, we ignore and only focus our analysis on the gradient of the norm terms, which we can further simplify:
As is common practice in statistics, we consider only the trace of the estimated covariance matrices.Arguably, the most prominent algorithm that follows this practice is principal component analysis (PCA). The trace of the covariance matrix (of a random variable) is also commonly referred to as the total variation (of a random variable).
Appendix D Symmetric Splitting CLD Sampler (SSCS)
In this section, we present a more complete derivation and analysis of our novel Symmetric Splitting CLD Sampler (SSCS).
Our derivation is inspired by methods from the statistical mechanics and molecular dynamics literature. In particular, we are leveraging symmetric splitting techniques as well as (Fokker–Planck) operator concepts. The high-level idea of symmetric splitting as well as the operator formalism are well-explained in Tuckerman (2010), in particular in their Section 3.10, which includes simple examples. Symmetric splitting methods for stochastic dynamics in particular are discussed in detail in Leimkuhler & Matthews (2015). We also recommend Leimkuhler & Matthews (2013), which discusses splitting methods for Langevin dynamics in a concise but insightful manner.
D.2 Derivation and Analysis
Generative SDE. From Sec. 3.3, recall that our generative SDE can be written as (with , , ):
Fokker–Planck Equation and Fokker–Planck Operators. The evolution of the probability distribution is described by the general Fokker–Planck equation (Särkkä & Solin, 2019):
For our SDE, we can write the Fokker–Planck equation in short form as
with the Fokker–Planck operators (defined via their action on functions of the variables ):
We are providing these formulas for transparency and completeness. We do not directly leverage them. However, working with these operators can be convenient. In particular, the operators describe the time evolution of states under the stochastic dynamics defined by the SDE. Given an initial state , we can construct a formal solution to the generative SDE via (Tuckerman, 2010; Leimkuhler & Matthews, 2015):
where the operator is known as the classical propagator that propagates states for time according to the dynamics defined by the combined Fokker–Planck operators (to avoid confusion, note that in Eq. (83) the operator is applied on in an element-wise or “vectorized” fashion on all elements of in parallel). The problem with that expression is that we cannot analytically evaluate it. However, we can leverage it to design an integration method.
Symmetric Splitting Integration. Using the symmetric Trotter theorem or Strang splitting formula as well as the Baker–Campbell–Hausdorff formula (Trotter, 1959; Strang, 1968; Tuckerman, 2010), it can be shown that:
Analyzing the Splitting Terms. Next, we need to analyze the two individual terms:
(i) Let us first analyze : This term describes the stochastic evolution of under the dynamics of an SDE like Eq. (75), but with set to zero. However, if is set to zero, the remaining SDE has affine drift and diffusion coefficients. In that case, if the input is Normal (or a discrete state corresponding to a Normal with variance) then the distribution is Normal at all times and we can calculate the evolution analytically. In particular, we can solve the differential equations for the mean and covariance of the Normal (see Sec. B.1), and obtain
The correctness of the proposed mean and covariance matrix can be verified by simply plugging them back in their respective ODEs; see App. G.2. Now, we can write the action of the the propagator on a state as:
(ii): Next, we need to analyze . Unfortunately, we cannot calculate the action of the propagator on analytically and we need to make an approximation. From Eq. (75), we can easily see that the propagator describes the evolution of the velocity component for time step under the ODE (this can be easily seen by noticing that the term in Eq. (75) only acts on the velocity component of the joint state ):
We propose to simply solve this ODE for the step via a simple step of Euler’s method, resulting in:
Error Analysis. It is now instructive to study the overall error of our proposed integrator. With the additional Euler integration in one of the splitting terms, we have
where we used and only kept the dominating error terms of lowest order in . We see that, just like Euler’s method, also our SSCS is a first-order integrator with local error and global error , which can be also seen from the last two lines of Eq. (95). This is expected, considering that we used an Euler step for the term. Nevertheless, as long as the dynamics is not dominated by the component, our proposed integration scheme is still expected to be more accurate than EM, since we split off the analytically tractable part and only use an Euler approximation for the term.
For large parts of the diffusion, is indeed close to , such that the term is very small (this cancellation is the reason why we pulled the term into the component). In Sec. 3, we have also seen that our neural network component can be much smoother than that of previous SGMs. Overall, this suggests that the error of SSCS indeed might be smaller than the error we would obtain when applying a naive Euler–Maruyama integrator to the full generative SDE. Our positive experimental results in Sec. 5.2 validate that. Only in the limit for very small steps, both our SSCS and EM make only very small errors and are expected to perform equally well, which is exactly what we observe in our experiments. Our SSCS turns out to be well-suited for integrating the generative SDE of CLD-SGMs with relatively few synthesis steps.
Note that error analysis of stochastic differential equation solvers is usually performed in terms of weak and strong convergence (Kloeden & Platen, 1992). Due to the use of Euler’s method for the component, as argued above, we expect our SSCS to formally have the same weak and strong convergence properties like EM, this is, weak convergence of order and strong convergence of order as well, since the noise is additive in our case (and assuming appropriate smoothness conditions for the drift and diffusion coefficients; furthermore, without additive noise, we would have strong convergence of order ). We leave a more detailed analysis to future work.
In practice, we do not use SSCS to integrate all the way from to , but only up to , and perform a denoising step, similar to previous works (Jolicoeur-Martineau et al., 2021a; Song et al., 2021c). It is worth noting that our SSCS scheme would also be applicable when we used time-dependent , as in our more general derivation of the CLD perturbation kernel in App. B. However, since we only used constant in the main paper, we also presented SSCS in that way.
A promising direction for future work would be to extend SSCS to adaptive step sizes and to use techniques to facilitate higher-order integration, while still leveraging the advantages of SSCS.
SSCS Algorithm. Finally, we summarize SSCS in terms of a concise algorithm:
Note that the algorithm uses the expressions in Eqs. (85) and (86) for and . Furthermore, in practice in the denoising step at the end, we usually only update the component of , since we are only interested in the data sample. This saves us the final neural network call during denoising, which only affects the component (also see App. E.2.4). However, we wrote the algorithm in the general way, which also allows to correctly generate the velocity sample . In Fig. 8, we show a conceptual visualization of our SSCS and contrast it to EM.
Also note that we could combine the second half-step from one iteration of SSCS with the first half-step from the next iteration of SSCS. This is commonly done in the Leapfrog integrator (Leimkuhler & Reich, 2005; Tuckerman, 2010; Neal, 2011; Leimkuhler & Matthews, 2015),The Leapfrog integrator corresponds to the velocity Verlet integrator in molecular dynamics. which follows a similar structure as our SSCS. However, it is not important in our case, as the only computationally costly operation is in the center full step, which involves the neural network evaluation. The first and last half-steps come at virtually no computational cost.
Appendix E Implementation and Experiment Details
In this section, we provide details for the experiments presented in Sec. 3.1. For both experiments, we consider a two-dimensional simple mixture of Normals of the form
where and
We empirically verify the reduced complexity of the score of , which is learned in CLD, compared to the score of , which is learned in VPSDE. To avoid scaling issues between VPSDE and CLD, we chose for CLD in this experiment; this results in an equilibrium distribution of (for both data and velocity components, which are independent at equilibrium), which is the same as the equilibrium distribution of the VPSDE. We then measure the difference of the respective scores at time and the equilibrium (or prior) scores, i.e. (recall that the score of a Normal distribution is simply ),
Therefore, to understand the above observations in terms of learning neural networks, we train a small ResNet architecture (less than 100k parameters) for each of the following four setups: both CLD and VPSDE each with and without a mixed score parameterization. The mixed score of the VPSDE simply assumes a standard Normal data distribution (which is also the equilibrium distribution of VPSDE) resulting in adding to the score function. Formally, is the score of a Normal distribution with unit variance.
E.2 Image Modeling Experiments
We perform image modeling experiments on CIFAR-10 as well as CelebA-HQ-256. We report FID scores on CIFAR-10 for our main model for various different solvers; see Tab. 3 and Tab. 3. We further present generated samples for both models in Sec. 5 using Euler–Maruyama with 2000 quadratic striding steps and Runge–Kutta 4(5) for CIFAR10 and CelebA-HQ-256, respectively. We present additional samples for various solver settings in App. F. All (average) NFEs for the Runge–Kutta solver are computed using a batch size of 128.
Our models are based on the NCSN++ and the DDPM++ architectures from Song et al. (2021c). Importantly, we changed the number of input channels from three to six to facilitate the additional velocity variables. Note that the number of additional neural network parameters due to this change is negligible.
For fair a comparison, we train our models using the same -sampling cutoff during training as is used for VESDE and VPSDE in Song et al. (2021c). Note, however, that this is not strictly necessary for CLD as we do not have any “blow-up” of the SDE due to unbounded scores as (also see Fig. 18 and Fig. 19).
We summarize our three model architectures as well as our SDE and training setups in Tab. 6.
E.2.2 CIFAR-10 Results for VESDE and VPSDE
The results reported for VESDE and VPSDE using the GGF sampler are taken from Jolicoeur-Martineau et al. (2021a). All other results for VESDE and VPSDE are generated using the provided PyTorch code as well as the provided checkpoints from Song et al. (2021c).https://github.com/yang-song/score_sde_pytorch We used EM and PC to sample from the VPSDE and VESDE models, respectively (see Sec. 5.2), since these choices correspond to their recommended settings.https://colab.research.google.com/drive/1dRR_0gNRmfLtPavX2APzUggBuXyjWW55
Furthermore, in App. F.2 we also used DDIM (Song et al., 2021a) to sample the VPSDE. DDIM’s update rule is
where , , and .
E.2.3 Quadratic Striding
When we simulate our generative SDE numerically, using for example EM or our SSCS, we need to choose a time discretization. Given a certain NFE budget , how do we choose time step sizes? The standard approach is to use an equidistant discretization, corresponding to a set of evaluation time steps with . However, prior work (Song et al., 2021a; Kong & Ping, 2021; Watson et al., 2021) has shown that it can be beneficial to focus function evaluations (neural network calls) on times “close to the data”. This is because the diffusion process distribution is most complex close to the data and almost perfectly Normal close to the prior. Among other techniques, these works used a useful heuristic, denoted as quadratic striding (QS), which discretizes the integration interval such that the evaluation times follow a quadratic schedule and the individual time steps follow a linear schedule. We also used this QS approach in our experiments.
We can formally define it as follows (assuming a time interval here for simplicity): Denote the evaluation times as (including and ) and define:
and to ensure that .
This describes the time steps as going from to . During synthesis, however, we are going backwards. Hence, we can define our time steps as
where now counts time steps in the other direction. Note that this can be easily adapted to general integration intervals .
E.2.4 Denoising
As has been pointed out in Jolicoeur-Martineau et al. (2021b), samples that are generated with models similar to ours can contain noise that is hard to detect visually but worsens FID scores significantly.
For a fair comparison we use the same denoising setup for all experiments we conducted (including VESDE (PC/ODE) and VPSDE (EM/ODE)) except for LSGM.Denoising has not been used in the original LSGM work (Vahdat et al., 2021) and is not needed in their case, since the output of the latent SGM lives in a smooth latent space and is further processed by a decoder. We simulate the underlying generative ODE/SDE until the time cutoff and then take a single denoising step of the form
This denoising step can be considered as an Euler–Maruyama step without noise injection. For SDEs acting on data directly (VESDE, VPSDE, etc.) the corresponding denoising formula is
For SDEs acting in the data space directly, it has been reported that this denoising step is crucial to obtain good FID scores Jolicoeur-Martineau et al. (2021b); Song et al. (2021c). When we simulate the generative probability flow ODE we found that denoising is important in order for the Runge–Kutta solver not to “blow-up” as . On the other hand, when simulating CLD using our new SSCS solver, we found that denoising only slightly influences FID (see Tab. 7). We believe that this might be because the neural network does not have any influence on the denoising step for CLD. More specifically, the neural network only denoises the velocity component. However, we are primarily interested in the data component. Putting the drift and diffusion coefficients of CLD in the denoising formula in Eq. (107), we obtain
E.2.5 Solver Error Tolerances for Runge–Kutta 4(5)
In Tab. 3, we report FID scores for a Runge–Kutta 4(5) solver (Dormand & Prince, 1980) as well as the “Gotta Go Fast” solver from Jolicoeur-Martineau et al. (2021a) (see their Table 1). For simulating CLD with Runge–Kutta 4(5) we chose the solver error tolerances to hit certain regimes of NFEs to facilitate comparisons with VPSDE and VESDE. We obtain a mean number of function evaluations of 312 and 137 using Runge–Kutta 4(5) solver error tolerances of and , respectively. For VESDE, VPSDE and LSGM we used as the ODE solver error tolerance, following the recommended default setups (Song et al., 2021c; Vahdat et al., 2021). These values are used for both relative and absolute error tolerances.
E.2.6 Ablation Experiments
The model architecture used for all ablation experiments can be found in Tab. 6. As pointed out in Sec. 5 we found that the hyperparameters and only have small effects on CIFAR-10 FID scores. On the other hand, we found that the mixed score parameterization helps significantly in obtaining competitive FIDs.
E.2.7 LSGM-100M Model
Our CLD-based SGM has parameters, while the original CIFAR-10 Latent SGM from Vahdat et al. (2021), to which we compare in Tab. 1, uses parameters. To establish a fairer comparison between our CLD-based SGMs and LSGM (Vahdat et al., 2021), we train another smaller LSGM model with parameters. To do this, we followed the exact setup of the “CIFAR-10 (balanced)” model from LSGM (see Table 7 in Vahdat et al. (2021)), with a few minor modifications: We used a VAE backbone model with only 10 groups instead of 20, which corresponds to a reduction in parameters by a factor of 2 in the encoder and decoder networks. We also reduced the convolutional channels in the latent space SGM from 512 to 256 and reduced the number of the residual cells per scale from 8 to 4. With these modifications the resulting “LSGM-100M” uses only parameters overall with approximately half of them in the encoder and decoder networks and the other half in the latent SGM. Other than these architecture modifications, our model is trained in exactly the same way as the bigger, original models in Vahdat et al. (2021).
For evaluation, we follow the recommended setting by Vahdat et al. (2021) and use the same Runge-Kutta 4(5) ODE solver with an error tolerance of to solve the probability flow ODE in LSGM’s latent space. LSGM-100M achieves an FID of 4.60, an NLL bound of 2.96 bpd, and requires on average 131 NFE for sampling new images. We report these results in Tabs. 1 and 3 in the main text.
Note that we also tried training a model following the “CIFAR-10 (best FID)” setup, but found training to be unstable (however, the orignal “CIFAR-10 (best FID)” model from Vahdat et al. (2021) only performs marginally better in FID than their “CIFAR-10 (balanced)” model anyway). Furthermore, we also tried training another small LSGM with a similar number of parameters but with more parameters in the latent SGM and less in the encoder and decoder networks, compared to the reported LSGM-100M. However, this model performed significantly worse.
Appendix F Additional Experiments
In order to test combinations of diffusions and numerical samplers in isolation, we consider a dataset for which we know the ground truth score function (for all ) analytically. In particular, we use the mixture of Normals introduced in App. E.1; see Fig. 10(a) for a visualization of the data distribution. In Fig. 11, we show samples for VPSDE (Euler–Maruyama (EM) sampler) and CLD (EM and SSCS samplers). For quantitative comparison, we also compute negative log-likelihoods for the three combinations (which can be done easily due to our access to the ground truth distribution): as can be seen in Tab. 8, for each number of steps CLD with SCSS outperforms both VPSDE and CLD with EM. As discussed in Sec. 3.3, we can see in Tab. 8 that EM is not well-suited for CLD. This is true, in particular, when using a small number of synthesis steps (function evaluations). In Fig. 11, we see that CLD with EM leads to sampling distributions which are too broad. These results are exactly in line with the “diverging” dynamics that is observed when solving Hamiltonian dynamics with a non-symplectic integrator, such as the standard Euler method (Neal, 2011). This problematic behavior of Euler-based techniques is more pronounced when using fewer steps with larger stepsizes, which is also what we observe in our experiments. These results further motivate the use of our novel SSCS, which addresses these challenges, for sampling from our CLD-based SGMs.
F.1.2 Maximum Likelihood Training
For maximum likelihood training, models based on overdamped Langevin dynamics such as VPSDE need to learn an unbounded score for . Our model, on the other hand, only ever needs to learn a bounded score even for . For our image data experiments, we use a reweighted objective function to improve visual quality of samples (as is general practice).
Here, we also study training towards maximum likelihood on toy dataset tasks. To explore this, we repeat the neural network complexity experiment from App. E.1 with maximum likelihood training (instead of the reweighted objective). Furthermore, we also train VPSDE-based and CLD-based SGMs on a challenging toy dataset and find that CLD significantly outperforms VPSDE. We leave the study of CLD with maximum likelihood training for high-dimensional (image) datasets to future work.
The setup of this experiment is equivalent to the setup in App. E.1 up to the training objective: in this experiment we do maximum likelihood learning, i.e., we train CLD models with the objective from Eq. (8) with .For the ML objective of the VPSDE, we refer the reader to Song et al. (2021b). Furthermore, we test CLD in this setup for three different values of . The results of this experiment can be found in Fig. 12. For CLD, we find that larger values of generally lead to less complex networks, in particular for smaller times . However, even for the learned neural network is still significantly smoother than the network learned for the VPSDE when a mixed score parameterization is used.The VPSDE-based model with mixed score parameterization did not converge to the target distribution, and therefore is not included in Fig. 12.
Using the same simple ResNet architecture (less than 100k parameters) from the above experiment, we trained a VPSDE-based as well as a CLD-based SGM to maximize the likelihood of a more challenging toy dataset (the dataset is essentially “multi-scale”, as it involves both large scale—the placement of the swiss rolls—and fine scale—the swiss rolls themselves—structure). Similar to the other toy datasets, the models are trained for 1M iterations using fresh data synthesized from the data distribution in each batch at a batch size of 512.
In Fig. 13, we compare samples of the models to the data distribution. Even with our simple model architecture, CLD is able to capture the multi-scale structure of the dataset: the five rolls are adequately resembled and only a few samples are in between modes. VPSDE, on the other hand, only captures the main modes, but not the fine structure. Furthermore, VPSDE has the undesired behavior of “connecting” the modes.
Overall, we conclude that also in the maximum likelihood training setting CLD is a promising diffusion showing superior behavior compared to the VPSDE in our toy experiments.
F.2 CIFAR-10 — Extended Results
In this section, we provide additional results on the CIFAR-10 image modeling benchmark.
An extended version of Tab. 3 (sampling the generative SDE with different fixed-step size solvers for different compute budgets) including additional baselines can be found in Tab. 9. Note that time stepping with quadratic striding (QS) improves sampling from VPSDE- and CLD-based models for all settings except for the combination of VPSDE and EM sampling in the setting . For the VESDE (using PC sampling), QS significantly worsens FID scores. The reason for this could be that the variance of the VESDE already follows an exponential schedule (see Fig. 5 in Song et al. (2021c)). We additionally present results for the VPSDE using the DDIM (Denoising Diffusion Implicit Models) sampler (Song et al., 2021a). As was observed by Song et al. (2021a), QS also helps for DDIM. Importantly, for any , our CLD with our novel SSCS (and QS) even outperforms DDIM. Only for , DDIM performs better. It needs to be mentioned, however, that the DDIM sampler was specifically designed for few-step sampling, whereas our CLD with SSCS is derived in a general fashion without this particular regime in mind. In particular, DDIM sampling can be interpreted as a non-Markovian sampling method and it is not clear how to calculate the log-likelihood of hold-out validation data under this non-Markovian synthesis approach. Nevertheless, it would be interesting to also explore non-Markovian DDIM-inspired techniques for CLD-SGMs to further improve sampling speed in CLD-SGMs.
Note that our DDIM results shown in Tab. 9 are better than those presented in Song et al. (2021a) itself, because we are relying on the DDPM++ model trained in Song et al. (2021c), whereas Song et al. (2021a) uses the DDPM model from Ho et al. (2020).
Finally, we present additional generated samples from our CLD-SGM model: see Fig. 14 and Fig. 15 for samples from EM-QS with 2000 evaluations and SSCS-QS with 150 evaluations, respectively.
F.3 CelebA-HQ-256 — Extended Results
In this section, we provide additional qualitative results on CelebA-HQ-256. For high quality samples using our new SSCS solver see Fig. 16.
Samples generated with an adaptive step size Runge–Kutta 4(5) solver at different solver tolerances can be found in Fig. 17. We found that our model still generates very good samples even for a solver error tolerance of using an average of 129 neural network evaluations.
Lastly, we show “generation paths” of samples from our CelebA-HQ-256 model: see Fig. 18 and Fig. 19 for samples from the probability flow ODE and the generative SDE, respectively. We visualize the continuous generation paths via snapshots of data and velocity variables at eight different time steps. Interestingly, we can see that the velocity variables “encode” the data at intermediate . On the other hand, at time , by construction, both data and velocity are distributed according to the “equilibrium distribution” of the diffusion, namely, . Furthermore, as the data variables approximately converge to the data distribution, while the velocity variables approximately converge to another Normal distribution (with in our experiments).
Recall that for CLD, the neural network approximates the score . We believe that the generation paths are further evidence that CLD-SGMs need to learn simpler models: for fixed the velocity variable appears to be a “noisy” version of the data , and therefore we believe to be relatively smooth and simple when compared to the marginal .
Finally, note that in Figs. 18 and 19, when visualizing the velocity variables, we used a colorization scheme that corresponds exactly to the inverse of the color scheme used for visualizing the images themselves. Alternatively, we could also interpret this in such a way that we are not actually visualizing velocities, but negative velocities with flipped signs. When using this inverse colorization scheme for the velocities, we see that at intermediate , where the velocities encode the data, the color values visualizing image data and velocities are, apart from the additional noise in the velocities, similar (i.e. the velocities appear as noisy versions of the actual images). This implies that image pixel values translate into corresponding negative velocities that pull the pixel values back towards the mean of the equilibrium distribution. This is a consequence of the Hamiltonian coupling between the data and velocity variables. In other words, it is a result of the negative sign in front of in the term in Eq. (5) (and analogously for the reverse-time generative SDE). Also see the visualizations on our project page (https://nv-tlabs.github.io/CLD-SGM).
Appendix G Proofs of Perturbation Kernels
In this section, we prove the correctness of the perturbation kernels of the forward diffusion (App. B.1) as well as for the analytical splitting term in our SSCS (App. D.2). All derivations are presented for general time-dependent .
We have the following ODEs describing the evolution of the mean and the covariance matrix
In App. B.1, we claim the following solutions:
where and as well as and are initial conditions.
Plugging the claimed solution (Eqs. (114)-(118)) back into the ODE (Eq. 110), we obtain
The above can be decomposed into two equations:
where we used the fact that . Eq. (127): Plugging the claimed solution into Eq. (127), we obtain:
Eq. (128): After simplification, plugging in the claimed solution into Eq. (128), we obtain:
This completes the proof of the correctness of the mean.
G.1.2 Proof of Correctness of the Covariance
Plugging the claimed solution (Eqs. (119)-(125)) back in the ODE (Eq. (111)), we obtain
as well as the fact that , we can decompose Eq. (131) into three equations:
Eq. (134): Plugging the claimed solution into Eq. (134), we obtain
Eq. (135): After simplification, plugging the claimed solution into Eq. (135), we obtain
Eq. (136): After simplification, plugging the claimed solution into Eq. (136), we obtain
This completes the proof of the correctness of the covariance.
G.2 Analytical Splitting Term of SSCS
We have the following ODEs describing the evolution of the mean and the covariance matrix
These ODEs are very similar to the ODEs of the forward diffusion in App. G.1, the only difference being flipped signs in the off-diagonal terms of (highlighted in red). In App. D.2, we claim the following solutions
where {\color[rgb]{1,0,0}\definecolor[named]{pgfstrokecolor}{rgb}{1,0,0}{\mathcal{B}}(t)=\int_{t^{\prime}}^{t}\beta(T-\hat{t})\,d\hat{t}} and is an initial condition. Differences of the above solution to the solutions of the forward diffusion are again highlighted in red. Note that by construction the initial covariance for the analytical splitting term of SSCS is the zero matrix, i.e., , since we always initialize from an “updated sample”, which itself does not have any uncertainty. Also note that in this derivation we use general initial (whereas in App. G.1 we set for simplicity).
Plugging the claimed solution (Eqs. (144)-(148)) into the ODE (Eq. (140)), we obtain
The above can be decomposed into two equations:
where we used the fact that . Eq. (157): Plugging the claimed solution into Eq. (157), we obtain:
Eq. (128): After simplification, plugging the claimed solution into Eq. (128), we obtain:
This completes the proof of the correctness of the mean.
G.2.2 Proof of Correctness of the Covariance
Plugging the claimed solution (Eqs. (149)-(155)) into the ODE (Eq. (141)), we obtain
as well as the fact , we can decompose Eq. (161) into three equations:
Eq. (164): Plugging the claimed solution into Eq. (164), we obtain
Eq. (165): After simplification, plugging the claimed solution into Eq. (165), we obtain
Eq. (166): After simplification, plugging the claimed solution into Eq. (166), we obtain
This completes the proof of the correctness of the covariance.
To connect back to the SSCS as presented in App. D.2, recall that in practice we use constant (and ) and that we solve for small time steps of size , such that , which leads to the expressions presented in App. D.2.