Conditional Image Generation with Score-Based Diffusion Models

Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb, Christian Etmann

Introduction

The goal of generative modelling is to learn a probability distribution from a finite set of samples. This classical problem in statistics has been studied for many decades, but until recently efficient learning of high-dimensional distributions remained impossible in practice. For images, the strong inductive biases of convolutional neural networks have recently enabled the modelling of such distributions, giving rise to the field of deep generative modelling.

Deep generative modelling became one of the central areas of deep learning with many successful applications. In recent years much progress has been made in unconditional and conditional image generation. The most prominent approaches are auto-regressive models , variational auto-encoders (VAEs) , normalizing flows and generative adversarial networks (GANs) .

Despite their success, each of the above methods suffers from important limitations. Auto-regressive models allow for likelihood estimation and high-fidelity image generation, but require a lot of computational resources and suffer from poor time complexity in high resolutions. VAEs and normalizing flows are less computationally expensive and allow for likelihood estimation, but tend to produce samples of lower visual quality. Moreover, normalizing flows put restrictions on the possible model architectures (requiring invertibility of the network and a Jacobian log-determinant that is computationally tractable), thus limiting their expressivity. While GANs produce state-of-the art quality samples, they don’t allow for likelihood estimation and are notoriously hard to train due to training instabilities and mode collapse.

Recently, score-based and diffusion-based generative models have been revived and improved in and . The connection between the two frameworks in discrete-time formulation has been discovered in . Recently in , both frameworks have been unified into a single continuous-time approach based on stochastic differential equations and called score-based diffusion models. These approaches have recently received a lot of attention, achieving state-of-the-art performance in likelihood estimation and unconditional image generation , surpassing even the celebrated success of GANs.

In addition to achieving state-of-the art performance in both image generation and likelihood estimation, score-based diffusion models don’t suffer from training instabilities or mode collapse . Moreover, their time complexity in high resolutions is much better than that of auto-regressive models . This makes score-based diffusion models very attractive for deep generative modelling.

In this work, we examine how score-based diffusion models can be applied to conditional image generation. We conduct a review and classification of existing approaches and perform a systematic comparison to find the best way of estimating the conditional score. We provide a proof of validity for the conditional denoising estimator (which has been used in without justification), and we thereby provide a firm theoretical foundation for using it in future research.

Moreover, we extend the original framework to support multi-speed diffusion, where different parts of the input tensor diffuse according to different speeds. This allows us to introduce a novel estimator of the conditional score and opens an avenue for further research.

The contributions of this paper are as follows:

We review and empirically compare score-based diffusion approaches to modelling conditional distributions of image data. The models are evaluated on the tasks of super-resolution, inpainting and edge to image translation.

We provide a proof of consistency for the conditional denoising estimator - one of the most successful approaches to estimating the conditional score.

We introduce a multi-speed diffusion framework which leads to conditional multi-speed diffusive estimator (CMDE), a novel estimator of conditional score, which unifies previous methods of conditional score estimation.

We provide an open-source library MSDiff, to facilitate further research on conditional and multi-speed diffusion models. The code will be released in the near future.

Notation

In this work we will use the following notation:

Probability distributions We denote the probability distribution of a random variable solely via the name of its density’s argument, e.g.

where xtx_{t} is a realisation of the random variable XtX_{t}.

Methods

In the following, we will provide details about the framework and estimators discussed in this paper.

has a reverse diffusion process governed by the following SDE:

where wˉ\bar{w} is a standard Wiener process in reverse time.

The forward diffusion process transforms the target distribution p(x0)p(x_{0}) to a diffused distribution p(xT)p(x_{T}). By appropriately selecting the drift and the diffusion coefficients of the forward SDE, we can make sure that after sufficiently long time TT, the diffused distribution p(xT)p(x_{T}) approximates a simple distribution, such as N(0,I)\mathcal{N}(0,I). We refer to this simple distribution as the prior distribution, denoted by π\pi.

If we have access to the score of the marginal distribution, ∇xtln⁡p(xt)\nabla_{x_{t}}{\ln{p(x_{t})}}, for all tt, we can derive the reverse diffusion process and simulate it to map pTp_{T} to p0p_{0}. In practice, we approximate the score of the time-dependent distribution by a neural network sθ(xt,t)≈∇xtln⁡p(xt)s_{\theta}(x_{t},t)\approx\nabla_{x_{t}}{\ln{p(x_{t})}} and map the prior distribution π≈p(xT)\pi\approx p(x_{T}) to pθ(x)≈p(x0)p_{\theta}(x)\approx p(x_{0}) by solving the reverse-time SDE from time TT to time . One can integrate the reverse SDE using standard numerical SDE solvers such Euler–Maruyama or other discretisation strategies. The authors propose to couple the standard integration step with a fixed number of Langevin MCMC steps to leverage the knowledge of the score of the distribution at each intermediate timestep. The MCMC correction step improves sampling; the combined algorithm is known as a predictor-corrector scheme. We refer to for details.

In order to fit a neural network model sθ(xt,t)s_{\theta}(x_{t},t) to approximate the score ∇xtln⁡p(xt)\nabla_{x_{t}}{\ln{p(x_{t})}}, we minimize the weighted Fisher’s divergence

The above quantity cannot be optimized directly since we don’t have access to the ground truth score ∇xtln⁡p(xt)\nabla_{x_{t}}{\ln{p(x_{t})}}. Therefore in practice, a different objective has to be used . In , the continuous denoising score-matching objective is chosen, which is equal to LSM(θ)\mathcal{L}_{SM}(\theta) up to an additive term, which does not depend on θ\theta and is defined as

2 Conditional generation

The continuous score-matching framework can be extended to conditional generation, as shown in . Suppose we are interested in p(x∣y)p(x|y), where xx is a target image and yy is a condition image. Again, we use the forward diffusion process (Equation 1) to obtain a family of diffused distributions p(xt∣y)p(x_{t}|y) and apply Anderson’s Theorem to derive the conditional reverse-time SDE

Now we need to learn the score ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y) in order to be able to sample from p(x∣y)p(x|y) using reverse-time diffusion.

In this work, we discuss the following approaches to estimating the conditional score ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y):

Multi-speed conditional diffusive estimators (our method)

We discuss each of them in a separate section.

In an additional approach to conditional score estimation was suggested: This method proposes learning ∇xtln⁡p(xt)\nabla_{x_{t}}\ln p(x_{t}) with an unconditional score model, and learning p(y∣xt)p(y|x_{t}) with an auxiliary model. Then, one can use

to obtain ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y). Unlike other approaches, this requires training a separate model for p(y∣xt)p(y|x_{t}). Appropriate choices of such models for tasks discussed in this paper have not been explored yet. Therefore we exclude this approach from our study.

The conditional denoising estimator (CDE) is a way of estimating p(xt∣y)p(x_{t}|y) using the denoising score matching approach . In order to approximate p(xt∣y)p(x_{t}|y), the conditional denoising estimator minimizes

This estimator has been shown to be successful in previous works , also confirmed in our experimental findings (cf. Section 4).

Despite the practical success, this estimator has previously been used without a theoretical justification of why training the above objective yields the desired conditional distribution. Since p(xt∣y)p(x_{t}|y) does not appear in the training objective, it is not obvious that the minimizer approximates the correct quantity.

By extending the arguments of , we provide a formal proof that the minimizer of the above loss does indeed approximate the correct conditional score p(xt∣y)p(x_{t}|y). This is expressed in the following theorem.

The proof for this statement can be found in Appendix B.1. Using the above theorem, the consistency of the estimator can be established.

Let θ∗\theta^{\ast} be a minimizer of a Monte Carlo approximation of (LABEL:CDN), then (under technical assumptions, cf. Appendix B.2) the conditional denoising estimator sθ∗(x,y,t)s_{\theta^{\ast}}(x,y,t) is a consistent estimator of the conditional score ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y), i.e.

as the number of Monte Carlo samples approaches infinity.

This follows from the previous theorem and the uniform law of large numbers. Proof in the Appendix B.2.

2.2 Conditional diffusive estimator (CDiffE)

Conditional diffusive estimators (CDiffE) have first been suggested in . The core idea is that instead of learning p(xt∣y)p(x_{t}|y) directly, we diffuse both xx and yy and approximate p(xt∣yt)p(x_{t}|y_{t}), using the denoising score matching. Just like learning diffused distribution ∇xtln⁡p(xt)\nabla_{x_{t}}\ln p(x_{t}) improves upon direct estimation of ∇xln⁡p(x)\nabla_{x}\ln p(x) , diffusing both the input xx and condition yy, and then learning ∇xtln⁡p(xt∣yt)\nabla_{x_{t}}\ln p(x_{t}|y_{t}) could make optimization easier and give better results than learning ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y) directly.

In order to learn p(xt∣yt)p(x_{t}|y_{t}), observe that

where zt:=(xt,yt)z_{t}:=(x_{t},y_{t}) and nxn_{x} is the dimensionality of xx. Therefore we can learn the (unconditional) score of the joint distribution p(xt,yt)p(x_{t},y_{t}) using the denoising score matching objective just like as in the unconditional case, i.e

We can then extract our approximation for the conditional score ∇xtln⁡p(xt∣yt)\nabla_{x_{t}}\ln p(x_{t}|y_{t}) by simply taking the first nxn_{x} components of sθ(xt,yt,t)s_{\theta}(x_{t},y_{t},t).

The aim is to approximate ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y) with ∇xtln⁡p(xt∣yt^)\nabla_{x_{t}}\ln p(x_{t}|\hat{y_{t}}), where y^t\hat{y}_{t} is a sample from p(yt∣y)p(y_{t}|y). Of course this approximation is imperfect and introduces an error, which we call the approximation error. CDiffE aims to achieve smaller optimization error by diffusing the condition yy and making the optimization landscape easier, at a cost of making this approximation error.

Now in order to obtain samples from the conditional distribution, we sample a point xT∼πx_{T}\sim\pi and integrate

from TT to , sampling y^t∼p(yt∣y)\hat{y}_{t}\sim p(y_{t}|y) at each time step.

2.3 Conditional multi-speed diffusive estimator (CMDE)

In this section we present a novel estimator for the conditional score ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y) which we call the conditional multi-speed diffusive estimator (CMDE).

Our approach is based on two insights. Firstly, there is no reason why xtx_{t} and yty_{t} in conditional diffusive estimation need to diffuse at the same rate. Secondly, by decreasing the diffusion rate of yty_{t} while keeping the diffusion speed of xtx_{t} the same, we can bring p(xt∣yt)p(x_{t}|y_{t}) closer to p(xt∣y)p(x_{t}|y), at the possible cost of making the optimization more difficult. This way we can interpolate between the conditional denoising estimator and the conditional diffusive estimator and find an optimal balance between optimization error and approximation error (cf. Figure 2). This can lead to a better performance, as indicated by our experimental findings (cf. Section 4).

In our conditional multi-speed diffusive estimator, xtx_{t} and yty_{t} diffuse according to SDEs with the same drift but different diffusion rates,

where v=∇ztln⁡p(zt∣z0)−sθ(zt,t)v=\nabla_{z_{t}}\ln{p(z_{t}|z_{0})}-s_{\theta}(z_{t},t), zt=(xt,yt)z_{t}=(x_{t},y_{t}).

In authors derive a likelihood weighting function λMLE(t)\lambda^{\text{MLE}}(t), which ensures that the objective of the score-based model upper-bounds the negative log-likelihood of the data, thus enabling approximate maximum likelihood training of score-based diffusion models. We generalize this result to the multi-speed diffusion case by providing a likelihood weighting matrix ΛMLE(t)\Lambda^{\text{MLE}}(t) with the same properties.

Let L(θ)\mathcal{L}(\theta) be the CMDE training objective (Equation 10) with the following weighting:

Then the joint negative log-likelihood is upper bounded (up to a constant in θ\theta) by the training objective of CMDE

Moreover we show that the mean squared approximation error of a multi-speed diffusion model is upper bounded and the upper bound goes to zero as the diffusion speed of the condition σy(t)\sigma^{y}(t) approaches zero.

Thus we see that the objective of CMDE approaches that of CDE as σy(t)→0\sigma^{y}(t)\to 0, and CMDE coincides with CDiffE when σy(t)=σx(t)\sigma^{y}(t)=\sigma^{x}(t) (cf. Figure 2).

We experimented with different configurations of σx(t)\sigma^{x}(t) and σy(t)\sigma^{y}(t) and found configurations that lead to improvements upon CDiffE and CDE in certain tasks. The experimental results are discussed in detail in Section 4.

2.4 MSDiff: Beyond multi-speed diffusion

Based on this work, we provide our open source library MSDiff, which generalizes the original framework of and allows to diffuse xtx_{t} and yty_{t} not only at different diffusion rates, but with two entirely different forward SDEs (i.e. with different diffusion coefficients and different drifts):

Moreover, the likelihood weighting of Theorem 2 holds in this more general case, allowing for principled training of multi-sde diffusion models (cf. Appendix B.3). This flexibility opens room for further research into multi-speed diffusion based training of score-based models, which we intend to examine in a future study.

Experiments

In this section we conduct a systematic comparison of different score-based diffusion approaches to modelling conditional distributions of image data. We evaluate these approaches on the tasks of super-resolution, inpainting and edge to image translation. Moreover, we compare the most successful score-based diffusion approaches for super-resolution with HCFlow – a state-of-the-art method in super-resolution.

Datasets In our experiments, we use the CelebA and Edges2shoes datasets. We pre-processed the CelebA dataset as in .

Models and hyperparameters In order to ensure the fair comparison, we separate the evaluation of a particular estimator of conditional score from the evaluation of a particular neural network model. To this end, we train the same neural network architecture for all estimators. The architecture is based on the DDPM model used in . We used the variance-exploding SDE given by:

Likelihood weighting was employed for all experiments. For CMDE, the diffusion speed of yy was controlled by picking an appropriate σmax⁡y\sigma^{y}_{\max}, which we found by trial-and-error. The performance of CMDE could be potentially improved by performing a systematic hyperparameter search for optimal σmax⁡y\sigma^{y}_{\max}. Details on hyperparameters and architectures used in our experiments can be found in Appendix LABEL:appendix:hyperparams.

Inverse problems The tasks of inpainting, super-resolution and edge to image translation are special cases of inverse problems . In each case, we are given a (possibly random) forward operator AA which maps our data xx (full image) to an observation yy (masked image, compressed image, sketch). The task is to come up with a high-quality reconstruction x^\hat{x} of the image xx based on an observation yy. The problem of reconstructing xx from yy is typically ill-posed, since yy does not contain all information about xx. Therefore, an ideal algorithm would produce a reconstruction x^\hat{x}, which looks like a realistic image (i.e. is a likely sample from p(x)p(x)) and is consistent with the observation yy (i.e. Ax^≈yA\hat{x}\approx y). Notice that if a conditional score model learns the conditional distribution correctly, then our reconstruction x^\hat{x} is a sample from the posterior distribution p(x∣y)p(x|y), which satisfies bespoke requirements. This strategy for solving inverse problems is generally referred to as posterior sampling.

Evaluation: Reconstruction quality Ill-posedness often means that we should not strive to reconstruct xx perfectly. Nonetheless reconstruction error does correlate with the performance of the algorithm and has been one of the most widely-used metrics in the community. To evaluate the reconstruction quality for each task, we measure the Peak signal-to-noise ratio (PSNR) , Structural similarity index measure (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS) between the original image xx and the reconstruction x^\hat{x}.

Evaluation: Consistency In order to evaluate the consistency of the reconstruction, for each task we calculate the PSNR between y:=Axy:=Ax and y^:=Ax^\hat{y}:=A\hat{x}.

Evaluation: Diversity We evaluate diversity of each approach by generating five reconstructions (x^)i=15(\hat{x})_{i=1}^{5} for a given observation y{y}. Then for each yy we calculate the average standard deviation for each pixel among the reconstructions (x^)i=15(\hat{x})_{i=1}^{5} . Finally, we average this quality over 5000 test observations.

Evaluation: Distributional distances If our algorithm generates realistic reconstructions while preserving diversity, then the distribution of reconstructions p(x^)p(\hat{x}) should be similar to the distribution of original images p(x)p(x). Therefore, we measure the Fréchet Inception Distance (FID) between unconditional distributions p(x)p(x) and p(x^)p(\hat{x}) based on 5000 samples. Moreover, we calculate the FID score between the joint distributions p(x^,y)p(\hat{x},y) and p(x,y)p(x,y), which allows us to simultaneously check the realism of the reconstructions and the consistency with the observation. We use abbreviation UFID to refer to FID between between unconditional distributions and JFID to refer to FID between joints. In our judgement, FID and especially the JFID is the most principled of the used metrics, since it measures how far pθ(x∣y)p_{\theta}(x|y) is from p(x∣y)p(x|y).

We perform the inpainting experiment using CelebA dataset. In inpainting, the forward operator AA is an application of a given binary mask to an image xx. In our case, we made the task more difficult by using randomly placed (square) masks. Then the conditional score model is used to obtain a reconstruction x^\hat{x} from the masked image yy. We select the position of the mask uniformly at random and cover 25%25\% of the image. The quantitative results are summarised in Table 1 and samples are presented in Figure 4. We observe that CDE and CMDE significantly outperform CDiffE in all metrics, with CDE having a small advantage over CMDE in terms of reconstruction error and consistency. On the other hand, CMDE achieves the best FID scores.

2 Super-resolution

We perform 8x super-resolution using the CelebA dataset. A high resolution 160x160 pixel image xx is compressed to a low resolution 20x20 pixels image yy. Here we use bicubic downscaling as the forward operator AA. Then using a score model we obtain a 160x160 pixel reconstruction image x^\hat{x}. The quantitative results are summarised in Table 1 and samples are presented in Figure 5. We find that CMDE and CDE perform similarly, while significantly outperforming CDiffE. CMDE achieves the smallest reconstruction error and captures the distribution most accurately according to FID scores.

3 Edge to image translation

We perform an edge to image translation task on the Edges2shoes dataset. The forward operator AA is given by a neural network edge detector , which takes an original photo of a shoe xx and transforms it into a sketch yy. Then a conditional score model is used to create an artificial photo of a shoe x^\hat{x} matching the sketch. The quantitative results are summarised in Table 1 and samples are presented in Figure 6. Unlike in inpainting and super-resolution where CDiffE achieved reasonable performance, in edge to image translation, it fails to create samples consistent with the condition (which leads to inflated diversity scores). CDE and CMDE are comparable, but CDE performed slightly better across all metrics. However, the performance of CMDE could be potentially improved by tuning the diffusion speed σy(t)\sigma^{y}(t).

Comparison with state-of-the-art

We compare score-based diffusion approaches with HCFlow – a state-of-the-art method in super-resolution. To ensure a fair comparison, we used the data pre-processing and hyperparameters for HCFlow exactly as in the original paper . We find that although HCFlow performs marginally better in terms of the reconstruction error, CDE and CMDE obtain a significantly better FID and diversity scores indicating better distribution coverage. We recall that a perfect reconstruction on a per-image basis is generally not desirable due to ill-posedness of the inverse problem and therefore in our view FID is the most principled of the used metrics. The FID scores suggest that CMDE was the most successful approach to approximating the posterior distribution.

Conclusions and future work

In this article, we conducted a systematic comparison of score-based diffusion models in conditional image generation tasks and provided an in-depth theoretical analysis of the estimators of conditional score. In particular, we proved the consistency of the conditional denoising estimator, thus providing a firm theoretical justification for using it in future research.

Moreover, we introduced a multi-speed diffusion framework, which led to CMDE, a novel estimator for conditional score which interpolates between conditional denoising estimator (CDE) and conditional diffusive estimator (CDiffE) by controlling the diffusion speed of the condition.

Our study showed that CMDE and CDE perform on par, while significantly outperforming CDiffE. This is particularly apparent in edge to image translation, where CDiffE fails to produce samples consistent with the condition image. Furthermore, CMDE outperformed CDE in terms of FID scores in inpainting and super-resolution tasks, which indicates that diffusing the condition at the appropriate speed can have beneficial effect on the optimization landscape, and yield better approximation of the posterior distribution.

We found that score-based diffusion models perform on par with prior state-of-the-art methods in super-resolution task and achieve better posterior approximation according to FID score.

Acknowledgements

GB acknowledges the support from GSK and the Cantab Capital Institute for the Mathematics of Information. JS acknowledges the support from Aviva and the Cantab Capital Institute for the Mathematics of Information. CBS acknowledges support from the Philip Leverhulme Prize, the Royal Society Wolfson Fellowship, the EPSRC advanced career fellowship EP/V029428/1, EPSRC grants EP/S026045/1 and EP/T003553/1, EP/N014588/1, EP/T017961/1, the Wellcome Innovator Award RG98755, the Leverhulme Trust project Unveiling the invisible, the European Union Horizon 2020 research and innovation programme under the Marie Skodowska-Curie grant agreement No. 777826 NoMADS, the Cantab Capital Institute for the Mathematics of Information and the Alan Turing Institute. CE acknowledges support from the Wellcome Innovator Award RG98755 for part of the work that was done at Cambridge.

References

Appendix A Variance schedule

Apart from continuous-time training (where tt is sampled from U(0,T)U(0,T)), one can also perform discrete-time training for the CMDE estimator. In discrete-time training, tt is sampled uniformly from

We set N=1000N=1000, T=1T=1 as in , σmaxy=1\sigma_{max}^{y}=1 and trained the score model on the Edges2shoes dataset. We discovered that under this training configuration, the conditional score based model yielded a modelled distribution pθ(x∣y)p_{\theta}(x|y) (under reverse diffusion) that is far away from the true distribution pθ(x∣y)p_{\theta}(x|y), as shown visually in Figure 7. Apparently, the optimisation problem is harder because of the combination of discrete training and big difference in the diffusion speeds of xx and yy. These factors can cause convergence to a bad local minimum.

For this reason, we decided to make the optimisation problem easier by gradually decreasing the diffusion rate of yy as the training proceeds. More specifically, the diffusion coefficient at training iteration nn of the VE SDE that governs the diffusion of the condition y is given by

We performed the following experiment to test whether VS-CMDE outperforms CMDE: We set M=125000M=125000, σmax⁡target=1\sigma_{\max}^{\text{target}}=1 and trained the score model for 125000125000 iterations. At the end of training σmaxy\sigma_{max}^{y} reaches 11, just as in the previous experiment, where the score model failed to reach a good approximation. However, training the score model with the variance reduction schedule led to a significantly improved approximation of the posterior distribution as shown by the samples in Figure 7. The intuition behind the variance reduction schedule is that the problem of estimating ∇xt,ytlog⁡p(xt,yt)\nabla_{x_{t},y_{t}}{\log{p(x_{t},y_{t})}} is easier if yy diffuses faster, since the distribution p(xt,yt)p(x_{t},y_{t}) is smoother, yielding an easier score approximation problem. The score model can fit the smooth score vector field more accurately at the early stage of the training. The initial fitting is used as a good initialisation for the harder task of estimating ∇xt,ytlog⁡p(xt,yt)\nabla_{x_{t},y_{t}}{\log{p(x_{t},y_{t})}} at the later stages of training where y diffuses at a slower rate.

When we switched to continuous training, CMDE performed well, as indicated by the visual samples in Figure 8 and the evaluation metrics presented in Table 6. However, VS-CMDE still outperformed vanilla CMDE. It is also worth mentioning that VS-CMDE yielded the same performance as CDE (in terms of JFID and LPIPS), while providing slightly more diverse samples. CDE outperformed CMDE based on the FID-scores in the Edges2shoes task, while performing on par with VS-CMDE. VS-CMDE performed on par with CMDE on the tasks of super-resolution and inpainting, as shown in Table 2.

Our experimental findings indicate, that in discrete training, VS-CMDE should be preferred over CMDE. However, in continuous training, VS-CMDE does not have a consistent competitive advantage over CMDE.

Appendix B Proofs

Since yy and tt are fixed, we may define ψ(xt):=sθ(xt,y,t)\psi(x_{t}):=s_{\theta}(x_{t},y,t), q(x0):=p(x0∣y)q(x_{0}):=p(x_{0}|y) and q(xt∣x0)=p(xt∣x0,y)q(x_{t}|x_{0})=p(x_{t}|x_{0},y). Therefore, by the Tower Law, the statement of the lemma is equivalent to

Which follows directly from [25, Eq. 11]. ∎

in θ\theta is the same as the minimizer of

First, notice that xtx_{t} is conditionally independent of yy given x0x_{0}. Therefore, by applying the Tower Law we obtain

Now fix yy and tt. By Lemma 1, it follows that

Since tt and yy were arbitrary, this is true for all tt and yy. Therefore, substituting back into (∗)(*) we get that

(1) Tower Law, (2) Conditional independence of xtx_{t} and yy given x0x_{0}, (3) Lemma 1. ∎

B.2 Consistency of CDE

In order to prove the consistency, in this subsection we make the following assumptions:

The space of parameters Θ\Theta and the data space X\mathcal{X} are compact.

There exists a unique θ∗∈Θ\theta^{\ast}\in\Theta such that sθ∗(x,y,t)=∇xtln⁡p(x,y,t)s_{\theta^{\ast}}(x,y,t)=\nabla_{x_{t}}\ln p(x,y,t).

First we state some technical, but well-known lemmas, which will be useful in proving our consistency result.

[16, Lemma 2.4] Let ziz_{i} be i.i.d from a distribution q(z)q(z) and suppose that:

f(z,θ)f(z,\theta) is continuous for all θ∈Θ\theta\in\Theta and almost all zz.

f(⋅,θ)f(\cdot,\theta) is a measurable function of zz for each θ\theta.

L(θ)\mathcal{L}(\theta) is uniquely minimized at θ∗\theta^{\ast}.

L(n)(θ)\mathcal{L}^{(n)}(\theta) converges uniformly in probability to L(θ)\mathcal{L}(\theta).

Let θn∗\theta_{n}^{\ast} be a minimizer of a nn-sample Monte Carlo approximation of

Then under assumptions 1 and 2, the conditional denoising estimator sθn∗(x,y,t)s_{\theta_{n}^{\ast}}(x,y,t) is a consistent estimator of the conditional score ∇xtln⁡p(xt∣y)\nabla_{x_{t}}\ln p(x_{t}|y), i.e.

as the number of Monte Carlo samples nn approaches infinity.

By conditional independence and the Tower Law, we get

Let z=(t,x0,xt,y)z=(t,x_{0},x_{t},y) and denote by q(z):=p(t,x0,xt,y)q(z):=p(t,x_{0},x_{t},y) the joint distribution. Moreover, define f(z,θ):=λ(t)∥∇xtln⁡p(xt∣x0)−sθ(xt,y,t)∥22f(z,\theta):=\lambda(t)\left\lVert\nabla_{x_{t}}\ln{p(x_{t}|x_{0})}-s_{\theta}(x_{t},y,t)\right\rVert_{2}^{2}. Since t∼U(0,T)t\sim U(0,T) is independent of (x0,xt,y)∼p(x0,xt,y)(x_{0},x_{t},y)\sim p(x_{0},x_{t},y), the above is equal to

B.3 Likelihood weighting for multi-speed and multi-sde models

In this section we derive the likelihood weighting for multi-sde models (Theorem 2). First using the framework in [23, Appendix A] we present the Anderson’s theorem for multi-dimensional SDEs with non-homogeneous covariance matrix (without assuming Σ(t)≠σ(t)I\Sigma(t)\not=\sigma(t)I) and generalize the main result of to this setting. Then, we cast the problem of multi-speed and multi-sde diffusion as a special case of multi-dimensional diffusion with a particular covariance matrix Σ(t)\Sigma(t) and thus obtain the likelihood weighting for multi-sde models (Theorem 2).

If we train a score-based diffusion model to approximate ∇xln⁡pXt(x)\nabla_{x}\ln p_{X_{t}}(x) with a neural network sθ(x,t)s_{\theta}(x,t) we will obtain the following approximate reverse-time sde

Now we generalize [21, Theorem 1] to multi-dimensional setting.

Let p(xt)p(x_{t}) and pθ(xt)p_{\theta}(x_{t}) denote marginal distributions of 12 and 13 respectively. Then under regularity assumptions of [21, Theorem 1] we have that

where v=∇xtln⁡p(xt)−sθ(xt,t)v=\nabla_{x_{t}}\ln{p(x_{t})}-s_{\theta}(x_{t},t).

We proceed in close analogy to the proof of [21, Theorem 1] but we use a more general diffusion matrix Σ(t)\Sigma(t). Let PP be the law of the true reverse-time sde and let PθP_{\theta} be the law of the approximate reverse-time sde. Then by [14, Theorem 2.4] (generalized chain rule for KL divergence) we have

Using the fact that pθ(xT)=πp_{\theta}(x_{T})=\pi and applying [14, Theorem 2.4] again, we obtain

Let Pz:=P(⋅∣xT=z)P^{z}:=P(\cdot|x_{T}=z) and Pθz:=Pθ(⋅∣xT=z)P_{\theta}^{z}:=P_{\theta}(\cdot|x_{T}=z)

Using Girsanov Theorem [17, Theorem 8.6.5] and the fact that Σ(t)\Sigma(t) is symmetric and invertible

where v(xt,t)=∇xtln⁡p(xt)−sθ(xt,t)v(x_{t},t)=\nabla_{x_{t}}\ln{p(x_{t})}-s_{\theta}(x_{t},t). Since ∫0TΣ(t)v(xt,t)dwt\int_{0}^{T}\Sigma(t)v(x_{t},t)dw_{t} is a martingale (Ito’s integral wrt Brownian motion)

Now we consider again the multi-speed and the more general multi-sde diffusion frameworks from Sections 3.2.3 and 3.2.4. Suppose that we have two tensors xx and yy which diffuse according to different SDEs

We may cast this system of two SDEs, as a single SDE

where z=(x,y)z=(x,y), μz(z,t)=(μx(x,t),μy(x,t))\mu^{z}(z,t)=(\mu^{x}(x,t),\mu^{y}(x,t)) and

If we train a score-based diffusion model for zt=(xt,yt)z_{t}=(x_{t},y_{t}), then by Theorem 5

where C1:=KL(p(xT)∣π(xT))C_{1}:=KL(p(x_{T})|\pi(x_{T})) does not depend on θ\theta. Because ΛMLE\Lambda_{MLE} (from Theorem 2) is equal to Σz(t)2\Sigma_{z}(t)^{2}, we may rewrite the above as

where C2C_{2} is another term constant in θ\theta. We conclude that

where C3:=C1+C2C_{3}:=C_{1}+C_{2}. Now recall that the term on the RHS is exactly the training objective of a multi-sde score-based diffusion model with likelihood weighting

B.4 Mean square approximation error

For this proof, we drop our convention of denoting the probability distribution of a random variable via the name of its density’s argument.

Since Yt∣YY_{t}|Y has normal distribution with mean yy and variance σy(t)2\sigma^{y}(t)^{2}:

where φσ\varphi_{\sigma} is a Gaussian kernel with variance σy(t)2\sigma^{y}(t)^{2}. Moreover, under the assumptions of the lemma we can exchange the differentiaion and integration. Therefore

Since ff is a C1C^{1} function on a compact domain, it is Lipschitz and bounded (in absolute value) by some constant MM. Fix ϵ>0\epsilon>0, and let LL denote the Lipschitz constant of ff. We have that ∣f(z)−f(y)∣<ϵ|f(z)-f(y)|<\epsilon whenever ∥z−y∥<ϵ/L\left\lVert z-y\right\rVert<\epsilon/L. Let By(ϵ/L):={z∈X:∥z−y∥<ϵ/L}B_{y}(\epsilon/L):=\{z\in\mathcal{X}:\left\lVert z-y\right\rVert<\epsilon/L\} be a ball of radius ϵ/L\epsilon/L around yy. Then

where ZσZ_{\sigma} is a normally-distributed random variable with mean zero and variance σ2\sigma^{2}. By the Chernoff bound, we have

Notice that the above minimum is achieved, since AA is compact and for a fixed σ\sigma, the function ϵ↦Eϵ(1/σ)\epsilon\mapsto E_{\epsilon}(1/\sigma) is continuous.

We will prove that EE is a monotonically decreasing to zero and upper-bounds ∥(f∗φσ)−f∥∞\left\lVert(f\ast\varphi_{\sigma})-f\right\rVert_{\infty}. Firstly, it is clear that E(x)→0E(x)\to 0 as x→∞x\to\infty, since for all ϵ∈A\epsilon\in A we have lim⁡x→∞E(x)≤lim⁡x→∞Eϵ(x)=ϵ\lim_{x\to\infty}E(x)\leq\lim_{x\to\infty}E_{\epsilon}(x)=\epsilon . Secondly, suppose a<ba<b, and let ϵa\epsilon_{a} be such that E(a)=Eϵa(a)E(a)=E_{\epsilon_{a}}(a). Then

Therefore EE is monotonically decreasing. Finally since for all ϵ>0\epsilon>0

Taking minimum over ϵ∈A\epsilon\in A on both sides we obtain

Let ff be a C1C^{1} function on a compact domain and let ZZ be a random variable with mean μ\mu and variance σ2\sigma^{2}. Then

where LL denotes the Lipschitz constant of ff.

Since ff is a C1C^{1} function on a compact domain it is Lipschitz with some Lipschitz constant LL. Therefore

For this proof, we drop our convention of denoting the probability distribution of a random variable via the name of its density’s argument.

To unclutter the notation, let p(y∣x):=pY∣Xt(y∣x)p(y|x):=p_{Y|X_{t}}(y|x) and pσ(y∣x):=pYt∣Xt(y∣x)p_{\sigma}(y|x):=p_{Y_{t}|X_{t}}(y|x). Applying this notation:

Adding and subtracting ∂xtln⁡p(yt∣xt)\partial_{x_{t}}\ln p(y_{t}|x_{t}) and using the triangle inequality:

We may bound the expectation by the supremum norm

We will bound each of the summands separately. Firstly, by Assumption 3 (yt,xt)→p(yt∣xt)(y_{t},x_{t})\to p(y_{t}|x_{t}) is C2C^{2} and therefore (yt,xt)→∂xtp(yt∣xt)(y_{t},x_{t})\to\partial_{x_{t}}p(y_{t}|x_{t}) is C1C^{1}. Moreover, since X\mathcal{X} is compact, yt→∂xtp(yt∣xt)y_{t}\to\partial_{x_{t}}p(y_{t}|x_{t}) is Lipschitz for some Lipschitz constant LL. Therefore, by Lemma 6,

Adding and subtracting ∂xtpσ(⋅∣xt)p(⋅∣xt)\frac{\partial_{x_{t}}p_{\sigma}(\cdot|x_{t})}{p(\cdot|x_{t})}:

By assumption 3 and 5 we have that ∂xtpσ(⋅∣xt)\partial_{x_{t}}p_{\sigma}(\cdot|x_{t}), pσ(⋅∣xt)p_{\sigma}(\cdot|x_{t}) and p(⋅∣xt)p(\cdot|x_{t}) are continuous functions on a compact domain. Therefore, ∂xtpσ(⋅∣xt)\partial_{x_{t}}p_{\sigma}(\cdot|x_{t}) is bounded from above by some constant MM. Moreover, by adding assumption 4 we obtain that pσ(⋅∣xt)p_{\sigma}(\cdot|x_{t}) and p(⋅∣xt)p(\cdot|x_{t}) are bounded from below by some ϵ>0\epsilon>0. Therefore

where E1E_{1} and E2E_{2} are monotonically decreasing to zero. The theorem follows with E(1/σy(t)2):=Mϵ2E1(1/σy(t)2)+1ϵE2(1/σy(t)2)+L2σy(t)2E(1/\sigma^{y}(t)^{2}):=\frac{M}{\epsilon^{2}}E_{1}(1/\sigma^{y}(t)^{2})+\frac{1}{\epsilon}E_{2}(1/\sigma^{y}(t)^{2})+L^{2}\sigma^{y}(t)^{2}, which monotonically decreases to zero as σy(t)2\sigma^{y}(t)^{2} decreases to zero. ∎

Appendix C Architectures and hyperparameters

We used almost the same neural network architecture across all tasks and all estimators, so that we can compare the estimators fairly. The only difference between the score model for the diffusive estimators and the score model for the CDE estimator is that the former contains 66 instead 33 filters in the final convolution to account for the joint score estimation. This difference in the final convolution leads to negligible difference in the number of parameters, which is highly unlikely to have impacted the final performance.

We used the basic version of the DDPM architecture with the following hyperparameters: channel dimension 9696, depth multipliers $,,2ResNetBlocksperscaleandattentioninthefinalResNet Blocks per scale and attention in the final3scales.Thetotalparametercountis43.5M.Songetal.reportimprovedperformancewiththeNCSN++architectureoverthebaselineDDPMwhentrainingwiththeVESDE.ThisclaimisalsosupportedbytheworkofSahariaetal..Therefore,adoptingthisarchitectureislikelytoimprovetheperformanceofallestimatorsandleadtoevenmorecompetitiveperformanceoverstate−of−the−artmethods.Forallestimators,weconcatenatetheconditionimagescales. The total parameter count is 43.5M. Song et al. report improved performance with the NCSN++ architecture over the baseline DDPM when training with the VE SDE. This claim is also supported by the work of Saharia et al. . Therefore, adopting this architecture is likely to improve the performance of all estimators and lead to even more competitive performance over state-of-the-art methods. For all estimators, we concatenate the condition imageyorory(t)withthediffusedtargetwith the diffused targetx(t)$ and pass the concatenated image as input to the score model for score calculation. In the super-resolution experiment, we first interpolate the condition to the same resolution as the target using nearest neighbours interpolation and then concatenate it with the target image.

We used exponential moving average (EMA) with rate 0.999 and the same optimizer settings as in . Moreover, we used a batch size of 5050 for the super-resolution and edge to image translation experiments and a batch size of 100100 for the inpainting experiments.

Appendix D Extended visual results

We provide additional samples in Figures 9, 10 and 11.

Appendix E Potential negative impact

The potential of negative impact of this work is the same as that of any work that advances generative modeling. Generative modeling can be used for the creation of deep-fakes which can be used for malicious purposes such as disinformation and blackmailing. However, research on generative modeling can indirectly or directly contribute to the robustification of deep-fake detection algorithms. Moreover, generative models have proven very useful in academic research and in industry. The potential benefits of generative modeling outweigh the potential threats. Therefore, the research community should continue to conduct research on generative modeling.