Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMC

Yilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum, Sander Dieleman, Rob Fergus, Jascha Sohl-Dickstein, Arnaud Doucet, Will Grathwohl

Introduction

In recent years, tremendous progress has been made in generative modeling across a variety of domains (Brown et al. 2020; Brock et al. 2018; Ho et al. 2020). These models now serve as powerful priors for downstream applications such as code generation (Li et al. 2022), text-to-image generation (Saharia et al. 2022), question-answering (Brown et al. 2020) and many more. However, to fit these complex data, generative models have grown inexorably larger (requiring 10’s or even 100’s of billions of parameters) (Kaplan et al. 2020) and require datasets containing non-negligible fractions of the entire internet, making it costly and difficult to train and or finetune such models. Despite this, some of the most compelling applications of large generative models do not rely on finetuning. For example, prompting (Brown et al. 2020) has been a successful strategy to selectively extract insights from large models. In this paper, we explore an alternative to finetuning and prompting, through which we may repurpose the underlying prior learned by generative models for downstream tasks.

Diffusion Models (Sohl-Dickstein et al. 2015; Song & Ermon 2019; Ho et al. 2020) are a recently popular approach to generative modeling which have demonstrated a favorable combination of scalability, sample quality, and log-likelihood. A key feature of diffusion models is the ability for their sampling to be “guided” after training. This involves combining the pre-trained Diffusion Model pθ(x)p_{\theta}(x) with a predictive model pθ(y∣x)p_{\theta}(y|x) to generate samples from pθ(x∣y)p_{\theta}(x|y). This predictive model can be either explicitly defined (such as a pre-trained classifier) (Sohl-Dickstein et al. 2015; Dhariwal & Nichol 2021) or an implicit predictive model defined through the combination of a conditional and unconditional generative model (Ho & Salimans 2022). These forms of conditioning are particularly appealing (especially the former) as they allow us to reuse pre-trained generative models for many downstream applications, beyond those considered at training time.

These conditioning methods are a form of model composition, i.e. combining probabilistic models together to create new models. Compositional models have a long history back to early work on Mixtures-Of-Experts (Jacobs et al. 1991) and Product-Of-Experts models (Hinton 2002; Mayraz & Hinton 2000). Here, many simple models or predictors were combined to increase their capacity. Much of this early work on model composition was done in the context of Energy-Based Models (Hinton 2002), an alternative class of generative model which bears many similarities to diffusion models.

In this work, we explore the ways that diffusion models can be reused and composed with one-another. First, we introduce a set of methods which allow pre-trained diffusion models to be composed, with one-another and with other models, to create new models without retraining. Second, we illustrate how existing methods for composing diffusion models are not fully correct, and propose a remedy to these issues with MCMC-derived sampling. Next, we propose the use of an energy-based parameterization for diffusion models, where the unnormalized density of each reverse diffusion distribution is explicitly modeled. We illustrate how this parameterization enables both additional ways to compose diffusion models, as well as the use of more powerful Metropolis-adjusted MCMC samplers. Finally, we demonstrate the effectiveness of our approach in settings from 2D data to high-resolution text-to-image generation. An illustration of our domains can be found in Figure 1.

Background

Diffusion models seek to model a data distribution q(x0)q(x_{0}). We augment this distribution with auxiliary variables {xt}t=1T\{x_{t}\}_{t=1}^{T} defining a Gaussian diffusion q(x0,…,xT)=q(x0)q(x1∣x0)…q(xT∣xT−1)q(x_{0},\ldots,x_{T})=q(x_{0})q(x_{1}|x_{0})\ldots q(x_{T}|x_{T-1}) where each transition is defined by q(xt∣xt−1)=N(xt;1−βtxt−1,βtI)q(x_{t}|x_{t-1})=\mathcal{N}(x_{t};\sqrt{1-\beta_{t}}x_{t-1},\beta_{t}I) for some 0<βt≤10<\beta_{t}\leq 1. This transition first scales down xt−1x_{t-1} by 1−βt\sqrt{1-\beta_{t}} and then adds Gaussian noise of variance βt\beta_{t}. For large enough TT, we will have q(xT)≈N(0,I)q(x_{T})\approx\mathcal{N}(0,I).

A useful feature of the diffusion distribution q(x0,...,xT)q(x_{0},...,x_{T}) is that we can analytically derive any time marginal q(xt∣x0)=N(xt;1−σt2x0,σt2I)q(x_{t}|x_{0})=\mathcal{N}(x_{t};\sqrt{1-\sigma_{t}^{2}}x_{0},\sigma^{2}_{t}I) where again σt\sigma_{t} is a function of {βt}t=1T\{\beta_{t}\}_{t=1}^{T}. We can sample xtx_{t} from this distribution using reparameterization, i.e xt(x0,ϵ)=1−σt2x0+σtϵx_{t}(x_{0},\epsilon)=\sqrt{1-\sigma^{2}_{t}}x_{0}+\sigma_{t}\epsilon where ϵ∼N(0,I)\epsilon\sim\mathcal{N}(0,I). Exploiting this, diffusion models are typically trained with the loss L(θ)=∑t=1TLt(θ)\mathcal{L}(\theta)=\sum_{t=1}^{T}\mathcal{L}_{t}(\theta), where

Once ϵθ(x,t)\epsilon_{\theta}(x,t) is trained, we recover μθ(x,t)\mu_{\theta}(x,t) with Equation 1 to parameterize pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) and perform ancestral sampling (also known as the reverse process) to reverse the diffusion, i.e sample xT∼N(0,I)x_{T}\sim\mathcal{N}(0,I), then for t=T−1→1t=T-1\rightarrow 1, sample xt−1∼pθ(xt−1∣xt)x_{t-1}\sim p_{\theta}(x_{t-1}|x_{t}). A more detailed description can be found in Appendix A.

2 Energy-Based Models and MCMC Sampling

When minimized, this ensures that efθ(x)∝qσ(x)e^{f_{\theta}(x)}\propto q_{\sigma}(x) and therefore ∇xfθ(x)=∇xlog⁡qσ(x)\nabla_{x}f_{\theta}(x)=\nabla_{x}\log q_{\sigma}(x). To estimate likelihoods or sample from our model, we must rely on approximate methods, such as MCMC sampling or numerical ODE integration. MCMC works by simulating a Markov chain beginning at x0∼p(x0)x_{0}\sim p(x_{0}) and using a transition distribution xt∼k(xt∣xt−1)x_{t}\sim k(x_{t}|x_{t-1}). If k(⋅∣⋅)k(\cdot|\cdot) has certain properties, namely invariance w.r.t. the target and ergodicity, then as t→∞t\rightarrow\infty, xtx_{t} converges to a sample from our target distribution.

Perhaps the most popular MCMC sampling algorithm for EBMs is Unadjusted Langevin Dynamics (ULA) (Roberts & Tweedie 1996; Du & Mordatch 2019; Nijkamp et al. 2020) which is defined by

This resembles a step of gradient ascent (with step-size σL22\frac{\sigma_{L}^{2}}{2}) with added Gaussian noise of variance σL2\sigma_{L}^{2}. This transition is based on a discretization of the Langevin SDE. In the limit of infinitesimally small σ\sigma this approach will draw exact samples. To handle the error accrued when using larger step sizes, a Metropolis correction can be added giving the Metropolis-Adjusted-Langevin-Algorithm (MALA) (Besag 1994). With Metropolis correction, we first generate a proposed update x^∼k(x∣xt−1)\hat{x}\sim k(x|x_{t-1}), then with probability min⁡(1,efθ(x^)efθ(xt−1)k(xt−1∣x^)k(x^∣xt−1))\min\left(1,\frac{e^{f_{\theta}(\hat{x})}}{e^{f_{\theta}(x_{t-1})}}\frac{k(x_{t-1}|\hat{x})}{k(\hat{x}|x_{t-1})}\right) we set xt=x^x_{t}=\hat{x}, otherwise xt=xt−1x_{t}=x_{t-1}.

Hamiltonian Monte Carlo (HMC) (Duane et al. 1987; Neal 1996) is a more advanced MCMC sampling method which augments the state-space with auxiliary momentum variables and numerically integrates energy-conserving Hamiltonian dynamics to explore the space. HMC is typically applied with a Metropolis correction, but an approximate variant can be used without it (U-HMC) (Geffner & Domke 2021). See Appendix B.1 for details of HMC variants we use.

3 Connection Between Diffusion Models and EBMs

Diffusion models and EBMs are closely related. For instance, Song & Ermon 2019 uses an EBM perspective to propose a close cousin to diffusion models. We can see from inspection that the training objective of diffusion models is identical (up to a constant) to the denoising score matching objective

where we have replaced ϵθ(x,t)\epsilon_{\theta}(x,t) with −σt∇xfθ(x+σtϵ)-\sigma_{t}\nabla_{x}f_{\theta}(x+\sigma_{t}\epsilon). Thus by training ϵθ(x,t)\epsilon_{\theta}(x,t) to minimize Equation 2, we can recover the diffused data distribution score with ∇xlog⁡qσt(x)≈−ϵθ(x,t)σt\nabla_{x}\log q_{\sigma_{t}}(x)\approx-\frac{\epsilon_{\theta}(x,t)}{\sigma_{t}}. From this, we can define ϵθ(x,t)=−∇xfθ(x,t)\epsilon_{\theta}(x,t)=-\nabla_{x}f_{\theta}(x,t) (the derivative of an explicitly defined scalar function) to learn a noise-conditional potential function fθ(x,t)f_{\theta}(x,t). We later demonstrate the benefits of this in two ways; it enables the use of more sophisticated sampling algorithms and more forms of composition.

4 Controllable Generation

It may be convenient to train a model of p(x)p(x) where xx is, say, the distribution of all images, but in practice we often want to generate samples from p(x∣y)p(x|y) where yy is some attribute, label, or feature. This can be accomplished within the framework of diffusion models by introducing a learned predictive model pθ(y∣x;t)p_{\theta}(y|x;t), i.e a time-conditional model of the distribution of some feature yy given xx. We can then exploit Bayes’ rule to notice that (for λ=1\lambda=1),

In practice, when using the right side of Equation 6 for sampling, it is beneficial to increase the ‘guidance scale’ λ\lambda to be >1>1 (Dhariwal & Nichol 2021). Thus, we can re-purpose the unconditional diffusion model and turn it into a conditional model.

If instead of a classifier, we have a both an unconditional diffusion model ∇xlog⁡pθ(x;t)\nabla_{x}\log p_{\theta}(x;t) and a conditional diffusion model ∇xlog⁡pθ(x∣y;t)\nabla_{x}\log p_{\theta}(x|y;t), we can again utilize Bayes’ rule to derive an implicit predictive model’s gradients

which can be used to replace the explicit model in Equation 6, giving what is known as classifier-free guidance (Ho & Salimans 2022). This method has led to incredible performance, but comes at a cost to modularity. This contrasts with the classifier-guidance setting, where we only need to train a single (costly) generative model. We can then attach any predictive model we would like to for conditioning. This is beneficial as it is often much easier and cheaper to train predictive models than a flexible generative model. In the classifier-free setting, we must know exactly which yy we would like to condition on, and incorporate these labels into model training. In both guidance settings, we use our (possibly implicit) predictive model to modify the learned score of our model. We then perform diffusion model sampling as we would in the unconditional setting. We will see later that even in toy settings, this is often not the optimal thing to do.

Compositional Generation Beyond Guidance

Most work on conditional diffusion models has come in the form of classifier or classifier-free guidance, but these are far from the only ways we can compose distributions to obtain new models. These ideas have been studied primarily in the context of EBMs because most compositional operators leave the resulting distribution unnormalized. We outline various options below.

Products: We can take a product of NN distributions and re-normalize to create a new distribution, roughly equivalent to the “intersection” of the composite distributions,

Regions of high probability under qprod(x)q^{\text{prod}}(x) will typically have high probability under all qi(x)q^{i}(x). A simple product model can be seen in Figure 2. These ideas were initially proposed to increase the capacity of weaker models by allowing individual “experts” to model specific features in the input (Hinton 2002), and were recently demonstrated at scale in the image domain using Deep Energy-Based Models (Du et al. 2020a).

The approaches to guidance discussed in Section 2.4 define product models with only two experts. The first models the relative density of the input data and the second models the conditional probability of yy. Combining these by a product models likely inputs which have the desired property yy. This form of composition has become popular for diffusion models since they do not directly model the probability, but instead the gradient of the log-probability which can also be composed in this way.

Mixtures: Complementary to the product or intersection is the mixture or union of multiple distributions. We can combine NN distributions through a mixture to create a new distribution equivalent to the union of the concepts captured in each distribution

where regions of high probability consist of regions of high probability under any qi(x)q^{i}(x). We cannot compose score-functions to define mixtures (unlike products). Instead, we need a model which specifies probability. Generating from mixtures of energy based models requires knowing the ratio of normalizers between the models. In our experiments, we assume this ratio is 1, though a different unknown ratio of normalizers would correspond to a weighted mixture. A simple compositional mixture model can be seen in Figure 2.

Negation: Finally, given two distributions p0(x)p_{0}(x) and p1(x)p_{1}(x), we can explicitly invert the density of p1(x)p_{1}(x) with respect to p0(x)p_{0}(x), which constructs a new distribution which assigns high likelihood to points in p0(x)p_{0}(x) that are not in p1(x)p_{1}(x) (Du et al. 2020a), where α\alpha controls the degree we invert p1(x)p_{1}(x) (we use α=0.5\alpha=0.5 in our experiments).

We can combine negation with our previous operators, in a nested manner to construct complex combinations of distributions (Figure 6).

In Section 2.3, we showed how diffusion models can be interpreted as approximating the gradient ∇xlog⁡q(x)\nabla_{x}\log q(x), but do not learn an explicit model of the log-likelihood log⁡q(x)\log q(x). This means with the standard ϵθ(x,t)\epsilon_{\theta}(x,t)-parameterization we can, in theory, utilize product and negation composition, but not mixture composition.

Scaling Compositional generation with Diffusion Models

While highly compositional, EBMs present many challenges. The lack of a normalized probability function makes training, evaluation, and sampling very difficult. Much progress has been made to scale these models (Du & Mordatch 2019; Nijkamp et al. 2020; Grathwohl et al. 2019; Du et al. 2020b; Grathwohl et al. 2021; Gao et al. 2021; Xie et al. 2016; Xie et al. 2021), but EBMs still lag behind other approaches in terms of efficiency and scalability. In contrast, diffusion models have demonstrated very impressive scalability. Fortuitously, diffusion models have similarities to EBMs, such as their training objective and their score-based interpretation, which makes many forms of composition readily applicable.

Unfortunately, when two diffusion models are composed into, for example, a product model qprod(x)∝q1(x)q2(x)q^{\textup{prod}}(x)\propto q^{1}(x)q^{2}(x), issues arise if the model which reverses the diffusion uses a score estimate obtained by adding the score estimates of the two models as done in prior work (Liu et al. 2022). We see in Figure 2 that composing two models in such a way leads indeed to sub-par samples. This is because to sample from this product distribution using standard reverse diffusion (Song et al. 2021), one would need to compute instead the score of the diffused target product distribution given by

For t>0t>0, this quantity is not equal to the sum of the scores of the two models which is given by

Therefore, plugging the composed score function into the standard diffusion ancestral sampling procedure discussed in Section 2.1, which we refer to as “reverse diffusion,” does not correspond to sampling from the composed model, and thus reverse diffusion sampling will generate incorrect samples from composed distributions. This effect can be seen in Figure 2, with details in Appendix C.

In order to sample from qprod(x)q^{\textup{prod}}(x) using the combined score function from Equation 12, we can use annealed MCMC sampling, described below in Algorithm 1. This method applies MCMC transition kernels to a sequence of distributions which begins with a known, tractable distribution and concludes at our target distribution. Annealed MCMC has a long history enabling sampling from very complex distributions (Neal 2001).

We explore two types of transition kernels kt(⋅∣⋅)k_{t}(\cdot|\cdot) based on Langevin Dynamics (Equation 4) and HMC. When using the standard ϵθ(x,t)\epsilon_{\theta}(x,t)-parameterization, we do not have access to an explicitly defined energy-function meaning we cannot utilize any MCMC sampler with Metropolis corrections. Thus, we only utilize the ULA and U-HMC samplers described in Section 2.2. These samplers are not exact, but can in practice generate good results. In the next section we detail how Metropolis corrections may be incorporated. Full details of our samplers can be found in Appendix B.1. A simple form of MCMC using the ULA sampler can directly be implemented by repeatedly running the diffusion reverse process at the same noise level with composed score functions, which we mathematically illustrate in Appendix B.2, providing a simple method to properly implement composition between diffusion models.

We can see again in Figure 2 that applying this MCMC sampling procedure allows samples from the composed distribution to be faithfully generated with no modification to the underlying diffusion models. Quantitative results can be found in Table 1 which further imply that the choice of sampler may be responsible for prior failures in compositional generation with diffusion models. Concurrent to our work, Geffner et al. 2022 also explores using MCMC to sample from diffusion product distributions.

2 Energy-Based Parameterization

As noted in Section 3, we are unable to use mixture composition without an explicit likelihood function. But, if we parameterize a potential function fθ(x,t)f_{\theta}(x,t) and implicitly define ϵθ(x,t)=−∇xfθ(x,t)\epsilon_{\theta}(x,t)=-\nabla_{x}f_{\theta}(x,t) we can recover an explicit estimate of the (unnormalized) log-likelihood – enabling us to utilize all presented forms for model composition.

Additionally, an explicit estimate of log-likelihood enables the use of more accurate samplers. As explained above, with the standard ϵθ(x,t)\epsilon_{\theta}(x,t)-parameterization we can only utilize unadjusted samplers. While they can perform well in practice, there exist many distributions from which they cannot generate decent samples (Roberts & Tweedie 1996) such as targets with lighter-than-Gaussian tails where the ULA chain is transient. Additionally, for an accurate approximation to the Langevin SDE, ULA will need increasingly small stepsizes as the curvature of the log-likelihood gradient increases which can lead to arbitrarily slow mixing (Durmus & Moulines 2019). In these settings a Metropolis correction can greatly improve sample quality and convergence. Again this issue can be solved by defining ϵθ(x,t)=−∇xfθ(x,t)\epsilon_{\theta}(x,t)=-\nabla_{x}f_{\theta}(x,t) for some explicitly defined scalar potential function fθ(x,t)f_{\theta}(x,t).

Energy-based parameterizations have been explored in the past (Salimans & Ho 2021) and were found to perform comparably to score-based models for unconditional generative modeling. In that setting the score parameterization is then preferable as computing the gradient of the energy requires more computation. In the compositional setting, however, the additional flexibility enabled by explicit (unnormalized) log-probability estimation motivates a re-exploration of the energy-parameterization.

We explored a number of energy-based parameterizations for diffusion models and ran a pilot study on ImageNet. In this study we found it best to parameterize the log probability as fθ(x,t)=−∣∣sθ(x,t)∣∣2f_{\theta}(x,t)=-||s_{\theta}(x,t)||^{2}, where sθ(x,t)s_{\theta}(x,t) is a vector-output neural network, like those used in ϵθ(x,t)\epsilon_{\theta}(x,t)-parameterized diffusion models. Full details on our study can be found in Appendix D. From here on, all energy-based diffusion models take the above form.

Our energy-parameterized models enable us to use MALA and HMC samplers which produce our best compositional generation results. An additional benefit of these samplers is that, through monitoring their acceptance rates, we are able to derive an effective automated method for tuning their hyper-parameters (a notoriously difficult task prior) which is not available for unadjusted samplers. Details of our samplers and tuning procedures can be found in Appendix B.1.

Experiments

We experiment with various model parameterizations and sampling schemes for compositional generation with diffusion models. We first investigate these ideas on some illustrative 2D datasets, then move to the image domain with an artificial dataset of shapes. Here, we compose a model conditioned on the location of a single shape with itself to condition on the location of all of the shapes in the image. After this we experiment with classifier guidance on the ImageNet dataset. Finally, we self-compose text-to-image models to generate from compositions of various text prompts and image tapestries. Full details of all experiments can be found in Appendix F. Throughout we compare our proposed improvements with a score-parameterized model using standard reverse diffusion sampling. We note that this baseline is exactly the approach of Liu et al. 2022.

We train diffusion models using both parameterizations and study the impact of various sampling approaches for compositional generation. Samples are evaluated using RAISE (Burda et al. 2015) (which gives lower bounds on log-likelihood) and MMD For the mixture we use MMD to replace RAISE likelihood based evaluation as we encountered numerical stability issues with RAISE when applying to the mixture., LL (log-likelihood of generated samples under composed distribution), and Var (L2 difference of variance of GMMs fit on generated samples compared to GMMs of the composed distribution). Results can be found in Table 1 and visualizations can be seen in Figure 2. All MCMC sampling methods improve sample quality and likelihood, with Metropolis adjusted methods performing the best. All MCMC experiments use the same number of score function evaluations. We include a baseline, labeled “Reverse (equal steps)” which is a diffusion model trained with more steps such that reverse diffusion sampling has the same cost as our MCMC samplers. We see that simply adding more time-steps does not solve compositional sampling.

2 Composing Cubes

Next, we train models on a dataset of images containing between 1 and 5 examples of various shapes taken from CLEVR (Johnson et al. 2017). We train our models to fit p(x∣y)p(x|y) where yy is the location of one of the shapes in the image. We then compose this conditional model with itself to create a product model which defines the distribution of images conditioned on cc shapes as a distribution log⁡pθ(x∣y1,…,yc)\log p_{\theta}(x|y_{1},\ldots,y_{c}) equal to

We then sample using various methods, where for each number of combination of cubes, the same number of score function evaluations are used, and evaluate each by the fraction of samples which have all objects placed in the correct location (as determined by a learned classifier). Results can be found in Table 2, where we see MCMC sampling leads to improvements and the Metropolis adjustment enabled by the energy-based parameterization leads to further improvements. We qualitatively illustrate results in Figure 3, and see more accurate generations with more steps of sampling, with more substantial increases with Metropolis adjustment.

3 Classifier conditioning

Next, we train unconditional diffusion models and a noise-conditioned classifier on ImageNet. We compose these models as

and sample using the corresponding score functions. We compare various samplers and model parameterizations on classifier accuracy, FID (Heusel et al. 2017) and Inception Score. Quantitative results can be seen in Table 3 and qualitative results seen in Figure 4. We find that MCMC improves performance over reverse sampling, with further improvements from Metropolis corrections.

4 Text-2-Image

Perhaps the most well-known results achieved with diffusion models are in text-to-image generation (Ramesh et al. 2022; Saharia et al. 2022). Here we model pθ(ximage∣ytext)p_{\theta}(x_{\text{image}}|y_{\text{text}}). While generated images generated are photo-realistic, they can fail to generate images from prompts which specify multiple concepts at a time (Liu et al. 2022) such as ytext=y_{\text{text}}= “A horse on a sandy beach or a grass plain on a not sunny day”. To deal with these issues we can dissect the prompt into smaller components y1,…,ycy_{1},\ldots,y_{c}, parameterize models conditioned on each component pθ(x∣yi)p_{\theta}(x|y_{i}) and compose these models using our introduced operators. We can parse the above example into

which can be used to define the following (unnormalized) distribution pθcomp(x∣ytext)p^{\text{comp}}_{\theta}(x|y_{\text{text}})

Liu et al. 2022 demonstrated that composing models this way can improve the efficacy of these kinds of generations, but was restricted to composition using classifier-free guidance. We train a energy-parameterized diffusion model for text conditional 64x64 image generation and illustrate composed results in Figure 6 (upsampled to 1024x1024). We find that composition enables more faithful generations of scenes in Figure 7 with more results in Appendix G.

5 Tapestries of Energy Functions

While text-to-image models generate images given natural language prompts, it is difficult to control the spatial location of different content, and difficult to generate images at higher resolutions than used during training. By composing multiple overlapping text-to-image models, at a variety of scales, we may construct an image tapestry with different specified content at different locations and scales. We illustrate in Figure 8 using this approach to generate images with content at specified spatial locations and scales. See Appendix F for details.

Discussion

Limitations. Our work demonstrates that diffusion models, in combination with MCMC-based sampling procedures, can be composed in novel ways capable of generating high-quality samples. However, our proposed solutions have a number of drawbacks. First, more sophisticated MCMC samplers come at a higher cost than the standard sampling approach and can take 5-times longer to generate samples than typical diffusion sampling. Second, we have shown that energy-parameterized models enable the use of more sophisticated sampling techniques, garnering further improvements. Unfortunately, this requires a second backward-pass through the model to compute the derivative implicitly, leading them to have double the memory and compute cost of score-parameterized models.

While these are considerable drawbacks, we note the focus of this work is to demonstrate that such things are possible within the framework of diffusion models. We believe there is much that can be done to achieve the benefits of our sampling procedures at less cost such as distillation (Salimans & Ho 2022) and easier-to-differentiate neural networks (Chen & Duvenaud 2019).

Conclusion. We have explored the ways that pretrained diffusion models can be composed to model new distributions. We demonstrate ways that naïve implementations fail, and present two ways that performance can be improved: MCMC sampling and energy-parameterized diffusion models. Our approach leads to notable improvement across a variety of domains, scales, and compositional operators. In future work, it would further be interesting to explore how such compositional operations may further be applied to related generative models (Wang et al. 2021).

Acknowledgements. We gratefully acknowledge support from NSF grant 2214177; from AFOSR grant FA9550-22-1-0249; from ONR MURI grant N00014-22-1-2740; from ARO grant W911NF-23-1-0034; from the MIT-IBM Watson Lab; and from the MIT Quest for Intelligence. Yilun Du is supported by a NSF Graduate Research Fellowship.

References

Appendix A Detailed Derivation of Diffusion Models

Diffusion models seek to model a data distribution q(x0)q(x_{0}) (written this way for notational convenience) and define a series of latent variables x1,…,xTx_{1},\ldots,x_{T} generated from a Markov process xt∼q(xt∣xt−1)x_{t}\sim q(x_{t}|x_{t-1}) where

A unique and useful property of this process is that all time marginals q(xt∣x0)q(x_{t}|x_{0}) can be computed in closed form and are Gaussian

where σt2=1−αˉt\sigma^{2}_{t}=1-\bar{\alpha}_{t} and αˉt=∏t=1T(1−βt)\bar{\alpha}_{t}=\prod_{t=1}^{T}(1-\beta_{t}). We can see that if all βt>0\beta_{t}>0, then as t→∞t\rightarrow\infty q(xt∣x0)→N(xt;0,I)q(x_{t}|x_{0})\rightarrow\mathcal{N}(x_{t};0,I).

and set p(xT)=N(0,I)p(x_{T})=\mathcal{N}(0,I). We train this model to maximize a variational bound on the marginal likelihood

The first term lacks parameters and the final is approximately 0 from the convergence of the qq-process so we focus on the middle terms.

Typically, the model μθ(xt−t,t)\mu_{\theta}(x_{t-t},t) is not parameterized to predict the mean of xtx_{t}. Instead it is parameterized to predict the noise added to xtx_{t} to arrive at xt−1x_{t-1}. This motivates the following parameterization

In this form we can rewrite the important terms in the objective as

where CtC_{t} is a time-dependent constant. Typically these are dropped and all objectives are weighted equally.

Appendix B MCMC Sampling Details

Hamiltonian Monte Carlo (Neal 1996) seeks to sample from an unnormalized probability distribution log⁡p(x)=f(x)+log⁡Z{\log p(x)=f(x)+\log Z}. To do this, we augment our distribution over xx with auxiliary variables vv and define the joint distribution p(x,v)=p(x)N(v;0,M)p(x,v)=p(x)\mathcal{N}(v;0,M) where the covariance MM is known as the “mass-matrix.” We now seek to draw samples x,v∼p(x,v)x,v\sim p(x,v) and since xx and vv are independent under the joint, we can simply throw away our vv samples leaving us with a sample x∼p(x)x\sim p(x).

Like other MCMC methods we sequentially update a particle (xi,vi)(x^{i},v^{i}) in such a way that as i→∞{i\rightarrow\infty} we arrive at a sample from p(x,v)p(x,v). For a step of HMC, starting at (xi,vi)(x^{i},v^{i}) we first sample vi′∼N(vi;0,M){v^{i^{\prime}}\sim\mathcal{N}(v^{i};0,M)} since the target distribution factorizes and p(v)p(v) is known and tractable. We then integrate a Hamiltonian-conserving ODE defined on x,vx,v known as “Hamiltonian Dynamics.” We can use the leapfrog integrator which will guarantee which is a symplectic integrator. Thus, the Metropolis acceptance probability simplifies to min⁡(1,p(x′,v′)p(x,v))\min\left(1,\frac{p(x^{\prime},v^{\prime})}{p(x,v)}\right). An overview of the HMC algorithm can be found in Algorithm 2. We refer the reader to Neal 1996 for a more complete description of the algorithm.

Since our ϵθ(x,t)\epsilon_{\theta}(x,t) parameterized models do not admit an explicit likelihood function, we use an unadjusted variant of HMC (U-HMC) where the accept/reject step is simply ignored.

We can see in Algorithm 2 that at every step, the momentum is re-sampled. This can be sub-optimal, as the momentum determines the initial direction of xx’s movement and if a good direction is found, it may be beneficial to continue in that direction. To deal with this, Neal 1996 presents a variant of HMC where the momentum vv is partially retained between sampling steps. We add an additional sampler parameter γ∈\gamma\in (known as the “damping-factor”) which controls the amount to which vv is retained. When γ\gamma is close to 1, vv is mostly kept and when it is near 00, vv is mostly refreshed. This variant is summarized in Algorithm 3. The (potentially confusing) momentum negations ensure the validity of the sampler. Intuitively, when the proposal is accepted, the momentum is retained and when it is rejected the momentum is flipped. For this reason, one should maintain a reasonably high acceptance rate when using this approach.

B.2 Implementing ULA Sampling with Multiple Reverse Diffusion Steps

We illustrate how a step of diffusion reverse sampling step at the same fixed noise level (from timestep tt to timestep tt) is equivalent to ULA MCMC sampling at the same fixed noise level. We use the αt\alpha_{t} and βt\beta_{t} formulation from (Ho et al. 2020). The reverse sampling step on an input xtx_{t} at a fixed noise level at timestep tt is given by a Gaussian with a mean

with the variance of βt\beta_{t} (using the variance small noise schedule in (Ho et al. 2020)). This corresponds to a sampling update,

Note that the expression scaled denoising function ϵθ(xt,t)1−αˉt\frac{\epsilon_{\theta}(x_{t},t)}{\sqrt{1-\bar{\alpha}_{t}}} corresponds to the score function for the composed distribution ∇xpt(x)\nabla_{x}p_{t}(x), through the denoising score matching objective (Vincent 2011). The reverse sampling step can be equivalently written as

The ULA sampler draws an MCMC sample from the composed EBM probability distribution pt(x)p_{t}(x) using the expression

where η\eta is the step size of sampling.

By substituting η=βt\eta=\beta_{t} in the ULA sampler, the sampler becomes

Note the similarity of ULA sampling in Eqn A10 and the diffusion reverse sampling procedure in Eqn A8, where there is a factor of 2\sqrt{2} scaling of the added Gaussian noise in the ULA sampling procedures. This means that we can implement the ULA sampling on a composed score function by running the standard diffusion reverse process with the variance small noise schedule, but by scaling the noise added in each timestep by a factor of 2\sqrt{2} (which is an intermediate noise variance level between small and large variance noise schedule). Alternatively, we can directly use the standard diffusion reverse process in Eqn A8 to run ULA, where this then corresponds to sampling a tempered variant of the composed distribution pt(x)p_{t}(x) with temperature 12\frac{1}{\sqrt{2}} (corresponding to less stochastic samples from the final composed probability distribution).

This equivalence allows us to easily implement MCMC sampling-based composition with existing diffusion codebases. Given a composed score function, we run the standard diffusion sampling procedure, but simply run the reverse sampling procedure multiple times at a noise level before transitioning to the next noise level.

B.3 MCMC Tuning

A crucial component to ensure successful MCMC sampling in diffusion models is the choice of step sizes for samplers. We initialize step sizes for all samplers at each distribution tt to be roughly proportional to the βt\beta_{t} noise values added to distribution tt in the diffusion process.

To tune step sizes across timesteps for both HMC and MALA samplers, to set step sizes at each timestep tt to be constant multiplied by βt\beta_{t}. We searched different constants to multiply βt\beta_{t}, and chose a value so that the average acceptance rate of MALA and HMC samplers across timesteps is approximately 60% and 70% respectively. For un-adjusted variants of these samplers, we set step sizes to be the same as adjusted samplers, and found limited gains when step sizes were specifically tuned towards the un-adjusted samplers. We utilize a mass matrix of βt\beta_{t} for HMC samplers.

Precise details on the exact MCMC steps sizes used in experiments can be detailed in Section F.

B.4 MCMC Implementation Details

When initially running MCMC sampling on diffusion models in the image domain, we found that our samplers tended to converge to images that had uniform textures. After experimentation, we found that the primary cause of this issue was the fact that by default, typical implementations of the reverse diffusion process clip samples at intermediate time-steps of sampling to be between -1 and 1. To enable proper MCMC sampling, we found that it was important to not clip intermediate values of diffusion sampling.

When running MCMC sampling on image domains, we further found that it was helpful for mixing to run a single step of the reverse process to initialize MCMC sampling, before running many steps of MCMC sampling at each timestep tt, and in all MCMC sampling settings on the image domain, we run one step of the reverse process before running MCMC sampling. This reverse sampling step serves to contract a sample from one noise level to another, and is close in form to Langevin sampling (Section B.2).

Appendix C Compositional Diffusions

In Equations 11 and 12, we demonstrate that for diffused distributions {qti(xt)}\{q^{i}_{t}(x_{t})\} where qti(xt)=∫qi(x0)q(xt∣x0)dx0q^{i}_{t}(x_{t})=\int q^{i}(x_{0})q(x_{t}|x_{0})dx_{0}, the diffusion of the product of qiq^{i}’s is not the same as the product of the diffusions, meaning plugging the product of diffusions into standard reverse diffusion sampling will not draw samples from the product model. We present similar results for tempering and predictive model composition.

It is tempting to believe that we can sample from a tempered/annealed version of the data distribution

using the tempered diffused data distribution λ∇log⁡tq(xt)\lambda\nabla\log_{t}q(x_{t}) but this is incorrect. For this procedure to be correct, we would need to have ∇log⁡qtλ(xt)=λ∇log⁡qt(xt)\nabla\log q_{t}^{\lambda}(x_{t})=\lambda\nabla\log q_{t}(x_{t}) for all tt. However, while we do have ∇log⁡q0λ(x0)=λ∇log⁡q0(x0)\nabla\log q_{0}^{\lambda}(x_{0})=\lambda\nabla\log q_{0}(x_{0}), this equality does not hold for t>0t>0

C.2 Guidance

For conditional generation, we should use in the reverse diffusion the score of the diffused conditional distribution ∇log⁡qt(xt∣y)\nabla\log q_{t}(x_{t}|y) where

allows you to do guidance without having to train say a classifier if yy is categorical.

In practice, it was found that using in the reverse time diffusion the score

generates much nicer images for λ>1\lambda>1. However, it is also often claimed that it samples from a modified posterior where the likelihood has been annealed. This is incorrect. For a modified posterior with annealed likelihood, we would have

C.3 Sampling from Composed Distributions

We can see that products, tempering, and guidance applied to diffused distributions do not give diffusions of the modified target distributions. Thus, we should not expect to arrive at our desired result by applying a sampling procedure which reverses a diffusion applied to the target distribution. Thankfully, as stated in Section 4, these operators do give us a sequence of distributions which anneals from N(0,I)\mathcal{N}(0,I) to the composed target which means we can utilize the family of annealed MCMC sampling methods mentioned in Section 4.1 to draw samples from our composed models in all of these settings, directly using the available score estimate.

Appendix D Energy-Based Parameterizations

As mentioned in section 4.2, when using the ϵθ(x,t)\epsilon_{\theta}(x,t) parameterization, we can recover an estimate of the time-conditional score function with ∇xlog⁡pt(x)≈−ϵθ(x,t)σt\nabla_{x}\log p_{t}(x)\approx-\frac{\epsilon_{\theta}(x,t)}{\sigma_{t}}. This estimate of the log-likelihood gradient can be used for MCMC sampling methods which only require the log-likelihood gradient – such as ULA or U-HMC. These methods can work well, but will never generate exact samples when using non-zero step-sizes. Exact samplers can be derived from approximate samplers like the above methods using Metropolis corrections. Unfortunately, even if the samplers’ transition distribution k(⋅,⋅)k(\cdot,\cdot) does not require log⁡pθ(xt)\log p_{\theta}(x_{t}) evaluation, the Metropolis correction probability:

does. Futhermore, when we only have an estimate of the score at our disposal, we are only able to compose models using products.

Much prior work on EBMs parameterizes fθ(x,t)f_{\theta}(x,t) using a feed-forward neural network, whose final layer has a single output (Nijkamp et al. 2020; Du & Mordatch 2019). Salimans & Ho 2021 compare this approach with the standard ϵθ(x,t)\epsilon_{\theta}(x,t) parameterization and find the ϵθ(x,t)\epsilon_{\theta}(x,t) parameterization to perform better for unconditional image generation. We believe this has to do with the relative sparsity of the gradients of feed-forward neural networks. This can cause difficulties when training to optimize a function of their implicitly computed gradients.

Intriguingly, Salimans & Ho 2021 also explore a more structured energy function definition inspired by denoising autoencoders:

In their study, this parameterization was found to perform near identically to the ϵθ(x,t)\epsilon_{\theta}(x,t) parameterization while admitting an explicit energy function. We believe this energy parameterization performs better because its gradients include the feed-forward network sθ(x,t)s_{\theta}(x,t), making optimization easier. Salimans & Ho 2021 conclude that the ϵθ(x,t)\epsilon_{\theta}(x,t) parameterization should be favored since the computing ∇xfθDAE(x,t)\nabla_{x}f^{DAE}_{\theta}(x,t) requires computing ∇sθ(x,t)\nabla s_{\theta}(x,t) which requires an extra backward pass through the neural network, increasing compute.

We reexamine this energy parameterization and two other choices now that our application motivates having access to an explicit energy function. The other parameterizations are based on different transformations of the sθ(x,t)s_{\theta}(x,t) architecture; the negative L2 norm (L2) and an inner product (IP). They are defined as:

We train models with each parameterization on ImageNet and compare using FID for unconditional sampling. Results can be seen in Table A1. We see that L2 and Inner-Product perform the best, but are both outperformed by the standard parameterization. We initially experimented with these two parameterizations but found that the L2-norm parameterization to be more stable for compositional sampling due, we believe, to the fact that the energy-function is bounded above meaning that MCMC sampling is incapable of running off to infinity to increase likelihood.

Appendix E Synthetic Distribution Compositions

Mixture We provide additional 2D illustrations of mixtures of two diffusion models in Figure A1. We find that HMC sampling enables more accurate mixtures of different synthetic distributions.

Negation We provide additional 2D illustrations of negating two diffusion model with respect to each other in Figure A3. We find that HMC sampling enables accurate negations of different sythetic distributions.

Failure Cases Next we illustrate a failure case of composition using our approach in Figure A3. Our approach fails to generate the product of two distribution when they are disjoint with respect to each other.

Appendix F Experimental Details

We provide detailed experimental details including underlying quantitative metrics, training details, and architectures on 2D synthetic, CLEVR, ImageNet, and test-to-image settings below. To enable stable training of energy-based diffusion models in image settings, we clip gradient norms to be less than 10, and initialize convolutional layers using zero-initialization (Zhang et al. 2019).

Synthetic Datasets For synthetic datasets, we train both score and energy based diffusion models using a small residual MLP model with 4 residual blocks, with a internal hidden dimension of 128 dimensions. We train models for 15000 iterations (10 minutes on a 8 TPUv2 cores) using the Adam optimizer with learning rate of 1e-3, and train diffusion models on 100 discrete timesteps with linear schedule of β\beta values.

When evaluating product of diffusion models, we generate two separate distributions, where train two separate diffusion models. In our first distribution, we construct a GMM of 8 Gaussians in a ring of radius 0.5 around the origin, with each Gaussian having a standard deviation of 0.3. In our second dataset, we construct a uniform distribution of points with xx between -0.1 and 0.1 and yy between -1 and 1. When evaluating mixture diffusion models, we generate one distribution consisting of a mixture of 3 Gaussian with standard deviation 0.03 and centers at (−0.25,0.5),(−0.25,0.0),(−0.25,−0.5)(-0.25,0.5),(-0.25,0.0),(-0.25,-0.5), and another distribution consisting of a mixture of 3 Gaussian with standard deviation 0.03 and centers at (0.25,0.5),(0.25,0.0),(0.25,−0.5)(0.25,0.5),(0.25,0.0),(0.25,-0.5).

To construct MCMC samplers from models on synthetic datasets, we run 3 steps of HMC per timestep, with 3 leapfrog steps per step of HMC. We run 10 steps of MALA sampling per timestep. We found that MCMC performed robustly in the 2D dimensional setting and set the step size of MALA to be 0.002 across all distributions and the step size of HMC to be 0.03 across all distributions (with a mass matrix of 1)

CLEVR For CLEVR, we generated a dataset of 200,000 64×6464\times 64 images with between 1 to 5 different cubes using dataset generation code in (Liu et al. 2021). To evaluate the accuracy in which generated images had cubes at each specified position, we trained a binary classifier on these images, and marked a cube as correctly generated if the confidence of the binary confidence of classifier is greater than 0.5.

To parameterize our diffusion architecture, we follow the architecture of (Ho et al. 2020), where we use a base hidden dimension of 128, and multiply the hidden dimensions by $atdifferentresolutionsoftheimage.Weutilize3residualblocksateachresolutionoftheimage.Wetraineddiffusionmodelswith100discretetimestepswithalinearat different resolutions of the image. We utilize 3 residual blocks at each resolution of the image. We trained diffusion models with 100 discrete timesteps with a linear\beta$ schedule. CLEVR models were trained for 20000 iterations with a batch size of 1024 using the Adam optimizer with step size 1e-4, corresponding to roughly 8 hours on 8 TPUv2 cores.

To initialize MCMC sampling on the CLEVR domain, at each timestep, before applying MCMC sampling, we run one step of the reverse process in the trained diffusion model. We run 40 steps of MCMC sampling per timestep for MALA samplers, and 13 steps of HMC sampling (with 3 leapfrog step per HMC step) (with the mass matrix of HMC samplers set to β\beta). We use HMC with partial momentum refreshment, and use a dampening coefficient of 0.9 across HMC iterations. MALA step sizes are set to 0.035∗βt0.035*\beta_{t} , and HMC step sizes are set to 0.1∗βt0.1*\beta_{t}

ImageNet For ImageNet, we train an unconditional diffusion model 128×128128\times 128 images. We train diffusion models for 1 million iterations of ImageNet with a batch size of 64 (3 days on 16 TPUv2 cores), using Adam optimizer with learning rate 1e-4, for 1 million iterations. We train diffusion models with 1000 discrete timesteps using the cosine beta schedule.

On the ImageNet dataset, we report three seperate metrics. To report classifier accuracy, we feed generated sample into a ImageNet classifier trained on clean images, and label a image as correctly generated if the classifier of a generated image having the specified class is greater than 50%. We further report the Inception Score and FID, which are calculated on 50000 generated samples.

We follow the architecture of (Ho et al. 2020), where we use a base hidden dimension of 128 and multiply the hidden dimensions by $$ at the different resolution of the image. We utilize 2 residual blocks at each resolution of the image.

To initialize MCMC sampling on the ImageNet domain, at each timestep, before applying MCMC sampling, we run one step of the reverse process in the trained diffusion model. We run 6 steps of MCMC sampling per timestep for MALA samplers, and 2 steps of HMC sampling (with 3 leapfrog steps per HMC step and with the mass matrix of HMC samplers set to β\beta). MALA step sizes are set to 0.5∗βt0.5*\beta_{t} , and HMC step sizes are set to 0.6∗βt1.50.6*\beta_{t}^{1.5}

Text-to-Image For text-to-image models, we train models for one week on an internal text/image dataset consisting of 400 million images using 32 TPUv3 cores, with a training data batch size of 256. We train our energy-based text-to-image model using a total of 1000 timesteps with a cosine beta schedule. We follow the architecture of (Ho et al. 2020), where we use a base hidden dimension of 256, and multiply the hidden dimensions by $$ at different resolution of the image. We utilize 3 residual blocks at each resolution of the image.

To upsample images from 64×6464\times 64 resolution to 1024×10241024\times 1024 resolution, we utilize two trained unconditional diffusion models, one trained to upsample from 64×6464\times 64 resolution to 256×256256\times 256 resolution and one trained to upsample from 256×256256\times 256 resolution to 1024×10241024\times 1024 resolution.

To initialize MCMC sampling on the text-to-image domain, at each timestep, before applying MCMC sampling, we run one step of the reverse process in the trained diffusion model. We ran 2 steps of HMC sampling per timestep, with 3 leapfrog step per HMC step and a mass matrix of HMC samplers set to β\beta). We use HMC with partial momentum refreshment, and use a dampening coefficient of 0.9 across HMC iterations. HMC step sizes are set to 0.1∗βt0.1*\beta_{t}.

Image Tapestries To construct image tapestries, we used the Imagen 64x64 base diffusion model. We did not use the Imagen super-res stages. We arranged overlapping image models, with one placed every 32 pixels on a grid horizontally and vertically (so each image model overlapped by 50% with its neighbors to the left, right, top, and bottom). We additionally applied 2x2 average-pooling to the canvas, and applied an image model to the resulting downsampled image, to capture global structure at a lower resolution. Each image model was given its own text prompt, as shown in Figures 1d and 7. All image models used a guidance weight of 30.

The full captions used in in Figure 1(d) are “A detailed oil painting of a fantastical ocean scene, with a mermaid, a ship, a lighthouse, and a whale”, “A detailed closeup of a fantastical oil painting showing a large sailing ship”, “A detailed closeup of a fantastical oil painting showing a mermaid sunning herself”, “A detailed closeup of a fantastical oil painting showing a lighthouse”, “A detailed closeup of a fantastical oil painting showing a curious whale surfacing”, and “A detailed closeup of a oil painting of the ocean”. The full captions used in Figure 8 are “Detailed movie still of an epic space battle. The starship Enterprise fights a giant mech robot over the planet Mars”, “Detailed closeup of a movie still. The starship Enterprise in a space battle, firing its phasers. NCC-1701”, “Detailed closeup of a movie still. A giant mecha robot is fighting in space, holding a glowing sword”, “Detailed closeup of a movie still of an epic space battle, showing glowing phaser beams”, “Detailed closeup of a movie still of an epic space battle. Sun with lens flare”, and “Detailed closeup of a movie still of an epic space battle, showing a portion of the red planet Mars. Debris falling into the atmosphere”.

Imagen models produce an score estimate corresponding to the Gaussian noise vector ϵt\epsilon_{t} between the image xtx_{t} at time tt and the predicted image x0x_{0} at time 00. In order to combine the score function from the overlapping models, we took a weighted sum of their estimated noise vectors, and normalized the resulting vector to have variance 1. The highest resolution models were given a weight of 1. The coarse resolution model was given a weight of [downsampling fraction]=0.25[\text{downsampling fraction}]=0.25. Since it is applied over 4 times as many pixels, this corresponds to an equal weighting of the coarse and fine scale energy functions that tile the canvas. In order to hide seams, weights were linearly tapered to 0 near the image edge, with the tapering performed over the outer 6 pixels.

We use 2 steps of LA sampling after each ancestral sampling step, so as to better approach the equilibrium distribution of the composed energy functions. We performed 256 ancestral sampling steps. We did not perform LA sampling for the final 16 ancestral sampling steps.

Appendix G Text-to-Image Results

We present additional use cases of composing models in different text-to-image domains. First, in Figure A4, we illustrate how composing two separate energy parameterized diffusion models enables us to more accurately generate images that have more detailed information in the caption. Next, in Figure A5, we illustrate how composing two separate energy parameterized diffusion models further enable us to accurately generate images with the correct colors assigned to each object. We further show in Figure A6 how composing the negation of one energy parameterized diffusion model with another other enables us to generate images where one commonly occurring co-founding factor does occur (i.e. a sandy beach without coastal water). Finally, we illustrate in Figure A7, how composing multiple diffusion models enables us to render the number of objects in a scene accurately.