Multistep Consistency Models

Jonathan Heek, Emiel Hoogeboom, Tim Salimans

Introduction

Diffusion models have rapidly become one of the dominant generative models for image, video and audio generation (Ho et al., 2020; Kong et al., 2021; Saharia et al., 2022). The biggest downside to diffusion models is their relatively expensive sampling procedure: whereas training uses a single function evaluation per datapoint, it requires many (sometimes hundreds) of evaluations to generate a sample.

Recently, Consistency Models (Song et al., 2023) have reduced sampling time significantly, but at the expense of image quality. Consistency models come in two variants: Consistency Training (CT) and Consistency Distillation (CD) and both have considerably improved performance compared to earlier works. TRACT (Berthelot et al., 2023) focuses solely on distillation with an approach similar to consistency distillation, and shows that dividing the diffusion trajectory in stages can improve performance. Despite their successes, neither of these works attain performance close to a standard diffusion baseline.

Here, we propose a unification of Consistency Models and TRACT, that closes the performance gap between standard diffusion performance and low-step variants. We relax the single-step constraint from consistency models to allow ourselves as much as 4, 8 or 16 function evaluations for certain settings. Further, we generalize TRACT to consistency training and adapt step schedule annealing and synchronized dropout from consistency modelling. We also show that as steps increase, Multistep CT becomes a diffusion model. We introduce a unifying training algorithm to train what we call Multistep Consistency Models, which splits the diffusion process from data to noise into predefined segments. For each segment a separate consistency model is trained, while sharing the same parameters. For both CT and CD, this turns out to be easier to model and leads to significantly improved performance with fewer steps. Surprisingly, we can perfectly match baseline diffusion model performance with only eight steps, on both Imagenet64 and Imagenet128.

Another important contribution of this paper that makes the previous result possible, is a deterministic sampler for diffusion models that can obtain competitive performance on more complicated datasets such as ImageNet128 in terms of FID score. We name this sampler Adjusted DDIM (aDDIM), which essentially inflates the noise prediction to correct for the integration error that produces blurrier samples.

In terms of numbers, we achieve performance rivalling standard diffusion approaches with as little as 8 and sometimes 4 sampling steps. These impressive results are both for consistency training and distillation. A remarkable result is that with only 4 sampling steps, multistep consistency models obtain performances of 1.6 FID on ImageNet64 and 2.3 FID on Imagenet128.

Background: Diffusion Models

Diffusion models are specified by a destruction process that adds noise to destroy data: zt=αtx+σtϵt{\bm{z}}_{t}=\alpha_{t}{\bm{x}}+\sigma_{t}{\bm{{\epsilon}}}_{t} where ϵt∼N(0,1){\bm{{\epsilon}}}_{t}\sim\mathcal{N}(0,1). Typically for t→1t\to 1, zt{\bm{z}}_{t} is approximately distributed as a standard normal and for t→0t\to 0 it is approximately x{\bm{x}}. In terms of distributions one can write the diffusion process as:

To sample from these models, one uses the denoising equation:

In contrast, consistency models (Song et al., 2023; Song & Dhariwal, 2023) aim to learn a direct mapping from noise to data. Consistency models are constrained to predict x=f(z0,0){\bm{x}}=f({\bm{z}}_{0},0), and are further trained by learning to be consistent, minimizing:

where zs=αsx+σsϵ{\bm{z}}_{s}=\alpha_{s}{\bm{x}}+\sigma_{s}{\bm{{\epsilon}}} and zt=αtx+σtϵ{\bm{z}}_{t}=\alpha_{t}{\bm{x}}+\sigma_{t}{\bm{{\epsilon}}}, (note both use the same ϵ{\bm{{\epsilon}}}) and ss is closer to the data meaning s<ts<t. When (or if) a consistency model succeeds, the trained model solves for the probability ODE path along time. When successful, the resulting model predicts the same x{\bm{x}} along the entire trajectory. At initialization it will be easiest for the model to learn ff near zero, because ff is defined as an identity function at t=0t=0. Throughout training, the model will propagate the end-point of the trajectory further and further to t=1t=1. In our own experience, training consistency models is much more difficult than diffusion models.

Consistency Training and Distillation

Consistency Models come in two flavours: Consistency Training (CT) and Consistency Distillation (CD). In the paragraph before, zs{\bm{z}}_{s} was given by the data which would be the case for CT. Alternatively, one might use a pretrained diffusion model to take a probability flow ODE step (for instance with DDIM). Calling this pretrained model the teacher, the objective for CD can be described by:

where DDIM now defines zs{\bm{z}}_{s} given the current zt{\bm{z}}_{t} and (possibly an estimate of) x{\bm{x}}.

An important hyperparameter in consistency models is the gap between the model evaluations at tt and ss. For CT large gaps incurs a bias, but the solutions are propagated through diffusion time more quickly. On the other hand, when s→ts\to t the bias tends to zero but it takes much longer to propagate information through diffusion time. In practice a step schedule N(⋅)N(\cdot) is used to anneal the step size t−s=1/N(⋅)t-s=1/N(\cdot) over the course of training.

DDIM Sampler

The DDIM sampler is a linearization of the probability flow ODE that is often used in diffusion models. In a variance preserving setting, it is given by:

In addition to being a sampling method, the DDIM⁡\operatorname{DDIM} equation will also prove to be a useful tool to construct an algorithm for our multistep diffusion models.

Another helpful equations is the inverse of DDIM (Salimans & Ho, 2022), originally proposed to find a natural way parameterize a student diffusion model when a teacher defines the sampling procedure in terms of zt{\bm{z}}_{t} to zs{\bm{z}}_{s}. The equation takes in zt{\bm{z}}_{t} and zs{\bm{z}}_{s}, and produces x{\bm{x}} for which DDIM⁡t→s(x,zt)=zs\operatorname{DDIM}_{t\to s}({\bm{x}},{\bm{z}}_{t})={\bm{z}}_{s}. It can be derived by rearranging terms from the DDIM⁡\operatorname{DDIM} equation:

Multistep Consistency Models

In this section we describe multi-step consistency models. First we explain the main algorithm, for both consistency training and distillation. Furthermore, we show that multi-step consistency converges to a standard diffusion training in the limit. Finally, we develop a deterministic sampler named aDDIM that corrects for the missing variance problem in DDIM.

For now it suffices to think aDDIM⁡\operatorname{aDDIM} as DDIM⁡\operatorname{DDIM}. It will be described in detail in section 3.2. In fact, one can drop-in any deterministic sampler in place of aDDIM⁡\operatorname{aDDIM} in the case of distillation.

A model can be trained on directly on this loss in zz space, however make the loss more interpretable and relate it more closely to standard diffusion, we re-parametrize the loss to xx-space using:

Finetuning Multistep CMs from a pretrained diffusion checkpoint will lead to quicker and more stable convergence.

As the number of steps increases, Multistep CMs will rival diffusion model performance, giving a direct trade-off between sample quality and duration.

What about training in continuous time?

2 The Adjusted DDIM (aDDIM) sampler.

Popular methods for distilling diffusion models, including the method we propose here, rely on deterministic sampling through numerical integration of the probability flow ODE. In practice, numerical integration of this ODE in a finite number of steps incurs error. For the DDIM integrator (Song et al., 2021a) used for distilling diffusion models in both consistency distillation (Song et al., 2023) and progressive distillation (Salimans & Ho, 2022; Meng et al., 2022) this integration error causes samples to become blurry. To see this quantitatively, consider a hypothetical perfect sampler that first samples x∗∼p(x∣zt){\bm{x}}^{*}\sim p({\bm{x}}|{\bm{z}}_{t}), and then samples zs{\bm{z}}_{s} using

If the initial zt{\bm{z}}_{t} is from the correct distribution p(zt)p({\bm{z}}_{t}), the sampled zs∗{\bm{z}}^{*}_{s} would then also be exactly correct. Instead, the DDIM integrator uses

Currently, the best sample quality is achieved with stochastic samplers, which can be tuned to add exactly enough noise to undo the oversmoothing caused by numerical integration. However, current distillation methods are not well suited to distilling these stochastic samplers directly. Here we therefore propose a new deterministic sampler that aims to achieve the norm increasing effect of noise addition in a deterministic way. It turns out we can do this by making a simple adjustment to the DDIM sampler, and we therefore call our new method Adjusted DDIM (aDDIM). Our modification is heuristic and is not more theoretically justified than the original DDIM sampler. However, empirically we find aDDIM to work very well leading to improved FID scores.

Instead of adding noise to our sampled zs{\bm{z}}_{s}, we simply increase the contribution of our deterministic estimate of the noise ϵ^=(zt−αtx^)/σt\hat{{\bm{{\epsilon}}}}=({\bm{z}}_{t}-\alpha_{t}\hat{{\bm{x}}})/\sigma_{t}. Assuming that x^\hat{{\bm{x}}} and ϵ^\hat{{\bm{{\epsilon}}}} are orthogonal, we achieve the correct norm for our sampling iterates using:

Related Work

Existing works closest to ours are Consistency Models (Song et al., 2023; Song & Dhariwal, 2023) and TRACT (Berthelot et al., 2023). Compared to consistency models, we propose to operate on multiple stages, which simplifies the modelling task and improves performance significantly. On the other hand, TRACT limits itself to distillation and uses the self-evaluation from consistency models to distill models over multiple stages. The stages are progressively reduced to either one or two stages and thus steps. The end-goal of TRACT is again to sample in either one or two steps, whereas we believe better results can be obtained by optimizing for a slightly larger number of steps. We show that this more conservative target, in combination with our improved sampler and annealed schedule, leads to significant improvements in terms of image quality that closes the gap between sample quality of standard diffusion and low-step diffusion-inspired approaches.

Earlier, DDIM (Song et al., 2021a) showed that deterministic samplers degrade more gracefully than the stochastic sampler used by Ho et al. (2020) when limiting the number of sampling steps. Karras et al. (2022) proposed a second order Heun sampler to reduce the number of steps (and function evaluations), while Jolicoeur-Martineau et al. (2021) studied different SDE integrators to reduce function evaluations. Zheng et al. (2023) use specialized architectures to distill the ODE trajectory from a pre-created noise-sample pair dataset. Progressive Distillation (Salimans & Ho, 2022; Meng et al., 2022) distills diffusion models in stages, which limits the number of model evaluations during training while exponentially reducing the required number of sampling steps with the number stages. Luo et al. (2023) distill the knowledge from the diffusion model into a single-step model.

Other methods inspired by diffusion such as Rectified Flows (Liu et al., 2023) and Flow Matching (Lipman et al., 2023) have also tried to reduce sampling times. In practice however, flow matching and rectified flows are generally used to map to a standard normal distribution and reduce to standard diffusion. As a consequence, on its own they still require many evaluation steps. In Rectified Flows, a distillation approach is proposed that does reduce sampling steps more significantly, but this comes at the expense of sample quality.

Experiments

Our experiments focus on a quantitative comparison using the FID score on ImageNet as well as a qualitative assessment on large scale Text-to-Image models. These experiments should make our approach comparable to existing academic work while also giving insight in how multi-step distillation works at scale.

For our ImageNet experiments we trained diffusion models on ImageNet64 and ImageNet128 in a base and large variant. We initialize the consistency models from the pre-trained diffusion model weights which we found to greatly increase robustness and convergence. Both consistency training and distillation are used. Classifier Free Guidance (Ho & Salimans, 2022) was used only on the base ImageNet128 experiments. For all other experiments we did not use guidance because it did not significantly improve the FID scores of the diffusion model. All consistency models are trained for 200,000200,000 steps with a batch size of 20482048 and an step schedule that anneals from 6464 to 12801280 in 100.000100.000 steps with an exponential schedule.

In Table 1 it can be seen that as the multistep count increases from a single consistency step, the performance considerably improves. For instance, on the ImageNet64 Base model it improves from 7.2 for one-step, to 2.7 and further to 1.8 for two and four-steps respectively. In there are generally two patterns we observe: As the steps increase, performance improves. This validates our hypothesis that more steps give a helpful trade-off between sample quality and speed. It is very pleasant that this happens very early: even on a complicated dataset such as ImageNet128, our base model variant is able achieve 2.1 FID in 8 steps, when consistency distilling.

To draw a direct comparison between Progressive Distillation (PD) (Salimans & Ho, 2022) and our approaches, we reimplement PD using aDDIM and we use same base architecture, as reported in Table 3. With our improvements, PD can attain better performance than previously reported in literature. However, compared to MultiStep CT and CD it starts to degrade in sample quality at low step counts. For instance, a 4-step PD model attains an FID of 2.4 whereas CD achieves 1.7.

Further we are ablating whether annealing the step schedule is important to attain good performance. As can be seen in Table 2, it is especially important for low multistep models to anneal the schedule. In these experiments, annealing always achieves better performance than tests with constant steps at 128,256,1024128,256,1024. As more multisteps are taken, the importance of the annealing schedule is less important.

Compared to existing works in literature, we achieve SOTA FID scores in both ImageNet64 on 2-step, 4-step and 8-step generation. Interestingly, we achieve approximately the same performance using single step CD compared to iCT-deep (Song & Dhariwal, 2023), which achieves this result using direct consistency training. Since direct training has been empirically shown to be a more difficult task, one could conclude that some of our hyperparameter choices may still be suboptimal in the extreme low-step regime. Conversely, this may also mean that multistep consistency is less sensitive to hyperparameter choices.

In addition, we compare on ImageNet128 to our reimplementation of Progressive Distillation. Unfortunately, ImageNet128 has not been widely adopted as a few-step benchmark, possibly because a working deterministic sampler has been missing until this point. For reference we also provide the recent result from (Kingma & Gao, 2023). Further, with these results we hope to put ImageNet128 on the map for few-step diffusion model evaluation.

2 Qualitative Evaluation on Text to Image modelling

In addition to the quantitative analysis on ImageNet, we study the effects on a text-to-image model by directly comparing samples. We first train a 20B parameter diffusion model on text-to-image pairs with a T5 XXL paper following (Saharia et al., 2022) for 1.3 million steps. Then, we distill a 16-step consistency model using the DDIM sampler. In Figure 2 and 3 we compare samples from our 16-step CD aDDIM distilled model to the original 100-step DDIM sampler. Because the random seed is shared we can easily compare the samples between these models, and we can see that there are generally minor differences. In our own experience, we often find certain details more precise, at a slight cost of overall construction. Another comparison in Figure 4 shows the difference between a DDIM distilled model (equivalent to η=0\eta=0 in aDDIM) and the standard DDIM sampler. Again we see many similarities when sharing the same initial random seed.

Conclusions

In conclusion, this paper presents Multistep Consistency Models, a simple unification between Consistency Models (Song et al., 2023) and TRACT (Berthelot et al., 2023) that closes the performance gap between standard diffusion and few-step sampling. Multistep Consistency gives a direct trade-off between sample quality and speed, achieving performance comparable to standard diffusion in as little as eight steps.

References