Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Simian Luo, Yiqin Tan, Longbo Huang, Jian Li, Hang Zhao

Introduction

Diffusion models have emerged as powerful generative models that have gained significant attention and achieved remarkable results in various domains (Ho et al., 2020; Song et al., 2020a; Nichol & Dhariwal, 2021; Ramesh et al., 2022; Song & Ermon, 2019; Song et al., 2021). In particular, latent diffusion models (LDMs) (e.g., Stable Diffusion (Rombach et al., 2022)) have demonstrated exceptional performance, especially in high-resolution text-to-image synthesis tasks. LDMs can generate high-quality images conditioned on textual descriptions by utilizing an iterative reverse sampling process that performs gradual denoising of samples. However, diffusion models suffer from a notable drawback: the iterative reverse sampling process leads to slow generation speed, limiting their real-time applicability. To overcome this drawback, researchers have proposed several methods to improve the sampling speed, which involves accelerating the denoising process by enhancing ODE solvers (Ho et al., 2020; Lu et al., 2022a; b), which can generate images within 10∼\sim20 sampling steps. Another approach is to distill a pre-trained diffusion model into models that enable few-step inference Salimans & Ho (2022); Meng et al. (2023). In particular, Meng et al. (2023) proposed a two-stage distillation approach to improving the sampling efficiency of classifier-free guided models. Recently, Song et al. (2023) proposed consistency models as a promising alternative aimed at speeding up the generation process. By learning consistency mappings that maintain point consistency on ODE-trajectory, these models allow for single-step generation, eliminating the need for computation-intensive iterations. However, Song et al. (2023) is constrained to pixel space image generation tasks, making it unsuitable for synthesizing high-resolution images. Moreover, the applications to the conditional diffusion model and the incorporation of classifier-free guidance have not been explored, rendering their methods unsuitable for text-to-image generation synthesis.

In this paper, we introduce Latent Consistency Models (LCMs) for fast, high-resolution image generation. Mirroring LDMs, we employ consistency models in the image latent space of a pre-trained auto-encoder from Stable Diffusion (Rombach et al., 2022). We propose a one-stage guided distillation method to efficiently convert a pre-trained guided diffusion model into a latent consistency model by solving an augmented PF-ODE. Additionally, we propose Latent Consistency Fine-tuning, which allows fine-tuning a pre-trained LCM to support few-step inference on customized image datasets. Our main contributions are summarized as follows:

We propose Latent Consistency Models (LCMs) for fast, high-resolution image generation. LCMs employ consistency models in the image latent space, enabling fast few-step or even one-step high-fidelity sampling on pre-trained latent diffusion models (e.g., Stable Diffusion (SD)).

We provide a simple and efficient one-stage guided consistency distillation method to distill SD for few-step (2∼\sim4) or even 1-step sampling. We propose the Skipping-Step technique to further accelerate the convergence. For 2- and 4-step inference, our method costs only 32 A100 GPU hours for training and achieves state-of-the-art performance on the LAION-5B-Aesthetics dataset.

We introduce a new fine-tuning method for LCMs, named Latent Consistency Fine-tuning, enabling efficient adaptation of a pre-trained LCM to customized datasets while preserving the ability of fast inference.

Related Work

Diffusion Models have achieved great success in image generation (Ho et al., 2020; Song et al., 2020a; Nichol & Dhariwal, 2021; Ramesh et al., 2022; Rombach et al., 2022; Song & Ermon, 2019). They are trained to denoise the noise-corrupted data to estimate the score of data distribution. During inference, samples are drawn by running the reverse diffusion process to gradually denoise the data point. Compared to VAEs (Kingma & Welling, 2013; Sohn et al., 2015) and GANs (Goodfellow et al., 2020), diffusion models enjoy the benefit of training stability and better likelihood estimation.

Accelerating DMs. However, diffusion models are bottlenecked by their slow generation speed. Various approaches have been proposed, including training-free methods such as ODE solvers (Song et al., 2020a; Lu et al., 2022a; b), adaptive step size solvers (Jolicoeur-Martineau et al., 2021), predictor-corrector methods (Song et al., 2020b). Training-based approaches include optimized discretization (Watson et al., 2021), truncated diffusion (Lyu et al., 2022; Zheng et al., 2022), neural operator (Zheng et al., 2023) and distillation (Salimans & Ho, 2022; Meng et al., 2023). More recently, new generative models for faster sampling have also been proposed (Liu et al., 2022; 2023).

Latent Diffusion Models (LDMs) (Rombach et al., 2022) excel in synthesizing high-resolution text-to-images. For example, Stable Diffusion (SD) performs forward and reverse diffusion processes in the data latent space, resulting in more efficient computation.

Consistency Models (CMs) (Song et al., 2023) have shown great potential as a new type of generative model for faster sampling while preserving generation quality. CMs adopt consistency mapping to directly map any point in ODE trajectory to its origin, enabling fast one-step generation. CMs can be trained by distilling pre-trained diffusion models or as standalone generative models. Details of CMs are elaborated in the following section.

Preliminaries

In this section, we briefly review diffusion and consistency models and define relevant notations.

By considering the reverse time SDE (see Appendix A for more details), one can show that the marginal distribution qt(x)q_{t}(\bm{x}) satisfies the following ordinary differential equation, called the Probability Flow ODE (PF-ODE) (Song et al., 2020b; Lu et al., 2022a):

In diffusion models, we train the noise prediction model ϵθ(xt,t)\bm{\epsilon}_{\theta}(\bm{x}_{t},t) to fit −∇log⁡qt(xt)-\nabla\log q_{t}(\bm{x}_{t}) (called the score function). Approximating the score function by the noise prediction model in 21, one can obtain the following empirical PF-ODE for sampling:

Consistency Models: The Consistency Model (CM) (Song et al., 2023) is a new family of generative models that enables one-step or few-step generation. The core idea of the CM is to learn the function that maps any points on a trajectory of the PF-ODE to that trajectory's origin (i.e., the solution of the PF-ODE). More formally, the consistency function is defined as f:(xt,t)⟼xϵ,{\bm{f}}:({\bm{x}}_{t},t)\longmapsto{\bm{x}}_{\epsilon}, where ϵ\epsilon is a fixed small positive number. One important observation is that the consistency function should satisfy the self-consistency property:

The key idea in (Song et al., 2023) for learning a consistency model fθ{\bm{f}}_{\bm{\theta}} is to learn a consistency function from data by effectively enforcing the self-consistency property in Eq. 4. To ensure that fθ(x,ϵ)=x{\bm{f}}_{\bm{\theta}}({\bm{x}},\epsilon)={\bm{x}}, the consistency model fθ{\bm{f}}_{\bm{\theta}} is parameterized as:

where cskip(t)c_{\text{skip}}(t) and cout(t)c_{\text{out}}(t) are differentiable functions with cskip(ϵ)=1c_{\text{skip}}(\epsilon)=1 and cout(ϵ)=0c_{\text{out}}(\epsilon)=0, and Fθ(x,t)\bm{F}_{\bm{\theta}}({\bm{x}},t) is a deep neural network. A CM can be either distilled from a pre-trained diffusion model or trained from scratch. The former is known as Consistency Distillation. To enforce the self-consistency property, we maintain a target model θ−{\bm{\theta}}^{-}, updated with exponential moving average (EMA) of the parameter θ{\bm{\theta}} we intend to learn, i.e., θ−←μθ−+(1−μ)θ{\bm{\theta}}^{-}\leftarrow\mu{\bm{\theta}}^{-}+(1-\mu){\bm{\theta}}, and define the consistency loss as follows:

where Φ\Phi denotes the one-step ODE solver applied to PF-ODE in Eq. 24. (Song et al., 2023) used Euler (Song et al., 2020b) or Heun solver (Karras et al., 2022) as the numerical ODE solver. More details and the pseudo-code for consistency distillation (Algorithm 2) are provided in Appendix A.

Latent Consistency Models

Consistency Models (CMs) (Song et al., 2023) only focused on image generation tasks on ImageNet 64×\times64 (Deng et al., 2009) and LSUN 256×\times256 (Yu et al., 2015). The potential of CMs to generate higher-resolution text-to-image tasks remains unexplored. In this paper, we introduce Latent Consistency Models (LCMs) in Sec 4.1 to tackle these more challenging tasks, unleashing the potential of CMs. Similar to LDMs, our LCMs adopt a consistency model in the image latent space. We choose the powerful Stable Diffusion (SD) as the underlying diffusion model to distill from. We aim to achieve few-step (2∼\sim4) and even one-step inference on SD without compromising image quality. The classifier-free guidance (CFG) (Ho & Salimans, 2022) is an effective technique to further improve sample quality and is widely used in SD. However, its application in CMs remains unexplored. We propose a simple one-stage guided distillation method in Sec 4.2 that solves an augmented PF-ODE, integrating CFG into LCM effectively. We propose Skipping-Step technique to accelerate the convergence of LCMs in Sec. 4.3. Finally, we propose Latent Consistency Fine-tuning to finetune a pre-trained LCM for few-step inference on a customized dataset in Sec 4.4.

Utilizing image latent space in large-scale diffusion models like Stable Diffusion (SD) (Rombach et al., 2022) has effectively enhanced image generation quality and reduced computational load. In SD, an autoencoder (E,D\mathcal{E},\mathcal{D}) is first trained to compress high-dim image data into low-dim latent vector z=E(x)z=\mathcal{E}(x), which is then decoded to reconstruct the image as x^=D(z)\hat{x}=\mathcal{D}(z). Training diffusion models in the latent space greatly reduces the computation costs compared to pixel-based models and speeds up the inference process; LDMs make it possible to generate high-resolution images on laptop GPUs. For LCMs, we leverage the advantage of the latent space for consistency distillation, contrasting with the pixel space used in CMs (Song et al., 2023). This approach, termed Latent Consistency Distillation (LCD) is applied to pre-trained SD, allowing the synthesis of high-resolution (e.g., 768×\times768) images in 1∼\sim4 steps. We focus on conditional generation. Recall that the PF-ODE of the reverse diffusion process (Song et al., 2020b; Lu et al., 2022a) is

where zt\bm{{\bm{z}}}_{t} are image latents, ϵθ(zt,c,t)\bm{\epsilon}_{\theta}\left(\bm{{\bm{z}}}_{t},\bm{c},t\right) is the noise prediction model, and c\bm{c} is the given condition (e.g text). Samples can be drawn by solving the PF-ODE from TT to . To perform LCD, we introduce the consistency function fθ:(zt,c,t)↦z0\bm{f_{\theta}}:(\bm{z_{t}},\bm{c},t)\mapsto\bm{z_{0}} to directly predict the solution of PF-ODE (Eq. 8) for t=0t=0. We parameterize fθ\bm{f_{\theta}} by the noise prediction model ϵ^θ\hat{\bm{\epsilon}}_{\theta}, as follows:

where cskip(0)=1,cout(0)=0c_{\text{skip}}(0)=1,c_{\text{out}}(0)=0 and ϵ^θ(z,c,t)\hat{\bm{\epsilon}}_{\theta}({\bm{z}},{\bm{c}},t) is a noise prediction model that initializes with the same parameters as the teacher diffusion model. Notably, fθ\bm{f_{\theta}} can be parameterized in various ways, depending on the teacher diffusion model parameterizations of predictions (e.g., x\bm{x}, ϵ\bm{\epsilon} (Ho et al., 2020), v\bm{v} (Salimans & Ho, 2022)). We discuss other possible parameterizations in Appendix D.

We assume that an efficient ODE solver Ψ(zt,t,s,c)\Psi({\bm{z}}_{t},t,s,{\bm{c}}) is available for approximating the integration of the right-hand side of Eq equation 8 from time tt to ss. In practice, we can use DDIM (Song et al., 2020a), DPM-Solver (Lu et al., 2022a) or DPM-Solver++ (Lu et al., 2022b) as Ψ(⋅,⋅,⋅,⋅)\Psi(\cdot,\cdot,\cdot,\cdot). Note that we only use these solvers in training/distillation, not in inference. We will discuss these solvers further when we introduce the skipping-step technique in Sec. 4.3. LCM aims to predict the solution of the PF-ODE by minimizing the consistency distillation loss (Song et al., 2023):

Here, z^tnΨ\bm{\hat{{\bm{z}}}}^{\Psi}_{t_{n}} is an estimation of the evolution of the PF-ODE from tn+1→tnt_{n+1}\rightarrow t_{n} using ODE solver Ψ\Psi:

where the solver Ψ(⋅,⋅,⋅,⋅)\Psi(\cdot,\cdot,\cdot,\cdot) is used to approximate the integration from tn+1→tnt_{n+1}\rightarrow t_{n}.

2 One-Stage Guided Distillation by solving augmented PF-ODE

Classifier-free guidance (CFG) (Ho & Salimans, 2022) is crucial for synthesizing high-quality text-aligned images in SD, typically needing a CFG scale ω\omega over 66. Thus, integrating CFG into a distillation method becomes indispensable. Previous method Guided-Distill (Meng et al., 2023) introduces a two-stage distillation to support few-step sampling from a guided diffusion model. However, it is computationally intensive (e.g. at least 45 A100 GPUs Days for 2-step inference, estimated in (Liu et al., 2023)). An LCM demands merely 32 A100 GPUs Hours training for 2-step inference, as depicted in Figure 1. Furthermore, the two-stage guided distillation might result in accumulated error, leading to suboptimal performance. In contrast, LCMs adopt efficient one-stage guided distillation by solving an augmented PF-ODE. Recall the CFG used in reverse diffusion process:

where the original noise prediction is replaced by the linear combination of conditional and unconditional noise and ω\omega is called the guidance scale. To sample from the guided reverse process, we need to solve the following augmented PF-ODE: (i.e., augmented with the terms related to ω\omega)

To efficiently perform one-stage guided distillation, we introduce an augmented consistency function fθ:(zt,ω,c,t)↦z0\bm{f_{\theta}}:(\bm{z_{t}},\omega,\bm{c},t)\mapsto\bm{z_{0}} to directly predict the solution of augmented PF-ODE (Eq. 13) for t=0t=0. We parameterize the fθ\bm{f_{\theta}} in the same way as in Eq. 9, except that ϵ^θ(z,c,t)\hat{\bm{\epsilon}}_{\theta}({\bm{z}},{\bm{c}},t) is replaced by ϵ^θ(z,ω,c,t)\hat{\bm{\epsilon}}_{\theta}({\bm{z}},\omega,{\bm{c}},t), which is a noise prediction model initializing with the same parameters as the teacher diffusion model, but also contains additional trainable parameters for conditioning on ω\omega. The consistency loss is the same as Eq. 10 except that we use augmented consistency function fθ(zt,ω,c,t)\bm{f_{\theta}}(\bm{z_{t}},\omega,\bm{c},t).

Again, we can use DDIM (Song et al., 2020a), DPM-Solver (Lu et al., 2022a) or DPM-Solver++ (Lu et al., 2022b) as the PF-ODE solver Ψ(⋅,⋅,⋅,⋅)\Psi(\cdot,\cdot,\cdot,\cdot).

3 Accelerating Distillation with Skipping Time Steps

Discrete diffusion models (Ho et al., 2020; Song & Ermon, 2019) typically train noise prediction models with a long time-step schedule {ti}i\{t_{i}\}_{i} (also called discretization schedule or time schedule) to achieve high quality generation results. For instance, Stable Diffusion (SD) has a time schedule of length 1,000. However, directly applying Latent Consistency Distillation (LCD) to SD with such an extended schedule can be problematic. The model needs to sample across all 1,000 time steps, and the consistency loss attempts to aligns the prediction of LCM model fθ(ztn+1,c,tn+1)\bm{f_{\theta}}(\bm{{\bm{z}}}_{t_{n+1}},\bm{c},t_{n+1}) with the prediction fθ(ztn,c,tn)\bm{f_{\theta}}(\bm{{\bm{z}}}_{t_{n}},\bm{c},t_{n}) at the subsequent step along the same trajectory. Since tn−tn+1t_{n}-t_{n+1} is tiny, ztn\bm{{\bm{z}}}_{t_{n}} and ztn+1\bm{{\bm{z}}}_{t_{n+1}} (and thus fθ(ztn+1,c,tn+1)\bm{f_{\theta}}(\bm{{\bm{z}}}_{t_{n+1}},\bm{c},t_{n+1}) and fθ(ztn,c,tn)\bm{f_{\theta}}(\bm{{\bm{z}}}_{t_{n}},\bm{c},t_{n})) are already close to each other, incurring small consistency loss and hence leading to slow convergence. To address this issues, we introduce the skipping-step method to considerably shorten the length of time schedule (from thousands to dozens) to achieve fast convergence while preserving generation quality.

Consistency Models (CMs) (Song et al., 2023) use the EDM (Karras et al., 2022) continuous time schedule, and the Euler, or Heun Solver as the numerical continuous PF-ODE solver. For LCMs, in order to adapt to the discrete-time schedule in Stable Diffusion, we utilize DDIM (Song et al., 2020a), DPM-Solver (Lu et al., 2022a), or DPM-Solver++ (Lu et al., 2022b) as the ODE solver. (Lu et al., 2022a) shows that these advanced solvers can solve the PF-ODE efficiently in Eq. 8. Now, we introduce the Skipping-Step method in Latent Consistency Distillation (LCD). Instead of ensuring consistency between adjacent time steps tn+1→tnt_{n+1}\rightarrow t_{n}, LCMs aim to ensure consistency between the current time step and kk-step away, tn+k→tnt_{n+k}\rightarrow t_{n}. Note that setting kk=1 reduces to the original schedule in (Song et al., 2023), leading to slow convergence, and very large kk may incur large approximation errors of the ODE solvers. In our main experiments, we set kk=20, drastically reducing the length of time schedule from thousands to dozens. Results in Sec. 5.2 show the effect of various kk values and reveal that the skipping-step method is crucial in accelerating the LCD process. Specifically, consistency distillation loss in Eq. 14 is modified to ensure consistency from tn+kt_{n+k} to tnt_{n}:

with z^tnΨ,ω\bm{\hat{{\bm{z}}}}^{\Psi,\omega}_{t_{n}} being an estimate of ztn{\bm{z}}_{t_{n}} using numerical augmented PF-ODE solver Ψ\Psi:

The above derivation is similar to Eq. 15. For LCM, we use three possible ODE solvers here: DDIM (Song et al., 2020a), DPM-Solver (Lu et al., 2022a), DPM-Solver++ (Lu et al., 2022b), and we compare their performance in Sec 5.2. In fact, DDIM (Song et al., 2020a) is the first-order discretization approximation of the DPM-Solver (Proven in (Lu et al., 2022a)). Here we provide the detailed formula of the DDIM PF-ODE solver ΨDDIM\Psi_{\text{DDIM}} from tn+kt_{n+k} to tnt_{n}. The formulas of the other two solver ΨDPM-Solver\Psi_{\text{DPM-Solver}}, ΨDPM-Solver++\Psi_{\text{DPM-Solver++}} are provided in Appendix E.

We present the pseudo-code for LCD with CFG and skipping-step techniques in Algorithm 1 The modifications from the original Consistency Distillation (CD) algorithm in Song et al. (2023) are highlighted in blue. Also, the LCM sampling algorithm 3 is provided in Appendix B.

4 Latent Consistency Fine-tuning for customized dataset

Foundation generative models like Stable Diffusion excel in diverse text-to-image generation tasks but often require fine-tuning on customized datasets to meet the requirements of downstream tasks. We propose Latent Consistency Fine-tuning (LCF), a fine-tuning method for pretrained LCM. Inspired by Consistency Training (CT) (Song et al., 2023), LCF enables efficient few-step inference on customized datasets without relying on a teacher diffusion model trained on such data. This approach presents a viable alternative to traditional fine-tuning methods for diffusion models. The pseudo-code for LCF is provided in Algorithm 4, with a more detailed illustration in Appendix C.

Experiment

In this section, we employ latency consistency distillation to train LCM on two subsets of LAION-5B. In Sec 5.1, we first evaluate the performance of LCM on text-to-image generation tasks. In Sec 5.2, we provide a detailed ablation study to test the effectiveness of using different solvers, skipping step schedules and guidance scales. Lastly, in Sec 5.3, we present the experimental results of latent consistency finetuning on a pretrained LCM on customized image datasets.

Datasets We use two subsets of LAION-5B (Schuhmann et al., 2022): LAION-Aesthetics-6+ (12M) and LAION-Aesthetics-6.5+ (650K) for text-to-image generation. Our experiments consider resolutions of 512×\times512 and 768×\times768. For 512 resolution, we use LAION-Aesthetics-6+, which comprises 12M text-image pairs with predicted aesthetics scores higher than 6. For 768 resolution, we use LAION-Aesthetics-6.5+, with 650K text-image pairs with aesthetics score higher than 6.5.

Model Configuration For 512 resolution, we use the pre-trained Stable Diffusion-V2.1-Base (Rombach et al., 2022) as the teacher model, which was originally trained on resolution 512×\times512 with ϵ\bm{\epsilon}-Prediction (Ho et al., 2020). For 768 resolution, we use the widely used pre-trained Stable Diffusion-V2.1, originally trained on resolution 768×\times768 with v\bm{v}-Prediction (Salimans & Ho, 2022). We train LCM with 100K iterations and we use a batch size of 72 for (512×512)(512\times 512) setting, and 16 for (768×768)(768\times 768) setting, the same learning rate 8e-6 and EMA rate μ=0.999943\mu=0.999943 as used in (Song et al., 2023). For augmented PF-ODE solver Ψ\Psi and skipping step kk in Eq. 17, we use DDIM-Solver (Song et al., 2020a) with skipping step k=20k=20. We set the guidance scale range [wmin,wmax]=[w_{\text{min}},w_{\text{max}}]=, consistent with (Meng et al., 2023). More training details are provided in the Appendix F.

Baselines & Evaluation We use DDIM (Song et al., 2020a), DPM (Lu et al., 2022a), DPM++ (Lu et al., 2022b) and Guided-Distill (Meng et al., 2023) as baselines. The first three are training-free samplers requiring more peak memory per step with classifier-free guidance. Guided-Distill requires two stages of guided distillation. Since Guided-Distill is not open-sourced, we strictly followed the training procedure outlined in the paper to reproduce the results. Due to the limited resource (Meng et al. (2023) used a large batch size of 512, requiring at least 32 A100 GPUs), we reduce the batch size to 7272, the same as ours, and trained for the same 100K iterations. Reproduction details are provided in Appendix G. We admit that longer training and more computational resources can lead to better results as reported in (Meng et al., 2023). However, LCM achieves faster convergence and superior results under the same computation cost. For evaluation, We generate 30K images from 10K text prompts in the test set (3 images per prompt), and adopt FID and CLIP scores to evaluate the diversity and quality of the generated images. We use ViT-g/14 for evaluating CLIP scores.

Results. The quantitative results in Tables 1 and 2 show that LCM notably outperforms baseline methods at 512512 and 768768 resolutions, especially in the low step regime (1∼\sim4), highlighting its efficency and superior performance. Unlike DDIM, DPM, DPM++, which require more peak memory per sampling step with CFG, LCM requires only one forward pass per sampling step, saving both time and memory. Moreover, in contrast to the two-stage distillation procedure employed in Guided-Distill, LCM only needs one-stage guided distillation, which is much simpler and more practical. The qualitative results in Figure 2 further show the superiority of LCM with 2- and 4-step inference.

2 Ablation Study

ODE Solvers & Skipping-Step Schedule. We compare various solvers Ψ\Psi (DDIM (Song et al., 2020a), DPM (Lu et al., 2022a), DPM++ (Lu et al., 2022b)) for solving the augmented PF-ODE specified in Eq 17, and explore different skipping step schedules with different kk. The results are depicted in Figure 3. We observe that: 1) Using Skipping-Step techniques (see Sec 4.3), LCM achieves fast convergence within 2,000 iterations in the 4-step inference setting. Specifically, the DDIM solver converges slowly at skipping step k=1k=1, while setting k=5,10,20k=5,10,20 leads to much faster convergence, underscoring the effectiveness of the Skipping-Step method. 2) DPM and DPM++ solvers perform better at a larger skipping step (k=50k=50) compared to the DDIM solver which suffers from increased ODE approximation error with larger kk. This phenomenon is also discussed in (Lu et al., 2022a). 3) Very small kk values (1 or 5) result in slow convergence and very large ones (e.g., 50 for DDIM) may lead to inferior results. Hence, we choose k=20k=20, which provides competitive performance for all three solvers, for our main experiment in Sec 5.1.

The Effect of Guidance Scale ω\omega. We examine the effect of using different CFG scales ω\omega in LCM. Typically, ω\omega balances sample quality and diversity. A larger ω\omega generally tends to improve sample quality (indicated by CLIP), but may compromise diversity (measured by FID). Beyond a certain threshold, an increased ω\omega yields better CLIP scores at the expense of FID. Figure 4 presents the results for various ω\omega across different inference steps. Our findings include: 1) Using large ω\omega enhances sample quality (CLIP Scores) but results in relatively inferior FID. 2) The performance gaps across 2, 4, and 8 inference steps are negligible, highlighting LCM's efficacy in 2∼\sim8 step regions. However, a noticeable gap exists in one-step inference, indicating rooms for further improvements. We present visualizations for different ω\omega in Figure 5. One can see clearly that a larger ω\omega enhances sample quality, verifying the effectiveness of our one-stage guided distillation method.

3 Downstream Consistency Fine-tuning Results

We perform Latent Consistency Fine-tuning (LCF) on two customized image datasets, Pokemon dataset (Pinkney, 2022) and Simpsons dataset (Norod78, 2022), to demonstrate the efficiency of LCF. Each dataset, comprised of hundreds of customized text-image pairs, is split such that 90% is used for fine-tuning and the rest 10% for testing. For LCF, we utilize pretrained LCM that was originally trained at the resolution of 768×\times768 used in Table 2. For these two datasets, we fine-tune the pre-trained LCM for 30K iterations with a learning rate 8e-6. We present qualitative results of adopting LCF on two customized image datasets in Figure 6. The finetuned LCM is capable of generating images with customized styles in few steps, showing the effectiveness of our method.

Conclusion

We present Latent Consistency Models (LCMs), and a highly efficient one-stage guided distillation method that enables few-step or even one-step inference on pre-trained LDMs. Furthermore, we present latent consistency fine-tuning (LCF), to enable few-step inference of LCMs on customized image datasets. Extensive experiments on the LAION-5B-Aesthetics dataset demonstrate the superior performance and efficiency of LCMs. Future work include extending our method to more image generation tasks such as text-guided image editing, inpainting and super-resolution.

References

Appendix A More Details on Diffusion and Consistency Models

Consider the forward process, described by the following SDE for t∈[0,T]t\in[0,T]:

where wt\bm{w}_{t} denotes the standard Brownian motion. Leveraging the classic result of Anderson (1982), Song et al. (2020b) show that the reverse process of the above forward process is also a diffusion process, specified by the following reverse-time SDE:

where w‾t\overline{\bm{w}}_{t} is a standard reverse-time Brownian motion. One can leverage the reverse SDE for data sampling from TT to , starting with qT(xT)q_{T}(\bm{x}_{T}), which follows a Gaussian distribution approximately. However, directly sampling from the reverse SDE requires a large number of discretization steps and is typically very slow. To accelerate the sampling process, prior work (e.g., (Song et al., 2020b; Lu et al., 2022a) leveraged the relation between the above SDE and ODE and designed ODE solvers for sampling. In particular, it is known that for SDE (Eq.20), the following ordinary differential equation (ODE), called the Probability Flow ODE (PF-ODE), has the same marginal distribution qt(x)q_{t}(\bm{x}) (Song et al., 2020b; Lu et al., 2022a):

The term −∇log⁡qt(xt)-\nabla\log q_{t}(\bm{x}_{t}) in Eq. 21 is typically called the score function of qt(xt)q_{t}(\bm{x}_{t}). In diffusion models, we train the noise prediction model ϵθ(xt,t)\bm{\epsilon}_{\theta}(\bm{x}_{t},t) to fit the scaled score function, via minimizing the following score matching objective:

where w(t)w(t) is the weight function, ϵ∼N(0,I)\bm{\epsilon}\sim N(0,I) and xt=α(t)x0+σ(t)ϵ\bm{x}_{t}=\alpha(t)\bm{x}_{0}+\sigma(t)\bm{\epsilon}. By substituting the score function with the noise prediction model in Eq. 21, we obtain the following ODE, which can be used for sampling:

A.2 More Details on Consistency Models in (Song et al., 2023)

In this subsection, we provide more details on the consistency models and consistency distillation algorithm in (Song et al., 2023). The pre-trained diffusion model used in (Song et al., 2023) adopts the continuous noise schedule from EDM (Karras et al., 2022), therefore the PF-ODE in Eq. 23 can be simplified as:

where the sϕ(xt,t)≈∇log⁡qt(xt)\bm{s}_{\phi}\left(\mathbf{x}_{t},t\right)\approx\nabla\log q_{t}(\bm{x}_{t}) is a score prediction model trained via score matching (Hyvärinen & Dayan, 2005; Song & Ermon, 2019). Note that different noise schedules result in different PF-ODE and the PF-ODE in Eq. 24 corresponds to the EDM noise schedule (Karras et al., 2022). We denote the one-step ODE solver applied to PF-ODE in Eq. 24 as Φ(xt,t;ϕ)\Phi({\bm{x}}_{t},t;\phi). One can either use Euler (Song et al., 2020b) or Heun solver (Karras et al., 2022) as the numerical ODE solver. Then, we use the ODE solver to estimate the evolution of a sample xtn\bm{x}_{t_{n}} from xtn+1\bm{x}_{t_{n+1}} as:

(Song et al., 2020b) used the same time schedule as in (Karras et al., 2022): ti=(ϵ1/ρ+i−1N−1(T1/ρ−ϵ1/ρ))ρt_{i}=(\epsilon^{1/\rho}+\frac{i-1}{N-1}(T^{1/\rho}-\epsilon^{1/\rho}))^{\rho}, and ρ=7\rho=7. To enforce the self-consistency property in Eq. 4, we maintain a target model θ−{\bm{\theta}}^{-}, which is updated with exponential moving average (EMA) of the parameter θ{\bm{\theta}} we intend to learn, i.e., θ−←μθ−+(1−μ)θ{\bm{\theta}}^{-}\leftarrow\mu{\bm{\theta}}^{-}+(1-\mu){\bm{\theta}}, and define the consistency loss as follows:

Appendix B Multistep Latent Consistency Sampling

Now, we present the multi-step sampling algorithm for latent consistency model. The sampling algorithm for LCM is very similar to the one in consistency models (Song et al., 2023) except the incorporation of classifier-free guidance in LCM. Unlike multi-step sampling in diffusion models, in which we predict zt−1{\bm{z}}_{t-1} from zt{\bm{z}}_{t}, the latent consistency models directly predicts the origin z0{\bm{z}}_{0} of augmented PF-ODE trajectory (the solution of the augmented of PF-ODE), given guidance scale ω\omega. This generates samples in a single step. The sample quality can be improved by alternating the denoising and noise injection steps. In particular, in the nn-th iteration, we first perform noise-injecting forward process to the previous predicted sample z{\bm{z}} as z^τn∼N(α(τn)z;σ2(τn)I)\hat{{\bm{z}}}_{\tau_{n}}\sim\mathcal{N}(\alpha(\tau_{n}){\bm{z}};\sigma^{2}(\tau_{n})\mathbf{I}), where τn\tau_{n} is a decreasing sequence of time steps. This corresponds to going back to point z^τn\hat{{\bm{z}}}_{\tau_{n}} on the PF-ODE trajectory. Then, we perform the next z0{\bm{z}}_{0} prediction again using the trained latent consistency function. In our experiments, one can see the second iteration can already refine the generation quality significantly, and high quality images can be generated in just 2-4 steps. We provide the pseudo-code in Algorithm 3.

Appendix C Algorithm Details of Latent Consistency Fine-tuning

In this section, we provide further details of Latent Consistency Fine-tuning (LCF). The pseudo-code of LCF is provided in Algorithm 4. During the Latent Consistency Fine-tuning (LCF) process, we randomly select two time steps tnt_{n} and tn+kt_{n+k} that are kk time steps apart and apply the same Gaussian noise ϵ\bm{\epsilon} to obtain the noised data ztn,ztn+k{\bm{z}}_{t_{n}},{\bm{z}}_{t_{n+k}} as follows:

Then, we can directly calculate the consistency loss for these two time steps to enforce self-consistency property in Eq.4. Notably, this method can also utilize the skipping-step technique to speedup the convergence. Furthermore, we note that latent consistency fine-tuning is independent of the pre-trained teacher model, facilitating direct fine-tuning of a pre-trained latent consistency model without reliance on the teacher diffusion model.

Appendix D Different ways to parameterize the consistency function

As previously discussed in Eq 9, we can parameterize our consistency model function fθ(z,c,t){\bm{f}}_{\bm{\theta}}({\bm{z}},{\bm{c}},t) in different ways, depending on the way the teacher diffusion model is parameterized. For ϵ-Prediction\bm{\epsilon}\text{-Prediction} (Song et al., 2020a), we use the following parameterization:

Recalling that zt=α(t)z0+σ(t)ϵ{\bm{z}}_{t}=\alpha(t){\bm{z}}_{0}+\sigma(t)\bm{\epsilon}, z^0\hat{{\bm{z}}}_{0} can be seen as a prediction of z0{\bm{z}}_{0} at time tt.

Next, we provide the parameterization of (x-Prediction)(\bm{x}\text{-Prediction}) (Ho et al., 2020; Salimans & Ho, 2022) with the following form:

where xθ(zt,c,t)\bm{x}_{\theta}({\bm{z}}_{t},{\bm{c}},t) corresponds to the teacher diffusion model with x\bm{x}-prediction.

Finally, for v\bm{v}-prediction (Salimans & Ho, 2022), the consistency function is parameterized as

where vθ(zt,c,t)\bm{v}_{\theta}({\bm{z}}_{t},{\bm{c}},t) corresponds to the teacher diffusion model with v\bm{v}-prediction.

As mentioned in Sec 5.1, we use the ϵ-Parameterization\bm{\epsilon}\text{-Parameterization} in Eq. 27 to train LCM at 512×\times512 resolution using the teacher diffusion model, Stable-Diffusion-V2.1-Base (originally trained with ϵ-Prediction\bm{\epsilon}\text{-Prediction} at 512 resolution). For resolution 768×\times768, we train the LCM using the v-Parameterization\bm{v}\text{-Parameterization} in Eq. 30, adopting the teacher diffusion model, Stable-Diffusion-V2.1 (originally trained with v-Prediction\bm{v}\text{-Prediction} at 768 resolution).

Appendix E Formulas of Other ODE Solvers

As discussed in Sec 4.3, we use the DDIM (Song et al., 2020a), DPM-Solver (Lu et al., 2022a) and DPM-Solver++ (Lu et al., 2022b) as the PF-ODE solvers. Proven in (Lu et al., 2022a), the DDIM-Solver is actually the first-order discretization approximation of the DPM-Solver.

For DDIM (Song et al., 2020a) , the detailed formula of DDIM PF-ODE solver ΨDDIM\Psi_{\text{DDIM}} from tn+kt_{n+k} to tnt_{n} is provided as follows.

For DPM-Solver (Lu et al., 2022a), we only consider the case for order=2order=2, and the detailed formula of PF-ODE solver ΨDPM-Solver\Psi_{\text{DPM-Solver}} is provided as follows. First we define some notations. We denote λtn=log⁡(αtnσtn)\lambda_{t_{n}}=\log(\frac{\alpha_{t_{n}}}{\sigma_{t_{n}}}), which is the Log-SNR, htn0=λtn−λtn+k,htn1=λtn−λtn+k/2h_{t_{n}}^{0}=\lambda_{t_{n}}-\lambda_{t_{n+k}},h_{t_{n}}^{1}=\lambda_{t_{n}}-\lambda_{t_{n+k/2}}, and rtn=htn1/htn0r_{t_{n}}=h_{t_{n}}^{1}/h_{t_{n}}^{0}.

where ϵ^\hat{\bm{\epsilon}} is the noise prediction model, and ztn+k/2Ψ{\bm{z}}^{\Psi}_{t_{n+k/2}} is the middle point between n+kn+k and nn, given by the following formula:

For DPM-Solver++ (Lu et al., 2022b), we consider the case for order=2order=2, DPM-Solver++ replaces the original noise prediction to data prediction (Lu et al., 2022b), with the detailed formula of ΨDPM-Solver++\Psi_{\text{DPM-Solver++}} provided as follows.

where x^\hat{\bm{x}} is the data prediction model (Lu et al., 2022a) and ztn+k/2Ψ{\bm{z}}^{\Psi}_{t_{n+k/2}} is the middle point between n+kn+k and nn, given by the following formula:

Appendix F Training Details of Latent Consistency Distillation

As mentioned in Section 5.1, we conduct our experiments in two resolution settings 512×\times512 and 768×\times768. For the former setting, we use the LAION-Aesthetics-6+ (Schuhmann et al., 2022) 12M dataset, consisting of 12M text-image pairs with predicted aesthetics scores higher than 6. For the latter setting, we use the LAIOIN-Aesthetic-6.5+ (Schuhmann et al., 2022), which comprise 650K text-image pairs with predicted aesthetics scores higher than 6.5.

For 512×\times512 resolution, we train the LCM with the teacher diffusion model Stable-Diffusion-V2.1-Base (SD-V2.1-Base) (Rombach et al., 2022), which is originally trained on 512×\times512 resolution images using the ϵ\bm{\epsilon}-Prediction (Ho et al., 2020). We train LCM (512×\times512) with 100K iterations on 8 A100 GPUs, using a batch size of 72, the same learning rate 8e-6 , EMA rate μ=0.999943\mu=0.999943 and Rectified Adam optimizer (Liu et al., 2019) used in (Song et al., 2023). We select the DDIM-Solver (Song et al., 2020a) and skipping step k=20k=20 in Eq. 17. We set the guidance scale range [ωmin,ωmax]=[\omega_{\text{min}},\omega_{\text{max}}]=, which is consistent with the setting in Guided-Distill (Meng et al., 2023). During training, we initialize the consistency function fθ(ztn,ω,c,tn)\bm{f_{\theta}}(\bm{{\bm{z}}}_{t_{n}},\omega,\bm{c},t_{n}) with the same parameters as the teacher diffusion model (SD-V2.1-Base). To encode the CFG scale ω\omega into the LCM, we applying Fourier embedding to ω\omega, integrating it into the origin LCM backbone by adding the projected ω\omega-embedding into the original embedding, as done in (Meng et al., 2023). We use a zero parameter initialization method mentioned in (Zhang & Agrawala, 2023) on projected ω\omega-embedding for better training stability. For training LCM (512×\times512), we use a augmented consistency function parameterized in ϵ\bm{\epsilon}-prediction as discussed in Appendix. D.

For 768×\times768 resolution, we train the LCM with the teacher diffusion model Stable-Diffusion-V2.1 (SD-V2.1) (Rombach et al., 2022), which is originally trained on 768×\times768 resolution images using the v\bm{v}-Prediction (Salimans & Ho, 2022). We train LCM (768×\times768) with 100K iterations on 8 A100 GPUs using a batch size of 16, while the other hyper-parameters remain the same as in 512×\times512 resolution setting.

Appendix G Reproduction Details of Guided-Distill

Guided-Distill (Meng et al., 2023) serves as a significant baseline for guided distillation but is not open-sourced. We adhered strictly to the training procedure described in the paper, reproducing the method for accurate comparisons. For 512×\times512 resolution setting, Guided-Distill (Meng et al., 2023) used a large batch size of 512, which requires at least 32 A100 GPUs for training. Due to limited resource, we reduced the batch size to 72 (512 resolution), while set the batchsize to 16 for 768 resolution, the same as ours, and trained for 100K iterations, also the same as in LCM.

Specifically, Guided Distill involves two stages of distillation. For the first stage, it use a student model to fit the outputs of the pre-trained guided diffusion model using classifier-free guidance scales ω\omega. The loss function is as follows:

where x^θ(zt)=(1+w)x^c,θ(zt)−wx^θ(zt),zt∼q(zt∣x)\hat{{\bm{x}}}_{\bm{\theta}}({\bm{z}}_{t})=(1+w)\hat{{\bm{x}}}_{c,{\bm{\theta}}}({\bm{z}}_{t})-w\hat{{\bm{x}}}_{\bm{\theta}}({\bm{z}}_{t}),{\bm{z}}_{t}\sim q({\bm{z}}_{t}|{\bm{x}}) and pw(w)=U[wmin,wmax]p_{w}(w)=\mathcal{U}[w_{\text{min}},w_{\text{max}}].

In our implementation, we follow the same training procedure in (Meng et al., 2023) except the difference of computation resources. For first stage distillation, we train the student model with 25,000 gradient updates (batch size 72), roughly the same computation costs in (Meng et al., 2023) (3,000 gradient updates, batch size 512), and we reduce the original learning rate 1e−41e-4 to 5e−55e-5 for smaller batch size. For second stage distillation, we progressively train the student model using the same schedule as in Guided-Distill (Meng et al., 2023) except for batch size difference. We train the student model with 2500 gradient updates except when the sampling step equals to 1,2, or 4, where we train for 20000 gradient updates, the same schedule used in (Meng et al., 2023). We trained until the total number of gradient iterations for the entire stage reached 100K the same as LCM training. The generation results of Guided Distill are shown in Figure 2. We can also see that the performances in Table 1 and Table 2 are similar, further verifying the correctness of our Guided-Distill implementation. Nevertheless, we acknowledge that longer training and more computational resources can lead to better results as reported in (Meng et al., 2023). However, LCM achieves faster convergence and superior results under the same computation cost (same batch size, same number of iterations), demonstrating its practicability and superiority.

Appendix H More few-step Inference Results

We present more images (768×\times768) generation results with LCM using 4 and 2-steps inference in Figure 7 and Figure 8. It is evident that LCM is capable of synthesizing high-resolution images with just 2, 4 steps of inference. Moreover, LCM can be derived from any pre-trained Stable Diffusion (SD) (Rombach et al., 2022) in merely 4,000 training steps, equivalent to around 32 A100 GPU Hours, showcasing the effectiveness and superiority of LCM.