Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

Yuxi Ren, Xin Xia, Yanzuo Lu, Jiacheng Zhang, Jie Wu, Pan Xie, Xing Wang, Xuefeng Xiao

Introduction

Diffusion models (DMs) have gained significant prominence in the field of Generative AI , but they are burdened by the computational requirements associated with multi-step inference procedures . To overcome these challenges and fully exploit the capabilities of DMs, several distillation methods have been proposed , which can be categorized into two main groups: trajectory-preserving distillation and trajectory-reformulating distillation.

Trajectory-preserving distillation techniques are designed to maintain the original trajectory of an ordinary differential equation (ODE) . The primary objective of these methods is to enable student models to make further predictions on the flow and reduce the overall number of inference steps. These techniques prioritize the preservation of similarity between the outputs of the distilled model and the original model. Adversarial losses can also be employed to enhance the accuracy of supervised guidance in the distillation process . However, it is important to note that, despite their benefits, trajectory-preserved distillation approaches may suffer from a decrease in generation quality due to inevitable errors in model fitting.

Trajectory-reformulating methods directly utilize the endpoint of the ODE flow or real images as the primary source of supervision, disregarding the intermediate steps of the trajectory . By reconstructing more efficient trajectories, these methods can also reduce the number of inference steps. Trajectory-reformulating approaches enable the exploration of the model’s potential within a limited number of steps, liberating it from the constraints of the original trajectory. However, it can lead to inconsistencies between the accelerated model and the original model’s output domain, often resulting in undesired effects.

To navigate these hurdles and harness the full potential of DMs, we present an advanced framework that adeptly combines trajectory-preserving and trajectory-reformulating distillation techniques. Firstly, we proposed trajectory segmented consistency distillation (TSCD), which divides the time steps into segments and enforces consistency within each segment while gradually reducing the number of segments to achieve all-time consistency. This approach addresses the issue of suboptimal consistency model performance caused by insufficient model fitting capability and accumulated errors in inference. Secondly, we leverage human feedback learning techniques to optimize the accelerated model, modifying the ODE trajectories to better suit few-step inference. This results in significant performance improvements, even surpassing the capabilities of the original model in some scenarios. Thirdly, we enhanced the one-step generation performance using score distillation , achieving the idealized all-time consistent model via a unified LORA. In summary, our main contributions are summarized as follows:

Accelerate: we propose TSCD that achieves a more fine-grained and high-order consistency distillation approach for the original score-based model.

Boost: we incorpoate human feedback learning to further enhance model performance in low-steps regime.

Unify: we provide a unified LORA as the all-time consistency model and support inference at all NTEs.

Performance: Hyper-SD achieves SOTA performance in low-steps inference for both SDXL and SD1.5.

Preliminaries

Diffusion models (DMs), as introduced by Ho et al. , consist of a forward diffusion process, described by a stochastic differential equation (SDE) , and a reverse denoising process. The forward process gradually adds noise to the data, transforming the data distribution pdata(x)p_{\text{data}}(x) into a known distribution, typically Gaussian. This process is described by:

where t∈[0,T]t\in[0,T], wtw_{t} represents the standard Brownian motion, μ(⋅,⋅)\mu(\cdot,\cdot) and σ(⋅)\sigma(\cdot) are the drift and diffusion coefficients respectively. The distribution of xtx_{t} sampled during the diffusion process is denoted as pt(x)p_{\text{t}}(x), with the empirical data distribution p0(x)≡pdata(x)p_{\text{0}}(x)\equiv p_{\text{data}}(x), and pT(x)p_{\text{T}}(x) being approximated by a tractable Gaussian distribution.

This SDE is proved to have the same solution trajectories as an ordinary differential equation (ODE) , dubbed as Probability Flow (PF) ODE, which is formulated as

2 Diffusion Model Distillation

As mentioned in Sec. 1, current techniques for distilling Diffusion Models (DMs) can be broadly categorized into two approaches: one that preserves the Ordinary Differential Equation (ODE) trajectory , and another that reformulates it .

Here, we provide a concise overview of some representative categories of methods. For clarity, we define the teacher model as fteaf_{tea}, the student model as fstuf_{stu}, noise as ϵ\epsilon, prompt condition as cc, off-the-shelf ODE Solver as Ψ(⋅,⋅,⋅)\Psi(\cdot,\cdot,\cdot), the total training timesteps as TT, the num of inference timesteps as NN, the noised trajectory point as xtx_{t} and the skipping-step as ss, where t0<t1⋯<tN−1=Tt_{0}<t_{1}\cdots<t_{N-1}=T, tn−tn−1=st_{n}-t_{n-1}=s, nn uniformly distributed over {1,2,…,N−1}\{1,2,\ldots,N-1\}.

Progressive Distillation. Progressive Distillation (PD) trains the student model fstuf_{stu} approximate the subsequent flow locations determined by the teacher model fteaf_{tea} over a sequence of steps.

Considering a 2-step PD for illustration, the target prediction x^tn−2\hat{x}_{t_{n-2}} by fteaf_{tea} is obtained through the following calculations:

Consistency Distillation. Consistency Distillation (CD) directly maps xtnx_{t_{n}} along the ODE trajectory to its endpoint x0x_{0}. The training loss is defined as :

where fstu−f^{-}_{stu} is the exponential moving average(EMA) of fstuf_{stu} and x^tn−1\hat{x}_{t_{n-1}} is the next flow location estimated by fteaf_{tea} with the same function as Eq. 3.

The Consistency Trajectory Model (CTM) was introduced to minimize accumulated estimation errors and discretization inaccuracies prevalent in multi-step consistency model sampling. Diverging from targeting the endpoint x0x_{0}, CTM targets any intermediate point xtendx_{t_{end}} within the range 0≤tend≤tn−10\leq t_{end}\leq t_{n-1}, thus redefining the loss function as:

Adversarial Diffusion Distillation. In contrast to PD and CD, Adversarial Distillation (ADD), proposed in SDXL-Turbo and SD3-Turbo , bypasses the ODE trajectory and directly focuses on the original state x0x_{0} using adversarial objective. The generative and discriminative loss components are computed as follows:

where DD denotes the discriminator, tasked with differentiating between x0x_{0} and Ψ(xtn,fstu(xtn,tn,c),0)\Psi(x_{t_{n}},f_{stu}(x_{t_{n}},t_{n},c),0). The target x0x_{0} can be sampled from real or synthesized data.

Score Distillation Sampling. Score distillation sampling(SDS) was integrated into diffusion distillation in SDXL-Turbo and Diffusion Matching Distillation(DMD). SDXL-Turbo utilizes fteaf_{tea} to estimate the score to the real distribution, while DMD further introduced a fake distribution simulator ffakef_{fake} to calibrate the score direction and uses the output distribution of the original model as the real distribution, thus achieving one-step inference.

Leveraging the DMD approach, the gradient of the Kullback-Leibler (KL) divergence between the real and fake distributions is approximated by the equation:

where zz is a random latent variable sampled from a standard normal distribution. This methodology enables the one-step diffusion model to refine its generative process, minimizing the KL divergence to produce images that are progressively closer to the teacher model’s distribution.

3 Human Feedback Learning

ReFL has been proven to be an effective method to learn from human feedback designed for diffusion models. It primarily includes two stages: (1) reward model training and (2) preference fine-tuning. In the first stage, given the human preference data pair, xwx_{w} (preferred generation) and xlx_{l} (unpreferred one), a reward model rθr_{\theta} is trained via the loss:

where D\mathcal{D} denotes the collected feedback data, σ(⋅)\sigma(\cdot) represents the sigmoid function, and cc corresponds to the text prompt. The reward model rθr_{\theta} is optimized to produce reward scores that align with human preferences. In the second stage, ReFL starts with an input prompt cc, and a randomly initialized latent xT=zx_{T}=z. The latent is then iteratively denoised until reaching a randomly selected timestep tn∈[tleft,tright]t_{n}\in[t_{left},t_{right}], when a denoised image x0′x^{\prime}_{0} is directly predicted from xtnx_{t_{n}}. The tleftt_{left} and trightt_{right} are predefined boundaries. The reward model is then applied to this denoised image, generating the expected preference score rθ(c,x0′)r_{\theta}(c,x^{\prime}_{0}), which is used to fine-tuned the diffusion model:

Method

In this study, we have integrated both the ODE-preserve and ODE-reformulate distillation techniques into a unified framework, yielding significant advancements in accelerating diffusion models. In Sec. 3.1, we propose an innovative approach to consistency distillation that employs a time-steps segmentation strategy, thereby facilitating trajectory segmented consistency distillation. In Sec. 3.2, we incorporate human feedback learning techniques to further enhance the performance of accelerated diffusion models. In Sec. 3.3, we achieve all-time consistency including one-step by utilizing the score-based distribution matching distillation.

Both Consistency Distillation (CD) and Consistency Trajectory Model (CTM) aim to transform a diffusion model into a consistency model across the entire timestep range [0,T][0,T] through single-stage distillation. However, these distilled models often fall short of optimality due to limitations in model fitting capacity. Drawing inspiration from the soft consistency target introduced in CTM, we refine the training process by dividing the entire time-steps range [0,T][0,T] into kk segments and performing segment-wise consistent model distillation progressively.

In the first stage, we set k=8k=8 and use the original diffusion model to initiate fstuf_{stu} and fteaf_{tea}. The starting timesteps tnt_{n} are uniformly and randomly sampled from {t1,t2,…,tN−1}\{t_{1},t_{2},\ldots,t_{N-1}\}. We then sample ending timesteps tend∈[tb,tn−1]t_{end}\in[t_{b},t_{n-1}] , where tbt_{b} is computed as:

where x^tn−1\hat{x}_{t_{n-1}} is computed as Eq. 3, and fstu−f^{-}_{stu} denotes the Exponential Moving Average (EMA) of fstuf_{stu}.

Subsequently, we resume the model weights from the previous stage and continue to train fstuf_{stu}, progressively reducing kk to $.Itisnoteworthythat. It is noteworthy thatk=1correspondstothestandardCTMtrainingprotocol.Forthedistancemetriccorresponds to the standard CTM training protocol. For the distance metricd,weemployahybridofadversarialloss,asproposedinsdxl−lightning,andMeanSquaredError(MSE)Loss.Empirically,weobservethatMSELossismoreeffectivewhenthepredictionsandtargetvaluesareproximate(e.g.,for, we employ a hybrid of adversarial loss, as proposed in sdxl-lightning, and Mean Squared Error (MSE) Loss. Empirically, we observe that MSE Loss is more effective when the predictions and target values are proximate (e.g., fork=8,4),whereasadversariallossprovesmorepreciseasthedivergencebetweenpredictionsandtargetsincreases(e.g.,for), whereas adversarial loss proves more precise as the divergence between predictions and targets increases (e.g., fork=2,1).Accordingly,wedynamicallyincreasetheweightoftheadversariallossanddiminishthatoftheMSElossacrossthetrainingstages.Additionally,wehaveintegratedanoiseperturbationmechanismtoreinforcetrainingstability.Takethetwo−stageTrajectorySegmentedConsistencyDistillation(TSCD)processasanexample.AsshowninFig.2,thefirststageexecutesindependentconsistencydistillationswithinthetimesegments). Accordingly, we dynamically increase the weight of the adversarial loss and diminish that of the MSE loss across the training stages. Additionally, we have integrated a noise perturbation mechanism to reinforce training stability. Take the two-stage Trajectory Segmented Consistency Distillation(TSCD) process as an example. As shown in Fig. 2, the first stage executes independent consistency distillations within the time segments[0,\frac{T}{2}]andand[\frac{T}{2},T]$. Based on the previous two-segment consistency distillation results, a global consistency trajectory distillation is then performed.

The TSCD method offers two principal advantages: Firstly, fine-grained segment distillation reduces model fitting complexity and minimizes errors, thus mitigating degradation in generation quality. Secondly, it ensures the preservation of the original ODE trajectory. Models from each training stage can be utilized for inference at corresponding steps while closely mirroring the original model’s generation quality. We illustrate the complete procedure of Progressive Consistency Distillation in Algorithm 1. It is worth noting that, by utilizing Low-Rank Adaptation(LoRA) technology, we train TSCD models as plug-and-play plugins that can be used instantly.

2 Human Feedback Learning

In addition to the distillation, we propose to incorporate feedback learning further to boost the performance of the accelerated diffusion models. In particular, we improve the generation quality of the accelerated models by exploiting the feedback drawn from both human aesthetic preferences and existing visual perceptual models. For the feedback on aesthetics, we utilize the LAION aesthetic predictor and the aesthetic preference reward model provided by ImageReward to steer the model toward the higher aesthetic generation as:

where rdr_{d} is the aesthetic reward model, including the aesthetic predictor of the LAION dataset and ImageReward model, cc is the textual prompt and αd\alpha_{d} together with ReLU function works as a hinge loss.

Beyond the feedback from aesthetic preference, we notice that the existing visual perceptual model embedded in rich prior knowledge about the reasonable image can also serve as a good feedback provider. Empirically, we found that the instance segmentation model can guide the model to generate entities with reasonable structure. To be specific, instead of starting from a random initialized latent, we first diffuse the noise on an image x0x_{0} in the latent space to xtx_{t} according to Eq. 1, and then, we execute denoise iteratively until a specific timestep dtd_{t} and directly predict a x0′x^{{}^{\prime}}_{0} similar to . Subsequently, we leverage perceptual instance segmentation models to evaluate the performance of structure generation by examining the perceptual discrepancies between the ground truth image instance annotation and the predicted results on the denoised image as:

where mIm_{I} is the instance segmentation model(e.g. SOLO ). The instance segmentation model can capture the structure defect of the generated image more accurately and provide a more targeted feedback signal. It is noteworthy that besides the instance segmentation model, other perceptual models are also applicable and we are actively investigating the utilization of advanced large visual perception models(e.g. SAM) to provide enhanced feedback learning. Such perceptual models can work as complementary feedback for the subjective aesthetic focusing more on the objective generation quality. Therefore, we optimize the diffusion models with the feedback signal as:

Human feedback learning can improve model performance but may unintentionally alter the output domain, which is not always desirable. Therefore, we also trained human feedback learning knowledge as a plugin using LoRA technology. By employing the LoRA merge technique with the TSCD LoRAs discussed in Section3.1, we can achieve a flexible balance between generation quality and output domain similarity.

3 One-step Generation Enhancement

One-step generation within the consistency model framework is not ideal due to the inherent limitations of consistency loss. As analyzed in Fig. 3, the consistency distilled model demonstrates superior accuracy in guiding towards the trajectory endpoint x0x_{0} at position xtx_{t}. Therefore, score distillation is a suitable and efficient way to boost the one-step generation of our TSCD models.

Specifically, we advance one-step generation with an optimized Distribution Matching Distillation (DMD) technique . DMD enhances the model’s output by leveraging two distinct score functions: freal(x)f_{real}(x) from the teacher model’s distribution and ffake(x)f_{fake}(x) from the fake model. We incorporate a Mean Squared Error (MSE) loss alongside the score-based distillation to promote training stability. The human feedback learning technique mentioned in Sec. 3.2 is also integrated, fine-tuning our models to efficiently produce images of exceptional fidelity.

After enhancing the one-step inference capability of the TSCD model, we can obtain an ideal global consistency model. Employing the TCD scheduler, the enhanced model can perform inference from 1 to 8 steps. Our approach eliminates the need for model conversion to x0-prediction, enabling the implementation of the one-step LoRA plugin. We demonstrated the effectiveness of our one-step LoRA in Sec 4.3. Additionally, smaller time-step inputs can enhance the credibility of the one-step diffusion model in predicting the noise . Therefore, we also employed this technique to train a dedicated model for single-step generation.

Experiments

Dataset. We use a subset of the LAION and COYO datasets following SDXL-lightning during the training procedure of Sec 3.1 and Sec 3.3. For the Human Feedback Learning in Sec 3.2, we generated approximately 140k artist-style text images for style optimization using the SDXL-Base model and utilized the COCO2017 train split dataset with instance annotations and captions for structure optimization.

Training Setting. For TSCD in Sec 3.1, we progressively reduced the time-steps segments number as 8→4→2→18\rightarrow 4\rightarrow 2\rightarrow 1 in four stages, employing 512 batch size and learning rate 1e−61e-6 across 32 NVIDIA A100 80GB GPUs. We trained Lora instead of Unet for all the distillation stages for convenience, and the corresponding Lora is loaded to process the human feedback learning optimization in Sec 3.2. For one-step enhancement in Sec 3.3, we trained the unified all-timesteps consistency Lora with time-step inputs T=999T=999 and the dedicated model for single-step generation with T=800T=800.

Baseline Models. We conduct our experiments on the stable-diffusion-v1-5(SD15) and stable-diffusion-xl-v1.0-base(SDXL) . To demonstrate the superiority of our method in acceleration, we compared our method with various existing acceleration schemes as shown in Tab. 1.

Evaluation Metrics. We use the aesthetic predictor pre-trained on the LAION dataset and CLIP score(ViT-B/32) to evaluate the visual appeal of the generated image and the text-to-image alignment. We further include some recently proposed metrics, such as ImageReward score , and Pickscore to offer a more comprehensive evaluation of the model performance. Note that we do not report the Fréchet Inception Distance(FID) as we observe it can not well reflect the actual generated image quality in our experiments. In addition to these, due to the inherently subjective nature of image generation evaluation, we conduct an extensive user study to evaluate the performance more accurately.

2 Main Results

Quantitative Comparison. We quantitatively compare our method with both the baseline and diffusion-based distillation approaches in terms of objective metrics. The evaluation is performed on COCO-5k dataset with both SD15 (512px) and SDXL (1024px) architectures. As shown in Tab. 2, our method significantly outperforms the state-of-the-art across all metrics on both resolutions. In particular, compared to the two baseline models, we achieve better aesthetics (including AesScore, ImageReward and PickScore) with only LoRA and fewer steps. As for the CLIPScore that evaluates image-text matching, we outperform other methods by +0.1 faithfully and are also closest to the baseline model, which demonstrates the effectiveness of our human feedback learning.

Qualitative Comparison. In Figs. 5, 4 and 6, we present comprehensive visual comparison with recent approaches, including LCM , TCD , PeRFLow , Turbo and Lightning . Our observations can be summarized as follows. (1) Thanks to the fact that SDXL has almost 2.6B parameters, the model is able to synthesis decent images in 4 steps after different distillation algorithms. Our method further utilizes its huge model capacity to compress the number of steps required for high-quality outcomes to 1 step only, and far outperforms other methods in terms of style (a), aesthetics (b-c) and image-text matching (d) as indicated in Fig. 4. (2) On the contrary, limited by the capacity of SD15 model, the images generated by other approaches tend to exhibit severe quality degradation. While our Hyper-SD consistently yields better results across different types of user prompts, including photographic (a), realistic (b-c) and artstyles (d) as depicted in Fig. 5. (3) To further release the potential of our methodology, we also conduct experiments on the fully fine-tuning of SDXL model following previous works . As shown in Fig. 6, our 1-Step UNet again demonstrates superior generation quality that far exceeds the rest of the opponents. Both in terms of colorization (a-b) and details (c-d), our images are more presentable and attractive when it comes to the real-world application scenarios.

User Study. To verify the effectiveness of our proposed Hyper-SD, we conduct an extensive user study across various settings and approaches. As presented in Fig. 7, our method (red in left) obtains significantly more user preferences than others (blue in right). Specifically, our Hyper-SD15 has achieved more than a two-thirds advantage against the same architectures. The only exception is that SD21-Turbo was able to get significantly closer to our generation quality in one-step inference by means of a larger training dataset of SD21 model as well as fully fine-tuning. Notably, we found that we obtained a higher preference with less inference steps compared to both the baseline SD15 and SDXL models, which once again confirms the validity of our human feedback learning. Moreover, our 1-Step UNet shows a higher preference than LoRA against the same UNet-based approaches (i.e. SDXL-Turbo and SDXL-Lightning ), which is also consistent with the analyses of previous quantitative and qualitative comparisons. This demonstrates the excellent scalability of our method when more parameters are fine-tuned.

3 Ablation Study

Unified LoRA. In addition to the different steps of LoRAs proposed above, we note that our one-step LoRA can be considered as a unified approach, since it is able to reason about different number of steps (e.g. 1,2,4,8 as shown in Fig. 8) and consistently generate high-quality results under the effect of consistency distillation. For completeness, Tab. 3 also presents the quantitative results of different steps when applying the 1-Step unified LoRA. We can observe that there is no difference in image-text matching between different steps as the CLIPScore evaluates, which means that user prompts are well adhered to. And as the other metrics show, the aesthetics rise slightly as the step increases, which is as expected after all the user can choose based on the needs for efficiency. This would be of great convenience and practicality in real-world deployment scenarios, since generally only one model can be loaded per instance.

Compatibility with Base Model. Fig. 9 shows that our LoRAs can be applied to different base models. Specifically, we conduct experiments on animehttps://civitai.com/models/112902, realistichttps://civitai.com/models/133005 and artstylehttps://civitai.com/models/119229 base models. The results demonstrate that our method has a wide range of applications, and the lightweight LoRA also significantly reduces the cost of acceleration.

Compatibility with ControlNet. Fig. 10 shows that our models are also compatible with ControlNet . We test the one-step unified SD15 and SDXL LoRAs on the scribblehttps://huggingface.co/lllyasviel/control_v11p_sd15_scribble and cannyhttps://huggingface.co/diffusers/controlnet-canny-sdxl-1.0 control images, respectively. And we can observe the conditions are well followed and the consistency of our unified LoRAs can still be demonstrated, where the quality of generated images under different inference steps are always guaranteed.

Discussion and Limitation

Hyper-SD demonstrates promising results in generating high-quality images with few inference steps. However, there are several avenues for further improvement:

Classifier Free Guidance: the CFG properties of diffusion models allow for improving model performance and mitigating explicit content, such as pornography, by adjusting negative prompts. However, most diffusion acceleration methods including ours, eliminated the CFG characteristics, restricting the utilization of negative cues and imposing usability limitations. Therefore, in future work, we aim to retain the functionality of negative cues while accelerating the model, enhancing both generation effectiveness and security.

Customized Human Feedback Optimization: this work employed the generic reward models for feedback learning. Future work will focus on customized feedback learning strategies designed specifically for accelerated models to enhance their performance.

Diffusion Transformer Architecture: Recent studies have demonstrated the significant potential of DIT in image generation, we will focus on the DIT architecture to explore superior few-steps generative diffusion models in our future work.

Conclusion

We propose Hyper-SD, a unified framework that maximizes the few-step generation capacity of diffusion models, achieving new SOTA performance based on SDXL and SD15. By employing trajectory-segmented consistency distillation, we enhanced the trajectory preservation ability during distillation, approaching the generation proficiency of the original model. Then, human feedback learning and variational score distillation stimulated the potential for few-step inference, resulting in a more optimal and efficient trajectory for generating models. We have open-sourced Lora plugins for SDXL and SD15 from 1 to 8 steps inference, along with a dedicated one-step SDXL model, aiming to further propel the development of the generative AI community.

References