UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs

Yanwu Xu, Yang Zhao, Zhisheng Xiao, Tingbo Hou

Introduction

Diffusion models has recently emerged as a powerful class of generative models, demonstrating unprecedented results in many generative modeling tasks . In particular, they have shown the remarkable ability to synthesize high-quality images conditioned on texts . Beyond the text-to-image synthesis tasks, large-scale text-to-image models serve as foundational building blocks for various downstream applications, including personalized generation , controlled generation and image editing . Yet, despite their impressive generative quality and wide-ranging utility, diffusion models have a notable limitation: they rely on iterative denoising to generate final samples, which leads to slow generation speeds. The slow inference and the consequential computational demands of large-scale diffusion models pose significant impediments to their deployment.

In the seminal work by Song et al. , it was revealed that sampling from a diffusion model is equivalent to solving the probability flow ordinary differential equation (PF-ODE) associated with the diffusion process. Presently, the majority of research aimed at enhancing the sampling efficiency of diffusion models centers on the ODE formulation. One line of work seeks to advance numerical solvers for the PF-ODE, with the intention of enabling the solution of the ODE with greater discretization size, ultimately leading to fewer requisite sampling steps . However, the inherent trade-off between step size and accuracy still exists. Given the highly non-linear and complicated trajectory of the PF-ODE, it would be extremely difficult to reduce the number of required sampling steps to a minimal level. Even the most advanced solvers can generate images within 10 to 20 sampling steps, and further reduction leads to a noticeable drop in image quality. An alternative approach seeks to distill the PF-ODE trajectory from a pre-trained diffusion model. For instance, progressive distillation tries to condense multiple discretization steps of the PF-ODE solver into a single step by explicitly aligning with the solver’s output. Similarly, consistency distillation works on learning consistency mappings that preserve point consistency along the ODE trajectory. These methods have demonstrated the potential to significantly reduce the number of sampling steps. However, due to the intrinsic complexity of the ODE trajectory, they still struggle in the extremely small step regime, especially for large-scale text-to-image diffusion models.

The pursuit of developing ultra-fast large-scale diffusion models that requires just one or two sampling steps, remains a challenging open problem. We assert that to achieve this ambitious objective, fundamental adjustments are necessary in the formulation of diffusion models, as the current ODE-based approach seems intrinsically constrained for very few steps sampling, as elucidated earlier. In this work, we introduce a novel one-step text-to-image generative model, representing a fusion of GAN and diffusion model elements. Our inspiration stems from previous work that successfully incorporated GANs into the framework of diffusion models , which have demonstrated the capacity to generate images in as few as four steps when trained on small-scale datasets. These models diverge from the traditional ODE formulation by leveraging adversarial loss for learning the denoising distribution, rather than relying on KL minimization. Section 3 offers a comprehensive review of existing diffusion-GAN hybrid models.

Despite the promising outcomes of earlier diffusion GAN hybrid models, achieving one-step sampling and extending their utility to text-to-image generation remains a non-trivial challenge. In this research, we introduce innovative techniques to enhance diffusion GAN models, resulting in an ultra-fast text-to-image model capable of producing high-quality images in a single sampling step. In light of this achievement, we have named our model UFOGen, an acronym denoting “You Forward Once” Generative model. A detailed exposition of UFOGen is presented in Section 4. Our UFOGen model excels at generating high-quality images in just one inference step. Notably, when initialized with a pre-trained Stable Diffusion model , our method efficiently transforms Stable Diffusion into a one-step inference model while largely preserving the quality of generated content. See Figure 1 for a showcase of text-conditioned images generated by UFOGen. To the best of our knowledge, our model stands among the pioneers to achieve a reduction in the number of required sampling steps for text-to-image diffusion models to just one.

Our work presents several significant contributions:

We introduce UFOGen, a powerful generative model capable of producing high-quality images conditioned on text descriptions in a single inference step.

We present an efficient and simplified training process, enabling the fine-tuning of pre-existing large-scale diffusion models, like Stable Diffusion, to operate as one-step generative models.

Our model’s versatility extends to applications such as image-to-image and controllable generation, thereby unlocking the potential for one-step inference across various generative scenarios.

Related Works

Text-to-image Diffusion Models Denoising diffusion models are trained to reconstruct data from corrupted inputs. The simplicity of the training objective makes denoising diffusion models well-suited for scaling up generative models. Researchers have made numerous efforts to train diffusion models on large datasets containing image-text pairs for the text-to-image generation task . Among these, latent diffusion models, such as the popular Stable Diffusion model , have gained substantial attention in the research community due to their simplicity and efficiency compared to pixel-space counterparts.

Accelerating Diffusion Models The notable issue of slow generation speed has motivated considerable efforts towards enhancing the sampling efficiency of diffusion models. These endeavors can be categorized into two primary approaches. The first focuses on the development of improved numerical solvers . The second approach explores the concept of knowledge distillation , aiming at condensing the sampling trajectory of a numerical solver into fewer steps . However, both of these approaches come with significant limitations, and thus far, they have not demonstrated the ability to substantially reduce the sampling steps required for text-to-image diffusion models to a truly minimal level.

Text-to-image GANs As our model has GAN as one of its component, we provide a brief overview of previous attempts of training GANs for text-to-image generation. Early GAN-based text-to-image models were primarily confined to small-scale datasets . Later, with the evolution of more sophisticated GAN architectures , GANs trained on large datasets have shown promising results in the domain of text-to-image generation . Comparatively, our model has several distinct advantages. Firstly, to overcome the well-known issues of training instability and mode collapse, text-to-image GANs have to incorporate multiple auxiliary losses and complex regularization techniques, which makes training and parameter tuning extremely intricate. This complexity is particularly exemplified by GigaGAN , currently regarded as the most powerful GAN-based models. In contrast, our model offers a streamlined and robust training process, thanks to the diffusion component. Secondly, our model’s design allows us to seamlessly harness pre-trained diffusion models for initialization, significantly enhancing the efficiency of the training process. Lastly, our model exhibits greater flexibility when it comes to downstream applications (see Section 5.3), an area in which GAN-based models have not explored.

Recent Progress on Few-step Text-to-image Generation While developing our model, we noticed some concurrent work on few-step text-to-image generation. Latent Consistency Model extends the idea of consistency distillation to Stable Diffusion, leading to 4-step sampling with reasonable quality. However, further reducing the sampling step results in significant quality drop. InstaFlow achieves text-to-image generation in a single sampling step. Similar to our model, InstaFlow tackles the slow sampling issue of diffusion models by introducing improvements to the model itself. Notably, they extend Rectified Flow models to create a more direct trajectory in the diffusion process. In direct comparison to InstaFlow, our model outperforms in terms of both quantitative metrics and visual quality. Moreover, our approach presents the added benefits of a streamlined training pipeline and improved training efficiency. InstaFlow requires multiple stages of fine-tuning, followed by a subsequent distillation stage. In contrast, our model only need one single fine-tuning stage with a minimal number of training iterations.

Background

Diffusion Models Diffusion models is a family of generative models that progressively inject Gaussian noises into the data, and then generate samples from noise via a reverse denoising process. Diffusion models define a forward process that corrupts data x0∼q(x0)x_{0}\sim q(x_{0}) in TT steps with variance schedule βt\beta_{t}: q(xt∣xt−1):=N(xt;1−βtxt−1,βtI)q(x_{t}|x_{t-1}):=\mathcal{N}(x_{t};\sqrt{1-\beta_{t}}x_{t-1},\beta_{t}\textbf{I}). The parameterized reversed diffusion process aims to gradually recover cleaner data from noisy observations: pθ(xt−1∣xt):=N(xt−1;μθ(xt,t),σt2I)p_{\theta}(x_{t-1}|x_{t}):=\mathcal{N}(x_{t-1};\mu_{\theta}(x_{t},t),\sigma_{t}^{2}\textbf{I}).

The model pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}) is parameterized as a Gaussian distribution, because when the denoising step size from tt to t−1t-1 is sufficiently small, the true denoising distribution q(xt−1∣xt)q(x_{t-1}|x_{t}) is a Gaussian . To train the model, one can minimize the negative ELBO objective :

where q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}) is Gaussian posterior distribution derived in .

Diffusion-GAN Hybrids The idea of combining diffusion models and GANs is first explored in . The main motivation is that, when the denoising step size is large, the true denoising distribution q(xt−1∣xt)q(x_{t-1}|x_{t}) is no longer a Gaussian. Therefore, instead of minimizing KL divergence with a parameterized Gaussian distribution, they parameterized pθ(xt−1′∣xt)p_{\theta}(x_{t-1}^{\prime}|x_{t}) as a conditional GAN to minimize the adversarial divergence between model pθ(xt−1′∣xt)p_{\theta}(x_{t-1}^{\prime}|x_{t}) and q(xt−1∣xt)q(x_{t-1}|x_{t}):

The objective of Denoising Diffusion GAN (DDGAN) in can be expressed as:

where DϕD_{\phi} is the conditional discriminator network, and the expectation over the unknown distribution q(xt−1∣xt)q(x_{t-1}|x_{t}) can be approximated by sampling from q(x0)q(xt−1∣x0)q(xt∣xt−1)q(x_{0})q(x_{t-1}|x_{0})q(x_{t}|x_{t-1}). The flexibility of a GAN-based denoising distribution surpasses that of a Gaussian parameterization, enabling more aggressive denoising step sizes. Consequently, DDGAN successfully achieves a reduction in the required sampling steps to just four.

Nonetheless, the utilization of a purely adversarial objective in DDGAN introduces training instability, as documented by the findings in . In response to this challenge, the authors in advocated matching the joint distribution q(xt−1,xt)q(x_{t-1},x_{t}) and pθ(xt−1,xt)p_{\theta}(x_{t-1},x_{t}), as opposed to the conditional distribution as outlined in Equation 2. further demonstrated that the joint distribution matching can be disassembled into two components: matching marginal distributions using adversarial divergence and matching conditional distributions using KL divergence:

The objective of adversarial divergence minimization is similar to Equation 3 except that the discriminator does not take xtx_{t} as part of its input. The KL divergence minimization translates into a straightforward reconstruction objective, facilitated by the Gaussian nature of the diffusion process (see Appendix A.1 for a derivation). This introduction of a reconstruction objective plays a pivotal role in enhancing the stability of the training dynamics. As observed in , which introduced Semi-Implicit Denoising Diffusion Models (SIDDMs), this approach led to markedly improved results, especially on more intricate datasets.

Methods

In this section, we present a comprehensive overview of the enhancements we have made in our diffusion-GAN hybrid models, ultimately giving rise to the UFOGen model. These improvements are primarily focused on two critical domains: 1) enabling one step sampling, as detailed in Section 4.1, and 2) scaling-up for text-to-image generation, as discussed in Section 4.2.

Diffusion-GAN hybrid models are tailored for training with a large denoising step size. However, attempting to train these models with just a single denoising step (i.e., xT−1=x0x_{T-1}=x_{0}) effectively reduces the training to that of a conventional GAN. Consequently, prior diffusion-GAN models were unable to achieve one-step sampling. In light of this challenge, we conducted an in-depth examination of the SIDDM formulation and implemented specific modifications in the generator parameterization and the reconstruction term within the objective. These adaptations enabled UFOGen to perform one-step sampling, while retaining training with several denoising steps.

Parameterization of the Generator In diffusion-GAN models, the generator should produce a sample of xt−1x_{t-1}. However, instead of directly outputting xt−1x_{t-1}, the generator of DDGAN and SIDDM is parameterized by pθ(xt−1∣xt)=q(xt−1∣xt,x0=Gθ(xt,t))p_{\theta}(x_{t-1}|x_{t})=q(x_{t-1}|x_{t},x_{0}=G_{\theta}(x_{t},t)). In other words, first x0x_{0} is predicted using the denoising generator Gθ(xt,t)G_{\theta}(x_{t},t), and then, xt−1x_{t-1} is sampled using the Gaussian posterior distribution q(xt−1∣xt,x0)q(x_{t-1}|x_{t},x_{0}) derived in . Note that this parameterization is mainly for practical purposes, as discussed in , and alternative parameterization would not break the model formulation.

We propose another plausible parameterization for the generator: pθ(xt−1)=q(xt−1∣x0=Gθ(xt,t))p_{\theta}(x_{t-1})=q(x_{t-1}|x_{0}=G_{\theta}(x_{t},t)). The generator still predicts x0x_{0}, but we sample xt−1x_{t-1} from the forward diffusion process q(xt−1∣x0)q(x_{t-1}|x_{0}) instead of the posterior. As we will show later, this design allows distribution matching at x0x_{0}, paving the path to one-step sampling.

The second term in the objective minimizes the KL divergence between pθ(xt∣xt−1′)p_{\theta}(x_{t}|x_{t-1}^{\prime}) and q(xt∣xt−1)q(x_{t}|x_{t-1}), which, as derived in Appendix A.1, can be simplified to the following reconstruction term:

Based on above analysis on xt−1′x_{t-1}^{\prime} and xt−1x_{t-1}, it is easy to see that minimizing this reconstruction loss will essentially matches x0x_{0} and x0′x_{0}^{\prime} as well (a straightforward derivation is provided in Appendix A.2.2).

Per our analysis, both terms in the SIDDM objective in Equation 3 implicitly matches the distribution at x0x_{0}, which suggests that one-step sampling is possible. However, empirically we observe that one-step sampling from SIDDM does not work well even on 2-D toy dataset (See Figure 2). We conjecture that this is due to the variance introduced in the additive Gaussian noise when sampling xt−1x_{t-1} with x0x_{0}. To reduce the variance, we propose to replace the reconstruction term in Equation 5 with the reconstruction at clean sample ∣∣x0−x0′∣∣2||x_{0}-x_{0}^{\prime}||^{2}, so that the matching at x0x_{0} becomes explicit. We observe that with this change, we can obtain samples in one step, as shown in Figure 2.

Training and Sampling of UFOGen To put things together, we present the complete training objective and strategy for the UFOGen model. UFOGen is trained with the following objective:

where γt\gamma_{t} is a time-dependent coefficient. The objective consists of an adversarial loss to match noisy samples at time step t−1t-1, and a reconstruction loss at time step . Note that the reconstruction term is essentially the training objective of diffusion models , and therefore the training of UFOGen model can also be interpreted as training a diffusion model with adversarial refinement. The training scheme of UFOGen is presented in Algorithm 1.

Despite the straightforward nature of the modifications to the training objective, these enhancements have yielded impressive outcomes, particularly evident in the context of one-step sampling, where we simply sample xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\textbf{I}) and produce sample x0′=Gθ(xT)x_{0}^{\prime}=G_{\theta}(x_{T}).

2 Leverage Pre-trained Diffusion Models

Our objective is developing an ultra-fast text-to-image model. However, the transition from an effective UFOGen recipe to web-scale data presents considerable challenges. Training diffusion-GAN hybrid models for text-to-image generation encounters several intricacies. Notably, the discriminator must make judgments based on both texture and semantics, which govern text-image alignment. This challenge is particularly pronounced during the initial stage of training. Moreover, the cost of training text-to-image models can be extremely high, particularly in the case of GAN-based models, where the discriminator introduces additional parameters. Purely GAN-based text-to-image models confront similar complexities, resulting in highly intricate and expensive training.

To surmount the challenges of scaling-up diffusion-GAN hybrid models, we propose the utilization of pre-trained text-to-image diffusion models, notably the Stable Diffusion model . Specifically, our UFOGen model is designed to employ a consistent UNet structure for both its generator and discriminator. This design enables seamless initialization with the pre-trained Stable Diffusion model. We posit that the internal features within the Stable Diffusion model contain rich information of the intricate interplay between textual and visual data. This initialization strategy significantly streamlines the training of UFOGen. Upon initializing UFOGen’s generator and discriminator with the Stable Diffusion model, we observe stable training dynamics and remarkably fast convergence. The complete training strategy of UFOGen is illustrated in Figure 3.

Experiments

In this section, we evaluate our proposed UFOGen model for the text-to-image synthesis problem. In Section 5.1, we start with briefly introducing our experimental setup, followed by comprehensive evaluations of UFOGen model on the text-to-image task, both quantitatively and qualitatively. We conduct ablation studies in Section 5.2, highlighting the effectiveness of our modifications introduced in Section 4. In Section 5.3, we present qualitative results for downstream applications of UFOGen.

Configuration for Training and Evaluation For experiments on text-to-image generation, we follow the scheme proposed in Section 4.2 to initialize both the generator and discriminator with the pre-trained Stable Diffusion 1.5https://huggingface.co/runwayml/stable-diffusion-v1-5 model . We train our model on the LAION-Aesthetics-6+ subset of LAION-5B . More training details are provided in Appendix A.3. For evaluation, we adopt the common practice that uses zero-shot FID on MS-COCO , and CLIP score with ViT-g/14 backbone .

Main Results To kick-start our evaluation, we perform a comparative analysis in Table 1, bench-marking UFOGen against other few-step sampling models that share the same Stable Diffusion backbone. Our baselines include Progressive Distillation and its variant , which are previously the state-of-the-art for few-step sampling of SD, as well as the concurrent work of InstaFlow . Latent Consistency Model (LCM) is excluded, as the metric is not provided in their paper. Analysis of the results presented in Table 1 reveals the superior performance of our single-step UFOGen when compared to Progressive Distillation across one, two, or four sampling steps, as well as the CFG-Aware distillation in eight steps. Furthermore, our method demonstrates advantages in terms of both FID and CLIP scores over the single-step competitor, InstaFlow-0.9B, which share the same network structure of SD with us. Impressively, our approach remains highly competitive even when compared to InstaFlow-1.7B with stacked UNet structures, which effectively doubles the parameter count.

The results depicted in Table 1 may suggest that InstaFlow remains a strong contender in one-step generation alongside UFOGen. However, we argue that relying solely on the MS-COCO zero-shot FID score for evaluating visual quality might not be the most reliable metric, a concern highlighted in prior research such as and discussed by . Consequently, we believe that qualitative assessments can provide more comprehensive insights. We present qualitative comparisons involving InstaFlow and LCMInstaFlow (https://huggingface.co/spaces/XCLiu/InstaFlow) and LCM (https://huggingface.co/spaces/SimianLuo/Latent_Consistency_Model) in Table 2. The comparisons allow for a clear-cut conclusion: UFOGen’s one-step image generation surpasses InstaFlow by a substantial margin in terms of image quality. Notably, UFOGen also demonstrates significant advantages when contrasted with the 2-step LCM, as showed by the evident blurriness present in LCM’s samples. Furthermore, even when compared to the samples generated by the 4-step LCM, our generated images exhibit distinct characteristics, including sharper textures and finer details. We do not present results of single-step LCM, as we observe that it fail to generate any textures (see Appendix A.5.1). Additional examples of the comparison are provided in Appendix A.5.2, where we display multiple images generated by each model for different prompts. We provide additional qualitative samples of UFOGen in Appendix A.6.

For completeness, we extend our comparison to encompass a diverse array of text-to-image generative models in Table 3. While the results in Table 3 are not directly comparable due to substantial variations in model architecture, parameter count, and training data, it is noteworthy that UFOGen is a competitive contender among the contemporary landscape of text-to-image models, offering the advantage of remarkable speed over auto-regressive or diffusion models, thanks to its inherent one-step generation capability.

Based on both quantitative and qualitative assessments, we assert that UFOGen stands as a powerful text-to-image generative model, capable of producing sharp and visually appealing images that align well with the provided text conditioning, all in a single step. Our evaluation underscores its capacity to produce superior sample quality when contrasted with competing diffusion-based methods designed for a few-step generation process.

2 Ablation Studies

Ablation studies have been conducted to offer deeper insights into the effectiveness of our training strategies. As outlined in Table 4, we compare the training of diffusion-GAN hybrid models using the SIDDM objective against the proposed UFOGen objective in Section 4.1. The results validate our assertions, demonstrating that the modifications in the UFOGen objective facilitate one-step sampling. We additionally provide qualitative samples, and an supplementary ablation study on the denoising step size during training in Appendix A.4.

3 Applications

A promising aspect of text-to-image diffusion models is their versatility as foundational components for various applications, whether fine-tuned or utilized as is. In this section, we showcase UFOGen’s ability to extend beyond text-to-image generation, while benefiting from its unique advantage of single-step generation. Specifically, we explore two applications of UFOGen: image-to-image generation and controllable generation .

Table 5 showcases UFOGen’s image-to-image generation outcomes. Following SDEdit , we introduce a suitable amount of noise to the input data, and let UFOGen to execute single-step generation based on the given prompt. Our observations affirm that UFOGen adeptly produces samples that adhere to the specified conditions of both the prompt and the input image.

To facilitate controllable generation, we conduct fine-tuning of UFOGen by incorporating an additional adapter network, akin to the approach outlined in . This adapter network takes control signals as input to guide the generation process. In our exploration, we employ two types of control signals: depth maps and canny edges. The results are presented in Table 6. Post fine-tuning, UFOGen exhibits the ability to generate high-quality samples that align with both the provided prompt and control signal.

Our results highlight UFOGen can work on diverse generation tasks in a single step, a distinctive feature that, to the best of our knowledge, sets our model apart. Unlike GAN-based text-to-image models , which lack the ability to handle zero-shot image-to-image generation tasks as they do not generate samples through denoising, UFOGen excels in this context. Moreover, our model succeeds in controllable generation, a domain that earlier GAN-based models have not explored due to the complexities of fine-tuning and adding supplementary modules to the StyleGAN architecture. Consequently, the flexibility of our model in addressing various downstream tasks positions it uniquely among one-step text-to-image models. Additional results of the applications are provided in Appendix A.7.

Conclusions

In this paper, we present UFOGen, a groundbreaking advancement in text-to-image synthesis that effectively addresses the enduring challenge of inference efficiency. Our innovative hybrid approach, combining diffusion models with a GAN objective, propels UFOGen to achieve ultra-fast, one-step generation of high-quality images conditioned on textual descriptions. The comprehensive evaluations consistently affirm UFOGen’s superiority over existing accelerated diffusion-based methods. Its distinct capability for one-step text-to-image synthesis and proficiency in downstream tasks underscore its versatility and mark it as a standout in the field. As a pioneer in enabling ultra-fast text-to-image synthesis, UFOGen paves the way for a transformative shift in the generative models landscape. The potential impact of UFOGen extends beyond academic discourse, promising to revolutionize the practical landscape of rapid and high-quality image generation.

References

Appendix A Appendices

In this section, we provide a derivation of obtaining a reconstruction objective from the KL term in Equation 3:

Note that q(xt∣xt−1)=N(1−βtxt−1,βtI)q(x_{t}|x_{t-1})=\mathcal{N}(\sqrt{1-\beta_{t}}x_{t-1},\beta_{t}\mathbf{I}) is a Gaussian distribution defined by the forward diffusion. For pθ(xt∣xt−1′)p_{\theta}(x_{t}|x_{t-1}^{\prime}), although the distribution on pθ(xt−1′)p_{\theta}(x_{t-1}^{\prime}) is quite complicated (because this depends on the generator model), given a specific xt−1′x_{t-1}^{\prime}, it follows the same distribution of forward diffusion: pθ(xt∣xt−1′)=N(1−βtxt−1′,βtI)p_{\theta}(x_{t}|x_{t-1}^{\prime})=\mathcal{N}(\sqrt{1-\beta_{t}}x_{t-1}^{\prime},\beta_{t}\mathbf{I}). Therefore, Equation 7 is the KL divergence between two Gaussian distributions, which we can computed in closed form. For two multivariate Gaussian distributions with means μ1,μ2\mu_{1},\mu_{2} and covariance Σ1,Σ2\Sigma_{1},\Sigma_{2}, the KL divergence can be expressed as

We can easily plug-in the means and variances for q(xt∣xt−1)q(x_{t}|x_{t-1}) and pθ(xt′∣xt−1′)p_{\theta}(x_{t}^{\prime}|x_{t-1}^{\prime}) into the expression. Note that Σ1=Σ2=βtI\Sigma_{1}=\Sigma_{2}=\beta_{t}\mathbf{I}, so the expression can be simplified to

where CC is a constant. Therefore, with the outer expectation over q(xt)q(x_{t}) in Equation 3, minimizing the KL objective is equivalent to minimizing a weighted reconstruction loss between xt−1′x_{t-1}^{\prime} and xt−1x_{t-1}, where xt−1x_{t-1} is obtained by sampling x0∼q(x0)x_{0}\sim q(x_{0}) and xt−1∼q(xt−1∣x0)x_{t-1}\sim q(x_{t-1}|x_{0}); xt−1′x_{t-1}^{\prime} is obtained from generating an x0′x_{0}^{\prime} from the generator followed by sampling xt−1′∼q(xt−1′∣x0′)x_{t-1}^{\prime}\sim q(x_{t-1}^{\prime}|x_{0}^{\prime}).

Note that in , the authors did not leverage the Gaussian distribution’s KL-divergence property. Instead, they decomposed the KL-divergence into an entropy component and a cross-entropy component, subsequently simplifying each aspect by empirically estimating the expectation. This simplification effectively converges to the same objective as expressed in Equation 8, albeit with an appended term associated with entropy. The authors of introduced an auxiliary parametric distribution for entropy estimation, which led to an adversarial training objective. Nevertheless, our analysis suggests that this additional term is dispensable, and we have not encountered any practical challenges when omitting it.

In this section, we offer a detailed explanation of why training the model with the objective presented in Equation 3 effectively results in matching x0x_{0} and x0′x_{0}^{\prime}. The rationale is intuitive: xt−1x_{t-1} and xt−1′x_{t-1}^{\prime} are both derived from their respective base images, x0x_{0} and x0′x_{0}^{\prime}, through independent Gaussian noise corruptions. As a result, when we enforce the alignment of distributions between xt−1x_{t-1} and xt−1′x_{t-1}^{\prime}, this implicitly encourages a matching of the distributions between x0x_{0} and x0′x_{0}^{\prime} as well. To provide a more rigorous and formal analysis, we proceed as follows.

Similarly, p(θ)(xt−1)=pθ(x0)∗k(x)p_{(}\theta)(x_{t-1})=p_{\theta}(x_{0})*k(x), where pθ(x0)p_{\theta}(x_{0}) is the implicit distribution defined by the generator GθG_{\theta}.

In the following lemma, we show that for a probability divergence DD, if p(x)p(x) and q(x)q(x) are convoluted with the same kernel k(x)k(x), then minimizing DD on the distributions after the convolution is equivalent to matching the original distributions p(x)p(x) and q(x)q(x).

Proof: The probability density of the summation between two variables is the convolution between their probability densities. Thus, we have:

where F\mathcal{F} denotes the Fourier Transform, and we utilize the invertibility of the Fourier Transform for the above derivation.

Thus, from Lemma 1, we can get q(x0)=pθ(x0)q(x_{0})=p_{\theta}(x_{0}) almost everywhere when JSD(q(xt−1)∣∣pθ(xt−1))=0\text{JSD}(q(x_{t-1})||p_{\theta}(x_{t-1}))=0. Notably, while training with the adversarial objective on xt−1x_{t-1} inherently aligns the distributions of q(x0)q(x_{0}) and ptheta(x0′)p_{theta}(x_{0}^{\prime}),it is crucial to acknowledge that we cannot directly employ GAN training on x0x_{0}. This is because the additive Gaussian noise, which serves to smooth the distributions, rendering GAN training more stable. Indeed, training GANs on smooth distributions is one of the essential components of all diffusion-GAN hybrid models, as highlighted in .

A.2.2 KL term

Here we show that minimizing the reconstruction loss in Equation 8 over the expectation of q(xt)q(x_{t}) as in Equation 3 is equivalent to minimizing the reconstruction loss between x0x_{0} and x0′x_{0}^{\prime}. According to the sampling scheme of xt−1x_{t-1} and xt−1′x_{t-1}^{\prime}, we have

Since the forward diffusion q(xt−1∣x0)q(x_{t-1}|x_{0}) has the Gaussian form

and similar form holds for pθ(xt−1′∣x0′)p_{\theta}(x_{t-1}^{\prime}|x_{0}^{\prime}), we can rewrite the expectation in Equation 10 over the distribution of simple Gaussian distribution p(ϵ)=N(ϵ;0,I)p(\epsilon)=\mathcal{N}\left(\epsilon;0,\mathbf{I}\right):

where xt−1′=αˉt−1x0′+(1−αˉt−1)ϵ′x_{t-1}^{\prime}=\sqrt{\bar{\alpha}_{t-1}}\mathbf{x}_{0}^{\prime}+\left(1-\bar{\alpha}_{t-1}\right)\epsilon^{\prime} and xt−1=αˉt−1x0+(1−αˉt−1)ϵx_{t-1}=\sqrt{\bar{\alpha}_{t-1}}\mathbf{x}_{0}+\left(1-\bar{\alpha}_{t-1}\right)\epsilon are obtained by i.i.d. samples ϵ′,ϵ\epsilon^{\prime},\epsilon from p(ϵ)p(\epsilon). Plug in the expressions to Equation 12, we obtain

where CC is a constant independent of the model. Therefore, we claim the equivalence of the reconstruction objective and the matching between x0x_{0} and x0′x_{0}^{\prime}.

However, it’s essential to emphasize that the matching between x0x_{0} and x0′x_{0}^{\prime} is performed with an expectation over Gaussian noises. In practical terms, this approach can introduce significant variance during the sampling of xt−1x_{t-1} and xt−1′x_{t-1}^{\prime}. This variance, in turn, may result in a less robust learning signal when it comes to aligning the distributions at clean data.As detailed in Section 4.1, we propose a refinement to address this issue. Specifically, we advocate for the direct enforcement of reconstruction between x0x_{0} and x0′x_{0}^{\prime}. This modification introduces explicit distribution matching at the level of clean data, enhancing the model’s robustness and effectiveness.

A.3 Experimental Details

For all the experiments, we initialize the parameters of both the generator and discriminator with the pre-trained Stable Diffusion 1.5 checkpoint. In consequence, we follow SD 1.5 to use the same VAE for image encoding/decoding and the frozen text encoder of CLIP ViT-L/14 for text conditioning. Note that both the generator and discriminator operates on latent space. In other words, the generator generates the latent variables and the discriminator distinguishes the fake and true (noisy) latent variables.

One important hyper-parameter is the denoising step size during training, which is the gap between t−1t-1 and tt. Note that in Section 4.1, we mentioned that the model is trained with multiple denoising steps, while it enables one-step inference. Throughout the experiments, we train the models using denoising step size 250, given the 10001000-step discrete time scheduler of SD. Specifically, during training, we sample tt randomly from 11 to 10001000, and the time step for t−1t-1 is max(0,t−250)max(0,t-250). We conduct ablation studies on this hyper-parameter in Section A.4.

Another important hyper-parameter is λKL\lambda_{KL}, the weighting coefficient for reconstruction term in the objective in Equation 4.1. We set λKL=1.0\lambda_{KL}=1.0 throughout the experiments. We found the results insensitive to slight variations of this coefficient.

We train our models on the LAION Aesthetic 6+ dataset. For the generator, we use AdamW optimizer with β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999; for the discriminator, we use AdamW optimizer with β1=0.0\beta_{1}=0.0 and β2=0.999\beta_{2}=0.999. We adopt learning rate warm-up in the first 1000 steps, with peak learning rate 1e−41e-4 for both the discriminator and the generator. For training the generator, we apply gradient norm clipping with value 1.01.0 for generator only. We use batch size 1024. For the generator, we apply EMA with coefficient 0.999. We observe quick convergence, typically in <50<50k steps.

A.4 Additional Results of Ablation Studies

In this section, we provide additional results for ablation studies, which are briefly covered in the main text due to the constraints of space. In Appendix A.4.1, we provide qualitative results corresponds to the ablation study conducted in Section 5.2. In Appendix A.4.2, we conduct an additional ablation experiment on the denoising step size during training.

We provide qualitative examples to contrast between the single-step sample generated by SIDDM and our proposed UFOGen. Results are shown in Table 7 and 8. We observe that when sampling from SIDDM in only one-step, the samples are blurry and over-smoothed, while UFOGen can produce sharp samples in single step. The observation strongly supports the effectiveness of our introduced modifications to the training objective.

A.4.2 Ablation on Denoising Step-size

One important hyper-parameter of training UFOGen is the denoising step size, which is the gap between tt and t−1t-1 during training. Note that although UFOGen can produce samples in one step, the training requires a meaningful denoising step size to compute the adversarial loss on noisy observations. Our model is based on Stable Diffusion, which adopts a discrete time scheduler with 1000 steps. Previous diffusion GAN hybrid models divides the denoising process into 2 to 4 steps. We explore denoising step size 125, 250, 500 and 1000, which corresponds to divide the denoising process to 8, 4, 2, and 1 steps. Note that during training, we sample tt uniformly in [1,1000)[1,1000), and when the sampled tt is smaller than the denoising step size, we set t−1t-1 to be . In other words, a denoising step size 10001000 corresponds to always setting t−1=0t-1=0 and hence the adversarial loss is computed on clean data x0x_{0}.

Quantitative results of the ablation study is presented in Table 9. We observe that a denoising step size 1000 fails, suggesting that training with the adversarial loss on noisy data is critical for stabilizing the diffusion-GAN training. This observation was made on earlier work as well. We also observe that denoising step size 250 is the sweet spot, which is also aligned with the empirical observations of . We conjecture that the reason for the performance degrade when reducing the denoising step size is that the discriminator does not have enough capacity to discriminate on many distinct noise levels.

A.5 Additional Results for Qualitative Comparisons

Consistency models try to learn the consistency mapping that maps every point on the PF-ODE trajectory to its boundary value, i.e., x0x_{0} , and therefore ideally consistency models should generate samples in one single step. However, in practice, due to the complexity of the ODE trajectory, one-step generation for consistency models is not feasible, and some iterative refinements are necessary. Notably, Latent consistency models (LCM) distilled the Stable Diffusion model into a consistency model, and we observe that single-step sampling fail to generate reasonable textures. We demonstrate the single-step samples from LCM in figure 4. Due to LCM’s ineffectiveness of single-step sampling, we only qualitatively compare our model to 2-step and 4-step LCM.

A.5.2 Extended Results of Table 2

In consideration of space constraints in the main text, our initial qualitative comparison of UFOGen with competing methods for few-step generation in Table 2 employs a single image per prompt. It is essential to note that this approach introduces some variability due to the inherent randomness in image generation. To provide a more comprehensive and objective evaluation, we extend our comparison in this section by presenting four images generated by each method for every prompt. This expanded set of prompts includes those featured in Table 2, along with additional prompts. The results of this in-depth comparison are illustrated across Table 10 to 17, consistently highlighting UFOGen’s advantageous performance in generating sharp and visually appealing images within an ultra-low number of steps when compared to competing methods.

Concurrent to our paper submission, the authors of LCM released updated LCM models trained with more resources. The models are claimed to be stronger than the initially released LCM model, which is used in our qualitative evaluation. For fairness in the comparison, we obtain some qualitative samples of the updated LCM model that shares the SD 1.5 backbone with ushttps://huggingface.co/latent-consistency/lcm-lora-sdv1-5, and show them in Table 18 and 19. We observe that while the new LCM model generates better samples than initial LCM model does, our single-step UFOGen is still highly competitive against 4-step LCM and significantly better than 2-step LCM.

A.6 Additional Qualitative Samples from UFOGen

In this section, we present supplementary samples generated by UFOGen models, showcasing the diversity of results in Table 20, 21 and 22. Through an examination of these additional samples, we deduce that UFOGen exhibits the ability to generate high-quality and diverse outputs that align coherently with prompts spanning various styles (such as painting, photo-realistic, anime) and contents (including objects, landscapes, animals, humans, etc.). Notably, our model demonstrates a promising capability to produce visually compelling images with remarkable quality within just a single sampling step.

In Table 23, we present some failure cases of UFOGen. We observe that UFOGen suffers from missing objects, attribute leakage and counting, which are common issues of SD based models, as discussed in [chefer2023attend, feng2022training].

A.7 Additional Results of UFOGen’s Applications

In this section, we provide extended results of UFOGen’s applications, including the image-to-image generation in Figure 5 and controllable generation in Figure 6.