Universal Guidance for Diffusion Models
Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Soumyadip Sengupta, Micah Goldblum, Jonas Geiping, Tom Goldstein
Introduction
Diffusion models are powerful tools for creating digital art and graphics. Much of their success stems from our ability to carefully control their outputs, customizing results for each user’s individual needs. Most models today are controlled through conditioning. With conditioning, the diffusion model is built from the ground up to accept a particular modality of input from the user, be it descriptive text, segmentation maps, class labels, etc. While conditioning is a powerful tool, it results in models that are handcuffed to a single conditioning modality. If another modality is required, a new model needs to be trained, often from scratch. Unfortunately, the high cost of training makes this prohibitive for most users.
A more flexible approach to controlling model outputs is to use guidance. In this approach, the diffusion model acts as a generic image generator, and is not required to understand a user’s instructions. The user pairs this model with a guidance function that measures whether some criterion has been met. For example, one could guide the model to minimize the CLIP score between the generated image and a text description of the user’s choice. During each iteration of image creation, the iterates are nudged down the gradient of the guidance function, causing the final generated image to satisfy the user’s criterion.
In this paper, we study guidance methods that enable any off-the-shelf model or loss function to be used as guidance for diffusion. Because guidance functions can be used without re-training or modification, this form of guidance is universal in that it enables a diffusion model to be adapted for nearly any purpose.
From a user perspective, guidance is superior to conditioning, as a single diffusion network is treated like a foundational model that provides universal coverage across many use cases, both commonplace and bespoke. Unfortunately, it is widely believed that this approach is infeasible. While early diffusion models relied on classifier guidance (Dhariwal & Nichol, 2021), the community quickly turned to classifier-free schemes (Ho & Salimans, 2022) that require a model to be trained from scratch on class labels with a particular frozen ontology that cannot be changed (Nichol et al., 2021; Rombach et al., 2022; Bansal et al., 2022).
The difficulty of using guidance stems from the domain shift between the noisy images used by the diffusion sampling process and the clean images on which the guidance models are trained. When this gap is closed, guidance can be performed successfully. For example, Nichol et al. (2021) successfully use a CLIP model as guidance, but only after re-training CLIP from scratch using noisy inputs. Noisy retraining closes the domain gap, but at a very high financial and engineering cost. To avoid the additional cost, we study methods for closing this gap by changing the sampling scheme, rather than the model.
To this end, our contributions are summarized as follows:
We propose an algorithm that enables universal guidance for diffusion models. Our proposed sampler evaluates the guidance models only on denoised images, rather than noisy latent states. By doing so, we close the domain gap that has plagued standard guidance methods. This strategy provides the end-user with the flexibility to work with a wide range of guidance modalities and even multiple modalities simultaneously. The underlying diffusion model remains fixed and no fine-tuning of any kind is necessary.
We demonstrate the effectiveness of our approach for a variety of different constraints such as classifier labels, human identities, segmentation maps, annotations from object detectors, and constraints arising from inverse linear problems.
Background
We first briefly review the recent literature on the core framework behind diffusion models. Then, we define the problem setting of controlled image generation and discuss previous related works.
Diffusion models are strong generative models that proved powerful even when first introduced for image generation (Song & Ermon, 2019; Ho et al., 2020). The approach has been successfully extended to a number of domains, such as audio and text generation (Kong et al., 2020; Huang et al., 2022; Austin et al., 2021; Li et al., 2022).
We introduce (unconditional) diffusion formally, as it is helpful in describing the nuances of different types of models. A diffusion model is defined as a combination of a -step forward process and a -step reverse process. Conceptually, the forward process gradually adds Gaussian noise of different magnitudes to a clean data point , while the reverse process attempts to gradually denoise a noisy input in hopes of recovering a clean data point. More concretely, given an array of scalars representing noise scales and an initial, clean data point , applying steps of the forward process to yields a noisy data point
A diffusion model is a learned denoising network . It is trained so that for any pair and any sample of ,
The reverse process takes the form with various detail definitions, where is generally parameterized as a Gaussian distribution. Different works also studied different approximations of the unknown used to perform sampling. For example, denoising diffusion implicit model (DDIM) (Song et al., 2021a) first computed a predicted clean data point
and sample from by replacing unknown with On the other hand, while the details of individual sampling methods vary, all sampling methods produce based on current sample , current time step and a predicted noise To ease the notation burden, we define a function as an abstraction of the sampling method, where
2 Controlled Image Generation
Prior work that studied controlled generative diffusion mainly falls into two categories. We refer to the first category as conditional image generation, and the second category as guided image generation. Next, we discuss the characteristics of each category and better situate our work among existing methods.
Methods from this category require training new diffusion models that accept the prompt as an additional input (Ho & Salimans, 2022; Bansal et al., 2022; Nichol et al., 2021; Whang et al., 2022; Wang et al., 2022a). For example, Ho & Salimans (2022) proposed classifier-free guidance using class labels as prompts, and trained a diffusion model by linear interpolation between unconditional and conditional outputs of the denoising networks. Bansal et al. (2022) studied the case where the guidance function is a known linear degradation operator, and trained a conditional model to solve linear inverse problems. Nichol et al. (2021) further extended classifier-free guidance to text-conditional image generation with descriptive phrases as prompts, and trained a diffusion model to enforce the similarity between the CLIP (Radford et al., 2021) representations of the generated images and the text prompts. These methods are successful across different types of constraints, however the requirement to retrain the diffusion model makes them computationally intensive.
Guided Image Generation.
Works in this category employed a frozen pre-trained diffusion model as a foundation model, but modify the sampling method to guide the image generation with feedback from the guidance function. Our method falls into this category. Prior work that studied guided image generation did so with a variety of restrictions and external guidance functions (Dhariwal & Nichol, 2021; Kawar et al., 2022; Wang et al., 2022b; Chung et al., 2022a; Lugmayr et al., 2022; Chung et al., 2022b; Graikos et al., 2022). For example, Dhariwal & Nichol (2021) proposed classifier guidance, where they trained a classifier on images of different noise scales as the guidance function , and included gradients of the classifier during the sampling process. However, a classifier for noisy images is domain-specific and generally not readily available – an issue our method circumvents. Wang et al. (2022b) assumed the external guidance functions to be linear operators, and generated the component of images residing in the null space of linear operators with the foundation model. Unfortunately, extending that method to handle non-linear guidance functions is non-trivial. Chung et al. (2022a) studied general guidance functions, and modified the sampling process with the gradient of guidance function calculated on the expected denoised images. Nevertheless, the authors only presented results with simpler non-linear guidance functions such as non-linear blurring.
In this work, we study universal guidance algorithms for guided image generation with diffusion models using any off-the-shelf guidance functions , such as object detection or segmentation networks.
Universal Guidance
We propose a guidance algorithm that augments the image sampling method of a diffusion model to include guidance from an off-the-shelf auxiliary network. Our algorithm is motivated by an empirical observation that the reconstructed clean image obtained by Equation 3, while naturally imperfect, is still appropriate for a generic guidance function to provide informative feedback to guide the image generation. In Section 3.1, we motivate our forward universal guidance by extending classifier guidance (Dhariwal & Nichol, 2021) to leverage this observation and handle generic guidance functions. In Section 3.2, we propose a supplementary backward universal guidance to help enforce the generated image to satisfy the constraint based on the guidance function . In Section 3.3, we discuss a simple yet helpful self-recurrence trick to empirically improve the fidelity of generated images.
To address the issue, we leverage the fact that predicts the noise added to the data point, and we can therefore obtain a predicted clean image by Equation 3. We propose to instead calculate the guidance based on the predicted clean data point as
where controls the guidance strength for each sampling step and
as in Equation 3. We term Equation 6 forward universal guidance, or forward guidance in short. In practice, applying forward guidance effectively brings the generated image closer to the prompt while keeping the generation trajectory in the data manifold. We note that a related approach is also studied in (Chung et al., 2022a), where the guidance step is computed based on . The approach drew inspiration from the score-based generative framework (Song et al., 2021b), but resulted in a different update method.
2 Backward Universal Guidance
As will be shown in Section 4.2, we observe that forward guidance sometimes over-prioritizes maintaining the “realness” of the image, resulting in an unsatisfactory match with the given prompt. Simply increasing the guidance strength is suboptimal, as this often results in instability as the image moves off the manifold faster than the denoiser can correct it.
Comparing to forward guidance, backward guidance (as Equation 9) produces an optimized direction for the generated image to match the given prompt, and hence prioritizes enforcing the constraint. Furthermore, calculation of a gradient step for Equation 7 is computationally cheaper than forward guidance (Equation 6), and we can therefore afford to solve Equation 7 with multiple gradient steps, further improving the match with the given prompt.
We note that the names “forward” and “backward” are used analogously to the forward and backward Euler methods.
3 Per-step Self-recurrence
Unfortunately, when we apply our universal guidance to standard generation pipelines, we often find images with artifacts and strange behaviors that clearly separate them from natural images. Similar observations have been made in (Lugmayr et al., 2022; Wang et al., 2022b), where linear guidance functions are studied. Our attempts to prioritize realness by decreasing proved ineffective; the sweet spot that both ensures the realness and guidance constraint satisfaction doesn’t always exist, especially for complex guidance functions. We conjecture that the guidance direction produced by our universal method is not always related to the realness of the images when the guidance function creates too much information loss, causing the image to stray from the natural image sampling trajectory.
Inspired by (Lugmayr et al., 2022; Wang et al., 2022b), we address the issue by applying per-step self-recurrence. More concretely, after is sampled, we re-inject random Gaussian noise to to obtain by
Equation 10 ensures to have proper noise scale for input at time step . We repeat the self-recurrence times before continuing the sampling for step . Intuitively, the self-recurrence allows exploration of different regions of the data manifold at the same noise scale, allowing more budget to find a solution that satisfies both guidance and image quality. Empirically, we find that our self-recurrence can keep the realness of the generated image with a proper guidance strength that ensures the match with the given prompt. We illustrate an example of how self-recurrence improves the harmony of generated images in Figure 2.
We summarize our universal guidance algorithm composed of forward universal guidance, backward universal guidance and per-step self-recurrence in Algorithm 1. For simplicity, the algorithm assumes only one guidance function, but can be easily adapted to handle multiple pair of . Additionally, the objectives of the forward and backward guidance do not have to be identical, allowing different ways to simultaneously utilize multiple guidance functions.
Experiments
In this section, we present results testing our proposed universal guidance algorithm against a wide variety of guidance functions. Specifically, we experiment with Stable Diffusion (Rombach et al., 2022), a diffusion model that is able to perform text-conditional generation by accepting text prompt as additional input, and experiment with a purely unconditional diffusion model trained on ImageNet (Deng et al., 2009), where we use pre-trained model provided by OpenAI (Dhariwal & Nichol, 2021). We note that Stable Diffusion, while being a text-conditional generative model, can also perform unconditional image generation by simply using an empty string for the text prompt. We first present the experiment on Stable Diffusion for different guidance functions in Section 4.1, and present the results on ImageNet diffusion model in Section 4.2.
In this section, we present the results of guided image generation using Stable Diffusion as the foundation model. The guidance functions we experiment with include the CLIP feature extractcor (Radford et al., 2021), a segmentation network, a face recognition network and an object detection network. For experiments on Stable Diffusion, we discover that applying forward guidance already produce high-quality images that match the given prompt, and hence set . To perform forward guidance on Stable Diffusion, we forward the predicted clean latent variable computed by Equation 3 through the image decoder of Stable Diffusion to obtain predicted clean images. We discuss the results and implementation details for each guidance function in its corresponding subsection.
CLIP (Radford et al., 2021) is a state-of-the-art text-to-image similarity model developed by OpenAI. To apply our algorithm to text-guided image generation, we use the image feature extractor of CLIP as the guidance function. We construct a loss function that calculates the negative cosine similarity between an image embedding and the CLIP text embedding produced by a given text prompt. We use and and use Stable Diffusion as an unconditional image generator.
We generate images guided by a number of text prompts. To further assess our universal guidance algorithm and compare guidance and conditioning, we also generate images using classical, text-conditional generation by Stable Diffusion with identical prompts as inputs, and summarize the results in Figure 3. The results in Figure 3 show that our algorithm can guide the generation to produce high-quality images that match the given text description, and are comparable with images generated by the specialized text-conditioning model.
Segmentation Map Guidance.
In our experiment, we combine segmentation maps that depict objects of different shapes with new text prompts. We use the text prompt as a fixed additional input to Stable Diffusion to perform text-conditional sampling, and guide the text-conditional generated images to match the given segmentation maps. Results are presented in Figure 4. From Figure 4, we see that the generated images show a clear separation between object and background that matches the given segmentation map nearly perfectly. The generated object and background also each match their descriptive text (i.e. dog breed and environment description). Furthermore, the generated images are overall highly realistic.
Face Recognition Guidance.
We explore different combinations of face guidance and text prompts. Similarly to the segmentation case, we use the text prompt as a fixed additional conditioning to Stable Diffusion and guide this text-conditional trajectory with our algorithm so that the face in the generated image looks similar to the face prompt. In Figure 5, we clearly see that the facial characteristics of a given face prompt are reproduced almost perfectly on the generated images. The descriptive text of either background, material, or style is also realized correctly and blends nicely with the generated faces.
Object Location Guidance
We again experiment with different combinations of text prompt and object location prompt, and similarly use the text prompt as a fixed conditioning to Stable Diffusion. Using our proposed guidance algorithm, we perform guided image generation that generates and matches the objects presented in the text prompt to the given object locations. The results are presented in Figure 6. We observe from Figure 6 that objects in the descriptive text all appear in the designated location with the appropriate size indicated by the given bounding boxes. Each location is filled with appropriate, high-quality generations that align with varied image content prompts, ranging from “beach” to “oil painting”.
Style Guidance
Finally, we conclude our experiments on Stable Diffusion by guiding the image generation based on a reference style given by a style image. To achieve so, we capture the reference style from the style image by the image feature extractor from CLIP, and use the resulting image embedding as prompts. The loss function calculates the negative cosine similarity between the embedding of generated images and the embedding of the style image. Similar to previous experiments, we control the content using text input as additional conditioning to the Stable Diffusion model.
We experiment with combinations of different style images and different text prompts, and present the results in Figure 7. From Figure 7, we can see that the generated images contain contents that match the given text prompts, while exhibiting style that matches the given style images. In this experiment we set and . Furthermore, in order to control the amount of content we set the scale , a parameter of Stable Diffusion that balances the text-conditional generation and unconditional generation, as 3.0, 3.0, and 4.0 respectively for each column.
2 Results for ImageNet Diffusion
In this section, we present results for guided image generation using an unconditional diffusion model trained on ImageNet. We experiment with CLIP guidance, object location guidance and a hybrid guided image generation task which we term segmentation-guided inpainting. We will discuss results and implementations of each guidance in its corresponding subsection.
Object Location Guidance.
We again experiment with different object location prompts using two configurations of our algorithm, namely (1) using only forward universal guidance and (2) using both forward and backward universal guidance. We observe from Figure 8 that applying both forward and backward guidance generates images that are realistic and the objects matches the prompt nicely. On the other hand, while images generated using only forward guidance remain realistic, they feature objects with mismatching categories and locations. The results demonstrate the effectiveness of our universal guidance algorithm, and also validate the necessity of our backward guidance.
Segmentation-Guided Inpainting.
Limitations
Generation using universal guidance is typically slower than standard conditional generation for several reasons. Empirically, multiple iterations of denoising are required at every noise level to generate high-quality images with complex guidance functions. However, the time complexity of our algorithm scales linearly with the number of recurrence steps , which slows down image generation when is large. Also, as demonstrated in the main paper, backward guidance is required in certain scenarios to help generate images that match the given constraint. Computing backward guidance requires performing minimization with a multi-step gradient descent inner loop. While proper choices of gradient-based optimization algorithms and learning rate schedules significantly speed up the convergence of minimization, the time it takes to compute backward guidance inevitably becomes longer when the guidance function is itself a very-large neural network. Finally, we note that, to get optimal results, sampling hyper-parameters must be chosen individually for each guidance network.
Conclusion
In this paper, we propose a universal guidance algorithm that is able to perform guided image generation with any off-the-shelf guidance function based on a fixed foundation diffusion model. Our algorithm only requires guidance and loss functions to be differentiable, and avoids any retraining to adapt either the guidance function or the foundation model to a specific type of prompt. We demonstrate promising results with our algorithm on complex guidance including segmentation, face recognition and object detection systems. Even multiple guidance functions can be combined and used in conjunction.
Acknowledgements
This work was made possible by the National Science Foundation (IIS-2212182), the AFOSR MURI Program, the Office of Naval Research (N000142112557), the ONR MURI program, IARPA WRIVA, and Capital One Bank.