More Control for Free! Image Synthesis with Semantic Diffusion Guidance

Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang, Arman Chopikyan, Yuxiao Hu, Humphrey Shi, Anna Rohrbach, Trevor Darrell

Introduction

Image synthesis has made great progress in recent years . In addition to the goal of generating high-quality photo-realistic images, fine-grained control over the generated images is also an important desideratum when assisting users with art creation and design.

Previous works have explored controllable image synthesis by adding different conditions, including language , attributes , scene graphs , and user sketches or scribbles . Specifically, text-to-image synthesis, as shown in Figure LABEL:fig:intro-(a), aims to generate images based on text instructions, by adding text embeddings as conditional information to the image generation network. However, most previous text-to-image synthesis methods require image-caption pairs for training, and cannot generalize to datasets without text annotations.

Besides text instructions, users may want to guide the image generation model with a reference image. E.g., a user might want to generate cat images which look similar to a given photo of a cat in terms of its appearance. This information cannot be easily described by language, but can be provided via a reference image, as shown in Figure LABEL:fig:intro-(b). Furthermore, a user may want to provide both language and image guidance. For example, a user might seek to generate “a woman with curly hair” that looks similar to a reference image of a woman with red hair, as shown in Figure LABEL:fig:intro-(c).

Current image-conditioned synthesis techniques either only transfer the “style” of a reference image to a target image or are restricted to the domains with well-defined structure, e.g. human or animal faces . They cannot generate diverse images with various poses, structures and layouts based on a single reference image.

We propose Semantic Diffusion Guidance (SDG), a unified framework for text-guided and image-guided synthesis that overcomes these limitations. Our model is based on denoising diffusion probabilistic models (DDPM) which generate an image from a noise map and iteratively removes noise to approach the data distribution of natural images.

We inject the semantic input by using a guidance function to guide the sampling process of an unconditional diffusion model. This enables more controllable generation in diffusion models and gives us a unified formulation for both language and image guidance. Specifically, our language guidance is based on the image-text matching score predicted by CLIP finetuned on noised images. As for the image guidance, depending on what information we seek in the image, we define two options: content and style guidance. The flexibility of the guidance module allows us to inject either language or image guidance alone or both at once into any unconditional diffusion model without the need for re-training. We propose a self-supervised scheme to finetune the CLIP image encoder without text annotations, from which we obtain the guidance model with minimal cost.

Our unified framework is flexible and allows fine-grained semantic control in image synthesis as shown in Figure LABEL:fig:intro. We show that our model can handle: (1) Text-guided image synthesis with fine-grained text queries on any dataset without language annotations; (2) Image-guided image synthesis with content or style control from the input image, which generates diverse images with different pose, structure, and layout; (3) Multi-modal guidance for image synthesis with both language and image input. Our flexible guidance network can be injected into off-the-shelf unconditional diffusion models, without the need for re-training the diffusion model. We further present a self-supervised efficient finetuning scheme for the CLIP guidance model which does not require textual annotations. We conduct experiments on FFHQ and LSUN datasets to validate the quality, diversity, and controllability of our generated images, and show various applications of our proposed Semantic Diffusion Guidance.

Related Work

Text-guided Synthesis Pioneered by GAN-INT-CLS and GAWWN , conditional generative adversarial networks (GANs) have been the dominant framework for text-based image synthesis. Various methods have been studied leading to significant improvements in editing quality and correctness . Recent work DALL-E shows promising results with transformers and discrete VAE by leveraging web-scale data. A concurrent work GLIDE adapts classifier-free guidance for large diffusion models and large-scale training for text-guided image synthesis. Despite great advancements, prior methods require paired image-text annotations which limits the application to certain datasets or requires large amount of data and computational resources for training. Our proposed framework is able to generate images on multiple domains given detailed text prompts, requiring neither image-text paired data from those domains nor large amount of compute to train the text-guided image synthesis model.

Image-guided Synthesis Image-guided synthesis aims to generate diverse images with the constraint that they all should resemble a given reference image in terms of content or style. Many style transfer works fall under this category where the content of the input image must be preserved while the style of the reference image is transferred , yet they struggle to generate diverse images. Some works study image synthesis guided by the content of the reference images. ILVR proposes a way to iteratively inject image guidance to a diffusion model, yet it exhibits limited structural diversity of the generated images. Instance-Conditioned GAN uses nearest neighbor images of a given reference for adversarial training to generate structurally diverse yet semantically relevant images. Yet, it requires training the GAN model with instance-conditioned techniques. Our approach demonstrates better controllability as different types of image guidance are proposed where users can decide how much semantic, structural, or style information to preserve by using different types and scales of guidance, while not requiring to re-train the unconditional diffusion model.

Diffusion Models Diffusion models are a new type of generative models consisting of a forward process (signal to noise) and a reverse process (noise to signal). The denoising diffusion probabilistic model (DDPM) is a latent variable model where a denoising autoencoder gradually transforms Gaussian noise into signal. Score-based generative model trains a neural network to predict the score function which is used to draw samples via Langevin Dynamics. Diffusion models have demonstrated comparable or superior image quality compared to GANs while exhibiting better mode coverage and training stability. They have also been explored for conditional generation such as class-conditional generation, image-guided synthesis, and super-resolution . Concurrent work explored text-guided image editing with diffusion models. Dhariwal et al. proposed classifier guidance for class-conditional image synthesis with diffusion models. Based on the guidance algorithm proposed in , we further explore whether diffusion models can be semantically guided by text or image, or both to synthesize images.

CLIP-guided Generation and Manipulation CLIP is a powerful vision-language joint embedding model trained on large-scale images and texts. Its representations have been shown to be robust and general enough to perform zero-shot classification and various vision-language tasks on diverse datasets. StyleCLIP and StyleGAN-NADA have demonstrated that CLIP enables text-guided image manipulation and domain adaptation of image generation without domain-specific image-text pairs. DiffusionCLIP uses CLIP for language-based image editing, and Blended Diffusion explores mask-guided image editing with CLIP. However, the application to image synthesis has not been explored. Our work investigates text and/or image guided synthesis using CLIP and unconditional DDPM.

Semantic Diffusion Guidance

We propose Semantic Diffusion Guidance (SDG), a unified framework that incorporates different forms of guidance into a pretrained unconditional diffusion model. SDG can leverage language guidance, image guidance, and both, enabling controllable image synthesis. The guidance module can be injected into any off-the-shelf unconditional diffusion model without re-training or finetuning it. We only need to finetune the guidance network, which is a CLIP model in our implementation, on the images with different levels of noise. We propose a self-supervised finetuning scheme, which is efficient and does not require paired language data to finetune the CLIP image encoder.

In Section 3.1, we review the preliminaries on diffusion models, and introduce our approach for injecting guidance for controllable image synthesis. In Section 3.2, we describe the language guidance which enables the unconditional diffusion model to perform text-to-image synthesis. In Section 3.3, we propose two types of image guidance, which take the content and style information from the reference image as the guidance signal, respectively. In Section 3.5, we explain how we finetune the CLIP network without requiring text annotations in the target domain.

Diffusion models define a Markov chain where random noise is gradually added to the data, known as the forward process. Formally, given a data point sampled from a real-data distribution x0∼q(x)x_{0}\sim q(x), the forward process sequentially adds Gaussian noise to the sample over TT timesteps:

where {β}t=1:T\{\beta\}_{t=1:T} denotes a constant or learned variance schedule that controls the noise step size. A property of the forward process is that we can sample xtx_{t} from x0x_{0} in a closed form:

where αt=1−βt\alpha_{t}=1-\beta_{t} and αt‾=∏s=1tαs\overline{\alpha_{t}}=\prod_{s=1}^{t}\alpha_{s}.

Generative modeling is done by learning the backward process where the forward process is reversed via a parameterized diagonal Guassian transition:

We choose the notation pθ(xt−1∣xt)=N(μθ,σθ2I)p_{\theta}(x_{t-1}|x_{t})=\mathcal{N}(\mu_{\theta},\sigma^{2}_{\theta}\mathbf{I}) for brevity. In order to learn the backward process, neural networks are trained to predict μθ\mu_{\theta} and σθ2\sigma^{2}_{\theta}.

The formulations above explain the unconditional backward process pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_{t}); with an extra guidance signal yy, the sampling distribution becomes:

where ZZ is a normalizing constant. It is proven that the new distribution after incorporating the guidance can be approximated by a Gaussian distribution with shifted mean:

where μ=μθ\mu=\mu_{\theta}, Σ=σθ2I\Sigma=\sigma^{2}_{\theta}\mathbf{I}, g=∇xt−1log⁡pϕ(y∣xt−1)g=\nabla_{x_{t-1}}\log p_{\phi}(y|x_{t-1}).

Class-guided synthesis was explored in where yy is a discrete class label, and pϕ(y∣xt−1)p_{\phi}(y|x_{t-1}) is the probability of xt−1x_{t-1} belonging to class yy. Here, we generalize yy to a continuous embedding for language, image or multimodal guidance. In the following, we introduce the guidance function Fϕ(xt,y,t)=log⁡pϕ(y∣xt)F_{\phi}(x_{t},y,t)=\log p_{\phi}(y|x_{t}) for different guidance types.

Figure 1 and Algorithm 1 summarize the proposed Semantic Diffusion Guidance. Note that there is an additional scaling factor ss for semantic guidance in Algorithm 1, a user-controlled hyperparameter that determines the strength of the guidance. We discuss its effect in Section 4.

2 Language Guidance

Language is one of the most intuitive ways that a user can control the generation model. In order to incorporate language information to the image synthesis process, we use a visual-semantic embedding model for image-text alignment. Specifically, given an image xx and a text prompt ll, the model embeds them into the joint embedding space using an image encoder EIE_{I} and a text encoder ELE_{L}, respectively. The similarity between the embeddings EI(x)E_{I}(x) and EL(l)E_{L}(l) is calculated as the cosine distance, and we utilize this to formulate the language guidance function.

However, the models for backward process and guidance in Equation 5 are time-dependent, and take noisy images as input. This means that the image encoder EIE_{I} needs to incorporate the timestep tt as input and be further trained on noisy images at different timesteps as well. We denote such time-dependent image encoder for noisy images as EI′E^{\prime}_{I}. Finally, the language guidance function can be defined as:

where E′E^{\prime} denotes the image encoder trained on noisy images with additional timestep input. In Section 3.5, we give details on adapting a CLIP model to become time-dependent with minimal architecture changes, and present a self-supervised finetuning strategy for noisy images.

3 Image Guidance

Sometimes, an image can convey information that is difficult to express in language. For example, users may want to generate a photo of a cat that looks similar to another cat, or want to generate a photo of a bedroom in the style of Van Gogh’s painting “The Starry Night”. They may also want to generate realistic images given an emoji or a painting. We thus propose an approach for image-guided diffusion that effectively controls the content or style information according to an image. We present two types of image guidance, namely image content guidance and image style guidance.

Image Content Guidance aims to control the content of the generated image, with or without structural constraints, based on a reference, and is formulated as the cosine similarity of the image feature embeddings. Let x0′x^{\prime}_{0} denote the noise-free reference image. We perturb x0′x^{\prime}_{0} per Equation 2 to get xt′x^{\prime}_{t}. Then, the guidance signal at timestep tt is,

Similar to language guidance, we use an image encoder finetuned with noised images to define the image guidance function and extract embeddings that mostly capture the high-level semantics. An interesting property of using image encoders for guidance is that one can control how much structural information such as pose and viewpoint is maintained from the reference image. For instance, the embeddings used in Equation 7 do not have spatial dimensions, resulting in samples with great variations in pose and layout. However, by utilizing spatial feature maps and forcing alignment between features in corresponding spatial locations, we can guide the generated image to additionally share similar structure with the reference image as follows.

where EI′()j∈RCj×Hj×WjE^{\prime}_{I}()_{j}\in\mathcal{R}^{C_{j}\times H_{j}\times W_{j}} denotes the spatial feature maps of the jj-th layer of the image encoder EI′E^{\prime}_{I}.

Image Style Guidance allows style transfer from the reference image. It is formulated similarly, except the alignment between the Gram matrices of the intermediate feature maps is enforced:

where GI′()jG^{\prime}_{I}()_{j} is the Gram matrix of the jj-th layer feature map of the image encoder EI′E^{\prime}_{I}.

4 Multimodal Guidance

In some application scenarios, image and language may contain complementary information, and allowing both image and language guidance at the same time provides further flexibility for user control. Our pipeline can easily incorporate both by a weighted sum of the two guidance functions, with their scaling factors as weights.

By adjusting the weighting factors of each modality, users can balance between the language and image guidance.

5 Self-supervised Finetuning of CLIP without Text Annotations

CLIP is a powerful vision and language model trained on large-scale image-text data. We leverage its semantic knowledge to achieve controllable synthesis for diffusion models. To act as a guidance function, CLIP is expected to handle noisy images xtx_{t} at any timestep tt. We make a minor architectural change to CLIP image encoder EIE_{I} to accept an additional input tt by converting batch normalization layers to adaptive batch normalization layers, where the prediction of scale and bias terms is conditioned on tt. We denote this modified CLIP image encoder as EI~\widetilde{E_{I}}. The parameters of EI~\widetilde{E_{I}} are initialized by the parameters of the pretrained CLIP model EIE_{I}, except for the parameters for the adaptive batch normalization layers.

To finetune EI~\widetilde{E_{I}}, we propose a self-supervised approach in which we force an alignment between features extracted from clean and noised images. Formally, given a batch of NN pairs of clean and noised images {x0i,xtii}i=1N\{x^{i}_{0},x^{i}_{t_{i}}\}_{i=1}^{N} where tit_{i} is the timestep sampled for the ii-th image that governs the amount of noise, we encode x0ix^{i}_{0} and xtiix^{i}_{t_{i}} with EIE_{I} and EI~\widetilde{E_{I}}, respectively. We rely on CLIP’s contrastive objective to maximize the cosine similarity of the NN positive pairs while minimizing the similarity of the remaining negative pairs. We fix the parameters of EIE_{I} and use the contrastive objective to finetune the parameters of EI~\widetilde{E_{I}}. With our finetuned CLIP model, the diffusion model can be guided by image or language information that users provide. Moreover, the CLIP model is finetuned in a self-supervised manner without requiring any language data for the target dataset.

Experiments

We conduct experiments on FFHQ and LSUN cat, horse, and bedroom subsets. FFHQ dataset contains 70,000 images of human faces. LSUN contains 3 million bedroom images, 2 million horse images, and 1.7 million cat images. We use unconditional DDPMs from , and finetune CLIP RestNet 50×\times16 models on noised images on each dataset with initial learning rate 10−410^{-4} and weight decay 10−310^{-3}, with a batch size of 256. When synthesizing images with our SDG, the scaling factor is a hyperparameter that we manually adjust for each guidance, which will be discussed in Sec. 4.3. The default scaling factor is 100 for image guidance and 120 for language guidance.

2 Quantitative Evaluation

Evaluation Setup Since our SDG is the first method that unifies text guidance and image guidance for image synthesis, there is no previous work on image synthesis with both image and language guidance. So we evaluate the language-guided image synthesis and image-guided image synthesis separately in order to compare with previous work. We evaluate the language-guided generation on FFHQ dataset. For that we define 400 text instructions based on combinations of gender and face attributes from CelebA-Attributes . For example, “A photo of a smiling man with glasses”. We generate 25 images for each text query, which results in 10,000 images in total. We compare our language-guided generation with StyleGAN+CLIPhttps://colab.research.google.com/drive/1br7GP_D6XCgulxPTAFhwGaV-ijFe084X, which uses CLIP loss to optimize the randomly initialized latent codes of StyleGAN for text-guided image synthesis. StyleGAN+CLIP removes the GAN inversion module of StyleCLIP so that it can be applied for language-based image synthesis. Since our model does not require text annotation for training, our text-guided image synthesis experiments are conducted on image-only datasets without paired text annotations. So our method cannot be directly compared with other text-based image synthesis methods which have to be trained on text-image paired datasets. To evaluate image-guided image synthesis, we randomly choose 10,000 images from each dataset as guidance and synthesize new images based on the guidance images. We compare our image-guided results to ILVR .

We present quantitative results and comparison with previous work in Table 1 with the following evaluation metrics.

FID for image quality evaluation. We report FID score calculated on 10,000 images for each dataset to evaluate the quality of generated images. Lower FID indicates better generation quality. Our SDG outperforms comapred methods for both image-guided synthesis and language-guided synthesis.

LPIPS for diversity evaluation. We calculate the LPIPS score between paired images generated from the same image guidance or the same text guidance, as shown in Table 1. Higher LPIPS indicates more diversity. Our model generates more diverse images compared to previous work ILVR and StyleGAN+CLIP. The images generated by ILVR follows the same structure and layout, with variations in details. While our method is able to generate diverse images with different pose, structure, and layout, as shown in Figure 6(a). The images generated by StyleGAN+CLIP also suffers from low diversity, as shown in Figure 6(b). The high FID score of StyleGAN+CLIP is also because of the low diversity of the generated images.

Retrieval accuracy to evaluate consistency with guidance. We use text-to-image retrieval or image retrieval by an original CLIP ResNet 50×\times16 model without finetuning to evaluate how well the generated images matches the guidance. For an image generated with text guidance, we randomly select 99 real images from the training set as negative images, and evaluate the text-to-image retrieval performance. Similarly, for an image synthesized with a reference image, we use the reference image to retrieve the generated image from the 99 randomly selected real imagesThe selected negative images are disjoint with the guidance images we used for synthesizing images.. StyleGAN+CLIP has a very high retrieval performance because the latent codes of the StyleGAN model are directly optimized to minimize the CLIP score calculated by the CLIP model used for retrieval. So the high retrieval performance of StyleGAN+CLIP comes at the cost of low generation diversity, as indicated by the high FID and low LPIPS scores.

3 Ablation Study

As demonstrated in Section 3.1 and Algorithm 1, the scaling factor ss is a user-controllable hyper-parameter that controls the strength of the guidance. We explore the effect of the scaling factor in Table 2 and Table 3. A visual example of the effect of different scaling factors is shown in the Appendix. We observe the trade-off between semantic correctness and diversity of generated images. As the scaling factor gets larger, the guidance signal has more control on the generation results, as indicated by the increased semantic consistency with the guidance. While larger scaling factor also leads to lower diversity of generated images. Users can adjust the scaling factor to control how diverse they expect the generated images to be.

4 Qualitative Results

Text-guided and image-guided synthesis results Our model combines the language and image guidance in a unified framework, and is easy to adapt to various applications. In Figure 2 we show the synthesis results with image content guidance (Equation 7). With the image guided diffusion, the model is able to synthesize new images with diverse structures that match the semantics of the guidance image. Figure 3 shows the language-guided diffusion results, where our model is able to handle complex and fine-grained descriptions, such as “A smiling woman with curly brown hair and lipstick.”, or “A bedroom with a wooden closet and a painting on the wall.” We can also incorporate language and image guidance jointly, as shown in Figure 4. The image and language guidance provide complementary information, and our semantic diffusion guidance is able to generate images that align with both. For example, we can generate a bedroom similar to the guidance bedroom image but with windows, or generate a woman according to a guidance image but with a new attribute defined the language guidance (e.g., “smiling” or “short hair” or “sunglasses”).

Comparison to prior work Since there is no prior work that incorporates text and image guidance in the same unified framework, we compare our approach to previous text-guided and image-guided synthesis work. In image-guided synthesis, the most related to our work is ILVR . As shown in Fig. 6(a), our model can generate images in different poses and structures, while ILVR can only generate images of the same pose and structure. We compare our language-guided image synthesis with StyleGAN+CLIP in Fig. 6(b). Although StyleGAN+CLIP is able to generate high-quality images, diversity is lacking in their results, while our model is able to generate high-quality and diverse results based on the language instructions.

Other applications In Fig. 6(a,b), we demonstrate the results of style (Equation 9) and structure-preserving (Equation 8) image guidance. With the style guidance, the model trained on LSUN bedroom is able to synthesize bedrooms in the unseen style. With the structure-preserving content guidance, the synthesized images preserve the structure, pose, and layout from the reference image. Fig. 6(c) shows that the model is able to take an out-of-domain image as guidance, and synthesize photo-realistic images which are semantically similar to the guidance cartoon image.

Conclusion and Discussions

We propose Semantic Diffusion Guidance (SDG), a unified framework for diffusion-based image synthesis with language, image, or multi-modal guidance. The flexible guidance module allows us to inject various types of guidance into any off-the-shelf unconditional diffusion model without re-training or finetuning the diffusion model. We further present a self-supervised efficient finetuning scheme for the CLIP guidance model which does not require textual annotations. However, image generation has as much potential for misuse as it has for beneficial applications. We should be aware of the potential negative social impact if image synthesis is used for generating fake images to mislead people.

Acknowledgements

This work was supported in part by DoD including DARPA’s SemaFor, PTG and/or LwLL programs, as well as BAIR’s industrial alliance programs.

References