Sketch-Guided Text-to-Image Diffusion Models

Andrey Voynov, Kfir Aberman, Daniel Cohen-Or

Introduction

Large text-to-image diffusion models have been an inspiring tool for content creation and editing, enabling synthesis of diverse images with unprecedented quality that follow a given text-prompt. Despite the semantic guidance provided by the text-prompt, these models still lack intuitive control handles that can guide spatial properties of the synthesized images. In particular, guiding a pertained text-to-image diffusion model during inference with a spatial map from another domain, such as sketch, is yet an open challenge.

A possible attempt is by training a dedicated encoder to map the guiding image into the latent space of the pretrained unconditional diffusion model . The trained encoder, however, performs well in-domain, but struggles with out-of-domain free-hand sketches.

In this work, we introduce a generic approach to guide the inference process of a pretrained text-to-image diffusion model with a spatial map. Our key idea is to use a small multi-layer perceptron (MLP) network that is trained to map latent features of noisy images to spatial maps, where the latent features are extracted from the core network of the diffusion model. The trained MLP serves as a latent guidance predictor, over which the loss with a target spatial map is computed and propagated back to push the intermediate image to agree with the map.

The latent guidance predictor is trained in a self-supervised fashion and learns to translate features of images with different noise levels into encoded spatial maps, where the noise scheduling corresponds to the noise scheduling of the diffusion process. Importantly, the latent guidance predictor is trained, and operates independently, on each latent pixel in the latent space, rather than on the whole image. Hence, it is sufficient to train it with a few thousand images only, which is a few orders of magnitude less than the required amount to train a dedicated image-to-image translation model. Yet, the latent guidance predictor is generic and domain-oblivious, in the sense that it can operate on out-of-domain guiding maps. Hence, our method can accept free-hand sketches inputs, as in Figure 1, and generate diverse results that correspond to the text-prompt and follow the spatial layout of the sketch.

In our experiments, we demonstrate sketch-guided text-to-image synthesis results on various domains, including free-hand style drawing. We conduct experiments and ablation studies to analyze the performance of various components of our method and present comparisons to other image translation approaches. In addition, we show that our general framework can be applied to other spatially guided text-to-image tasks such as saliency-guided inpainting and horizon control. Throughout our examples in the paper, we demonstrate that our method can be applied to a rich variety of sketch styles from diverse domains, which is the key advantage of our approach. Project page: sketch-guided-diffusion.github.io

Related Work

Image-to-image translation has been a long-standing task in the computer vision domain with a myriad of explorations and prior works . One common architecture to perform image translation is conditional generative adversarial networks that translate an input domain to an output domain with a discriminator to bridge the gap between real images and fake ones. Commonly, these methods tackle each task independently and use task-specific datasets and models. Considering the commonality between tasks, some research efforts aim to learn a unified model for diverse translation tasks via multi-task training.

A special case of image-to-image translation is the task of sketch-to-photo task . SketchyGAN uses edge-preserving image augmentations to train a Generative Adversarial Network (GAN), ContextualGAN leverages conditional GAN with joint image-sketch representation. CoGS minimizes the distances between the embeddings of the input sketch and the corresponding ground truth real image in the vector-quantized space of a VQ-GAN . In contrast, our work leverages a pretrained text-to-image generative prior of general images and treats the image translation problems as downstream tasks.

2 Diffusion Models

DDPM introduced unprecedented quality of conditional and unconditional image synthesis. , rivaling GAN-based methods both in visual quality and sampling diversity. In particular, uses diffusion models to solve various image translation tasks, but tackles each task independently and trains a model from scratch for each task.

More recently, diffusion models have demonstrated unprecedented quality for text-to-image synthesis and editing tasks, when large models are trained on pairs of text and images . Our approach builds on these key advances, and we show how a pretrained text-to-image diffusion model can be guided by a spatial map from a different domain and serve as a universal generative prior that facilitates various image translation tasks. Note that models like Make-a-Scene and eDiffi allow the user a particular type of control by enabling them to provide a semantic segmentation map to control the composition of the elements in the synthesized image.

ILVR proposes to iteratively refine the diffusion process using a noisy reference image at each time-step during inference, enabling control of the amount of high-level semantics being adapted from the images. In addition, SDEdit suggests to add noise to the input guiding image, halfway of the forward diffusion process, then denoise it in a reverse process with a guiding text. Both approaches enable to guide the model with an image where the guiding image should lay in the RGB domain and the fidelity to the spatial property of the guiding image is limited and random.

A closely related approach is the recent work of Wang et al. that suggests to use a pretrained unconditional diffusion model for various image translation tasks, by training a specialized, per-task, encoder to map spatial maps into the latent space of the diffusion model. While their approach requires dedicated large-scale training to train the encoder, we use a light weight training (only a few thousand images are required) of a small MLP which is trained per-pixel, and thus offers a generalization that extends beyond the domain defined by the training data.

Method

In this section, we describe the main steps of the proposed spatially-guided text-to-image synthesis approach. Although our method is generic, for the ease of reading, in the method description we are focused on the sketch-to-image task, and in Section 4, we show how the very same approach also works for different tasks.

The key idea of our method is to guide the inference process of a pretrained text-to-image diffusion model with an edge predictor that operates on the internal activations of the core network of the diffusion model, encouraging the edge of the synthesized image to follow a reference sketch. Our edge predictor is an MLP network, that operates per-pixel, and is trained to map features of noisy images into spatial edge maps. The training procedure which is performed only one time, requires a few thousand images only, and takes only about an hour on a single GPU.

Our first goal is to train an MLP that guides the image generation process with a target edge map. The MLP is trained to map the internal activations of a denoising diffusion model network into spatial edge maps, as depicted in Figure 2. Inspired by , we extract our activations from a fixed sequence of intermediate layers in the core U-net network UU of the diffusion model. Formally, for an input tensor ww, we denote F(w∣c,t)=[l1(w∣c,t),…,ln(w∣c,t)]{\bf F}(w|c,t)=\left[l_{1}(w|c,t),\dots,l_{n}(w|c,t)\right] as the concatenated activations of selected internal layers {l1,…,ln}\{l_{1},\dots,l_{n}\}, when ww is processed by the network with a conditioning text-prompt cc and noise level tt. Since activations from different layers may have different spatial resolution, we resize them to match the spatial dimensions of the input ww and concatenate them alongside the channel dimension. The input dimension of the MLP is then the sum of the number of channels of the selected activations.

Our training corpus D\mathcal{D} is formed by triplets (x,e,c)(x,e,c) of an image, edge map, and a corresponding text caption, respectively. Since our work is implemented with latent diffusion models (specifically Stable Diffusion), we use the model encoder EE to preprocess the images and the edge maps. In order to encode the edge map, we convert it into a 3-channel image by replicating its intensity channel. Thus, in practice the input tensor is the encoded image with additive Gaussian noise, zt=αt⋅E(x)+μt⋅ξz_{t}=\alpha_{t}\cdot E(x)+\mu_{t}\cdot\xi, where 0≤αt,μt≤10\leq\alpha_{t},\mu_{t}\leq 1 are the blending scalars that is dictated by the noise scheduling of the diffusion model, and the MLP is trained to map the concatenated features F(zt∣c,t){\bf F}(z_{t}|c,t) to the encoded edge map E(e)E(e).

In order to consider the noise level of the input, the MLP also receives tt and its positional encoding as sin⁡(2πt⋅2−l), l=0,…,9\sin(2\pi t\cdot 2^{-l}),\ l=0,\dots,9. The output dimension of the MLP is equal to the number of output channels of EE (4 in the case of Stable Diffusion). Each spatial position (i,j)(i,j), of the latent pixel F(z∣c,t)ij{\bf F}(z|c,t)_{ij} is translated to the corresponding latent edge E(e)ijE(e)_{ij} by PP, thus, the training objective of our latent edge predictor PP is

where PP is applied to each latent pixel independently.

Once optimized with the objective L\mathcal{L}, the model PP constitutes a per-spatial location differential predictor of encoded edges for an encoded image with noise level tt. Due to the per-pixel nature of the architecture, the MLP is trained to predict edges in a local manner, being agnostic to the domain of the image. In addition, it enables training on a relatively small corpus (a few thousand images), in reasonable training time (One hour on a single A100 GPU).

We next show how such a component can serve as a guidance through the diffusion process.

2 Sketch-Guided Text-to-Image Synthesis

Given a sketch image ee and a caption cc, our goal is to generate a corresponding highly detailed image that follows the sketch outline. Figure 3 illustrates the proposed latent features-based guidance described in detail below.

We start with a latent image representation zTz_{T} sampled from a uniform Gaussian. Normally, the DDPM synthesis consists of TT consecutive denoising steps zt→zt−1z_{t}\to z_{t-1} which constitute the reverse diffusion process, with z0z_{0} being an encoded output image. The reverse diffusion process, on each of the denoising steps t=T,…,1t=T,\dots,1, evaluates a density score gradient estimation ε(zt,t,c)\varepsilon(z_{t},t,c), and based on it, depending on a sampler algorithm, computes the next sample zt−1z_{t-1}. Notably, the score gradient computation consists of the forward pass of the main denoising U-net model. Thus, once the quantity ε(zt,t,c)\varepsilon(z_{t},t,c) is computed, we may also collect the intermediate activations l1(zt∣t,c),…,ln(zt∣t,c)l_{1}(z_{t}|t,c),\dots,l_{n}(z_{t}|t,c).

with β\beta being a constant throughout the synthesis process. Normally β\beta takes values of order O(1)O(1). Once being synthesized with the guidance from the objective L\mathcal{L}, the model produces a natural image aligned with the desired sketch.

In practice, as the final steps of the reverse denoising process commonly do not affect the geometric layout of the final generated image, we perform the edge guidance only for the steps t=T,...,S>1t=T,...,S>1, where commonly S=0.5TS=0.5T. We further discuss the choice of the edge guidance stop step SS in the following section.

Experiments

In this section, we discuss the implementation details of our approach, show sketch-guided text-to-image synthesis results, conduct experiments and ablation studies to analyze the performance of various components in our framework, and present comparisons to state-of-the-art image translation techniques. Figure 4 shows a gallery of results which demonstrate the ability of our framework to convert sketches to images with an input text-prompt.

In all of our experiments, we use Imagenet samples with their class names as captions (e.g., “shoes”). The corresponding edge maps were generated with the edge prediction model of and then thresholded with 0.5. The latent edge predictor consists of 4 fully-connected layers with ReLU activations, batch normalization, and hidden dimensions 512512, 256256, 128128, 6464, and output dimension 44. The denoising model’s features are taken from 9 different layers across the network: input block - layers 2, 4, 8, middle block - layers 0, 1, 2, output block - layers 2, 4, 8. The training is performed for 3000 steps with Adam optimizer and batch size 16 which takes less than an hour on a single A100 GPU.

For inference, we found a set of reliable parameters for the edge guidance scale β=1.6\beta=1.6, guidance stop step S=0.5TS=0.5T, and prompt-conditioning equal to 88 (classifier-free guidance scale in DDPM), though these parameters can be modified based on the user requirement, to balance between edge fidelity and realism (see Section 4.3 for more details).

2 Comparisons

We compare our method to three types of baseline approaches: SDEdit , pix2pix , and PITI , each of which can be used for sketch-to-image synthesis.

A possible attempt to solve the sketch-to-image task with a pretrained text-to-image diffusion model, is by adding noise to the input sketch, for tt steps in the forward diffusion process, then denoise it in a reverse process with a text prompt, as suggested in SDEdit . This process enables to implicitly guide the model with a spatial map. However, as can be seen in Figure 5, the model expects that the guiding image lays in the RGB domain, hence, resulting in unnatural, black and white images that follow the input sketch (text-prompt condition used: “A photograph of a bike made of wood”). For low values of tt, the system struggles to add texture to the model, and when tt is increased, the fidelity to the input sketch significantly decreases.

We next compare our approach to pix2pix , and PITI . Pix2pix is a self-supervised method that is trained on pairs of images (in this case real images and their corresponding sketches) using a reconstruction loss that is enhanced by the adversarial loss that is applied to pairs. Figure 6 demonstrates that this approach works well on sketches that lay within the domain of the training data, for example, realistic shoes sketch, while failing on out-of-domain hand-drawn sketches.

PITI trains a dedicated encoder to map the guiding image into the latent space of the pretrained unconditional diffusion model. Figure 7 shows that while PITI performs well on realistic sketch samples, it struggles to create realistic outputs on free-hand sketches that are out of their training data domain. In addition, notably, our method provides significantly more color and style variability compared to their approach.

3 Ablations and Parameter Tuning

We next discuss the different parameters in our system and their effect on edge fidelity. First, our method demonstrates a trade-off between the realism level of a generated image and its alignment with the edges of the target sketch. The trade-off, which can be controlled by the edge-guidance scale β\beta, is depicted in Figure 8. It can be seen that for small values of β\beta, we get a more realistic image with details and textures that cover regions in the entire image, while the larger value of β\beta favors edge alignment but generates less realistic, piece-wise smooth, results.

We also quantitatively measured the edge-fidelity (Mean Squared Error between the target edge map and the edges of the synthesized image), as a function of the guidance stop step SS and depicted the result in Figure 10. As expected, the quality of edge reconstruction is improved for larger values of SS. However, since high edge fidelity comes at the expense of realism, we want to find a sweet spot that will enable us to balance these two factors. For that, we conducted an experiment that measures the reconstruction error of our Latent Edge Predictor for different values of tt. Figure 11 depicts the loss in Equation 1 as a function of tt . Notably, starting from t≈0.5Tt\approx 0.5T, the error stabilizes, indicating that for t<0.5Tt<0.5T, the model does not receive new information on the edges. Hence, we use this stabilization point as the guidance stop SS to mitigate the trade-off, which is highly aligned with the segmentation errors as a function of tt that were reported in , that uses latent features of DDPM for the few-shots semantic image segmentation task.

Since our latent edge predictor works in a local, per-pixel, manner, we also demonstrate its insensitivity to the stroke style. We generated samples produced with the same sketch geometry but with different stroke styles. Figure 9 shows that such a setting yields the same shape, but with variation in colors and textures. This observation explains that the stroke style affects the inner synthesis process only, which accumulates into varying colors of the output image.

Applications

We demonstrated our spatially-guided text-to-image synthesis approach on the sketch-to-image application, however, our approach is generic and can be applied to different image-to-image translation tasks. In this section, we show how can we use saliency maps as a guiding map for text-to-image models, and how it can be used for natural enhancement of image regions, as well as background inpainting . More applications and examples can be found in the supplementary material.

Saliency prediction models can be used to detect the most attention grabbing regions within an image. Recent works have shown that saliency can be also used as a guiding component for image editing, to reduce distraction in images . We next show that saliency maps can also guide text-to-image diffusion models such that the saliency in specific areas of the generated image is high or low. In this case, we train our latent guidance predictor to directly predict the original (not encoded) downsampled saliency maps from our noisy latent features, and we optimize the model with the binary cross-entropy loss instead of mean squared error. To supervise the MLP we use the saliency model from . We then run the inpainting model with out-of-mask conditioning, and with an empty prompt. Figure 13. demonstrates how this technique can be used for background inpainting. For a given image of a bird and a mask that covers the bird’s body without the tail, it can be seen that the model fills the hole with a new bird due to the semantic hint that the tail provides to it. In contrast, when the model is guided by a saliency map with low values in the mask region, it simply removes the bird body and fills it in with background inpainting as expected. In addition, we can guide the model to generate high saliency values within a region. Figure 12 shows how the marked region is highly illuminated by the sun due to the requirement for high saliency values in this region. Notably, it is sufficient to apply the latent guidance predictor for only the first 20%20\% of the steps.

Conclusions

We presented a technique to guide a pre-trained text-to-image model diffusion model with a spatial map. We have focused on sketch-guidance, and showed that the technique can handle well out-of-domain sketches, which may have a large variety of styles completely different than those seen in the training time. The gist of the technique is the per-pixel training of a lightweight MLP component that is trained on rather small training data. The per-pixel training acts more like a differential edge-detector and unlike common per-image training, it is not bound to a particular global sketching style.

Our technique piggybacks on a pertained text-to-image model diffusion model and thus offers a strong multi-modal sketch-guidance technique to users. In a sense, the technique accepts a rich variety of sketching styles and at the same time provides a rich variety of outputs, where the user has intuitive control over the input, and semantic control over the output.

Still, our presented technique is only a step toward gaining more control over the output of generative text-image models. The technique has its limitations. Currently, the technique is vulnerable to the local style of the strokes. The technique still struggles with complex and cluttered sketches as it treats all of the strokes equally without prioritizing them according to their saliency or semantics. Also, since the text-image diffusion model is stochastic, there might be conflicts between the random seed and the input sketch, which may lead to a generation of an output that does not agree well with the sketch. Figure 14 shows representative examples where the model fails to provide satisfying results. The quality of the results may drop for different initialization, and complex scenes with mixed and ambiguous semantics.

In the future, we would like to advance and improve the technique by adding a sketch inversion step to yield a stronger seed to the diffusion process, to better push the output toward the outline of the input sketch. Another direction is to quickly learn a personalized style using just a few shots. With a quick training session, the latent sketch predictor can accommodate the artist’s stroke style.

Acknowledgements

We thank Chu Qinghao, Yael Vinker, Yael Pritch, Dani Valevski and David Salesin for their valuable inputs that helped improve this work.

References

Appendix A Societal Impact

This work provides a powerful tool to convert simple sketches to detailed, highly realistic images with full control over style and content. As reported by professional artists we interviewed, while a simple sketch drawing takes just a few minutes, a detailed colored picture based on it normally takes way more time and counts in hours. Thus, this work can potentially significantly speed up the process of artistic creation, enabling the democratization of creativity. This work is inspired by the idea of not replacing an artist, but giving an artist a tool where AI takes all the technical parts of the creativity process while leaving the imagination and inspiration to a human. As the back side of the proposed method, this also gives a tool for deep-fake and misleading material creation, this once more stands the challenge before the community to create the safety mechanism to prevent the generative models to be used in with controversial intentions.

Appendix B Ablation studies

We start by performing a test to highlight the ”out-of-domain” virtue of the proposed technique. Our Latent Edge Predictor (LEP) appears to perform well out of its training domain samples due to its per-pixel training nature. We examine it with an extremely tiny training set. We train the edges predictor on the subset of Imagenet validation set, consisting of the dogs’ classes only. The qualitative evaluation and comparison of a model trained with this single-domain protocol is depicted in Figure 20. Notably, even in this minimal setup, it still performs reasonably well.

To highlight the importance of using the deep features of the diffusion network, rather than the intermediate states, we perform the following experiment. We train an independent noise level-conditioned edge predictor that operates over intermediate states ztz_{t}, instead of the internal features of the network. Though this is a straightforward generalization of the classifier-guidance technique, we have not succeeded to train a plausible predictor, as it commonly collapses to the prediction of empty maps. We argue that this is due to the fact that edge prediction based on a noisy image is almost as complex as the original DDPM denoising, which requires a comprehensive image understanding. Thus, making this approach work might take a significant computational effort, while our proposed per-pixel training works out of the box for a rich variety of tasks. In addition, our approach takes only an hour to train and requires a limited amount of data.

We also tried to perform guidance over the intermediate z0z_{0} predictions of the DDPM model – a more intuitive input to the edge predictor which operates better on clean images. In each of the denoising steps zt→zt−1z_{t}\to z_{t-1}, the model simultaneously predicts the end result z0=z0(t)z_{0}=z_{0}(t). Given this prediction, we compute the current edges prediction as e(E−1(z0(t)))e(E^{-1}(z_{0}(t))) where E−1E^{-1} states for the VQVAE decoder, and e(⋅)e(\cdot) is the edge prediction model we use for the edges labeling. Then we guide the denoising process with the gradients of the similarity of the predicted edges and target edges. Despite the fact that the edge predictor operates on a noise-free image, it still struggles to predict edges from those images that are directly estimated from fully noisy images, rather than passing through the entire diffusion process. Hence, the entire process of sketch guidance fails. Figure 15 demonstrates the guidance performed that way. It can be seen that samples with low guidance weight fails to produce an image that matches the sketch, while high guidance weight produces nearly adversarial images.

Appendix C Additional tests

Our approach enables to use inputs from two different modalities - a spatial sketch map and a text. Figure 18 visualizes the effect of different prompts guidance. While an empty prompt commonly induces unsatisfying results, a minimal relevant prompt induces plausible generation. Commonly, a detailed description of a sketch subject induces more realistic and detailed generation (column 4: ”…a cow on a snowy field…”). Once a prompt contains objects non-presented in the sketch, it may confuse the generation process (in column 5, ”…a cow surrounded by trees.”: there are no tree edges presented in the sketches. While the model succeeded in generating a proper environment for the second sketch, for the first sketch it makes the trees to be formed by the sketch subject shape). Adding an artistic prefix (column 6: ”An oil painting…”) always induces high alignment with the sketch as in that case the prompt guidance is less concerned about the generated image realism. The rightmost column shows conflicting text prompt (”a dog”). The output is indeed a dog that admits to the guiding sketch. This clearly highlights the competence of our edge-guiding technique.

Note that due to stochastic sampling of DDPM, a single sketch and a prompt can generate a variety of different samples, as shown in Figure 16.

Spatial Labels Guidance

We next show another application of our generic approach and use a soft $grayscalemaptoguidethediffusionprocess,wherethemaximalandminimalvaluesofthemaprepresenttwoclasses−dayandnight.Wetrainourper−pixelMLPtopredictwhetherapixelbelongstoadayornightscene.Namely,basedontheaggregatednoisedstackednoisedfeaturesgrayscale map to guide the diffusion process, where the maximal and minimal values of the map represent two classes - day and night. We train our per-pixel MLP to predict whether a pixel belongs to a day or night scene. Namely, based on the aggregated noised stacked noised features\mathbf{F}(z_{t}|c,t)_{i,j},theMLPmodel, the MLP modelPpredictseitherthespatiallocationpredicts either the spatial location(i,j)ontheoriginalencodedimageon the original encoded imagezrepresentasceneatthedaytime,oratthenighttime.Wetrainrepresent a scene at the day time, or at the night time. We trainPwithonly500dayand500nightimages,whereallpixelsfromtheoneimage(dayornight)correspondedtothesameclass.Asthedataislimited,weperformoptimizationwith1000stepsonly.Thetrainingschemeremainsunchangedexceptforthelosswherenowweusethecross−entropybetweenwith only 500 day and 500 night images, where all pixels from the one image (day or night) corresponded to the same class. As the data is limited, we perform optimization with 1000 steps only. The training scheme remains unchanged except for the loss where now we use the cross-entropy betweenP(\mathbf{F}(z_{t}|c,t)_{i,j})andtheground−truthclassatthelocationand the ground-truth class at the location(i,j)$. We also always use the null prompt for feature extraction. Similarly, throughout the generation, we guide with the cross-entropy loss. Now, when the guidance is performed with a constant 1 or 0 spatial labeling map, the produced images are either attributed to day or night (Figure 17, top). When the labeling map is formed by the interpolation between labels probabilities, the generated image also interpolates the scene between night and day (Figure 17, bottom).

The proposed method induces a computational overhead that is mostly induced by the backpropagation of the edge prediction loss from the inner features to the input image. Once the guidance is applied for the first half of all reverse diffusion steps, the relative sampling time overhead is approximately 80%.

Appendix D Models details and data

Dataset: Sketches presented in Figure 4, Figure 3 (third row), and Figure 14 are provided by the authors. The sketch presented in Figure 6 (first row) is taken from edge2shoes dataset , and the rest of the samples are taken from Sketchy dataset . All synthetic quantitative results are based on Sketchy dataset, and the real numbers are based on Imagenet with class names used as prompts.

Prompts used in Figure 4 (b): ”A skull of a monster.”, ”A macro photograph of a snail.”, ”A photograph of a windmill.”, ”A hot air balloon.”, ”A photograph of a barn owl.”, ”A photograph of a big wave.” (left in the last row), ”William Turner’s picture of a big wave.”. A prompt used in Figure 6: ”A shoe.”. Prompts used in Figure 7: ”A photograph of a giraffe.”, ”A photograph of an elephant.”, ”A mountain in clouds.”, ”An oil painting of a mountain in clouds.”, ”A photograph of a green mountain in clouds.”. Figure 8: ”A photograph of a wooden house on a hill in the winter.”. Figure 9: ”A hydrant.”.

In all the experiments except inpainting, we use the Stable Diffusion checkpoint stable-diffusion-v-1-4-original. As for inpainting, we use the checkpoint stable-diffusion-inpainting. We always sample in the non-deterministic mode with 250 reverse diffusion steps.