Inversion-Based Style Transfer with Diffusion Models

Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, Changsheng Xu

Introduction

If a photo speaks 1000 words, every painting tells a story. A painting contains the engagement of an artist’s own creation. The artistic style of a painting can be the personalized textures and brushstrokes, the portrayed beautiful moment or some particular semantic elements. All those artistic factors are difficult to be described by words. Therefore, when we wish to utilize a favorite painting to create new digital artworks which can imitate the original idea of the artist, the task turns to example-guided artistic image generation.

Generating artistic image(s) from given example(s) has attracted many interests in recent years. A typical task is style transfer , which can create a new artistic image from arbitrary input pair of natural image and painting image, by combining the content of the natural image and the style of the painting image. However, particular artistic factors such as object shape and semantic elements are difficult to be transferred (see Figures 2(b) and 2(e)). Text guided stylization produces an artistic image from a natural image and a text prompt, but usually the text prompt for target style can only be a rough description of material (e.g., “oil”, “watercolor”, “sketch”), art movement(e.g. “Impressionism”,“Cubism”), artist (e.g., “Vincent van Gogh”, “Claude Monet”) or a famous artwork (e.g., “Starry Night”, “The Scream”). Diffusion-based methods generate high-quality and diverse artistic images based on a text prompt, with or without image examples. In addition to the input image, a detailed auxiliary textual input is required to guide the generation process if we want to reproduce some vivid contents and styles, which may be still difficult to reproduce the key idea of a specific painting in the result.

In this paper, we propose a novel example-guided artistic image generation framework InST which related to style transfer and text-to-image synthesis, to alleviate all the above problems. Given only a single input painting image, our method can learn and transfer its style to a natural image with a very simple text prompt (see Figures 1 and 2(f)). The resulting image exhibit very similar artistic attributes of the original painting, including material, brushstrokes, colors, object shapes and semantic elements, without losing diversity. Furthermore, we can also control the content of the resulting image by giving a text description (see Figure 2(c)).

To achieve this merit, we need to obtain the representation of image style, which refers to the set of attributes that appear in the high-level textual description of the image. We define the textual descriptions as “new words” that do not exist in the normal language and get the embeddings via inversion method. We benefit from the recent success of diffusion models and inversion . We adapt diffusion models in our work as a backbone to be inverted and as a generator in image-to-image and text-to-image synthesis Specifically, we propose an efficient and accurate textual inversion based on the attention mechanism, which can quickly learn key features from an image, and a stochastic inversion to maintain the semantic of the content image. We use CLIP image embedding to obtain high-quality initial points, and learn key information in the image through multi-layer cross-attention. Taking an artistic image as a reference, the attention-based inversion module is fed with its CLIP image embedding and then gives its textual embedding. The diffusion models conditioning on the textual embedding can produce new images with the learned style of the reference.

To demonstrate the effectiveness of InST, we conduct comprehensive experiments, applying our method to numerous images of various artists and styles. All of the experiments show that our InST produces outstanding results, generating artistic images that both imitate well the style attributes to a high degree, and achieve content consistent with the input natural images or text descriptions. We demonstrate much improved visual quality and artistic consistency as compared to state-of-the-art approaches. These outcomes demonstrate the generality, precision and adaptability of our method.

Related Work

Image style transfer has been widely studied as a typical mechanism of example guided artistic image generation. Traditional style transfer methods use low-level hand-crafted features to match the patches between content image and style image . In recent years pre-trained deep convolutional neural networks are used to extract the statistical distribution of features which can capture style patterns effectively. Arbitrary style transfer methods use unified models to handle arbitrary inputs by building feed-forward architectures .

Liu et al. learn spatial attention score from both shallow and deep features by an adaptive attention normalization module (AdaAttN). An et al. alleviate content leak by reversible neural flows and an unbiased feature transfer module (ArtFlow). Chen et al. apply internal-external scheme to learn feature statistics (mean and standard deviation) as style priors (IEST). Zhang et al. learn style representation directly from image features via contrastive learning to achieve domain enhanced arbitrary style transfer (CAST). Besides CNN, visual transformer has also been used for style transfer tasks. Wu et al. perform content-guided global style composition by a transformer-driven style composition module (StyleFormer). Deng et al. propose a transformer-based method (StyTr2) to avoid the biased content representation in style transfer by taking long-range dependencies of input images into account. Image style transfer methods mainly focus on learning and transferring colors and brushstrokes, but have difficulty on other artistic creativity factors such as object shape and decorative elements.

Text-to-image synthesis

Text guided synthesis methods can also be used to generate artistic images . CLIPDraw synthesizes artistic images from text by using CLIP encoder to maximize similarity between the textual description and generated drawing. VQGAN-CLIP uses CLIP-guided VQGAN to generate artistic images of various styles from text prompts. Rombach et al. train diffusion models in the latent space to reduce complexity and generate high quality artistic images from texts. Those models only use text guidance to generate an image, without fine-grained content or style control. Some methods add image prompt to increase controllability to the content of the generate image. CLIPstyler transfers an input image to a desired style with a text description by using CLIP loss and PatchCLIP loss. StyleGAN-NADA use CLIP to adaptively train the generator, which can transfer a photo to artistic domain by text description of the target style. Huang et al. propose a diffusion-based artistic image generation approach by utilizing multimodal prompts as guidance to control the classifier free diffusion model. Hertz et al. change an image to artistic style by using the text prompt with style description and injecting the source attention maps. Those methods are still difficult to generate images with complex or special artistic characteristics which cannot be described by normal texts. StyleCLIPDraw jointly optimizes text description and style image for artistic image generation. Liu et al. extract style description from CLIP model by a contrastive training strategy, which enables the network to perform style transfer between content image and a textual style description. These methods utilize the aligned image and text embedding of CLIP to achieve style transfer via narrowing the distance between the generated image and the style image, while we obtain the image representation straight from the artistic image.

Inversion of diffusion models

Inversion of diffusion models is to find a noise map and a conditioning vector corresponding to a generated image. It is a potential way for improving the quality of example guided artistic image generation. However, naively adding noise to an image and then denoising it may yield an image with significantly different content. Choi et al. perform inversion by using noised low-pass filter data from the target image as the basis for the denoising process. Dhariwal et al. invert the deterministic DDIM sampling process in closed form to obtain a latent noise map that will produce a given real image. Ramesh develop a text-conditional image generator based on the diffusion models and the inverted CLIP. The above methods are difficult to generate new instances of a given example while maintaining fidelity. Gal et al. presents a textual inversion method to find a new pseudo-word to describe visual concept of a specific object or artistic style in the embedding space of a fixed text-to-image model. They use optimization-based methods to directly optimize the embedding of the concept. Ruiz et al. implant a subject into the output domain of the text-to-image diffusion model so that it can be synthesized in novel views with a unique identifier. Their inversion method is based on the fine-tuning of diffusion models, which demands high computational resources. Both methods learn concepts from pictures through textual inversion, while they need a small (3-5) image set to depict the concept. The concept they aim to learn is always an object. Our method can learn the corresponding textual embedding from a single image and use it as a condition to guide the generation of artistic images without fine-tuning the generative model.

Method

In this work, we use inversion as the basis of our InST framework and SDM as the generative backbone. Note that our framework is not restricted to a specific generative model. As shown in Figure 3, our method involves pixel space, latent space, and textual space. During training, image xx is the same as image yy. The image embedding of image xx is obtained by the CLIP image encoder and then sent to the attention-based inversion module. By multi-layer cross-attention, the key information of the image embedding is learned. The inversion module gives text embedding vv, which is converted into the standard format of caption conditioning SDMs. Conditioned on the input textual information, the generative model obtains a series of latent codes ztz_{t} through sequence denoise process from to the random noise zTz_{T} and finally gives the latent code zz corresponding to the artistic image. The inversion module is optimized by the simple loss of LDMs computed on the “latent noise” of forward process and the reverse process (see Sec. 3.2). In the inference process, xx is the content image, and yy is the reference image. The textual embedding vv of the reference image yy guides the generative model to generate a new artistic image.

2 Textual Inversion

We aim to get the intermediate representation of a pre-trained text-to-image model for a specific painting. SDMs utilize CLIP text embedding as the condition in text-to-image generation. The CLIP text encoding contains two processes of tokenization and parameterization. An input text is first transformed into a token, which is an index in a pre-defined dictionary, for each word or sub-word. After that, each token is associated with a distinct embedding vector that can be located using an index. We set the concept of a picture as a placeholder “[C]”, and its tokenized corresponding text embedding as a learnable vector v^\hat{v}. [C][C] is in the normal language domain, and v^\hat{v} is in the textual space. By assuming a [C][C] that does not exist in real language, we create a “new word” for a certain artistic image that cannot be expressed in normal language. To obtain v^\hat{v}, we need to design constraints as supervision that relies on a single image. An instinctive way to learn v^\hat{v} is by direct optimization , which is minimizing the LDM loss of a single image:

where yy denotes the artistic image, vθ(y)v_{\theta}(y) is a learnable vector, z∼E(x),ϵ∼N(0,1)z\sim E(x),\epsilon\sim\mathcal{N}(0,1). However, this optimization-based approach is inefficient, and it is difficult to obtain accurate embeddings without overfitting with a single image as training data.

Thanks to CLIP’s aligned latent space of image embedding and text embedding, it provides powerful guidance for our optimization process. We propose a learning method based on multi-layer cross attention. The input artistic image is first sent into the CLIP image encoder and gives image embeddings. By performing multi-layer attention on these image embeddings, the key information of the image can be quickly obtained. The CLIP image encoder τθ\tau_{\theta} projects yy to an image embedding τθ(y)\tau_{\theta}(y). The multi-layer cross attention starts with v0=τθ(y)v_{0}=\tau_{\theta}(y). Then each layer is implementing Attention⁡(Q,K,V)=softmax⁡(QKTd)⋅V\operatorname{Attention}(Q,K,V)=\operatorname{softmax}\left(\frac{QK^{T}}{\sqrt{d}}\right)\cdot V with:

During training, the model is conditioned by the corresponding text embedding only. To avoid overfitting, we apply a dropout strategy in each cross-attention layer which set to 0.05.

Our optimization goal can finally be defined as:

where z∼E(x),ϵ∼N(0,1)z\sim E(x),\epsilon\sim\mathcal{N}(0,1). τθ\tau_{\theta} and ϵθ\epsilon_{\theta} are fixed during training. In this way, v^\hat{v} can be optimized to the target area efficiently.

3 Stochastic Inversion

We observe that in addition to the text description, the random noise controlled by the random seed is also important for the representation of the image. As demonstrated in , the changes of random seed results in obvious changes of visual differences. We define pre-trained text-to-image diffusion model-based image representation into two parts: holistic representation and detail representation. The holistic representation refers to the text conditions, and the detail representation is controlled by the random noise. We define the process from an image to noise maps as an inversion problem, and propose stochastic inversion to preserve the semantics of the content image. We first add random noise to the content image, and then use the denoising U-Net in the diffusion model to predict the noise in the image. The predicted noise is used as the initial input noise during generation to preserve content. Specifically, for each image zz, the stochastic inversion module takes the image latent code z=E(y)z=E(y) as input. Set ztz_{t}, the noisy version of zz, as computable parameters, then ϵt\epsilon_{t} is obtained by:

We illustrate the stochastic inversion in Figure 3.

Experiments

In this section, we provide visual comparisons and applications to demonstrate the effectiveness of our approach.

We retain the original hyper-parameter choices of SDMs. The training process takes about 20 minutes each image on one NVIDIA GeForce RTX3090 with a batch size of 1. The base learning rate was set to 0.001. The synthesis process takes the same time as SDM, which depends on the steps.

1 Comparison with Style Transfer Methods

We compare our method with the state-of-the-arts image style transfer methods, including ArtFlow , AdaAttN , StyleFormer , IEST , StyTr2 and CAST to show the effectiveness of our method. From the results, we can see apparent advantages of our method on transferring the semantics and artistic techniques of the reference images to the content images over traditional style transfer methods. For example, our method can better transfer the shapes of important objects, such as the facial forms and eyes (the 1st, 3rd, 4th and 5th rows), the mountain (the 7th row) and the sun (the 8th row). Our method can capture some special semantics of the reference images and reproduce the visual effects in the results, such as the stars on the background (the 1st row), the flower headwear (the 3rd row) and the roadsters (the 6th row, the cars in the content image are changed to roadsters). Those effects are very difficult for traditional style transfer methods to achieve.

2 Comparison with Text-Guided Methods

We compare our method with textual inversion and SDM guided by a human caption. Following , we measure accuracy and editability by the similarities between the CLIP embeddings of style images and generated images, and the similarities of guide texts and generated images, respectively.

We begin by demonstrating the effectiveness of our attention-based inversion on learning and transferring style. In Figure 5, we show the optimization process of and ours. Our method can quickly optimize to the target text embedding in about 1000 iterations, while usually takes 10 times the iterations of ours due to its simple optimization-based scheme. In Figure 6, we demonstrate our superior generality and editability by giving additional semantic descriptions that do not appear in the reference image. It can be easily observed that our method is more robust to those additional descriptions and is able to generate results that match both the textual description and the reference image. However, loses adaptability to these texts and is not able to depict the specific artistic visual effect. As shown in Table. 1, our method outperforms in both accuracy and editability.

Comparison with SDMs

We compare with the state-of-the-art text-to-image generative model SDM . SDM can generate high-quality images from text descriptions. However, it is difficult to describe the style of a specific painting with only text as a condition, so satisfactory results cannot be obtained. As shown in Figure 6 and Figure 7, our method better captures the unique artistic attributes of the reference images. As shown in Table 1, our method outperforms in accuracy.

3 Ablation Study

As shown in Figure 8(a), the buildings in the content images are turned into trees or mountains without stochastic inversion, while the full model can maintain the content information and reduce the impact of the semantic of the style image.

Hyper-parameter Strength

For image synthesis, the most related hyper-parameter is the strength of changes and its impacts are shown in Figure 8(a). The larger the Strength, the stronger the influence of the style image on the generated result, vice versa, the generated image is closer to the content image

Attention module

We show the ablation study of multi-layer attention Figures 8(b). By interacting with CLIP embedding multiple times, the multi-layer attention helps the learned concept be consistent with the CLIP feature space, such improve editability, as shown in Table. 1.

Dropout

Dropout is added in the linear layer of attention module to prevent overfitting. As shown in Figure 8(b) and Table 1, by dropping the parameters of the latent embeddings, both the accuracy and the editability are improved.

4 User Study

We compare our method with several SOTA image style transfer methods (i.e., ArtFlow , AdaAttN , StyleFormer , IEST , StyTr2 , and CAST ), and text-to-image generation method (i.e., Textual Inversion ). All the baselines are trained using publicly available implementations with default configurations.

For each participant, 26 content-reference pairs are randomly selected and the generated results of ours and one of the other methods are displayed randomly. Participants were suggested that the artistic consistency between the generated image and the reference image was the main metric. Then, they were invited to select the better result of each content-reference pair. Finally, we collect 2,262 votes from 87 participants. The percentage of votes for each approach is shown in Table 2, demonstrating that our method achieves the best visual characteristics transfer results.

Furthermore, we conducted a survey of 60 participants on the preferences of the content image guidance strength and artistic visual effects. In the case of a content image existing, users tend to consider that “To depict the artistic style, the details of the content should be embellished appropriately”. We then invite the participants to rank the factors of their expected visual effect. The average comprehensive score of the options in the sorting question is automatically calculated based on the ranking of the options by all the participants. The higher the score, the higher the comprehensive ranking. The scoring rule is:

where scorescore denotes the average comprehensive score of the options, participantesparticipantes denotes the number of people who complete this question, frequencyfrequency denotes the frequency that the option is selected by users, weightweight denotes the weight which is determined by the option’s ranking. The ranking results (rank by score from highest to lowest): (1) Similar artistic effect on semantic corresponding subjects (s=5.4); (2) With the same paint material (scorescore=3.65); (3) Having similar brushstrokes (scorescore=3.2); (4) Having typical shapes (scorescore=2.65); (5) With the same decorative elements (scorescore=2.1); (6) Sharing the same color (scorescore=1.4).

5 Discussions and Limitations

Although our method can transfer typical colors to some extent, when there is a significant difference between the colors of the content image and the reference image, our method may fail to transfer the color in a one-to-one correspondence semantically. For example, the green hair of the content images in the 1st row of Figure 4 is not transferred into brown. As shown in Figure 9, we employ an additional tone transfer module to align the color of content and reference images. However, we observe that different users have different preferences on whether the colors of the content image should be retained. We believe that the colors of a photograph is crucial, so we choose to respect the tone of the original content image in some conditions.

Conclusion

We introduce a novel example-guided artistic image generation framework called InST, which refers to learning the high-level textual descriptions of a single painting image and then guiding the text-to-image generative model in creating images of specific artistic appearance. We propose an attention-based textual inversion method to invert a painting into the corresponding textual embeddings, which benefits from the aligned text and image feature spaces of CLIP. The extensive experimental results demonstrate that our method achieves superior image-to-image and text-to-image generation results compared with state-of-the-art approaches. Our approach is intended to pave the way for upcoming unique artistic image synthesis tasks.

This work was supported in part by National Key R&D Program of China under no. 2020AAA0106200, by National Natural Science Foundation of China under nos. 61832016, U20B2070, and 62102162, and in part by Beijing Natural Science Foundation under no. L221013.

References