DiffStyler: Controllable Dual Diffusion for Text-Driven Image Stylization

Nisha Huang, Yuxin Zhang, Fan Tang, Chongyang Ma, Haibin Huang, Yong Zhang, Weiming Dong, Changsheng Xu

I Introduction

Image stylization is appealing in practice as it allows amateurs to turn real-world photos into renderings that mimic the style of artwork without requiring any professional skills. Based on the extracted style textures, the style transfer methods migrate the semantic textures of the style images to the content images and generate vivid artwork. They are followed and improved by later works that have shown substantial value in creative visual design.

While the above image stylization methods are capable of delivering attractive and stable results for artwork, they all require users to provide a style image as a reference for stylization, which leads to the following problems. First, providing an appropriate style image itself increases the complexity and redundancy of the user operation. In addition, style images themselves have strong limitations for color and texture. As a result, the user’s need for stylization and artistic creation cannot be adequately expressed through a single image either. In contrast, expressing the user’s artistic needs and aesthetic preferences through text is more in line with the original intention of artistic creation. For illustration, the text conveys arbitrary artists, styles, and artistic movements. This more natural and intuitive way of textual guidance can create more imaginative and unrestricted digital artworks.

Current approaches for text-driven stylization tasks are mainly based on GAN models, which results in limited modeling and generation capabilities between text and style. The GAN-based approaches, which rely on the patch style discriminator, are unable to generate satisfactory results in the case of inappropriate patches to sample. In our investigation, we found that textual information can easily guide the diffusion model for stable stepwise diffusion without relying on additional random style patches. Meanwhile, with the development of diffusion-based methods , they have demonstrated phenomenal results on visual tasks, especially in generating artworks . Therefore, we introduce diffusion into stylization tasks to guide the image global sampling process. This improves the problem that GAN-based methods repeatedly have the same stylized patches at different locations of the generated results thus causing artifacts and failure to highlight the main content.

Moreover, the stylization task requires keeping the content image structure , but the current work of diffusion does not preserve the global content of the input image. There are several works that use the content image as the initial input to the diffusion model. However, the gradual addition of noise, which can be destructive to the image content, cannot achieve similar stylization applications. Specifically, DreamBooth can maintain the content of a particular subject and embed it into the output domain. However, the method requires fine-tuning for each new set of content images. In addition, the forward diffusion entails introducing Gaussian noise to the image during the training phase. This noise gradually transforms the image, making it progressively more featureless and resembling a process of entropy increase. As the noise level increases, the order and regularity within the image weaken, leading to significant changes in the content representation. Therefore it remains challenging to maintain the global content for arbitrary images based on the diffusion model.

To solve the image content preservation problem of the diffusion model, we propose a novel controllable double-diffusion text-driven image stylization method, the working schematic is shown in Fig. 1. Unlike traditional diffusion models for image generation, we make the following three improvements. First, we replace random noise with learnable noise in the free diffusion process of the content image, preserving the main structure of the content image. Second, for each step of the inference process, we adopt dual diffusion architectures so that the stylized results are both explicit in content and abstract in aesthetics. Furthermore, we optimize the simulation of the diffusion process in terms of numerical methods, so that the generation process can guarantee quality while speeding up the sampling speed. Experiments show that our diffusion model-based image content retention and style transfer is effective and achieves excellent results on both quantitative metrics and manual evaluation (see Fig. 2). In summary, our main contributions are as follows:

We present a new dual diffusion-based text-driven image stylization framework that generates output matching the text prompt of the desired style while preserving the main structure of the input content image.

We apply learnable noise based on the content image to overcome the destructive effect of adding random noise to the image content during traditional diffusion.

Numerous experimental and qualitative examples show that DiffStyler outperforms baseline methods and achieves outstanding results with desirable content structures and style patterns.

II Related Work

Image style transfer. Style transfer aims to migrate the style of a painting to a photograph, maintaining its original content. Initially, Gatys et al. find that hierarchical layers in CNNs can be used to extract image content structures and style texture information and propose an optimization-based iterative method to generate stylized images. An increasing number of methods have been developed thereafter to advance the quality of stylization. Arbitrary style transfer in real-time is improved by minimizing perceptual loss which is the combination of feature reconstruction loss as well as the style reconstruction loss. More generally, arbitrary style transfer has gained more attention in recent years.

Huang et al. propose an adaptive instance normalization (AdaIN) to replace the mean and variance of content with that of style. AdaIN is widely adopted in image generation tasks to fuse the content and style features. Deng et al. propose StyTr2 which contains two different transformer encoders to generate domain-specific sequences for content and style, respectively. Zhang et al. present contrastive arbitrary style transfer (CAST) to learn style representation directly from image features by analyzing the similarities and differences between multiple styles and taking the style distribution into account. In conclusion, while traditional image style transfer is stable, it requires the provision of additional style images and the results lack artistry and creativity.

Text-driven image manipulation. In the existing text-guided image synthesis , the encoders for text embedding work as guide conditions for generative models. OpenAI proposes the high-performance text-image embedding model CLIP , based on which several methods manipulate images with textual conditions. StyleCLIP performed attribute manipulation by exploring the learned latent space of StyleGAN . They control the generative process toward a given textual condition by finding the appropriate vector direction. Therefore, StyleGAN-NADA proposed a model modification method using text conditions only and modulated the trained model into a novel domain. Although these methods have been successful in specific domains, they are difficult to apply to arbitrary data.

Based on the above works, Kwon et al. proposed CLIPstyler that contains a patch-wise text-image matching loss with multiview augmentations for realistic texture transfer. Fu et al. propose a contrastive language visual artist (CLVA) that learns to extract visual semantics from style instructions and accomplish LDAST by the patch-wise style discriminator. The stylization effect of CLIPstyler and LDAST is dependent on the patch style discriminator, and the quality of the randomly sampled patch will have a critical impact on the results of the transfer.

Diffusion models for image synthesis. The diffusion model is a generative model learning to generate images by removing noise from random signals in a stepwise manner, which has received a great deal of attention recently. Sohl-Dickstein et al. were the first to implement image generation using the diffusion model, and with continued research in subsequent approaches , the diffusion model is generating high-resolution images with unprecedented quality, often surpassing GANs.

Although denoising diffusion probability models (DDPMs) can produce high-quality samples, they require hundreds to thousands of iterations to produce the final samples. Some previous methods have successfully accelerated DDPMs by adjusting variance schedules (e.g., IDDPMs ) or denoising equations (e.g., DDIMs ). However, these acceleration methods cannot maintain the quality of the samples. Liu et al. proposed the idea that DDPMs should be considered as solving differential equations on manifolds and proposed pseudo-numerical methods for diffusion models (PNDMs) to accelerate the inference process while maintaining the sample quality.

Several large-scale text-image models have recently emerged, such as DALL·E 2 , GLIDE , and Imagen , demonstrating unprecedented image generation results. Notably, some studies have focused on enhancing the control of the progressive inference process , thereby endowing diffusion models with remarkable controllability. Noteworthy advancements in the field of style transfer have been realized through the remarkable text-to-image diffusion models developed by GLIDE and stable diffusion , enabling diffusion-based approaches to achieve superior outcomes.

Since the above works are not designed to maintain the structure of the input image, this does not quite satisfy the problem setting of traditional style transfer tasks . If the above works are applied to an image-to-image application with additional steps, the results will differ significantly from the input image. Kim et al. introduced the method known as DiffusionCLIP, which focuses on performing global changes in images. However, it should be noted that DiffusionCLIP is limited in its applicability to specific domains, and is not designed to operate on arbitrary images. In contrast to existing approaches, our objective is to develop a novel method that enables the stylization of arbitrary images guided by text prompts through the utilization of the diffusion method. By leveraging the power of diffusion, we aim to provide users with precise and controllable stylization capabilities, expanding beyond the constraints of specific domains.

Noise impact. Given the uniqueness of the effect of noise on the diffusion process, many researchers have conducted in-depth works on input noises. For instance, Video ControlNet proposes that the initial noisy samples used in the denoising process have a significant impact on texture synthesis. Since spatial translation or distortion of the noise can change the output semantics, Video ControlNet introduces a well-designed input noise combination scheme. By employing a sliding window approach to maximize the temporal consistency of the generated video. However, discrepancies and inaccuracies in the estimation of the optical flow result in some temporal discontinuities remaining in the generated video.

PFB-Diff utilizes initial random noise and combines progressive feature blending and attention masking mechanisms to generate intermediate noisy images, which are then progressively denoised to obtain the final edited image. One limitation is that the model may struggle to generate desired scenes in background replacement. Another limitation is that the ability to describe desired objects through text remains limited, making personalized editing challenging.

More recently, VideoFusion resolves the per-frame noise into two parts, namely base noise and residual noise, where the base noise is shared by consecutive frames. While VideoFusion sharing the underlying noise between consecutive frames helps to utilize temporal correlation better, it also limits the motion in the generated video and does not apply to video generation with large differences between frames. This allows the image priors of the pre-trained model to be efficiently shared by all frames and thereby facilitate the learning of video data.

In contrast to the aforementioned research, our study investigates the influence of noise on content preservation within the domain of style transfer. In addition, an attempt is made to employ learnable noise within a diffusion modeling framework to facilitate a style stylization task.

III Method

We now formally introduce DiffStyler, a controllable dual diffusion framework for text-driven image stylization. Given a content image x0\textbf{x}_{0} and the target text prompt T\mathcal{T}, DiffStyler can transfer x0\textbf{x}_{0} into a stylized one with the desired style. As illustrated in Fig. 3, DiffStyler consists of three main stages. First, we input the content image x0\textbf{x}_{0} into the diffusion model and let it perform the T1T_{1} steps of free-guided diffusion to obtain learnable noise. Second, the result x^T\hat{\textbf{x}}_{T} at the TT steps of free-guided diffusion is used as the input xT′\textbf{x}_{T}^{\prime} to the dual diffusion model for the reverse sampling process. Third, in the reverse sampling process, we perform the relevant optimization. The process is depicted in Algorithms 1 and 2. We investigate the inherent properties of the diffusion model and propose novel content-preserving learnable noise and a dual diffusion pathway for style/content decoupling, and successfully enable the diffusion model to break through the limitations of the diffusion model in style transfer tasks.

Denoising Diffusion Probabilistic Models. Recently, denoising diffusion probability models (DDPMs) have been shown to generate high-quality images . DDPMs learn the denoising process for parametrized Markov noise images. The isotropic Gaussian noise samples are converted to samples from the training distribution. In the following, we provide a brief overview of DDPMs .

At the core of DDPMs lies the forward noising process, which involves adding Gaussian noise with variance βt∈(0,1)\beta_{t}\in(0,1) at each time step tt to an initial data distribution x0∼q(x0)\textbf{x}_{0}\sim q\left(\textbf{x}_{0}\right). This creates a sequence of images x1,…,xT\textbf{x}_{1},\ldots,\textbf{x}_{T} governed by the following equations:

where q(x1,…,xT∣x0)q\left(\textbf{x}_{1},\ldots,\textbf{x}_{T}\mid\textbf{x}_{0}\right) denotes the joint distribution of the image sequence given the initial image x0\textbf{x}_{0}, and q(xt∣xt−1)q\left(\textbf{x}_{t}\mid\textbf{x}_{t-1}\right) represents the conditional distribution of image xt\textbf{x}_{t} given the previous image xt−1\textbf{x}_{t-1}. It is important to note that as the number of steps TT increases, the final output xT\textbf{x}_{T} tends to approximate an isotropic Gaussian distribution.

A key feature of the forward noising is that each step xt\textbf{x}_{t} may be sampled straight from x0\textbf{x}_{0}, eliminating the need to construct intermediary stages.

where ϵ∼N(0,I)\epsilon\sim\mathcal{N}(\textbf{0},\mathbf{I}) denotes a Gaussian noise sample, αt=1−βt\alpha_{t}=1-\beta_{t} represents the complementary noise level at time tt, and αˉt=∏s=0tαs\bar{\alpha}_{t}=\prod_{s=0}^{t}\alpha_{s} captures the accumulated noise level up to time tt. Importantly, this process can be inverted to generate a fresh sample from the distribution q(x0)q\left(\textbf{x}_{0}\right). The Markovian process is reversed to generate a fresh sample from the distribution q(x0)q\left(\textbf{x}_{0}\right). The posteriors, q(xt−1∣xt)q\left(\textbf{x}_{t-1}\mid\textbf{x}_{t}\right), which were demonstrated to also be Gaussian distributions . They are sampled beginning from a Gaussian noise sample, xT∼N(0,I)\textbf{x}_{T}\sim\mathcal{N}(\textbf{0},\mathbf{I}), to produce a reverse sequence. q(xt−1∣xt)q(\textbf{x}_{t-1}\mid\textbf{x}_{t}) is determined by the unknown data distribution q(x0)q(\textbf{x}_{0}).

To predict the mean and covariance of xt−1\textbf{x}_{t-1} using the input xt\textbf{x}_{t}, a deep neural network pθp_{\theta} is employed. This network enables the sampling of xt−1\textbf{x}_{t-1} from a normal distribution parameterized by these statistics:

Since it is very difficult to infer μθ(xt,t)\mu_{\theta}\left(\textbf{x}_{t},t\right) directly, we refer to and first compute the noise prediction ϵθ(xt,t)\epsilon_{\theta}\left(\textbf{x}_{t},t\right) and then substitute it into Bayes’ theorem (refer to Equation 2) to derive:

Pseudo numerical methods for DDPM. Diffusion models convert Gaussian data to images iteratively. They require many iterations for high-quality samples, slowing large sample generation. Numerical methods have limitations within a limited range. Another problem arises when using numerical methods with the diffusion model equation. The neural network and the equation are well-defined only within a limited range. To address this, a reverse process is used to calculate the derivative of the generated data.

The diffusion model equation is unbounded in most cases, which means that the generation process can produce samples far away from the well-defined area, introducing new errors. The related differential equations of the diffusion model may be derived directly and self-consistently to provide a theoretical relationship between diffusion processes and numerical methods:

The equation of the numerical method is redefined as:

Here, ff is Equation 5. It has been experimentally demonstrated that the pseudo-linear multi-step method is the most efficient method for the diffusion model with similar generation quality.

Learnable noise. The diffusion model is inspired by nonequilibrium thermodynamics. They define a Markov chain of diffusion steps that adds random noise slowly to the data and then learns to reverse the diffusion process by reconstructing the desired data samples from the noise. In previous diffusion work, either a random Gaussian noise image or an input image superimposed with a corresponding amount of random Gaussian noise is usually used as the starting image xT′\textbf{x}_{T}^{\prime} for the sampling process. This operation does not protect the content structure of the input image, making the generated image differ significantly from the input image in terms of content, as shown in the results of stable diffusion in Fig. 4. Despite the remarkable progress of diffusion models for image generation, the destructive nature of the stochastic noising process and the random nature of the reverse denoising process make it difficult to preserve the content of the original image, leading that style transfer by diffusion methods still not well explored. In this context, we investigate and exploit the inherent properties of the diffusion model. Our findings indicate that the final denoising outcome of xT\textbf{x}_{T} is achieved by performing TT steps of free diffusion on the input image without any textual guidance. Specifically, we introduce a zero embedding as a guiding condition for the denoising network, represented by

The denoising network ϵθ(x0)\epsilon_{\theta}(\textbf{x}_{0}) represents the parametrized conditional model. Typically, the number of forward and sampling process steps in diffusion methods are equal. However, we discover that increasing the number of free diffusion steps T1T_{1} beyond the number of reverse diffusion steps TT better preserves the content structure of the input image (as observed in Fig. 9). Therefore, we adopt x^T\hat{\textbf{x}}_{T} as the initial image xT′\textbf{x}_{T}^{\prime} for the sampling process in our experiments, as shown by the pink arrow in pipeline Fig. 3. T1T_{1} and TT are set to 150150 and 5050, respectively. This approach ensures enhanced preservation of the content structure of the input image, rendering it suitable for image stylization tasks. This process is summarized in Algorithms 1 and 2.

Basic architecture of DiffStyler. We found that if the diffusion model is trained only on the natural image dataset, the results will be close to real images and lack artistic appearance. If the diffusion model is trained only on the artistic image training set, the results will be too abstract to preserve the input content. Therefore, we used light denoising networks ϵθ1\boldsymbol{\epsilon}_{\theta 1} and ϵθ2\boldsymbol{\epsilon}_{\theta 2} trained on the Conceptual 12M and WikiArt datasets, respectively. Then, we perform sampling using the following linear combination of the dual channel score estimates:

where 0<w<10<w<1. We utilize it for each step of the inference process, resulting in stylized outputs that are both explicit in content and abstract in aesthetics.

III-B Network Optimization

Instruction loss. We leverage a pre-trained ViT-B/16 CLIP model to stylize the content image according to the text prompt. The cosine distance between the CLIP embedding of the transferred image xt\textbf{x}_{t} during the diffusion process and the CLIP embedding of the text prompt T\mathcal{T} may be used to specify the CLIP-based loss, or Linst\mathcal{L}_{inst}. We define the language guidance function using the cosine distance, which measures the similarity between the embeddings ExtE_{\textbf{x}_{t}} and ETE_{\mathcal{T}}. The text guidance function can be defined as:

where DCLIP\mathcal{D}_{CLIP} is the cosine distance of the CLIP embeddings.

Content perceptual loss. To further enhance the alignment with the content of input images, we use feature maps extracted to compute the content loss and patchwise contrastive content loss :

where ϕi(⋅)\phi_{i}(\cdot) denotes features extracted from the ii-th layer in a pre-trained VGG19 and Nl{N_{l}} is the number of layers.

Aesthetic loss. We also employ an aesthetic loss to make the model’s representation of style more consistent with human preferences. We leverage a model to fit and inference code for CLIP aesthetic regressions RR trained on Simulacra Aesthetic Captions . It is a dataset of over 238,000 synthetic images generated by diffusion models as well as human ratings. We use it as a scorer to evaluate the results generated by DiffStyler, and return the weighted aesthetic exploring loss for optimization:

Laes\mathcal{L}_{aes} improves the overall visual quality.

Total variation loss. Total variation loss (Ltv\mathcal{L}_{tv}) serves as a regularization technique extensively utilized in image processing and computer vision applications. Its primary objective is to promote the presence of smooth transitions and minimize noise by penalizing sudden alterations in intensity or gradients. The mathematical representation for the Ltv\mathcal{L}_{tv} is defined as:

In the equation, the symbol II corresponds to the image under consideration, while Ii,jI_{i,j} denotes the pixel intensity located at the spatial coordinates (i,j)(i,j). This loss computation involves the absolute differences between neighboring pixel intensities in both the horizontal and vertical directions. By summing up these differences across the entire image, the Ltv\mathcal{L}_{tv} encourages smoothness by discouraging abrupt variations in pixel values.

As a consequence, the total loss is formulated as follows:

We set λd\lambda_{d}, λc1\lambda_{c1}, λc2\lambda_{c2}, λaes\lambda_{aes}, and λtv\lambda_{tv} to 5050, 3.03.0, 1.01.0, 1010, and 8080, respectively. The impacts of loss weights are discussed in Section IV-E.

IV Experiments

For the CLIP model, we use ViT-B/16 as the vision transformer . For the diffusion model, we use the primary model of a resolution 256×256256\times 256 trained on the Conceptual 12M dataset. In addition, we use a pre-trained model on the WikiArt dataset as a secondary diffusion model to enable a superior artistic representation of the entire model. Our approach to the combination of the two pre-trained models can be extended to the combination of any diffusion model according to the demands of the application scenario. This results in significant savings in training resources. The output resolution of DiffStyler is 256×256256\times 256. The U-Net architecture we utilize is based on Wide-ResNet . To ensure the quality of the results and to maintain consistency of the parameters, the diffusion step and the time step used for the experiments in our work are set to 5050.

IV-B Computing Consumptions

DiffStyler takes 1818 seconds to generate a 256×256256\times 256 image on a single NVIDIA GeForce RTX 30903090. It takes 1414 seconds with a single channel. This is a moderate amount of time compared to related work. DiffusionCLIP and Stable diffusion are two other diffusion-based methods that take significantly longer to reason about compared to our model. This difference is because they require longer inference steps to obtain relatively desirable results. Although there is still a gap between our current speed compared to the non-denoising method , given the continuous progress in accelerating the denoising inference process, it can be expected that the efficiency of the current inference process will be further improved in the foreseeable future.

The inference process for one image consumes about 4 GB (6 GB and 16 GB for DiffusionCLIP and Stable Diffusion , respectively), which is much smaller than the general diffusion model . The WikiArt model and the CC12M model, trained using 4 NVIDIA GeForce RTX 3090s, required approximately 40 minutes and 50 hours of training time, respectively.

IV-C Qualitative Evaluation

We compare DiffStyler with SOTA text-driven image stylization methods including LDAST , CLIPstyler , DiffusionCLIP , and Stable Diffusion . There are two versions of CLIPstyler. CLIPstyler (opti) requires real-time optimization on each content and each text. CLIPstyler (fast) requires real-time optimization on each text. We can observe that LDAST has relatively similar style performance results in different text-driven cases and is less artistic. CLIPstyler (opti) texture patches used to convey style are substantially corrupting the content of the input image, e.g. (the 3rd row) the dog’s eyes are shifted to other locations in the result. Both CLIPstyler (opti) and CLIPstyler (fast) are more focused on details and lack macro lines, color blocks, and textures. DiffusionCLIP and Stable Diffusion do not express style as prominently (e.g., the 1st, 2nd, and 4th rows) and result in a severe content loss (e.g., the 3rd row) under some textual conditions.

Comparison with style transfer

We indirectly compare our results with existing style transfer methods, by comparing the proposed DiffStyler with the “retrieval-stylization” baselines (see Fig. 5). The text retrieved art images are firstly used as style images and then stylized using the state-of-the-art available artistic style transfer methods, including Gatys et al. , AdaIN , AdaAttN , MCCNet , ArtFlow+AdaIN , StyTr2 , CAST , CAP-VSTNet , and UCAST .

The results of earlier baseline models demonstrate that their portrayal of the intended texture style was limited, focusing only on color variations. And the content information is influenced by the content of the style image. Recent advances in traditional style transfer methods show promise in capturing richer, more intricate textural nuances while better maintaining the structural integrity of the source content. Nevertheless, these state-of-the-art methodologies appear to struggle when the defining brushstrokes or stylistic elements within the reference artistic image are less prominently discernible, as exemplified by the inferior second-row outputs relative to other approaches.

In contrast, our method yields results with more intricate brushstrokes that closely resemble real art paintings. The major advantage of the diffusion approach lies in its capability to handle complex and non-stationary stylistic patterns. Unlike some general deep learning stylization methods that struggle with intricate or rapidly changing artistic styles, diffusion-based methods can effectively adapt to such variations. Diffusion methods are effective in preserving fine details and structural information and flexible in achieving desired artistic effects compared to general deep learning methods. These advantages make diffusion-based methods well-suited for producing high-quality, faithfully stylized images that capture the essence of various artistic styles.

Content modifications

DiffStyler’s dual diffusion is concerned with decoupling the representation of content and style. As a result, we can create more realistic, fine-grained structures in a variety of styles. As shown in Fig. 4, our approach maintains better overall content compared to other approaches (e.g., the 1st1^{st} and 3rd3^{rd} rows of Fig. 4). In addition, to further verify the effectiveness of our method in content preservation, we also compare it with traditional style transfer methods. As can be seen in Fig. 5, DIffStyler can preserve both the detailed content and has a very harmonious overall effect. Balancing content and style preservation can be challenging for image style transfer. a significant advantage of DiffStyler is that there are no obvious artificial traces of content preservation of other traditional style transfer methods. Finally, we experimentally found that the content modification is related to the chosen style, and is specified in Section IV-F.

IV-D Quantitative Evaluation

For quantitative evaluation, we randomly selected 2020 content images and defined text prompts containing 2525 styles and artistic movement descriptions, generating 500500 stylized images for each method.

To measure the correspondence between text prompts and stylized textures, we computed the CLIP score, which is the cosine similarity between the target texts derived from the CLIP encoder. Since we employ the ViT-B/16 CLIP model throughout the inference phase, we compute the CLIP score using the ViT-B/32 CLIP model for fairness.

The results (the first row of Table I) show that our method achieves the highest score of 0.2869. Stable diffusion follows closely with a score of 0.2773, marking the second-best result. This indicates that our approach demonstrates superior capability in expressing textual requirements through stylized image outputs.

LPIPS metric

The learned perceptual image patch similarity (LPIPS) metric is widely used to measure the difference between two images based on learned perceptual image patch similarity, which aligns more closely with human perception compared to traditional methods. The lower the LPIPS value, the higher the similarity between the images.

Analyzing the second row of Table I, it can be observed that both the CLIPstyler (fast) and the DiffStyler methods outperform the other approaches in terms of content retention, as indicated by their lower LPIPS values. The proposed methods effectively preserve the content of the input images while applying the desired style. This performance is desirable, as it ensures that the stylized images maintain the key features and characteristics of the original content.

Content loss

To perform a quantitative evaluation of our approach, we leverage content loss alongside establishing the state-of-the-art stylization techniques employed for test ratings. Table II presents the reported scores for both our proposed DiffStyler method and the baseline models. Notably, our approach showcases several advantages in terms of content preservation, even when compared to traditional stylization methods. The employed learnable noise methods successfully retain the essential content of the input images while effectively applying the desired style.

User study

Since the evaluation of artistic aspects is relatively subjective, we conducted a user perception assessment to further compare our methods. We compared DiffStyler with several state-of-the-art text-guided style transfer methods (i.e., including LDAST , CLIPstyler , DiffusionCLIP , and Stable diffusion ) and multimodal driven image generation method MGAD . All baselines we used to perform inference were publicly available implementations with default configurations. For each participant, 2828 input-output pairs were randomly selected. Participants were requested to rate the following metrics on a 1−51-5 scale based on the stylized results: (1) quality of stylized results, (2) content preservation, and (3) consistency of stylized results with text prompts. Finally, we collected 67206720 scored results from 8080 participants. Table III indicates that DiffStyler scores are high for all three evaluation metrics, which indicates that our findings excel at the level of text-guided stylization.

IV-E Ablation Study

The guidance scale ww of the diffusion model determines how strongly each of the two-path diffusion guides the final stylization outputs. As shown in Fig. 6, when the CC12M model is not utilized and only the WikiArt model is employed (as shown in column 4), the stylization effect closely aligns with the text guidance. However, this approach results in significant content deformation while expressing the desired style. On the other hand, when solely relying on the CC12M model without the WikiArt model (as shown in column 5), the generated image retains clearer content and exhibits a more pronounced style expression. Nevertheless, the generated image, when solely relying on the CC12M model without the WikiArt model, deviates from the style explicitly specified in the textual guidelines. Additionally, it displays issues such as over-saturated colors, unnatural textures, and noticeable artifacts. Furthermore, columns 6 and 7 of Fig. 6 highlight that the impact of the loss on image generation outcomes cannot be equated with the impact of the U-Net model alone. These results demonstrate the importance of employing the dual diffusion model architecture.

Loss terms

We ablate the different loss terms in our objective by qualitatively comparing our results when training with our full objective (Equation 15) and with a specific loss removed. The results are shown in Fig. 7. We set up four control groups for ablation experiments to illustrate the effectiveness of loss. λd\lambda_{d} is used as a control variable for Linst \mathcal{L}_{\text{inst }}, and the generated results are more consistent with the text as λd\lambda_{d} increases. λc1\lambda_{\text{c1}} is the control variable of Lc\mathcal{L}_{\text{c}}, which makes the overall features of the generated image more consistent with the content map. λc2\lambda_{\text{c2}} is the control variable of Lcpatch\mathcal{L}_{\text{c}_{patch}} for better performance of the detailed features. As shown in Table IV, without Lc\mathcal{L}_{\text{c}} and Lcpatch\mathcal{L}_{\text{c}_{patch}}, the LPIPS increased by 0.0880.088 and 0.0160.016. λaes\lambda_{aes} makes the results more consistent with human aesthetic preferences. The Ltv\mathcal{L}_{tv}, which is common in image generation work, can remove artifacts and improve smoothing as λtv\lambda_{tv} increases. Thus, when we use all the proposed loss functions, we can obtain better-stylized results in the perceptual domain.

Diffusion steps

The number of steps determines the quality of the generation and the length of the generation time. We found experimentally (see Fig. 8) that the number of inference steps TT is set to 5050, and we can obtain good results. When the number of inference steps is larger than 5050, the result changes are no longer significant and converge to the upper mass limit. To maintain the structure of the input content, we find it beneficial when the number of free diffusion steps T1T_{1} of learnable noise is more than the number of reverse sampling steps TT. Ablation study results (Fig. 9) show that at T=50T=50. As shown in Table IV, LPIPS increased by 0.1690.169 in the absence of learnable noise. The content of the stylized results is better maintained as T1T_{1} increases gradually, and the image quality no longer changes significantly after a certain number of diffusion steps. To ensure the quality of the results and the efficiency of the generation process, the number of free diffusion steps T1T_{1} used in the experiments of this work is 150150 and the number of inference process steps TT is 5050.

Effect of the CLIP Guidance Scale

CLIP guidance scale λd\lambda_{d} determines how consistently the results match the text prompts. We compared the generated results by taking different λd\lambda_{d} values to verify the impact of text guidance. As shown in Fig. 10, the generated results are more consistent with the input content image when using a smaller λd\lambda_{d}. As λd\lambda_{d} is gradually increased, the generated results are brought closer to the semantic content of the prompt and the user’s requirements. However, when the value of λd\lambda_{d} is too large, the discrepancy between the stylized result and the content of the input image increases, which will cause the stylization to fail. Therefore, with an appropriate CLIP guidance scale value, the result can be made to match the text prompt content while keeping the content of the input image.

The mixed dataset results

As shown in Fig. 11, we conducted ablation experiments on a mixed dataset. The inclusion of both the CC12M dataset, primarily composed of nature images, and the WikiArt dataset, consisting of art images, introduced a substantial scale disparity between the two datasets. Consequently, combining these datasets resulted in an imbalance in the expression of artistic features, wherein the art features were inadequately represented, and the potential of the art image dataset remained underutilized. This challenge was further exacerbated by the limited availability of online art images, hindering the comprehensive representation of diverse art features, as demonstrated by the aforementioned findings in our research. To tackle this issue, we propose a two-tier architectural model that aims to overcome the aforementioned limitations and improve the incorporation of art features in the stylization process.

IV-F Discussion

Further experiments and comparisons revealed that the modification of the content is related to the chosen style. As Fig. 12 shows, the content of the input image in the traditional styling task remains influenced by the style image, and the content remains influenced by the input text in the text-guided styling task. Compared with the SOTA image/text-guided style transfer approaches , the dual diffusion pathway in DiffStyler focuses on the decoupling of content and style representation. As a result, we can generate more natural and fine-grained structures in different styles that are significantly better than other styles (e.g., castle windows and outlines). For a long time, balancing content and style preservation has been challenging for image style transfer. Nevertheless, with text input, the user can further improve the results by changing the text prompt.

IV-G Failure cases

While our approach demonstrates high-quality results for text-guided image stylization, it is important to acknowledge its limitations. As illustrated in Fig. 13, we present several instances where our model fails to produce satisfactory results. In the first row, there is a partial disappearance of the cat’s fur; in the second row, the distinction between “river”, “land” and “mountain” becomes less apparent; and in the third row, the windows of the house vanish. We have observed that the generated results primarily emphasize the elements mentioned in the text, potentially causing other elements to be overlooked and subsequently disappear.

In conclusion, these failure cases highlight the potential drawback of our approach, which relies on textual guidance. This reliance may lead to a limited focus on elements that are not explicitly mentioned or implicitly referenced. Consequently, important visual details or contextual relationships may be omitted, distorted, or inadequately represented. To address these limitations, it is crucial to enhance the model’s capacity to capture and incorporate relevant contextual information beyond explicit textual guidance. This improvement would enable a more comprehensive and faithful stylization process.

V Conclusions and Future Work

In this work, we propose a novel controllable dual diffusion framework, namely DiffStyler, for text-driven image stylization. Our method contains a dual diffusion processing architecture to achieve a trade-off between abstraction and realism of the stylization results. To eliminate the destructive effect of noise on the content of the input image, we use the results of the learnable diffusion process based on the content image to replace the random noise. As the novel baseline for single text-driven style transfer using a diffusion model, the DiffStyler model has a highly distinguished ability to represent artistic style and provides new insights into the challenging problem of style transfer. At present, the sample-time speed of our method is not as fast as some GAN-based approaches. Performing diffusion processes in latent space to speed up the computation will be a possible future research direction. Although we present high-quality results under a single text condition, our method is not without limitations. In particular, the input content may not be fully preserved in the results (see Section IV-G). We plan to investigate more fine-grained control over stylization results in the future.

In addition, DiffStyler has not been able to implement video styling at this time, so we consider it very promising to extend the approach further into the video data domain. We envision that possible future implementations are as follows: one approach is that noise can be used as crucial information for generating each frame in a video to achieve spatio-temporal consistency in video stylization. Another approach is that more realistic and high-quality video reconstruction can be achieved by introducing noise constraints in the video reconstruction process. The potential for utilizing our approach to enhance the visual aesthetics and artistic expression of moving images is an exciting direction for future exploration.

References