Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided Diffusion

Nisha Huang, Fan Tang, Weiming Dong, Changsheng Xu

Introduction

Many people like to appreciate paintings, but not everyone has the expertise and labor time to create an ideal artwork. Therefore, a tool that can create paintings with high quality and a wide variety from simple inputs would be helpful for novice people to experience art creation easily. In recent years, research on this topic has drawn extensive attention because of its scientific and artistic values.

Existing works use computational algorithms (Wang et al., 2014; Zhang and Yu, 2016; Alvarez et al., 2021) or style transfer approaches (Johnson et al., 2016; Huang and Belongie, 2017; Liu et al., 2021b; Deng et al., 2022) for art creation. The task of style transfer involves transferring the style of an image or a collection of images (a pictorial artwork or the artworks of a painter) to another image (usually a photograph) or video clip. Style transfer (Deng et al., 2020b, 2021; Wei et al., 2022; Zhang et al., 2022) can generate high-quality art images/videos but does not closely resemble the actual paintings. The resulting content depends on the content of the input photograph, hence limiting the controllability of the creation process and the diversity of the results. Image-to-image translation scheme was also used for artistic image generation (Zhu et al., 2017; Kotovenko et al., 2019; Lin et al., 2021). Nevertheless, these methods can only generate results with the style of one or a limited number of artists. Diversified style transfer methods (Wang et al., 2020; Chen et al., 2021b, a) were proposed to generate multiple results from the same input, but all the contents of the results are still the same as the input content image. Recently, some new works used text to control the generation of paintings (Frans et al., 2021; Schaldenbrand et al., 2022; Jain, 2021), based on dual language-image encoders, such as CLIP (Radford et al., 2021). However, the quality and diversity of the paintings produced by the text guidance still need to be improved urgently. Therefore, we strive to create user-controlled, realistic, high-quality, and diverse artworks.

Meanwhile, outstanding works in image generation have been produced (Ramesh et al., 2021; Zhao et al., 2018; Xu et al., 2018; Ruan et al., 2021; Huo and Yoon, 2021; Xue et al., 2022). However, few of them were employed to generate digital paintings. Diffusion models that emerged lately (Gu et al., 2021; Rombach et al., 2021; Dhariwal and Nichol, 2021) achieve state-of-the-art quality and outstanding diversity, and have the potential to create excellent digital artworks. Restricted by the guidance conditions of existing diffusion models (Ho et al., 2020; Song et al., 2021a), generating digital artworks freely is challenging. For example, ADM-G (Dhariwal and Nichol, 2021) can only generate natural images of a certain category based on the category label. Therefore, CLIP (Radford et al., 2021), which enables multimodal prompts as guidance conditions to broaden the application of diffusion models to generate digital artworks, was used. A freshly proposed form of guidance known as classifier-free guidance (Ho and Salimans, 2021) could produce similar results without the use of a separate classifier. Therefore, we combine the classifier-free diffusion model (Ho and Salimans, 2021) and CLIP (Radford et al., 2021) for digital art generation, which has the advantage of high quality and superb diversity.

Taking advantage of the capabilities of guided diffusion models to generate images and the abilities of text-to-image or image-to-image models to handle prompts, guided diffusion is applied to the issue of multimodal-conditional digital artwork synthesis. In this paper, we propose a multimodal guided artwork diffusion (MGAD) model, which adopts CLIP to assist multimodal prompts for generating digital artworks guided by classifier-free diffusion models (Figure 2). We begin by utilizing a fine-tuned unconditional model with 512×512512\times 512 resolution for the diffusion model from OpenAI’s class-conditional ImageNet diffusion model (Dhariwal and Nichol, 2021). Then, we employed a secondary diffusion model (Crowson and AI, 2022) that was previously trained on the Yahoo Flickr Creative Commons 100 million (yfc, 2022) dataset to produce higher-quality outputs. Findings show that the samples (Figure 1) from our model generated pleasing and artistic results. The main contributions of this work are summarized as follows:

We propose MGAD, a method used to refine each transition in the generative process by matching each latent variable with the given multimodal guidance.

We enable the user to control the semantic content of the prompts with respect to the similarity of the results by using CLIP and the classifier-free diffusion model.

The aggregate experimental results show that MGAD outperforms the baseline approaches and achieves excellent results in terms of diversity and quality of the generated digital artworks.

RELATED WORK

Previous research has focused on the combination of natural language and image, including tasks, such as text-guided (Liu et al., 2020; Kim et al., 2021) or image-guided synthesis (Wang et al., 2021; Choi et al., 2021). Consequently, some prominent works for vision and language– representations (Desai and Johnson, 2021; Chen et al., 2020; Li et al., 2020b) have been studied in depth to integrate image–text embedding. The pre-trained text–image embedding model CLIP (Radford et al., 2021) can be efficiently transferred non–trivially to most tasks, generally without any specific dataset training compared with fully supervised baselines. Its representations have been proven robust and comprehensive enough to perform zero-shot classification and various vision-language tasks on different datasets. The combination of CLIP and GANs, which utilize CLIP to guide the optimization of a latent code for a required image manipulated or generated, has emerged in subsequent studies. For instance, StyleCLIP (Patashnik et al., 2021) used CLIP embedding vectors to tune the latent codes. StyleGAN-NADA (Gal et al., 2021) utilized CLIP frame to adjust the zero-shot domain. Our goal is to recognize these descriptions automatically. CLIPstyler (Kwon and Ye, 2022) proposed patchCLIP for transferring semantic texture information on text conditions.

GLIDE (Nichol et al., 2021) and DALL-E 2 (Ramesh et al., 2022) focus on open domain image synthesis. Both of them are implemented with the idea of integrating image generators and joint text-image encoders into their architectures. They all contain pre-trained models with large-scale datasets of numerous text-image pairs. In contrast, we do not intend to spend such expensive training resources and time. We aim to only use CLIP to measure the similarity between the prompts and the generated results to guide the reverse direction. In addition, the above CLIP-based image editing methods only allowed the users to supply a textual description as the style condition. To improve controllability, we attempt to input multiple modality information as control conditions for art image synthesis.

Diffusion Models

Diffusion models (Sohl-Dickstein et al., 2015a), which consist of one forward process (signal to noise) and one reverse process (noise to signal), have been newly demonstrated to produce high-quality images (Sohl-Dickstein et al., 2015b; Song and Ermon, 2019; Ho et al., 2020; Dhariwal and Nichol, 2021). Denoising diffusion probabilistic models (DDPM) (Ho et al., 2020) and score-based generative models (Song and Ermon, 2019; Song et al., 2021b), have recently gained remarkable success in the field of image generation (Ho et al., 2020; Song et al., 2021a, b; Jolicoeur-Martineau et al., 2021). Ablated diffusion model (ADM) (Dhariwal and Nichol, 2021) has demonstrated higher image synthesis quality than variational autoencoders (VAEs) (Razavi et al., 2019), flow-based models (Kingma and Dhariwal, 2018), auto-regressive models (Menick and Kalchbrenner, 2018), and GANs (Goodfellow et al., 2014; Karras et al., 2018, 2020). The generative power of these models (Dhariwal and Nichol, 2021; Ho et al., 2020; Song et al., 2021b) stems from a natural adaptation to the inductive biases of image-like data when their underlying neural skeleton is implemented as a U-Net (Ronneberger et al., 2015). Recently, diffusion models have been explored for conditional generation, such as class-conditional generation (Choi et al., 2021), image-guided synthesis (Dhariwal and Nichol, 2021), text-guided synthesis (Gu et al., 2021; Kim et al., 2021), semantics-guided synthesis (Liu et al., 2021c), and super-resolution (Rombach et al., 2021).

Art Painting Synthesis

Artwork analysis (Deng et al., 2020a; Hall et al., 2015; Deng et al., 2019) and synthesis (Tang et al., 2018; Huang et al., 2022) are the most challenging tasks that enable effective engagement of the public with art, balancing between sophisticated computational/engineering techniques and artistic purposes. Tan et al. (Tan et al., 2017, 2019) proposed ArtGAN, where the label information was propagated back to the generator for more efficient learning. More recently, CLIPDraw (Frans et al., 2021) and StyleCLIPDraw (Schaldenbrand et al., 2022) began generating paintings from randomized Bézier curves that fit a given text and style. In contrast, these two models mainly capture larger features, such as shapes or outlines, rather than fine-grained textures. Yi et al. (Yi et al., 2021) explored the generation of fine art painting that used diffusion models without any guidance.

Inspired by the aforementioned works, we propose a novel MGAD model, which concentrates on synthesizing drawings rather than realistic pictures. We explore whether ADM (Dhariwal and Nichol, 2021) can be guided by text and image prompts to synthesize high-quality and fantasy art paintings rather than only using text prompts, which makes generating results less controllable.

METHODS

Figure 3 shows the overall framework of the proposed MGAD for art painting synthesis. MGAD is a new unified framework that incorporates different modality guidance into a pre-trained classifier-free diffusion model. Our goal is to generate the required art paintings according to one or both images and text prompts to take advantage of multimodal guidance. The multimodal guidance enables the controllable art painting synthesis and realizes complementarity between modes. First, we utilize a pre-trained diffusion model ϵθ\epsilon_{\theta} to transform the noise image x0x_{0} into latent xt0(θ)x_{t_{0}}\left(\theta\right). Second, the target ytary_{tar} leads the diffusion model to generate samples guided by the CLIP loss at the reverse process. Third, the deterministic forward processes are based on DDPM (Ho et al., 2020). For translation among unseen domains, the image generation is also performed by combining multimodal guidance and diffusion models (Dhariwal and Nichol, 2021; Ho and Salimans, 2021).

2. Multimodal Guided Diffusion

Diffusion models (Sohl-Dickstein et al., 2015a) are inspired by non-equilibrium thermodynamics. They define a Markov chain of diffusion steps to add random noise to data slowly and then learn to reverse the diffusion process to construct the desired data samples from noise. Unlike VAEs (Razavi et al., 2019) or flow-based (Kingma and Dhariwal, 2018) models, diffusion models are learned using a fixed procedure, and the latent variable has high dimensionality (same as the original data).

A diffusion model includes one forward and one reverse diffusion process. The forward process is fixed to a Markov chain trained using variational inference, which gradually adds noise to the data. The data distribution is defined as x0∼q(x0)x_{0}\sim q\left(x_{0}\right) and a Markov chain forward process named as qq. In addition, the data generates noised samples from x1x_{1} to xTx_{T}. At each step of the forward process, Gaussian noise is added to the data accordingly by a variance schedule β1,...,βT\beta_{1},...,\beta_{T}:

The training goal is to optimize the negative log likelihood of the usual variational bound:

Forward process samples xt\mathbf{x}_{t} at time step tt:

where αt:=1−βt\alpha_{t}:=1-\beta_{t} and αˉt:=∏s=0tαs\bar{\alpha}_{t}:=\prod_{s=0}^{t}\alpha_{s}. The variance of the noise for an arbitrary time step is defined as 1−αˉt1-\bar{\alpha}_{t} to determine the noise schedule. The form of the diffusion models (Sohl-Dickstein et al., 2015a) can be represented by pθ(x0):=∫pθ(x0:T)dx1:Tp_{\theta}\left(\mathbf{x}_{0}\right):=\int p_{\theta}\left(\mathbf{x}_{0:T}\right)d\mathbf{x}_{1:T}. pθ(x0:T)p_{\theta}\left(\mathbf{x}_{0:T}\right) is the expression of the reverse process that is defined as a learning Gaussian distribution Markov chain that initiates at p(xT)=N(xT;0,I)p\left(\mathbf{x}_{T}\right)=\mathcal{N}\left(\mathbf{x}_{T};\mathbf{0},\mathbf{I}\right).

If the magnitude 1−αt1-\alpha_{t} of the noise added at each step is small enough, then the posterior q(xt−1∣xt)q\left(x_{t-1}\mid x_{t}\right) could be well-approximated by a diagonal Gaussian. Furthermore, if the magnitude 1−α1…αT1-\alpha_{1}\ldots\alpha_{T} of the total noise added throughout the chain is large enough, then xTx_{T} could be well-approximated by N(0,I)\mathcal{N}(0,\mathcal{I}). These properties suggest learning a model pθ(xt−1∣xt)p_{\theta}\left(x_{t-1}\mid x_{t}\right) to approximate the true posterior:

which can be used to produce samples x0∼pθ(x0)x_{0}\sim p_{\theta}\left(x_{0}\right) by starting with Gaussian noise xT∼N(0,I)x_{T}\sim\mathcal{N}(0,\mathcal{I}) and gradually reducing the noise in a sequence of steps xT−1,xT−2,…,x0x_{T-1},x_{T-2},\ldots,x_{0}. To compute this surrogate objective, Ho et al. (Ho et al., 2020) found that predicting ϵ\epsilon worked best, especially when combined with a reweighted loss function:

The DDPM model (Ho et al., 2020) shows how to derive μθ(xt)\mu_{\theta}\left(x_{t}\right) from ϵθ(xt,t)\epsilon_{\theta}\left(x_{t},t\right), and fix Σθ\Sigma_{\theta} to a constant. Results from the DDPM show that they can rapidly sample and achieve better log-likelihoods used parameterization and simplified training objective to improve the learning of Σθ{\Sigma_{\theta}}.

Guided Diffusion. To explicitly incorporate class information into the diffusion process, Dhariwal et al. (Dhariwal and Nichol, 2021) trained a classifier fϕ(y∣xt,t)f_{\phi}\left(y\mid\mathbf{x}_{t},t\right) on noisy image xtx_{t} and use gradients ∇xtlog⁡pϕ(y∣xt)\nabla_{x_{t}}\log p_{\phi}\left(y\mid x_{t}\right) to guide the diffusion sampling process toward the target class label yy. The new resulting perturbed mean μ^θ(xt∣y)\hat{\mu}_{\theta}\left(x_{t}\mid y\right) is given by

where mean μθ(xt∣y)\mu_{\theta}\left(x_{t}\mid y\right) and variance Σθ(xt∣y)\Sigma_{\theta}\left(x_{t}\mid y\right) is perturbed additively by the gradient of the log-probability log⁡pϕ(y∣xt)\log p_{\phi}\left(y\mid x_{t}\right) of a target class yy predicted by a classifier. The ADM and the one with additional classifier guidance (ADM-G) can achieve results that are better than those of state-of-the-art generative models (Brock et al., 2018).

Classifier-free guidance. The disadvantage of classifier guidance is that it needs an additional classifier model, thereby complicating the training process. Ho and Salimans (Ho and Salimans, 2021) presented classifier-free guidance, a methodology for guiding diffusion models that do not demand the training of a separate classifier model. Throughout the training, the tag yy in a class-conditional diffusion model ϵθ(xt∣y)\epsilon_{\theta}\left(x_{t}\mid y\right) is substituted with a null tag ∅\emptyset with a defined likelihood for classifier-free guidance. The output of the model is further extended in the direction of ϵθ(xt∣y)\epsilon_{\theta}\left(x_{t}\mid y\right) and away from ϵθ(xt∣∅)\epsilon_{\theta}\left(x_{t}\mid\emptyset\right) during sampling:

The guidance scale is s≥1s\geq 1. The latent classifier inspired this equation.

where gradient is expressed as a function of the true scores ϵ∗\epsilon^{*}

As shown in Algorithm 1, the amended prediction ϵ^\hat{\epsilon} is then used to steer us towards the multimodal prompts cc:

Overall, classifier-free guidance has two advantages. First, rather than depending on the information of a separate (and possibly smaller) classification model, it enables a single model to exploit its expertise during guiding. Second, when conditioned on information that is difficult to anticipate using a classifier, it simplifies guiding.

3. CLIP-based Multimodal Guidance

CLIP (Radford et al., 2021) was presented to acquire visual concepts with natural language supervision and can provide the similarity scores between texts and images. Several works have used CLIP to steer generative models, such as GANs (Patashnik et al., 2021; Gal et al., 2021; Kwon and Ye, 2022), toward user-defined text prompts. In this paper, we leverage a pre-trained CLIP model for text-driven and image-driven art paintings synthesis. Both prompts are formulated as the cosine similarity of the created paintings’ features to control the semantic content of the generated digital art paintings.

Text prompt ll and image prompt xx are embedded into the joint embedding space. The image encoder EIE_{I} is time-dependent and trained on noisy pictures. EI′E_{I}^{\prime} is the label given to a time-dependent image encoder for noisy pictures. The text guidance function can be defined as:

To define the image guiding function, we employ an image encoder finetuned using denoised painting synthesis, similar to how text guidance is generated. The guidance signal at time step tt is:

The ability to unify image and text guidance simultaneously increases user control flexibility and controllability. Both can be readily included in our pipeline (Fig. 3).

Users may modify the balance between the two by adjusting each modality’s weighting and scale factors.

We replace the classifier with a CLIP model. As illustrated in Algorithm 1, the gradient of the dot product of the multimodal prompts and generated image encoding affect the reverse-process mean:

For MGAD, we employ noised CLIP model, which has been specifically trained to be noise-aware.

EXPERIMENTS

For the diffusion model, we used an unconditional model of resolution 512×512512\times 512 fine-tuned from OpenAI’s class-conditional ImageNet diffusion model (Dhariwal and Nichol, 2021). Additionally, we use a secondary diffusion model (Crowson and AI, 2022) pre-trained on Yahoo Flickr Creative Commons 100 Million (YFCC100m) (yfc, 2022) dataset to achieve a better performance. For the CLIP model, we used ViT-B/32 released by OpenAI for the Vision Transformer (Dosovitskiy et al., 2020). The output size of MGAD is 512×512512\times 512. For sampling, we set w1w_{1}, w2w_{2} and ss to 1.01.0, 1.01.0 and 5,0005,000, respectively. We utilize the U-Net (Ronneberger et al., 2015) architecture based on Wide-ResNet (Zagoruyko and Komodakis, 2016). The model has seven downsampling and seven upsampling layers. The 4×44\times 4 feature is generated from the 512×512512\times 512 input image via one input convolution and fine Resblocks. From the 32×3232\times 32 to 4×44\times 4 resolution, self-attention blocks are added to the Resblocks. The clip model for generating is ViT-B/32. To ensure the quality of the results and to maintain the consistency of the parameters, the diffusion step and the time step used for the experiments in this work are set to 2,0002,000. Remarkable results are generated when the time step is set to 500500 or more. One upsampling step takes approximately 6363 milliseconds on a single A40 GPU. For more hyperparameter settings, please refer to the Supplementary Material.

2. Qualitative Evaluation

To demonstrate the performance of our method, we compare MGAD with SOTA digital painting synthesis works, including VQGAN-CLIP (Rodent, 2022), FuseDream (Liu et al., 2021a), BigSleep (Adverb, 2022), StyleCLIPDraw (Schaldenbrand et al., 2022), CLIPDraw (Frans et al., 2021), and VectorAscent (Jain, 2021). Figure 4 shows the digital art generation results. VQGAN-CLIP (Esser et al., 2021) and StyleCLIPDraw (Schaldenbrand et al., 2022) use two modal prompts. The others are only able to use text prompts. VQGAN-CLIP (Esser et al., 2021), FuseDream (Liu et al., 2021a), and BIGGAN-CLIP (Adverb, 2022) utilize CLIP-loss between the prompts and the generated results to realize latent code optimization. Still, problems are encountered in generating paintings with similar brush strokes and structures. Their results have similar repeated textures in different image locations. VQGAN-CLIP (Esser et al., 2021) can make use of the two modal prompts. However, it may not mimic the painter’s strokes well (e.g., the 4th row in Figures 4a and 4d-4i) during the generation process, and the generated results do not resemble the image prompt very much. FuseDream and BigSleep use pre-trained BigGANs (Brock et al., 2018) and CLIP to achieve high-resolution text-to-image generation. The weights of the generator are frozen; only the latent ZZ vectors are optimized, so they often generate repetitive patterns in the results (e.g., the 6th and 7th rows in Figures 4g-4h). CLIPDraw (Frans et al., 2021) and VectorAscent (Jain, 2021) rely on CLIP because of its aligned text and image encoders and diffvg(Li et al., 2020a), which is a differentiable vector graphics rasterizer used to generate raster images from vector paths. By using gradient ascent, they can optimize for a vector graphic whose rasterization has high similarity with a user-provided caption, backpropagating through CLIP (Radford et al., 2021) and diffvg (Li et al., 2020a) to the vector graphic parameters. Their results have resemblances to the text prompt. However, this type of method can only produce discontinuous outlines that resemble the touch of a watercolor brush or blocks of color, unlike human paintings (e.g., the 5th, 8th and 9th rows in Figure 4). The StyleCLIPDraw (Schaldenbrand et al., 2022) adds a stylized module to the CLIPDraw (Frans et al., 2021), thereby making the stroke styles more diverse, but it still falls short in approaching the realism of the content described by the text prompts (e.g., the 5th and 8th rows in Figure 4).

Compared with baselines, the proposed MGAD can combine the semantic content of visual and language modalities to generate the requested digital artworks. CLIP (Radford et al., 2021) provides a good understanding of the different modal prompts at the level of semantic content. Diffusion models (Dhariwal and Nichol, 2021) can balance high quality and diversity in the generation of images. Therefore, using the MAGD, generating complete and sophisticated artworks with skills that meet the user’s requirements for the content and the texture and style that closely resembles the paintings created by the mentioned artist is possible.

Strong Understanding for Wide Prompts

By observing the content of the prompts in Figure 1, we find that MGAD can meet the prompts that present a variety of requirements. There is a good understanding of the painter, painting style, color, and texture, and can show their differences. Image prompt can generate good results regardless of whether it is a real image or a painting. MGAD’s ability to blend the two modalities and find commonalities between them is outstanding.

3. Quantitative Evaluation

To demonstrate the advantages of our method in terms of the quality and diversity of digital art synthesis, we compare it with the existing methods, namely, VectorAscent (Jain, 2021), FuseDream (Liu et al., 2021a), BigSleep (Adverb, 2022), CLIPDraw (Frans et al., 2021), StyleCLIPDraw (Schaldenbrand et al., 2022) and VQGAN-CLIP (Rodent, 2022) using LPIPS (Zhang et al., 2018), and dHash (Buchner, 2021). User study is also conducted.

LPIPS (Zhang et al., 2018) measures the perceptual similarity between two images. We calculate the LPIPS score between paired images generated from the same prompts for guidance (Table 1). Intuitively, the lower the LPIPS score is, the lesser the similarity of the generated images and the higher the diversity of the generated paintings will be. The results show that our model outperforms baseline models by showing the best LPIPS score and surpassing others by a margin. This finding indicates that the LPIPS score results have almost the same tendencies as the score of the user study.

dHash

dHash (Buchner, 2021) is a hash mapping method based on pixel point level to measure the similarity between images. As an implementation, dHash is nearly identical to aHash, but it performs much better. A value of 0 indicates the same hash and likely a similar picture. The comparison between the different model dHash values is shown in the bottom row of Table 1. Our model scores significantly superior to other models on this evaluation metric. This result demonstrates that MGAD is notable in terms of diversity.

User Study

We conduct a user study to further compare our method. The user study was divided into two groups. The first group compares the models using text as guidance, and the second group uses text and image as guidance. VectorAscent (Jain, 2021), FuseDream (Liu et al., 2021a), BigSleep (Adverb, 2022), and CLIPDraw (Frans et al., 2021) are in the first group, and the other models are in the second group. We invited 8282 users to participate in the first study and invited 8989 users to participate in the second study to evaluate the results of different approaches.

Given the guidance, for each set, we show the result generated by our approach and the output from another randomly selected method for comparison and ask the user to select which digital artwork has better effects. Participants in the first group wrote 2727 questions and collected 2,2142,214 votes. Participants in the second group reported 1515 questions and gathered 1,3351,335 votes. We calculated the percentage of votes where the existing methods are superior to ours and show the statistical results in Table 2. The results show that most people prefer our approach compared with other methods.

4. Ablation Study

The MGS decides how strongly the result should match the multimodal prompts. We compared the generated results for different MGS values to verify the impact of multimodal guidance. As shown in Figure 6, the generated results are more creative and uncertain when using a lower MGS. Meanwhile, using a higher MGS brings the generated results closer to the semantic content of the prompts and the user requirements. With the appropriate scale of multimodal guidance, the results can be made to fit the prompt content while exploiting the creativity of the model. Users can adjust the scaling factor to control how diverse they expect the generated images to be.

Impact of Diffusion Steps

To investigate the effect of the number of diffusion steps on the quality and generation speed of the generated images, we tested the generation at different diffusion steps. The experiments show that the sampling process is faster when the number of diffusion steps is small, although the generated images are unclear. After the diffusion steps gradually increase, the image quality improves, and the sampling time increases. However, as shown in Fig. 7, after a certain number of diffusion steps, the image quality no longer changes significantly. In the trade-off between generation time and generation quality, we usually set the number of diffusion steps to 2000.

Balance between Text and Image Prompts

As shown in Figure 8, we can adjust the degree of guidance of both for the final generated results by setting the values of the weight w1w_{1} for the image prompt and the weight w2w_{2} for the text prompt. During our experiments, we found that if the difference between image prompt and text prompt in semantic content is too large, then it will lead to the generation of unsatisfactory results. For this reason, we propose the adaptive prompt discarding module, which calculates the similarity between image prompt and text prompt and then compares the weights of the two to decide whether one of the modalities should be discarded to obtain a better generation result.

CONCLUSIONS AND FUTURE WORK

In this paper, we propose the MGAD model, which is a digital artwork generation model that combines multimodal prompt guidance with classifier-free diffusion model guidance. To enable MGAD to control the generation of digital artwork with the semantic content of multimodal prompts, we propose multimodal guidance loss to constrain the diversity of models and the content similarity with prompts. Adequate experiments show that our approach can achieve a balance between the quality and diversity of digital artwork generation again. In future work, we will focus more on the disentangling and fusion of multimodal prompts and involving art stroke/model-based techniques to make the model more general, creative, and user friendly.

References