ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, Wangmeng Zuo

Introduction

Recently, large-scale diffusion models have demonstrated impressive superiority in text-to-image generation. By training with billions of image-text pairs, large text-to-image diffusion models have exhibited excellent semantic understanding ability, and generate diverse and photo-realistic images being accordant to the given text prompts. Owning to their unprecedentedly creative capabilities, these models have been applied to various tasks, such as image editing , data augmentation , and even artistic creation .

However, despite diverse and general generation, users may expect to create imaginary instantiates with undescribable personalized concepts , e.g., “corgi” in Fig. 1. To this end, many recent studies have been conducted for customized text-to-image generation , which aims to learn a specific concept from a small set of user-provided images (e.g., 3∼\sim5 images). Then, users can flexibly compose the learned concepts into new scenes, e.g., A S* wearing sunglasses in Fig. 1. Given a small image set depicting the target concept, Textual Inversion learned a new pseudo-word (i.e., S*) in the well-editable textual word embedding space of text encoder to represent the user-defined concept. DreamBooth finetuned the entire diffusion model to accurately align the target concept with a unique identifier. Custom Diffusion balanced the fidelity and memory by selectively finetuning K, V mapping parameters in cross attention layers.

Albeit flexible generation has been achieved by , the computational efficiency remains a challenge to obtain the textual embedding of a visual concept. Existing methods usually adopt the per-concept optimization formulation, which requires several or tens of minutes to learn a single concept. As shown in Fig. 1, the Custom Diffusion , which is among the fastest existing algorithms, still takes around 6 minutes to learn one concept, which is infeasible for online applications. In contrast, in GAN inversion, many efficient learning-based methods have been proposed to accelerate the optimization process. An encoder can be trained to infer the latent codes, which only needs one step forward inference.

Driven by the above analysis, we propose a learning-based encoder for Encoding visuaL concepts Into Textual Embeddings, termed as ELITE. As shown in Fig. 2, our ELITE adopts a pre-trained CLIP image encoder for feature extraction, followed by a global mapping network and a local mapping network to encode visual concepts into textual embeddings. Firstly, we train a global mapping network to map the CLIP image features into the textual word embedding space of the CLIP text encoder, which has superior editing capacity . Since a given image contains both the subject and irrelevant disturbances, encoding them as a single word embedding severely degrades the editability of subject concept. Thus, we propose to separately learn them with a well-editable primary word and several auxiliary words. Using the hierarchical features from CLIP intermediate layers, the word learned from the deepest features naturally links to the primary concept (i.e., the subject), while auxiliary words learned from other features describe the irrelevant disturbances (as shown in Fig. 5). When deploying to customized generation, we only use the primary word to avoid editability degradation from auxiliary words.

Usually, a visual concept is worth more than one word, and describing it with a single word may result in the inconsistency of local details . For higher fidelity of the learned concept without sacrificing its editability, we further propose a local mapping network to inject finer details. From Fig. 2, our local mapping network encodes the CLIP features into the textual feature space (i.e., the output space of the text encoder). Compared with the textual word embeddings learned by global mapping network, the textual feature embeddings focus on the local details of each patch in the given image. Then, the obtained textual feature embeddings are injected through additional cross attention layers, and the output feature is fused with the global part to improve the local details. Experiments show that our ELITE can encode the target concept efficiently and faithfully, while keeping control and editing abilities. The contributions of this work are summarized as follows:

We propose a learning-based encoder, namely ELITE, for fast and accurate customized text-to-image generation. It adopts a global and a local mapping networks to encode visual concepts into textual embeddings.

Multi-layer features are adopted in global mapping to learn a well-editable primary word embedding, while the local mapping improves the consistency of details without sacrificing editability.

Experimental results show that our ELITE can faithfully recover the target concept with higher visual fidelity, and enable more robust editing.

Related Work

Deep generative models have achieved tremendous success on text-conditioned image generation and have recently attracted intensive attention. They can be categorized into three groups: GAN-based, VAE-based, and diffusion-based models. Albeit GAN-based and VAE-based models can synthesize images with promising quality and diversity, they still cannot match user descriptions very well. Recently, diffusion models have shown unprecedentedly high-quality and controllable imaginary generation and been broadly applied to text-to-image generation . By training with massive corpora, these large text-to-image diffusion models, such as DALLE-2 , Imagen , and Stable Diffusion have demonstrated excellent semantic understanding, and can generate diverse and photo-realistic images according to a given text prompt. However, despite the superior performance on general synthesis, they still struggle to express the specific or user-defined concepts, e.g., “corgi” in Fig. 1. Our method focuses on making pre-trained diffusion models to learn these new concepts efficiently.

2 GAN Inversion

GAN inversion refers to projecting real images into latent codes so that images can be faithfully reconstructed and edited with pre-trained GAN models . Generally speaking, there are two types of GAN inversion algorithms in the literature: i) optimization-based: directly optimize latent code to minimize the reconstruction error , and ii) encoder-based: train an encoder to invert an image into latent space . The optimization-based methods usually require hundreds of iterations to obtain promising results, while encoder-based methods greatly accelerate this process via one feed-forward pass only. To improve image fidelity without compromising editability, HFGI embeds the omitted information into high-rate features. Similarly, our ELITE adopts a local mapping network that encodes the concept images into textual feature space to improve details consistency.

3 Diffusion-based Inversion

The inversion of text-to-image diffusion models can be performed in two types of latent spaces: the Textual Word Embedding (TWE) space of text encoder or the Image-based Noise Map (INM) space . INM-based inversion methods, such as DDIM or Null Text find the initial noise to reconstruct image faithfully, while it suffers from a degraded editing ability. In contrast, the TWE space has shown superior editing capacity, which is well suitable for textual inversion and customized generation . For example, Textual Inversion and DreamArtist optimize the embedding of new “words” using a few user-provided images to recover the target concept. DreamBooth finetunes the entire text-to-image model to learn high-fidelity new concept with a unique identifier. To improve the computation efficiency, Custom Diffusion only updates the key and value mapping parameters in cross attention layers with better performance. Albeit flexible generation has been achieved, existing methods still require several or tens of minutes to learn a single concept.

In this work, we choose the TWE space as target for inversion, while proposing a learning-based encoder ELITE for fast and accurate customized text-to-image generation. With the proposed global and local mapping networks, our method can learn the new concept quickly and faithfully with one single image.

Proposed Method

Given a pretrained text-to-image model ϵθ\epsilon_{\theta} and an image xx indicating the target concept (usually an object), customized text-to-image generation aims to learn a pseudo-word (S*) in word embedding space to describe the concept faithfully, while keeping the editability. To achieve fast and accurate customized text-to-image generation, we propose an encoder ELITE to encode the visual concept into textual embeddings. As illustrated in Fig. 2, our ELITE first adopts a global mapping network to encode visual concepts into the textual word embedding space. The obtained word embeddings can be composed with texts flexibly for customized generation. To address the information loss in word embedding, we further propose a local mapping network to encode the visual concept into textual feature space to improve the consistency of the local details. In the following, we begin by presenting an overview of the text-to-image model utilized in our approach (Sec. 3.1). Then, we will introduce the details of the proposed global mapping network (Sec. 3.2) and local mapping network(Sec. 3.3).

In this work, we employ the Stable Diffusion as our text-to-image model, which is trained on large-scale data and consists of two components. First, the autoencoder (E(⋅)\mathcal{E}(\cdot), D(⋅)\mathcal{D}(\cdot)) is trained to map an image xx to a lower dimensional latent space by the encoder z=E(x)z=\mathcal{E}(x). While the decoder D(⋅)\mathcal{D}(\cdot) learns to map the latent code back to the image so that D(E(x))≈xD(\mathcal{E}(x))\approx x. Then, the conditional diffusion model ϵθ(⋅)\epsilon_{\theta}(\cdot) is trained on the latent space to generate latent codes based on text condition yy. We simply adopt the mean-squared loss to train the diffusion model:

where ϵ\epsilon denotes the unscaled noise, tt is the time step, ztz_{t} is the latent noise at time tt, and τθ(⋅)\tau_{\theta}(\cdot) represents the pretrained CLIP text encoder . During inference, a random Gaussian noise zTz_{T} is iteratively denoised to z0z_{0}, and the final image is obtained through the decoder x′=D(z0)x^{\prime}=\mathcal{D}(z_{0}).

To incorporate text information in the process of image generation, cross attention is adopted in Stable Diffusion. Specifically, the latent image feature ff and text feature τθ(y)\tau_{\theta}(y) are first transformed by the projection layers to obtain the query Q=WQ⋅fQ=W_{Q}\cdot f, key K=WK⋅τθ(y)K=W_{K}\cdot\tau_{\theta}(y) and value V=WV⋅τθ(y)V=W_{V}\cdot\tau_{\theta}(y). WQW_{Q}, WKW_{K}, and WVW_{V} are weight parameters of query, key, and value projection layers, respectively. Attention is conducted by a weighted sum over value features:

where d′d^{\prime} is the output dimension of key and query features. The latent image feature is then updated with the output of the attention block.

2 Global Mapping

Following , we choose the textual word embedding space of CLIP text encoder as the target for inversion. To improve the computation efficiency, we propose a global mapping network that encodes the given concept image into word embeddings directly. As illustrated in Fig. 3(a), to facilitate the embedding learning, the pretrained CLIP image encoder ψθ(⋅)\psi_{\theta}(\cdot) is adopted as feature extractor, and our global mapping network Mg(⋅)M^{g}(\cdot) projects the CLIP features as word embeddings vv:

Since the image xx contains both the desired subject and irrelevant disturbances, encoding them into one word (i.e., N=1N=1) results in an entangled word embedding vv with poor editability. To obtain a more informative and editable word embedding, we adopt a multi-layer approach to learn NN (NN >11) words from the image xx separately. Specifically, we select NN layers from the CLIP image encoder, and each layer ψθLi(⋅)\psi_{\theta}^{Li}(\cdot) learns one word wiw_{i} independently. All words [w0,⋯ ,wN][w_{0},\cdots,w_{N}] are concatenated together to form the textual word embedding vv. A pseudo-word S* is further introduced to represent the learned new concept, and vv is associated with its word embeddings. To train our global mapping network, we adopt Eqn. (1), and regularize the obtained word embeddings as follows:

where λglobal\lambda_{global} is a trade-off hyperparameter. Analogous to , we randomly sample a text from the CLIP ImageNet templates as text input during training, such as a photo of a S*. The full template list is provided in the Suppl. Besides, following , the key and value projection layers in the cross attention layer are finetuned with Mg(⋅)M^{g}(\cdot), and the obtained new projections are denoted as Kg=WKg⋅τθ(y)K^{g}=W^{g}_{K}\cdot\tau_{\theta}(y) and Vg=WVg⋅τθ(y)V^{g}=W^{g}_{V}\cdot\tau_{\theta}(y).

Benefiting from the hierarchical semantics learned by different layers in the CLIP image encoder, the feature from the deepest layer (i.e., layer 24) possesses the highest comprehension of the image. The word embedding associated with the deepest feature is naturally learned to describe the primary concept (i.e., the subject), while keeping superior editability. In contrast, word embeddings from shallower features are learned to describe the irrelevant disturbances (see Sec. 4.2 for more details). Note that, we use only the word embedding of the deepest feature during local training (Sec. 3.3) and image generation stages for better editability.

3 Local Mapping

Usually, a single word embedding is not sufficient to faithfully describe the details of the given concept, while multiple word embeddings may suffer from degraded editing capacity. To improve the consistency between the given concept and synthesized image without sacrificing editability, we further propose a local mapping network. As shown in Fig. 3(b), the local mapping network Ml(⋅)M^{l}(\cdot) encodes the multi-layer CLIP features into the textual feature space (i.e., the output space of text encoder):

where λ\lambda is a hyperparameter and set as 11 during training. To emphasize more on the object region, the obtained attention map QKlTQ{K^{l}}^{T} is reweighted by QKgiTQ{K^{g}}^{T}_{i}, where ii is the index of w0w_{0} in the text prompt (more details will be provided in the Suppl). To train the local mapping network, we also adopt Eqn. (1) while regularizing the local values VlV_{l}:

where λlocal\lambda_{local} is a trade-off hyperparameter.

Experiments

Datasets. To train our local and global mapping networks, we use the testset of OpenImages as our training dataset. It contains 125k images with 600 object classes. During training, we crop and resize the object image to 512×\times512 according to the bounding box annotations. While for local mapping training, mask annotations are also used to extract foreground objects. For customized generation, we adopt concept images from existing works with 20 subjects, including dog, cat, and toy, etc. The subject masks can be obtained by a pretrained segmentation model . For quantitative evaluation, we employ the editing prompts from , which contains 25 editing prompts for each subject. We randomly generate five images for each subject-prompt pair, obtaining 2,500 images in total. More details can be found in the Suppl.

Evaluation metrics. Following Dreambooth , we evaluate our method with three metrics: CLIP-I, CLIP-T, and DINO-I. For CLIP-I, we calculate the CLIP visual similarity between the generated and target concept images. For CLIP-I, we calculate the CLIP text-image similarity between the generated images and the given text prompts. The pseudo-word (S*) in the text prompt is replaced with the proper object category for extracting CLIP text feature. For DINO-I, we calculate cosine similarity between the ViTS/16 DINO embeddings of generated and concept images. Moreover, we adopt the optimization time as a metric to evaluate the efficiency of each method.

Implementation Details. We use the V1-4 version of Stable Diffusion in our experiments, and the mapping network is implemented with three-layer MLP (for both global mapping network and local mapping network). To extract multi-layer CLIP features, features from five layers are selected, and the layer indexes are {24,4,8,12,16}\{24,4,8,12,16\} in order. To train the global mapping network, we use the batch size of 16 and λglobal=0.01\lambda_{global}=0.01. The learning rate is set to 1e-6. To train the local mapping network, we adopt the batch size of 8 and λlocal=0.0001\lambda_{local}=0.0001. The learning rate is set to 1e-5. All experiments are conducted on 4×\timesV100 GPUs. During image generation, we use 100 steps of the LMS sampler, and the scale of classifier-free guidance is 5. Unless mentioned otherwise, we use λ=0.8\lambda=0.8 for concept generation (e.g., a photo of a S*), and λ=0.6\lambda=0.6 for concept editing (e.g., a S* wearing sunglasses).

2 Ablation Study

We first conduct the ablation studies to evaluate the effects of various components in our method, including the multi-layer features in the global mapping network, local mapping network, and the value of λ\lambda.

Effect of Multi-layer Features. Fig. 5 gives the visualization of words learned by multi-layer features in the global mapping network. For each word visualization, we use the text A photo of a [wiw_{i}]. One can see that, the word embedding of the deepest feature (i.e., w0w_{0}) describes the primary concept (i.e., corgi, teddybear), while other words describe some irrelevant details. Meanwhile, the obtained w0w_{0} maintains superior editability. To further demonstrate this, we conducted experiments with several variants: i) Single-layer Single-word: learning a single word embedding from the deepest feature. ii) Single-layer Multi-words: learning multiple word embeddings from the deepest feature separately. iii) Multi-layers Multi-words: our setting, learning multiple word embeddings from the multiple layer features separately. Fig. 6 illustrates the results of concept generation and editing for each variant. For multiple word settings, we show the results of the full embeddings (i.e., [vv] ) and the primary word embedding (denoted as [ww]). As shown in the figure, encoding concept image into one single word embedding leads to entangled embedding with poor editability. When learning multiple words from the deepest feature, the obtained full word embeddings vv and the primary word embedding ww are not editable either. In contrast, the primary word ww learned by our multi-features describes the object concept while maintaining superior editing capacity. That’s why we only keep it during image generation. Since a single ww is not sufficient to describe the details of the given concept faithfully, a local mapping network is further proposed to address this.

Effect of Local Mapping. We further conduct the ablation to evaluate the effect of the proposed local mapping network. As shown in Fig. 6, with the local mapping network, our ELITE generates images with higher consistency with the concept image. Meanwhile, from Table 1 we see that the introduction of a local mapping network does not compromise the editable capability, demonstrating its superiority over learning multiple words. Though its image alignment may not be the best, it has a good trade-off between image alignment and text alignment.

Effect of λ\lambda. In Eqn. 6, λ\lambda is introduced to control the fusion of information from the global mapping network and the local mapping network. To evaluate its effect, we vary its value from 0 to 1.2, and the generated results are shown in Fig. 7. We see that with the increase of λ\lambda, the consistency between the synthesized image and concept image is improved. However, when the value of λ\lambda is too large, it may lead to degenerated editing results. Therefore, for a good trade-off between inversion and editability, we set λ=0.6\lambda=0.6 for editing prompts and λ=0.8\lambda=0.8 for generating prompts. We find these parameters work well for most cases.

More ablations are provided in the Suppl.

3 Qualitative Results

To demonstrate the effectiveness of our ELITE, we compare it with existing optimization-based methods, including Textual Inversion , DreamBooth , and Custom Diffusion . For a fair comparison, we trained all models using their official codesSince the official code of Dreambooth is not publicly available, we use the code implemented by https://github.com/XavierXiao/Dreambooth-Stable-Diffusion. For Textual Inversion, we use its stable diffusion version. and default hyperparameters on a single image. Fig. 4 illustrates the images generated with the text prompt A photo of a S*. With only one concept image, Textual Inversion cannot learn a word embedding to accurately describe the target concept. Although Dreambooth and Custom Diffusion learn the concept with detail consistency, their diversity may be limited. In comparison, our ELITE is capable of faithfully capturing the details of the target concept and generating diverse images. We also conduct evaluation with editing prompts and compare our method with existing methods. As shown in Fig. 8, Dreambooth and Custom Diffusion exhibit degraded editing ability, and in some cases, the editing prompts fail to produce the desired results (first row). In contrast, our method demonstrates superior editing performance. Fig. 10 illustrates more qualitative results obtained by our method. One can see that our ELITE can generate various subjects with different contexts, accessories, and properties consistently, demonstrating its effectiveness.

4 Quantitative Results

In addition to the qualitative comparisons, we further conduct the quantitative evaluation to validate the performance of our ELITE. From Table 2, one can see that our method achieves better text-alignment compared to the state-of-the-art methods, demonstrating its superior editability. Moreover, our method achieves comparable detail consistency and image quality, indicating its capability to generate high-quality images. Furthermore, our method provides a significant merit in terms of computational efficiency. Unlike optimization-based methods that require several or tens of minutes to obtain the concept embedding, our method can finish it in just 0.05s. This makes our method highly practical and efficient for real-world applications where speed is a critical factor.

User Study. We then perform the user study to compare with competing methods. Given a subject, a text prompt and two synthesized images (ours v.s. competitor), the users are asked to select the better one from three views: i) Text alignment: “Which image is more consistent with the text?”. ii) Image alignment: “Which image better represents the objects in target images?”. iii) Editing alignment: “Which image is more consistent with both the target subject and the text?”. For each evaluated view, we employ 60 users, and each user is asked to answer 30 randomly selected questions, i.e., 1800 responses in total. As shown in Table 3, our method receives comparable preference to others.

5 Limitations

As shown in Fig. 9, our ELITE inherits the weakness from stable diffusion, i.e., failing to deal with images involving text characters.

Conclusion

In this paper, we proposed a novel learning-based encoder, namely ELITE, for fast and accurate customized text-to-image generation. Compared with existing optimization-based methods, our ELITE directly encoded visual concepts into textual embeddings, significantly reducing the computational and memory burden of learning new concepts. Moreover, our method demonstrates superior flexibility in editing the learned concepts into new scenes while preserving the image-specific details, making it a valuable tool for customized text-to-image generation. In future work, we will explore to leverage multiple concept images for better inversion, and investigate effective methods for composing multiple concepts in ELITE.

Acknowledgement. This work was supported in part by National Key R&D Program of China under Grant No. 2020AAA0104500, the National Natural Science Foundation of China (NSFC) under Grant No.s U19A2073 and 62006064, and the Hong Kong RGC RIF grant (R5001-18).

References

A More Ablation Studies

From Eqn. ( 6) in main paper, λ\lambda is introduced to control the fusion of information from the global mapping network and the local mapping network. To evaluate its effect, we vary its value from to 1.21.2 during customized generation. As shown in Fig. 7 in main paper, with the increasing of λ\lambda, the consistency between the synthesized image and concept image is improved. Meanwhile, from Fig. 12, the image alignment (i.e., CLIP-I and DINO-I) improves as λ\lambda increases. However, when the value of λ\lambda is too large, it may lead to degenerated editing results, resulting decreased text alignment (i.e., CLIP-T). Therefore, for a trade-off between inversion and editability, we set λ=0.6\lambda=0.6 for editing prompts and λ=0.8\lambda=0.8 for generating prompts. We find these parameters work well for most cases.

A.2 Effect of the layer indexes

In our experiments, we select the features of the five layers from CLIP image encoder to learn multiple word embeddings, whose indexes are {24,4,8,12,1624,4,8,12,16} in order. We have further conducted the ablation studies by putting the deepest layer (i.e., layer 2424) in different orders. Specifically, we compare four variants, i) Single-layer Multi-words: learning multiple word embeddings from the deepest feature separately. ii) Multi-layers Multi-words First: our setting, learning multiple word embeddings from the multiple layer features separately, and the layer indexes are {24,4,8,12,1624,4,8,12,16} in order. iii) Multi-layers Multi-words Middle: learning multiple word embeddings from the multiple layer features separately, and the layer indexes are {4,8,24,12,164,8,24,12,16} in order. iv) Multi-layers Multi-words Last: learning multiple word embeddings from the multiple layer features separately, and the layer indexes are {4,8,12,16,244,8,12,16,24} in order. Fig. 13 illustrates the visualization of words learned by each variant.

As shown in the figure, each variant contains one primary word that describes the subject concept. When learning multiple word embeddings from multi-layer features, we observe that the primary word is naturally linked to the features from the deepest layer, regardless of the position indices of layers. Besides, as illustrated in Fig. 14, in contrast to the primary word obtained by the single layer feature, the primary word learned by multi-layer features is well-editable. Among them, our setting achieves better editability, which is shown in Table 4.

A.3 Effect of the local attention map reweighting

The local mapping network aims to inject the fine-grained details of given subject during generation. To further emphasize its effect on the subject region rather than irrelevant areas (e.g., background), we reweight the obtained local attention map by multiplying it with the attention map of primary word (refer to Sec. 3.3 in main paper),

where Al=Softmax(QKlTd′)A^{l}=\text{Softmax}\left(\frac{Q{K^{l}}^{T}}{\sqrt{d^{\prime}}}\right) denotes the attention map of local mapping network and Ag=Softmax(QKgTd′)A^{g}=\text{Softmax}\left(\frac{Q{K^{g}}^{T}}{\sqrt{d^{\prime}}}\right) denotes the attention map of global mapping network. d′d^{\prime} is the output dimension of key and query features. Aw0gA^{g}_{w_{0}} is the attention map of primary word w0w_{0}. To verify its effectiveness, we firstly visualize the cross-attention map of each word in input text prompt in Fig. 15. As one can see, the learned primary word w0w_{0} is associated with the subject concept and its attention map Aw0gA^{g}_{w_{0}} accurately delineates the subject region, so we can leverage it to reweight the local attention map AlA^{l}. Furthermore, as illustrated in Fig. 16, without the local attention reweighting, the features of local mapping network may affect the subject-irrelevant areas, resulting in degraded editability. In contrast, our ELITE with reweighting strategy reduces the disturbances on subject-irrelevant areas, and achieves better editability.

A.4 Effect of the Global Mapping

We have further conducted the ablation study to evaluate the effect of our global mapping network. For comparison, we remove the global mapping network, while replace the pseudo word S* with a ground-truth category word to learn the local mapping network (e.g., S* →\rightarrow dog in Fig. 17). As shown in Fig. 17, without the global mapping, using the local mapping network only provides a few fine-grained details (e.g., fur color), yet fails to keep the structure of the concept (e.g., ear). In contrast, by adding global mapping network to encode a suitable primary word embedding, our ELITE faithfully recovers the target concept with higher visual fidelity while enabling robust editing.

B More Experimental Details

Textual Inversion . We use the official stable diffusion version of Textual Inversionhttps://github.com/rinongal/textual_inversion. For each subject, experiment is conducted with the batch size of 1 and a learning rate of 0.005 for 5,000 steps. The new token is initialized with the category word, e.g., “cat”.

Custom Diffusion . We use the official implementation of Custom Diffusionhttps://github.com/adobe-research/custom-diffusion. We train it with a batch size of 1 for 300 training steps. The learning rate is set as 1e-5. The regularization images are generated with 50 steps of the DDIM sampler with the text prompt “A photo of a [category]”.

DreamBooth . We use the third-party implementation of DreamBoothhttps://github.com/XavierXiao/Dreambooth-Stable-Diffusion. Training is done with finetuning both the U-net diffusion model and the text transformer. The training batch size is 1 and the learning rate is set as 1e-6. The regularization images are generated with 50 steps of the DDIM sampler with the text prompt “A photo of a [category]”. For each subject, we train it for 800 steps.

B.2 Testing Datasets

For customized generation, we adopt concept images from existing works with 20 subjects, including dog, cat, and toy, etc. Fig. 11 illustrates the full image samples.

B.3 Text prompts

We adopt the text prompt list used in Textual Inversion for training, which is provided as below:

For qualitative evaluation, we employ the editing templates used in . For quantitative evaluation, we employ the editing prompts from DreamBench , which contains 25 editing prompts for each subject. The full prompts are listed in Table 5.