MultiBooth: Towards Generating All Your Concepts in an Image from Text

Chenyang Zhu, Kai Li, Yue Ma, Chunming He, Xiu Li

Introduction

The advent of diffusion models has ignited a new wave in the text-to-image (T2I) task, leading to the proposal of numerous models . Despite the broad capabilities of these models, users often desire to generate specific concepts such as beloved pets or personal items. These personal concepts are not captured during the training of large-scale T2I models due to their subjective nature, emphasizing the need for customized generation . Customized generation aims to create new variations of a given concept, including different contexts (e.g., beaches, forests) and styles (e.g., painting), based on just a few user-provided images (typically fewer than 5).

Recent customized generation methods either learn a concise token representation for each subject or adopt an efficient fine-tuning strategy to adapt the T2I model specifically for the subject . While these methods have achieved impressive results, they primarily focus on single-concept customization and struggle when users want to generate customized images for multiple subjects (see Fig. 2). This motivates the study of multi-concept customization (MCC).

Existing methods for MCC commonly employ joint training approaches. However, this strategy often leads to feature confusion, as illustrated in the third column of Fig. 2. Furthermore, these methods require training distinct models for each combination of subjects and are hard to scale up as the number of subjects grows. An alternative method addresses MCC by adjusting attention maps with residual token embeddings during inference. While this approach shows promise, it incurs a notable inference cost. Furthermore, the method encounters difficulties in attaining high fidelity due to the restricted learning capacity of a single residual embedding.

To address the aforementioned issues, we introduce MultiBooth, a two-phase MCC solution that accurately and efficiently generates customized multi-concept images based on user demand, as demonstrated in the Fig. 2. MultiBooth includes a discriminative single-concept learning phase and a plug-and-play multi-concept integration phase. In the former phase, we learn each concept separately, resulting in a single-concept module for every concept. In the latter phase, we effectively combine these single-concept modules to generate multi-concept images without any extra training.

More concretely, we propose the Adaptive Concept Normalization (ACN) to enhance the representative capability of the generated customized embedding in the single-concept learning phase. We employ a trainable multi-model encoder to generate customized embeddings, followed by the ACN to adjust the L2 norm of these embeddings. Finally, by incorporating an efficient concept encoding technique, all detailed information of a new concept is extracted and stored in a single-concept module which contains a customized embedding and the efficient concept encoding parameters.

In the plug-and-play multi-concept integration phase, we further propose a regional customization module to guide the inference process, allowing the correct combination of different single-concept modules for multi-concept image generation. Specifically, we divide the attention map into different regions within the cross-attention layers of the U-Net, and each region’s attention value is guided by the corresponding single-concept module and prompt. Through the proposed regional customization module, we can generate multi-concept images via any combination of single-concept modules while bringing minimal cost during inference.

Our approach is extensively validated with various representative subjects, including pets, objects, scenes, etc. The results from both qualitative and quantitative comparisons highlight the advantages of our approach in terms of concept fidelity and prompt alignment capability. Our contributions are summarized as follows:

We propose a novel framework named MultiBooth. It allows plug-and-play multi-concept generation after separate customization of each concept.

The adaptive concept normalization is proposed in our MultiBooth to mitigate the problem of domain gap in the embedding space, thus learning a representative customized embedding. We also introduce the regional customization module to effectively combine multiple single-concept modules for multi-concept generation.

Our method consistently outperforms current methods in terms of image quality, faithfulness to the intended concepts, and alignment with the text prompts.

Related Work

Text-to-image generation is an extensively researched problem, with numerous early studies focusing on Generative Adversarial Networks (GANs). Noteworthy examples include AttnGAN, StackGAN, StackGAN++, Mirrorgan. While these methods have proven to be effective on specific datasets such as faces and landscapes, they have limited generalization ability to larger-scale datasets. Moreover, training GANs poses instability issues and is susceptible to mode collapse. With the advancements of diffusion models, more and more methods have started exploring text-to-image generation based on diffusion models. By training on large-scale text-image datasets such as LAION, these methods have achieved superior performance. Notable recent works include DALLE2, Imagen, GLIDE, and Stable Diffusion, which are capable of generating images based on open-vocabulary text prompts.

2 Customized Text to Image Generation

The goal of customized text-to-image generation is to acquire knowledge of a novel concept from a limited set of examples and subsequently generate images of these concepts in diverse scenarios based on text prompts. By leveraging the aforementioned diffusion-based methodologies, it becomes possible to employ the comprehensive text-image prior to customizing the text-to-image process.

Textual Inversion achieves customization by creating a new embedding in the tokenizer and associating all the details of the newly introduced concept to this embedding. DreamBooth binds the new concept to a rare token followed by a class noun. This process is achieved by finetuning the entire diffusion model. Additionally, DreamBooth addresses the issue of language drift through a prior preservation loss. ELITE utilizes multi-layer embeddings from CLIP image encoder and employs a mapping network to acquire customized embeddings. Similar approaches are adopted in InstantBooth and E4T. Such encoder-based methods require less training time and yield better image generation results compared to directly optimizing the embedding in Textual Inversion. Custom Diffusion links a novel concept to a rare token by adjusting specific parameters of the diffusion model. To alleviate overfitting, Custom Diffusion also incorporates a regularization dataset. Through joint training, Custom Diffusion explores the problem of multi-concept customization for the first time. Cones2 learns new concepts by adding a residual embedding on top of the base embedding and generates multi-concept images through attention map manipulation.

In this work, we utilize a multi-modal model and LoRA to discriminatively and concisely encode every single concept. Then, we introduce the regional customization module to efficiently and accurately produce multi-concept images.

Method

Given a series of images S={Xs}s=1S\mathcal{S}=\{X_{s}\}^{S}_{s=1} that represent SS concepts of interest, where {Xs}={xi}i=1M\{X_{s}\}=\{x_{i}\}^{M}_{i=1} denotes the MM images belonging to the concept ss which is usually very small (e.g., M<=5M<=5), the goal of multi-concept customization (MCC) is to generate images that include any number of concepts from S\mathcal{S} in various styles, contexts, layout relationship as specified by given text prompts. MCC poses significant challenges for two primary reasons. Firstly, learning a concept with a limited number of images is inherently difficult. Secondly, generating multiple concepts simultaneously and coherently within the same image while faithfully adhering to the provided text is even harder.

To tackle these challenges, Custom Diffusion simultaneously finetunes the model with all subjects of interest. This approach can lead to issues like feature confusion and inaccurate attribution (e.g., attributing dog features to a cat). For Cone2, the process requires repetitive strengthening and weakening of the same concept features during the inference phase, which can result in a significant decrease in fidelity. To avoid these problems, we propose a method called MultiBooth. Our MultiBooth initially performs high-fidelity learning of a single concept. We employ a multi-modal encoder and the adaptive concept normalization strategy to obtain text-aligned representative customized embeddings. Additionally, the efficient concept encoding technique is employed to further improve the fidelity of single-concept learning. To generate multi-concept images, we employ the regional customization module. This module serves as a guide for multiple single-concept modules and utilizes bounding boxes to indicate the positions of each generated concept.

In this paper, the foundational model utilized for text-to-image generation is Stable Diffusion . It takes a text prompt PP as input and generates the corresponding image xx. Stable diffusion consists of three main components: an autoencoder(E(⋅),D(⋅))(\mathcal{E}(\cdot),\mathcal{D}(\cdot)), a CLIP text encoder τθ(⋅)\tau_{\theta}(\cdot) and a U-Net ϵθ(⋅)\epsilon_{\theta}(\cdot). Typically, it is trained with the guidance of the following reconstruction loss:

where ϵ∼N(0,1)\epsilon\sim\mathcal{N}\left(0,1\right) is a randomly sampled noise, t denotes the time step. The calculation of ztz_{t} is given by zt=αtz+σtϵz_{t}=\alpha_{t}z+\sigma_{t}\epsilon, where the coefficients αt\alpha_{t} and σt\sigma_{t} are provided by the noise scheduler.

Given MM images {Xs}={xi}i=1M\{X_{s}\}=\{x_{i}\}_{i=1}^{M} of a certain concept ss, previous works associate a unique placeholder string S∗S^{*} with concept ss through a specific prompt PsP_{s} like “a photo of a S∗S^{*} dog”, with the following finetuning objective:

The result of minimizing Eq. 2 is to encourage the U-Net ϵθ(⋅)\epsilon_{\theta}(\cdot) to accurately reconstruct the images of the concept ss, effectively binding the placeholder string S∗S^{*} to the concept ss.

2 Single-Concept Learning

Existing methods mainly utilize a single image encoder to encode the concepts of interest. However, this approach may also encode irrelevant information, such as unrelated objects in the images. To remedy this, we employ a multi-modal encoder that takes as input both the images and the subject name (e.g., “dog”) to learn a concise and discriminative representation for each concept. Inspired by MiniGPT4 and BLIP-Diffusion , we utilize the QFormer, a light-weighted multi-modal encoder, to generate customized concept embeddings. The QFormer encoder EE has three types of inputs: visual embeddings ξ\xi of an image, text description ll, and learnable query tokens W=[w1,⋯ ,wK]W=[w_{1},\cdots,w_{K}] where KK is the number of query tokens. The outputs of QFormer are tokens O=[o1,⋯ ,oK]O=[o_{1},\cdots,o_{K}] with the same dimensions as the input query tokens.

As shown in Fig. 3, given an image xi∈Xsx_{i}\in X_{s}, we employ a frozen CLIP image encoder to extract the visual embeddings ξ\xi of the image. Subsequently, we set the input text ll as the subject name for the image and input it into the encoder EE. The learnable query tokens WW interact with the text description ll through a self-attention layer and with the visual embedding ξ\xi through a cross-attention layer, resulting in text-image aligned output tokens O=E(ξ,l,W)O=E(\xi,l,W). Finally, we obtain the initial customized embedding viv_{i} by taking the average of these tokens:

2.2 Adaptive Concept Normalization

During the reconstruction with the prompt mentioned above, we have observed a domain gap between our customized embedding viv_{i} and other word embeddings in the prompt. As shown in Tab. 1, the L2 norm of our customized embedding is considerably larger than that of other word embeddings in the prompt. Notably, these word embeddings, belonging to the same order of magnitude, are predefined within the embedding space of the CLIP text encoder τθ(⋅)\tau_{\theta}(\cdot). This significant difference in quantity weakens the model’s ability for multi-concept generation. To remedy this, we further apply the Adaptive Concept Normalization (ACN) strategy to the customized embedding viv_{i}, adjusting its L2 norm to obtain the final customized embedding vi^\hat{v_{i}}

2.3 Efficient Concept Encoding

The whole single-concept learning framework can be trained by optimizing Eq. 2 with a regularization term, as shown in the following equation:

where λ\lambda denotes a balancing hyperparameter and is consistently set to 0.01 across all experiments. During training, we randomly select text prompts PsP_{s} from the CLIP ImageNet templates following the Textual Inversion . The complete templates can be found in the SupplSuppl. Through the above methods, the detailed information of a new concept can be extracted and stored in a customized embedding and corresponding LoRA parameters, which can be called a single-concept module. The storage requirement of a concept is less than 7MB, which is significantly advantageous compared to 3.3GB in and 72MB in Custom Diffusion.

3 Multi-Concept Integration

Our key insight is to restrict the generation of each concept within a given region. As illustrated in the right part of Fig. 3, we propose the regional customization module in cross-attention layers, which makes it possible to integrate multiple LoRAs for multi-concept generation. Given a series of images S={Xs}s=1S\mathcal{S}=\{X_{s}\}^{S}_{s=1} that represent SS concepts of interest, users can define the corresponding bounding boxes B={bi}i=1SB=\{b_{i}\}_{i=1}^{S} and region prompts Pr={pi}i=1SP_{r}=\{p_{i}\}_{i=1}^{S} for each concept, along with a base prompt pbasep_{base}. The text embeddings C={ci}i=1SC=\{c_{i}\}_{i=1}^{S} of each concept can be acquired using the CLIP text encoder, calculated as follows:

The attention operation is then applied to the query, key, and value vectors to derive the image feature as follows:

The benefit of the proposed regional customization module is that each given region only interacts with the specific concept’s content during the cross-attention operation. This avoids the issue of mixing multi-concept features during cross-attention and allows each single-concept module to generate specific concepts following different text prompts in their respective regions. Moreover, the regional customization module brings minimal cost during inference, as evidenced in Tab. 2.

Experiment

All of our experiments are based on Stable Diffusion v1.5 and are conducted on a single RTX3090. We set the rank of LoRA to be 16. We use the AdamW optimizer with a learning rate of 8×10−58\times 10^{-5} and a batch size of 1, optimizing for 900 steps. During the inference stage, we utilize the DPM-Solver for sampling 100 steps, with the guidance scale ω=7.5\omega=7.5.

1.2 Datasets.

Following Custom Diffusion, we conduct experiments on twelve subjects selected from the DreamBooth dataset and CustomConcept101. They cover a wide range of categories including two scene categories, two pets, and eight objects.

1.3 Evaluation Metrics.

As shown in Fig. 6, we have observed that when redundant elements dominate the calculation of the CLIP image score, it can result in high CLIP image scores for low-quality images. To address this issue, we suggest masking the objects in the source image that are truly relevant before computing the CLIP image score. We refer to this modified metric as Seg CLIP-I. As a result, we assess all the methods using three evaluation metrics: CLIP-I, Seg CLIP-I, and CLIP-T. (1) CLIP-I measures the average cosine similarity between the CLIP embeddings of the generated images and the source images. (2) Seg CLIP-I is similar to CLIP-I, but all of the source images are processed with segmentation on the relevant objects. (3) CLIP-T calculates the average cosine similarity between the prompt CLIP embeddings and image CLIP embeddings.

1.4 Baselines.

We conduct comparisons between our method and four existing methods: Textual Inversion (abbreviated as TI), DreamBooth (abbreviated as DB), Custom Diffusion (abbreviated as Custom), and Cones2. To ensure consistency, we implement Textual Inversion, DreamBooth, and Custom Diffusion using their respective Diffusers versions. For Cones2, we implement it using its official code. All experimental settings follow the official recommended settings of each method.

2 Qualitative Comparison

We validate our method and all the comparison methods on various prompts. Specifically, these prompts include illustrating the relative positions of concepts, generating concepts in new situations, creating multi-concept images with a specific style, and associating different concepts with different prompts. In Fig. 5, the results of all comparative methods are generated by the base prompt which is presented under the images, while the results of our method are jointly generated through both the base prompt and region prompts. These generated images demonstrate that our method excels in generating high-quality multi-concept images while effectively adhering to different prompts.

3 Quantitative Comparison

As presented in Tab. 2, our method demonstrates superior image alignment compared to other methods in the single-concept setting. Additionally, our method achieves comparable text alignment, showcasing its adaptability to various complex prompts. In the multi-concept setting, our method outperforms all the compared methods in the three selected metrics, notably excelling in CLIP-I and Seg CLIP-I. Moreover, with excellent image fidelity and prompt alignment ability, our method does not incur significant training and inference costs. This indicates the effectiveness of our efficient concept encoding and regional customization module regarding multi-concept customization.

4 Ablation Study

We conduct ablation experiments to validate the effectiveness of the regional customization module. In an attempt to directly load two LoRAs into the U-Net and generate images without utilizing the regional customization module, we find that

this approach leads to feature confusion between different concepts and often only generates one object. As shown in Fig. 7, the features of the candle and teapot have fused to some extent under the prompt of “a S∗S^{*} candle and a V∗V^{*} teapot with a city in the background”, resulting in an output that contains both candle and teapot features. Moreover, such a generation method without the regional customization module produces only one object most of the time, which significantly lowers its CLIP-T in comparison to our approach as presented in Tab. 3.

4.2 Without Training QFormer (Ours w/o QFormer).

In MiniGPT4, the QFormer is completely frozen, and a linear layer is trained to project the output of the QFormer. Following this, we experiment by freezing the QFormer and training only a linear layer. As illustrated in Tab. 3, this approach results in a decrease in the learning capacity of concepts and consequently a decline in the representation capability of the generated customized embeddings. As a result, there is a noticeable decrease in all three selected metrics. Images generated by this approach also have lower fidelity compared to our method as shown in Fig. 7.

4.3 Without adaptive concept normalization (Ours w/o ACN).

We conduct an ablation experiment to verify the effectiveness of the Adaptive Concept Normalization (ACN), which is used to mitigate the domain gap between the customized embedding and the other word embedding. As illustrated in Fig. 7 and Tab. 3, this domain gap leads to a decline in the customization ability, resulting in a notable decrease in image fidelity and prompt alignment ability compared to our method.

5 User Study

We conduct user studies for all methods. For text alignment, we display an image of each method alongside its prompt and inquire, “Which images accurately represent the text description?”. For image alignment, we exhibit various training and generated images and ask, “Which images align best with the target images?”. Each questionnaire comprises 50 such inquiries. From a total of 460 received questionnaires, 72 are deemed invalid, resulting in 388 valid responses. As shown in Tab. 4, users prefer our approach for multi-concept generation regarding both text and image alignment.

6 Challenge Cases

Our method theoretically enables the combination of an unlimited number of concepts, facilitating true multi-concept customization. In Fig. 8, we present the results of customizing three and four concepts, including a total of 10 concepts. The visualization demonstrates that our method consistently generates high-quality customization images when facing a larger number of concepts, further demonstrating the superiority of our approach.

7 Extensions

Our regional customization module can guide multiple LoRAs to conduct multi-concept generation, and it is adaptable to any LoRA-related method. An illustration of transferring the regional customization module to the LoRA-based DreamBooth method has been provided. As shown in Fig. 9, with the support of the regional customization module, DreamBooth is capable of the generation of multiple concepts. However, the absence of customized embeddings results in a decrease in fidelity.

7.2 Apply ControlNet to our MultiBooth.

Our method is compatible with ControlNet to achieve structure-controlled multi-concept generation. As illustrated in Fig. 9, our model inherits the architecture of the original U-Net model, resulting in satisfactory generations through seamless integration with pre-trained ControlNet without additional training.

Conclusion

In this paper, we introduce MultiBooth, a novel and efficient framework for multi-concept customization(MCC). Compared with existing MCC methods, our MultiBooth allows plug-and-play multi-concept generation with high image fidelity while bringing minimal cost during training and inference. By conducting qualitative and quantitative experiments, we robustly demonstrate our superiority over state-of-the-art methods within diverse multi-subject customization scenarios. Since current methods still require training to learn new concepts, in the future, we will investigate the task of training-free multi-concept customization based on our MultiBooth.

References