Attention Calibration for Disentangled Text-to-Image Personalization
Yanbing Zhang, Mengping Yang, Qin Zhou, Zhe Wang
Introduction
Recently developed large-scale text-to-image models have shown unprecedented capabilities in synthesizing high-quality and diverse images based on a target text prompt. Built on these models, personalized techniques are further introduced to customize the models for synthesizing personal concepts with sufficient fidelity.
Given as input just a few images of the personal concepts (e.g., family, friends, pets, or individual objects), personalized text-to-image models aim to learn a new word embedding to represent a specific concept . However, existing methods still lack the flexibility to render all existing concepts in a given image, or only focus on a specific concept . Given a unique photo from a user (which could be people rarely seen together or uncommon furniture pieces), with multiple concepts occurring in the complex scene, the user naturally desires the ability to freely synthesize the concepts by composing multiple objects or focusing on only one of them. For example, two specific individuals at a beach, or alternatively, one of them in Times Square, as shown in Fig. 1.
To achieve flexible renditions of the concepts, instead of using a single new word to represent one concept , we employ multiple new words to represent multiple concepts. For example, considering an image containing a distinct chair and lamp (as shown in Fig. 2), we utilize the prompt “ chair and lamp” to distinguish between them, with “” serving as the modifier for “chair” and “” as the modifier for “lamp”. This intuitive formulation poses two key challenges. Firstly, the new word embeddings are likely to map confusing information, failing to maintain visual-fidelity to the target concepts. Secondly, with a relatively small training set (e.g., only one image), the model is prone to synthesizing multiple subjects, even when the target prompt pertains to a single concept. For example, as depicted in Fig. 2, the ideal output should exclusively feature the specified lamp when the target text is “A lamp”. Nonetheless, the image generated by the current state-of-the-art model not only includes a lamp that doesn’t match the color and texture of the input image but also involves a chair that shouldn’t be present.
In this paper, we propose a novel personalized T2I model, referred to as DisenDiff (i.e., Disentangled Diffusion), to address the above-mentioned issues. To preserve the good generalization ability in pre-trained large-scale models, we follow to only update the light-weight modules ( and matrices) within the cross-attention units along with new token embeddings to extend concepts. Our key insight is that current methods lack the necessary guidance for the optimization process, resulting in cluttered attention maps (as shown in Fig. 4, the first row). Consequently, existing methods struggle to synthesize each concept effectively.
Based on the above observations, we strive to generate precise attention maps from the following two aspects. Building on the discovery that the attention map of the class token can roughly align with the location of the concept, then we propose a modifier-class alignment term to bind the attention map of each new modifier with its corresponding class token, correcting attention to focus on the region of the related concept. However, the attention maps of different class tokens often exhibit overlaps, leading to the incorrect attribute binding and mutual entanglement. To achieve effective decoupling, we introduce the separate and strengthen (s&s) strategy to allow flexibly synthesizing each concept independently. By minimizing the overlapping regions between the attention maps of different class tokens, we can effectively mitigate the co-occurring issue when targeting at a specific concept. To further enhance the independence of concepts, we introduce a suppression technique to sharpen the boundaries of class tokens’ attention maps. Our contributions are summarized below:
We propose DisenDiff to comprehend multiple personal concepts from only a single image. By using diverse target texts, it can render combined/independent concepts in imaginary contexts while preserving high fidelity to the input image.
We employ two key constraints to attain precise attention maps for crucial tokens. The binding constraint locates new modifiers to different concepts, while the s&s constraint decouples these concepts.
We conduct experiments on various datasets and demonstrate that our method outperforms the current state of the art in quantitative and qualitative aspects. Additionally, we show the flexibility of our approach by applying it to extended tasks.
Related Work
Text-to-image generative models. The objective of text-to-image (T2I) tasks is to generate an image corresponding to a given textual description. Thanks to large-scale datasets and advancements in language models , T2I models have witnessed remarkable progress. While Generative adversarial networks (GANs) and autoregressive (AR) transformers have delivered impressive results, diffusion models have taken the lead in T2I generation. These models employ denoising processes in image space or latent space , resulting in unprecedented image generation quality. However, they encounter challenges when generating specific objects, such as custom furniture, even with detailed prompts. We aim to augment these models to accurately capture the appearances of novel concepts from real-world images.
Text-guided image editing. With the surge of powerful T2I models, numerous studies have delved into enhancing the controllability of diffusion models to cater to diverse user demands. Approaches such as refine the cross-attention units to encompass all subject tokens, motivating the model to fully convey the semantics in the input prompt. Techniques like implement region control in T2I generation by using bounding boxes and paired object labels as inputs. Additionally, and harness pre-trained diffusion models for image-to-image translation. A substantial body of work also focuses on local or global modifications of single images using existing T2I models. Notable examples include SINE and UniTune , which achieve image editing by fine-tuning the diffusion model. Other methods like prompt-to-prompt , null-text inversion , and impose constraints on latent noise during inference time without model training. While our objectives share some common ground with these methods, our primary focus is optimizing the model to seamlessly extend personalized concepts into new prompts.
T2I personalization Personalization techniques adapt diffusion models to learn new concepts from user-provided images, often relying on a small dataset of 3-5 images or even a single image. Textual Inversion uses pseudo-words to represent new concepts through a visual reconstruction objective. To leverage semantic priors from pre-trained models, DreamBooth utilizes a unique identifier and class name within the input text to represent new concepts. Custom-Diffusion and Perfusion compose multiple new concepts by updating only the cross-attention Keys and Values along with new token embeddings. When working with a dataset containing just a single image, current methods typically begin with additional domain-specific pre-training on a large dataset before adapting to the new concept. In contrast to these methods, we aim to address the more challenging problem of acquiring multiple concepts from a single image without domain-specific pre-training.
Method
Our objective is to understand multiple concepts within a single image. To this end, we propose a novel attention calibration mechanism to help generate accurate cross-attention maps in our T2I model. Firstly, the cross-attention maps are calculated as the activation responses between each word of the input text and the intermediate visual features. Then, we impose constraints on the cross-attention maps between both the modifier-class token pairs and class-class token pairs to bind the cross-attention maps of each modifier with its corresponding class (modifier-class constraint), as well as to ensure full comprehension of each class and separation between different classes (class-class constraint). To further mitigate the cross-interference issue in our T2I model, we introduce a suppression technique to obtain a sharper attention map for each class token. A schematic workflow of our method is presented in Fig. 3.
Stable Diffusion. In our experiments, we use Stable Diffusion as our backbone model, inheriting the structure of the Latent Diffusion Model (LDM) . It primarily consists of three components: a pre-trained text encoder from CLIP , a VAE model , and a U-Net diffusion model trained on the latent space of the pre-trained VAE. Given the noisy latent code at timestep, the diffusion model predicts the random added noise . The training objective of the diffusion model is formulated as follows:
where denotes the input image, is the input text. Following , prior knowledge in CLIP is integrated via the cross-attention mechanism.
Integrating textual features via cross-attention. Formally, the intermediate spatial representation of the denoiser U-Net is mapped to a query matrix , while text embeddings are mapped to a key matrix and a value matrix , using learnable projection matrices , , and . Then, the cross-attention maps are obtained as:
Text encoding. Generally, during training of a T2I system, a suitable text prompt is required in addition to the selected single image. In this paper, we adopt a manner similar to , incorporating new modifiers and the classes to be modified into the input text. For example, if the target image contains a cat and a dog, the text prompt would be “ cat and dog”. The modifier tokens “” are initialized with rare vocabulary. Given only a single training image, the T2I model will likely lack the diversity of generation, known as the language drift problem. Using our text prompt, we can easily select regularized images with the same caption to mitigate the issue of language drift, enabling our model to generate a variety of cats and dogs (not limited to the ones present in the target images, as shown in Fig. 5, left of the second row).
Current methods are prone to overfitting when the training data only consists of a single image, resulting in ambiguous attention maps for each token (as shown in the first row of Fig. 4). As demonstrated in P2P , the spatial layout and geometry of the generated images depend on the cross-attention maps. Therefore, our primary focus is to optimize the model to produce accurate cross-attention maps, elaborated in the following part.
2 Coherent binding of modifiers with classes
Based on the cross-attention maps () obtained by a previous method (shown in Fig. 4, the first row), we can observe that while of new modifiers are chaotic ( and ), cross-attention of class tokens can roughly capture the semantic boundaries ( and ). We attribute it to the fact that the majority of parameters in the T2I models are frozen, preserving the category information of class tokens. To aid the new modifiers in understanding their responsibilities, we define the constraint to bind the cross-attention maps of modifiers with their corresponding class tokens as
where and represent the attention map of the -th modifier and the -th class at timestep, respectively. The loss is formulated to reduce the intersection over union (IoU) between these two attention maps, encouraging a close alignment between the activations of the modifiers and the class tokens. To prevent substantial influence on , we detach its gradient during the loss computation.
Nonetheless, there are two potential issues when we directly apply this constraint. Given that is the result of the Softmax operation (i.e., , where denotes the activation of the -th token at pixel ), input tokens would contend for attention at the same position. Consequently, a precise pixel-to-pixel correspondence between and can not be established. Furthermore, our intention is for the activations of to fully encompass the corresponding object, thereby capturing all its attributes comprehensively. However, as depicted in Fig. 4, it is evident that within the object region, certain activations of exhibit high values, while others appear considerably lower. This poses a challenge for the attention to sustain a comprehensive focus on the object. To address these challenges, we employ a Gaussian filter on , which leads to the generation of smooth attention maps referred to as . This smoothing process helps to alleviate the pixel-wise competition among tokens and facilitates more comprehensive attention to the object. Consequently, by using the loss function , we encourage to have coherent attention areas with , while achieving a broader coverage of the object, without the need for precise point-to-point binding. For simplicity, in the subsequent sections of this paper, unless explicitly specified otherwise, we apply a Gaussian filter to .
3 Separating and strengthening attention maps for multiple classes
Given a single image as the training set, it’s inevitable for one class token to attend to multiple concepts simultaneously. For instance, in the first row of Fig. 4, specifically in , the attention dedicated to the “cat” token is not solely limited to the “cat” concept. It also exhibits some degree of attention towards the “dog” concept. Thus, incorporates attributes associated with the “dog” concept due to its binding with . To ensure independent editing of concepts without interference, it is necessary to separate the attention regions of different objects (i.e., and ). A straightforward approach is to minimize the overlap between attention maps of different object tokens as
The utilization of effectively prevents the activations of class tokens from overlapping. However, it may come with a side effect of reducing the area of , potentially leading to a loss of identity for the corresponding class, which can be found in the supplement. To simultaneously minimize the overlap among attention maps and preserve the class identity, we design the following constraint,
where “s&s” stands for “separate and strengthen.” The loss strikes a balance between avoiding overlap with other objects and ensuring comprehensive coverage of the target object, thus improving the accuracy and fidelity of the attention mechanism.
Suppression. The utilization of the loss can potentially lead to another issue where the attention map captures a significant portion of the activations, while exhibits very few activations. This imbalance in activation distribution between different class tokens can result in an uneven emphasis on certain classes. To address it, we introduce a suppression mechanism. Specifically, before computing the , we apply an element-wise multiplication operation to (i.e., ). Given that activations fall within the range of $f_{m}(A_{t}^{c_{i}})L_{\text{s\&s}}(f_{m}(A_{t}^{c_{i}}),f_{m}(A_{t}^{c_{j}}))A_{t}^{m_{i}}A_{t}^{c_{i}}$.
In summary, the total training loss is formulated as:
where is the number of classes in the input image, and is the base loss of the T2I model in Eq. 1. is responsible for refining the attention maps related to class tokens, while is responsible for constraining new modifier tokens to acquire correct attributes. The auxiliary functions and facilitate the optimization process. The synergy among these constraints results in the generation of precise and interpretable attention maps for input tokens, shown in the second row of Fig. 4.
Experiments
Datasets. We conducted experiments on ten datasets spanning a large range of categories including people, animals, furniture, and people with pets/toys. Please note, instead of concentrating only on one concept, our datasets contain two distinct concepts within each image. During the inference phase, we test 30 different prompts for each image: 10 for combined concepts, 10 specifically targeting the first concept, and 10 focusing on the second concept.
Compared methods. We compare with three personalized T2I methods, which all utilize new word embeddings to represent novel concepts. (1) Textual Inversion (TI): In TI, only the new token embedding representing the novel concept is updated, while the other parameters remain frozen. (2) DreamBooth (DB): DB updates all layers of the T2I model to maintain visual fidelity and employs a prior preservation loss to mitigate language drift. (3) Custom-Diffusion (CD): CD updates the most relevant weights related to the input textual features, including and within the cross-attention units, as well as the new token embedding. The implementation details are provided in the supplement.
Evaluation metrics. The synthetic images should faithfully capture the visual characteristics of the input image while accurately conveying all elements of the target text. We employ two key metrics: (1) The image-alignment metric evaluates the reconstruction of concepts, which measures the pairwise CLIP-space cosine similarity between the generated images and the corresponding real images. (2) The text-alignment metric assesses the editing effectiveness of the fine-tuned model by calculating the text-image similarity between the generated images and the provided prompts using CLIP . Notably, these two indicators often conflict with each other . For each concept, we synthesize 16 samples per prompt, using 50 DDIM steps and a guidance scale of 6. For comparison, we provide scores for combined concepts, the first concept, the second concept, and their average (referred to as Combined, Concept1, Concept2, and Mean in Fig. 6). For instance, if the training image caption is “ cat and dog”, the test prompts of the Combined, Concept1 and Concept2 settings are “ cat and dog in a garden”, “ cat wearing a hat”, “A pink dog”, respectively. When testing on independent concepts, we calculate the image-alignment metric between the synthesized images and the segmented image containing only the corresponding subject.
Implementation details. We fine-tune the Stable Diffusion model for 250 steps, with a batch size of 8 and a learning rate of . Similar to , we employ clip-retrieval to select samples from LAION-5B dataset as regularization images. Captions of these selected images exhibit a similarity of over 0.85 in the CLIP textual embedding space with the input text. Meanwhile, we use the data augmentation in . In our experiments, we apply the proposed cross-attention calibration to the attention units, which have been shown to contain the most semantic information .
2 Comparison Results
Quantitative comparisons. Fig. 6a illustrates the results averaged across ten datasets. As shown, we outperform all the compared methods, especially on the image-alignment scores. Specifically, despite Textual Inversion (TI) achieving the highest text-alignment score, it has the lowest image-alignment score, indicating its struggle to maintain the appearance of concepts. DreamBooth (DB) outperforms TI in image-alignment score but falls significantly short compared to our approach in both metrics. Custom Diffusion (CD) maintains a better balance between the two metrics and competes with ours in combined concepts scores and Concept1 scores. However, there is a noticeable performance gap in the scores for Concept2. In summary, we achieve the highest image fidelity while maintaining strong text editing effectiveness. Detailed results for each dataset can be found in the supplement.
Qualitative comparisons. We visually demonstrate the favorable outcomes in Fig. 5. Concretely, we design diverse target prompts to assess the learned independent concepts and combined concepts in different editing scenarios, including scene changes, object addition, style transfer, property change, accessory addition, interactions between multiple concepts, concept decoupling, and the ability to address the language drift (e.g., generating a specific cat consistent with the input and a dog with a breed distinct from the one present in the input). As shown in Fig. 5, images synthesized by DB either lack key attributes of the concepts or suffer from severe overfitting to the input image. With most of its parameters frozen, CD improves editability and reconstruction compared to DB. However, it still struggles to preserve concepts’ appearances or decouple from the input image, especially as shown in the first and last rows of Fig. 5. By incorporating cross-attention calibration, our method achieves high visual fidelity and maintains effective cross-concept disentanglement during T2I generation. For the sake of space efficiency, additional results including Textual Inversion are provided in the supplement.
3 Ablation Studies
We conduct ablation studies to show the effectiveness of each component and analyze the influence of different design choices, adopting the same setup described in Sec. 4.1.
To assess the necessity of each component, we set up the following experiment settings: (1) Removing the loss, (2) removing the loss, (3) removing the suppression strategy, (4) removing the Gaussian filter, (5) applying twice suppression (in contrast to one-time). Detailed results are presented in Fig. 6b. As shown, our full model achieves a balanced performance between visual fidelity and editing effectiveness for both combined and independent concepts. Removing either the or loss results in a significant decrease in image-alignment for both Concept1 and Concept2. Similarly, the removal of the Gaussian filter leads to a notable reduction in image-alignment for combined concepts. No suppression significantly harms image-alignment for Concept2, confirming the benefits of sharper boundaries in for understanding multiple concepts (as explained in Sec. 3.3). Meanwhile, this also leads to lower text-alignment for both Concept1 and Concept2. Furthermore, applying twice suppression has detrimental effects on image-alignment as it filters out important information.
On the other hand, there are two design choices worth considering. As indicated in , averaging all scales of attention layers, instead of just using the scale, could potentially yield improved attribution maps for each input word. Therefore, we explore (1) impose constraints on the average of all scales attention layers. Additionally, we investigate releasing more parameters, specifically (2) updating the , , and matrices within the cross-attention units (in contrast to our approach, which only updates the and ). As depicted in Fig. 6b, operating on all scales of attention layers resulted in the model’s inability to reconstruct Concept2. Updating , , and does help the model remember the appearances of concepts but leads to a significant decrease in text-alignment. This suggests that updating more parameters does not preserve the good features of the pre-trained model.
4 Applications
Personalized concept inpainting. With any image and its corresponding mask, our method can seamlessly integrate learned concepts into the masked region while preserving the rest of the image, as shown in Fig. 7. Users can effortlessly perform inpainting by simply modifying the text prompt, thanks to our method’s conversion of concepts into new word embeddings.
Compatible with LoRA . LoRA techniques, actively discussed in the community, such as CivitAI , have gained popularity for enhancing specific capabilities of T2I models, such as improving the ability to refine images. LoRA adds small, trainable parameters to the frozen T2I models for fine-tuning, and our method is orthogonal with it. Therefore, we combine the LoRA with our trained model to unlock a wider range of applications, as shown in Fig. 8. This combination is akin to domain-specific pre-training on a large dataset before personalization , with the added benefit of having access to a wealth of readily available LoRA parameters in the community.
Extending to three concepts. We explore the application of our method to the more challenging task of capturing three concepts from a single image, as shown in Fig. 9. In this scenario, we employ the loss for each pair of the three class tokens to disentangle these concepts.
Conclusions and Limitations
We propose the DisenDiff to mimic multiple concepts from a single image. We introduce constraints on the cross-attention units to attain precise attention maps for crucial tokens, mitigating the overfitting to the single image and accurately capturing concept appearances. Consequently, our method enables diverse edits involving combined or independent concepts while enhancing the visual similarity between the synthesized images and the input image. Furthermore, we show the flexibility of our method by evaluating several applications.
Limitations. Disentangling fine-grained categories becomes notably challenging when two subjects from the same category co-exist in a single image, such as Golden Retriever and Border Collie dogs. Additionally, while our method can handle images with three concepts, its performance degrades considerably. This can be attributed to the limitations of existing T2I models in such scenarios, as well as the need for algorithm adjustments to address these specific challenges. We believe that there is considerable room to enhance the performance in these complex tasks.
Acknowledgment. This work is supported by Shanghai Science and Technology Program ”Federated based cross-domain and cross-task incremental learning” under Grant No. 21511100800, Natural Science Foundation of China under Grant No. 62076094 and No. 62201341.
References
Experiments
Additional qualitative results. Further comparisons, including Textual Inversion (TI) , are illustrated in Figure 11 (independent concepts) and Figure 12 (combined concepts). Evidently, the concepts synthesized by TI differ significantly from the input image, affirming the quantitative analysis in Sec. 4.2.
Detailed quantitative results on ten datasets. As shown in Tab. 1, our method consistently attains the highest image-alignment across most datasets while maintaining favorable text-alignment compared to the three baselines.
Attention map visualization of ablation studies. The attention maps for the component ablations are presented in Fig. 13, encompassing the following scenarios: (1) Removing the loss, (2) removing the loss, (3) using (i.e., in Sec. 3.3) instead of , (4) removing the suppression strategy, (5) applying twice suppression, (6) removing the Gaussian filter. Observing Fig. 13 reveals the following insights: (1) Without , new modifiers tend to focus on incorrect classes or vague regions; (2) Absence of results in interdependence among learned class tokens, especially the “cat” token; (3) Sole reliance on leads to tiny activation areas for crucial tokens; (4) Removal of the suppression strategy introduces unnecessary activations for new modifiers, apart from their corresponding class regions; (5) Applying twice suppression causes the loss of vital information for new modifiers, (e.g., the attention of is obviously smaller than the “dog”); (6) The absence of the Gaussian filter may cause new modifiers to lack specific attributes related to the concepts, such as the attention on the mouth part for in the specific dog instance. In summary, our full method generates independent and comprehensive attention maps for crucial tokens.
Implementation and Experiment Details
Datasets. We present each training image in Fig. 10.
Textual Inversion . We utilized the implementation from with 5000 training steps, a batch size of 4, and a learning rate of 0.0005. The input prompt, originally “A photo of ” in Textual Inversion, is modified to “A photo of and ”. The two new words ( and ) are initialized with the classes from the input image. For example, if the image contains a cat and a dog, and token embeddings are initialized as the pre-trained “cat” and “dog” token embeddings.
DreamBooth . We employ the implementation from with 250 training steps, a batch size of 2, and a learning rate of . The input prompt is “ [class1] and [class2]”, consistent with our setting in Sec. 3.1. Additionally, we generate 1000 “a [class1] and a [class2]” images using the pre-trained model . New modifiers are initialized as rare token embeddings.
Custom Diffusion . We employ the official implementation with 250 training steps, a batch size of 8, and a learning rate of . The input prompt is also “ [class1] and [class2]”, and modifiers are also initialized as rare token embeddings. For regularization, 200 images are selected using clip-retrieval with the caption “a [class1] and a [class2]”. We apply the default data augmentation in Custom Diffusion.
DisenDiff (ours). Implementation details are described in Sec. 4.1. For the total loss in Eq. 6, the weight of is set to in all experiments. The weight of defaults to and occasionally adjusts to for specific cases.