Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

Bingyan Liu, Chengyu Wang, Tingfeng Cao, Kui Jia, Jun Huang

Introduction

Text-to-Image Synthesis (TIS) models, such as Stable Diffusion , DALL-E 2 , and Imagen , have demonstrated remarkable visual effects for text-to-image generation, capturing substantial attention from both academia and industry . These TIS models are trained on vast amounts of image-text pairs, such as Laion , and employ cutting-edge techniques, including large-scale pre-trained language models , variational auto-encoders , and diffusion models to achieve success in generating realistic images with vivid details. Specifically, Stable Diffusion stands out as a popular and extensively studied model, making significant contributions to the open-source community.

In addition to image generation, these TIS models possess powerful image editing capabilities, which hold great importance as they aim to modify images while ensuring realism, naturalness, and meeting human preferences. Text-guided Image Editing (TIE) involves modifying an input image based on a descriptive prompt. Existing TIE methods achieve remarkable effects in image translation, style transfer, and appearance replacement, as well as preserving the input structure and scene layout. To this end, Prompt-to-Prompt (P2P) modifies image regions by replacing cross-attention maps corresponding to the target edit words in the source prompt. Plug-and-Play (PnP) first extracts the spatial features and self-attention of the original image in the attention layers and then injects them into the target image generation process. Among these methods, attention layers play a crucial role in controlling the image layout and the relationship between the generated image and the input prompt. However, inappropriate modifications to attention layers can yield varied editing outcomes and even lead to editing failures. For example, as depicted in Figure 1, editing authentic images on cross-attention layers can result in editing failures; converting a man into a robot or changing the color of a car to red fails. Moreover, some operations in the above-mentioned methods can be revised and optimized.

In our paper, we explore attention map modification to gain comprehensive insights into the underlying mechanisms of TIE using diffusion-based models. Specifically, we focus on the attribution of TIE and ask the fundamental question: how does the modification of attention layers contribute to diffusion-based TIE? To answer this question, we carefully construct new datasets and meticulously investigate the impact of modifying the attention maps on the resulting images. This is accomplished by probe analysis and systematic exploration of attention map modification with different blocks in the diffusion model. We find that (1) editing cross-attention maps in diffusion models is optional for image editing. Replacing or refining cross-attention maps between the source and target image generation process is dispensable and can result in failed image editing. (2) The cross-attention map is not only a weight measure of the conditional prompt at the corresponding positions in the generated image but also contains the semantic features of the conditional token. Therefore, replacing the target image’s cross-attention map with the source image’s map may yield unexpected outcomes. (3) Self-attention maps are crucial to the success of the TIE task, as they reflect the association between image features and retain the spatial information of the image. Based on our findings, we propose a simplified and effective algorithm called Free-Prompt-Editing (FPE). FPE performs image editing by replacing the self-attention map in specific attention layers during denoising, without needing a source prompt. It is beneficial for real image editing scenarios. The contributions of our paper are as follows:

We conduct a comprehensive analysis of how attention layers impact image editing results in diffusion models and answer why TIE methods based on cross-attention map replacement can lead to unstable results.

We design experiments to prove that cross-attention maps not only serve as the weight of the corresponding token on the corresponding pixel but also contain the characteristic information of the token. In contrast, self-attention is crucial in ensuring that the edited image retains the original image’s layout information and shape details.

Based on our experimental findings, we simplify currently popular tuning-free image editing methods and propose FPE, making the image editing process simpler and more effective. In experimental tests over multiple datasets, FPE outperforms current popular methods.

Overall, our paper contributes to the understanding of attention maps in Stable Diffusion and provides a practical solution for overcoming the limitations of inaccurate TIE.

Related Works

Text-guided Image Editing (TIE) is a crucial task involving the modification of an input image with requirements expressed by texts. These approaches can be broadly categorized into two groups: tuning-free methods and fine-tuning based methods.

Tuning-free TIE methods aim to control the generated image in the denoising process. To achieve this goal, SDEdit uses the given guidance image as the initial noise in the denoising step, which leads to impressive results. Other methods operate in the feature space of diffusion models to achieve successful editing results. One notable example is P2P , which discovers that manipulating cross-attention layers allows for controlling the relationship between the spatial layout of the image and each word in the text. Null-text inversion further employs an optimization method to reconstruct the guidance image and utilizes P2P for real image editing. DiffEdit automatically generates a mask by comparing different text prompts to help guide the areas of the image that need editing. PnP focuses on spatial features and self-affinities to control the generated image’s structure without restricting interaction with the text. Additionally, MasaCtrl converts self-attention in diffusion models into a mutual and mask-guided self-attention strategy, enabling pose transformation. In this paper, we aim to provide in-depth insights into the attention layers of diffusion models and further propose a more streamlined tuning-free TIE approach.

2 Fine-tuning Based Methods

The core idea of fine-tuning-based TIE methods is to synthesize ideal new images by model fine-tuning over the knowledge of domain-specific data or by introducing additional guidance information . DreamBooth fine-tunes all the parameters in the diffusion model while keeping the text transformer frozen and utilizes generated images as the regularization dataset. Textual Inversion optimizes a new word embedding token for each concept. Imagic learns the approximate text embedding of the input image through tuning and then edits the posture of the object in the image by interpolating the approximate text embedding and the target text embedding. ControlNet and T2I-Adapter allow users to guide the generated images through input images by tuning additional network modules. Instructpix2pix fully fine-tunes the diffusion model by constructing image-text-image triples in the form of instructions, enabling users to edit authentic images using instruction prompts, such as “turn a man into a cyborg”. In contrast to these works, our method focuses on tuning-free techniques without the fine-tuning process.

Analysis on Cross and Self-Attention

In this section, we analyze how cross and self-attention maps in Stable Diffusion contribute to the effectiveness of TIE.

2 Self-Attention in Stable Diffusion

where dselfd_{self} is the dimension of KselfK_{self} and QselfQ_{self}. MselfM_{self} determines the weights assigned to the relevance of the ii-th and jj-th spatial features in the image and can affect the spatial layout and shape details of the generated image. Consequently, the self-attention map can be utilized to preserve the spatial structure characteristics of the original image throughout the image editing process.

3 Probing Analysis

Yet, the semantics of cross and self-attention maps remain unclear. Are these attention maps merely weight matrices, or do they contain feature information of the image? To answer these questions, we aim to explore the meaning of attention maps in diffusion models. Inspired by probing analysis methods in the field of NLP, we propose building datasets and training classification networks to explore the properties of attention maps. Our fundamental idea is that if a trained classifier can accurately classify attention maps from different categories, then the attention map contains meaningful feature representation of the category information. Therefore, we introduce a task-specific classifier on top of the diffusion model’s cross-attention and self-attention layers. This classifier is a two-layer MLP designed to predict specific semantic properties of the attention maps. To present the analysis results more visually, we utilize color adjectives and animal nouns to form prompt datasets, each containing ten categories. For the color adjective, there are two prompt formats: a <<color>> car and a <<color>> <<object>>. The prompt format for animal nouns is a/an <<animal>> standing in the park. After generating the prompts, we employ the probing method to extract the cross-attention maps corresponding to the words <<color>> and <<animal>>, along with the self-attention maps in the attention layers. Finally, by training and evaluating the performance of the classifiers, we gain insights into the semantic knowledge captured by the attention maps.

4 Probing Results on Cross-Attention Maps

What does the cross-attention map learn? We directly visualize the attention maps, as demonstrated in Figure 3. Each word in the prompt has a corresponding attention map associated with the image, indicating that the information related to the word exists in specific areas of the image. However, is this information exclusive to these areas? Referring to Equation 2, we observe that McrossM_{cross} is derived from KcrossK_{cross} and QcrossQ_{cross}, indicating that McrossM_{cross} carries information from both. To validate this hypothesis, we conduct probing experiments on McrossM_{cross}, with the results presented in Table 1. Due to space limitations, we show only the probing results for five colors and five animals from the last layer of the down, middle, and up blocks. As evident in Table 1, the trained classifier achieves high accuracy in both the color and animal classification tasks. For instance, the average accuracy for classifying “sheep” reaches 98%, and that for “orange” reaches 93%. These results demonstrate that the cross-attention map acts as a reliable category representation, indicating that it reflects not only weight information but also contains category-related features. This explains the failure of image editing using cross-attention map replacement. The upper part of Figure 4 illustrates the editing results obtained by replacing the cross-attention map of the corresponding word (“rabbit” and “coral”) at different cross-attention layers. It is apparent that when all layers are replaced, the editing results are the least satisfactory. The dog fails to transform completely into a rabbit, and the black car cannot turn into a coral car. Conversely, when the cross-attention map is left unaltered, correct editing results can be achieved. The complete and more additional experimental results are available in Section 8 in the Supplementary Material.

5 Probing Results on Self-Attention Maps

What does the self-attention map learn? Table 2 presents the results of the probing experiments. The results indicate that the trained classifier struggles to classify the self-attention map generated from images containing color prompts. For animals, the results are better, although not as precise as those using cross-attention maps. This discrepancy may be attributed to the irregular spatial structure present in the self-attention map corresponding to the color prompt. Conversely, the self-attention map corresponding to the animal prompt contains structural information of different animals, enabling the learning of category information through recognizing structural or contour features. As shown in the lower part of Figure 3, the first component of the horse’s self-attention map clearly expresses the outline information of the horse. The lower part of Figure 4 showcases our experimental results of operating on the self-attention map across different attention layers. When the self-attention map of all layers in the source image is replaced during the generation process of the target image, the resulting target image retains all the structural information from the original image but hinders successful editing. Conversely, if we do not replace the self-attention map, we obtain an image identical to that generated directly using the target prompt. As a compromise, replacing the self-attention map in Layers 4 to 14 allows for preserving the structural information of the original image to the greatest extent while ensuring successful editing. This experimental result further supports the idea that the self-attention map in Layers 4 to 14 does not serve as a reliable category representation but does contain valuable spatial structure information of the image.

6 Probing Results for Other Tokens

Do cross-attention maps corresponding to non-edited words contain category information? Furthermore, we explore the attention maps associated with non-edited words. This is relevant because within a text sequence, the text embedding for each word retains the contextual information of the sentence, particularly when a transformer-based text encoder is utilized. We employ the prompt data in the format of a <<color>> car for our probing experiments. The experimental results are presented in Table 3. The findings demonstrate that the article “a” does not encompass any category information of color. In contrast, the noun “car,” when modified by the color adjective, does contain color category information. Consequently, if we replace the cross-attention map corresponding to a non-edited word with the cross-attention map of the target image, color information may be introduced, ultimately resulting in editing failures. This observation is also evident from the experimental results in Figure 5, where replacing the cross-attention maps of non-edited words likewise leads to editing failures.

Our Approach

Based on our exploration of attention layers, we propose a more straightforward yet more stable and efficient approach named Free-Prompt-Editing (FPE). Let IsrcI_{src} be the image to be edited. Our goal is to synthesize a new desired image IdstI_{dst} based on the target prompt PdstP_{dst} while preserving the content and structure of the original image IsrcI_{src}. Current editing methods like P2P replace the cross-attention map in the source and target image generation process. This requires modifying the original prompt to find the corresponding attention map for replacement. However, this limitation prevents the direct application of P2P to editing real images, as they do not come with an original prompt.

Based on our exploration of attention layers, our core idea is to combine the layout and contents of IsrcI_{src} with the semantic information synthesized with the target prompt PdstP_{dst} to synthesize the desired image IdstI_{dst} that retains the structure and content information of the original image IsrcI_{src}. To achieve this, we adapt the self-attention hijack mechanism in the diffusion model’s attention layers 4 to 14 during the denoising process between the source and target images. For generated image editing, we substitute the target image’s self-attention map with the source image’s self-attention map during the diffusion denoising process. When working with actual images, we first obtain the necessary latents for reconstructing the real image by employing the inversion operation . Subsequently, during the editing process, we replace the self-attention map of the real image within the generation process of the target image. We can accomplish the TIE task for the following reasons: 1) the cross-attention mechanism facilitates the fusion of the synthetic image and the target prompt, allowing the target prompt and the image to be automatically aligned even without introducing the cross-attention map of the source prompt; 2) the self-attention map contains spatial layout and shape details of the source image, and the self-attention mechanism allows for the injection of structural information from the original image into the generated target image. Algorithms 1 and 2 present the pseudocode for our simplified method applied to generated and real images, respectively. FPE can also be combined with null text inversion for real image editing (refer to Section 10).

Experiments

Since there are no publicly available datasets specifically designed to verify the effectiveness of image editing algorithms, we construct two types of image-prompt pairs datasets: one for generated images and one for real images. The generated images dataset includes Car-fake-edit and ImageNet-fake-edit, where Car-fake-edit contains 756 prompt pairs, and ImageNet-fake-edit contains 1182 prompt pairs sampled from FlexIT and ImageNet . The real image datasets include Car-real-edit, sampled from the Stanford Car (CARS196) dataset , containing 3321 image-prompt pairs, and ImageNet-real-edit, which contains 1092 pairs. For more details, see section 7.2 in the Supplementary Material. In addition, we also use benchmarks constructed by PnP . These benchmarks contain two datasets: Wild-TI2I and ImageNet-R-TI2I. For generated images, Wild-TI2I contains 70 prompt pairs, and ImageNet-R-TI2I contains 150 pairs. For real images, Wild-TI2I contains 78 image-prompt pairs, and ImageNet-R-TI2I includes 150 pairs.

We utilize Clip Score (CS) and Clip Directional Similarity (CDS) to quantitatively analyze and compare our method with currently popular image editing algorithms. The underlying model for our experiments is Stable Diffusion 1.5https://huggingface.co/runwayml/stable-diffusion-v1-5. The experimental results of comparative methods are produced using the publicly disclosed codes from their original papers with unified random seeds.

2 Image Editing Results

We evaluate our method through quantitative and qualitative analyses. As illustrated in Figure 6, we showcase the editing outcomes of our method, demonstrating that it successfully transforms various attributes, styles, scenes, and categories of the original images.

In this section, we compare our work with state-of-the-art image editing methods, including (i) P2P (with null text inversion for the real image scene), (ii) PnP , (iii) SDEdit under two noise levels (0.5 and 0.75), (iv) DiffEdit , (v) MasaCtrl , (vi) Pix2pixzero , (vii) Shape-guided , and (viii) InstructPix2Pix . We further present the image editing results using other Stable Diffusion-based models to demonstrate the universality of our method, including Realistic-V2https://huggingface.co/SG161222/Realistic_Vision_V2.0, Deliberatehttps://huggingface.co/XpucT/Deliberate, and Anything-V4https://huggingface.co/xyn-ai/anything-v4.0.

Comparison to P2P We first compare our method with P2P for synthetic image editing scenes and P2P combined with null text inversion for real image scenes, both denoted as P2P. The experimental results are shown in Figure 7 and Table 4. In Figure 7, it is evident that when performing color transformation on a real image by modifying the cross-attention map, the editing fails. The editing results of P2P for car color tend to replicate the color (white) of the original image. Regarding the category conversion results for generated images, we observe that while P2P can accurately transform different animals, the edited results still retain appearances of sheep. This leads to an incomplete conversion for patterned animals such as giraffes, leopards, and tigers. Unlike P2P, our method operates only at the self-attention layers and is not susceptible to editing failures caused by modifications to the cross-attention map.

Comparison to Other Methods Further, we compare our method with other state-of-the-art (SOTA) image editing methods over the Wild-TI2I and ImageNet-R-TI2I benchmarks. The experimental results are presented in Figure 8 and Table 5. As shown in Figure 8, our method successfully converts different inputs for both real and synthetic images. In all examples, our method achieves high-fidelity editing that aligns with the target prompt while preserving the original image’s structural information to the greatest extent possible. In contrast, SDEdit and InstructPix2Pix struggle to preserve the structural information of the original image. SDEdit aligns the editing results better with the target prompt when there is high-level noise but fails in the presence of low-level noise. InstructPix2Pix retains consistency with the target prompt but loses the original structural information. DiffEdit and Pix2pix-zero also struggle to perform better editing based on the target prompt. Similarly, PnP achieves good editing results, but it is a two-step method that leads to significant computational overhead; editing a single image in a generated image editing scenario takes approximately 335.65 seconds. In contrast, our method only requires around 6.30 seconds on an A100 GPU with 40GB memory, as Table 5 indicates.

Table 5 presents the quantitative experimental results of different editing algorithms on the Wild-TI2I and ImageNet-R-TI2I benchmarks. From Table 5, it is evident that our method outperforms all others in terms of the CDS metric. This indicates that our method excels in preserving the spatial structure of the original image and performing editing according to the requirements of the target prompt, yielding superior results. Meanwhile, our method achieves a good balance between time consumption and effectiveness, as demonstrated in Table 5.

2.2 Results in Other TIS Models

We have applied our method to other TIS models based on Stable Diffusion-style frameworks to demonstrate its transferability. Figure 9 showcases the editing results of our method on Realistic-V2, Deliberate, and Anything-V4 TIS models. From these results, it can be observed that our method is capable of effectively editing images on other diffusion models as well. For example, it can transform a girl into a boy, change a boy’s age to 10 or 80, modify hairstyles, change hair colors, alter backgrounds, and switch categories.

3 Limitations and Discussion

Although our method employs probe analysis to elucidate the role of attention layers in the TIS model and proposes a novel method for editing images in multiple scenarios without complex operations, it still has some limitations. Firstly, our method is constrained by the generative capabilities of the TIS model. Our editing method will fail if the generative model cannot produce images consistent with the target prompt description. When editing real images, the original image must first be reconstructed. Some detailed information, especially facial details, may be lost during the reconstruction process, primarily due to the limitations of the VQ autoencoder . Optimizing the VQ autoencoder is beyond the scope of this paper, as our objective is to provide a simple and universal editing framework. Addressing these challenges will be part of our future work.

Conclusion

In this work, we utilized probe analysis and conducted experiments to elucidate the following insights on TIS models: the cross-attention map carries the semantic information of the prompt, which leads to the ineffectiveness of image editing methods that rely on it. On the contrary, the self-attention map captures the spatial structural information of the original image, playing an essential role in preserving the image’s inherent structure during editing. Based on our comprehensive analysis and empirical evidence, we have streamlined current image editing algorithms and proposed an innovative image editing approach. Our approach does not require additional tuning or the alignment of target and source prompts to achieve effective object or background editing in images. In extensive experiments across multiple datasets, our simplified method has outperformed existing image editing algorithms. Furthermore, our algorithm can be seamlessly adapted to other TIS models.

This work is partially supported by Alibaba Cloud through the Research Talent Program with South China University of Technology, and the Program for Guangdong Introducing Innovative and Entrepreneurial Teams (No. 2017ZT07X183).

References

Details of Data Collection

Cross-Attention Map Data: We have constructed five datasets by saving the cross-attention maps of the target word in the prompt, each containing 2,000 samples, for cross-attention map analysis. The prompts include color adjectives and animal nouns. Specifically, for color adjectives, we used two types of prompts: “a car” and “a ”, to construct the data. The prompt “a car” consists of ten categories of color words using 200 random seeds. We obtained the cross-attention maps by averaging the steps during the generation process. Each cross-attention map consists of 16 attention maps corresponding to the 16 attention layers of the diffusion model. The same procedure was followed for the construction of cross-attention maps in the format “a ”, but with two random seeds and 100 everyday objects. Similarly, we used the prompt format “a standing in the park”, sampled 200 random seeds, and constructed 2,000 samples for animal nouns. For the cross-attention maps corresponding to non-editing words, we used the prompt format “a car” and saved the cross-attention maps for the words “a” and “car.” The same approach was applied to the complex text templates.

Self-Attention Map Data: For the self-attention maps, we constructed two datasets, each containing 2,000 samples, using prompt formats “a car” and “a standing in the park” and sampling 200 random seeds. However, due to the large size of the self-attention maps, which are 4096×40964096\times 4096, 1024×10241024\times 1024, 256×256256\times 256, 64×6464\times 64, and 8×88\times 8 in dimensions, we resized the layers with dimensions larger than or equal to 256×256256\times 256 to 256×256256\times 256 for the probing analysis experiments.

The 10 color, 10 animal, 100 object categories, and 12 more complex text templates are as follows:

2 Data Collection for Editing Experiments

Car-fake-edit: Using the prompt format “a car” and 28 color words, we constructed Car-fake-edit, which contains 756 prompt pairs.

Car-real-edit: We sampled 123 real images from the Stanford Car dataset with image sizes ranging from 512 to 768. We then used CLIP to align these images with the 28 color words, resulting in the source prompts for the original images in the format “a car”. We constructed 27 target prompts for each image, resulting in 3,321 image-text pairs.

ImageNet-fake-edit: The paper FlexIT proposes constructing a validation set from the ImageNet validation set. The method for constructing the test set is as follows: A subset of 273 labeled categories is taken from ImageNet, and these categories are manually divided into 47 clusters. During testing, only transformations within the same cluster are allowed; for example, a cat to a dog, but not a laptop to a butterfly. For each label T, eight random categories are sampled from the cluster to serve as queries, resulting in 2,184 queries for the 273 categories. We utilized their test dataset, which consists of 1,092 queries, and added ten animal categories to construct 1,182 prompt pairs using prompt templates.

ImageNet-real-edit: Based on ImageNet-fake-edit, we used the prompt format “a photo of a/an {}” and constructed 1,092 image-text pairs for the real images. The color words, animal list, and prompt templates are as follows:

Probing Analysis Results

In this section, we present the complete experimental results for probe analysis and supplementary experiments, shown in Tables 6, 7, 8 and 9.

Table 6 presents the probe analysis experiment results for cross-attention maps with four sub-tables. The first three sub-tables, from top to bottom, correspond to the prompt formats “a/an standing in the park,” “a ,” and “a car.” The bottom sub-table corresponds to the experiment with training data “a car” and test data “a .”

For the experiments within the distribution, regardless of the prompt format, the cross-attention maps corresponding to the words can be accurately classified by the classifier. Even on out-of-distribution data, an average accuracy of around 50% can be achieved. This indicates that the diffusion model’s 16 layers of cross-attention maps can serve as good feature representations, containing semantic information about the corresponding words. This is consistent with the conclusion in the main text that the cross-attention maps are both weight matrices and rich in semantic information. Table 7 presents the probe experiment results for non-target words, aiming to verify whether the attention maps corresponding to words other than the target word in the prompt contain the semantics of the target word. We conducted experiments using the simple prompt format “a car.”

Table 8 presents the complete probe experiment results for self-attention maps. Compared to cross-attention maps, self-attention maps are not directly usable as feature representations for classification, especially for prompts with color adjectives where the classification performance could be improved. However, compared to color adjective prompts, higher classification accuracy is observed for prompts with animal adjectives, which may be related to the self-attention maps’ ability to represent objects’ appearance contours in animal class images.

We expanded our investigation through probing analyses utilizing a set of twelve intricate text templates to eliminate the potential experimental bias that may arise from using consistently simple and regular templates in previous experiments. Examples of these templates include phrases such as “a painting of a wooden car” and “a photo of a car and a dog”, among others. The findings, as presented in Table 9, corroborate the conclusions drawn from experiments using simpler text prompts, indicating consistency in results across varying levels of template complexity.

Impact of Replacement Steps

In this section, we conduct ablation experiments on different attention layers of cross-attention and self-attention maps under various denoising steps. The experimental results are presented in Figure 10 and Figure 11.

When only the cross-attention map is replaced during editing, the target image loses the structural information of the original image. As shown in Figure 10, although the leopard in the target image bears a resemblance to the dog in the original image, notable modifications in the background, particularly the grass, are observed. Similarly, when performing a color conversion on a car using only the cross-attention map, the original image’s structural information is lost, leading to a car that lacks its original structure and takes on a brown appearance.

When both the cross-attention map and the self-attention map are replaced simultaneously, the results depicted in Figures 10 and 11 are obtained by keeping the cross-replace ratio fixed at 0.8 while varying the self-replace ratio. The replacement of the cross-attention map aids in swiftly identifying the target region and reconstructing the structure of the original image. However, it also introduces the original image’s feature information, particularly when replacing attention maps in all layers, which significantly includes the original features. As illustrated in Figures 10 and 11, the leopard exhibits attributes similar to a dog, while the car retains its blue color.

When only the self-attention map is replaced with a low self-replace ratio, such as 0.1, the resulting target image closely resembles the one obtained using the target prompt directly. However, when the self-attention map is replaced in all attention layers and for 90% of the denoising steps, a target image that closely matches the original image is generated, as depicted in the top left corner of Figure 10. A more balanced approach involves replacing the self-attention map in layers 4–14 with replacement ratios ranging from 0.4 to 0.8, resulting in more favorable outcomes.

Real Image Editing with Null-Text Inversion

Algorithm 3 outlines the pseudo-code for editing real images using the Freeprompt Editing (FPE) combined with Null-Text Inversion . Figure 12 presents the experimental results of editing real images using DDIM Inversion and Null-Text Inversion. Both methods effectively modify the original image based on the target text.