Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection Guidance
Wenhao Sun, Benlei Cui, Xue-Mei Dong, Jingqun Tang
Introduction
The widespread adoption of diffusion models (DMs) (Ho, Jain, and Abbeel 2020; Song et al. 2021; He et al. 2024; Liu et al. 2024c) in recent years has enabled the generation of high-quality images that match the quality of real photos and provide a realistic visualization based on user specifications. This raises a natural question of whether the image-generating capabilities of these models can be harnessed to remove objects of interest from images. Such a task, termed object removal (Yu et al. 2018; Suvorov et al. 2022), represents a specialized form of image inpainting, and requires addressing two critical aspects. Firstly, the user-specified object (usually given as a binary mask) must be successfully and effectively removed from the image. Secondly, the mask area must be filled with content that is realistic, plausible, and appropriate to maintain overall coherence within the image.
Traditional approaches for object removal are the patch-based methods (Guo et al. 2018; Lu et al. 2018), which fill in the missing regions after removal by searching for well-matched replacement patches ( candidate patches) in the undamaged part of the image and copying them to the corresponding removal locations. However, such processing methods often lead to inconsistency and unnaturally between the removed region and its surroundings. In recent years, convolutional neural networks (CNNs) have demonstrated considerable potential for object removal tasks. However, CNNs-based methods (Yan et al. 2018; Oleksii 2019; Suvorov et al. 2022) typically utilize a fixed-size convolutional kernel or network structure, which constrains the perceptual range of the model and the utilization of contextual information (Fang et al. 2023a; Xu et al. 2024; Fang et al. 2025). Consequently, the model’s performance is sub-optimal when confronted with large-scale removal or complex scenes.
With the rapid development of generative models (Shen et al. 2024c; Zhang et al. 2024b; Wei and Zhang 2024; Yuan et al. 2024; Zhang et al. 2024a; Wang et al. 2025) in deep learning(Tang et al. 2022a; Shen et al. 2023a; Fang et al. 2024a, d; Liu et al. 2024b, 2025; Li et al. 2025), a proliferation of generative models has been applied to object removal. Among these, the most common are generative adversarial network (GAN) (Goodfellow et al. 2014)-based methods and DMs-based methods. GAN-based methods (Chen and Hu 2019; Shin et al. 2020) employ neural networks of varying granularity, with the context-focused module exhibiting robust performance and efficacy in image inpainting. However, their training is inherently slow and unstable, and they are susceptible to issues such as mode collapse or failure to converge (Salimans et al. 2016).
In current times, DMs have made new waves in the field of deep generative models, broken the long-held dominance of GANs, and achieved new state-of-the-art performance in many computer vision tasks (Shen et al. 2024a, b, c; Shen and Tang 2024; Zhao et al. 2024c). The most prevalent open-source pre-trained model in DMs is Stable Diffusion (SD) (Rombach et al. 2022), which is a pre-trained latent diffusion model. To apply SD to the object removal task, fine-tuned from SD, SD-inpainting (Rombach et al. 2022) was developed into an end-to-end model with a particular focus on inpainting, to incorporate a mask as an additional condition within the model. However, even after spending a considerable cost in terms of resources, its object removal ability is not stable, and it often fails to completely remove the object or generates random artifacts(as shown in Figure 4). An additional methodology entails guiding the model to perform object removal via prompt instruction (Yildirim et al. 2023; Brooks, Holynski, and Efros 2023). The downside of this method is that to achieve a satisfactory result, these models often necessitate a considerable degree of prompt engineering and fail to allow for accurate interaction even with a mask. Additionally, they often necessitate substantial resources for fine-tuning.
To address these problems, we propose a tuning-free method, Attentive Eraser, a simple yet highly effective method for mask-guided object removal. This method ensures that during the reverse diffusion denoising process, the content generated within the mask tends to focus on the background rather than the foreground object itself. This is achieved by modifying the self-attention mechanism in the SD model and utilizing it to steer the sampling process. We show that when Attentive Eraser is combined with the prevailing diffusion-based inpainting pipelines (Couairon et al. 2023; Avrahami, Fried, and Lischinski 2023), these pipelines enable stable and reliable object removal, fully exploiting the massive prior knowledge in the pre-trained SD model to unleash its potential for object removal (as shown in Figure 1). The main contributions of our work are presented as follows:
We propose a tuning-free method Attentive Eraser to unleash DM’s object removal potential, which comprises two components: (1) Attention Activation and Suppression (AAS), a self-attention-modified method that enables the generation of images with enhanced attention to the background while simultaneously reducing attention to the foreground object. (2) Self-Attention Redirection Guidance (SARG), a novel sampling guidance method that utilizes the proposed AAS to steer sampling towards the object removal direction.
Experiments and user studies demonstrate the effectiveness, robustness, and scalability of our method, with both removal quality and stability surpassing SOTA methods.
Related Works
Existing diffusion model-based object removal methods can be classified into two categories, tuning-free (Zhao et al. 2024b) vs. training-based (Fang et al. 2023b), depending on whether they require fine-tuning or not. In the case of the training-based methods, DreamInpainter (Xie et al. 2023b) captures the identity of an object and removes it by introducing the discriminative token selection module. Powerpaint (Zhuang et al. 2023) introduces learnable task prompts for object removal tasks. Inst-Inpaint (Yildirim et al. 2023) constructs a dataset for object removal, and uses it to fine-tune the pre-trained diffusion model. There are other instruction-based methods achieving object removal via textual commands (Huang et al. 2024; Yang et al. 2024b; Geng et al. 2024). In the case of the tuning-free methods, Blended Diffusion (Avrahami, Fried, and Lischinski 2023) and ZONE (Li et al. 2024) perform local text-guided image manipulations by introducing text conditions to the diffusion sampling process. Magicremover (Yang et al. 2023) implements object removal by modifying cross-attention to direct diffusion model sampling. SuppressEOT (Li et al. 2023) suppresses negative target generation by focusing on the manipulation of text embeddings. However, these methods can lead to artifacts in the final result or incomplete removal of the target due to the stochastic nature of the diffusion model itself and imprecise guiding operations. To address the above issues and to avoid consuming resources for training, we propose a tuning-free method SARG to gradually steer the diffusion process towards object removal.
Sampling guidance for diffusion models
Sampling guidance for diffusion models involves techniques that steer the sampling process toward desired outcomes. Classifier guidance (Dhariwal and Nichol 2021) involves the incorporation of an additional trained classifier to generate samples of the desired category. Unlike the former, Classifier-free Guidance (Ho and Salimans 2021) does not rely on an external classifier but instead constructs an implicit classifier to guide the generation process. There are two methods that combine self-attention with guidance, SAG (Hong et al. 2023) and PAG (Ahn et al. 2024), which utilize or modify the self-attention mechanism to guide the sampling process, thereby enhancing the quality of the generated images. Our work is similar to PAG in that it modifies the self-attention map to guide sampling, but the purpose and approach to modification are different.
Preliminaries
DMs are a class of probabilistic generative models that learn a given data distribution by progressively adding noise to the data to destroy its structure and then learning a corresponding inverse process of a fixed Markov chain of length T to denoise it. Specifically, given a set of data , the forward process could be formulated by
where denotes the time step of diffusion process, is the noisy data at step , is the variance schedule at step and represents the level of noise.
Starting from , the reverse process aims to obtain a true sample by iterative sampling from . Unfortunately, this probability is intractable, therefore, a deep neural network with parameter is used to fit it:
proposed by Ho(Ho, Jain, and Abbeel 2020), a U-net (Ronneberger, Fischer, and Brox 2015) is trained to predict the noise that is introduced to to obtain , by minimizing the following object:
After training, a sample can be generated following the reverse process from .
Self-Attention in Stable Diffusion
Guidance
A key advantage of diffusion models is the ability to integrate additional information into the iterative inference process for guiding the sampling process, and the guidance can be generalized as any time-dependent energy function from the score-based perspective. Modifying with this energy function can guide the sampling process towards generating samples from a specifically conditioned distribution, formulated as:
where represents conditional information, is an energy function and represents the imaginary labels for the desirable sample and is the guidance scale. There are many forms of (Nichol et al. 2021; Dhariwal and Nichol 2021; Ho and Salimans 2021; Bansal et al. 2023; Epstein et al. 2023; Mo et al. 2024), the most prevalent of which is classifier-free guidance (Ho and Salimans 2021), where represents textual information (Liu et al. 2023; Fang et al. 2024b, c), and .
Methodology
The overall framework diagram of the proposed method is depicted in Figure 2. There are two principal components: AAS and SARG, which will be elucidated in more detail in the following sections.
Attention Activation and Suppression
We define to reflect the relevance of the content to be generated in the foreground object area to the background, while information about the appearance of the foreground object is reflected in .
In the object removal task, we are dealing with foreground objects, and the background should remain the same. As shown in Figure 3, after DDIM inversion (Song, Meng, and Ermon 2020), we utilize PCA (Maćkiewicz and Ratajczak 1993) and clustering to visualize the average self-attention maps over all time steps for different layers during the reverse denoising process. It can be observed that self-attention maps resemble a semantic layout map of the components of the image (Yang et al. 2024a), and there is a clear distinction between the self-attention corresponding to the generation of the foreground object and background. Consequently, to facilitate object removal during the generation process, an intuitive approach would be to ”blend” the self-attention of foreground objects into the background, thus allowing them to be clustered together. In other words, the region corresponding to the foreground object should be generated with a greater degree of reference to the background region than to itself during the generation process. This implies that the attention of the region within the mask to the background region should be increased and to itself should be decreased. Furthermore, the background region is fixed during the generation process and should remain unaffected by the changes in the generated content of the foreground area. Thus, the attention of the background region to the foreground region should also be decreased.
Combining the above analysis, we propose an approach that is both simple and effective: AAS (as shown in Figure 2(a)). Activation refers to increasing , which serves to enhance the attention of the foreground-generating region to the background. In contrast, Suppression refers to decreasing and , which entails the suppression of the foreground region’s information about its appearance and its effect on the background. Given the intrinsic characteristics of the Softmax function, AAS can be simply achieved by assigning to , thereby the original semantic information of the foreground objects is progressively obliterated throughout the denoising process. In practice, the aforementioned operation is achieved by the following equation:
where represents the corresponding value matrix for the time step of layer .
Nevertheless, one of the limitations of the aforementioned theory is that if the background contains content that is analogous to the foreground object, due to the inherent nature of self-attention, the attention in that particular part of the generative process will be higher than in other regions, while the above theory exacerbates this phenomenon, ultimately leading to incomplete object removal (see an example on the right side of Figure 2(a)). Accordingly, to reduce the attention devoted to similar objects and disperse it to other regions, we employ a straightforward method of reducing the variance of , which is referenced in this paper as SS. To avoid interfering with the process of generating the background, we address the foreground and background generation in separate phases:
where is the suppression factor less than 1. Finally, to guarantee that the aforementioned operations are executed on the appropriate corresponding foreground and background regions, we integrate the two outputs and to obtain the final output according to :
To ensure minimal impact on the subsequent generation process, we apply SS at the beginning of the denoising process timesteps, for , and still use Eq.(11), Eq.(12) to get output for , where denotes the diffusion steps and signifies the final time-step of SS. In the following, we denote the U-net processed by the AAS approach as .
Self-Attention Redirection Guidance
To further enhance the capability of object removal as well as the overall quality of the generated images, inspired by PAG (Ahn et al. 2024), can be seen as a form of perturbation during the epsilon prediction process, we can use it to steer the sampling process towards the desirable direction. Therefore, the final predicted noise at each time step can be defined as follows:
where is the removal guidance scale. Subsequently, the next time step output latent is obtained by sampling using the modified noise . In this paper, we refer to the aforementioned guidance process as SARG.
Through the iterative inference guidance, the sampling direction of the generative process will be altered, causing the distribution of the noisy latent to shift towards the object removal direction we have specified, thereby enhancing the capability of removal and the quality of the final generated images. For a more detailed analysis refer to Appendix A.
Experiments
We apply our method on all mainstream versions of Stable Diffusion (1.5, 2.1, and XL1.0) with two prevailing diffusion-based inpainting pipelines (Couairon et al. 2023; Avrahami, Fried, and Lischinski 2023) to evaluate its generalization across various diffusion model architectures. Based on the randomness, we refer to pipelines as the stochastic inpainting pipeline (SIP) and the deterministic inpainting pipeline (DIP), respectively. Detailed descriptions of SIP and DIP are provided in Appendix B, with further experimental details available in Appendix C.
Baseline
We select the state-of-the-art image inpainting methods as our baselines, including two mask-guided approaches SD-Inpaint (Rombach et al. 2022), LAMA (Suvorov et al. 2022) and two text-guided approaches Inst-Inpaint (Yildirim et al. 2023), Powerpaint (Zhuang et al. 2023), to demonstrate the efficacy of our method, we have also incorporated SD2.1 with SIP into the baseline for comparative purposes.
Testing Datasets
We evaluate our method on a common segmentation dataset OpenImages V5 (Kuznetsova et al. 2018), which contains both the mask information and the text information of the corresponding object of the mask. This facilitates a comprehensive comparison of the entire baseline. We randomly select 10000 sets of data from the OpenImages V5 test set as the testing datasets, a set of data including the original image and the corresponding mask, segmentation bounding box, and segmentation class labels.
Evaluation Metrics
We first use two common evaluation metrics FID and LPIPS to assess the quality of the generated images following LAMA(Suvorov et al. 2022) setup, which can indicate the global visual quality of the image. To further assess the quality of the generated content in the mask region, we adopt the metrics Local-FID to assess the local visual quality of the image following (Xie et al. 2023a). To assess the effectiveness of object removal, we select CLIP consensus as the evaluation metric following (Wasserman et al. 2024), which enables the evaluation of the consistent diversity of the removal effect. High diversity is often seen as a sign of failed removal, with random objects appearing in the foreground area. Finally, to indicate the degree of object removal, we calculate the CLIP score (Radford et al. 2021; Lu et al. 2024; Liu, Li, and Yu 2024) by taking the foreground region patch and the prompt ”background”. The greater the value, the greater the degree of alignment between the removed region and the background, effectively indicating the degree of removal.
Qualitative and Quantitative Results
The quantitative analysis results are shown in Table 1. For global quality metrics FID and LPIPS, our method is at an average level, but these two metrics do not adequately reflect the effectiveness of object removal. Subsequently, we can observe from the local FID that our method has superior performance in the local removal area. Meanwhile, the CLIP consensus indicates the instability of other diffusion-based methods, and the CLIP score demonstrates that our method effectively removes the object and repaints the foreground area that is highly aligned with the background, even reaching a competitive level with LAMA, which is a Fast Fourier Convolution-based inpainting model. Qualitative results are shown in Figure 4, where we can observe the significant differences between our method and others. LAMA, due to its lack of generative capability, successfully removes the object but produces noticeably blurry content. Other diffusion-based methods share a common issue: the instability of removal, which often leads to the generation of random artifacts. To further substantiate this issue, we conducted experiments on the stability of removal. Figure 5 presents the results of removal using three distinct random seeds for each method. It can be observed that our method achieves stable erasure across various SD models, generating more consistent content, whereas other methods have struggled to maintain stable removal of the object.
User Study and GPT-4o Evaluation
Due to the absence of effective metrics for the object removal task, the metrics mentioned above may not be sufficient to demonstrate the superiority of our method. Therefore, to further substantiate the effectiveness of our approach, we conduct a user preference study. Table 2 presents the user preferences for various methods, revealing consistent results with the quantitative results and highlighting that our method is strongly preferred over other methods. Furthermore, we design fairly and reasonably prompts, utilizing GPT-4o (OpenAI 2024) to conduct a further assessment of object removal performance between our method and the runner-up method LAMA. The results also indicate that our method significantly outperforms LAMA, demonstrating exceptional performance. Please refer to Appendix D for more details and visualizations of user study and GPT evaluation.
Ablations
To validate the effectiveness of the proposed Attentive Eraser, we conduct ablation studies. We use SD2.1 with SIP as the baseline for comparison, Figure 6 provides a visual representation of the ablation study concerning our method’s components. Figure 6(a) shows that the application of AAS alone cannot completely remove the foreground object, but integrating it with the sampling process through SARG can effectively remove the object and generate content consistent with the background. At the same time, we also verify the impact of SS, and it can be seen that SS effectively suppresses the generation of similar objects while maintaining the removal efficacy of the general image. As shown in In Figure 6(b), we visualize the heatmaps of the top-1 component of the self-attention maps at each step of the denoising process after SVD (Kalman 1996), demonstrating that SARG gradually, as previously stated, ”blends” the foreground objects’ self-attention into the background to remove objects. In Figure 6(c), we discuss the effect of two parameters (removal guidance and suppression factor ) upon the removal process. It is depicted that as decreases, the generation of similar objects decreases progressively, thereby reaffirming the efficacy of SS. On the other hand, the intensity of the removal process escalates with an increase in . This suggests that acts as a pivotal control in modulating the strength of the removal, allowing for a more nuanced and tailored approach to removing objects.
Conclusion
We present a novel tuning-free method Attentive Eraser, which adeptly harnesses the rich repository of prior knowledge embedded within pre-trained diffusion models for the object removal task. Extensive experiments and user studies demonstrate the stability, effectiveness, and scalability of our proposed method, and also reveal that our method significantly outperforms existing methods.
Acknowledgments
This work is supported in part by the Summit Advancement Disciplines of Zhejiang Province (Zhejiang Gongshang University - Statistics) and ”Digital+” discipline construction management project of Zhejiang Gongshang University (SZJ2022C011).
References
Appendix
For a more thorough comprehension of our method, we have expanded on the details in the ensuing sections.
Similar to the derivation in PAG (Ahn et al. 2024), we introduce an implicit discriminator, denoted by , that distinguishes desirable samples that follow the real data distribution from undesirable samples in the diffusion process. In our work, the original model predictions are regarded as undesirable samples and the AAS-processed model predictions are regarded as desirable samples. The implicit discriminator can be defined as:
where and represent the imaginary labels for the desirable sample and the undesirable sample, respectively.
Subsequently, analogous to WGAN (Arjovsky, Chintala, and Bottou 2017; Wu et al. 2018), our generator loss for the implicit discriminator, , is established as our energy function and its derivative is calculated as:
The diffusion sampling process can then be defined as:
where the pre-trained score estimation network and the AAS processed network are approximations of and , respectively.
B. Detailed Description of SIP and DIP
When inpainting real images, two pipelines are commonly employed, which were proposed by BLD (Avrahami, Fried, and Lischinski 2023) and DiffEdit (Couairon et al. 2023) respectively. We refer to them as the stochastic inpainting process (SIP) and the deterministic inpainting process (DIP) based on the randomness inherent in their processes. The SIP introduces randomness into the generation process by incorporating Gaussian noise. However, the DIP retains the original image information through DDIM inversion (Song, Meng, and Ermon 2020; Dhariwal and Nichol 2021), which, like DDIM sampling, is a deterministic process and thus does not involve randomness in the generation process. A schematic diagram of SIP and DIP is shown in Figure 7. The algorithms of SIP and DIP after applying SARG for object removal are presented in Algorithms 1 and 2, respectively.
C. Additional Experimental Details
Firstly, we present the implementation details of our theory: in all SD models, we adopt DDIM sampling as the default sampling method and apply SARG to all time steps. Concurrently, based on previous research that has established the significant impact of the decoder part in U-net on the appearance information of generated images (Zhang, Xiao, and Huang 2023; Tumanyan et al. 2023; Jiang et al. 2024), we integrate AAS into the decoder of U-net. We have also provided a visual comparison to substantiate this, as shown in Figure 8. Additionally, starting from the perspective of ensuring the erasure capability and quality as much as possible, we provide the corresponding evaluative parameter settings of SIP and DIP with SARG in Table 3.
Below, we will provide a brief overview of the comparative methods in the baseline as well as the corresponding experimental setup:
SD-Inpaint is finetuned from Stable Diffusion and is capable of accepting a mask as input for inpainting. In the experiments, we integrate this model with varying input conditions into the baseline, corresponding to scenarios with only mask input and those with both mask and text input (Here ”background” is designated as the prompt, while the object label serves as the negative prompt, and the guidance scale is set to 7.5).
LAMA leverages Fast Fourier Convolutions (FFCs) to expand the receptive field, along with an effective perceptual loss and aggressive training mask generation strategy, achieving high-quality inpainting on large missing areas. In the experiments, we incorporate the most powerful model as per the official documentation, Big-LAMA, into the baseline.
Inst-Inpaint trains a novel conditional diffusion model to implement object removal based on the instructions given as text prompts. In the experiments, we adhere to the original paper settings by designating the text instruction as: ”Remove the [object label] at the [location].”
PowerPaint achieves versatile and high-quality image inpainting by utilizing learnable task prompts and specialized fine-tuning strategies. In the experiments, we employ the section of the official code related to object removal and following the suggestions provided in their demo, set the guidance scale to 12.
During the evaluation process, aside from Inst-Inpaint requiring 256256 image inputs, we utilize 512512 images as the input for each method. When calculating the metrics, all output images are resized to 512512. Given the necessity for the CLIP consensus metric to assess image generation across various random seeds and subsequently calculate the standard deviation of the CLIP embeddings within the foreground object region. We extract 6000 images from the testing dataset of 10000 and generate corresponding results using the random seeds (123, 321, 777) for the computation of this metric. The remaining metrics are calculated using results from the testing dataset of 10000 images generated with the random seed 123.
D. User Study and GPT-4o Evaluation
In our user study experiment, we recruited 10 participants to assess each image and determine which one had the best object removal effect based on the provided reference evaluation criteria. Each participant was assigned 100 comparison images obtained through random sampling from the final results. At the same time, we ensured that each round of evaluation was conducted with randomized order and anonymous selection. Finally, we calculated the average user preference percentage for each method. A print screen is provided in Figure 9.
In the comparative experiment utilizing GPT-4o against the runner-up LAMA, we tasked GPT with selecting the image with the best object removal effect based on a fairly and reasonably designed prompt. The prompt was as follows: ”You are an expert in evaluating generated images. There are two images with their corresponding masks. Please assess the following aspects: 1. Whether the object within the mask has been effectively removed and consistent content with the background has been generated within the mask area. 2. The realism of the generated content within the mask. Based on these criteria, please tell me which image is better.” We conducted experiments with three different random seeds, randomly selecting 1000 pairs of images each time, and finally provided the selection rate based on the results of these 3000 image pairs.
E. Robustness and Scalability Analysis
Furthermore, we demonstrate the robustness of our method to input masks and its scalability to other pre-trained models. As shown in Figure 10, we utilize three mask types varying in refinement levels to assess the robustness of our method: instance segmentation masks, segmentation bounding box masks, and hand-drawn masks. It can be observed that even with the coarse hand-drawn masks, our method effectively removes the target and generates a plausible background, demonstrating that the performance of our method is not hindered by the mask’s level of refinement. Additionally, as shown in Figure 11, our method is not only applicable to the pre-trained models generating natural images ( SD1.5, 2.1) but can also be extended to models for anime and cartoon images, such as solarsync (Civital 2024).
F. Limitations
Our method has two primary drawbacks. Firstly, it shares a common issue with guidance methods, namely the increased inference time due to the necessity of two times of U-net predicting noise. Secondly, when the mask area is too large, the scarcity of referable background areas may result in the poor reconstruction of the removal region, leading to the generation of artifacts, as shown in Figure 12. We will endeavour to overcome these limitations in the development of generative AI (Feng et al. 2025; Tang et al. 2022b, 2023, a, 2024a, 2024b; Zhao et al. 2024a).
G. Additional Results
In this section, we provide more samples of the object removal results in Figure 13 and 14. By applying the Attentive Eraser to various inpainting pipelines and across different SD models, we demonstrate the robustness, effectiveness, and extensibility of our method. It successfully unleashes the potential for object removal in a multitude of pre-trained diffusion models.