InstantSwap: Fast Customized Concept Swapping across Sharp Shape Differences

Chenyang Zhu, Kai Li, Yue Ma, Longxiang Tang, Chengyu Fang, Chubin Chen, Qifeng Chen, Xiu Li

Introduction

We explore the task of Customized Concept Swapping (CCS), a subtask of text-to-image (T2I) generation, which aims to replace a concept in a source image with a highly customized new concept. Combined with diffusion models , recent CCS methods demonstrate widespread applicability in areas such as selfie enhancement, photo blog creation, and comic creation.

Early work in CCS primarily relies on copy-paste techniques, which are rough and unreliable. By integrating powerful customization techniques with image editing methods , a series of works have been proposed. Although achieving remarkable success, these approaches still face the problems of inconsistency and inefficiency as shown in Fig. 2. (1) Inconsistency: Attention-based methods such as PhotoSwap and P2P maintain background consistency well but struggle with shape differences between source and target concepts, resulting in foreground inconsistency. Score distillation based methods such as SDS , DDS , and CDS fail to generate foreground concepts precisely and alter the background significantly, causing both foreground and background inconsistency. (2) Inefficiency: Attention-based methods require an inefficient training phase on source image to maintain background consistency. While score distillation based methods are training-free, they still require redundant calculations of forward passes at each timestep, leading to inference inefficiency.

To address the aforementioned issues, we propose InstantSwap, a training-free framework that efficiently performs customized concept swapping across shape differences while maintaining both foreground and background consistency. Specifically, we extract the bounding box (bbox) that indicates the position of the source concept from the enhanced cross-attention map of the source image. With this bbox, we perform the background gradient masking (BGM) strategy to prevent modifications outside the bbox, thus ensuring background consistency. Moreover, to improve foreground consistency, we leverage the semantic information to highlight the cross-attention maps of source and target concepts respectively within the bbox. This strategy leads to the semantic-enhanced concept representation (SECR), which facilitates precise foreground swapping. Finally, we introduce the Step-skipping Gradient Updating (SSGU) strategy, which only performs forward passes at certain timesteps to calculate gradients. For the timesteps without direct gradient computations, we reuse the previously obtained gradients for updates. Through this strategy, we reduce the total number of forward passes and improve the efficiency of our method.

Since the CCS is a recently proposed task, no dedicated evaluation benchmark currently exists. To address this gap, we introduce ConSwapBench, the first benchmark dataset specifically designed for CCS. ConSwapBench comprises two sub-benchmarks: ConceptBench and SwapBench. ConceptBench contains images representing target concepts, while SwapBench includes images with one or more concepts to be swapped, serving as source images.

Through extensive qualitative and quantitative comparisons, we demonstrate the effectiveness and superiority of our InstantSwap. We also conduct comprehensive ablation studies to verify the effectiveness of each component of our approach. Additionally, we further extend our InstantSwap to related tasks, proving its efficacy and versatility. Our contributions are summarized as follows:

We propose InstantSwap, a novel training-free customized concept swapping (CCS) framework, which enables efficient concept swapping across sharp shape differences.

We design the background gradient masking (BGM) strategy and semantic-enhanced concept representation (SECR) to improve the background and foreground consistency respectively. Moreover, we adopt a step-skipping gradient updating (SSGU) strategy to reduce redundant computation and improve efficiency.

To provide a comprehensive evaluation for CCS, we introduce ConSwapBench, the first benchmark for customized concept swapping. Extensive qualitative and quantitative evaluations demonstrate the effectiveness and superiority of our InstantSwap.

Related work

Image Editing is a fundamental and popular topic in computer vision. Previous works based on Generative Adversarial Networks (GAN) only focus on specific object domains, which limits the application. With the emergence of diffusion model , image editing is now able to modify various objects through prompts. These methods are mainly divided into five categories: instruction-based methods, blending-based, attention-based, inversion-based, and score distillation based methods. Instruction-based methods typically require an instruction editing dataset to train the diffusion model. Blending-based methods merge the source and target prompts to guide the editing process, while attention-based methods inject the attention feature of the source image. Both methods have lower editing costs but poorer background preservation and prompt alignment. Inversion-based methods aim to reverse the fixed trajectory generated by the forward pass to reproduce the source image. These methods can serve as an extra training phase to enhance the background consistency of attention-based methods. Finally, score distillation based methods draw on the optimization process of SDS , using score distillation-based loss to optimize the source image for editing. These methods are more flexible than the previous ones but still face challenges with background preservation.

2 Concept Swapping

Concept swapping, a subtask of general image editing, focuses on replacing the source concept in an image with a user-specified target concept. This task is first proposed by PbE , which employs a CLIP encoder to extract features of the target concept and inject them into the UNet through a cross-attention layer. After that, concurrent works extend concept swapping into the customization field . They combine attention-based editing methods , with tuning-based customization methods to achieve customized concept swapping. Building on Photoswap, SwapAnything further obtains masks with external modules to specify the locations of objects in the source image. We improved the existing method in three key aspects. Firstly, we employ bounding boxes instead of masks as spatial indicators of the source concept, allowing greater flexibility for shape variation during concept swapping. Secondly, we use bounding boxes to prevent background changes via gradient masking, thus ensuring background consistency. Third, we augment concept representation with semantic information to maintain foreground consistency. Finally, rather than executing forward passes at every timestep, we execute them only at specific intervals to enhance efficiency.

Method

Given a set of images, Xt={xi}i=1M\mathcal{X}_{t}=\{x_{i}\}_{i=1}^{M} representing a specific concept OtO_{t}, along with an image xsx_{s} and a prompt psp_{s} describing a source concept OsO_{s}, the objective of CCS is to “seamlessly” replace OsO_{s} in xsx_{s} with OtO_{t} according to a target prompt PtP_{t}, resulting in a final target image xtx_{t}. An ideal customized concept swapping should handle the shape differences between source and target concepts to preserve swapping consistency while maintaining satisfactory efficiency. We introduce InstantSwap to achieve this. InstantSwap is based on Stable Diffusion and extends from the score distillation based image editing methods .

In this paper, the foundational model utilized for text-to-image generation is Stable Diffusion . It takes a text prompt PP as input and generates the corresponding image xx. Stable Diffusion consists of three main components: an autoencoder(E(⋅),D(⋅))(\mathcal{E}(\cdot),\mathcal{D}(\cdot)), a CLIP text encoder τ(⋅)\tau(\cdot) and a U-Net ϵϕ(⋅)\epsilon_{\phi}(\cdot). Typically, it is trained with the guidance of the following reconstruction loss:

where ϵ∼N(0,1)\epsilon\sim\mathcal{N}\left(0,1\right) is a randomly sampled noise, t denotes the time step. The calculation of ztz_{t} is given by zt=αtz+σtϵz_{t}=\alpha_{t}z+\sigma_{t}\epsilon, where the coefficients αt\alpha_{t} and σt\sigma_{t} are provided by the noise scheduler.

1.2 Score Distillation Based Image Editing

Different from traditional attention-based image editing, score distillation based methods achieve image editing through iterative optimization with a score distillation loss. Given the latent feature zz of source image and a denoising U-Net ϵϕ(⋅)\epsilon_{\phi}(\cdot), SDS can optimize the latent feature zz of the image to align with the target prompt PtP_{t} by employing the following loss:

where ϵ\epsilon and tt are randomly sampled noise and timestep.

The resulting image SDS is very blurry and only contains foreground objects in the target prompt PtP_{t}. To address this issue, DDS expresses the gradient of Eq. 2 as

where δtgt\delta_{{tgt}} indicates the direction aligned with the target prompt and δbias\delta_{{bias}} refers to undesired part that makes the image blurry. Based on this, DDS further utilizes the fixed latent zt^\hat{z_{t}} of the source image and the source prompt PsP_{s} to approximate the bias component in Eq. 3:

Finally, DDS is represented by the difference of Eq. 3 and Eq. 4:

Based on Sec. 3.1.2, the loss of DDS is given by

2 InstantSwap

Directly extending score distillation based editing methods to the task of CCS encounters the challenge of inconsistency. These methods optimize the background and foreground simultaneously, causing cross-interference and leading to undesirable inconsistency. To address these limitations, we first propose a strategy to automatically locate objects to be edited, resulting in the object bounding box (bbox). With this bbox, we propose a background gradient masking technique to remove gradients in the background region and confine swapping to the foreground region. To further enhance foreground swapping consistency, we propose to learn Semantic-enhanced concept representations for both source and target concepts based on an attention map feature injection mechanism. An overview of our method is presented in Fig. 3.

We first automatically obtain the bbox to indicate the position of the concept OsO_{s} in the source image. Given the source image xsx_{s} and the source prompt PsP_{s}, we perform a forward pass with the U-Net ϵϕ(⋅)\epsilon_{\phi}(\cdot) and obtain the cross-attention map AcA^{c} and self-attention map AsA^{s} through:

where QQ is the query vector projected from the image features, d′d^{\prime} represents the output dimension of key and query features. KK is the key vector and VV is the value vector. For cross-attention maps AcA^{c}, KK and VV are projected from the text embeddings τ(P)\tau(P). For self-attention maps AsA^{s}, KK, and VV are projected from the image features. Directly applying a threshold on the AcA^{c} can yield a coarse-grained mask, which cannot accurately reflect the location of OsO_{s}. Inspired by , we modify the AcA^{c} as follows:

Based on Eq. 7, all values in AcA^{c} range between 0 and 1. Therefore, element-wise exponentiation of AcA^{c} by α\alpha can weaken the activation of non-target regions. Additionally, as mentioned in , AsA^{s} contains rich structural information. This information can effectively assist A^c\hat{A}^{c} in better activating the target regions. Finally, we apply the threshold β\beta to A^c\hat{A}^{c} to obtain the mask. Subsequently, we converted the mask into the bbox BsB_{s} based on the minimum and maximum coordinates of all foreground points within the mask. This strategy allows us to obtain the bbox BsB_{s} without any additional modules. We intentionally set a relatively loose constraint on the mask to obtain a bbox that completely covers the source concept. We discuss the effectiveness of our automatically obtained bboxes in Sec. 4.5.

2.2 Background Gradient Masking

With the object bbox, we propose a background gradient masking (BGM) approach to ensure that the concept swap is confined to the foreground region. Given the latent feature z^\hat{z} of the source image and the latent feature zz of the target image, where zz is initialized to z^\hat{z} and is continuously optimized to obtain the final target image xtx_{t}. Based on Eq. 6, we first obtain the gradient of zz:

As stated in , the mid term is a U-Net Jacobian term and can be omitted, and αt=∂zt /∂z\alpha_{t}=\partial z_{t}\ /\partial z is a constant which can be represented as w(t)w(t):

This gradient shares the same dimension as zz, which means it can update zz in a pixel-wise manner. However, this will update the foreground and background simultaneously, producing inconsistent background. To remedy this, we apply the bbox BsB_{s} on Eq. 10 to mask the gradients related to the background before back propagation and obtain our BGM:

This simple masking strategy prevents the background from being updated and thus ensures background consistency.

2.3 Semantic-enhanced Concept Representation

The BGM module maintains the background consistency during swapping. However, whether the source concept can be replaced with the target concept cannot be guaranteed. This limitation arises because the optimization of Eq. 11 is still carried out at the entire feature map level of both the source latent zt^\hat{z_{t}} and target latent ztz_{t} without distinguishing between the foreground and the background. To address this, we propose to obtain Semantic-enhanced concept representations for both source and target concepts and emphasize their locations within the foreground region during concept swapping.

Let FsF_{s} be the source image feature and psp_{s} represent the prompt of source concept (e.g. “rose”), the semantic embedding csc_{s} can be acquired through cs=τ(ps)c_{s}=\tau(p_{s}). We first resize the previously obtained object bbox to fit the dimensions of the source image feature FsF_{s}, resulting in the feature bbox BfB_{f}. We then crop FsF_{s} with BfB_{f} to get a regional image feature fs{f}_{s}. With fs{f}_{s}, we calculate the query vector through Qs=Wq⋅fsQ_{s}=W^{q}\cdot{f}_{s}. After that, we can obtain the key and value vectors through:

Then the final partial attention output is calculated as follows:

where d′d^{\prime} represents the output dimension of key and query features. In this way, we inject the semantic information of the source concept into the cross-attention map, resulting in regional concept representation fs^\hat{f_{s}}. We then map fs^\hat{f_{s}} back to the original feature map FsF_{s} to get a Semantic-enhanced representation F^s\hat{F}_{s} for the entire source image. Fig. 4 illustrates the process. In the target branch, we first convert the target concept into semantic space with DreamBooth, using a specific rare token (e.g., “sks”) to represent the concept. With the target prompt ptp_{t} (e.g., “sks teapot”) and the feature bbox BfB_{f}, we similarly apply this process for the target image feature FtF_{t} and obtain the semantic-enhanced representation F^t\hat{F}_{t} for the target image.

Through proactive injection of semantic guidance, we provide the source and target branches with Semantic-enhanced concept representation within the foreground region. Consequently, SECR transforms the target branch into a target concept adder and the source branch into a source concept remover. Their collaboration results in precise and seamless concept swapping, thus enhancing the foreground consistency. Moreover, SECR can also facilitate concept insertion and removal, which is further discussed in Sec. 4.6.

2.4 Step-skipping Gradient Updating

After addressing the problem of inconsistency, we turn our attention to the challenge of inefficiency. As illustrated in Fig. 5, previous methods calculate the gradient at each timestep. However, the success of DDIM in accelerating DDPM motivates us to consider: Can we skip the calculation of gradients at certain timesteps? During concept swapping, we observe that the effect of gradient updates on the target image is similar across adjacent timesteps (see detailed results in Supp.). Based on this observation, we propose the step-skipping Gradient Updating (SSGU) strategy. The key insight of SSGU is that skipping some gradient calculations does not significantly sacrifice the swapping consistency while considerably improving efficiency. As a result, our SSGU calculates gradients at interval timesteps and reuses the previously calculated gradients during the intervening timesteps.

where η\eta is the learning rate. Our SSGU periodically retains some anchor gradients and skips the forward passes between two anchor gradients. The step-skipping period is controlled by the SSGU factor λ\lambda. The set of anchor gradients can be defined as:

For the next timestep tt, another anchor gradient gtg_{t} is used to update xtx_{t}. As a result, our SSGU reduces the number of forward passes during the entire concept swapping process to 1/λ1/\lambda of the original count. Since the forward pass accounts for approximately 95% of the total inference time (see detailed analysis in Supp.), our SSGU can improve the overall inference speed of our method by approximately λ\lambda times, with minimal effect on swapping consistency (see Tab. 4). Furthermore, our SSGU can be transferred to other score distillation based methods to improve their efficiency in the same way, which is further discussed in Tab. 5.

Experiments

We conduct the experiments with Stable Diffusion v2.1-base on a single RTX3090. We use the customized checkpoint from DreamBooth to introduce concepts. We set the SSGU factor λ\lambda to 5, α\alpha to 2, β\beta to 0.5 and the guidance scale to 7.5. The bbox is obtained through the first three steps. Subsequently, we use SGD with a learning rate of 0.1 to optimize for 550 steps of iterations.

2 ConSwapBench

Despite the significant application potential of customized concept swapping, there is currently no dedicated evaluation benchmark. To meet the needs of comprehensive evaluation, we introduce ConSwapBench, the first benchmark dataset specifically designed for customized concept swapping. ConSwapBench consists of two sub-benchmarks: ConceptBench and SwapBench. ConceptBench comprises 62 images covering 10 different target concepts used for customization, while SwapBench includes 160 real images containing one or more objects to be swapped, serving as source images. For each image in SwapBench, we use Grounding SAM to acquire the bbox of the foreground concepts as the ground truth for evaluation purposes. We apply each customized concept from ConceptBench to perform concept swaps on each image in SwapBench, ultimately generating a total of 1,600 images for evaluation. More details can be found in Supp.

3 Qualitative Comparison

Since customized concept swapping is a relatively novel task, there are limited methods available for direct comparison. Consequently, we include SOTA image editing methods and adapt them for customized concept swapping. We include the following methods: (1) Score distillation based methods: SDS , DDS , and CDS ; (2) Attention-based methods: PhotoSwap , PnPInv , and P2P . We excluded SwapAnything as it is not publicly available. The qualitative results are illustrated in Fig. 6. We find that score distillation based methods can accommodate shape variations during concept swapping. However, they exhibit poor foreground fidelity (3rd and 6th rows) and lead to unnecessary modifications on the background (1st and 2nd rows). Attention-based methods are unable to manage shape variations (4th row) and also struggle with maintaining background consistency (5th row). In contrast, our method demonstrates superior performance in addressing shape variations and maintaining swapping consistency.

4 Quantitative Comparison

We also conduct a thorough quantitative comparison on ConSwapBench. For each generated image, we first use the ground truth bbox in SwapBench to obtain their foreground and background respectively. We use seven different metrics to evaluate the methods from three aspects: (1) Foreground consistency: We calculate the CLIP Image Score between the foreground of generated images and the images of customized concepts. (2) Background consistency: We use the four metrics, PSNR, LPIPS , MSE, SSIM to evaluate the background consistency. (3) Overall consistency and efficiency: We calculate the CLIP Text Score between generated images and target prompts to evaluate the overall prompt consistency. We also report the inference time of each method to evaluate their efficiency. As shown in Tab. 1, our method outperforms other methods on all seven metrics.

5 Ablation Study

BGM. To verify the effectiveness of BGM in background preservation, we conduct an ablation study by removing BGM. As illustrated in the second column of Fig. 7, while our method can still achieve concept swapping without BGM, it causes serious modifications on the background. In contrast, our full method not only maintains high foreground fidelity but also effectively preserves the background consistency. We further conduct a quantitative analysis of the background consistency and prompt consistency, as shown in Tab. 2. Our full method outperforms in all metrics.

Automatic bounding box detection mechanism. We further verify the effectiveness of our boxes by using ground truth (GT) bboxes from SwapBench to replace the automatically obtained bboxes As shown in Fig. 7, although GT bboxes accurately indicate the location of the source concept, they prevent our method from fully swapping the source concept. Compared to GT bboxes, our bboxes are relatively larger and can fully cover the source concept, thereby facilitating complete concept swapping. We also provide quantitative comparisons of different bboxes. As shown in Tab. 3, Gen stands for the generation bboxes, while Eva stands for the evaluation bboxes. The two types of bboxes do not significantly affect background preservation in our method, whereas our bboxes perform better than GT bboxes on the foreground metric. More detailed comparisons can be found in Supp..

SECR. To verify the effectiveness of SECR, we conduct ablation studies including removing SECR from (1) source branch (w/o source), (2) target branch (w/o target), (3) both (w/o source & target). The visualization results are illustrated in columns 3 to 5 of Fig. 7. Although all methods preserve the background well, they show reduced foreground fidelity. Additionally, we perform a quantitative analysis of their foreground consistency and prompt consistency. The results presented in Tab. 4 indicate that our full method exhibits superior performance.

SSGU. To verify the effectiveness of our proposed SSGU, we first visualize the images generated under different SSGU factors. As shown in Fig. 8, λ=1\lambda=1 indicates that SSGU is not used. When λ≤9\lambda\leq 9, the SSGU can preserve foreground and background consistency well while improving the efficiency of our method. As λ\lambda increases, the images exhibit more artifacts due to excessive neglect of gradients. Therefore, identifying an optimal SSGU factor λ\lambda is crucial. We further conduct a detailed quantitative analysis of different λ\lambda values on foreground consistency and efficiency (see complete results in Supp.), as illustrated in Fig. 11, where the xx-axis represents different λ\lambda values and the yy-axis represents the respective metric outcomes. When SSGU is not used, our method achieves the best swapping consistency but the lowest efficiency. As λ\lambda increases, our SSGU sacrifices certain swapping consistency but significantly improves efficiency. To balance consistency and efficiency, we ultimately select λ=5\lambda=5 for our final model.

6 Applications of InstantSwap

Multi-concept swapping. Our InstantSwap can be easily extended to facilitate multi-concept swaps by sequentially performing multiple single-concept swaps. As shown in the left of Fig. 10, our method can swap each concept within the image with both foreground and background consistency.

Human face swapping. InstantSwap demonstrates exceptional capabilities in human face swapping. As shown in the middle of Fig. 10, with customized face models from CivitAI , users can seamlessly replace the face in a source image with a customized target face.

Concept insertion and removal. In addition to concept swapping, our method also supports concept insertion and removal. For concept insertion, we employ the same procedure as concept swapping. For concept removal, we adjust the target prompt and the target semantic input ptp_{t} of SECR to a null prompt. The results in the right of Fig. 10 further demonstrate the versatility of our method.

Accelerating other methods. SSGU can be transferred to other score distillation based methods to enhance their efficiency. We select three representative methods: SDS , DDS for image editing, and CoSD for video editing. As shown in Fig. 9, combining these methods with SSGU can significantly improve their efficiency while almost not altering the generation quality. We further conduct a quantitative analysis to assess the transferability of the proposed SSGU, as presented in Tab. 5.

Conclusion

This paper introduces InstantSwap, a novel framework for precise and efficient customized concept swap. Our BGM and SECR collaborate to maintain both background and foreground consistency. Furthermore, we propose the SSGU to eliminate redundant computation and improve efficiency. Finally, we introduce ConSwapBench, a comprehensive benchmark dataset for customized concept swapping. The impressive performance of InstantSwap demonstrates its effectiveness. We hope our InstantSwap can inspire future research, particularly in efficiently managing concept swapping with obvious shape variance. Future work could focus on (1) extending image-based customized concept swapping to the video domain; (2) enhancing the images of target concepts with low-level methods ; and (3) achieving more lightweight and precise concept swapping.

References