AnyDoor: Zero-shot Object-level Image Customization

Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, Hengshuang Zhao

Introduction

Image generation is flourishing with the booming advancement of diffusion models . Humans could generate favored images by giving text prompts, scribbles, skeleton maps, or other conditions. The power of these models also brings the potential for image editing. For example, some works learn to edit the posture, styles, or content of an image via instructions. Other works explore re-generating a local image region with the guidance of text prompts.

In this paper, we investigate “object teleportation”, which means accurately and seamlessly placing the target object into the desired location of the scene image. Specifically, we re-generate a box-marked local region of a scene image by taking the target object as the template. This ability is in significant requirement in practical applications, like image composition, effect-image rendering, poster-making, virtual try-on, etc.

Although strongly in need, this topic is not well explored by previous researchers. Paint-by-Example and Objectstitch take a target image as the template to edit a specific region of the scene image, but they could not generate ID (identity)-consistent contents, especially for untrained categories. Customized synthesis methods are able to conduct generations for the new concepts but could not be specified for a location of a given scene. Besides, most customization methods need finetuning on multiple target images for nearly an hour, which largely limits their practicability for real applications.

We address this challenge by proposing AnyDoor. Different from previous methods, AnyDoor is able to generate ID-consistent compositions with high quality in zero-shot. To achieve this, we represent the target object with identity- and detail-related features, then composite them with the interaction of the background scene. Specifically, we use an ID extractor to produce discriminative ID tokens and delicately design a frequency-aware detail extractor to get detail maps as a supplement. We inject the ID tokens and the detail maps into a pre-trained text-to-image diffusion model as guidance to generate the desired composition. To learn customized object generation with high diversities, we collect image pairs for the same object from videos to learn the appearance variations, and also leverage large-scale statistic images to guarantee the scenario diversity. To further take advantage of both videos and images, we design an adaptive timestep sampler to make different denoising steps to benefit from different sourced training data.

Equipped with these techniques, AnyDoor demonstrates extraordinary abilities for zero-shot customization. As in Fig. 1, AnyDoor shows promising performance for the synthesis of the new concept and could serve as a powerful solution for virtual try-on (top row). Besides, since AnyDoor owns the high controllability for editing the specific local regions of the scene image, it is easy to be extended to multi-subject composition (middle row), which is a hot and challenging topic explored by many customized generation methods . Moreover, the high generation fidelity and quality of AnyDoor unlock the possibilities for more fantastic applications like object moving and swapping (bottom row). We hope that AnyDoor could serve as a foundation solution for various image generation and editing tasks with image input, and act as the basic ability to energize more fancy applications.

Related Work

Local image editing. Most of the previous works focus on editing local image regions with text guidance. Blended Diffusion conducts multi-step blending in the masked region to generate more harmonized outputs. InpaintAnything involves SAM and Stabble Diffusion to replace any object in the source image with text described target. Paint-by-Example uses CLIP image encoder to convert the target image as an embedding for guidance, thus painting a semantic consistency object on the scene image. ObjectStitch proposes a similar solution with , it trains a content adaptor to align the outputs of the CLIP image encoder to the text encoder to guide the diffusion progress. However, those methods could only give coarse guidance for generations and often fails to synthesize ID-consistent results for untrained new concepts.

Customized image generation. Customized or termed subject-driven generation aims to generate images for specific objects given several target images and relevant text prompts. Some works finetune a “vocabulary” to describe the target concepts. Cones finds the corresponding neurons for the referred object. Although they could generate high-fidelity images, the user could not specify the scenario and the location of the target object. Besides, the time-consuming finetuning impedes them to be used in large-scale applications. Recently, BLIP-Diffusion leverages BLIP-2 to align image and text, thus supporting using zero-shot subject-driven generation. Some methods explore large-scale upstream training for the finetune-free subject-driven generation. Fastcomposer binds the image representation with certain text embeddings to do multiple-person generation. However, these zero-shot explorations are still in the initial stages with unsatisfactory performances or limited application scenarios.

Image harmonization. A classical image composition pipeline is cutting the foreground object and pasting it on the given background. Image harmonization could furtherly adjust the pasted region for more reasonable lighting and color. DCCF designs pyramid filters to better harmonize the foreground. CDTNet leverages dual transformers. HDNet proposes a hierarchical structure to consider both global and local consistency and reaches the state-of-the-art. However, these methods only explore the low-level changes, editing the structure, view, and pose of the foreground objects, or generating the shadows and reflections are not taken into consideration.

Method

The pipeline of AnyDoor is demonstrated in Fig. 2. Given the target object, the scene, and the location, AnyDoor generates the object-scene composition with high fidelity and diversity. The core idea is representing the object with identity- and detail-related features, and recomposing them in the given scene by injecting those features into a pre-trained diffusion model. To learn the appearance changes, we leverage large-scale data including both videos and images for training.

We leverage the pre-trained visual encoders to extract the identity information of the target object. Previous works choose CLIP image encoder to embed the target object. However, as CLIP is trained with text-image pairs with coarse descriptions, it could only embed semantic-level information but struggles to give discriminative representations that preserve the object identity. To overcome this challenge, we made the following updates.

Background removal. Before feeding the target image into the ID extractor, we remove the background with a segmentor and align the object to the image center. The segmentor model could be either automatic or interactive . This operation is proven helpful to extract more neat and discriminative features.

Self-supervised representation. In this work, we find the self-supervised models show a strong ability to preserve more discriminative features. Pretrained on large-scale datasets, self-supervised models are naturally equipped with the instance-retrieval ability and could project the object into an augmentation-invariant feature space. We choose the currently strongest self-supervised model DINO-V2 as the backbone of our ID extractor, which encodes image as a global token Tg1×1536\mathbf{T}_{\text{g}}^{1\times 1536}, and patch tokens Tp256×1536\mathbf{T}_{\text{p}}^{256\times 1536}. We concatenate the two types of tokens to preserve more information. We find that using a single linear layer as a projector could align these tokens to the embedding space of the pre-trained text-to-image UNet. The projected tokens TID257×1024\mathbf{T}_{\text{ID}}^{257\times 1024} are noted as our ID tokens.

2 Detail Feature Extraction

We consider that, as the ID tokens lose the spatial resolution, it would be hard for them to adequately maintain the fine details of the target object. Thus we need extra guidance for the detail generation in complementary.

Collage representation. Inspired by , using collage as controls could provide strong priors, we attempt to stitch the “background removed object” to the given location of the scene image. With this collage, we observe a significant improvement in the generation fidelity, but the generated results are too similar to the given target which lacks diversity. Facing this problem, we explore setting an information bottleneck to prevent the collage from giving too many appearance constraints. Specifically, we design a high-frequency map to represent the object, which could maintain the fine details yet allow versatile local variants like the gesture, lighting, orientation, etc.

High-frequancy map. We extract the high-frequency map of the target object with

where Kh,Kv\mathbf{K}_{h},\mathbf{K}_{v} denote horizontal and vertical Sobel kernels, acting as high-pass filters. ⊗,⊙\otimes,\odot refer to convolution and Hadamard product. Given an Image I\mathbf{I}, we first extract the high-frequency regions using these high-pass filters, then extract the RGB colors using the Hadamard product. We also add an eroded mask Merode\mathbf{M}_{\text{erode}} to filter out the information near the outer contour of the target object. After getting the high-frequency map, we stitch it onto the scene image according to the given locations and then pass the collage to the detail extractor. The detail extractor is a ControlNet-style UNet encoder, which produces a series of detail maps with hierarchical resolutions.

Focus region visualization. As visualized in Fig. 3, the tokens produced by DINO-V2 focus more on the overall structure, leaving it hard to encode the fine details like the logos of the backpack in the first row. In contrast, the high-frequency map could help take care of these details as a complementary.

3 Feature Injection

After getting the ID tokens and detail maps, we inject them into a pre-trained text-to-image diffusion model to guide the generation. We pick Stable Diffusion , which projects the images into latent space and conducts the probabilistic sampling using a UNet. We note the pre-trained UNet as x^θ\hat{\mathbf{x}}_{\theta}, it starts denoising from an initial latent noise ϵ∼U()\mathbf{\epsilon}\sim\mathcal{U}() and takes the text embedding c\mathbf{c} as the condition to generate new image latent zt=αtx^θ(ϵ,c)+σtϵ\mathbf{z}_{t}=\alpha_{t}\hat{\mathbf{x}}_{\theta}(\mathbf{\epsilon},\mathbf{c})+\sigma_{t}\mathbf{\epsilon}. The training supervision is a mean square error loss as

x\mathbf{x} is the ground-truth image latent, tt is the diffusion timestep, αt,σt\alpha_{t},\sigma_{t} are denoising hyperparameters.

In this work, we replace the text embedding c\mathbf{c} as our ID tokens, which are injected into each UNet layer via cross-attention. For the detail maps, we concatenate them with UNet decoder features at each resolution. During training, we freeze the pre-trained parameters of the UNet encoder to preserve the priors and tune the UNet decoder to adapt it to our new task.

4 Training Strategies

Image pair collection. The ideal training samples are image pairs for “the same object in different scenes”, which are not directly provided by existing datasets. As alternatives, previous works leverage single images and apply augmentations like rotation, flip, and elastic transforms. However, these naive augmentations could not well represent the realistic variants of the poses and views.

To deal with this problem, in this work, we utilize video datasets to capture different frames containing the same object. The data preparation pipeline is demonstrated in Fig. 4, where we leverage video segmentation/tracking data as examples. For a video, we pick two frames and extract the masks for the foreground object. Then, we mask the background for one image and crop it around the mask as the target object. For the other frame, we generate the box and mask the box region to get the scene image, the unmasked image could serve as the training ground truth. The full data used is listed in Tab. 1, which covers a large variety of domains like nature scenes, virtual try-on, saliency, and multi-view objects.

Adaptive timestep sampling. Although the video data would be beneficial for learning the appearance variation, the frame qualities are usually unsatisfactory due to the low resolution or motion blur. In contrast, images could provide high-quality details and versatile scenarios but lack appearance changes.

To take advantage of both video data and image data, we develop adaptive timestep sampling to make different modalities of data to benefit different stages of denoising training. The original diffusion model evenly samples the timestep (T) for each training data. However, it is observed that the initial denoising steps mainly focus on generating the overall structure, the pose, and the view; and the later steps cover the fine details like the texture and colors . Thus, for the video data, we increase the possibility of sampling early denoising steps (large T) during training to better learn the appearance changes. For images, we increase the probabilities of the late steps (small T) to learn how to cover the fine details.

Experiments

Hyperparameters. We choose Stable Diffusion V2.1 as the base generator. During training, we process the image resolution to 512×512512\times 512. We choose Adam optimizer with an initial learning rate of 1e−51e^{-5}.

Zoom-in strategy. During inference, given a scene image and a location box, we expand the box into a square with an amplifier ratio of 2.0. Then, we crop the square and resize it to 512×512512\times 512 as the input for our diffusion model. Thus, we could deal with scene images with arbitrary aspect ratios and boxes for extremely small or large areas.

Benchmarks. For quantitative results, we construct a new benchmark with 30 new concepts provided by DreamBooth for the target images. For the scene image, we manually pick 80 images with boxes in COCO-Val . Thus we generate 2,400 images for the object-scene combinations. We also make qualitative analysis on VitonHD-test to valid of performance for virtual try-on.

Evaluation metrics. On our constructed DreamBooth dataset, we follow DreamBooth to calculate the CLIP-Score and DINO-Score, as these metrics could reflect the similarity between the generated region and the target object. In addition, we organize user studies with a group of 15 annotators to rate the generated results from the perspective of fidelity, quality, and diversity.

2 Comparisons with Existing Alternatives

Reference-based methods. In Fig. 5, we present the visualization results compared with previous reference-based methods. Paint-by-Example and Graphit support the same input format as ours, they take a target image as input to edit a local region of a scene image without parameter tuning. We also compare Stable Diffusion , which is a text-to-image model, we use its inpainting version and give detailed text descriptions as the condition to conduct the generation for the text-described target.

Results show that previous reference-based methods could only keep the semantic consistency with distinguishing features like the dog face on the backpack, and coarse granites of patterns like the color of the sloth toy. However, as those new concepts are not included in the training category, their generation results are far from ID-consistent. In contrast, our AnyDoor shows promising performance for zero-shot image customization with highly-faithful details.

Tuning-based methods. Customized generation is extensively explored. Previous works usually fine-tune a subject-specific text inversion to present the target object, thus making generations with arbitrary text prompts. They could better preserve the fidelity compared with previous reference-based methods, but have the following drawbacks: first, the fine-tuning usually requires 4-5 target images and takes near an hour; second, they could not specify the background scene and target locations; third, when it comes to multi-subject composition, the attributes of different subjects often mix together.

In Fig. 6, we include tuning-based methods for comparisons and also use Paint-by-Example as the representative for previous reference-based methods. Results show that Paint-by-Example performs well for trained categories like dog and cat (in row 3) but performs poorly for new concepts (row 1-2). DreamBooth , Custom Diffusion , and Cones give better fidelity for new concepts but still suffer from the problem of “multi-subject confusion”. In contrast, AnyDoor owns the advantages of both reference- and tuning-based methods, which could generate high-fidelity results for mult-subject composition without the need for parameter tuning.

User study. We organize a user study to compare Paint-by-Example , Graphit , and our model. We let 15 annotators rate 30 groups of images. For each group, we provide one target image and one scene image; and make each of the three models generates four predictions. We prepare detailed regulations and templates to rate the images for scores of 1 to 4 from three perspectives: “Fidelity”, “quality”, and “diversity”. “Fidelity” measures the ability of ID preserving. “Quality” counts for whether the generated image is harmonized without considering fidelity. As we do not encourage “copy-paste” style generation, we use “diversity” to measure the differences among the four generated proposals. The user-study results are listed in Tab. 2. It shows that our model owns obvious superiorities for fidelity and quantity, especially for fidelity. However, as only keeps the semantic consistency but our methods preserve the instance identity, they naturally have larger space for the diversity. In this case, AnyDoor still gets higher rates than and competitive results with , which verifies the effectiveness of our method.

3 Ablation Studies

We carry out extensive ablation studies to verify the effectiveness of our designs. We first validate the core components, then we dive into the details of the ID extractor and detail extractor to give an in-depth analysis.

Core components. As demonstrated in Fig. 7, given the same target object, scene, and location, we analyze the generated results with different model designs. We demonstrate the generation results of AnyDoor in the last column and remove each core component individually to observe the influences. We first change the backbone of our ID extractor from the DINO-V2 to CLIP image encoder , which is widely used in previous counterparts like . We find the generated results lose the identity features, and could only keep the semantic consistency. Then, we set the collage region from the high-frequency map to an all-zero map like the inpainting baselines . We find that the fine details degenerate compared with our full model (last column), like the logo of the bag (row 1), and the eye shape of the toy sloth (row 2). It shows that our frequency map effectively guides the generation of fine structural details. We also make ablation for our adaptive timestep sampling (ATS) strategy. We replace ATS with an even distribution sampler and find the results present better diversity but are inferior for both image quality and fidelity.

The quantitative results are shown in Tab. 3, where we construct a baseline solution using the CLIP image encoder like Paint-by-Example . We add each component step-by-step on this baseline. Each component makes contributions to the CLIP score and DINO score.

ID extractor. We explore the key factors for designing the ID extractor. In Fig. 8, we compare CLIP , DINO-V2 , VGG to extract the ID tokens. We conclude that DINO-V2 shows a dominant superiority for keeping the target identity. We also verify that it is significant to filter out the background information for the target object, thus DINO-V2 could extract cleaner and more discriminative features. Quantitative results are listed in Tab. 4, which are consistent with our visual analysis.

Detail extractor. We make multiple explorations for the collaged image. The CLIP and DINO scores are reported in Tab. 5, compared with non-college, all these collaging methods bring notable improvements. To make better comparisons, we give visualization results in Fig. 9, which shows comparisons for no collage, pasting of the original target object, the noised inversion of the target object, the shuffled patches, and our high-frequency map. We observe a trade-off between fidelity and diversity. “Original image” presents the highest fidelity for both the robot and the dog, but the generated images seem like a copy-paste of the target. “None” shows the best diversity for the poses of the dog, but it lacks details like the badge of the dog and the whole shape of the robots. Among those methods, the high-frequency map shows a satisfactory trade-off, which keeps the majority of the details but adjusts the dog and robot with proper poses and views.

4 More Applications

Virtual try-on. As shown in Fig. 10, only trained with a small portion of task-specific data , AnyDoor could give satisfactory performance for virtual try-on. AnyDoor could preserve the color, texture, and patterns of the target clothes and performs well for large human gestures. It should be noticed that the traditional GAN-based try-on methods require more restricted inputs like human parsing maps. However, AnyDoor only needs a box to indicate the upper-body’s position, which is a much more relaxed condition.

Flexible interactions. When incorporating an inpainting model and an interactive segmentation model , we could realize more fantastic functions by clicking and dragging. As demonstrated in Fig. 11, in the first column, the user could click on the dog that appears near the image boundary and drag it to the image center. In the second column, users could swap the location of the two objects by clicking. In the third column, we could adjust the shape of the dog by dragging the corner points of the box. The pipeline uses an inpainting model to fill the object’s original position according to the scene background and apply the AnyDoor to re-generate it at the new location.

Conclusion

In this work, we present AnyDoor, a diffusion-based generator that could conduct object teleportation. The core contribution of our research is using a discriminative ID extractor and a frequency-aware detail extractor to characterize the target object. Trained on a large combination of video and image data, we composite the object at the specific location of the scene image. AnyDoor provides a universal solution for general region-to-region mapping tasks and could be profitable for various applications.

References