SINE: SINgle Image Editing with Text-to-Image Diffusion Models

Zhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris Metaxas, Jian Ren

Introduction

Automatic real image editing is an exciting direction, enabling content generation and creation with minimal effort. Although many works have been conducted in this area, achieving high-fidelity semantic manipulation on an image is still a challenging problem for the generative models, considering the target image might be out of the training data distribution . The recently introduced large-scale text-to-image models, e.g., DALL·E 2 , Imagen , Parti , and StableDiffusion , can perform high-quality and diverse image generation with natural language guidance. The success of these works has inspired many subsequent efforts to leverage the pre-trained large-scale models for real image editing . They show that, with properly designed prompts and a limited number of fine-tuning steps, the text-to-image models can manipulate a given subject with text guidance.

On the downside, the recent text-guided editing works that build upon the diffusion models suffer several limitations. First, the fine-tuning process might lead to the pre-trained large-scale model overfit on the real image, which degrades the synthesized images’ quality when editing. To tackle these issues, methods like using multiple images with the same content and applying regularization terms on the same object have been introduced . However, querying multiple images with identical content or object might not be an available choice; for instance, there is only one painting for Girl with a Pearl Earring. Directly editing the single image brings information leakage from the pre-trained large-scale models, generating images with different content (examples in Fig. 5); therefore, the application scenarios of these methods are greatly constrained. Second, these works lack a reasonable understanding of the object geometry for the edited image. Thus, generating images with different spatial size as the training data cause undesired artifacts, e.g., repeated objects, and incorrectly modified geometry (examples in Fig. 5). Such drawbacks restrict applying these methods for generating images with an arbitrary resolution, e.g., synthesizing high-resolution images from a single photo of a castle (as in Fig. 4), again limiting the usage of these methods.

In this work, we present SINE, a framework utilizing pre-trained text-to-image diffusion models for SINgle image Editing and content manipulation. We build our approach based upon existing text-guided image generation approaches and propose the following novel techniques to solve overfitting issues on content and geometry, and language drift :

First, by appropriately modifying the classifier-free guidance , we introduce model-based classifier-free guidance that utilizes the diffusion model to provide the score guidance for content and structure. Taking advantage of the step-by-step sampling process used in diffusion models, we use the model fine-tuned on a single image to plant a content “seed” at the early stage of the denoising process and allow the pre-trained large-scale text-to-image model to edit creatively conditioned with the language guidance at a later stage.

Second, to decouple the correlation between pixel position and content, we propose a patch-based fine-tuning strategy, enabling generation on arbitrary resolution.

With a text descriptor describing the content that is aimed to be manipulated and language guidance depicting the desired output, our approach can edit the single unique image to the targeted domain with details preserved in arbitrary resolution. The output image keeps the structure and background intact while having features well-aligned with the target language guidance. As shown in Fig. 1, trained on a painting Girl with a Pearl Earring with resolution as 512×512512\times 512, we can sample an image of a sculpture of the girl at the resolution of 640×512640\times 512 with the identity features preserved. Moreover, our method can successfully handle various edits such as style transfer, content addition, and object manipulation (more examples in Fig. 3). We hope our method can further boost creative content creation by opening the door to editing arbitrary images.

Related Work

Text-guided image synthesis has drawn considerable attention in the generative model context . The recent development of diffusion models introduced new solutions to this problem and produced impressive results . With the significant improvement of these models, rather than training a large-scale text-to-image model from scratch, a leading line of works focuses on taking advantage of the existing pre-trained model and manipulating images according to given natural language guidance . In these works, studies explore text-based interfaces for image editing , style transfer , and generator domain adaption .

The development of the diffusion model provides a giant and flexible design space for this task. Many works utilize pre-trained diffusion models as generative priors and are training-free. ILVR guides the denoising process by replacing the low-frequency part of the sample with that of the target reference image. SDEdit applies the diffusion process first on an image or a user-created semantic map and then conducts the denoising procedure conditioned with the desired output. Blended diffusion performs language-guided inpainting with a given mask.

Another line of research showed great potential and semantic editing ability of fine-tuning. DiffusionCLIP leverages the CLIP model to provide gradients for image manipulation and delivers impressive results on style transfer. Textual-Inversion and DreamBooth fine-tune the text embedding or the full diffusion model using a few personalized images (typically 3∼53\sim 5) to synthesize images of the same object in a novel context. These methods, however, either drastically change the layout of the original image when dealing with a single image or can not fully leverage the generalization ability of the pre-trained model for editing due to overfitting or language drift. Notably, Prompt-to-Prompt controls the editing of synthesized images by manipulating the cross-attention maps; however, its editing ability is limited when applied to real images.

This work introduces a solution to achieve image fidelity and text alignment simultaneously. Our method can perform high-quality semantic editing globally and locally on one single image. On the other hand, previous works lack an understanding of the object geometry of the edited image. When editing the image at an arbitrary resolution, the artifacts in the results will be obvious. Prior works have investigated generating images at arbitrary resolution using positional encoding as inductive bias so that the correlation between content and position can be eliminated. Anyres-GAN adopt a patch training mechanism to leverage high-resolution data to help the generation of images in the low-resolution domain. We propose a patch-based fine-tuning method to achieve arbitrary resolution editing.

Methods

For one arbitrary in-the-wild image, our goal is to edit the image via language while preserving the maximal amount of details from the original image. To do so, we leverage the generalization ability of pre-trained large-scale text-to-image models . An intuitive approach is to fine-tune the diffusion models with the single image and text description, similar to DreamBooth . Ideally, it should provide a model that can reconstruct the input image using the given text descriptor and synthesize new images when given other language guidance. Unfortunately, we find the model can easily overfit the single trained image and its corresponding text description. Thus, although the fine-tuned model can still reconstruct the input image perfectly, it can no longer synthesize diverse images according to the given language guidance (as shown in Fig. 5). Moreover, it struggles to generate arbitrary resolution images due to the lack of positional information (as in Fig. 4).

To solve the above issues, we propose a test-time model-based classifier-free guidance and a patch-based fine-tuning technique. An overview of our method is illustrated in Fig. 2. In the following sections, we review the backbone model used in our approach (Sec. 3.1). Then, we describe how to overcome the overfitting problem with model-based guidance (Sec. 3.2). Lastly, we present how to address the problem of limited resolution generation (Sec. 3.3).

where tt is the time step, zt\mathbf{z}_{t} is the latent noised to time tt, ϵ\boldsymbol{\epsilon} is the unscaled noise sample, ϵθ\boldsymbol{\epsilon}_{\theta} is the denoising model, yy is the conditioning input, and τθ\boldsymbol{\tau}_{\theta} maps yy to a conditioning vector. During training time, ϵθ\boldsymbol{\epsilon}_{\theta} and τθ\boldsymbol{\tau}_{\theta} are jointly optimized. A random noise tensor is sampled and denoised at inference time based on the conditioning input, e.g., text prompt, to produce a new latent. Inspired by DreamBooth , we construct the text prompt for fine-tuning a single image as “a photo/painting of a [∗\ast] [class noun]”, where “[∗\ast]” is a unique identifier and “[class noun]” is a coarse class descriptor (e.g., “castle”, “lake”, “car”, etc.).

2 Model-Based Classifier-Free Guidance

With the above-presented LDMs, we introduce our approach, inspired by classifier-free guidance, to overcome overfitting when fine-tuning LDMs with one image.

Classifier-free guidance is a technique widely adopted by prior text-to-image diffusion models . A single diffusion model is trained using conditional and unconditional objectives by randomly dropping the condition during training. When sampling, a linear combination of the conditional and unconditional score estimation is used:

Since we only have one image as the training data, e.g., painting of Mona Lisa, and one corresponding text descriptor of that image, the diffusion model suffers from overfitting, and severe language drifts after fine-tuning . As a result, the fine-tuned model fails to synthesize images containing features from other language guidance. The overfitting issue might be due to only one repeated prompt used during fine-tuning, making other text prompts no longer accurate enough to control editing (see examples in Fig. 6).

Model-based classifier-free guidance. Existing “personalized” text-guided real image editing works only use one fine-tuned model for image generation and editing , ignoring the capacity of pre-trained large-scale text-to-image models. Instead, to alleviate the overfitting of the fine-tuned model, we leverage the pre-trained text-to-image model for image generation with the provided language guidance and use the fine-tuned model to provide content features in a fashion of combining scores from the two models, similar to classifier-free guidance.

Specifically, let ϵ^θ\boldsymbol{\hat{\epsilon}}_{\theta} denote the fine-tuned denoising model, and ϵθ\boldsymbol{\epsilon}_{\theta} denote the pre-trained text-to-image model. During sampling, at specified steps, we use our fine-tuned model to guide the pre-trained one by using a linear combination of the scores from each model. Thus, the score estimation in Eqn. 2 becomes:

where vv stands for the model guidance weight, c^\mathbf{\hat{c}} is the language guidance token obtained from the fine-tuned diffusion model with the text prompt used during fine-tuning, and c\mathbf{c} is the target language conditioning obtained from the target prompt.

To prevent artifacts from the over-fitted model and maintain the fidelity of the generated image, we propose to sample using Eqn. 3 with t>Kt>K and sample using Eqn. 2 for t≤Kt\leq K. From KK to 0, the denoising process only depends on the pre-trained model. Following this approach, we can fully leverage the generalization ability of the pre-trained model (examples in Fig. 6). Also note that this method could be generalized to include multiple prompts or even multiple modalities.

3 Patch-Based Fine-Tuning

With model-based classifier-free guidance, we are able to edit and manipulate a single image with given language guidance. Here, we further show how to improve the fine-tuning process for a single training image so that the fine-tuned model can better understand the content and geometry of the image. Thus, it can provide better content guidance for the large-scale text-to-image model during the sampling time and unleash the potential for generating arbitrary-resolution images .

Limited-resolution generation. We first review the limitations of the current fine-tuning process. Given an input image I\mathcal{I} with resolution as H×WH\times W, we can obtain a downsampled latent code z\mathbf{z} from the pre-trained encoder. Since the text-to-image diffusion model is pre-trained at a fixed resolution, i.e., p×pp\times p, we need to resize the input image to a corresponding resolution sp×spsp\times sp, where ss represents the scaling factor of the encoder, to match the resolution for reducing the fine-tuning cost. In essence, prior knowledge of the correlation between the position and content information is learned by the diffusion model. Thus, when sampling from a higher-resolution noise tensor, the generated latent code leads to artifacts like duplicates or position shifting (visual examples in Fig. 5). To tackle such drawbacks, we propose a simple yet effective fine-tuning method.

After fine-tuning, the model can generate latent code at different resolutions by giving the positional information directly to the model. The arbitrary resolution image editing is conducted by feeding two inputs to the model: the positional embedding of the whole image; and a randomly sampled noisy latent with the dimension corresponding to the resolution we want. When sampling in an arbitrary resolution, the model can still keep the structure of the original image intact (examples in Fig. 4). It is worth noting that with or without the correct position encoding, our framework still naturally permits retargeting, i.e., maintaining the aspect ratio of salient objects, like SinGAN, InGAN, and Drop-the-GAN.

Experiments

Implementation Details While our method can be generally applied to different frameworks, we implement it based on the recently released text-to-image LDM, Stable Diffusion . The pre-trained model was pre-trained on 512×512512\times 512 images from LAION dataset . The spatial size of the latent code from the pre-trained model is 64×6464\times 64.

For patch-based fine-tuning, we randomly crop images to patches with height and width uniformly in the range of [0.1H,H]×[0.1W,W][0.1H,H]\times[0.1W,W] and resize them to 512×512512\times 512. Experiments are conducted using 1×RTX 80001\times\text{RTX 8000} GPU with a batch size of 11. The base learning rate is set to 1×10−61\times 10^{-6}. The number of time steps for the diffusion model, TT, is 10001000. Experiments without and with patch-based fine-tuning are created after 800800 and 10,00010,000 optimization steps, respectively. Unless otherwise noted, we adopt other hyperparameter choices from Stable Diffusion , and the results are generated with image resolution 512×512512\times 512 and with latent dimension 64×6464\times 64. For sampling parameters, we choose K=400K=400 and v=0.7v=0.7.

To better understand various approaches, we collect images from a wide range of domains, i.e., free-to-use high-resolution images from Flickrhttps://www.flickr.com/ and Unsplashhttps://unsplash.com/. During fine-tuning, we apply a coarse class descriptor to the content we want to preserve, e.g., dog, cat, castle, etc. After optimization, we edit each image with diverse editing prompts. We randomly generate 44 edit results for each image and editing prompt and choose the best one (such a process is also applied to other comparison methods). Our work shows impressive editing ability when applied to various images with different language guidance.

As presented in Fig. 3, using the model-based classifier-free guidance (Sec. 3.2) enables us to apply various editing via text prompts on the single real images. Each image has two text prompts describing different features we want to edit, e.g., image style, background content, the texture of the content, etc. Our method can edit the related features while keeping the content intact. We further show our editing results on arbitrary resolution generation in Fig. 4. For each source image, we edit it with different prompts at various resolutions. As can be seen, our patch-based fine-tuning schedule (Sec. 3.3) successfully preserves the original portion and geometry features of the single source image, even on highly challenging resolution such as 512×1024512\times 1024.

2 Comparisons

We compare our method to concurrent leading techniques, Textual-Inversion and DreamBooth , that can be used for single-image editing. Considering no official implementation has been released for DreamBooth, we adopt an unofficial but well-adopted and highly competitive implementation based on Stable Diffusion. We compare these techniques strictly according to the detailed guidance provided with the implementations.

Fig. 5 shows the comparison results. As can be noticed, our method maintains the fidelity of the images while applying changes as desired. Furthermore, our approach has high authenticity and structural integrity even for higher-resolution editing. For example, in the last row of Fig. 5, when the target prompt is “… standing on grass”, our method generates results by modifying the texture of the land on which the dog stands with other features intact. However, other methods result in a dramatic change in the structure of the whole image. Moreover, in the second row, when modifying the painting Mona Lisa, both DreamBooth and Textual-Inversion fail to edit the image. Our work also shows clear advantages over the approaches on training-free editings, such as ILVR , SDEdit , and Prompt-to-prompt , with the qualitative comparisons presented in the Appendix.

3 Ablation Analysis

Patch-based fine-tuning. In Fig. 5, we show the results of editing images in higher resolution when fine-tuned without or with the proposed patch-based fine-tuning technique (w/o pos vs. w/ pos). When sampling at a higher resolution, as in the right part of Fig. 5, the denoising model fine-tuned without the patch-based training mechanism performs poorly. In the first row, the castle towers get duplicated to meet the resolution, and in the third row, the bench gets stretched disproportionately. In essence, our patch-based fine-tuning technique enables the diffusion model to leverage the super-resolution ability of the decoder and edit images at arbitrary resolution during testing time.

Analysis of model-based classifier-free guidance. We generate editing results by directly sampling from the fine-tuned model using Eqn. 2. We fine-tune the model without the patch-based schedule for 800800 steps to retain more generalization ability. In this case, the model can perfectly reconstruct the source image while preserving as much editing ability as possible. We denote the setting as w/o gudiance. Using the same fine-tuned model, we conduct experiments under model-based classifier-free guidance, which we denote as w/ guidance. As shown in Fig. 6, sampling without our model-based classifier-free guidance fails to react to the prompt, while our method can successfully edit images to match the target language guidance. We further analyze two hyper-parameters (KK and vv) in model-based classifier-free guidance.

Analysis on guidance step KK in Sec. 3.2. In Fig. 7, we show our results on editing with different settings of KK. We conduct this set of experiments by editing one single image with the same language guidance at 768×768768\times 768 resolution. We set v=0.7v=0.7. When K=0K=0, the model-based classifier-free guidance is applied for each step of the denoising process. Since the generalization ability of the fine-tuned model is limited, in this case, the model fails to apply the desired property to the single source image. When K=1,000K=1,000, the model-based classifier-free guidance is not applied to any step. Thus, the structure of the image is not preserved, and the generated result becomes a random sample of the pre-trained model.

We further show the quantitative results Fig. 8(a). We repeat the abovementioned procedure over different KK and randomly sample 2020 images for each KK. We calculate two metrics. To understand the editing result, the image fidelity that is measured by the LPIPS distance between the original and the edited image. The text alignment calculated by the CLIP score to understand the alignment between our generated images and target text. As can be seen, the image fidelity drops with the increase of KK, indicating more details provided by the pre-trained model instead of the fine-tuned one. The text alignment measurement improves since the more details generated by the pre-trained model, the better editing results align with the target domain. To preserve the edit result’s authenticity and fidelity to the source image, we set KK as 400400.

Analysis on guidance weight vv in Sec. 3.2. We further study the impact of the guidance weight (vv) in Fig. 9. We set K=400K=400 and resolution as 768×768768\times 768 for each edit and use the same random seed to generate the result. As can be observed, the value of vv controls the fidelity of the edit result. However, since the pre-trained model is trained at the resolution 512×512512\times 512, the generated image contains many artifacts. When v=1v=1, the synthesized image entirely depends on the results from the pre-trained model. Additionally, we conduct quantitative experiments with LPIPS score measuring the image fidelity and CLIP score for the text alignment in Fig. 8(b). When vv is close to 11, the fidelity decreases while the edited feature decreases. When vv is close to , the model relies mainly on the fine-tuned model for the output when t>Kt>K. However, since the fine-tuned model contains poor generalization ability, there is a significant amount of artifacts in the generated results, which leads to a poor LPIPS score. We choose 0.70.7 for each edit in this work as a trade-off between fidelity and creativity.

4 More Editing Tasks

Face manipulation. Our method demonstrates promising editing ability for in-the-wild human faces. As shown in Fig. 10, our approach can edit locally and globally on human faces for various facial manipulation tasks, e.g., image stylization, adding accessories, and age changing.

Content removal. In Fig. 11(a), we show the content removal using our method. We fine-tune the pre-trained large-scale text-to-image model with the language descriptor as “a [∗\ast] dog with a flower in mouth”. For sampling, we use text prompts such as “a dog” and “a [∗\ast] dog” for the pre-trained and fine-tuned models. The pre-trained model successfully removes the flower held in the mouth of the dog.

Style generation. Our method can also be employed to learn the underlying style of an image. As shown in Fig. 11(b), the model is fine-tuned with the text, “a painting in the [∗\ast] style”. When sampling results, we feed the pre-trained model a prompt as “painting of a forest”. The model can successfully synthesize images with the specified content in the style of the given real image.

Style transfer. Our model-based classifier-free guidance can be leveraged to combine multiple models for providing the guidance. We show the result in Fig. 11(c) by doing a style transfer task with dual-model guidance. We fine-tune two models using prompts: “picture of a [∗\ast] dog” and “painting in [∗\ast] style”. During inference, we give the pre-trained model the prompt “painting of a dog” and fine-tuned models with prompts the same as training. With guidance from two separate models, our method can generate images with the content from one and style from the other and achieve stylized generation.

Conclusion

This work introduces SINE, a method for single-image editing. With only one image and a brief description of the object in the image, our approach can enable a wide range of editing for arbitrary resolution, followed by the information depicted in the language guidance. To achieve such results, we leverage the pre-trained large-scale text-to-image diffusion model. Specifically, we first fine-tune the pre-trained model with our patch-based fine-tuning method until it overfits the single image. Then, during sampling time, we use the overfitted model to guide the pre-trained diffusion model for image synthesis, which maintains the fidelity of the results while taking advantage of the generalization ability of the pre-trained model. Compared with other methods, our approach has a better geometrical understanding of the image and thus can conduct complex editing to the images besides style transfer.

However, in some cases where confusing editing guidance is given for the diffusion model, e.g., a chair-shaped dog, our method could fail. In cases where drastic changes are to be applied, e.g., changing a dog to a tiger in the same posture, there are also noticeable artifacts. We show more examples in the Appendix.

One future direction is improving the fidelity of the editing results, which could be achieved by alleviating the overfitting problem of the fine-tuned model.

References

Appendix

In the Appendix, we provide the following:

More comparisons with exiting works on editing single real image (Appendix A).

More results for applying SINE on editing a single image and the novel image manipulation tasks that can be enabled by our approach (Appendix B).

Discussion about the limitation of our method and possible future work (Appendix D).

Appendix A More Comparisons

Besides the comparison with existing works shown in the main paper, we provide more results by comparing our approach with Prompt-to-Promt . In addition, we compare our methods with training-free single-image editing approaches, including SDEdit and ILVR .

We first show the technical differences between our works and training-free methods in Tab. 1. SDEdit applies the diffusion process on an image or a user-created semantic map to conduct the denoising procedure, conditioned with the desired output. ILVR guides the denoising process by replacing the low-frequency part of the sample with that of the target reference image.

The visual comparisons are illustrated in Fig. 12. As can be seen, our approach significantly outperforms other methods for generating high-fidelity images with the maximal keeping of the details in the source image.

Appendix B More Editing Results

We provide more editing results in Fig. 13, Fig. 14, Fig. 15, Fig. 16, Fig. 17, and Fig. 18. All results are obtained by fine-tuning the large-scale text-to-image model using our proposed patch-based method at the resolution of 512×512512\times 512 and sampling with our introduced model-based classifier-free guidance at a higher resolution, e.g., 768×1024768\times 1024. Images on the top-left corner of these results are the real images utilized for fine-tuning. We specify the hyper-parameters used during sampling in the caption of each image.

Appendix C More Ablations

Analysis on guidance step KK and guidance weight vv. We conduct experiments by varying the guidance step KK and guidance weight vv in Fig. 19, Fig. 20, and Fig. 21. We use the same random seed and generate results with specific text prompts at a fixed resolution by varying the parameters. These experiments show the same behavior of our approach as mentioned in Sec 4.3. By adjusting these two parameters, we can find an optimal combination specifically for the image and the target language guidance. In most cases, we adopt the parameters setting of K=400K=400 and v=0.7v=0.7. However, we want our model to maintain more fidelity or apply a stronger edit in some instances. For example, the “optimal” setting we decide for experiments in Fig. 19 is v=0.5v=0.5 and K=400K=400.

Analysis on regularization loss. Dreambooth proposes to leverage Prior-Preservation Loss(PPL) to address the issues of overfitting and language drift. They propose to generate 200200 samples with the pre-trained model using the prompt “a [class noun]”. Then, during fine-tuning, they use these samples to regulate the model with the Prior-Preservation Loss to maintain the generalization ability of the model. However, in our experiments, as shown in Fig. 23, this loss does not improve the final results due to the uniqueness of certain pictures/paintings. On the contrary, more artifacts are introduced to the results, and the fidelity of the editing results decreases. Therefore, given the motivation of editing unique images, we forfeit the generalization ability provided by regularizing the model with the samples generated by the pre-trained model. We encourage our model to overfit a single image for the fidelity of the editing results.

Appendix D Limitations

We present some failure cases in Fig. 23. As mentioned in the main paper, when confusing guidance is given to the model or drastic change is to be applied, our method produces unsatisfying results. The language comprehension limitation of the pre-trained model and the over-fitting issue of our fine-tuned model can cause this. It would be an interesting future direction to explore how to over-fit on one single image without “forgetting” prior knowledge.

Also, as can be noticed in the second row of Fig. 10, the color of the sweater is changed in most cases. Also, the background letters are twisted after editing. Even though our method can perform editing with maximal protection of the details in the source image, editing strictly on a specific part of an image is also worth further exploration.