Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models

Daiki Miyake, Akihiro Iohara, Yu Saito, Toshiyuki Tanaka

Introduction

Diffusion models are known to yield high-quality results in the fields of image generation , video generation , and text-to-speech conversion . Text-guided diffusion models are diffusion models conditional on given texts (“prompts”), which can generate data with various modalities that fit well with the prompts. It is known that by strengthening the text conditioning through classifier-guidance or classifier-free guidance (CFG) , the fidelity to the text can be improved further. In image editing using text-guided diffusion models, elements in images, such as objects and styles, can be changed with high quality and diversity guided by text prompts.

In applications based on image editing methods, one must be able to generate images that are of high fidelity to original images in the first place, including reproduction of their details, and then one will be able to perform appropriate editing of images according to the prompts therefrom. To achieve high-fidelity image generation, most existing research exploits optimization of parameters such as model weights, text embeddings, and latent variables, which results in high computational costs and memory usage.

In this paper, we propose a method that can obtain latent variables and text embeddings yielding high-quality reconstruction of real images while using only forward computations. Our method does not require optimization, enabling faster processing and reducing computational costs. The proposed method is based on null-text inversion , which has the denoising diffusion implicit model (DDIM) inversion and CFG as its principal building blocks. DDIM inversion computes the diffusion process by inverting the reverse diffusion process calculated by the unconditional diffusion model. CFG computes the reverse diffusion process using a linear combination of the text-conditional prediction and the unconditional prediction. In practice, a specific embedding is used as the input for the unconditional prediction, which is the null-text embedding. Null-text inversion improves the reconstruction accuracy by optimizing the embedding vector of an empty string so that the diffusion process calculated by DDIM inversion aligns with the reverse diffusion process calculated using CFG. We discovered that the optimal solution obtained by this method can be approximated by the embedding of the conditioning text prompt.

Figure 1 shows a comparison between the proposed method and existing ones. Our method generated high-quality reconstructions when a real image and a corresponding prompt were given. DDIM inversion had noticeably lower reconstruction accuracy. Null-text inversion achieved high-quality results, nearly indistinguishable from the input image, but required more computation time. The proposed method, which we call negative-prompt inversion, allows for computation at the same speed as DDIM inversion, while achieving accuracy comparable to null-text inversion. Furthermore, combining our method with image editing methods such as prompt-to-prompt allows fast single-image editing (Editing).

We summarize our contributions as follows:

We propose a method for fast reconstruction of real images with diffusion models without optimization.

We experimentally demonstrated that our method achieves visually equivalent reconstruction quality to existing methods while enabling a more than 30-fold increase in processing speed.

Combining our method with existing image editing methods like prompt-to-prompt allows fast real image editing.

Related work

In the field of image editing using diffusion models such as Imagen and Stable Diffusion , Imagic , UniTune , and SINE are models for editing compositional structures, as well as states and styles of objects, in a single image. These methods ensure fidelity to original images via fine-tuning models and/or text embeddings.

Prompt-to-prompt , another image editing methods based on diffusion models, reconstructs original images via making use of null-text inversion. Null-text inversion successfully reconstructs real images by optimizing the null-text embedding (the embedding for unconditional prediction) at each prediction step. All these methods attempt to reconstruct real images by incorporating an optimization process, which typically takes several minutes for editing a single image.

Plug-and-Play edits a single image without optimization. It obtains latent variables corresponding to the input image using DDIM inversion and reconstructs it according to the edited prompt, inserting attention and feature maps to preserve image structures. Our inversion method is independent of editing methods, allowing for the freedom to choose an editing method while maintaining a high-quality image structure regardless of the editing method.

Image reconstruction by diffusion models.

Textual Inversion and DreamBooth are methods that reconstruct common concepts from a few real images by fine-tuning the model. On the other hand, ELITE and Encoder for Tuning (E4T) seek text embeddings that reconstruct real images using an encoder. The former ones are aimed at concept acquisition, making them difficult to reconstruct the original image with high fidelity. Although the latter ones require less computation time compared with the former ones, the ease of editing operations is limited, as the corresponding text is not explicitly obtained.

The proposed method realizes nearly the same reconstruction as null-text inversion, but with only forward computation, enabling image editing in just a few seconds. By combining our method with image editing methods such as prompt-to-prompt, it becomes possible to achieve flexible and advanced editing using text prompts.

Method

In this section, we describe our method for obtaining latent variables and text embeddings which reconstruct a real image using diffusion models without optimization. Our goal is that when given a real image II and an appropriate prompt PP, we calculate latent variables (zt)(\bm{z}_{t}), where tt is the index for the diffusion steps, in the reverse diffusion process so as to reconstruct II.

2 DDIM inversion

A diffusion model has a forward diffusion process over diffusion steps from to TT (e.g., T=1000T=1000 in ), which degrades the representation z0\bm{z}_{0} of an original sample into a pure noise zT\bm{z}_{T}, and an associated reverse diffusion process, which is used in the inference phase to generate z0\bm{z}_{0} from zT\bm{z}_{T}. In the training process, a latent variable zt\bm{z}_{t} for t∈{1,⋯ ,T}t\in\{1,\cdots,T\} is calculated by adding noise ϵ\bm{\epsilon} to z0\bm{z}_{0} at a scale defined at each diffusion step tt, and the model is trained to predict the noise ϵ\bm{\epsilon} from the latent variables zt\bm{z}_{t}. In text-guided diffusion models, the model is further conditioned by a feature representation (an embedding) CC of a text prompt PP, which is obtained via a text encoder like CLIP . The loss function for training the model is the mean squared error (MSE) between the predicted noise ϵθ\bm{\epsilon}_{\theta} and the actual noise ϵ\bm{\epsilon},

where U(1,T)U(1,T) denotes the uniform distribution on the set {1,⋯ ,T}\{1,\cdots,T\}, and where N(μ,Σ)\mathcal{N}(\bm{\mu},\bm{\Sigma}) denotes the multivariate Gaussian distribution with mean μ\bm{\mu} and covariance Σ\bm{\Sigma}.

Stable Diffusion considers diffusion processes in a latent space: during the training process, a feature representation z0\bm{z}_{0} in the latent space is obtained by passing a sample x0x_{0} through an encoder. In the inference stage, a sample x0x_{0} is generated by passing the generated representation z0\bm{z}_{0} through a decoder.

CFG is used to strengthen text conditioning. During the computation of the reverse diffusion process, the null-text embedding ∅\varnothing, which corresponds to the embedding of a null text “”, is used as a reference for unconditional prediction to enhance the conditioning:

where the guidance scale w≥0w\geq 0 controls strength of the conditional prediction ϵθ(zt,t,C)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,C) against the unconditional prediction ϵθ(zt,t,∅)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,\varnothing).

Although the reverse diffusion process of the denoising diffusion probabilistic models (DDPMs) is stochastic, the reverse process of the DDIM sampling method becomes deterministic and still generates the same data distribution. Therefore, we will use the DDIM for the sampling method in the following discussion. DDIM calculates from the latent variable zt\bm{z}_{t} at the diffusion step tt the latent variable zt−1\bm{z}_{t-1} at the diffusion step (t−1)(t-1) via

3 Null-text inversion

DDIM is known to work well: Given an original sample, by performing the forward process starting from the representation z0\bm{z}_{0} of the sample to obtain zT\bm{z}_{T} and then by inverting the forward process, one can reconstruct the original sample with high quality when we do not use CFG (i.e., w=1w=1). However, since CFG is useful to strengthen the text conditioning, it is necessary to reconstruct original samples well even when one uses CFG (i.e., w>1w>1). Null-text inversion enables us to faithfully reconstruct given samples even when using CFG by optimizing the null-text embedding ∅\varnothing at each diffusion step tt.

First, we calculate the sequence of latent variables (zt∗)t∈{1,⋯ ,T}(\bm{z}^{*}_{t})_{t\in\{1,\cdots,T\}} from z0\bm{z}_{0} via DDIM inversion. Next, we do initialization with zˉT=zT∗\bar{\bm{z}}_{T}=\bm{z}^{*}_{T} and ∅T=∅\varnothing_{T}=\varnothing. We then iteratively optimize ∅t\varnothing_{t} and calculate zˉt−1\bar{\bm{z}}_{t-1} from zˉt\bar{\bm{z}}_{t} for t=Tt=T to 11 as follows: At each diffusion step tt, assuming that we have zˉt\bar{\bm{z}}_{t}, one can calculate zt−1\bm{z}_{t-1} via DDIM (2) and CFG (1) with the null-text embedding ∅t\varnothing_{t} as

Then, we optimize ∅t\varnothing_{t} to minimize the MSE between the predicted zt−1(zˉt,t,C,∅t)\bm{z}_{t-1}(\bar{\bm{z}}_{t},t,C,\varnothing_{t}) and zt−1∗\bm{z}^{*}_{t-1}:

with the initialization ∅t=∅t+1\varnothing_{t}=\varnothing_{t+1}. After several updates (e.g., 10 iterations), we fix ∅t\varnothing_{t} and set zˉt−1=zt−1(zˉt,t,C,∅t)\bar{\bm{z}}_{t-1}=\bm{z}_{t-1}(\bar{\bm{z}}_{t},t,C,\varnothing_{t}). By performing the optimization at t=T,…,1t=T,\ldots,1 sequentially, we can reconstruct the original image with high quality even when using CFG with w>1w>1.

4 Negative-prompt inversion

The proposed method, negative-prompt inversion, utilizes the text prompt embeddings CC instead of the optimized null-text embeddings (∅t)t∈{1,…,T}(\varnothing_{t})_{t\in\{1,\ldots,T\}} in null-text inversion. As a result, we can perform reconstruction with only forward computation without optimization, significantly reducing computation time.

Let us assume that at diffusion step tt in null-text inversion one has zˉt\bar{\bm{z}}_{t} that is close enough to zt∗\bm{z}_{t}^{*}, so that one can regard zˉt=zt∗\bar{\bm{z}}_{t}=\bm{z}^{*}_{t} to hold. In null-text inversion, one obtains zt−1\bm{z}_{t-1} from zˉt\bar{\bm{z}}_{t} by moving one diffusion step backward using (4). Recall that zt∗\bm{z}^{*}_{t} was calculated from zt−1∗\bm{z}^{*}_{t-1} by moving one diffusion step forward in the diffusion process using (3):

As we have assumed zˉt=zt∗\bar{\bm{z}}_{t}=\bm{z}_{t}^{*}, one can substitute the above into (4), yielding

It implies that the discrepancy between zˉt−1\bar{\bm{z}}_{t-1} and zt−1∗\bm{z}^{*}_{t-1} in null-text inversion will be minimized when the predicted noises are equal:

If furthermore we are allowed to assume that the predicted noises at adjacent diffusion steps are equal, i.e., ϵθ(zt−1∗,t−1,C)=ϵθ(zt∗,t,C)=ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bm{z}^{*}_{t-1},t-1,C)=\bm{\epsilon}_{\theta}(\bm{z}_{t}^{*},t,C)=\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C), then we can deduce that at the optimum the conditional and unconditional predictions are equal:

Therefore, the optimized ∅t\varnothing_{t} can be approximated by the prompt embedding CC, so that we can discard the optimization of the null-text embedding ∅t\varnothing_{t} in null-text inversion altogether, simply by replacing the null-text embedding ∅t\varnothing_{t} with CC. The argument so far has the following two consequences:

For editing, optimizing ∅t\varnothing_{t} in null-text inversion can be replaced by the simple substitution ∅t=C\varnothing_{t}=C, thereby avoiding the computationally-heavy optimization.

Experiments

In this section, we evaluate the proposed method qualitatively and quantitatively. We experimented it using Stable Diffusion v1.5 in Diffusers implemented with PyTorch . Our code used in the experiments is provided in SM. Following , we used 100 images and captions, randomly selected from validation data in COCO dataset , in our experiments. The images were trimmed to make them square and resized to 512×512512\times 512. Unless otherwise specified, in both DDIM inversion and sampling we set the number of the sampling steps to be 50 via using the stride of 20 over the T=1000T=1000 diffusion steps.

We compared our method with DDIM inversion with CFG and null-text inversion, and evaluated its quality by peak signal-to-noise ratio (PSNR) and learned perceptual image patch similarity (LPIPS) . See Appendix B in SM for our setting of null-text inversion. The inference speed was measured on one NVIDIA RTX A6000 connected to one AMD EPYC 7343 (16 cores, 3.2 GHz clockspeed).

2 Reconstruction

Table 1 shows PSNR, LPIPS, and inference time of reconstruction by the three methods compared. In terms of PSNR (higher is better) and LPIPS (lower is better), the reconstruction quality of the proposed method was slightly worse than that of null-text inversion but far better than that of DDIM inversion with CFG. On the other hand, the inference speed was 30 times as fast as null-text inversion. This acceleration is achieved since the iterative optimization and backpropagation processing required for null-text inversion are not necessary for our method.

In Figure 3, the left four columns display examples of reconstruction by the three methods. DDIM inversion with CFG reconstructed images with noticeable differences from the input images, such as object position and shape. In contrast, null-text inversion and negative-prompt inversion (Ours) were capable of reconstructing images that are nearly identical to the input images, and the proposed method achieved a high reconstruction quality comparable to that of null-text inversion. See Appendix C.1 in SM for additional reconstruction examples. These results suggest that the proposed method can achieve reconstruction quality nearly equivalent to null-text inversion, with a speedup of over 30 times.

3 Editing

We next demonstrate the feasibility of editing real images by combining our inversion method with existing image editing methods. Our method is independent of the image editing approach and is principally compatible with any method that uses CFG, allowing for the selection of an appropriate image editing method depending on the objective. Here, we verify the effectiveness of our method for real-image editing using prompt-to-prompt in the same manner as in .

The rightmost column of Fig. 3 shows examples of real-image editing via prompt-to-prompt using the proposed method. The proposed method managed to maintain the composition while editing the image according to the modified prompt, such as replacing the objects and changing the background. Additional editing examples are provided in Appendices C.2 and C.3 in SM. These observations show that our inversion method can be combined with editing methods like prompt-to-prompt to enable rapid real-image editing.

4 Number of sampling steps

As the proposed method allows fast reconstruction/editing, one may be able to use a larger number of sampling steps to further improve quality, at the expense of reduced speed. To investigate the relationship between the number of sampling steps and reconstruction quality, we measured the PSNR and LPIPS using four different sampling steps: 20, 50, 100, and 200.

Figure 4 shows PSNR, LPIPS, and speed versus the number of sampling steps by the three methods. Although results with high enough quality were obtained with 50 sampling steps, increasing the number of sampling steps further improved the reconstruction quality of the proposed method, approaching that by null-text inversion. It should be noted that the total execution time is roughly given by the product of the execution time per sampling step and the number of sampling steps so that even if the proposed inversion method is performed with 200 sampling steps, it would still take less time than executing null-text inversion with 50 sampling steps thanks to the 30×30\times speedup. In fact, Figure 4 right shows the time taken for inversion; with 200 sampling steps, it took 18.1 seconds, which is approximately seven times faster than the null-text inversion at 50 sampling steps, which took 130 seconds. We would like to note that in Fig. 4 right the execution time of null-text inversion was not proportional to the number of sampling steps, since in our experimental setting the early stopping employed in the null-text optimization was more effective as the number of sampling steps became larger.

Figure 5 describes how the reconstructed image changed as the number of sampling steps was increased. Even with a small number of sampling steps, such as 20, the input image’s objects and composition were successfully reconstructed. Focusing on the finer details, for example, the head of the bed and the desk in the first row, and the wall color and pipes on the wall in the second row, we observe that the reconstruction quality improved as the number of sampling steps was increased. This improvement is generally imperceptible at first glance, suggesting that conventionally adopted the number of sampling steps, such as 20 or 50 sampling steps, yields sufficiently satisfactory reconstruction results for practical purposes.

Limitations

A limitation of the proposed method is that the average reconstruction quality does not reach that of null-text inversion. As demonstrated in the previous section, the difference is generally imperceptible at first glance; however, there were instances where our inversion method failed significantly. Figure 6 shows an image example which the proposed method failed to reconstruct. We observed that the proposed method tends to fail in reconstructing persons. Such failures could be attributed to characteristics of Stable Diffusion’s AutoEncoder, which struggles to reconstruct human faces. In such cases, employing a more effective encoder-decoder pair may result in improvements. Moreover, some failure cases were improved by increasing the number of sampling steps. Additional failure cases are presented in Appendix C.4 in SM.

Although failures in post-reconstruction image editing may occur, our inversion method is independent of editing methods, making the related discussion beyond the scope of this paper.

Conclusions

We have proposed negative-prompt inversion, which enables real-image inversion in diffusion models without the need for optimization. Experimentally, it produced visually high-quality reconstruction results comparable to inversion methods requiring optimization, while achieving a remarkable speed-up of over 30 times. Furthermore, we discovered that increasing the number of sampling steps further improved the reconstruction quality while maintaining faster computational time than existing methods.

On the basis of these results, our method provides a practical approach for real-image reconstruction. This utility excels in high-computational-cost scenarios, such as video editing, where our method proves to be even more beneficial. Although the proposed approach reduces computational costs and is available to any user, it does not encourage socially inappropriate use.

References

Appendix A Justifying arguments

In this appendix, we firstly show that in the non-Markovian forward process of DDIM one can regard that z1:T\bm{z}_{1:T} are all on a straight line, that is, there exists d\bm{d} such that zt=αtz0+1−αtd\bm{z}_{t}=\sqrt{\alpha_{t}}\bm{z}_{0}+\sqrt{1-\alpha_{t}}\bm{d} for every t=1,…,Tt=1,\ldots,T. On the basis of this observation, we next show a detailed derivation of the proposed method. An issue of conditioning on z0\bm{z}_{0} will be discussed at the end of this section.

for all t=1,…,Tt=1,\ldots,T. The forward process in DDIM is obtained from the above family via taking the deterministic limit σ→0\bm{\sigma}\to\mathbf{0}.

Let dt:=(zt−αtz0)/1−αt\bm{d}_{t}:=(\bm{z}_{t}-\sqrt{\alpha_{t}}\bm{z}_{0})/\sqrt{1-\alpha_{t}} be the normalized noise component in zt\bm{z}_{t} relative to αtz0\sqrt{\alpha_{t}}\bm{z}_{0}. We study the joint distribution of d1:T\bm{d}_{1:T} conditional on z0\bm{z}_{0}, induced by the above non-Markovian forward process. One has

for all t=1,…,Tt=1,\ldots,T. Conditional on z0\bm{z}_{0}, d1:T\bm{d}_{1:T} are jointly Gaussian with mean zero.

Given z0\bm{z}_{0}, let z1:T\bm{z}_{1:T} be generated from z0\bm{z}_{0} via the forward process of DDIM, and let d1:T\bm{d}_{1:T} be defined via dt=(zt−αtz0)/1−αt\bm{d}_{t}=(\bm{z}_{t}-\sqrt{\alpha_{t}}\bm{z}_{0})/\sqrt{1-\alpha_{t}}. Then, for any m=0,1,…,T−1m=0,1,\ldots,T-1 and any t=m+1,…,Tt=m+1,\ldots,T, one has, conditional on z0\bm{z}_{0},

where the convergence is in the mean-square sense.

For any m=0,1,…,T−1m=0,1,\ldots,T-1 and any t=m+1,…,Tt=m+1,\ldots,T, the cross-covariance between dt−m\bm{d}_{t-m} and dm\bm{d}_{m} conditional on z0\bm{z}_{0} is given by

One then has, in the deterministic limit σ→0\bm{\sigma}\to 0,

where DD is the dimension of zt\bm{z}_{t}. This proves that, in the deterministic limit σ→0\bm{\sigma}\to\mathbf{0} of DDIM, dt−m−dt\bm{d}_{t-m}-\bm{d}_{t} converges to 0\mathbf{0} in the mean-square sense. ∎

In other words, in DDIM one can regard that, given z0\bm{z}_{0}, dt=(zt−αtz0)/1−αt\bm{d}_{t}=(\bm{z}_{t}-\sqrt{\alpha_{t}}\bm{z}_{0})/\sqrt{1-\alpha_{t}} is independent of tt, and that z1:T\bm{z}_{1:T} are all aligned along the straight line connecting z0\bm{z}_{0} and zT\bm{z}_{T}.

Assuming that zt\bm{z}_{t} is available, the model ϵθ(zt,t)\bm{\epsilon}_{\theta}(\bm{z}_{t},t) attempts to estimate dt\bm{d}_{t} from zt\bm{z}_{t}, which in turn yields an estimate fθ(t)(zt):=(zt−1−αtϵθ(zt,t))/αtf_{\theta}^{(t)}(\bm{z}_{t}):=(\bm{z}_{t}-\sqrt{1-\alpha_{t}}\bm{\epsilon}_{\theta}(\bm{z}_{t},t))/\sqrt{\alpha_{t}} of z0\bm{z}_{0}, and then one can use it to estimate zt−1\bm{z}_{t-1} by plugging it into qσ(zt−1∣zt,z0)q_{\bm{\sigma}}(\bm{z}_{t-1}\mid\bm{z}_{t},\bm{z}_{0}). Specifically, zt−1\bm{z}_{t-1} is estimated via

where nt∼N(0,I)\bm{n}_{t}\sim\mathcal{N}(\mathbf{0},\bm{I}) is a standard Gaussian random noise vector independent of z0\bm{z}_{0}. In the generative process of DDIM the noise variances σt\sigma_{t} are sent to zero, so that the above formula is reduced to

which corresponds to (2) in the main text. As ϵθ(zt,t)\bm{\epsilon}_{\theta}(\bm{z}_{t},t) is expected to approximate dt\bm{d}_{t} via training, in the deterministic limit one can expect ϵθ(zt,t)=ϵθ(zt−1,t−1)\bm{\epsilon}_{\theta}(\bm{z}_{t},t)=\bm{\epsilon}_{\theta}(\bm{z}_{t-1},t-1) to hold once the training has been done successfully.

We next show a derivation of the formula (3) in the main text for DDIM inversion. The joint distribution of dt−1\bm{d}_{t-1} and dt\bm{d}_{t} conditional on z0\bm{z}_{0} is a zero-mean Gaussian distribution with covariance matrix Σ\bm{\Sigma} given by

where ⊗\otimes denotes the Kronecker product. It immediately follows that the distribution of dt\bm{d}_{t} conditional on dt−1\bm{d}_{t-1} and z0\bm{z}_{0} is given by

It can be translated into the distribution of zt\bm{z}_{t} conditional on zt−1\bm{z}_{t-1} and z0\bm{z}_{0}, as

As above, assuming that zt−1\bm{z}_{t-1} is available, ϵθ(zt−1,t−1)\bm{\epsilon}_{\theta}(\bm{z}_{t-1},t-1) provides an estimate of dt−1\bm{d}_{t-1} using zt−1\bm{z}_{t-1}, which in turn yields an estimate (zt−1−1−αt−1ϵθ(zt−1,t−1))/αt−1(\bm{z}_{t-1}-\sqrt{1-\alpha_{t-1}}\bm{\epsilon}_{\theta}(\bm{z}_{t-1},t-1))/\sqrt{\alpha_{t-1}} of z0\bm{z}_{0}. One can then use it to estimate zt\bm{z}_{t} by plugging it into qσ(zt∣zt−1,z0)q_{\bm{\sigma}}(\bm{z}_{t}\mid\bm{z}_{t-1},\bm{z}_{0}), as

which corresponds to the formula (3) for DDIM inversion. It should be noted that in deriving (23) we have not assumed that the diffusion steps are small enough.

In what follows, we provide a justifying argument for the proposed method, via extending the argument so far by incorporating conditioning into the model. It is straightforward to incorporate conditioning in the DDIM inversion formula (23) and the DDIM sampling formula (18), by replacing the model ϵθ(zt,t)\bm{\epsilon}_{\theta}(\bm{z}_{t},t) without conditioning with the conditional model ϵθ(zt,t,C)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,C), where CC is the prompt embedding. In various applications, on the other hand, the reverse process using the DDIM sampling formula (3) is often combined with CFG to strengthen the effects of the conditioning, where the conditional model ϵθ(zt,t,C)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,C) is further replaced with

where w≥0w\geq 0 is the guidance scale, which controls the strength of the conditioning, and where ∅\varnothing is the null-text embedding.

The first step of null-text inversion is to obtain zt∗\bm{z}^{*}_{t} for t=1,…,Tt=1,\ldots,T by initializing z0∗=z0\bm{z}^{*}_{0}=\bm{z}_{0} and successively applying the forward process derived as the DDIM inversion formula:

where the model ϵθ(zt−1,t−1)\bm{\epsilon}_{\theta}(\bm{z}_{t-1},t-1) in (23) without conditioning has been replaced with the conditional model ϵθ(zt−1∗,t−1,C)\bm{\epsilon}_{\theta}(\bm{z}^{*}_{t-1},t-1,C).

Next, starting from zˉT=zT∗\bar{\bm{z}}_{T}=\bm{z}_{T}^{*}, we calculate the reverse diffusion process to obtain zˉt\bar{\bm{z}}_{t} in the backward direction, while optimizing the null-text embedding ∅t\varnothing_{t} at each diffusion step so that zˉt\bar{\bm{z}}_{t} well reproduces zt∗\bm{z}^{*}_{t}. More specifically, for t=T,T−1,…,1t=T,T-1,\ldots,1, zˉt−1\bar{\bm{z}}_{t-1} is calculated via combining the DDIM sampling (18) and CFG (24) as

The null-text embedding ∅t\varnothing_{t} is optimized to minimize the MSE between zt−1∗\bm{z}^{*}_{t-1} and zˉt−1\bar{\bm{z}}_{t-1} as

where zˉt−1\bar{\bm{z}}_{t-1} is dependent on ∅t\varnothing_{t} via (26). The following proposition shows that the choice ∅t=C\varnothing_{t}=C does minimize the MSE between zt−1∗\bm{z}^{*}_{t-1} and zˉt−1\bar{\bm{z}}_{t-1} under an ideal situation.

Assume that the guidance scale ww in CFG is not equal to 1. For any t=1,…,Tt=1,\ldots,T, if the model ϵ(zt,t,C)\bm{\epsilon}(\bm{z}_{t},t,C) is able to correctly predict the noise and if zt∗=zˉt\bm{z}^{*}_{t}=\bar{\bm{z}}_{t} holds true, then the difference between zt−1∗\bm{z}_{t-1}^{*} and zˉt−1\bar{\bm{z}}_{t-1} in null-text inversion is made equal to zero if and only if ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}) is equal to ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C).

The difference between zt−1∗\bm{z}^{*}_{t-1} and zˉt−1\bar{\bm{z}}_{t-1} is expressed as

In the second line of the above equation we used the assumption zt∗=zˉt\bm{z}^{*}_{t}=\bar{\bm{z}}_{t}, and in the third line we substituted (25) into zt∗\bm{z}_{t}^{*} above.

As described above, the model ϵθ(zt,t,C)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,C) attempts to estimate noise dt\bm{d}_{t} from zt\bm{z}_{t}, and the assumption that the model correctly predicts the noise, together with Proposition 1, implies that ϵθ(zt∗,t,C)=ϵθ(zt−1∗,t−1,C)\bm{\epsilon}_{\theta}(\bm{z}^{*}_{t},t,C)=\bm{\epsilon}_{\theta}(\bm{z}^{*}_{t-1},t-1,C) should hold. One therefore has

As we have assumed w≠1w\not=1, zt−1∗−zˉt−1\bm{z}^{*}_{t-1}-\bar{\bm{z}}_{t-1} is proportional to ϵθ(zˉt,t,C)−ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C)-\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}), and it is made equal to zero if and only if ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C) and ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}) are equal. ∎

Since we initialize zˉT=zT∗\bar{\bm{z}}_{T}=\bm{z}^{*}_{T} at diffusion step TT, recursive application of Proposition 2 shows, under the ideal situation that the model has learned perfectly, that one will have zˉt=zt∗\bar{\bm{z}}_{t}=\bm{z}_{t}^{*} for all tt via letting ∅t=C\varnothing_{t}=C. In other words, one can regard that null-text inversion optimizes the unconditional prediction to approach the conditional prediction at each diffusion step.

The argument presented so far is based on conditioning on sample z0\bm{z}_{0}, which is not justifiable in the actual process of DDIM sampling where there exists more than one sample and where the model does not look at z0\bm{z}_{0}. We thus extend the above argument via assuming z0\bm{z}_{0} to be generated according to a certain probability distribution p(z0)p(\bm{z}_{0}). More concretely, we assume z0∼p(z0)\bm{z}_{0}\sim p(\bm{z}_{0}) and d∼N(0,I)\bm{d}\sim\mathcal{N}(\mathbf{0},\bm{I}), which induces the diffusion path zt=αtz0+1−αtd\bm{z}_{t}=\sqrt{\alpha_{t}}\bm{z}_{0}+\sqrt{1-\alpha_{t}}\bm{d}, t∈{0,…,T}t\in\{0,\ldots,T\}, in DDIM according to the above discussion. Consequently, at position z\bm{z} and at timestep tt, the noise ϵ(z,t)\bm{\epsilon}(\bm{z},t) to be learned by the model ϵθ(z,t)\bm{\epsilon}_{\theta}(\bm{z},t) is not determined by a single sample z0\bm{z}_{0} but given by the posterior mean of d=(z−αtz0)/1−αt\bm{d}=(\bm{z}-\sqrt{\alpha_{t}}\bm{z}_{0})/\sqrt{1-\alpha_{t}} with respect to the posterior distribution of z0\bm{z}_{0} given z\bm{z}, which is obtained from the prior distributions z0∼p(z0)\bm{z}_{0}\sim p(\bm{z}_{0}) and d∼N(0,I)\bm{d}\sim\mathcal{N}(\mathbf{0},\bm{I}), as well as the likelihood p(z∣z0,d)=δ(z−αtz0−1−αtd)p(\bm{z}\mid\bm{z}_{0},\bm{d})=\delta(\bm{z}-\sqrt{\alpha_{t}}\bm{z}_{0}-\sqrt{1-\alpha_{t}}\bm{d}).

Assume z0∼p(z0)\bm{z}_{0}\sim p(\bm{z}_{0}) and d∼N(0,I)\bm{d}\sim\mathcal{N}(\mathbf{0},\bm{I}). Then the noise ϵ(z,t)\bm{\epsilon}(\bm{z},t) in DDIM at position z\bm{z} and at timestep tt, which is to be learned by the model ϵθ(z,t)\bm{\epsilon}_{\theta}(\bm{z},t), is given by

denotes the probability density function of the DD-dimensional standard Gaussian distribution, and where ⟨⋅⟩z0\langle\cdot\rangle_{\bm{z}_{0}} denotes expectation with respect to z0∼p(z0)\bm{z}_{0}\sim p(\bm{z}_{0}).

The joint distribution of z0\bm{z}_{0} and z\bm{z} is given by

from which the posterior distribution of z0\bm{z}_{0} given z\bm{z} is obtained as

The noise ϵ(z,t)\bm{\epsilon}(\bm{z},t) at z\bm{z} and tt, to be learned by the model, is given by the posterior mean of d=(z−αtz0)/1−αt\bm{d}=(\bm{z}-\sqrt{\alpha_{t}}\bm{z}_{0})/\sqrt{1-\alpha_{t}}, which is represented as (30), proving the proposition. ∎

Despite its complex appearance, one can see that ϵ(z,t)\bm{\epsilon}(\bm{z},t) in (30) is continuous in z\bm{z} and αt\alpha_{t} when αt∈(0,1)\alpha_{t}\in(0,1). This continuity implies that when αt\alpha_{t} and αt−1\alpha_{t-1} are close enough one can expect zt∗≈zt−1∗\bm{z}_{t}^{*}\approx\bm{z}_{t-1}^{*} and consequently ϵ(zt∗,t)≈ϵ(zt−1∗,t−1)\bm{\epsilon}(\bm{z}_{t}^{*},t)\approx\bm{\epsilon}(\bm{z}_{t-1}^{*},t-1) to hold.

A.2 Empirical evaluations

The assumption of perfect learning of the model adopted in Proposition 2 in the previous section is certainly too strong to be applied to practical situations. We have already discussed the issue of conditioning on z0\bm{z}_{0} in the previous section. Another reason is that it is almost always the case that the model learns only approximately. Accordingly, what one can expect in practice would be that zˉt=zt∗\bar{\bm{z}}_{t}=\bm{z}_{t}^{*} holds only approximately, which would then make the validity of the optimality of ∅t=C\varnothing_{t}=C in null-text inversion rather questionable. In this section, we investigate empirically how good the prompt-text embedding CC is compared with the optimized null-text embedding ∅t\varnothing_{t}, in terms of the noise prediction by the model, as well as their representation in the embedding space. In the experiments in this section, we used the same 100 image-prompt pairs from the COCO dataset as those used in the experiments in the main text.

We first investigated how close the noise prediction ϵθ(zt,t,∅t)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,\varnothing_{t}) using the optimized null-text embedding ∅t\varnothing_{t} and the prediction ϵθ(zt,t,C)\bm{\epsilon}_{\theta}(\bm{z}_{t},t,C) using the prompt embedding CC are. More specifically, we performed null-text inversion, starting from zT∗\bm{z}_{T}^{*} obtained via DDIM inversion using the embedding CC, and with the resulting sequences (zˉt)t∈{1,…,T}(\bar{\bm{z}}_{t})_{t\in\{1,\ldots,T\}} and (∅t)t∈{1,…,T}(\varnothing_{t})_{t\in\{1,\ldots,T\}} we evaluated the L1L_{1} distance between ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}) and ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C). For comparison, we also calculated the L1L_{1} distance between ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}) and the noise prediction ϵθ(zˉt,t,C′)\bm{\epsilon}_{\theta}(\bar{\bm{\bm{z}}}_{t},t,C^{\prime}) obtained using the embeddings C′C^{\prime} of the prompts associated with images other than the target image, as well as the L1L_{1} distance between ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C) and ϵθ(zˉt,t,C′)\bm{\epsilon}_{\theta}(\bar{\bm{\bm{z}}}_{t},t,C^{\prime}). Figure 7 left shows the mean L1L_{1} distance of the predicted noises. The predicted noises using the optimized embeddings (∅t)t∈{1,…,T}(\varnothing_{t})_{t\in\{1,\ldots,T\}} were closer to those using CC than those using C′C^{\prime}, with a smaller distance than the distance between the predicted noise using CC and that using C′C^{\prime}. One observes that the distance between the noise predictions using ∅t\varnothing_{t} and CC became larger as tt became smaller, which would be ascribed to the accumulation of optimization errors. One can also notice that the distance between the noise predictions using (∅t)t∈{1,…,T}(\varnothing_{t})_{t\in\{1,\ldots,T\}} and CC were larger than that between those using CC and C′C^{\prime} near t=0t=0. Noise predictions near t=0t=0, however, would have almost no impact on generated samples since they are added at very small scales. The results suggest that the predicted noise ϵθ(zˉt,t,∅t)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,\varnothing_{t}) using the optimized embedding ∅t\varnothing_{t} in null-text inversion can be well approximated by the noise prediction ϵθ(zˉt,t,C)\bm{\epsilon}_{\theta}(\bar{\bm{z}}_{t},t,C) using the embedding CC of the input prompt in (6).

We next calculated the cosine similarity in the 768-dimensional embedding space between the embeddings CC for 100 prompts and optimized embeddings (∅t)t∈{1,…,T}(\varnothing_{t})_{t\in\{1,\ldots,T\}} for each image. For each embedding sequence we took its average along the length of the sequence, and we centered the resulting average 768-dimensional prompt embeddings by subtracting the mean of 25,014 prompt embeddings, which are all the prompts included in the COCO validation dataset, and took a mean of embeddings over all tokens included in each prompt as the prompt embedding. Figure 7 right shows the mean cosine similarity. As tt became smaller, the similarity between the optimized null-text embedding (∅t)(\varnothing_{t}) and the embedding CC of the given prompt became positive, whereas the similarity between (∅t)(\varnothing_{t}) and embeddings C′C^{\prime} of the prompts for images other than the target image, as well as that between CC and C′C^{\prime}, remained around zero. (We postulate that the small negative values of the similarity between CC and C′C^{\prime} throughout the entire range of tt are due to the bias induced from the centering.) This suggests that, although the implicit “meaning” represented by the optimized null-text embedding was almost orthogonal to the “meanings” of those of randomly-chosen prompts, it was closer to the “meaning” represented by the input prompt embedding CC in the region distant from t=Tt=T, as can be observed by the larger values of similarity between the optimized null-text embedding and the embedding of the input prompt (Optimized vs Input prompt). In the region distant from t=Tt=T, except the region near t=0t=0, the model is thought to generate detailed information about the image, which should be crucial in obtaining a high-quality reconstruction, so that the higher values of similarity in this region would suggest that embeddings that would be good in the sense of yielding a good reconstruction are closer to the embedding CC of the target prompt. In the large-tt region, on the other hand, the optimized null-text embedding (∅t)(\varnothing_{t}) had small similarity with the embedding CC of the given prompt, which can be ascribed to the fact that the null-text optimization is initialized with the same null-text embedding ∅\varnothing, and is performed from t=Tt=T down to t=1t=1. Note that, in the large-tt region, the similarity values were around zero because early stopping in optimizing ∅t\varnothing_{t} was effective and optimization barely progressed.

From these results, we can say that the optimized embedding ∅t\varnothing_{t} becomes semantically similar to the input prompt embedding CC as the optimization progresses. Therefore, it has been confirmed that our inversion method approximates null-text inversion.

Appendix B Implementation details

In our experiments, for the null-text inversion, we used the same settings at 50 sampling steps as those in the implementation available on the GitHub page of . Optimization was performed with the Adam optimizer, and the learning rate was set to reach 5×10−35\times 10^{-3} at the last sampling step, changing linearly by the factor of 10−410^{-4} with the number of sampling steps. We further employed early stopping, and the threshold for early stopping was increased linearly in the number of sampling steps from 10−510^{-5} by the factor of 2×10−52\times 10^{-5}. We observed that when scheduling the learning rate and threshold with a function of diffusion steps, the reconstruction quality was getting worse. See our code included in SM for details for more detailed implementation settings of our experiments.

Appendix C Additional experimental results

Figure 8 shows additional images reconstructed by the three methods compared. All the results show that DDIM inversion produced reconstructions that were not similar to the input images, while null-text inversion almost perfectly reconstructed the input images, and that our method also yielded results which were close to the reconstructions by null-text inversion.

C.2 Comparison of edited images using prompt-of-prompt

Figure 9 shows additional images edited by prompt-to-prompt. As can be seen, DDIM inversion failed to perform editing while maintaining the details of the original images. On the other hand, null-text inversion and the proposed method are both capable of editing while maintaining details of the original images, including object replacement and style changes.

C.3 Comparison of edited images using SDEdit

We demonstrate the advantage of the proposed method that it can be combined with various editing methods. For this purpose, we performed editing experiments by combining the proposal with another editing method, SDEdit . In SDEdit, a certain ratio t0t_{0} is used as a hyperparameter to add noise to the sample z0\bm{z}_{0}, and the latent variable zt\bm{z}_{t} at the diffusion step t=t0⋅Tt=t_{0}\cdot T is obtained, which is then reconstructed by tracing the inverse diffusion process. For image editing, z0\bm{z}_{0} is obtained from the original image and an edited prompt is used during the inverse diffusion process calculation. We set the noisy sample zt\bm{z}_{t} calculated by DDIM inversion for null-text inversion and our negative-prompt inversion since they assume starting the sampling from zT\bm{z}_{T} calculated by DDIM inversion.

Figure 10 shows images edited by SDEdit. As can be observed, SDEdit could not reconstruct the input images, while negative-prompt inversion and the proposed method were able to reconstruct details of the input images and appropriately edit them as specified by the prompts.

C.4 Additional failure cases

Figure 11 shows additional failure cases of our method. In all the cases shown, our method failed to reconstruct the images in 50 sampling steps, whereas null-text inversion successfully reconstructed them. The first two rows show failures due to the disappearance of people, where the objects were either reconstructed as non-human or as different persons. The third and fourth rows show failures due to the color gradient being reconstructed as separate objects, such as a single duck being reconstructed as scattered pieces, and a tree trunk being reconstructed as a different object. The last row shows a failure due to the disappearance of a tiny object, where one of the ski poles was missing. In the duck example, the reconstruction quality improved by increasing the number of sampling steps.