EDICT: Exact Diffusion Inversion via Coupled Transformations

Bram Wallace, Akash Gokul, Nikhil Naik

Introduction

Using the iterative denoising diffusion principle, denoising diffusion models (DDMs) trained with web-scale data can generate highly realistic images conditioned on input text, layouts, and scene graphs . After image generation, the next important application of DDMs being explored by the research community is that of image editing. Models such as DALL-E-2 and Stable Diffusion can perform inpainting, allowing users to edit images through manual annotation. Methods such as SDEdit have demonstrated that both synthetic and real images can be edited using stroke or composite guidance via DDMs. However, the goal of a holistic image editing tool that can edit any real/artificial image using purely text has still not been achieved.

The generative process of DDMs starts with an initial noise vector (xTx_{T}) and performs iterative denoising (typically with a guidance signal e.g. in the form of text-conditional denoising), ending with a realistic image sample (x0x_{0}). To solve the image editing problem, running the reverse of this generative process is necessary. Formally, this problem is known as “inversion” i.e., finding the initial noise vector that produces the input image when passed through the diffusion process.

A naïve approach for inversion is to add Gaussian noise to the input image and perform a predefined number of diffusion steps, which typically results in significant distortions . A more robust method is adapting Denoising Diffusion Implicit Models (DDIMs) . Unlike the commonly used Denoising Diffusion Probabilistic Models (DDPMs) , the generative process in DDIMs is defined in a non-Markovian manner, which results in a deterministic denoising process. DDIM can also be used for inversion, deterministically noising an image to obtain the initial noise vector (x0→xTx_{0}\rightarrow x_{T}).

DDIM inversion has been used for editing real images through text methods such as DDIBs and Prompt-to-Prompt (P2P) image editing . After DDIM inversion, P2P edits the original image by running the generative process from the noise vector and injecting conditioning information from a new text prompt through the cross-attention layers in the diffusion model, thus generating an edited image that maintains faithfulness to the original content while incorporating the edit. However, as noted in the original P2P work , the DDIM inversion is unstable in many cases—encoding from x0x_{0} to xTx_{T} and back often results in inexact reconstructions of the original image as in Fig. 2. These distortions limit the ability to perform significant manipulations through text as increase in the corruption is correlated with the strength of the conditioning.

To improve the inversion ability of DDMs and enable robust real image editing, we diagnose the problems in DDIM inversion, and offer a solution: Exact Diffusion Inversion via Coupled Transformations (EDICT). EDICT is a re-formulation of the DDIM process inspired by coupling layers in normalizing flow models that allows for mathematically exact inversion. By maintaining two coupled noise vectors in the diffusion process, EDICT enables recovery of the original noise vector in the case of model-generated images; and for real imagery, initial noise vectors that are guaranteed to map to the original image when the EDICT generative process is run. While EDICT doubles the computation time of the diffusion process, it can be combined with any pretrained DDM model and does not require any computationally-expensive model finetuning, prompt tuning, or multiple images.

For the standard generative process, EDICT approximates DDIM well, resulting in nearly identical generations given equal initial conditions For real images, EDICT can recover a noise vector which yields an exact reconstruction when used as input to the generative process. Experiments with the COCO dataset show that EDICT can recover complex image features such as detailed textures, thin objects, subtle reflections, faces, and text, while DDIM fails to do so consistently. Finally, using the initial noise vectors derived from a real image with EDICT, we can sample from a DDM and perform complex edits or transformations to real images using textual guidance. We show editing capabilities including local and global modifications of objects and background and object transformations (Fig. 1).

Related Work

Diffusion Models and Normalizing Flows: Denoising diffusion models (DDMs), drawing on nonequilibrium thermodynamics , have emerged at the forefront of image generation. Models such as GLIDE, DALLE-2, Imagen (Video), Latent/Stable Diffusion, and eDiffi all utilize concepts borrowed from thermodynamics to hallucinate an image from pure noise by training on intermediately noised images. While a commonly used sampling process in DDMs is the stochastic Denoising Diffusion Probabilistic Models (DDPMs) method , a deterministic sampling method was introduced in Denoising Diffusion Implicit Models (DDIM) . Multi-step or higher-order methods that parallel DDIM have also been proposed , including methods that aim to reducing the computational time for generation . Another class of generative models relevant to our work are normalizing flow models . In these models, an invertible mapping is learned between a latent gaussian distribution and image space. The methods of invertibility, specifically coupling layers, are used as an inspiration for our method. Invertible neural networks have been studied in areas outside of normalizing flows as well. Neural ODEs have many parallels with the diffusion process and can be inverted using a variety of ODE solvers.

Editing in Diffusion Models: The seminal work in applying DDMs to image editing is SDEdit where coarse layouts are used to guide the generative process by noising the layout to resemble an intermediately noised image. Prompt-to-Prompt combines query-key pairs from one prompt with values from another in the attention layers of a DDM to enable prompt-guided image editing from intermediate latents obtained by sampler inversion. DiffEdit edits real/synthetic images using automatically generated masks for regions of an input image that should be edited given a text query. Kwon et al. introduce style and structure losses to guide the sampling process to enable text-guided image translation. CycleDiffusion uses a deterministic DPM encoder to enable zero-shot image-to-image translation. Another set of methods finetune the model with the target image and/or learn a new conditioning prompt to enable indirect image editing via sampling. EDICT, our proposed approach, does not require any specialized model training/finetuning or losses, and can be paired with any pretrained DDM.

Background

DDMs are trained on a simple denoising objective. A set of timesteps index a monotonic strictly increasing noising schedule {αt}t=0T,αT=0,α0=1\{\alpha_{t}\}^{T}_{t=0},\alpha_{T}=0,\alpha_{0}=1. Images (or autoencoded latents) x∈Xx\in X are noised with draws ϵ∼N(0,1)\epsilon\sim N(0,1) according to the noising schedule following the formula

The time-aware DDM Θ\Theta is trained on the objective MSE(Θ(xt,t,C),ϵ)MSE(\Theta(x_{t},t,C),\epsilon) to predict the noise added to the original image where CC is a conditioning signal (typically in the form of a text embedding) with some degree of dropout to the null conditioning ∅\emptyset. To generate a novel image from a gaussian draw ϵT∼N(0,1)\epsilon_{T}\sim N(0,1), partial denoising is applied at each tt. The most common sampling scheme is that of DDIM where intermediate steps are calculated as

In practice, for text-to-image models to hallucinate from random noise an x0x_{0} that matches conditioning CC to desired levels, the model has to be biased more heavily towards generations aligned with CC. To do so, a pseudo-gradient G⋅(Θ(xt,t,C)−Θ(xt,t,∅))G\cdot(\Theta(x_{t},t,C)-\Theta(x_{t},t,\emptyset)) is added to the unconditional prediction Θ(xt,t,∅)\Theta(x_{t},t,\emptyset) to up-weight the effect of conditioning, where GG is a weighting parameter, Substituting Φ(xt,t,C,G)=Θ(xt,t,∅)+G⋅(Θ(xt,t,C)−Θ(xt,t,∅))\Phi(x_{t},t,C,G)=\Theta(x_{t},t,\emptyset)+G\cdot(\Theta(x_{t},t,C)-\Theta(x_{t},t,\emptyset)) into the prior equation for the Θ\Theta term, we simplify the notation Φ(xt,t,C,G)→ϵ(xt,t)\Phi(x_{t},t,C,G)\xrightarrow{}\epsilon(x_{t},t) and rewrite the previous equation as xt−1=atxt+btϵ(xt,t)x_{t-1}=a_{t}x_{t}+b_{t}\epsilon(x_{t},t) where

2 Denoising Diffusion Implicit Model (DDIM)

As noted in DDIM , the above denoising process is approximately invertible; that is xtx_{t} is approximately recoverable from xt−1x_{t-1}

where the approximation is a linearization assumption that ϵ(xt,t)≈ϵ(xt−1,t)\epsilon(x_{t},t)\approx\epsilon(x_{t-1},t) (necessary due to the discrete nature of both computation and the underlying noise schedule). This corresponds with reversing the Euler integration which is a first-order ODE solver. More sophisticated solvers such as multi-step Euler have been shown to stabilize the generative process, and correspondingly the deterministic inversion process, with fewer time steps. However, such methods are also approximations where the inversion accuracy ultimately relies on the strength of the linearization assumption and the reconstruction is not exactly equal. This assumption is largely accurate for unconditional DDIM models, but the pseudo-gradient of classifier-free guidance G⋅(Θ(xt,t,C)−Θ(xt,t,∅))G\cdot(\Theta(x_{t},t,C)-\Theta(x_{t},t,\emptyset)) is inconsistent across time steps as shown in the Supplementary.

While unconditional reconstructions have relatively insignificant errors (Fig. 2), conditional reconstructions are extremely distorted when noised to high levels. This phenomenon was noted in , where the guidance scale must be heavily downweighted in order for inversions on real-world images to be stable, thus limiting the strength of edits. Obtaining an xtx_{t} from x0x_{0} allows for the generative process to be run with novel conditioning. In SDEdit , this process is done stochastically to obtain broad sample diversity, at the cost of controllability and faithfulness to the original image contents. In contrast, the inverse DDIM process produces a unique xtx_{t} from a single x0x_{0} in a deterministic manner, yielding only one sample but enabling higher strength edits while preserving finer-grain structure and content.

3 Affine Coupling Layers

Affine Coupling Layers (ACL) are invertible neural network layers introduced in and used in other normalizing flow models such as Glow . The layer input zz, is split into two equal-dimensional halves zaz_{a} and zbz_{b}. A modified version of zaz_{a} is then calculated, according to:

where Ψ\Psi and ψ\psi are neural networks. The layer output z′z^{\prime} is the concatenation of za′z^{\prime}_{a} and zbz_{b} in accordance with the original splitting function. ACL can parameterize complex functions and zz can be exactly recovered given z′z^{\prime}:

Noting the similarity of equations 6 and 7 to the simplified form of Eq. 2, we parallel this construction in our method (described next) where two separate quantities are tracked and alternately modified by transformations that are affine with respect to the original modified quantity and a non-linear transformation of its counterpart.

Exact Diffusion Inversion via Coupled Transformations (EDICT)

As summarized in Sec. 3.3, affine coupling layers track two quantities which can then be used to invert each other. These two quantities are partitions of a latent representation with a network specifically designed to operate in a fitting alternating manner. Without training of a new DDM, this method can not be applied out-of-the-box to the forward diffusion process. We consider the simplified form of the forward step equation from Sec. 3.1 below

If the noise prediction term, ϵ(xt,t)=ε\epsilon(x_{t},t)=\varepsilon was independent of xtx_{t}, this would be an affine function in both xtx_{t}, and ε\varepsilon. Paralleling Eq. 6, by creating a new variable yt=xty_{t}=x_{t} the stepping equation fits the desired form. Consider performing this computation, so we have the variables xt,  yt=xt,  xt−1=atxt+btϵ(yt,t)x_{t},\ \ y_{t}=x_{t},\ \ x_{t-1}=a_{t}x_{t}+b_{t}\epsilon(y_{t},t). xtx_{t} can be recovered exactly from xt−1x_{t-1} in the non-trivial form:

yt=xty_{t}=x_{t} is trivial and will not be true in the general case. Now consider the initialization of the reverse (denoising) diffusion process, where xT∼N(0,1)x_{T}\sim\mathcal{N}(0,1), we similarly initialize yT=xTy_{T}=x_{T}. Following the above process, we define the update rule

Note that the noise prediction term in the second line is a function of the other sequence value at the next timestep. Only one member of each sequence (xi,yjx_{i},y_{j}) must be held in memory at any given time. The sequences can be recovered exactly according to

As illustrated in Fig. 3, the entire sequence can be reconstructed from any two adjacent xix_{i} and yiy_{i}. In sum, our method re-uses the linearization assumption of Euler DDIM inversion, that ϵ(xt,t)≈ϵ(xt−1,t)\epsilon(x_{t},t)\approx\epsilon(x_{t-1},t), but crucially does not rely on it for invertibility, guaranteeing recovery up to machine precision. We note that the functional form of our approach bears similarities to the ODE solver Heun’s method where derivative values at initial predictions are used to refine predictions, but due to the need for invertibility we cannot exploit this for more accurate/faster sampling.

2 Stabilization

While our method assures invertibility by design, realism and faithfulness to the original diffusion process are not automatic. When naively applied for a typical, low number of DDIM steps (e.g., T=50T=50), the sequences xtx_{t} and yty_{t} can diverge (Fig. 4). This is a result of the strong linearization assumption not holding in practice. To alleviate this problem, we introduce intermediate mixing layers after each diffusion step computing weighted averages of the form

which are invertible affine transformations. Note that this averaging layer becomes a dilating layer during deterministic noising; the inversion being

A high (near 1) value of pp results in the averaging layer not being strong enough to prevent divergence of the xx and yy series during denoising, while a low value of pp results in a numerically unstable exponential dilation in the backwards pass (Fig. 5). Note that in both cases that our generative process remains exactly mathematically invertible with no further assumptions, but there is a degradation in utility of the results. Typically we employ p=0.93p=0.93, with values in the interval [0.9,0.97][0.9,0.97] generally being effective for 50 steps.

3 Complete Summary of the Method

We dub our presented process EDICT: Exact Diffusion Inversion via Coupled Transformations. In sum, EDICT uses a combination of “coupling” and “averaging/dilating” steps for exact inversion of the diffusion process. Given xtx_{t} and yty_{t}, we calculate the denoising process by

and the deterministic noising inversion process by:

Recall that the conditioning CC is implicitly included in the ϵ\epsilon terms. In practice, we alternate the order in which the xx and yy series are calculated at each step in order to symmetrize the process with respect to both sequences. We cast all operations to double floating point precision (from the native half precision) to mitigate roundoff floating point errors.

4 Image Editing

Given an image II, we edit the semantic contents to match text conditioning CtargetC_{target}. We describe the current content by text CbaseC_{base} in a parallel manner to CtargetC_{target}. We compute an autoencoder latent x0=VAEenc(I)x_{0}=VAE_{enc}(I), initializing y0=x0y_{0}=x_{0}. We run the deterministic noising process of Eq. 15 on (x0,y0)(x_{0},y_{0}) using text conditioning CbaseC_{base} for s⋅Ss\cdot S steps, where SS is the number of global timesteps and ss is a chosen editing strength. This yields partially “noised” latents (xt,yt)(x_{t},y_{t}) which are not necessarily equal and, in practice, tend to diverge by a small amount due to linearization error and the dilation of the mixing layers. These intermediate representations are then used as input to Eq. 14 using text condition CtargetC_{target}, and identical step parameters, (s,S)(s,S). The resulting image outputs (VAEdec(x0edit),VAEdec(y0edit))(VAE_{dec}(x^{edit}_{0}),VAE_{dec}(y^{edit}_{0})) are empirically nearly identical as seen in Fig. 4, this is by design of the method and in particular, the mixing layers. For all methods we find that a guidance scale of 3 performs well (as opposed to the standard 7.5 for generation).

Experiments

We now describe results using the Stable Diffusion 1.4 latent diffusion model .

We demonstrate the exact invertibility of EDICT using the MS-COCO-2017 validation set (n=5,000n=5,000), which contains both simple object-centric images and complex scene images . Given an image-caption pair, inverted latents are calculated and used to reconstruct the image. Mean-square error is calculated on pixels normalized to $$ and averaged across all images in the dataset. This process is performed both with and without the text as conditioning (C vs. UC). For COCO, we use the first listed prompt as conditioning. The LDM autoencoder reconstruction error serves as a lower bound. EDICT maintains complete latent recovery in all examples for both 50 and 200 steps, with error 50-75% that of DDIM (UC) (Tab. 1). DDIM (C) is unstable for inversions (also noted in ), which results in an error an order of magnitude greater than any other, with failure to reconstruct as shown in Fig. 2.

2 Image Editing

We show EDICT’s ability to perform complex editing tasks on real images in Fig. 6. In the first row, a diverse set of objects are added to a lake scene, demonstrating object addition. In the giraffe and car examples, we see interaction between introduced and original objects. Additionally, both the giraffe and castle examples capture reflections of the added objects, while the pattern of the water is maintained. Throughout all edits, details such as the cloud patterns and patches of tree color are preserved. The second row shows object-preserving global changes. A chair is placed into a variety of settings while keeping near-perfect identity and detail, even when occlusions are generated (grass and snow). In all examples the chair keeps a realistic footing, despite ground changes.

In the third row, we demonstrate that EDICT is able to perform object deformations, a challenging class of edits for DDM methods, as many broad compositional components are determined very early in the generation . The previously most successful method for these types of edits, Imagic requires model finetuning. EDICT makes the sculpture assume a broad set of poses, spatially changing the semantic map in a refined way. Novel views, such as “A statue from behind” are able to be plausibly rendered Across these large-scale edits, fine-grained details such as the face and dress of the statue –as well as the foliage and path– are preserved. In the fourth row, we show that EDICT is able to perform global style changes while maintaining layout and details where appropriate. The layout is nearly identical across images (note the preserved cloud pattern), but can be changed when needed (e.g., the lack of trees in the Mars example). Specific art-styles are capable of being generated, including challenging concepts such as cubism.

In Fig. 1, EDICT makes semantic entity edits to and from a variety of dog breeds. It proves adept at holding the original subject pose, including in the third row where the dog is viewed in a very atypical position. The realism in the Chihuahua examples is particularly interesting, due to the relatively small size of the breed. We again highlight the preservation of details such as the background foliage or ground in the upper three rows. The fourth row is of specific interest as small text is preserved across examples (a typical failure case for DDIM).

Baseline Comparison: In Fig. 7, we demonstrate EDICT’s superior performance to other DDM sampling-based methods for image editing: conditional and unconditional DDIM inversion, prompt-to-prompt image editing, and SDEdit (as a stochastic baseline). All methods are run with 50 steps (we do not observe improvement of baselines when the number of steps are increased, see Supplementary). Since methods that require model finetuning or prompt tuning can be combined with EDICT, we view them as complementary to, rather than competitive with, EDICT, and do not perform a direct comparison with them. We quantitatively compare visual metrics of edits in Fig. 8.

Discussion

Limitations and Future Work: EDICT is deterministic unlike methods such as SDEdit and only outputs one generation per image-prompt pair. Also, the computational time is approximately twice that of a baseline DDIM process. As with all editing methods, performance can vary across inputs in a hard-to-predict manner and sometimes requires careful prompt selection. For future work, we note that the induced latent space constructed by the inversion process admits operations such latent interpolation which has not been widely applied to real images. As in prompt tuning, formalizing the process of prompt selection could further improve EDICT. Adding a controllable degree of randomness to EDICT could yield multiple candidate generations that still satisfy the desired properties.

Ethics: Like other image generation and editing models, EDICT will produce images that may reflect the socioeconomic biases of the training data or images that could be considered inappropriate. Image editing methods can also utilized for malicious purposes, including harassment and misinformation spread. Practitioners utilizing EDICT or its methodologies in a production setting should consider these limitations. An in-depth discussion on the ethics of image generation can be found in Imagen .

References

Appendix A Quantitative Experiment Details

We sample from 5 ImageNet classes (African Elephant, Ram, Egyptian Cat, Brown Bear, and Norfolk Terrier, validation set). Four experiments are performed, one swapping the pictured animal’s species to each of the other classes (20 species editing pairs in total), two contextual changes (A [animal] in the snow and A [animal] in a parking lot), and one stylistic (An impressionistic painting of a [animal]). The prompt for the species edit is simply A [animal]. Throughout, base prompts are of form A [animal]. Edits are performed with inversion strength s=0.8s=0.8 and steps S=50S=50.

For computing CLIP score in the species example, the 5 text queries are of identical form A [animal]. CLIP text queries for other edits are as follows:

We plot the mean and median metrics for baselines on each individual benchmark experiment as well as the mean-average and median-average across experiments in Fig. S 9.

Appendix B Misalignment of Pseudo-Gradient

In Section 3.2 of the main paper we claim that the pseudo-gradient of classifier-free guidance G⋅(Θ(xt,t,C)−Θ(xt,t,∅))G\cdot(\Theta(x_{t},t,C)-\Theta(x_{t},t,\emptyset)) is inconsistent across time steps which drives the instability of vanilla DDIM inversion and reconstruction results. We demonstrate and analyze this instability in Fig. S 10 (see caption). We show that similar behavior holds for higher steps (Fig. S 11).

Appendix C Edits

In Fig. S 12 and Fig. S 13 we display further editing results.

C.2 Baselines with More Steps

In Fig. S 14 and Fig. S 15 we re-run the experiments of Figure 7 from the main paper with 100 and 250 global steps instead of the default 50. We observe minimal changes besides some instability in the final row. Note that we follow a scaling rule of p=0.9350/Sp=0.93^{50/S} to maintain the same aggregate dilation/contraction factor of 0.93500.93^{50} from the original experiments.

C.3 Dog Breeds: Extended Results

In Fig. S 16–Fig. S 22 we display additional results of dog breed editing with baselines included. EDICT xx vs yy are the two sequence outputs of the EDICT process to demonstrate the visually-identical convergence. We observe that EDICT consistently matches the desired output while preserving background details that baseline methods erase or alter. The base prompt is A dog and the target prompt is A [target dog breed].

Appendix D Reconstruction Results

In an extension of Table 1 from the main paper, we provide higher precision MSEs as well as reconstruction errors for 1000 steps in Tab. 2.