Inversion by Direct Iteration: An Alternative to Denoising Diffusion for Image Restoration
Mauricio Delbracio, Peyman Milanfar
Introduction
Recovering a high-quality image from a low-quality observation is a fundamental problem in computer vision and computational imaging. Single image restoration is a highly ill-posed inverse problem where multiple plausible sharp and clean images could lead to the very same degraded observation. The typical supervised approach is to formulate image restoration as a problem of inferring the underlying image given a low-quality version of it, by training a model with paired examples of the relevant degradation (Ongie et al., 2020). One of the most common approaches is to directly minimize a pixel reconstruction error using the or loss; an approach that correlates well with the popular PSNR (peak signal-to-noise-ratio) metric. However, it has been observed often in recent literature that measures such as PSNR (and in general point-distortion metrics) do not correlate well to human perception (Blau & Michaeli, 2018; Delbracio et al., 2021b; Freirich et al., 2021). Despite these shortcomings, much of the recent research work has been focused on improving deep architectures and optimizing a variety of point-loss formulations, resulting in general models that give an aggregate improved image in one step of inference.
To see the issues more concretely, let’s assume that we are given image pairs where represents a target high-quality image, and represents the respective degraded observation. For instance, may be pristine images degraded by a combination of blur/compression and noise to yield . A typical regression approach would predict directly from using a trained model , by minimizing the expected pixel error in some (e.g. ) metric as follows:
This evidently results in an image that is the (weighted) average of all plausible reconstructionsA similar statement is true for other in which case the mean is replaced by another aggregation operator (e.g. median for ). This resulting image will not have a natural appearance as the details have been wiped out due to the effect aggregation (i.e. “regression to the mean”) effect. The problem is compounded for more ill-posed problems. That is, the more ill-posed the inverse problem, the larger the set of plausible reconstructions and therefore the more severe the effect of the aggregation implied by the expectation of the posterior. To mitigate this problem, recent works have introduced additional loss terms (Gatys et al., 2016; Mechrez et al., 2018b; a; Delbracio et al., 2021b; Kupyn et al., 2018) that seek a balance in the formulation so the final image has improved perceptual quality (more on this in the next section).
In this work, we explicitly address this problem by avoiding single-step prediction of the clean image, and instead iterating a series of inferences, where at each step we solve an ‘easier’ (i.e., less ill-posed) inverse problem than the original. Specifically, we generate a sequence of intermediate restorations where at each step the goal is to reconstruct only a slightly less corrupted image. The core observation underlying this approach is this: a small-step restoration largely avoids the regression-to-the-mean effect because the set of plausible ‘slightly-less-bad’ images is relatively small. The core technical component enabling our approach is still a single deep model, but one that is trained to predict a better image given one with an intermediate level of degradation in the previous step, as summarized by Algorithm 1.
Background
Recently, much work on imaging inverse problems has been focused on using generative formulations (Bora et al., 2017; Kawar et al., 2022). Generative adversarial formulations train restoration networks with an adversarial loss that forces the restored image to be on the distribution of high-quality signals (Kupyn et al., 2018; 2019; Asim et al., 2020). GANs are in general hard to train and also hard to control image hallucinations since the two terms play an antagonic role (Lugmayr et al., 2021).
Image priors, including generative ones, can be used to solve inverse problems in an unsupervised fashion where the degradation operator is only know at inference time (Rudin & Osher, 1994; Venkatakrishnan et al., 2013; Delbracio et al., 2021a; Romano et al., 2017; Ongie et al., 2020). DDPMs have been recently adapted for unsupervised model-based image restoration (Kawar et al., 2021a; 2022; Kadkhodaie & Simoncelli, 2021; Jalal et al., 2021a; Laumont et al., 2022; Chung et al., 2022; Kawar et al., 2022).
How this approach compares to Denoising Diffusion. Denoising Diffusion Probabilistic Models (DDPMs) (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2021a) and Score-based models (Song & Ermon, 2019; 2020; Song et al., 2021b) have emerged as two powerful classes of generative models that produce high-quality samples by inverting a known diffusion (degradation) process. The standard Gaussian denoising formulation has been extended to more general corruption process (Bansal et al., 2022; Hoogeboom & Salimans, 2022; Daras et al., 2023; Deasy et al., 2021; Hoogeboom et al., 2022a; b; Nachmani et al., 2021; Johnson et al., 2021; Lee et al., 2022; Ye et al., 2022). The main common idea is to analytically define a known degradation process that is reversed to generate new samples starting from a fully degraded image (e.g., pure noise). The inference procedure makes use of the known analytical degradation at every step.
By contrast, in our formulation we do not require knowledge of any analytic form of the degradation process, we directly learn an iterative restoration process from low-quality/high-quality paired examples. This implies that we can apply our iterative procedure to virtually any degradation as long as we are given image pairs. Additionally, our formulation is motivated only from the idea of splitting the original inverse problem into multiple smaller ones. We do not require any knowledge of the underlying probability distributions, or the conditional distributions, at any step. Our inference procedure is solely based on the idea of restoring the signal a little bit at each step. This approach, with minimal assumptions, gives a unified formulation for any supervised image restoration problem under the same framework.
The natural extension of DDPMs to image restoration tasks is through the use of a Conditional DDPM models (cDDPM) (Li et al., 2021; Saharia et al., 2021; 2022; Whang et al., 2022). The goal of a cDDPM is to generate plausible reconstructions given the low-quality input (e.g., by generating samples from the posterior distribution). The idea is to train a supervised denoising diffusion model using paired examples that is conditioned on the low-quality input. The denoising network learns to generate a valid restored image (sample) by repeatedly denoising an initial image of pure noise. Our formulation has some similarity to conditional diffusion models, but contrary to the denoising formulation, we directly proceed by iteratively restoring the input image.
Overall, our method is straightforward to implement and train, and produces high-quality results. We evaluate the formulation on four different restoration tasks using different perceptual quality metrics. As shown, our method produces samples of higher quality than the state-of-the-art regression formulations while maintaining high-fidelity with respect to the original sample.
Related Work
The goal of image restoration is to generate a high-quality image from its degraded low-quality measurement (e.g., low-resolution, compressed, noisy, blurry). Since the seminal super-resolution work of Dong et al. (2015) many recent image restoration methods adopt an end-to-end supervised formulation where a deep neural network is trained to directly produce a point estimate (Zhao et al., 2016; Lim et al., 2017; Tao et al., 2018; Chen et al., 2018) These methods rely on low-quality high-quality image pairs to train a regression model. Most of the work has been focused on developing better and more powerful network architectures (Zamir et al., 2022; Chen et al., 2022; Tu et al., 2022; Zamir et al., 2021) so we can achieve better pixel-level reconstruction. While this formulation leads to state-of-the-art PSNR, the image generated is at best an average of all plausible solutions (regression to the mean). In the limit case where the low-quality image is completely obfuscated the best prediction in terms of PSNR is the average of the distribution.
Generative adversarial networks (Goodfellow et al., 2014; Arjovsky et al., 2017), and adversarial formulations (Ledig et al., 2017; Isola et al., 2017; Kupyn et al., 2018; 2019) have been introduced to push the generated image towards the manifold of natural images. GANs suffer from unstable training (Arora et al., 2017; Salimans et al., 2016; Arjovsky et al., 2017), while being prone to introduce significant image hallucinations. This is a direct consequence of a non-reference formulation that directly tries to minimize the distance of the range of the generator to the manifold of natural images (Cohen et al., 2018).
Blau & Michaeli (2018) proved that there is a trade-off between image perceptual quality and distortion. It is not possible to minimize both distortion and perceptual quality simultaneously. In fact, minimizing the average point distortion (e.g., PSNR) can be only done in detriment of the perceptual quality (Blau & Michaeli, 2018; Freirich et al., 2021).
A powerful way of avoiding the regression to the mean is to formulate the problem as one of sampling from the posterior distribution (Kawar et al., 2021b; a; 2022; Ohayon et al., 2021; Kadkhodaie & Simoncelli, 2021; Whang et al., 2022). An additional benefit of this formulation is to be able to generate multiple different plausible solutions that can be used for uncertainty quantification (Whang et al., 2021) or improving fairness (Jalal et al., 2021b).
Variational auto-encoders (Prakash et al., 2020), Normalizing flows (Lugmayr et al., 2020; 2021), and Diffusion probabilistic models (DPMs) (Saharia et al., 2021; Li et al., 2021; Whang et al., 2022) have been successfully applied to different image restoration tasks, where a diverse set of candidates can be generated from the learned posterior (Prakash et al., 2020).
Denoising Diffusion Probabilistic Models (DDPMs) (Sohl-Dickstein et al., 2015; Ho et al., 2020; Song et al., 2021a), Score-based models (Song & Ermon, 2019; 2020; Song et al., 2021b) and their recent generalizations (Bansal et al., 2022; Hoogeboom & Salimans, 2022; Daras et al., 2023; Deasy et al., 2021; Hoogeboom et al., 2022a; b; Nachmani et al., 2021; Johnson et al., 2021; Lee et al., 2022; Ye et al., 2022) generate high-quality samples by inverting a known degradation process. The main strategy is to analytically define a known degradation process that is reversed to generate new samples starting from a fully degraded image (e.g., pure noise). Bansal et al. (2022) introduced Cold Diffusion a generative framework to generate images by reverting arbitrary (known) degradations. They show promising results even with non-stochastic degradations such as blur, masking or pixelization. The strategy is to define intermediate analytical degradations (diffusion) and then revert them step by step. Daras et al. (2023) presented Soft Diffusion a generalization of diffusion models to linear degradations. The authors argue that noise is a fundamental component that is needed to provably learn the score.
In this work, we adopt a similar strategy and propose to decompose the image restoration problem into a sequence of intermediate steps each of them being a much easier problem to solve (i.e., less ill-posed). This path of intermediate reconstructions takes us from a low-quality input to a high-quality reconstruction through a series of slightly less corrupted signals. Different than in traditional generative diffusion formulations, the degradation is only implicitly given through a series of pair images (low-quality and high-quality). In our formulation, we define the intermediate steps as a convex combination of the target/input signals. This induces a simple linear propagation from the high-quality sample to the low-quality one.
An alternative formulation of the supervised image restoration problem is to use a conditional denoising diffusion models to generate samples from the posterior distribution (Li et al., 2021; Saharia et al., 2021; 2022; Whang et al., 2022). The overall idea is to train a denoising diffusion model that is conditioned on the low-quality input. The denoising network learns to generate a valid restored image by repeatedly denoising an initial image of pure noise. Our formulation has some similarity to conditional diffusion models, but contrary to the denoising formulation, we directly proceed by iteratively restoring the input image.
The very recent work by Luo et al. (2023a) and Welker et al. (2022), and the concurrent work by Luo et al. (2023b); Song et al. (2023) introduce related image restoration techniques based on ODE/SDE diffusion formulations. A major difference with our work is that InDI is completely motivated and formulated from elementary principles by splitting the restoration tasks into multiple smaller ones. There is also significant amount of recent related work that analyzes the connection of diffusion image generation with Bridge and Flow matching, Optimal Transport, and Schrodinger bridges (Albergo & Vanden-Eijnden, 2023; Albergo et al., 2023; Lipman et al., 2023; Liu et al., 2023a; b; Shi et al., 2023). Finally, the concurrent work of Heitz et al. (2023) adopts a linear diffusion scheme for image generation similar to, but less general than, the one in InDI.
InDI: Our Proposed Formulation
Given , we define a continuous forward degradation process by
The idea of this forward process is that it starts from a clean sharp image at time , and then degrades it to the blurry/noisy observation at time . Here, indexed by , represents an intermediate degraded image between the low-quality input (i.e., ) and the high-quality sharp target (i.e., ). Following the common notation in diffusion models, we will refer to the index as the time-step.
Let be given from equation 1, where . Then,
The proof is a direct consequence of change of variables and is given in Appendix A.
According to this proposition, the posterior mean (e.g. MMSE estimate) at time can be deduced from the estimate at time by first estimating the clean image (), and then doing a convex combination with the estimate at time . We can then apply the following scheme to move from to ,
The process starts from , and the step controls the “speed” of the reverse process (e.g., at constant “speed”, , where controls the total number of steps).
An intuitive motivation for the requirement in Remark 4.2 is that we need to move through a path of plausible samples at every step . One simple way to guarantee this is by adding a small amount of noise to . Then , and therefore will be non-zero everywhere. More discussion about this is presented at the end of this section, but first, we present a toy example to motivate our approach.
Let be the intermediate degraded samples according to equation 1. In this simple example, there is a closed form expression for all posterior distributions and conditional means; namely,
where and is a Gaussian kernel with identity covariance and . Then,
Figure 1 shows the results of applying the iterative regression scheme given by the above equation in two different examples. The iterative regression converges to one of the four possible modes (shown in orange), while the regression to the mean is always a weighted average of all possible modes (i.e., a blurry reconstruction, shown in red).
where is a predefined distribution for (e.g., uniform). The model allows us to do incremental reconstruction where from time step we predict the slightly less corrupted signal at time as given in equation 2. Thus, the iterative scheme becomes:
where . Although could be a function of time, in practice we use a constant time step, , where is the number of steps.
In the limit, as , equation 5 leads to an ordinary differential equation (ODE); namely
Another use of the continuous formulation is to understand the behavior of the proposed iterative procedure in terms of concrete examples. In Appendix B we show how the residual flow can be used to analyze the specific case where the prior is Gaussian and the restoration task is denoising.
An interesting connection emerges when the degradation is Gaussian noise (standard deviation ). In this case, InDI’s ODE in equation 6 boils down to the score-matching probabilistic ODE of Song et al. (2021b). More specifically, let , so the noise level in is . The probabilistic flow ODE (Eq(13) in Song et al. (2021b)) is given by
According to the denoising score-matching (DSM) approximation (Vincent (2011)),
Stochastic Perturbation: To make sure we have the regularity requirements for the iterative procedure from equation 2 to be well defined (Remark 4.2), we add a small amount of white noise to the low-quality input. As shown in Section 5 this leads to a significant improvement in image quality in certain tasks (in particular those that are restorations from deterministic degradations).
Our model with this noise perturbation becomes:
where , is a small constant (e.g., , where image values are in ${\bm{n}}\sim\mathcal{N}(0,Id)$.
A slightly more general formulation incorporates the perturbation as a general Brownian motion, where we can explicitly control the level of noise at each step. That is,
where is a non-negative function, and is the standard Brownian motion having zero mean and covariance at index .
In this more general setting, the base training objective becomes
And as a result, the general inference procedure in equation 2 becomes:
where the reconstruction process starts from , and . At each step, a new is sampled and noise is added to the current state. The added Gaussian noise is such that the noise at time has variance as required by equation 8. To be well defined needs to be a non-negative non-increasing function of . In the limit case where we are in the simplified case given by equation 7, while if the noise perturbation is a pure Brownian motion.
Our full iterative restoration inference scheme is given in Algorithm 1.
Experiments
We train and evaluate our framework on four widely popular image restoration tasks: motion deblurring, defocus deblurring, compression artifacts removal and single image super-resolution. Our formulation is generative-based, and we show that can be used for image generation even if this is not the main focus of the present work. To evaluate the quality of the proposed method we compute several distortion based and perceptual metrics: PSNR, LPIPS (Zhang et al., 2018), FID (Fréchet Inception Distance) (Heusel et al., 2017), and KID (Kernel Inception Distance) (Bińkowski et al., 2018).
Perception–Distortion tradeoff (Blau & Michaeli, 2018): To illustrate the potential of the method, we present results when using different number of steps for the reconstruction. This has a direct impact on the perception–distortion tradeoff. In general, a single step reconstruction with our model, will lead to an estimate that minimizes the average point distortion (e.g., PSNR) but this can be only done to the detriment of the perceptual quality.
Model Architecture and Training: We adopt a U-Net-like architecture similar to the ones in diffusion strategies (Saharia et al., 2021; Whang et al., 2022). Following Whang et al. (2022) we removed attention layers and group normalization to have a fully-convolutional architecture. The size of the model varies for each evaluated task (in general we chose a model size proportional to the size of the dataset to avoid significant overfitting). The model is trained on image crops using ADAM optimizer. Learning rates and other hyper-parameters are given in Appendix C. For each experiment, we train a single model that is conditioned on the parameter . The model is trained using the loss function of equation 9 with . We found that the distribution of , plays an important role. In Section 6.3 we present an empirical analysis of its impact.
Motion deblurring is a very challenging restoration task. Motion is intrinsically random in the sense that a priori we don’t have a known degradation model. The current best end-to-end deep learning solution is to train regression models using paired data sharp, blurry frames. One of the most adopted training datasets is the GoPro motion deblurring dataset (Nah et al., 2017) containing 3214 pairs of clean and blurry images ( are reserved for evaluation). The blurry frames are generated by recording high-frame rate video clips and then averaging consecutive frames to simulate blurs caused due to longer exposure. We follow the standard setup (Nah et al., 2017; Kupyn et al., 2019; Chen et al., 2021a; Cho et al., 2021; Suin et al., 2020; Zhang et al., 2019) and perform training data augmentation with random horizontal/vertical flips and 90/180/270 rotations. We did not introduce additional noise to the blurry inputs ( in equation 7).
Figure 2 shows a visual comparison of our iterative image restoration and current state-of-the-art deblurring models. The iterative scheme produces images with much more details than regression based solutions (Restormer (Zamir et al., 2022), MAXIM (Tu et al., 2022)). Our results are similar to the ones generated by current conditional diffusion models (DvSR (Whang et al., 2022)). Quantitative results on the GoPro dataset are presented in Table 1. The proposed iterative reconstruction procedure achieves a new state-of-the-art performance across perceptual metrics while maintaining competitive PSNR to existing methods.
Number of steps. Figure 3 shows the impact of the number of inference steps on the Perception–Distortion trade-off (Blau & Michaeli, 2018). While doing a reconstruction on a single step (e.g., direct regression) produces the best PSNR, the perceptual metrics are significantly improved when the number of steps is larger than one. Both metrics can’t be optimized simultaneously (Blau & Michaeli, 2018).
2 Single-Image Super-resolution
We evaluated the iterative restoration methodology on single-image super-resolution on the div2k dataset (Agustsson & Timofte, 2017). This dataset contains 1000 2K-resolution images (800 for training, 100 images for validation, 100 testing). We compare to other state-of-the art models that span from regression models having powerful architectures (Wang et al., 2018; Chen et al., 2021b; Liang et al., 2022) and/or generative formulations: GAN based, i.e., LDL (Liang et al., 2022), ESRGAN (Wang et al., 2018), BSRGAN (Zhang et al., 2021); and also based on Normalizing Flows, SRFLOW (Lugmayr et al., 2020).
Figure 5(b) summarizes the quantitative results on SR div2k validation dataset. Figure 4 shows a selection of results. Our proposed framework leads to upscaled images with more defined structure than regression based formulations producing larger PSNR, e.g., RRDB (Wang et al., 2018). The recently introduced adversarial formulation LDL (Liang et al., 2022) produces slightly better fine grain details. This could indicate that in the situation where there is limited training data, careful adversarial formulation may be more data efficient.
The importance of adding noise in deterministic super-resolution. In our formulation of super-resolution, the degradation is a deterministic linear (blurring plus subsampling) operator. Figure 5(a) shows the importance of adding a small amount of noise to the input image. Directly applying the original iterative procedure (without adding noise to the input) leads to a blurry reconstruction (high PSNR but low FID score, in Figure 5 (a)). Adding a small amount of noise ( in Figure 5(a)) leads to significant better results in terms of perceptual quality (e.g., FID score).
3 Defocus deblurring
Defocus deblurring is the task of reducing the blur due to limited depth-of-field or misfocus. For such purposes we used the Canon dual-pixel (DP) defocus dataset (DDPD) provided by Abuolaim & Brown (2020), and train a defocus deblurring model only using single image input (i.e., we don’t use the dual-pixel images given in the dataset). The DDPD dataset contains 1000 pairs of sharp and blurry images, of which 30% are reserved for validation and testing. The blurry/sharp frames are generated by capturing two consecutive snapshots by changing the camera parameters (lens aperture). high-frame rate video clips and then averaging consecutive frames to simulate blurs caused due to longer exposure.
Figure 6 shows a visual comparison of our iterative image restoration when a different number of inference steps is used. Increasing the number of steps has a direct impact on the quality of the result. Quantitative results on the DDPD dataset are summarized in Table 3 in Appendix. As in the other experiments the best PSNR is obtained with a single step (direct regression), while the best perceptual metrics are obtained when the restoration is done in multiple steps. We did not introduce additional noise to the blurry inputs ( in equation 7).
4 Compression artifact removal
JPEG compression introduces blocking artifacts and lack of high-frequency details. We evaluated the proposed method on the task of removing strong JPEG compression artifacts (quality factor 15). To generate the training data we use the 1000 div2k high-quality images (Agustsson & Timofte, 2017). We evaluated the model on div2k validation set.
Figure 7 show some visual results of restored images with the model applying a different number of steps. As more inference steps are used the restored images have more details. More results on JPEG compression removal are discussed in the next section.
Discussion
A natural question to ask is whether the proposed approach is also generative in the spirit of diffusion formulations. Namely, if we take our formulation to the limit where the low-quality image is fully degraded, then could we potentially generate new samples from scratch?
To test the idea we trained a restoration model that starts from pure Gaussian noise paired to a celebA image and then proceed as described above. Figure 8 shows some generated samples with this formulation. The generated samples have a FID=9.19, which is not state-of-the-artIt is competitive with other methods from a couple years ago but illustrate the point. In this specific case, our proposed methodology leads to a similar denoising training loss as the one in DDPM (Ho et al., 2020). Despite this similarity, the two methods come from different motivations/formulations and therefore have different inference strategies. We didn’t fine-tune architecture or hyper-parameters to boost the performance since our goal was to present the idea and show that the formulation, at its core, can be generative as well.
2 Comparison of Inference Algorithms
In what follows we discuss different alternatives for recovering the clean sample with the trained models.
Naive Procedure and Relevance to Cold Diffusion: Given equation 1, one may be tempted to directly replace the clean image by the current estimate. This would lead to , and the inference iterative rule would become
Cold Diffusion (Bansal et al., 2022) proposes to generate images by inverting an arbitrary known degradation , where controls the strength. Our formulation is more general in the sense that we don’t require an explicit knowledge of . To apply Cold Diffusion sampling in our context, we define (given by equation 1). This leads to Cold Diffusion’s naive sampling (Algorithm 1 in Bansal et al. (2022)),
Note that this sampling scheme is the same as the one in Eq. 11. Cold diffusion improved sampling (Algorithm 2 in Bansal et al. (2022)) is given by,
In Figure 9 we compare our inference algorithm (equation 5), the naive inference algorithm (equation 11), and our adaptation of the Cold Diffusion sampler to our formulation (equation 15). In general, the naive sampler produces good results with very few steps (N=2,3) but then diverges. Our adaptation of Cold Diffusion sampler produces competitive results, while leading to slightly worse FID scores for the same distortion level than our proposed algorithm. In the limit, as the number of steps becomes very large, Cold Diffusion sampler seems to converge to a stable point, while ours after a certain large number of steps, deteriorates.
Figure 9. shows the FID score of CelebA 64x64 generated images when using the three different variants of the inference algorithm (sampler).
3 Impact of distribution p(t)
The impact of the distribution of used during training has a clear impact on performance. We evaluated several different options that are summarized in Figure 10 (a). Figure 10 (b) shows the results when different distributions are adopted. The best results are obtained when the model is trained with a bias towards (more degradation). Intuitively, this could imply that the iterative procedure needs to be more certain of the direction to move at the very early steps of the procedure. Nonetheless, the best distribution can depend on a combination of model capacity and restoration task so we are not drawing general conclusions.
4 Impact of adding noise on inverting deterministic degradations
JPEG compression is a non-linear, but deterministic, degradation. We empirically verified that adding a small amount of noise helps to improve the results as shown in Figure 11. We tested the variant of the inference algorithm that adds noise at each step (so the noise becomes a Brownian motion, e.g., ), and adding a constant noise level at the initial step (). We did not observe any practical difference in the two approaches.
5 Comparison to a Conditional Denoising Diffusion Model
We compare InDI to a vanilla conditional DDPM (Ho et al., 2020). We trained a vanilla conditional DDPM, using the (continuous) noise level as an additional input, similarly as done in Saharia et al. (2021); Whang et al. (2022). The model architecture is the same as in InDI but the auxiliary noise image (needed in any DDPM) is concatenated with the low-quality input at each step. Figure 12 shows a comparison between InDI and the conditional DDPM. To generate the DDPM plot we merged several possible noise schedules using different number of steps that span the perception-distortion tradeoff. InDI produces comparable results using significant less number of steps than the vanilla DDPM.
Conclusions and Limitations
We presented a novel formulation of image restoration that circumvents the regression-to-the mean problem. This allows us to get restored images with superior realism and perceptual quality, while still having a low distortion error. Our method is motivated by the observation that restoration from a small distortion is a better-conditioned problem. We therefore break a restoration task into many small ones – each of them easier (and less ill-posed) than the larger problem we solve overall. This enables our iterative approach to transforming the degraded image into a high-quality image, in spirit similar to current generative diffusion models.
Limitations. The present formulation is a supervised one, requiring paired training data. As such, for each type of degradation we need to train a specialized model, in contrast to unsupervised formulations such as RED (Romano et al., 2017), PnP (Venkatakrishnan et al., 2013; Kamilov et al., 2022), or DDRM (Kawar et al., 2022). Additionally, given the dependence to paired training data, its performance for out-of-distribution samples is not guaranteed. This question requires more in-depth analysis. Finally, while the proposed iterative inference algorithm produces high-quality restorations, in some tasks performance degrades after a certain number of steps. This is likely due to the accumulation of errors, and will likely require a more robust inference scheme.
For future work, we would like to better characterize the limiting points of the proposed inference procedure. Other possible research avenues are developing robust formulations that can successfully handle out-of-domain input.
The authors would like to thank our colleagues Jon Barron, Tim Salimans, Jascha Sohl-dickstein, Ben Poole, José Lezama, Sergey Ioffe, and Jason Baldridge for helpful discussions.
References
Appendix A Proof of Proposition 4.1
We have , and , so by substituting from one to the other we get,
where we have applied the fact that and . ∎
Appendix B Denoising with a Gaussian Prior
We are interested in solving this equation at , with boundary condition at . This is a separable ODE having general solution: , where . The solution at , is
It is worth noting that in this case the MMSE and MAP estimates coincide,
and are in fact different from InDI’s estimate.
Appendix C Model and Training Details
In all our restoration experiments we use a U-Net-like architecture (Ronneberger et al., 2015) similar to the one in SR3 (Saharia et al., 2021) and DvSR (Whang et al., 2022). We followed the same adaptations as the ones introduced in Whang et al. (2022) to make it fully-convolutional (removed self-attention layers and group normalization). Our U-Net has an adaptive number of resolutions each of them having an arbitrary number of channels (given by a multiplication factor from a base set of channels).
Table 2 summarizes the model definition for each of the tested applications.
All models are trained for 500K steps using 32 TPUv3 cores. We used the Adam optimizer with a fixed learning rate, and EMA decay rate of 0.9999. Models were trained using the respective indicated distribution for . For the super-resolution model, low-resolution crops of size are upscaled using bilinear interpolation to before feeding them into the model.