Deep Image Prior

Dmitry Ulyanov, Andrea Vedaldi, Victor Lempitsky

Introduction

State-of-the-art approaches to image reconstruction problems such as denoising burger2012image ; lefkimmiatis2016non and single-image super-resolution Ledig17sr ; Tai17sr ; Lai17sr are currently based on deep convolutional neural networks (ConvNets). ConvNets also work well in “exotic” inverse problems such as reconstructing an image from its activations within a deep network or from its HOG descriptor dosovitskiy16inverting . Popular approaches for image generation such as generative adversarial networks goodfellow2014generative , variational autoencoders kingma2013auto and direct pixel-wise error minimization dosovitskiy2015learning ; Bojanowski17 also use ConvNets.

ConvNets are generally trained on large datasets of images, so one might assume that their excellent performance is due to the fact that they learn realistic data priors from examples, but this explanation is insufficient. For instance, the authors of zhang16understanding recently showed that the same image classification network that generalizes well when trained on a large image dataset can also overfit the same images when labels are randomized. Hence, it seems that obtaining a good performance also requires the structure of the network to “resonate” with the structure of the data. However, the nature of this interaction remains unclear, particularly in the context of image generation.

In this work, we show that, in fact, not all image priors must be learned from data; instead, a great deal of image statistics are captured by the structure of generator ConvNets, independent of learning. This is especially true for the statistics required to solve certain image restoration problems, where the image prior must supplement the information lost in the degradation processes.

To show this, we apply untrained ConvNets to the solution of such problems (fig. 1). Instead of following the standard paradigm of training a ConvNet on a large dataset of example images, we fit a generator network to a single degraded image. In this scheme, the network weights serve as a parametrization of the restored image. The weights are randomly initialized and fitted to a specific degraded image under a task-dependent observation model. In this manner, the only information used to perform reconstruction is contained in the single degraded input image and the handcrafted structure of the network used for reconstruction.

We show that this very simple formulation is very competitive for standard image processing problems such as denoising, inpainting, super-resolution, and detail enhancement. This is particularly remarkable because no aspect of the network is learned from data and illustrates the power of the image prior implicitly captured by the network structure. To the best of our knowledge, this is the first study that directly investigates the prior captured by deep convolutional generative networks independently of learning the network parameters from images.

In addition to standard image restoration tasks, we show an application of our technique to understanding the information contained within the activations of deep neural networks trained for classification. For this, we consider the “natural pre-image” technique of mahendran15understanding , whose goal is to characterize the invariants learned by a deep network by inverting it on the set of natural images. We show that an untrained deep convolutional generator can be used to replace the surrogate natural prior used in mahendran15understanding (the TV norm) with dramatically improved results. Since the new regularizer, like the TV norm, is not learned from data but is entirely handcrafted, the resulting visualizations avoid potential biases arising form the use of learned regularizers dosovitskiy16inverting . Likewise, we show that the same regularizer works well for “activation maximization”, namely the problem of synthesizing images that highly activate a certain neuron erhan09visualizing .

Method

A deep generator network is a parametric function x=fθ(z)x=f_{\theta}(z) that maps a code vector zz to an image xx. Generators are often used to model a complex distribution p(x)p(x) over images as the transformation of simple distribution p(z)p(z) over the codes, such as a Gaussian distribution goodfellow2014generative .

One might think that knowledge about the distribution p(x)p(x) is encoded in the parameters θ\theta of the network, and is therefore learned from data by training the model. Instead, we show here that a significant amount of information about the image distribution is contained in the structure of the network even without performing any training of the model parameters.

Without training on a dataset, we cannot expect the a network fθf_{\theta} to know about specific concepts such as the appearance of certain objects classes. However, we demonstrate that the untrained network does capture some of the low-level statistics of natural images — in particular, the local and translation invariant nature of convolutions and the usage of a sequence of such operators captures the relationship of pixel neighborhood at multiple scales. This is sufficient for it to model conditional image distributions p(x∣x0)p(x|x_{0}) of the type that arise in image restoration problems, where xx has to be determined given a corrupted version x0x_{0} of itself. The latter can be used to solve inverse problems such as denoising burger2012image , super-resolution dong2014learning and inpainting.

Rather than working with distributions explicitly, we formulate such tasks as energy minimization problems of the type

where E(x;x0)E(x;x_{0}) is a task-dependent data term, x0x_{0} is the noisy/low-resolution/occluded image, and R(x)R(x) is a regularizer.

The choice of data term E(x;x0)E(x;x_{0}) is often directly dictated by the application and is thus not difficult. The regularizer R(x)R(x), on the other hand, is often not tied to a specific application because it captures the generic regularity of natural images. A simple example is Total Variation (TV), which encourages images to contain uniform regions, but much research has gone into designing and learning good regularizers.

In this work, we drop the explicit regularizer R(x)R(x) and use instead the implicit prior captured by the neural network parametrization, as follows:

The (local) minimizer θ∗\theta^{*} is obtained using an optimizer such as gradient descent, starting from a random initialization of the parameters θ\theta (see fig. 2). Hence, the only empirical information available to the restoration process is the noisy image x0x_{0}. Given the resulting (local) minimizer θ∗\theta^{*}, the result of the restoration process is obtained as x∗=fθ∗(z)x^{*}=f_{\theta^{*}}(z).Equation 2 can also be thought of as a regularizer R(x)R(x) in the style of eq. 1, where R(x)=0R(x)=0 for all images that can be generated by a deep ConvNet of a certain architecture with the weights being not too far from random initialization, and R(x)=+∞R(x)=+\infty for all other signals. This approach is schematically depicted in fig. 3 (left).

Since no aspect of the network fθf_{\theta} is learned from data beforehand, such deep image prior is effectively handcrafted, just like the TV norm. The contribution of the paper is to show that this hand-crafted prior works very well for various image restoration tasks, well beyond standard handcrafted priors, and approaching learning-based approaches in many cases.

As we show in the experiments, the choice of architecture does have an impact on the results. In particular, most of our experiments are performed using a U-Net-like “hourglass” architecture with skip connections, where zz and xx have the same spatial dimensions and the network has several millions of parameters. Furthermore, while it is also possible to optimize over the code zz, in our experiments we do not do so. Thus, unless noted otherwise, zz is a fixed randomly-initialized 3D3D tensor.

One may wonder why a high-capacity network fθf_{\theta} can be used as a prior at all. In fact, one may expect to be able to find parameters θ\theta recovering any possible image xx, including random noise, so that the network should not impose any restriction on the generated image. We now show that, while indeed almost any image can be fitted by the model, the choice of network architecture has a major effect on how the solution space is searched by methods such as gradient descent. In particular, we show that the network resists “bad” solutions and descends much more quickly towards naturally-looking images. The result is that minimizing (2) either results in a good-looking local optimum (fig. 3 — left), or, at least, that the optimization trajectory passes near one (fig. 3 — right).

In order to study this effect quantitatively, we consider the most basic reconstruction problem: given a target image x0x_{0}, we want to find the value of the parameters θ∗\theta^{*} that reproduce that image. This can be setup as the optimization of (2) using a data term such as the L2L^{2} distance that compares the generated image to x0x_{0}:

Plugging eq. 3 in eq. 2 leads us to the optimization problem

Figure 4 shows the value of the energy E(x;x0)E(x;x_{0}) as a function of the gradient descent iterations for four different choices for the image x0x_{0}: 1) a natural image, 2) the same image plus additive noise, 3) the same image after randomly permuting the pixels, and 4) white noise. It is apparent from the figure that the optimization is much faster for cases 1) and 2), whereas the parametrization presents significant “inertia” for cases 3) and 4). Thus, although in the limit the parametrization can fit noise as well, it does so very reluctantly. In other words, the parametrization offers high impedance to noise and low impedance to signal.

To use this fact in some of our applications, we restrict the number of iterations in the optimization process (2). The resulting prior then corresponds to projection onto a reduced set of images that can be produced from zz by ConvNets with parameters θ\theta that are not too far from the random initialization θ0\theta_{0}. The use of deep image prior with the restriction on the number of iterations in the optimization process is schematically depicted in fig. 3 (right).

2 “Sampling” from the deep image prior

The prior defined by eq. 2 is implicit and does not define a proper probability distribution in the image space. Nevertheless, it is possible to draw “samples” (in the loose sense) from this prior by taking random values of the parameters θ\theta and looking at the generated image fθ(z)f_{\theta}(z). In other words, we can visualize the starting points of the optimization process eq. 2 before fitting the parameters to the noisy image. Figure 5 shows such “samples” from the deep priors captured by different hourglass-type architectures. The samples exhibit spatial structures and self-similarities, whereas the scale of these structures depends on the depth of the network. Adding skip connections results in images that contain structures of different characteristic scales, as is desirable for modeling natural images. It is therefore natural that such architectures are the most popular choice for generative ConvNets. They have also performed best in our image restoration experiments described next.

Applications

We now show experimentally how the proposed prior works for diverse image reconstruction problems. More examples and interactive viewer can be found on the project webpage https://dmitryulyanov.github.io/deep_image_prior.

As our parametrization presents high impedance to image noise, it can be naturally used to filter out noise from an image. The aim of denoising is to recover a clean image xx from a noisy observation x0x_{0}. Sometimes the degradation model is known: x0=x+ϵx_{0}=x+\epsilon where ϵ\epsilon follows a particular distribution. However, more often in blind denoising the noise model is unknown.

Here we work under the blindness assumption, but the method can be easily modified to incorporate information about noise model. We use the same exact formulation as eqs. 3 and 4 and, given a noisy image x0x_{0}, recover a clean image x∗=fθ∗(z)x^{*}=f_{\theta^{*}}(z) after substituting the minimizer θ∗\theta^{*} of eq. 4.

Our approach does not require a model for the image degradation process that it needs to revert. This allows it to be applied in a “plug-and-play” fashion to image restoration tasks, where the degradation process is complex and/or unknown and where obtaining realistic data for supervised training is difficult. We demonstrate this capability by several qualitative examples in fig. 7, where our approach uses the quadratic energy (3) leading to formulation (4) to restore images degraded by complex and unknown compression artifacts. Figure 6 (top row) also demonstrates the applicability of the method beyond natural images (a cartoon in this case).

We evaluate our denoising approach on the standard datasethttp://www.cs.tut.fi/~foi/GCF-BM3D/index.html#ref_results, consisting of 9 colored images with noise strength of σ=25\sigma=25. We achieve a PSNR of 29.2229.22 after 1800 optimization steps. The score is improved up to 30.4330.43 if we additionally average the restored images obtained in the last iterations (using exponential sliding window). If averaged over two optimization runs our method further improves up to 31.0031.00 PSNR. For reference, the scores for the two popular approaches CMB3D dabov2007image and Non-local means buades2005non , that do not require pretraining, are 31.4231.42 and 30.2630.26 respectively.

To validate if the deep image prior is suitable for denoising images corrupted with real-world non-Gaussian noise we use the benchmark of plotz2017benchmarking . Using the same architecture and hyper-parameters as for fig. 6 we get 41.9541.95 PSNR, while CBM3D’s score is only 30.1330.13. We also use the deep image prior with different network architectures and get 35.0535.05 PSNR for UNet and 31.9531.95 for ResNet. The details of each architecture are described in section 4. Our hour-glass architecture resembles UNet, yet has less number of skip connections and additional BatchNorms before concatenation operators. We speculate that the overly wide skip-connections within UNet lead to a prior that are somewhat too weak and the fitting happens too fast; while the lack of skip-connections in ResNet leads to slow fitting and a prior that is too strong. Overall, this stark difference in the performance of different architectures emphasizes that different architectures impose rather different priors leading to very different results.

2 Super-resolution

Following eq. 2, we regularize the problem by considering the re-parametrization x=fθ(z)x=f_{\theta}(z) and optimizing the resulting energy w.r.t. θ\theta. Optimization still uses gradient descent, exploiting the fact that both the neural network and the most common downsampling operators, such as Lanczos, are differentiable.

We evaluate super-resolution ability of our approach using Set5 set5 and Set14 set14 datasets. We use a scaling factor of 44 and 88 to compare to other works in fig. 8.

Qualitative comparison with bicubic upsampling and state-of-the art learning-based methods SRResNet Ledig17sr , LapSRN Tai17sr is presented in fig. 8. Our method can be fairly compared to bicubic, as both methods never use other data than a given low-resolution image. Visually, we approach the quality of learning-based methods that use the MSE loss. GAN-based goodfellow2014generative methods SRGAN Ledig17sr and EnhanceNet Sajjadi17sr (not shown in the comparison) intelligently hallucinate fine details of the image, which is impossible with our method that uses absolutely no information about the world of HR images.

We compute PSNRs using center crops of the generated images (tables 2 and 1). While our method is still outperformed by learning-based approaches, it does considerably better than the non-trained ones (bicubic, glasner2009super , Huang15sr ). Visually, it seems to close most of the gap between non-trained methods and state-of-the-art trained ConvNets (c.f. figs. 1 and 8).

In fig. 9 we compare our deep prior to non-regularized solution and a vanilla TV prior. Our result do not have both ringing artifacts and cartoonish effect.

3 Inpainting

In image inpainting, one is given an image x0x_{0} with missing pixels in correspondence of a binary mask m∈{0,1}H×Wm\in\{0,1\}^{H\times W}; the goal is to reconstruct the missing data. The corresponding data term is given by

where ⊙\odot is Hadamard’s product. The necessity of a data prior is obvious as this energy is independent of the values of the missing pixels, which would therefore never change after initialization if the objective was optimized directly over pixel values xx. As before, the prior is introduced by optimizing the data term w.r.t. the re-parametrization (2).

In the first example (fig. 11) inpainting is used to remove text overlaid on an image. Our approach is compared to the method of RenXYS15 specifically designed for inpainting. Our approach leads to almost perfect results with virtually no artifacts, while for RenXYS15 the text mask remains visible in some regions.

Next, fig. 13 considers inpainting with masks randomly sampled according to a binary Bernoulli distribution. First, a mask is sampled to drop 50%50\% of pixels at random. We compare our approach to a method of PapyanRSE17 based on convolutional sparse coding. To obtain results for PapyanRSE17 we first decompose the corrupted image x0x_{0} into low and high frequency components similarly to GuZXMFZ15 and run their method on the high frequency part. For a fair comparison we use the version of their method, where a dictionary is built using the input image (shown to perform better in PapyanRSE17 ). The quantitative comparison on the standard data set heide2015fast for our method is given in fig. 13, showing a strong quantitative advantage of the proposed approach compared to convolutional sparse coding. In fig. 13 we present a representative qualitative visual comparison with PapyanRSE17 .

We also apply our method to inpainting of large holes. Being non-trainable, our method is not expected to work correctly for “highly-semantical” large-hole inpainting (e.g. face inpainting). Yet, it works surprisingly well for other situations. We compare to a learning-based method of IizukaSIGGRAPH2017 in fig. 10. The deep image prior utilizes context of the image and interpolates the unknown region with textures from the known part. Such behavior highlights the relation between the deep image prior and traditional self-similarity priors.

In fig. 14, we compare deep priors corresponding to several architectures. Our findings here (and in other similar comparisons) seem to suggest that having deeper architecture is beneficial, and that having skip-connections that work so well for recognition tasks (such as semantic segmentation) is highly detrimental for the deep image prior.

4 Natural pre-image

The natural pre-image method of mahendran15understanding is a diagnostic tool to study the invariances of a lossy function, such as a deep network, that operates on natural images. Let Φ\Phi be the first several layers of a neural network trained to perform, say, image classification. The pre-image is the set

of images that result in the same representation Φ(x0)\Phi(x_{0}). Looking at this set reveals which information is lost by the network, and which invariances are gained.

Finding pre-image points can be formulated as minimizing the data term

However, optimizing this function directly may find “artifacts”, i.e. non-natural images for which the behavior of the network Φ\Phi is in principle unspecified and that can thus drive it arbitrarily. More meaningful visualization can be obtained by restricting the pre-image to a set X\mathcal{X} of natural images, called a natural pre-image in mahendran15understanding .

In practice, finding points in the natural pre-image can be done by regularizing the data term similarly to the other inverse problems seen above. The authors of mahendran15understanding prefer to use the TV norm, which is a weak natural image prior, but is relatively unbiased. On the contrary, papers such as dosovitskiy16inverting learn to invert a neural network from examples, resulting in better looking reconstructions, which however may be biased towards the learned data-driven inversion prior. Here, we propose to use the deep image prior (2) instead. As this is handcrafted like the TV-norm, it is not biased towards a particular training set. On the other hand, it results in inversions at least as interpretable as the ones of dosovitskiy16inverting .

For evaluation, our method is compared to the ones of mahendran16visualizing and dosovitskiy16inverting . Figure 15 shows the results of inverting representations Φ\Phi obtained by considering progressively deeper subsets of AlexNet AlexNet : conv1, conv2, …, conv5, fc6, fc7, and fc8. Pre-images are found either by optimizing (2) using a structured prior.

As seen in fig. 15, our method results in dramatically improved image clarity compared to the simple TV-norm. The difference is particularly remarkable for deeper layers such as fc6 and fc7, where the TV norm still produces noisy images, whereas the structured regularizer produces images that are often still interpretable. Our approach also produces more informative inversions than a learned prior of dosovitskiy16inverting , which have a clear tendency to regress to the mean. Note that dosovitskiy16inverting has been followed-up by DosovitskiyB16 where they used a learnable discriminator and a perceptual loss to train the model. While the usage of a more complex loss clearly improved their results, we do not compare to their method here as our goal is to demonstrated what can be achieved with a prior not obtained from a training set.

We perform similar experiment and invert layers of VGG-19 Simonyan14c in fig. 16 and also observe an improvement.

5 Activation maximization

Along with the pre-image method, the activation maximization method is used to visualize internals of a deep neural network. It aims to synthesize an image that highly activates a certain neuron by solving the following optimization problem:

where mm is an index of a chosen neuron. Φ(x)m\Phi(x)_{m} corresponds to mm-th output if Φ\Phi ends with fully-connected layer and central pixel of the mm-th feature map if the Φ(x)\Phi(x) has spatial dimensions.

We compare the proposed deep prior to TV prior from mahendran15understanding in fig. 17, where we aim to maximize activations of the last fc8 layer of AlexNet and VGG-16. For AlexNet deep image prior leads to more natural and interpretable images, while the effect is not as clear in the case of VGG-16. In fig. 18 we show more examples, where we maximize the activation for a certain class.

6 Image enhancement

We also use the proposed deep image regularization to perform high frequency enhancement in an image. As demonstrated in section 2.1, the noisy image is reconstructed starting from coarse low-frequency details and finishing with fine high frequency details and noise. To perform enhancement we use the objective (4) setting the target image to be x0x_{0}. We stop the optimization process at a certain point, obtaining a coarse approximation xcx_{c} of the image x0x_{0}. The fine details are then computed as

We then construct an enhanced image by boosting the extracted fine details xfx_{f}:

In fig. 19 we present coarse and enhanced versions of the same image, running the optimization process for different number of iterations. At the start of the optimization process (corresponds to low number of iteration) the resulted approximation does not precisely recreates the shape of the objects (c.f. blue halo in the bottom row of fig. 19). While the shapes become well-matched with the time, unwanted high frequency details also start to appear. Thus we need to stop the optimization process in time.

7 Flash-no flash reconstruction

While in this work we focus on single image restoration, the proposed approach can be extended to the tasks of the restoration of multiple images, e.g. for the task of video restoration. We therefore conclude the set of application examples with a qualitative example demonstrating how the method can be applied to perform restoration based on pairs of images. In particular, we consider flash-no flash image pair-based restoration PetschniggSACHT04 , where the goal is to obtain an image of a scene with the lighting similar to a no-flash image, while using the flash image as a guide to reduce the noise level.

In general, extending the method to more than one image is likely to involve some coordinated optimization over the input codes zz that for single-image tasks in our approach was most often kept fixed and random. In the case of flash-no-flash restoration, we found that good restorations were obtained by using the denoising formulation (4), while using flash image as an input (in place of the random vector zz). The resulting approach can be seen as a non-linear generalization of guided image filtering he2013guided . The results of the restoration are given in the fig. 20.

Technical details

We use encoder-decoder (“hourglass”) architecture (possibly with skip-connections) for fθf_{\theta} in all our experiments except noted otherwise (fig. 21), varying a small number of hyper-parameters. Although the best results can be achieved by carefully tuning an architecture for a particular task (and potentially for a particular image), we found that wide range of hyper-parameters and architectures give acceptable results.

We use LeakyReLU he2015delving as a non-linearity. As a downsampling technique we simply use strides implemented within convolution modules. We also tried average/max pooling and downsampling with Lanczos kernel, but did not find a consistent difference between any of them. As an upsampling operation we choose between bilinear upsampling and nearest neighbor upsampling. An alternative upsampling method could be to use transposed convolutions, but the results we obtained using them were worse. We use reflection padding instead of zero padding in convolution layers everywhere except for the feature inversion and activation maximization experiments.

During fitting of the networks we often use a noise-based regularization. I.e. at each iteration we perturb the input zz with an additive normal noise with zero mean and standard deviation σp\sigma_{p}. While we have found such regularization to impede optimization process, we also observed that the network was able to eventually optimize its objective to zero no matter the variance of the additive noise (i.e. the network was always able to adapt to any reasonable variance for sufficiently large number of optimization steps).

We found the optimization process tends to destabilize as the loss goes down and approaches a certain value. Destabilization is observed as a significant loss increase and blur in generated image fθ(z)f_{\theta}(z). From such destabilization point the loss goes down again till destabilized one more time. To remedy this issue we simply track the optimization loss and return to parameters from the previous iteration if the loss difference between two consecutive iterations is higher than a certain threshold.

Finally, we use ADAM optimizer Kingma14adam in all our experiments and PyTorch as a framework. The proposed iterative optimization requires repeated forward and backward evaluation of a deep ConvNet and thus takes several minutes per image.

Below, we provide the remaining details of the network architectures. We use the notation introduced in fig. 21.

For 8×\times super-resolution (fig. 8) we have changed the standard deviation of the input noise to σp=120\sigma_{p}=\frac{1}{20} and the number of iterations to 40004000.

Text inpainting (fig. 11). We used the same hyper-parameters as for super-resolution but optimized the objective for 60006000 iterations.

Large hole inpainting (fig. 10). We used the same hyper-parameters as for super-resolution, but used meshgrid as an input, removed skip connections and optimized for 50005000 iterations.

Denoising (fig. 7). Hyper-parameters were set to be the same as in the case of super-resolution with only difference in iteration number, which was set to 18001800. We used the following implementations of referenced denoising methods: BM3Dcode for CBM3D and NLMcode for NLM. We used exponential sliding window with weight γ=0.99\gamma=0.99.

Image reconstruction (fig. 13). We used the same setup as in the case of super-resolution and denoising, but set num_iter = 11000, LR = 0.001.

Image enhancement (fig. 19). We used the same setup as in the case of super-resolution and denoising, but set σp=0\sigma_{p}=0.

Related work

Our approach is related to image restoration and synthesis methods based on learnable ConvNets and referenced above. Here, we review other lines of work related to our approach.

Modelling “translation-invariant” statistics of natural images using filter responses has a very long history of research. The statistics of responses to various non-random filters (such as simple operators and higher-order wavelets) have been studied in seminal works Field87 ; Mallat89 ; Simoncelli96 ; Zhu97 . Later, huang2000statistics noted that image response distribution w.r.t. random unlearned filters have very similar properties to the distributions of wavelet filter responses.

Our approach is closely related to a group of restoration methods that avoid training on the hold-out set and exploit the well-studied self-similarity properties of natural images Ruderman94 ; Turiel98 . This group includes methods based on joint modeling of groups of similar patches inside corrupted image buades2005non ; dabov2007image ; glasner2009super , which are particularly useful when the corruption process is complex and highly variable (e.g. spatially-varying blur bahat2017non ).

In this group, an interesting parallel work with clear links to our approach is the zero-shot super-resolution approach Shocher18 , which trains a feed-forward super-resolution ConvNet based on synthetic dataset generated from the patches of a single image. While clearly related, the approach Shocher18 is somewhat complementary as it exploit self-similarities across multiple scales of the same image, while our approach exploits self-similarities within the same scale (at multiple scales).

Several lines of work use dataset-based learning and modeling images using convolutional operations. Learning priors for natural images that facilitate restoration by enforcing filter responses for certain (learned) filters is behind an influential field-of-experts model Roth09 . Also in this group are methods based on fitting dictionaries to the patches of the corrupted image mairal2010online ; set14 as well as methods based on convolutional sparse coding Grosse07 ; Bristow13 . The connections between convolutional sparse coding and ConvNets are investigated in Papyan17jmlr in the context of recognition tasks. More recently in PapyanRSE17 , a fast single-layer convolutional sparse coding is proposed for reconstruction tasks. The comparison of our approach with PapyanRSE17 (figs. 11 and 13) however suggests that using deep ConvNet architectures popular in modern deep learning-based approaches may lead to more accurate restoration results.

Deeper convolutional models of natural images trained on large datasets have also been studied extensively. E.g. deconvolutional networks zeiler2010deconvolutional are trained by fitting hierarchies of representations linked by convolutional operators to datasets of natural images. The recent work lefkimmiatis2016non investigates the model that combines ConvNet with a self-similarity based denoising and thus bridges learning on image datasets and exploiting within-image self-similarities.

Our approach is also related to inverse scale space denoising Scherzer01 ; Burger05 ; Marquina09 . In this group of “non-deep” image processing methods, a sequence of solutions (a flow) that gradually progresses from a uniform image to the noisy image, while progressively finer scale details are recovered so that early stopping yields a denoised image. The inverse scale space approaches are however still driven by a simple total variation (TV) prior, which does not model self-similarity of images, and limits the ability to denoise parts of images with textures and gradual transitions. Note that our approach can also use the simple stopping criterion proposed in Burger05 , when the level of noise is known.

Finally, we note that this manuscript expands the conference version Ulyanov18 in multiple ways: 1) It gives more intuition, provides more visualizations and explanation for the presented method altogether with extensive technical details. 2) It contains a more thorough experimental evaluation and shows an application to activation maximization and high frequency enhancement. Since the publication of the preliminary version of our approach, it has also been used by other groups in different ways. Thus, Veen18CompressedSensing proposes a novel method for compressed sensing recovery using deep image prior. The work Athar18LCM learns a latent variable model, where the latent space is parametrized by a convolutional neural network. The approach Shedligeri18 aims to reconstruct an image from an event-based camera and utilizes deep image prior framework to estimate sensor’s ego-motion. The method Ilyas17 successively applies deep image prior to defend against adversarial attacks. Deep image prior is also used in Boominathan18 to perform phase retrieval for Fourier ptychography.

Discussion

We have investigated the success of recent image generator neural networks, teasing apart the contribution of the prior imposed by the choice of architecture from the contribution of the information transferred from external images through learning. In particular, we have shown that fitting a randomly-initialized ConvNet to corrupted images works as a “Swiss knife” for restoration problems. This approach is probably too slow to be useful for most practical applications, and for each particular application, a feed-forward network trained for that particular application would do a better job and do so much faster. Thus, the slowness and the inability to match or exceed the results of problem specific methods are the two main limitations of our approach, when practical applications are considered. While of limited practicality, the good results of our approach across a wide variety of tasks demonstrate that an implicit prior inside deep convolutional network architectures is an important part of the success of such architectures for image restoration tasks.

Why does this prior emerge, and, more importantly, why does it fit the structure of natural images so well? We speculate that generation by convolutional operations naturally tends impose self-similarity of the generated images (c.f. fig. 5), as convolutional filters are applied across the entire visual field thus imposing certain stationarity on the output of convolutional layers. Hourglass architectures with skip connections naturally impose self-similarity at multiple scales, making the corresponding priors suitable for the restoration of natural images.

We note that our results go partially against the common narrative that explain the success of deep learning in image restoration (and beyond) by the ability to learn rather than by hand-craft priors; instead, we show that properly hand-crafted network architectures correspond to better hand-crafted priors, and it seems that learning ConvNets builds on this basis. This observation also validates the importance of developing new deep learning architectures.

DU and VL are supported by the Ministry of Education and Science of the Russian Federation (grant 14.756.31.0001) and AV is supported by ERC 638009-IDIU.

References