Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise
Arpit Bansal, Eitan Borgnia, Hong-Min Chu, Jie S. Li, Hamid Kazemi, Furong Huang, Micah Goldblum, Jonas Geiping, Tom Goldstein
Introduction
Diffusion models have recently emerged as powerful tools for generative modeling (Ramesh et al., 2022). Diffusion models come in many flavors, but all are built around the concept of random noise removal; one trains an image restoration/denoising network that accepts an image contaminated with Gaussian noise, and outputs a denoised image. At test time, the denoising network is used to convert pure Gaussian noise into a photo-realistic image using an update rule that alternates between applying the denoiser and adding Gaussian noise. When the right sequence of updates is applied, complex generative behavior is observed.
The origins of diffusion models, and also our theoretical understanding of these models, are strongly based on the role played by Gaussian noise during training and generation. Diffusion has been understood as a random walk around the image density function using Langevin dynamics (Sohl-Dickstein et al., 2015; Song and Ermon, 2019), which requires Gaussian noise in each step. The walk begins in a high temperature (heavy noise) state, and slowly anneals into a “cold” state with little if any noise. Another line of work derives the loss for the denoising network using variational inference with a Gaussian prior (Ho et al., 2020; Song et al., 2021a; Nichol and Dhariwal, 2021).
The existence of cold diffusions that require no Gaussian noise (or any randomness) during training or testing raises questions about the limits of our theoretical understanding of diffusion models. It also unlocks the door for potentially new types of generative models with very different properties than conventional diffusion seen so far.
Background
Generative models exist for a range of modalities spanning natural language (Brown et al., 2020) and images (Brock et al., 2019; Dhariwal and Nichol, 2021), and they can be extended to solve important problems such as image restoration (Kawar et al., 2021a, 2022). While GANs (Goodfellow et al., 2014) have historically been the tool of choice for image synthesis (Brock et al., 2019; Wu et al., 2019), diffusion models (Sohl-Dickstein et al., 2015) have recently become competitive if not superior for some applications (Dhariwal and Nichol, 2021; Nichol et al., 2021; Ramesh et al., 2021; Meng et al., 2021).
Both the Langevin dynamics and variational inference interpretations of diffusion models rely on properties of the Gaussian noise used in the training and sampling pipelines. From the score-matching generative networks perspective (Song and Ermon, 2019; Song et al., 2021b), noise in the training process is critically thought to expand the support of the low-dimensional training distribution to a set of full measure in ambient space. The noise is also thought to act as data augmentation to improve score predictions in low density regions, allowing for mode mixing in the stochastic gradient Langevin dynamics (SGLD) sampling. The gradient signal in low-density regions can be further improved during sampling by injecting large magnitudes of noise in the early steps of SGLD and gradually reducing this noise in later stages.
Kingma et al. (2021) propose a method to learn a noise schedule that leads to faster optimization. Using a classic statistical result, Kadkhodaie and Simoncelli (2021) show the connection between removing additive Gaussian noise and the gradient of the log of the noisy signal density in deterministic linear inverse problems. Here, we shed light on the role of noise in diffusion models through theoretical and empirical results in applications to inverse problems and image generation.
Iterative neural models have been used for various inverse problems (Romano et al., 2016; Metzler et al., 2017). Recently, diffusion models have been applied to them (Song et al., 2021b) for the problems of deblurring, denoising, super-resolution, and compressive sensing (Whang et al., 2021; Kawar et al., 2021b; Saharia et al., 2021; Kadkhodaie and Simoncelli, 2021).
Although not their focus, previous works on diffusion models have included experiments with deterministic image generation (Song et al., 2021a; Dhariwal and Nichol, 2021) and in selected inverse problems (Kawar et al., 2022). Here, we show definitively that noise is not a necessity in diffusion models, and we observe the effects of removing noise for a number of inverse problems.
Despite prolific work on generative models in recent years, methods to probe the properties of learned distributions and measure how closely they approximate the real training data are by no means closed fields of investigation. Indirect feature space similarity metrics such as Inception Score (Salimans et al., 2016), Mode Score (Che et al., 2016), Frechet inception distance (FID) (Heusel et al., 2017), and Kernel inception distance (KID) (Bińkowski et al., 2018) have been proposed and adopted to some extent, but they have notable limitations (Barratt and Sharma, 2018). To adopt a popular frame of reference, we will use FID as the feature similarity metric for our experiments.
Generalized Diffusion
Standard diffusion models are built around two components. First, there is an image degradation operator that contaminates images with Gaussian noise. Second, a trained restoration operator is created to perform denoising. The image generation process alternates between the application of these two operators. In this work, we consider the construction of generalized diffusions built around arbitrary degradation operations. These degradations can be randomized (as in the case of standard diffusion) or deterministic.
In the standard diffusion framework, adds Gaussian noise with variance proportional to . In our generalized formulation, we choose to perform various other transformations such as blurring, masking out pixels, downsampling, and more, with severity that depends on . We explore a range of choices for in Section 4.
We also require a restoration operator that (approximately) inverts . This operator has the property that
In practice, this operator is implemented via a neural network parameterized by . The restoration network is trained via the minimization problem
2 Sampling from the model
After choosing a degradation and training a model to perform the restoration, these operators can be used in tandem to invert severe degradations by using standard methods borrowed from the diffusion literature. For small degradations (), a single application of can be used to obtain a restored image in one shot. However, because is typically trained using a simple convex loss, it yields blurry results when used with large . Rather, diffusion models (Song et al., 2021a; Ho et al., 2020) perform generation by iteratively applying the denoising operator and then adding noise back to the image, with the level of added noise decreasing over time. This corresponds to the standard update sequence in Algorithm 1.
When the restoration operator is perfect, i.e. when for all one can easily see that Algorithm 1 produces exact iterates of the form . But what happens for imperfect restoration operators? In this case, errors can cause the iterates to wander away from , and inaccurate reconstruction may occur.
We find that the standard sampling approach in Algorithm 1 works well for noise-based diffusion, possibly because the restoration operator has been trained to correct (random Gaussian) errors in its inputs. However, we find that it yields poor results in the case of cold diffusions with smooth/differentiable degradations as demonstrated for a deblurring model in Figure 2. We propose Algorithm 2 for sampling, which we find to be superior for inverting smooth, cold degradations.
This sampler has important mathematical properties that enable it to recover high quality results. Specifically, for a class of linear degradation operations, it can be shown to produce exact reconstruction (i.e. ) even when the restoration operator fails to perfectly invert . We discuss this in the following section.
3 Properties of Algorithm 2
It is clear from inspection that both Algorithms 1 and 2 perfectly reconstruct the iterate for all if the restoration operator is a perfect inverse for the degradation operator. In this section, we analyze the stability of these algorithms to errors in the restoration operator.
For small values of and , Algorithm 2 is extremely tolerant of error in the restoration operator . To see why, consider a model problem with a linear degradation function of the form for some vector . While this ansatz may seem rather restrictive, note that the Taylor expansion of any smooth degradation around has the form where HOT denotes higher order terms. Note that the constant/zeroth-order term in this Taylor expansion is zero because we assumed above that the degradation operator satisfies .
For a degradation of the form (3.3) and any restoration operator , the update in Algorithm 2 can be written
By induction, we see that the algorithm produces the value for all regardless of the choice of . In other words, for any choice of , the iteration behaves the same as it would when is a perfect inverse for the degradation .
By contrast, Algorithm 1 does not enjoy this behavior. In fact, when is not a perfect inverse for , is not even a fixed point of the update rule in Algorithm 1 because If does not perfectly invert we should expect Algorithm 1 to incur errors, even for small values of . Meanwhile, for small values of , the behavior of approaches its first-order Taylor expansion and Algorithm 2 becomes immune to errors in . We demonstrate the stability of Algorithm 2 vs Algorithm 1 on a deblurring model in Figure 2.
Generalized Diffusions with Various Transformations
In this section, we take the first step towards cold diffusion by reversing different degradations and hence performing conditional generation. We will extend our methods to perform unconditional (i.e. from scratch) generation in Section 5. We emprically evaluate generalized diffusion models trained on different degradations with our improved sampling Algorithm 2. We perform experiments on the vision tasks of deblurring, inpainting, super-resolution, and the unconventional task of synthetic snow removal. We perform our experiments on MNIST (LeCun et al., 1998), CIFAR-10 (Krizhevsky, 2009), and CelebA (Liu et al., 2015). In each of these tasks, we gradually remove the information from the clean image, creating a sequence of images such that retains less information than . For these different tasks, we present both qualitative and quantitative results on a held-out testing dataset and demonstrate the importance of the sampling technique described in Algorithm 2. For all quantitative results in this section, the Frechet inception distance (FID) scores (Heusel et al., 2017) for degraded and reconstructed images are measured with respect to the testing data. Additional information about the quantitative results, convergence criteria, hyperparameters, and architecture of the models presented below can be found in the appendix.
We consider a generalized diffusion based on a Gaussian blur operation (as opposed to Gaussian noise) in which an image at step has more blur than at . The forward process given the Gaussian kernels and the image at step can thus be written as
where denotes the convolution operator, which blurs an image using a kernel.
We train a deblurring model by minimizing the loss (1), and then use Algorithm 2 to invert this blurred diffusion process for which we trained a DNN to predict the clean image . Qualitative results are shown in Figure 3 and quantitative results in Table 1. Qualitatively, we can see that images created using the sampling process are sharper and in some cases completely different as compared to the direct reconstruction of the clean image. Quantitatively we can see that the reconstruction metrics such as RMSE and PSNR get worse when we use the sampling process, but on the other hand FID with respect to held-out test data improves. The qualitative improvements and decrease in FID show the benefits of the generalized sampling routine, which brings the learned distribution closer to the true data manifold.
In the case of blur operator, the sampling routine can be thought of adding frequencies at each step. This is because the sampling routine involves the term which in the case of blur becomes . This results in a difference of Gaussians, which is a band pass filter and contains frequencies that were removed at step . Thus, in the sampling process, we sequentially add the frequencies that were removed during the degradation process.
2 Inpainting
We define a schedule of transforms that progressively grays-out pixels from the input image. We remove pixels using a Gaussian mask as follows: For input images of size we start with a 2D Gaussian curve of variance discretized into an array. We normalize so the peak of the curve has value 1, and subtract the result from 1 so the center of the mask as value 0. We randomize the location of the Gaussian mask for MNIST and CIFAR-10, but keep it centered for CelebA. We denote the final mask by .
Input images are iteratively masked for steps via multiplication with a sequence of masks with increasing . We can control the amount of information removed at each step by tuning the parameter. In the language of Section 3, , where the operator denotes entry-wise multiplication.
Figure 4 presents results on test images and compares the output of the inpainting model to the original image. The reconstructed images display reconstructed features qualitatively consistent with the context provided by the unperturbed regions of the image. We quantitatively assess the effectiveness of the inpainting models on each of the datasets by comparing distributional similarity metrics before and after the reconstruction. Our results are summarized in Table 2. Note, the FID scores here are computed with respect to the held-out validation set.
3 Super-Resolution
For this task, the degradation operator downsamples the image by a factor of two in each direction. This takes place, once for each values of , until a final resolution is reached, 44 in the case of MNIST and CIFAR-10 and 22 in the case of Celeb-A. After each down-sampling, the lower-resolution image is resized to the original image size, using nearest-neighbor interpolation. Figure 5 presents example testing data inputs for all datasets and compares the output of the super-resolution model to the original image. Though the reconstructed images are not perfect for the more challenging datasets, the reconstructed features are qualitatively consistent with the context provided by the low resolution image.
Table 3 compares the distributional similarity metrics between degraded/reconstructed images and test samples.
4 Snowification
Apart from traditional degradations, we additionally provide results for the task of synthetic snow removal using the offical implementation of the snowification transform from ImageNet-C (Hendrycks and Dietterich, 2019). The purpose of this experiment is to demonstrate that generalized diffusion can succeed even with exotic transforms that lack the scale-space and compositional properties of blur operators. Similar to other tasks, we degrade the images by adding snow, such that the level of snow increases with step . We provide more implementation details in Appendix.
We illustrate our desnowification results in Figure 6. We present testing examples, as well as their snowified images, from all the datasets, and compare the desnowified results with the original images. The desnowified images feature near-perfect reconstruction results for CIFAR-10 examples with lighter snow, and exhibit visually distinctive restoration for Celeb-A examples with heavy snow. We provide quantitative results in Table 4.
Cold Generation
Diffusion models can successfully learn the underlying distribution of training data, and thus generate diverse, high quality images (Song et al., 2021a; Dhariwal and Nichol, 2021; Jolicoeur-Martineau et al., 2021; Ho et al., 2022). We will first discuss deterministic generation using Gaussian noise and then discuss in detail unconditional generation using deblurring. Finally, we provide a proof of concept that the Algorithm 2 can be extended to other degradations.
Here we discuss image generation using noise-based degradation. We consider “deterministic” sampling in which the noise pattern is selected and frozen at the start of the generation process, and then treated as a constant. We study two ways of applying Algorithm 2 with fixed noise. We first define
as the (deterministic) interpolation between data point and a fixed noise pattern , for increasing , as in Song et al. (2021a). Algorithm 2 can be applied in this case by fixing the noise used in the degradation operator . Alternatively, one can deterministically calculate the noise vector to be used in step of reconstruction by using the formula
The second method turns out to be closely related to the deterministic sampling proposed in Song et al. (2021a), with some differences in the formulation of the training objective. We discuss this relationship in detail in Appendix A.6. We present quantitative results for CelebA and AFHQ datasets using the fixed noise method and the estimated noise method (using ) in Table 5.
2 Image generation using blur
The forward diffusion process in noise-based diffusion models has the advantage that the degraded image distribution at the final step is simply an isotropic Gaussian. One can therefore perform (unconditional) generation by first drawing a sample from the isotropic Gaussian, and sequentially denoising it with backward diffusion.
When using blur as a degradation, the fully degraded images do not form a nice closed-form distribution that we can sample from. They do, however, form a simple enough distribution that can be modeled with simple methods. Note that every image degenerates to an that is constant (i.e., every pixel is the same color) for large . Furthermore, the constant value is exactly the channel-wise mean of the RGB image , and can be represented with a 3-vector. This 3-dimensional distribution is easily represented using a Gaussian mixture model (GMM). This GMM can be sampled to produce the random pixel values of a severely blurred image, which can be deblurred using cold diffusion to create a new image.
Our generative model uses a blurring schedule where we progressively blur each image with a Gaussian kernel of size 27x27 over 300 steps. The standard deviation of the kernel starts at 1 and increases exponentially at the rate of 0.01. We then fit a simple GMM with one component to the distribution of channel-wise means. To generate an image from scratch, we sample the channel-wise mean from the GMM, expand the 3D vector into a image with three channels, and then apply Algorithm 2.
Empirically, the presented pipeline generates images with high fidelity but low diversity, as reflected quantitatively by comparing the perfect symmetry column with results from hot diffusion in Table 5. We attribute this to the perfect correlation between pixels of sampled from the channel-wise mean Gaussian mixture model. To break the symmetry between pixels, we add a small amount of Gaussian noise (of standard deviation ) to each sampled . As shown in Table 5, the simple trick drastically improves the quality of generated images. We also present the qualitative results for cold diffusion using blur transformation in Figure 7, and further discuss the necessity of Algorithm 2 for generation in Appendix A.7.
3 Generation using other transformations
In this section, we further provide a proof of concept that generation can be extended to other transformations. Specifically, we show preliminary results on inpainting, super-resolution, and animorphosis. Inspired by the simplicity of the degraded image distribution for the blurring routine presented in the previous section, we use degradation routines with predictable final distributions here as well.
To use the Gaussian mask transformation for generation, we modify the masking routine so the final degraded image is completely devoid of information. One might think a natural option is to send all of the images to a completely black image , but this would not allow for any diversity in generation. To get around this maximally non-injective property, we instead make the mask turn all pixels to a random, solid color. This still removes all of the information from the image, but it allows us to recover different samples from the learned distribution via Algorithm 2 by starting off with different color images. More formally, a Gaussian mask is created in a similar way as discussed in the Section 4.2, but instead of multiplying it directly to the image , we create as follows:
where is an image of a randomly sampled color.
For super-resolution, the routine down-samples to a resolution of , or values in each channel. These degraded images can be represented as one-dimensional vectors, and their distribution is modeled using one Gaussian distribution. Using the same methods described for generation using blurring described above, we sample from this Gaussian-fitted distribution of the lower-dimensional degraded image space and pass this sampled point through the generation process trained on super-resolution data to create one output.
Additionally to show one can invert nearly any transformation, we include a new transformation deemed animorphosis, where we iteratively transform a human face from CelebA to an animal face from AFHQ. Though we chose CelebA and AFHQ for our experimentation, in principle such interpolation can be done for any two initial data distributions.
More formally, given an image and a random image sampled from the AFHQ manifold, can be written as follows:
Note this is essentially the same as the noising procedure, but instead of adding noise we are adding a progressively higher weighted AFHQ image. In order to sample from the learned distribution, we sample a random image of an animal and use Algorithm 2 to reverse the animorphosis transformation.
We present results for the CelebA dataset, and hence the quantitative results in terms of FID scores for inpainting, super-resolution and animorphosis are 90.14, 92.91 and 48.51 respectively. We further show some qualitative samples in Figure 8, and in Figure LABEL:fig:all_transforms_cover.
Conclusion
Existing diffusion models rely on Gaussian noise for both forward and reverse processes. In this work, we find that the random noise can be removed entirely from the diffusion model framework, and replaced with arbitrary transforms. In doing so, our generalization of diffusion models and their sampling procedures allows us to restore images afflicted by deterministic degradations such as blur, inpainting and downsampling. This framework paves the way for a more diverse landscape of diffusion models beyond the Gaussian noise paradigm. The different properties of these diffusions may prove useful for a range of applications, including image generation and beyond.
References
Appendix A Appendix
For the deblurring experiments, we train the models on different datasets for 700,000 gradient steps. We use the Adam [Kingma and Ba, 2014] optimizer with learning rate . The training was done on the batch size of 32, and we accumulate the gradients every 2 steps. Our final model is an Exponential Moving Average of the trained model with decay rate 0.995 which is updated after every 10 gradient steps. For the MNIST dataset, we blur recursively 40 times, with a discrete Gaussian kernel of size 11x11 and a standard deviation 7. In the case of CIFAR-10, we recursively blur with a Gaussian kernel of fixed size 11x11, but at each step , the standard deviation of the Gaussian kernel is given by . The blur routine for CelebA dataset involves blurring images with a Gaussian kernel of 15x15 and the standard deviation of the Gaussian kernel grows exponentially with time at the rate of .
Figure 9 shows an additional nine images for each of MNIST, CIFAR-10 and CelebA. Figures 19 and 20 show the iterative sampling process using a deblurring model for ten example images from each dataset. We further show 400 random images to demonstrate the qualitative results in the Figure 21.
A.2 Inpainting
For the inpainting transformation, models were trained on different datasets with 60,000 gradient steps. The models were trained using Adam [Kingma and Ba, 2014] optimizer with learning rate . We use batch size 64, and the gradients are accumulated after every 2 steps. The final model is an Exponential Moving Average of the trained model with decay rate 0.995. This EMA model is updated after every 10 gradient steps. For all our inpainting experiments we use a randomized Gaussian mask and with and .
To avoid potential leakage of information due to floating point computation of the Gaussian mask, we discretize the masked image before passing it through the inpainting model. This was done by rounding all pixel values to the eight most significant digits.
Figure 11 shows nine additional inpainting examples on each of the MNIST, CIFAR-10, and CelebA datasets. Figure 10 demonstrates an example of the iterative sampling process of an inpainting model for one image in each dataset.
A.3 Super-Resolution
We train the super-resolution model per Section 3.1 for 700,000 iterations. We use the Adam [Kingma and Ba, 2014] optimizer with learning rate . The batch size is 32, and we accumulate the gradients every 2 steps. Our final model is an Exponential Moving Average of the trained model with decay rate 0.995. We update the EMA model every 10 gradient steps.
The number of time-steps depends on the size of the input image and the final image. For MNIST and for CIFAR10, the number of time steps is 3, as it takes three steps of halving the resolution to reduce the initial image down to . For CelebA, the number of time steps is 6 to reduce the initial image down to . For CIFAR10, we apply random crop and random horizontal flip for regularization.
Figure 13 shows an additional nine super-resolution examples on each of the MNIST, CIFAR-10, and CelebA datasets. Figure 12 shows one example of the progressive increase in resolution achieved with the sampling process using a super-resolution model for each dataset.
A.4 Colorization
Here we provide results for the additional task of colorization. Starting with the original RGB-image , we realize colorization by iteratively desaturating for steps until the final image is a fully gray-scale image. We use a series of three-channel convolution filters with the form
and obtain via a schedule defined as for each respective step. Notice that a gray image is obtained when .
We can tune the ratio to control the amount of information removed in each step. For our experiment, we schedule the ratio such that for every we have
This schedule ensures that color information lost between steps is smaller in earlier stage of the diffusion and becomes larger as increases.
We train the models on different datasets for 700,000 gradient steps. We use Adam [Kingma and Ba, 2014] optimizer with learning rate . We use batch size 32, and we accumulate the gradients every steps. Our final model is an exponential moving average of the trained model with decay rate 0.995. We update the EMA model every 10 gradient steps. For CIFAR-10 we use and for CelebA we use .
We illustrate our recolorization results in Figure 14. We present testing examples, as well as their grey scale images, from all the datasets, and compare the recolorization results with the original images. The recolored images feature correct color separation between different regions, and feature various and yet semantically correct colorization of objects. Our sampling technique still yields minor differences in comparison to the direct reconstruction, although the change is not visually apparent. We attribute this to the shape restriction of colorization task, as human perception is rather insensitive to minor color change. We also provide quantitative measurement for the effectiveness of our recolorization results in terms of different similarity metrics, and summarize the results in Table 6.
A.5 Image Snow
and clip each entry of into the range $S_{C}c_{2}SS^{\prime}x_{0}+S+S^{\prime}h(x_{0},S_{A},c_{0},c_{1})$.
To create a series of images with increasing snowification, we linearly interpolate and between and respectively, to create and , . Then for each , a seed matrix is sampled, the motion blur direction is randomized, and we construct each related by . Visually, dictates the severity of the snow, while determines how “windy" the snowified image seems.
For both CIFAR-10 and Celeb-A, we use the same Gaussian distribution with parameters and to generate the seed matrix. For CIFAR-10, we choose , , and , which generates a visually lighter snow. For Celeb-A, we choose , , and , which generates a visually heavier snow.
We train the models on different datasets for 700,000 gradient steps. We use Adam [Kingma and Ba, 2014] optimizer with learning rate . We use batch size 32, and we accumulate the gradients every steps. Our final model is an exponential moving average of the trained model with decay rate 0.995. We update the EMA model every 10 gradient steps. For CIFAR-10 we use and for CelebA we use . We note that the seed matrix is resampled for each individual training batch, and hence the snow pattern varies across the training stage.
A.6 Generation using noise : Further Details
Here we will discuss in further detail on the similarity between the sampling method proposed in Algorithm 2 and the deterministic sampling in DDIM [Song et al., 2021a]. Given the image at step , we have the restored clean image from the diffusion model. Hence given the estimated and , we can estimate the noise (or ) as
Thus, the and can be written as
using which the sampling process in Algorithm 2 to estimate can be written as,
which is same as the sampling method as described in [Song et al., 2021a].
A.7 Generation using blur transformation: Further Details
The Figure 16, shows the generation without breaking any symmetry within each channel are quite promising as well.
Necessity of Algorithm 2: In the case of unconditional generation, we observe a marked superiority in quality of the sampled reconstruction using Algorithm 2 over any other method considered. For example, in the broken symmetry case, the FID of the directly reconstructed images is 257.69 for CelebA and 214.24 for AFHQ, which are far worse than the scores of 49.45 and 54.68 from Table 5. In Figure 17, we also give a qualitative comparison of this difference. We can also clearly see from Figure 18 that Algorithm 1, the method used in Song et al. [2021b] and Ho et al. , completely fails to produce an image close to the target data distribution.