SRDiff: Single Image Super-Resolution with Diffusion Probabilistic Models

Haoying Li, Yifan Yang, Meng Chang, Huajun Feng, Zhihai Xu, Qi Li, Yueting Chen

Introduction

Over the years, single image super-resolution (SISR) has drawn active attention due to its wide applications in computer vision such as object recognition Fookes et al. (2012); Sajjadi et al. (2017), remote sensing Li et al. (2009), surveillance monitoring Fang et al. (2019); Park et al. (2020) and so on. SISR aims to recover high-resolution (HR) images from low-resolution (LR) ones, which is an ill-posed problem, for multiple HR images can be degenerated to one LR image as shown in Figure 1.

To establish the mapping between HR and LR images, lots of deep learning-based methods emerge and could be categorized into three main types: PSNR-oriented, GAN-driven and flow-based methods. PSNR-oriented methods Dong et al. (2015); Lim et al. (2017); Zhang et al. (2018b); Qiu et al. (2019); Guo et al. (2020) are trained with simple distribution assumption-based losses (e.g., Laplacian for L1L_{1} and Gaussian for L2L_{2}) and achieve excellent PSNR. However, these losses tend to drive the SR result to an average of several possible SR predictions, causing over-smoothing images with high-frequency information loss. One ground-breaking solution to tackle the over-smoothing problem is GAN-driven methods Cheon et al. (2018); Kim et al. (2019); Ledig et al. (2017); Wang et al. (2018), which combine content losses (e.g., L1L_{1} and L2L_{2}) and adversarial losses to obtain sharper SR images with better perceptual quality. However, GAN-driven methods are easy to fall into mode collapse, which leads to a single generated SR sample without diversity. Additionally, GAN-based training process is not easy to converge and requires an extra discriminator which is not used in inference. Flow-based methods Lugmayr et al. (2020) directly account for the ill-posed problem with an invertible encoder, which maps HR images to the flow-space latents conditioned on LR inputs. Trained with a negative log-likelihood loss, flow-based methods avoid training instability but suffer from extremely large footprint and high training cost due to the strong architectural constraints to keep the bijection between latents and data.

Lately, the successful adoptions of diffusion probabilistic models (diffusion models for short) in image synthesis Ho et al. (2020) and speech synthesis Kong et al. (2020); Chen et al. (2021) witness the power of diffusion models in generative tasks. The diffusion models use a Markov chain to convert data x0x_{0} to latent variable xTx_{T} in simple distribution (e.g., Gaussian) by gradually adding noise ϵ\epsilon in the diffusion process, and predict the noise ϵ\epsilon in each diffusion step to recover the data x0x_{0} through a learned reverse process. Diffusion models are trained by optimizing a variant of the variational lower bound, which is efficient and avoids the mode collapse encountered by GAN.

In this paper, we propose a novel single image super-resolution diffusion probabilistic model (SRDiff) to tackle the over-smoothing, mode collapse and huge footprint problems in previous SISR models. Specifically, 1) to extract the image information in LR image, SRDiff exploits a pretrained low-resolution encoder to convert LR image into hidden condition. 2) To generate the HR image conditioned on LR image, SRDiff employs a conditional noise predictor to recover x0x_{0} iteratively. 3) To speed up convergence and stabilize training, SRDiff introduces residual prediction by taking the difference between the HR and LR image as the input x0x_{0} in the first diffusion step, making SRDiff focus on restoring high-frequency details. To the best of our knowledge, SRDiff is the first diffusion-based SR model and has several advantages:

Diverse and high-quality outputs: SRDiff converts Gaussian white noise into an SR prediction through a Markov chain, which does not suffer from mode collapse and can generate diverse and high-quality SR results.

Stable and efficient training with small footprint: Although the data distribution of HR image is hard to estimate, SRDiff admits a variant of the variational bound maximization and applies residual prediction. Compared with GAN-driven methods, SRDiff is stably trained with a single loss and does not need any extra module (e.g., discriminator, which is only used in training). Compared with flow-based methods, SRDiff has no architectural constraints and thus enjoys benefits from small footprint and fast training.

Flexible image manipulation: SRDiff can perform flexible image manipulation including latent space interpolation and content fusion using both diffusion process and reverse process, which shows broad application prospects.

Our extensive experiments on CelebA Liu et al. (2015a) and DIV2K Timofte et al. (2018) datasets show that 1) SRDiff can reconstruct multiple SR results given one LR input and outperform state-of-the-art SISR methods; 2) SRDiff only has 1/4 parameters and is stable and fast to train (about 30 hours on 1 GPU until convergence) compared with SRFlow; 3) we can manipulate the generated SR images in latent space to obtain more diverse outputs.

Related Works

Recently, deep learning methods have become widely adopted to single image super-resolution (SISR) and we categorize them into three types: PSNR-oriented and GAN-driven and flow-based methods. The training goal of PSNR-oriented methods is to minimize the mean squared error (MSE) between the ground truth and the SR image recovered from the LR image. SRCNN Dong et al. (2015) sets a precedent of end-to-end mapping between the LR and HR images. Kim et al. Kim et al. (2016b, a) then apply residual neural network techniques to SR tasks and deepen the network. Some worksLim et al. (2017); Zhang et al. (2018b); Qiu et al. (2019); Guo et al. (2020) enhance the SR performance by carefully adjusting network structures and losses. GAN-driven methods Cheon et al. (2018); Kim et al. (2019) solve the over-smoothing problem towards perceptual restrictions Rad et al. (2019); Zhang et al. (2019). The pioneer work is SRGAN Ledig et al. (2017), using SRResNet Ledig et al. (2017) and perceptual loss as well as the adversarial loss to improve the naturalness of the recovered image. ESRGAN Wang et al. (2018) further enhances it with network adjustments of structure and loss function. The first flow-based methods is SRFlow Lugmayr et al. (2020), which builds an invertible neural network to transform a Gaussian distribution into an HR image space instead of modeling one single output and inherently resolves the pathology of the original ”one-to-many” SR problem.

2 Diffusion models

Diffusion probabilistic models Sohl-Dickstein et al. (2015); Ho et al. (2020) are a kind of generative models using a Markov chain to transform latent variables in simple distributions (e.g., Gaussian) to data in complex distributions. Researchers find it useful to tackle ”one-to-many” problems and synthesize high-quality results in speech synthesis tasks Kong et al. (2020) and image synthesis fields Ho et al. (2020). However, to the best of our knowledge, diffusion models have not yet been used in image reconstruction fields like super-resolution. In this paper, we propose our impressive work, SRDiff, building on diffusion models to generate diverse SR images with a single LR input, and solving over-smoothing, mode collapse and large footprint issues together.

Diffusion Model

In this section, to provide a basic understanding of diffusion probabilistic models (diffusion model for short) Ho et al. (2020), we first briefly review its formulation.

The posterior q(x1,⋯ ,xT∣x0)q(x_{1},\cdots,x_{T}|x_{0}), named as the diffusion process, converts the data distribution q(x0)q(x_{0}) to the latent variable distribution q(xT)q(x_{T}), and is fixed to a Markov chain which gradually adds Gaussian noise ϵ\epsilon to the data according to a variance schedule β1,⋯ ,βT\beta_{1},\cdots,\beta_{T}:

where βt\beta_{t} is a small positive constant and could be regarded as constant hyperparameters. Setting αt:=1−βt, αˉt:=∏s=1tαs\alpha_{t}:=1-\beta_{t},~{}\bar{\alpha}_{t}:=\prod_{s=1}^{t}\alpha_{s}, the diffusion process allows sampling xtx_{t} at an arbitrary timestep tt in closed form:

The reverse process transforms the latent variable distribution pθ(xT)p_{\theta}(x_{T}) to the data distribution pθ(x0)p_{\theta}(x_{0}) parameterized by θ\theta. It is defined by a Markov chain with learned Gaussian transitions beginning with p(xT)=N(xT;0,I)p(x_{T})=\mathcal{N}(x_{T};\textbf{0},\textbf{I}):

In training phase, we maximize the variational lower bound (ELBO) on negative log likelihood and introduce KL divergence and variance reduction Ho et al. (2020):

For simplicity, the training procedure minimizes the variant of the ELBO with x0x_{0} and tt as inputs:

where ϵθ\epsilon_{\theta} is a noise predictor.

In inference, we first sample an xT∼N(xT;0,I)x_{T}\sim\mathcal{N}(x_{T};\textbf{0},\textbf{I}), and then sample xt−1∼pθ(xt−1∣xt)x_{t-1}\sim p_{\theta}(x_{t-1}|x_{t}) according to Eq. (3), where

SRDiff

As depicted in Figure 2, SRDiff is built on a T-step diffusion model which contains two processes: diffusion process and reverse process. Instead of predicting the HR image directly, we apply residual prediction to predict the difference between the HR image xHx_{H} and the upsampled LR image up(xL)up(x_{L}) and denote the difference as input residual image x0x_{0}. Diffusion process converts the x0x_{0} into a latent xTx_{T} in Gaussian distribution by gradually adding Gaussian noise ϵ\epsilon as implied in Eq. (2). According to Eq. (3) and (7), the reverse process is determined by ϵθ\epsilon_{\theta}, which is a conditional noise predictor with an RRDB-based Wang et al. (2018) low-resolution encoder (LR encoder for short) D\mathcal{D}, as shown in Figure 3. The reverse process converts a latent variable xTx_{T} to a residual image xrx_{r} by iteratively denoising in finite step TT using the conditional noise predictor ϵθ\epsilon_{\theta}, conditioned on the hidden states encoded from LR image by the LR encoder D\mathcal{D}. The SR image is reconstructed by adding the generated residual image xrx_{r} to the upsampled LR image up(xL)up(x_{L}). Therefore, the goal of ϵθ\epsilon_{\theta} is to predict the noise ϵ\epsilon added at each diffusion timestep in the diffusion processInstead of the original L2L_{2} in Eq. (6), we use L1L_{1} for better training stability following Chen et al. Chen et al. (2021)..

In the following subsections, we will introduce architectures of the conditional noise predictor, LR encoder, training and inference.

The conditional noise predictor ϵθ\epsilon_{\theta} predicts noise added in each timestep of the diffusion process conditioned on the LR image information, according to Eq. (6) and (7). As shown in Figure 3, we use U-Net as the main body, taking 3-channel xtx_{t}, the diffusion timestep t∈{1,2,...,T−1,T}t\in\{1,2,...,T-1,T\} and the output of LR encoder as inputs. First, xtx_{t} is transformed to hidden through a 2D-convolution block which consists of one 2D-convolutional layer and Mish activation Misra (2019). Then the LR information is fused with the 2D-convolution block output hidden. Following Ho et al. Ho et al. (2020), we transform the timestep tt to the timestep embedding tet_{e} using the Transformer sinusoidal positional encoding Vaswani et al. (2017). Then the last output hidden and tet_{e} are fed into the contracting path, one middle step and the expansive path successively. The contracting path and expansive path both consist of four steps, each of which successively applies two residual blocks and one downsampling/upsampling layer. To reduce the model size, we only double the channel size in the second and the fourth contracting steps and halve the spatial sizes of the feature map in each contracting step. The downsampling layer in contracting path is a two-stride 2D convolution and the upsampling layer in expansive path is 2D transposed convolution. The middle step consists of two residual blocks, which is inserted between the contracting and expansive paths. Besides, the inputs of each expansive step concatenate the corresponding feature map from the contracting path. Finally, a 2D-convolution block is applied to generate ϵ^\hat{\epsilon} in timestep t−1t-1 as the predicted noise, which is then used to recover xt−1x_{t-1} according to Eq. (3) and (7). Our conditional noise predictor is easy and stable to train due to the multi-scale skip connection. Moreover, it combines local and global information through the contracting and expansive path.

LR Encoder

An LR encoder encodes the LR information xex_{e}, which is added to each reverse step to guide the generation to the corresponding HR space. In this paper, we choose the RRDB architecture following SRFlow Lugmayr et al. (2020), which employs the residual-in-residual structure and multiple dense skip connections without batch normalization layers. In particular, we abandon the last convolution layer of the RRDB architecture because we do not aim at the concrete SR results but the hidden LR image information.

Training

In the training phase, as illustrated in Algorithm 1, the input LR-HR image pairs in the training set are used to train SRDiff with the total diffusion step TT (Line 1). We randomly initialize the conditional noise predictor ϵθ\epsilon_{\theta} and the RRDB based LR encoder D\mathcal{D} is pretrained by L1L_{1} loss function (Line 2). We then sample a mini-batch of LR-HR image pairs from the training set (Line 4) and compute the residual image xrx_{r} (Line 5). The LR images are encoded by the LR encoder as xex_{e} (Line 6), which is fed into the noise predictor ϵθ\epsilon_{\theta} together with tt and xTx_{T}. Then we sample ϵ\epsilon from the standard Gaussian distribution and tt from the integer set {1,⋯ ,T}\{1,\cdots,T\} (Line 7). We optimize the noise predictor by taking gradient step on Eq. (6) (Line 8).

Inference

A T-step SRDiff inference takes an LR image xLx_{L} as input (Line 1), as illustrated in Algorithm 2. We sample a latent variable xTx_{T} from the standard Gaussian distribution (Line 3) and upsample xLx_{L} with bicubic kernel (Line 4). Different from the training procedure, we encode the LR image xLx_{L} to xex_{e} by the LR encoder only once before the iteration begins (Line 5) and apply it in every iteration, which speeds up the inference. The iterations start from t=Tt=T (Line 6), and each iteration outputs a residual image with a different noise level, which gradually declines as tt decreases. For t>1t>1, we sample zz from standard Gaussian distribution (Line 7) and compute xt−1x_{t-1} using the noise predictor ϵθ\epsilon_{\theta} with xtx_{t}, xex_{e} and tt as inputs (Line 8). Then for t=1t=1, we set z=0z=0 (Line 6) and x0x_{0} is the final residual prediction (Line 10). An SR image is recovered by adding the residual image x0x_{0} on the upsampled LR image up(xL)up(x_{L}).

Experiments

In this section, we first describe the experimental settings including datasets, model configurations and details in training and inference. Then we report experimental results and conduct some analyses.

SRDiff is trained and evaluated on face SR (8×8\times) and general SR (4×4\times) tasks. For face SR, we use CelebFaces Attributes Dataset (CelebA) Liu et al. (2015b), which is a large-scale face attributes dataset with more than 200K celebrity images. The images in this dataset cover large pose variations and background clutter. In this paper, we use the whole training set which consists of 162,770 images for training and evaluate on 5000 images from the test split following SRFlow Lugmayr et al. (2020). We central-crop the aligned patcheshttps://drive.google.com/drive/folders/0B7EVK8r0v71pWEZsZE9oNnFzTm8 and resize them to 160×160160\times 160 as HR ground truth using standard MATLAB bicubic kernel. To obtain the corresponding LR images, we downsample the HR images with bicubic kernel. For ProgFSR Kim et al. (2019), we use progressive bilinear kernel introduced in its original paper for a fair comparison.

For General SR, we use the DIV2K Timofte et al. (2018) and Flickr2K Timofte et al. (2018). These datasets consist of high-resolution RGB images with a large diversity of contents. In training, we use the whole training data (800 images) in DIV2K and whole images in Flickr2K (2650 images). Then we crop each image into patches with a size of 160×160160\times 160 as HR ground truth following SRFlow. To obtain the corresponding LR images, we downsample the HR images with bicubic kernel. For evaluation, we use the whole validation data (100 images) in DIV2K. We downsample the HR images with bicubic kernel to obtain the LR images and directly apply SISR methods on the LR images to obtain the SR predictions without cropping.

Model Configuration

Our SRDiff model consists of a 4-step conditional noise predictor and an LR encoder with multiple RRDB blocks. The number of channels cc in the first contracting step is set to 64. The numbers of RRDB blocks in the LR encoder are set to 8 and 15 for CelebA and DIV2K respectively and the channel size is set to 32. For diffusion process and reverse process, we set the diffusion step TT to 100100 and our noise schedule β1,...,βT\beta_{1},...,\beta_{T} follows Nichol et al. Nichol and Dhariwal (2021), which is proved to be beneficial for training. We also explore the model performance under different TT and cc in Section 5.3.

Training and Evaluation

Firstly, we pretrain the LR encoder D\mathcal{D} using an L1L_{1} loss for 100k iterations for the sake of efficiency. The training of the conditional noise predictor uses Eq. (6) as loss term and Adam Kingma and Ba (2014) as optimizer, with batch size 1616 and learning rate 2×10−42\times 10^{-4}, which is halved every 100100k steps. The entire SRDiff takes about 34/45 hours (300k/400k steps) to train on 1 GeForce RTX 2080Ti with 11GB memory for CelebA/DIV2K respectively.

Beside the well-known evaluation metrics PSNR and SSIM Wang et al. (2004), we also evaluate our SRDiff on LPIPS Zhang et al. (2018a), LR-PSNR Lugmayr et al. (2020) and the pixel standard deviation σ\sigma. LPIPS is recently introduced as a reference-based image quality evaluation metric, which computes the perceptual similarity between the ground truth and the SR images. LR-PSNR is computed as the PSNR between the downsampled SR image and the LR one indicating the consistency with the LR image. The pixel standard deviation σ\sigma indicates diversity in the SR output.

2 Performance

In this subsection, we evaluate SRDiff by comparing with several state-of-the-art SR methods on face SR (8×\times) and general SR (4×4\times) tasks. The detailed configurations of baseline models can be found in their original papers.

We compare SRDiff with RRDB Wang et al. (2018), ESRGAN Wang et al. (2018), ProgFSR Kim et al. (2019) and SRFlow (τ=0.8\tau=0.8) Lugmayr et al. (2020)Due to inconsistent patch size, we retrain all these baseline models from scratch on our pre-processed CelebA dataset with released codes. SRFlow uses the same patch size as our model, but we cannot obtain the same patch with its released image example, and therefore, we also have to re-train it.. RRDB is trained by only L1L_{1} loss and can be regarded as a PSNR-oriented method. The evaluation results are shown in Table 1, which reveals that SRDiff outperforms previous works in term of most of the evaluation metrics, and can generate high-quality and diverse SR images with strong LR-consistency. Specifically, 1) as shown in Figure 4, compared with PSNR-oriented methods, SRDiff reconstructs clearer textures, and compared with GAN-driven methods, SRDiff avoids artifacts and the results look more natural; and 2) as shown in Figure 1, SRDiff provides diverse and realistic SR predictions given only one LR input. Every SR prediction is a complete portrait of a human face with rich details and maintains consistency with the input LR image. Moreover, SRDiff uses fewer model parameters (12M) than SRFlow (40M) and only takes about 30 hours until converge as described in Section 5.1, while SRFlow needs 5 days, which demonstrates that SRDiff is training-efficient and can achieve comparable performance with a small model footprint since SRDiff does not impose any architectural constraints to guarantee bijection. Compared with GAN-driven methods, SRDiff does not need any extra module (e.g., discriminator) in training.

General SR

We also evaluate SRDiff on general SR (4×4\times) compared with EDSR Lim et al. (2017), RRDB, ESRGAN, RankSRGAN Zhang et al. (2019) and SRFlow (τ=0.9\tau=0.9) with their official released pretrained modelsExcept RRDB which is trained from scratch with L1L_{1} loss as that in Face SR.. As shown in Table 2, SRDiff achieves better quantitative results than the previous methods for most evaluation metrics (PSNR, SSIM and LR-PSNR) and comparable LPIPS, which reveals the effectiveness and great potential of our method. Figure 5 shows that SRDiff balances sharpness and naturalness well and produces strong consistency with the LR image. In contrast, PSNR-oriented methods (EDSR and RRDB) and SRFlow, smear the edges of the objects, and GAN-driven methods (ESRGAN and RankSRGAN) introduce more artifacts.

3 Ablation Study

To probe the influences of the total diffusion step TT, noise predictor channel size cc and the effectiveness of residual prediction, we conduct some ablation studies as illustrated in Table 3. From row 1, 2, 3 and 4, we can see that the image quality is enhanced as total diffusion steps increases. From row 1, 5 and 6, we can see that larger model width results in better performance. However, larger total diffusion steps and model width both lead to slower inference, and therefore, we choose T=100T=100 and c=64c=64 as the default setting after trading off. Row 1 and 7 indicate that the residual prediction not only greatly enhances the image quality but also speeds up the training, which demonstrates the effectiveness of residual prediction.

4 Extensions

In this subsection, we explore some extended applications including content fusion and latent space interpolation.

SRDiff is applicable in content fusion tasks, which aim to generate an image by fusing contents from two source images, e.g., an eye-source image and a face-source image providing the eye and face contents respectively. In this paragraph, we use SRDiff to conduct face content fusion by a demonstration of fusing one’s eyes with another one’s face. The procedure of content fusion is shown in Algorithm 3. First, we directly fuse a face image xfx_{f} by replacing the eye region of the face-source image xfacex_{face} with that of the eye-source image xeyex_{eye} and compute the differences between xfx_{f} and the upsampled LR face-source image up(xL)up(x_{L}) to get the residual xrx_{r}. Second, xrx_{r} goes through a tˉ\bar{t}-step diffusion process, which outputs the xtˉx_{\bar{t}} in latent space. Then xtˉx_{\bar{t}} is denoised to an HR residual using the conditional noise predictor iterated from tˉ\bar{t} to with the LR face-source information encoded by the LR encoder, which ensures the compatibility of the two contents. Then we get the fused SR image by adding the SR residual to up(xL)up(x_{L}). Finally, we replace the eye region of the face-source image with that of the SR face image and preserve the non-manipulated face. As shown in Figure 6(a), we set different timesteps t∈{30,50,70}t\in\{30,50,70\} and find that the eye region of the fusion result is more similar to the eye-source image when tt is small, and is closer to the face-source image as tt becomes larger.

Latent Space Interpolation

Given an LR image, SRDiff can manipulate its prediction by latent space interpolation, which linearly interpolates the latents of two SR predictions and generate a new one. Let xtˉ,xtˉ′∼q(xtˉ∣x0)x_{\bar{t}},x^{\prime}_{\bar{t}}\sim q(x_{\bar{t}}|x_{0}) and we decode the latent xˉtˉ=λxtˉ+(1−λ)xtˉ′\bar{x}_{\bar{t}}=\lambda x_{\bar{t}}+(1-\lambda)x^{\prime}_{\bar{t}} by the reverse process, which feeds xˉtˉ\bar{x}_{\bar{t}} into the noise predictor with the LR information encoded by LR encoder iteratively. Then, we add the output residual result to the up(xL)up(x_{L}) to obtain the interpolated SR prediction. We set tˉ=50\bar{t}=50 and λ∈{0.0,0.4,0.8,1.0}\lambda\in\{0.0,0.4,0.8,1.0\}. It could be observed in Figure 6(b) that with λ\lambda approaching to 1.0, the woman’s expression becomes closer to xtx_{t}, which is the top right image holding a big laugh. In the same way, the man’s mouth turns wider and bigger from λ=0.0\lambda=0.0 to λ=1.0\lambda=1.0. The trend of the interpolated images shows the effectiveness of SRDiff in latent space interpolation. The detailed algorithm of latent space interpolation is shown in Algorithm 4.

Conclusion

In this paper, we proposed SRDiff, which is the first diffusion-based model for single image super-resolution to the best of our knowledge. Our work exploits a Markov chain to convert HR images to latents in simple distribution and then generate SR predictions in the reverse process which iteratively denoises the latents using a noise predictor conditioned on LR information encoded by the LR encoder. To speed up convergence and stabilize training, SRDiff introduces residual prediction. Our extensive experiments on both face and general datasets demonstrate that SRDiff can generate diverse and realistic SR images and avoids over-smoothing and mode collapse issues that occurred in PSNR-oriented methods and GAN-driven methods respectively. Moreover, SRDiff is stable to train with small footprint and without an extra discriminator. Besides, SRDiff allows for flexible image manipulation including latent space interpolation and content fusion.

In the future, we will further improve the performance of the diffusion-based SISR model and speed up the inference. We will also extend our work to more image restoration tasks (e.g., image denoising, deblurring and dehazing) to verify the potential of diffusion models in the image restoration domain.

References