DR2: Diffusion-based Robust Degradation Remover for Blind Face Restoration

Zhixin Wang, Xiaoyun Zhang, Ziying Zhang, Huangjie Zheng, Mingyuan Zhou, Ya Zhang, Yanfeng Wang

Introduction

Blind face restoration aims to restore high-quality face images from their low-quality counterparts suffering from unknown degradation, such as low-resolution , blur , noise, compression , etc. Great improvement in restoration quality has been witnessed over the past few years with the exploitation of various facial priors. Geometric priors such as facial landmarks , parsing maps , and heatmaps are pivotal to recovering the shapes of facial components. Reference priors of high-quality images are used as guidance to improve details. Recent research investigates generative priors and high-quality dictionaries , which help to generate photo-realistic details and textures.

Despite the great progress in visual quality, these methods lack a robust mechanism to handle degraded inputs besides relying on pre-defined degradation to synthesize the training data. When applying them to images of severe or unseen degradation, undesired results with obvious artifacts can be observed. As shown in Fig. 1, artifacts typically appear when 1) the input image lacks high-frequency information due to downsampling or blur (1st1^{st} row), in which case restoration networks can not generate adequate information, or 2) the input image bears corrupted high-frequency information due to noise or other degradation (2nd2^{nd} row), and restoration networks mistakenly use the corrupted information for restoration. The primary cause of this inadaptability is the inconsistency between the synthetic degradation of training data and the actual degradation in the real world.

Expanding the synthetic degradation model for training would improve the models’ adaptability but it is apparently difficult and expensive to simulate every possible degradation in the real world. To alleviate the dependency on synthetic degradation, we leverage a well-performing denoising diffusion probabilistic model (DDPM) to remove the degradation from inputs. DDPM generates images through a stochastic iterative denoising process and Gaussian noisy images can provide guidance to the generative process . As shown in Fig. 2, noisy images are degradation-irrelevant conditions for DDPM generative process. Adding extra Gaussian noise (right) makes different degradation less distinguishable compared with the original distribution (left), while DDPM can still capture the semantic information within this noise status and recover clean face images. This property of pretrained DDPM makes it a robust degradation removal module though only high-quality face images are used for training the DDPM.

Our overall blind face restoration framework DR2E consists of the Diffusion-based Robust Degradation Remover (DR2) and an Enhancement module. In the first stage, DR2 first transforms the degraded images into coarse, smooth, and visually clean intermediate results, which fall into a degradation-invariant distribution (4th4^{th} column in Fig. 1). In the second stage, the degradation-invariant images are further processed by the enhancement module for high-quality details. By this design, the enhancement module is compatible with various designs of restoration methods in seeking the best restoration quality, ensuring our DR2E achieves both strong robustness and high quality.

We summarize the contributions as follows. (1) We propose DR2 that leverages a pretrained diffusion model to remove degradation, achieving robustness against complex degradation without using synthetic degradation for training. (2) Together with an enhancement module, we employ DR2 in a two-stage blind face restoration framework, namely DR2E. The enhancement module has great flexibility in incorporating a variety of restoration methods to achieve high restoration quality. (3) Comprehensive studies and experiments show that our framework outperforms state-of-the-art methods on heavily degraded synthetic and real-world datasets.

Related Work

Blind Face Restoration Based on face hallucination or face super-resolution , blind face restoration aims to restore high-quality faces from low-quality images with unknown and complex degradation. Many facial priors are exploited to alleviate dependency on degraded inputs. Geometry priors, including facial landmarks , parsing maps , and facial component heatmaps help to recover accurate shapes but contain no information on details in themselves. Reference priors of high-quality images are used to recover details or preserve identity. To further boost restoration quality, generative priors like pretrained StyleGAN are used to provide vivid textures and details. PULSE uses latent optimization to find latent code of high-quality face, while more efficiently, GPEN , GFP-GAN , and GLEAN embed generative priors into the encoder-decoder structure. Another category of methods utilizes pretrained Vector-Quantize codebooks. DFDNet suggests constructing dictionaries of each component (e.g. eyes, mouth), while recent VQFR and CodeFormer pretrain high-quality dictionaries on entire faces, acquiring rich expressiveness.

Diffusion Models Denoising Diffusion Probabilistic Models (DDPM) are a fast-developing class of generative models in unconditional image generation rivaling Generative Adversarial Networks (GAN) . Recent research utilizes it for super-resolution. SR3 modifies DDPM to be conditioned on low-resolution images through channel-wise concatenation. However, it fixes the degradation to simple downsampling and does not apply to other degradation settings. Latent Diffusion performs super-resolution in a similar concatenation manner but in a low-dimensional latent space. ILVR proposes a conditioning method to control the generative process of pretrained DDPM for image-translation tasks. Diffusion-based methods face a common problem of slow sampling speed, while our DR2E adopts a hybrid architecture like to speed up the sampling process.

Methodology

Our proposed DR2E framework is depicted in Fig. 3, which consists of the degradation remover DR2 and an enhancement module. Given an input image y{\bm{y}} suffering from unknown degradation, diffused low-quality information yt−1{\bm{y}}_{t-1} is provided to refine the generative process. As a result, DR2 recovers a coarse result x^0\hat{{\bm{x}}}_{0} that is semantically close to y{\bm{y}} and degradation-invariant. Then the enhancement module maps x^0\hat{{\bm{x}}}_{0} to the final output with higher resolution and high-quality details.

Denoising Diffusion Probabilistic Models (DDPM) are a class of generative models that first pre-defines a variance schedule {β1,β2,...,βT}\{\beta_{1},\beta_{2},...,\beta_{T}\} to progressively corrupt an image x0{\bm{x}}_{0} to a noisy status through forward (diffusion) process:

Moreover, based on the property of the Markov chain, for any intermediate timestep t∈{1,2,...,T}t\in\{1,2,...,T\}, the corresponding noisy distribution has an analytic form:

where αˉt:=∏s=1t(1−βs)\bar{\alpha}_{t}:=\prod_{s=1}^{t}(1-\beta_{s}) and ϵ∼N(0,I){\bm{\epsilon}}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}). Then xT∼N(0,I){\bm{x}}_{T}\sim{\mathcal{N}}({\bm{0}},{\bm{I}}) if TT is big enough, usually T=1000T=1000.

The model progressively generates images by reversing the forward process. The generative process is also a Gaussian transition with the learned mean μθ{\bm{\mu}}_{\theta}:

where σt\sigma_{t} is usually a pre-defined constant related to the variance schedule, and μθ(xt,t){\bm{\mu}}_{\theta}({\bm{x}}_{t},t) is usually parameterized by a denoising U-Net ϵθ(xt,t){\bm{\epsilon}}_{\theta}({\bm{x}}_{t},t) with the following equivalence:

2 Framework Overview

Suppose the low-quality image y{\bm{y}} is degraded from the high-quality ground truth x∼X(x){\bm{x}}\sim\mathcal{X}({\bm{x}}) as y=T(x,z){\bm{y}}={\mathcal{T}}({\bm{x}},{\bm{z}}) where z{\bm{z}} describes the degradation model. Previous studies constructs the inverse function T−1(⋅,z){\mathcal{T}}^{-1}(\cdot,{\bm{z}}) by modeling p(x∣y,z)p({\bm{x}}|{\bm{y}},{\bm{z}}) with a pre-defined z{\bm{z}} . It meets the adaptation problem when actual degradation z′{\bm{z}}^{\prime} in the real world is far from z{\bm{z}}.

To overcome this challenge, we propose to model p(x∣y)p({\bm{x}}|{\bm{y}}) without a known z{\bm{z}} by a two-stage framework: it first removes degradation from inputs and get x^0\hat{x}_{0}, then maps degradation-invariant x^0\hat{x}_{0} to high-quality outputs. Our target is to maximize the likelihood:

pϕ(x^0∣y)p_{\bm{\phi}}(\hat{{\bm{x}}}_{0}|{\bm{y}}) corresponds to the degradation removal module, and pψ(x∣x^0)p_{\bm{\psi}}({\bm{x}}|\hat{{\bm{x}}}_{0}) corresponds to the enhancement module. For the first stage, instead of directly learning the mapping from y{\bm{y}} to x^0\hat{{\bm{x}}}_{0} which usually involves a pre-defined degradation model z{\bm{z}}, we come up with an important assumption and propose a diffusion-based method to remove degradation.

Assumption. For the diffusion process defined in Eq. 2, (1) there exists an intermediate timestep τ\tau such that for t>τt>\tau, the distance between q(xt∣x)q({\bm{x}}_{t}|{\bm{x}}) and q(yt∣y)q({\bm{y}}_{t}|{\bm{y}}) is close especially in the low-frequency part; (2) there exists ω>τ\omega>\tau such that the distance between q(xω∣x)q({\bm{x}}_{\omega}|{\bm{x}}) and q(yω∣y)q({\bm{y}}_{\omega}|{\bm{y}}) is eventually small enough, satisfying q(xω∣x)≈q(yω∣y){q({\bm{x}}_{\omega}|{\bm{x}})}\approx{q({\bm{y}}_{\omega}|{\bm{y}})}.

Note this assumption is not strong, as paired x{\bm{x}} and y{\bm{y}} would share similar low-frequency contents, and for sufficiently large t≈Tt\approx T, q(xt∣x)q({\bm{x}}_{t}|{\bm{x}}) and q(yt∣y)q({\bm{y}}_{t}|{\bm{y}}) are naturally close to the standard N(0,I){\mathcal{N}}({\bm{0}},{\bm{I}}). This assumption is also qualitatively justified in Fig. 2. Intuitively, if x{\bm{x}} and y{\bm{y}} are close in distribution (implying mild degradation), we can find ω\omega and τ\tau in a relatively small value and vice versa.

Then we rewrite the objective of the degradation removal module by applying the assumption q(xω∣x)≈q(yω∣y){q({\bm{x}}_{\omega}|{\bm{x}})}\approx{q({\bm{y}}_{\omega}|{\bm{y}})}:

By replacing variable from yω{\bm{y}}_{\omega} to xω{\bm{x}}_{\omega}, Eq. 7 and Eq. 8 naturally yields a DDPM model that denoises xω{\bm{x}}_{\omega} back to xτ{\bm{x}}_{\tau}, and we can further predict x^0\hat{{\bm{x}}}_{0} by the reverse of Eq. 2. x^0\hat{{\bm{x}}}_{0} would maintain semantics with x{\bm{x}} if proper conditioning methods like is adopted. So by leveraging a DDPM, we propose Diffusion-based Robust Degradation Remover (DR2) according to Eq. 6.

3 Diffusion-based Robust Degradation Remover

Consider a pretrained DDPM pθ(xt−1∣xt)p_{\theta}({\bm{x}}_{t-1}|{\bm{x}}_{t}) (Eq. 3) with a denoising U-Net ϵθ(xt,t){\bm{\epsilon}}_{\theta}({\bm{x}}_{t},t) pretrained on high-quality face dataset. We respectively implement q(yω∣y)q({\bm{y}}_{\omega}|{\bm{y}}), pθ(xτ∣yω)p_{\theta}({\bm{x}}_{\tau}|{\bm{y}}_{\omega}) and p(x^0∣xτ)p(\hat{{\bm{x}}}_{0}|{\bm{x}}_{\tau}) in Eq. 6 by three steps in below.

(1) Initial Condition at ω\mathbf{\omega}. We first “forward” the degraded image y{\bm{y}} to an initial condition yω{\bm{y}}_{\omega} by sampling from Eq. 2 and use it as xω{\bm{x}}_{\omega}:

ω∈{1,2,...,T}\omega\in\{1,2,...,T\}. This corresponds to q(xω∣y)q({\bm{x}}_{\omega}|{\bm{y}}) in Eq. 6. Then the DR2 denoising process starts at step ω\omega. This reduces the samplings steps and helps to speed up as well.

(2) Iterative Refinement. After each transition from xt{\bm{x}}_{t} to xt−1{\bm{x}}_{t-1} (τ+1⩽t⩽ω\tau+1\leqslant t\leqslant\omega), we sample yt−1{\bm{y}}_{t-1} from y{\bm{y}} through Eq. 2. Based on Assumption (1), we replace the low-frequency part of xt−1{\bm{x}}_{t-1} with that of yt−1{\bm{y}}_{t-1} because they are close in distribution, which is fomulated as:

where ΦN(⋅)\Phi_{N}(\cdot) denotes a low-pass filter implemented by downsampling and upsampling the image with a sharing scale factor NN. We drop the high-frequency part of y{\bm{y}} for it contains little information due to degradation. Unfiltered degradation that remained in the low-frequency part would be covered by the added noise. These conditional denoising steps correspond to pθ(xτ∣yω)p_{\bm{\theta}}({\bm{x}}_{\tau}|{\bm{y}}_{\omega}) in Eq. 6, which ensure the result shares basic semantics with yy.

Iterative refinement is pivotal for preserving the low-frequency information of the input images. With the iterative refinement, the choice of ω\omega and the randomness of Gaussian noise affect little to the result. We present ablation study in the supplementary for illustration.

(3) Truncated Output at τ\mathbf{\tau}. As tt gets smaller, the noise level gets milder and the distance between q(xt∣x)q({\bm{x}}_{t}|{\bm{x}}) and q(yt∣y)q({\bm{y}}_{t}|{\bm{y}}) gets larger. For small tt, the original degradation is more dominating in q(yt∣y)q({\bm{y}}_{t}|{\bm{y}}) than the added Gaussian noise. So the denoising process is truncated before tt is too small. We use predicted noise at step τ\tau (0<τ<ω)(0<\tau<\omega) to estimate the generation result as follows:

This corresponds to p(x^0∣xτ)p(\hat{{\bm{x}}}_{0}|{\bm{x}}_{\tau}) in Eq. 6. x^0\hat{{\bm{x}}}_{0} is the output of DR2, which maintains the basic semantics of y{\bm{y}} and is removed from various degradation.

Selection of NN and τ\tau. Downsampling factor NN and output step τ\tau have significant effects on the fidelity and “cleanness” of x^0\hat{{\bm{x}}}_{0}. We conduct ablation studies in Sec. 4.4 to show the effects of these two hyper-parameters. The best choices of NN and τ\tau are data-dependent. Generally speaking, big NN and τ\tau are more effective to remove the degradation but lead to lower fidelity. On the contrary, small NN and τ\tau leads to high fidelity, but may keep the degradation in the outputs. While ω\omega is empirically fixed to τ+0.25T\tau+0.25T.

4 Enhancement Module

With outputs of DR2, restoring the high-quality details only requires training an enhancement module pψ(x∣x^0)p_{\bm{\psi}}({\bm{x}}|\hat{{\bm{x}}}_{0}) (Eq. 5). Here we do not hypothesize about the specific method or architecture of this module. Any neural network that can be trained to map a low-quality image to its high-quality counterpart can be plugged in our framework. And the enhancement module is independently trained with its proposed loss functions.

Backbones. In practice, without loss of generality, we choose SPARNetHD that utilized no facial priors, and VQFR that pretrain a high-quality VQ codebook as two alternative backbones for our enhancement module to justify that it can be compatible with a broad choice of existing methods. We denote them as DR2 + SPAR and DR2 + VQFR respectively.

Training Data. Any pretrained blind face restoration models can be directly plugged-in without further finetuning, but in order to help the enhancement module adapt better and faster to DR2 outputs, we suggest constructing training data for the enhancement module using DR2 as follows:

Given a high-quality image x{\bm{x}}, we first use DR2 to reconstruct itself with controlling parameters (N,τ)(N,\tau) then convolve it with an Gaussian blur kernel kσk_{\sigma}. This helps the enhancement module adapt better and faster to DR2 outputs, which is recommended but not compulsory. Noting that beside this augmentation, no other degradation model is required in the training process as what previous works do by using Eq. 13.

Experiments

Implementation. DR2 and the enhancement module are independently trained on FFHQ dataset , which contains 70,000 high-quality face images. We use pretrained DDPM proposed by for our DR2. As introduced in Sec. 3.4, we choose SPARNetHD and VQFR as two alternative architectures for the enhancement module. We train SPARNetHD backbone from scratch with training data constructed by Eq. 12. We set N=4N=4 and randomly sample τ\tau, σ\sigma from {50,100,150,200}\{50,100,150,200\}, {1:7}\{1:7\}, respectively. As for VQFR backbone, we use its official pretrained model.

Testing Datasets. We construct one synthetic dataset and four real-world datasets for testing. A brief introduction of each is as followed:

∙\bullet CelebA-Test. Following previous works , we adopt a commonly used degradation model as follows to synthesize testing data from CelebA-HQ :

A high-quality image x{\bm{x}} is first convolved with a Gaussian blur kernel kσk_{\sigma}, then bicubically downsampled with a scale factor rr. nδn_{\delta} represents additive noise and is randomly chosen from Gaussian, Laplace, and Poisson. Finally, JPEG compression with quality qq is applied. We use rr = 16, 8, and 4 to form three restoration tasks denoted as 16×\mathbf{16\times}, 8×\mathbf{8\times}, and 4×\mathbf{4\times}. For each upsampling factor, we generate three splits with different levels of degradation and each split contains 1,000 images. The mild split randomly samples σ\sigma, δ\delta and qq from {3:5}\{3:5\}, {5:20}\{5:20\}, {60:80}\{60:80\}, respectively. The medium from {5:7}\{5:7\}, {15:40}\{15:40\}, {40:60}\{40:60\}. And the severe split from {7:9}\{7:9\}, {25:50}\{25:50\}, {30:40}\{30:40\}.

∙\bullet WIDER-Normal and WIDER-Critical. We select 400 critical cases suffering from heavy degradation (mainly low-resolution) from WIDER-face dataset to form the WIDER-Critical dataset and another 400 regular cases for WIDER-Normal dataset.

∙\bullet CelebChild contains 180 child faces of celebrities collected from the Internet. Most of them are only mildly degraded.

∙\bullet LFW-Test. LFW contains low-quality images with mild degradation from the Internet. We choose 1,000 testing images of different identities.

During testing, we conduct grid search for best controlling parameters (N,τ)(N,\tau) of DR2 for each dataset. Detailed parameter settings are presented in the suplementary.

2 Comparisons with State-of-the-art Methods

We compare our method with several state-of-the-art face restoration methods: DFDNet , SPARNetHD , GFP-GAN , GPEN , VQFR , and Codeformer . We adopt their official codes and pretrained models.

For evaluation, we adopt pixel-wise metrics (PSNR and SSIM) and the perceptual metric (LPIPS ) for the CelebA-Test with ground truth. We also employ the widely-used non-reference perceptual metric FID .

Synthetic CelebA-Test. For each upsampling factor, we calculate evaluation metrics on three splits and present the average in Tab. 1. For 16×16\times and 8×8\times upsampling tasks where degradation is severe due to low resolution, DR2 + VQFR and DR2 + SPAR achieve the best and the second-best LPIPS and FID scores, indicating our results are perceptually close to the ground truth. Noting that DR2 + VQFR is better at perceptual metrics (LPIPS and FID) thanks to the pretrained high-quality codebook, and DR2 + SPAR is better at pixel-wise metrics (PSNR and SSIM) because without facial priors, the outputs have higher fidelity to the inputs. For 4×4\times upsampling task where degradation is relatively milder, previous methods trained on similar synthetic degradation manage to produce high-quality images without obvious artifacts. But our methods still obtain superior FID scores, showing our outputs have closer distribution to ground truth on different settings.

Qualitative comparisons from are presented in Fig. 4. Our methods produce fewer artifacts on severely degraded inputs compared with previous methods.

Real-World Datasets. We evaluate FID scores on different real-world datasets and present quantitative results in Tab. 2. On severely degraded dataset WIDER-Critical, our DR2 + VQFR and DR2 + SPAR achieve the best and the second best FID. On other datasets with only mild degradation, the restoration quality rather than robustness becomes the bottleneck, so DR2 + SPAR with no facial priors struggles to stand out, while DR2 + VQFR still achieves the best performance.

Qualitative results on WIDER-Critical are shown in Fig. 5. When input images’ resolutions are very low, previous methods fail to complement adequate information for pleasant faces, while our outputs are visually more pleasant thanks to the generative ability of DDPM.

3 Comparisons with Diffusion-based Methods

Diffusion-based super-resolution methods can be grouped into two categories by whether feeding auxiliary input to the denoising U-Net.

SR3 typically uses the concatenation of low-resolution images and xt{\bm{x}}_{t} as the input of the denoising U-Net. But SR3 fixes degradation to bicubic downsampling during training, which makes it highly degradation-sensitive. For visual comparisons, we re-implement the concatenation-based method based on . As shown in Fig. 6, minor noise in the second input evidently harm the performance of this concatenation-based method. Eventually, this type of method would rely on synthetic degradation to improve robustness like , while our DR2 have good robustness against different degradation without training on specifically degraded data.

Another category of methods is training-free, exploiting pretrained diffusion methods like ILVR . It shows the ability to transform both clean and degraded low-resolution images into high-resolution outputs. However, relying solely on ILVR for blind face restoration faces the trade-off problem between fidelity and quality (realness). As shown in Fig. 7, ILVR Sample 1 has high fidelity to input but low visual quality because the conditioning information is over-used. On the contrary, under-use of conditions leads to high quality but low fidelity as ILVR Sample 2. In our framework, fidelity is controlled by DR2 and high-quality details are restored by the enhancement module, thus alleviating the trade-off problem.

4 Effect of Different N𝑁N and τ𝜏\tau

In this section, we explore the property of DR2 output in terms of the controlling parameter (N,τ)(N,\tau) so that we can have a better intuitions for choosing appropriate parameters for variant input data. To avoid the influence of the enhancement modules varying in structures, embedded facial priors, and training strategies, we only evaluate DR2 outputs with no enhancement.

In Fig. 8, DR2 outputs are generated with different combinations of NN and τ\tau. Bigger NN and τ\tau are effective to remove degradation but tent to make results deviant from the input. On the contrary, small NN and τ\tau lead to high fidelity, but may keep the degradation in outputs.

We provide quantitative evaluations on CelebA-Test (8×8\times, medium split) dataset in Fig. 9. With bicubically downsampled low-resolution images used as ground truth, we adopt pixel-wise metric (PSNR↑\uparrow) and identity distance (Deg↓\downarrow) based on the embedding angle of ArcFace for evaluating the quality and fidelity of DR2 outputs. For scale NN = 4, 8, and 16, PSNR first goes up and Deg goes down because degradation is gradually removed as τ\tau increases. Then they hit the optimal point at the same time before the outputs begin to deviate from the input as τ\tau continues to grow. Optimal τ\tau is bigger for smaller NN. For N=2N=2, PSNR stops to increase before Deg reaches the optimality because Gaussian noise starts to appear in the output (like results sampled with (N,τ)=(2,350)(N,\tau)=(2,350) in Fig. 8). This cause of the appearance of Gaussian noise is that yt{\bm{y}}_{t} sampled by Eq. 2 contains heavy Gaussian noise when tt (t>τ)(t>\tau) is big and most part of yt{\bm{y}}_{t} is utilized by Eq. 10 when NN is small.

5 Discussion and Limitations

Our DR2 is built on a pretrained DDPM, so it would face the problem of slow sampling speed even we only perform 0.25T0.25T steps in total. But DR2 can be combined with diffusion acceleration methods like inference every 10 steps. And keep the output resolution of DR2 relatively low (2562256^{2} in our practice) and leave the upsampling for enhancement module for faster speed.

Another major limitation of our proposed DR2 is the manual choosing for controlling parameters NN and τ\tau. As a future work, we are exploring whether image quality assessment scores (like NIQE) can be used to develop an automatic search algorithms for NN and τ\tau.

Furthermore, for inputs with slight degradation, DR2 is less necessary because previous methods can also be effective and faster. And in extreme cases where input images contains very slight degradation or even no degradation, DR2 transformation may remove details in the inputs, but that is not common cases for blind face restoration.

Conclusion

We propose the DR2E, a two-stage blind face restoration framework that leverages a pretrained DDPM to remove degradation from inputs, and an enhancement module for detail restoration. In the first stage, DR2 removes degradation by using diffused low-quality information as conditions to guide the generative process. This transformation requires no synthetically degraded data for training. Extensive comparisons demonstrate the strong robustness and high restoration quality of our DR2E framework.

Acknowledgements

This work is supported by National Natural Science Foundation of China (62271308), Shanghai Key Laboratory of Digital Media Processing and Transmissions (STCSM 22511105700, 18DZ2270700), 111 plan (BP0719010), and State Key Laboratory of UHD Video and Audio Production and Presentation.

References

Appendix

In the appendix, we provide additional discussions and results complementing Sec. 4. In Appendix A, we conduct further ablation studies on initial condition and iterative refinement to show the control and conditioning effect these two mechanisms bring to the DR2 generative process. In Appendix B, we provide (1) our detailed settings of DR2 controlling parameters (N,τ)(N,\tau) for each testing dataset, and (2) show more qualitative comparisons on each split of CelebA-Test dataset in this section to illustrate how our methods and previous state-of-the-art methods perform over variant levels of degradation.

Appendix A More Ablation Studies

In this section, we explore the effect of initial condition and iterative refinement in DR2. To avoid the influence of the enhancement modules varying in structures, embedded facial priors, and training strategies, we only conduct experiments on DR2 outputs with no enhancement. To evaluate the degradation removal performance and fidelity of DR2 outputs, we use bicubic downsampled images as ground truth low-resolution (GT LR) image. This is intuitive as DR2 is targeted to produce clean but blurry middle results.

During DR2 generative process, diffused low-quality inputs is provided through initial condition and iterative refinement. The latter one yields stronger control to the generative process because it is performed at each step, while initial condition only provides information in the beginning with heavy Gaussian noise attached. To quantitatively evaluate the effect of initial condition, we follow the settings of SSec. 4.4 by calculating the pixel-wise metric (PSNR) and identity distance (Deg) between DR2 outputs and ground truth low-resolution images on CelebA-Test (8×8\times, medium split) dataset. Quantitative results are shown in Tab. A1. We fix (N,τ)=(4,300)(N,\tau)=(4,300) and change the value of ω\omega. When ω=1000=T\omega=1000=T, no initial condition is provided because y1000\mathbf{y}_{1000} is pure Gaussian noise. As shown in the table, with iterative refinement providing strong control to DR2 generative process, the quality and fidelity of DR2 outputs are not evidently affected as ω\omega varies.

Qualitative results are provided in Fig. A1. With fixed iterative refinement controlling parameters, ω\omega has little visual effect on DR2 outputs. Although the initial condition provides limited information compared with iterative refinement, it significantly reduces the total steps of DR2 denoising process.

A.2 Conditioning Effect of Initial Condition with Iterative Refinement Disabled

We conduct experiments without iterative refinement in this section to show that generative results bear less fidelity to the input without it. Without iterative refinement, DR2 generative process relies solely on the initial condition to utilize information of low-quality inputs, and generate images through DDPM denoising steps stochastically from initial condition. ω\omega now becomes an important controlling parameter determining how much conditioning information is provided. We also calculate PSNR and Deg between DR2 outputs and ground-truth low-resolution images on CelebA-Test (8×8\times, medium split) dataset. Quantitative results with different (ω,τ)(\omega,\tau) are provided in Tab. A2. Note that PSNR and Deg are all worse than those in Tab. A1, and have a negative correlation with ω\omega because less information of inputs is used as ω\omega increases.

Qualitative results are shown in Fig. A2. When ω⩾400\omega\geqslant 400, added noise in initial condition is strong enough to cover the degradation in inputs so the output tends to be smooth and clean. But as ω\omega increases, the outputs become more irrelevant to the input because the initial conditions are weakened. Compared with results that were sampled with iterative refinement, the importance of it on preserving semantic information is obvious.

Appendix B Detailed Settings and Comparisons

As introduced in Sec. 4.1, to evaluate the performance on different levels of degradation, we synthesize three splits (mild, medium, and severe) for each upsampling task (16×16\times, 8×8\times, and 4×4\times) together with four real-world datasets. During the experiment in Sec. 4.2, different controlling parameters (N,τ)(N,\tau) are used for each dataset or split. Generally speaking, big NN and τ\tau are more effective to remove the degradation but lead to lower fidelity and vice versa. We provide detailed settings we employed in Tab. A3

B.2 More Qualitative Comparisons

For more comprehensive comparisons with previous methods on different levels of degraded dataset, we provide qualitative results on each split of CelebA-Test dataset under each upsampling factor in Figs. A1, A2 and A3. As shown in the figures, for inputs with slight degradation, DR2 transformation is less necessary because previous methods can also be effective. But for severe degradation, previous methods fail since they never see such degradation during training. While our method shows great robustness even though no synthetic degraded images are employed for training.