DR2: Diffusion-based Robust Degradation Remover for Blind Face Restoration
Zhixin Wang, Xiaoyun Zhang, Ziying Zhang, Huangjie Zheng, Mingyuan Zhou, Ya Zhang, Yanfeng Wang
Introduction
Blind face restoration aims to restore high-quality face images from their low-quality counterparts suffering from unknown degradation, such as low-resolution , blur , noise, compression , etc. Great improvement in restoration quality has been witnessed over the past few years with the exploitation of various facial priors. Geometric priors such as facial landmarks , parsing maps , and heatmaps are pivotal to recovering the shapes of facial components. Reference priors of high-quality images are used as guidance to improve details. Recent research investigates generative priors and high-quality dictionaries , which help to generate photo-realistic details and textures.
Despite the great progress in visual quality, these methods lack a robust mechanism to handle degraded inputs besides relying on pre-defined degradation to synthesize the training data. When applying them to images of severe or unseen degradation, undesired results with obvious artifacts can be observed. As shown in Fig. 1, artifacts typically appear when 1) the input image lacks high-frequency information due to downsampling or blur ( row), in which case restoration networks can not generate adequate information, or 2) the input image bears corrupted high-frequency information due to noise or other degradation ( row), and restoration networks mistakenly use the corrupted information for restoration. The primary cause of this inadaptability is the inconsistency between the synthetic degradation of training data and the actual degradation in the real world.
Expanding the synthetic degradation model for training would improve the models’ adaptability but it is apparently difficult and expensive to simulate every possible degradation in the real world. To alleviate the dependency on synthetic degradation, we leverage a well-performing denoising diffusion probabilistic model (DDPM) to remove the degradation from inputs. DDPM generates images through a stochastic iterative denoising process and Gaussian noisy images can provide guidance to the generative process . As shown in Fig. 2, noisy images are degradation-irrelevant conditions for DDPM generative process. Adding extra Gaussian noise (right) makes different degradation less distinguishable compared with the original distribution (left), while DDPM can still capture the semantic information within this noise status and recover clean face images. This property of pretrained DDPM makes it a robust degradation removal module though only high-quality face images are used for training the DDPM.
Our overall blind face restoration framework DR2E consists of the Diffusion-based Robust Degradation Remover (DR2) and an Enhancement module. In the first stage, DR2 first transforms the degraded images into coarse, smooth, and visually clean intermediate results, which fall into a degradation-invariant distribution ( column in Fig. 1). In the second stage, the degradation-invariant images are further processed by the enhancement module for high-quality details. By this design, the enhancement module is compatible with various designs of restoration methods in seeking the best restoration quality, ensuring our DR2E achieves both strong robustness and high quality.
We summarize the contributions as follows. (1) We propose DR2 that leverages a pretrained diffusion model to remove degradation, achieving robustness against complex degradation without using synthetic degradation for training. (2) Together with an enhancement module, we employ DR2 in a two-stage blind face restoration framework, namely DR2E. The enhancement module has great flexibility in incorporating a variety of restoration methods to achieve high restoration quality. (3) Comprehensive studies and experiments show that our framework outperforms state-of-the-art methods on heavily degraded synthetic and real-world datasets.
Related Work
Blind Face Restoration Based on face hallucination or face super-resolution , blind face restoration aims to restore high-quality faces from low-quality images with unknown and complex degradation. Many facial priors are exploited to alleviate dependency on degraded inputs. Geometry priors, including facial landmarks , parsing maps , and facial component heatmaps help to recover accurate shapes but contain no information on details in themselves. Reference priors of high-quality images are used to recover details or preserve identity. To further boost restoration quality, generative priors like pretrained StyleGAN are used to provide vivid textures and details. PULSE uses latent optimization to find latent code of high-quality face, while more efficiently, GPEN , GFP-GAN , and GLEAN embed generative priors into the encoder-decoder structure. Another category of methods utilizes pretrained Vector-Quantize codebooks. DFDNet suggests constructing dictionaries of each component (e.g. eyes, mouth), while recent VQFR and CodeFormer pretrain high-quality dictionaries on entire faces, acquiring rich expressiveness.
Diffusion Models Denoising Diffusion Probabilistic Models (DDPM) are a fast-developing class of generative models in unconditional image generation rivaling Generative Adversarial Networks (GAN) . Recent research utilizes it for super-resolution. SR3 modifies DDPM to be conditioned on low-resolution images through channel-wise concatenation. However, it fixes the degradation to simple downsampling and does not apply to other degradation settings. Latent Diffusion performs super-resolution in a similar concatenation manner but in a low-dimensional latent space. ILVR proposes a conditioning method to control the generative process of pretrained DDPM for image-translation tasks. Diffusion-based methods face a common problem of slow sampling speed, while our DR2E adopts a hybrid architecture like to speed up the sampling process.
Methodology
Our proposed DR2E framework is depicted in Fig. 3, which consists of the degradation remover DR2 and an enhancement module. Given an input image suffering from unknown degradation, diffused low-quality information is provided to refine the generative process. As a result, DR2 recovers a coarse result that is semantically close to and degradation-invariant. Then the enhancement module maps to the final output with higher resolution and high-quality details.
Denoising Diffusion Probabilistic Models (DDPM) are a class of generative models that first pre-defines a variance schedule to progressively corrupt an image to a noisy status through forward (diffusion) process:
Moreover, based on the property of the Markov chain, for any intermediate timestep , the corresponding noisy distribution has an analytic form:
where and . Then if is big enough, usually .
The model progressively generates images by reversing the forward process. The generative process is also a Gaussian transition with the learned mean :
where is usually a pre-defined constant related to the variance schedule, and is usually parameterized by a denoising U-Net with the following equivalence:
2 Framework Overview
Suppose the low-quality image is degraded from the high-quality ground truth as where describes the degradation model. Previous studies constructs the inverse function by modeling with a pre-defined . It meets the adaptation problem when actual degradation in the real world is far from .
To overcome this challenge, we propose to model without a known by a two-stage framework: it first removes degradation from inputs and get , then maps degradation-invariant to high-quality outputs. Our target is to maximize the likelihood:
corresponds to the degradation removal module, and corresponds to the enhancement module. For the first stage, instead of directly learning the mapping from to which usually involves a pre-defined degradation model , we come up with an important assumption and propose a diffusion-based method to remove degradation.
Assumption. For the diffusion process defined in Eq. 2, (1) there exists an intermediate timestep such that for , the distance between and is close especially in the low-frequency part; (2) there exists such that the distance between and is eventually small enough, satisfying .
Note this assumption is not strong, as paired and would share similar low-frequency contents, and for sufficiently large , and are naturally close to the standard . This assumption is also qualitatively justified in Fig. 2. Intuitively, if and are close in distribution (implying mild degradation), we can find and in a relatively small value and vice versa.
Then we rewrite the objective of the degradation removal module by applying the assumption :
By replacing variable from to , Eq. 7 and Eq. 8 naturally yields a DDPM model that denoises back to , and we can further predict by the reverse of Eq. 2. would maintain semantics with if proper conditioning methods like is adopted. So by leveraging a DDPM, we propose Diffusion-based Robust Degradation Remover (DR2) according to Eq. 6.
3 Diffusion-based Robust Degradation Remover
Consider a pretrained DDPM (Eq. 3) with a denoising U-Net pretrained on high-quality face dataset. We respectively implement , and in Eq. 6 by three steps in below.
(1) Initial Condition at . We first “forward” the degraded image to an initial condition by sampling from Eq. 2 and use it as :
. This corresponds to in Eq. 6. Then the DR2 denoising process starts at step . This reduces the samplings steps and helps to speed up as well.
(2) Iterative Refinement. After each transition from to (), we sample from through Eq. 2. Based on Assumption (1), we replace the low-frequency part of with that of because they are close in distribution, which is fomulated as:
where denotes a low-pass filter implemented by downsampling and upsampling the image with a sharing scale factor . We drop the high-frequency part of for it contains little information due to degradation. Unfiltered degradation that remained in the low-frequency part would be covered by the added noise. These conditional denoising steps correspond to in Eq. 6, which ensure the result shares basic semantics with .
Iterative refinement is pivotal for preserving the low-frequency information of the input images. With the iterative refinement, the choice of and the randomness of Gaussian noise affect little to the result. We present ablation study in the supplementary for illustration.
(3) Truncated Output at . As gets smaller, the noise level gets milder and the distance between and gets larger. For small , the original degradation is more dominating in than the added Gaussian noise. So the denoising process is truncated before is too small. We use predicted noise at step to estimate the generation result as follows:
This corresponds to in Eq. 6. is the output of DR2, which maintains the basic semantics of and is removed from various degradation.
Selection of and . Downsampling factor and output step have significant effects on the fidelity and “cleanness” of . We conduct ablation studies in Sec. 4.4 to show the effects of these two hyper-parameters. The best choices of and are data-dependent. Generally speaking, big and are more effective to remove the degradation but lead to lower fidelity. On the contrary, small and leads to high fidelity, but may keep the degradation in the outputs. While is empirically fixed to .
4 Enhancement Module
With outputs of DR2, restoring the high-quality details only requires training an enhancement module (Eq. 5). Here we do not hypothesize about the specific method or architecture of this module. Any neural network that can be trained to map a low-quality image to its high-quality counterpart can be plugged in our framework. And the enhancement module is independently trained with its proposed loss functions.
Backbones. In practice, without loss of generality, we choose SPARNetHD that utilized no facial priors, and VQFR that pretrain a high-quality VQ codebook as two alternative backbones for our enhancement module to justify that it can be compatible with a broad choice of existing methods. We denote them as DR2 + SPAR and DR2 + VQFR respectively.
Training Data. Any pretrained blind face restoration models can be directly plugged-in without further finetuning, but in order to help the enhancement module adapt better and faster to DR2 outputs, we suggest constructing training data for the enhancement module using DR2 as follows:
Given a high-quality image , we first use DR2 to reconstruct itself with controlling parameters then convolve it with an Gaussian blur kernel . This helps the enhancement module adapt better and faster to DR2 outputs, which is recommended but not compulsory. Noting that beside this augmentation, no other degradation model is required in the training process as what previous works do by using Eq. 13.
Experiments
Implementation. DR2 and the enhancement module are independently trained on FFHQ dataset , which contains 70,000 high-quality face images. We use pretrained DDPM proposed by for our DR2. As introduced in Sec. 3.4, we choose SPARNetHD and VQFR as two alternative architectures for the enhancement module. We train SPARNetHD backbone from scratch with training data constructed by Eq. 12. We set and randomly sample , from , , respectively. As for VQFR backbone, we use its official pretrained model.
Testing Datasets. We construct one synthetic dataset and four real-world datasets for testing. A brief introduction of each is as followed:
CelebA-Test. Following previous works , we adopt a commonly used degradation model as follows to synthesize testing data from CelebA-HQ :
A high-quality image is first convolved with a Gaussian blur kernel , then bicubically downsampled with a scale factor . represents additive noise and is randomly chosen from Gaussian, Laplace, and Poisson. Finally, JPEG compression with quality is applied. We use = 16, 8, and 4 to form three restoration tasks denoted as , , and . For each upsampling factor, we generate three splits with different levels of degradation and each split contains 1,000 images. The mild split randomly samples , and from , , , respectively. The medium from , , . And the severe split from , , .
WIDER-Normal and WIDER-Critical. We select 400 critical cases suffering from heavy degradation (mainly low-resolution) from WIDER-face dataset to form the WIDER-Critical dataset and another 400 regular cases for WIDER-Normal dataset.
CelebChild contains 180 child faces of celebrities collected from the Internet. Most of them are only mildly degraded.
LFW-Test. LFW contains low-quality images with mild degradation from the Internet. We choose 1,000 testing images of different identities.
During testing, we conduct grid search for best controlling parameters of DR2 for each dataset. Detailed parameter settings are presented in the suplementary.
2 Comparisons with State-of-the-art Methods
We compare our method with several state-of-the-art face restoration methods: DFDNet , SPARNetHD , GFP-GAN , GPEN , VQFR , and Codeformer . We adopt their official codes and pretrained models.
For evaluation, we adopt pixel-wise metrics (PSNR and SSIM) and the perceptual metric (LPIPS ) for the CelebA-Test with ground truth. We also employ the widely-used non-reference perceptual metric FID .
Synthetic CelebA-Test. For each upsampling factor, we calculate evaluation metrics on three splits and present the average in Tab. 1. For and upsampling tasks where degradation is severe due to low resolution, DR2 + VQFR and DR2 + SPAR achieve the best and the second-best LPIPS and FID scores, indicating our results are perceptually close to the ground truth. Noting that DR2 + VQFR is better at perceptual metrics (LPIPS and FID) thanks to the pretrained high-quality codebook, and DR2 + SPAR is better at pixel-wise metrics (PSNR and SSIM) because without facial priors, the outputs have higher fidelity to the inputs. For upsampling task where degradation is relatively milder, previous methods trained on similar synthetic degradation manage to produce high-quality images without obvious artifacts. But our methods still obtain superior FID scores, showing our outputs have closer distribution to ground truth on different settings.
Qualitative comparisons from are presented in Fig. 4. Our methods produce fewer artifacts on severely degraded inputs compared with previous methods.
Real-World Datasets. We evaluate FID scores on different real-world datasets and present quantitative results in Tab. 2. On severely degraded dataset WIDER-Critical, our DR2 + VQFR and DR2 + SPAR achieve the best and the second best FID. On other datasets with only mild degradation, the restoration quality rather than robustness becomes the bottleneck, so DR2 + SPAR with no facial priors struggles to stand out, while DR2 + VQFR still achieves the best performance.
Qualitative results on WIDER-Critical are shown in Fig. 5. When input images’ resolutions are very low, previous methods fail to complement adequate information for pleasant faces, while our outputs are visually more pleasant thanks to the generative ability of DDPM.
3 Comparisons with Diffusion-based Methods
Diffusion-based super-resolution methods can be grouped into two categories by whether feeding auxiliary input to the denoising U-Net.
SR3 typically uses the concatenation of low-resolution images and as the input of the denoising U-Net. But SR3 fixes degradation to bicubic downsampling during training, which makes it highly degradation-sensitive. For visual comparisons, we re-implement the concatenation-based method based on . As shown in Fig. 6, minor noise in the second input evidently harm the performance of this concatenation-based method. Eventually, this type of method would rely on synthetic degradation to improve robustness like , while our DR2 have good robustness against different degradation without training on specifically degraded data.
Another category of methods is training-free, exploiting pretrained diffusion methods like ILVR . It shows the ability to transform both clean and degraded low-resolution images into high-resolution outputs. However, relying solely on ILVR for blind face restoration faces the trade-off problem between fidelity and quality (realness). As shown in Fig. 7, ILVR Sample 1 has high fidelity to input but low visual quality because the conditioning information is over-used. On the contrary, under-use of conditions leads to high quality but low fidelity as ILVR Sample 2. In our framework, fidelity is controlled by DR2 and high-quality details are restored by the enhancement module, thus alleviating the trade-off problem.
4 Effect of Different N𝑁N and τ𝜏\tau
In this section, we explore the property of DR2 output in terms of the controlling parameter so that we can have a better intuitions for choosing appropriate parameters for variant input data. To avoid the influence of the enhancement modules varying in structures, embedded facial priors, and training strategies, we only evaluate DR2 outputs with no enhancement.
In Fig. 8, DR2 outputs are generated with different combinations of and . Bigger and are effective to remove degradation but tent to make results deviant from the input. On the contrary, small and lead to high fidelity, but may keep the degradation in outputs.
We provide quantitative evaluations on CelebA-Test (, medium split) dataset in Fig. 9. With bicubically downsampled low-resolution images used as ground truth, we adopt pixel-wise metric (PSNR) and identity distance (Deg) based on the embedding angle of ArcFace for evaluating the quality and fidelity of DR2 outputs. For scale = 4, 8, and 16, PSNR first goes up and Deg goes down because degradation is gradually removed as increases. Then they hit the optimal point at the same time before the outputs begin to deviate from the input as continues to grow. Optimal is bigger for smaller . For , PSNR stops to increase before Deg reaches the optimality because Gaussian noise starts to appear in the output (like results sampled with in Fig. 8). This cause of the appearance of Gaussian noise is that sampled by Eq. 2 contains heavy Gaussian noise when is big and most part of is utilized by Eq. 10 when is small.
5 Discussion and Limitations
Our DR2 is built on a pretrained DDPM, so it would face the problem of slow sampling speed even we only perform steps in total. But DR2 can be combined with diffusion acceleration methods like inference every 10 steps. And keep the output resolution of DR2 relatively low ( in our practice) and leave the upsampling for enhancement module for faster speed.
Another major limitation of our proposed DR2 is the manual choosing for controlling parameters and . As a future work, we are exploring whether image quality assessment scores (like NIQE) can be used to develop an automatic search algorithms for and .
Furthermore, for inputs with slight degradation, DR2 is less necessary because previous methods can also be effective and faster. And in extreme cases where input images contains very slight degradation or even no degradation, DR2 transformation may remove details in the inputs, but that is not common cases for blind face restoration.
Conclusion
We propose the DR2E, a two-stage blind face restoration framework that leverages a pretrained DDPM to remove degradation from inputs, and an enhancement module for detail restoration. In the first stage, DR2 removes degradation by using diffused low-quality information as conditions to guide the generative process. This transformation requires no synthetically degraded data for training. Extensive comparisons demonstrate the strong robustness and high restoration quality of our DR2E framework.
Acknowledgements
This work is supported by National Natural Science Foundation of China (62271308), Shanghai Key Laboratory of Digital Media Processing and Transmissions (STCSM 22511105700, 18DZ2270700), 111 plan (BP0719010), and State Key Laboratory of UHD Video and Audio Production and Presentation.
References
Appendix
In the appendix, we provide additional discussions and results complementing Sec. 4. In Appendix A, we conduct further ablation studies on initial condition and iterative refinement to show the control and conditioning effect these two mechanisms bring to the DR2 generative process. In Appendix B, we provide (1) our detailed settings of DR2 controlling parameters for each testing dataset, and (2) show more qualitative comparisons on each split of CelebA-Test dataset in this section to illustrate how our methods and previous state-of-the-art methods perform over variant levels of degradation.
Appendix A More Ablation Studies
In this section, we explore the effect of initial condition and iterative refinement in DR2. To avoid the influence of the enhancement modules varying in structures, embedded facial priors, and training strategies, we only conduct experiments on DR2 outputs with no enhancement. To evaluate the degradation removal performance and fidelity of DR2 outputs, we use bicubic downsampled images as ground truth low-resolution (GT LR) image. This is intuitive as DR2 is targeted to produce clean but blurry middle results.
During DR2 generative process, diffused low-quality inputs is provided through initial condition and iterative refinement. The latter one yields stronger control to the generative process because it is performed at each step, while initial condition only provides information in the beginning with heavy Gaussian noise attached. To quantitatively evaluate the effect of initial condition, we follow the settings of SSec. 4.4 by calculating the pixel-wise metric (PSNR) and identity distance (Deg) between DR2 outputs and ground truth low-resolution images on CelebA-Test (, medium split) dataset. Quantitative results are shown in Tab. A1. We fix and change the value of . When , no initial condition is provided because is pure Gaussian noise. As shown in the table, with iterative refinement providing strong control to DR2 generative process, the quality and fidelity of DR2 outputs are not evidently affected as varies.
Qualitative results are provided in Fig. A1. With fixed iterative refinement controlling parameters, has little visual effect on DR2 outputs. Although the initial condition provides limited information compared with iterative refinement, it significantly reduces the total steps of DR2 denoising process.
A.2 Conditioning Effect of Initial Condition with Iterative Refinement Disabled
We conduct experiments without iterative refinement in this section to show that generative results bear less fidelity to the input without it. Without iterative refinement, DR2 generative process relies solely on the initial condition to utilize information of low-quality inputs, and generate images through DDPM denoising steps stochastically from initial condition. now becomes an important controlling parameter determining how much conditioning information is provided. We also calculate PSNR and Deg between DR2 outputs and ground-truth low-resolution images on CelebA-Test (, medium split) dataset. Quantitative results with different are provided in Tab. A2. Note that PSNR and Deg are all worse than those in Tab. A1, and have a negative correlation with because less information of inputs is used as increases.
Qualitative results are shown in Fig. A2. When , added noise in initial condition is strong enough to cover the degradation in inputs so the output tends to be smooth and clean. But as increases, the outputs become more irrelevant to the input because the initial conditions are weakened. Compared with results that were sampled with iterative refinement, the importance of it on preserving semantic information is obvious.
Appendix B Detailed Settings and Comparisons
As introduced in Sec. 4.1, to evaluate the performance on different levels of degradation, we synthesize three splits (mild, medium, and severe) for each upsampling task (, , and ) together with four real-world datasets. During the experiment in Sec. 4.2, different controlling parameters are used for each dataset or split. Generally speaking, big and are more effective to remove the degradation but lead to lower fidelity and vice versa. We provide detailed settings we employed in Tab. A3
B.2 More Qualitative Comparisons
For more comprehensive comparisons with previous methods on different levels of degraded dataset, we provide qualitative results on each split of CelebA-Test dataset under each upsampling factor in Figs. A1, A2 and A3. As shown in the figures, for inputs with slight degradation, DR2 transformation is less necessary because previous methods can also be effective. But for severe degradation, previous methods fail since they never see such degradation during training. While our method shows great robustness even though no synthetic degraded images are employed for training.