DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior
Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Bo Dai, Fanghua Yu, Wanli Ouyang, Yu Qiao, Chao Dong
Introduction
Image restoration aims at reconstructing a high-quality image from its low-quality observation. Typical image restoration problems, such as image denoising, deblurring and super-resolution, are usually defined under a constrained setting, where the degradation process is simple and known (e.g., Gaussian noise and Bicubic downsampling). They have successfully promoted a vast number of excellent restoration algorithms [14; 65; 36; 7; 58; 63; 8], but are born to have limited generalization ability. To deal with real-world degraded images, blind image restoration (BIR) comes into view and becomes a promising direction. The ultimate goal of BIR is to realize realistic image reconstruction on general images with general degradations. BIR does not only extend the boundary of classic image restoration tasks, but also has a wide practical application field (e.g., old photo/film restoration).
The research of BIR is still in its primary stage, thus requiring more explanations of its current state. According to the problem settings, existing BIR methods can be roughly grouped into three research topics, namely blind image super-resolution (BSR), zero-shot image restoration (ZIR) and blind face restoration (BFR). They all have achieved remarkable progress, but also have apparent limitations. BSR is initially proposed to solve real-world super-resolution problems, where the low-resolution image contains unknown degradations. According to the recent BSR survey , the most popular solutions may be BSRGAN and Real-ESRGAN . They formulate BSR as a supervised large-scale degradation overfitting problem. To simulate real-world degradations, a degradation shuffle strategy and high-order degradation modeling are proposed separately. Then the adversarial loss [31; 17; 56; 41; 49] is incorporated for learning the reconstruction process in an end-to-end manner. They have indeed removed most degradations on general images, but cannot generate realistic details. Furthermore, their degradation settings are limited to super-resolution, which is not complete for the BIR problem. The second group ZIR is a newly emerged direction. Representative works are DDRM , DDNM , and GDP . They incorporate the powerful diffusion model as the additional prior, thus having greater generative ability than GAN-base methods. With a proper degradation assumption, they can achieve impressive zero-shot restoration on classic IR tasks. However, the problem setting of ZIR is not in accordance with BIR. Their methods can only deal with clearly defined degradations (linear or non-linear), but cannot generalize well to unknown degradations. In other words, they can achieve realistic reconstruction on general images, but not on general degradations. The third group is BFR, which focuses on human face restoration. State-of-the-art methods can refer to CodeFormer and VQFR . They have a similar solution pipeline as BSR methods, but are different in the degradation model and generation network. Due to a smaller image space, these methods can utilize VQGAN and Transformer to achieve surprisingly good results on real-world face images. Nevertheless, BFR is only a sub-domain of BIR. It usually assumes a fixed input size and restricted image space, thus cannot be applied to general images. According to the above analysis, we can see that existing BIR methods cannot achieve (1) realistic image reconstruction on (2) general images with (3) general degradations, simultaneously. Therefore, we desire a new BIR method to overcome these limitations.
In this work, we propose DiffBIR to integrate the advantages of previous works into a unified framework. Specifically, DiffBIR (1) adopts an expanded degradation model that can generalize to real-world degradations, (2) utilizes the well-trained Stable Diffusion as the prior to improve generative ability, (3) introduces a two-stage solution pipeline to ensure both realness and fidelity. We also make dedicated designs to realize these strategies. First, to increase generalization ability, we combine the diverse degradation types in BSR and the wide degradation ranges in BFR to formulate a more practical degradation model. This helps DiffBIR to handle diverse and extreme degradation cases. Second, to leverage Stable Diffusion, we introduce an injective modulation sub-network – LAControlNet that can be optimized for our specific task. Similar to ZIR, the pre-trained Stable Diffusion is fixed during finetuning to maintain its generative ability. Third, to realize faithful and realistic image reconstruction, we first apply a Restoration Module (i.e., SwinIR) to reduce most degradations, and then finetune the Generation Module (i.e., LAControlNet) to generate new textures. Without this pipeline, the model may either produce over-smoothed results (remove Generation Module) or generate wrong details (remove Restoration Module). In addition, to meet users’ diverse requirements, we further propose a controllable module that could achieve continuous transition effects between restoration result in stage one and generation result in stage two. This is achieved by introducing the latent image guidance during the denoising process without re-training. The gradient scale that applies to the latent image distance can be tuned to trade off realness and fidelity.
Equipped with the above components, the proposed DiffBIR demonstrates excellent performance in both BSR and BFR tasks on synthetic and real-world datasets. It is worth noting that DiffBIR achieves a great performance leap in general image restoration, outperforming existing BSR and BFR methods (e.g., BSRGAN , Real-ESRGAN , CodeFormer , et.al). We can observe the differences of these methods in some aspects. For complex textures, BSR methods tend to generate unrealistic details, while DiffBIR can produce visually pleasant results, see Figure 1(first row). For semantic regions, BSR methods tend to achieve over-smoothed effects, while DiffBIR can reconstruct semantic details, see Figure 1(second row). For tiny stripes, BSR methods tend to erase those details, while DiffBIR can still enhance their structures, see Figure 1(third row). Moreover, DiffBIR is able to deal with extreme degradations and regenerate realistic and vivid semantic content, see Figure 1(the last row). All these show that DiffBIR has successfully broken the bottlenecks of existing BSR methods. For blind face restoration, DiffBIR shows superiority in dealing with some hard cases, such as maintaining good fidelity on facial area occluded by other objects (see first row in 1 (b)), achieving successful restoration beyond facial areas (see first row in 1 (b)). In conclusion, our DiffBIR could obtain competitive performance for both BSR and BFR tasks in a unified framework for the first time. Extensive and intensive experiments have demonstrated the superiority of our proposed DiffBIR over the existing state-of-the-art BSR and BFR methods.
Related Work
Blind Image Super-Resolution. Latest advances on BSR have explored more complex degradation models to approximate real-world degradations. In particular, BSRGAN aims to synthesize more practical degradations based on a random shuffling strategy, and Real-ESRGAN exploits "high-order" degradation modeling. They both utilize GANs [17; 41; 49; 31; 56] to learn the image reconstruction process under complex degradations. SwinIR-GAN uses the new prevailing backbone Swin Transformer to achieve better image restoration performance. FeMaSR formulates SR as a feature-matching problem based on pre-trained VQ-GAN . Although BSR methods can be useful to remove degradations in the real world, they are not good at generating realistic details. In addition, they typically assume the low-quality image input is downsampled by some certain scales (e.g. ), which is limited for BIR problem.
Zero-shot Image Restoration. ZIR aims to achieve image restoration by leveraging a pre-trained prior network in an unsupervised manner. Earlier works [2; 10; 40; 44] mainly concentrate on searching a latent code within a pre-trained GAN’s latent space. Recent advancements in this field embrace the utilization of Denoising Diffusion Probabilistic Models [21; 51; 52; 46; 45; 48]. DDRM introduces an SVD-based approach to handle linear image restoration tasks efficiently. Meanwhile, DDNM analyzes the range-null space decomposition of a vector theoretically and then designs a sampling schedule based on the null space. Inspired by classifier guidance , GDP introduces a more convenient and effective guidance approach, in which the degradation model can be estimated during inference. Although these works contribute to the advancement of zero-shot image restoration techniques, ZIR methods still cannot achieve satisfactory restoration results in low-quality images from real world.
Blind Face Restoration. As a specific sub-domain of general images, the face image typically carries more structural and semantic information. Early attempts utilize geometric priors (e.g. facial parsing maps , facial landmarks[9; 27], and facial component heatmaps ) or reference priors[34; 33; 32; 13] as auxiliary information to guide the face restoration process. With the rapid development of generative networks, many BFR approaches incorporate powerful generative-prior to reconstruct faces in great realness. Representative GAN-prior-based methods [54; 61; 19; 4] have demonstrated their capability in achieving both high-quality and high-fidelity face reconstruction. State-of-the-art works [68; 18; 59] introduce the HQ codebook to generate surprisingly realistic face details by exploiting Vector-Quantized (VQ) dictionary learning [53; 15].
Methodology
In this work, we aim to exploit a powerful generative prior – Stable Diffusion to solve blind restoration problems for both general and face images. Our proposed framework adopts a two-stage pipeline that is effective, robust, and flexible. First, we employ a Restoration Module to remove corruptions, such as noises or distortion artifacts, using regression loss. As the lost local textures and coarse/fine details are still absent, we then leverage Stable Diffusion to remedy the information loss. The overall framework is illustrated in Figure 2. Specifically, we first pretrain a SwinIR on large-scale dataset to achieve the preliminary degradation removal across diversified degradations (Section 3.1). Then, the generative prior is leveraged for producing realistic restoration results (Section 3.2). In addition, a controllable module based on latent image guidance is introduced for trade-off between realness and fidelity (Section 3.3).
Degradation Model. BIR aims to restore clean images from low-quality (LQ) ones with unknown and complex degradations. Typically, blur, noise, compression artifacts, and low-resolution are often involved. In order to better cover the degradation space of the LQ images, we employ a comprehensive degradation model that considers diversified degradation and high-order degradation. Among all degradations, blur, resize, and noise are the three key factors in real-world scenarios . Our diversified degradation involves blur: isotropic Gaussian and anisotropic Gaussian kernels; resize: area resize, bilinear interpolation and bicubic resize; noise: additive Gaussian noise, Poisson noise, and JPEG compression noise. Regarding high-order degradation, we follow to use the second-order degradation, which repeats the classical degradation model: blur-resize-noise process twice. Note that our degradation model is designed for image restoration, thus all the degraded images will be resized back to their original size.
Restoration Module. To build a robust generative image restoration pipeline, we adopt a conservative yet feasible solution by first removing most of the degradations (especially the noise and compression artifacts) in the LQ images, and then use the subsequent generative module to reproduce the lost information. This design will promote the latent diffusion model to focus more on textures/details generation without the distraction of noise corruption, and achieve more realistic/sharp results without wrong details (see Section 4.3). We modify SwinIR as our restoration module. Specifically, we utilize the pixel unshuffle operation to downsample the original low-quality input with a scale factor of 8. Then, a convolutional layer is adopted for shallow feature extraction. All the subsequent transformer operations are performed in low resolution space, which is similar to latent diffusion model. The deep feature extraction adopts several Residual Swin Transformer Blocks (RSTB), and each RSTB has several Swin Transformer Layers (STL). The shallow and deep features will be added for maintaining both low-frequency and high-frequency information. For upsampling the deep features back to the original image space, we perform nearest interpolation for three times, and each interpolation is followed by one convolutional layer as well as one Leaky ReLU activation layer. We optimize the parameters of the restoration module by minimizing the pixel loss. The formulation is as follows:
where and denote the high-quality image and the low-quality counterpart, respectively. is obtained by regression learning and will be used for the finetuning on latent diffusion model.
2 Leverage Generative Prior for Image Reconstruction
Preliminary: Stable Diffusion. In this paper, we implement our method based on the large-scale text-to-image latent diffusion model – Stable Diffusion. Diffusion models learn to generate data samples through a denoising sequence that estimate the score of the data distribution. In order to achieve better efficiency and stabilized training, Stable Diffusion pretrains an autoencoder that converts an image into a latent with encoder and reconstructs it with decoder . This latent representation is learned by using hybrid objectives of VAE , Patch-GAN , and LPIPS . The diffusion and denoising processes are performed in the latent space. In diffusion process, Gaussian noise with variance at time is added to the encoded latent for producing the noisy latent:
where , and . When is large enough, the latent is nearly a standard Gaussian distribution.
A network is learned by predicting the noise conditioned on (i.e., text prompts) at a randomly picked time-step . The optimization of latent diffusion model is defined as follows:
where are sampled from the dataset and , is uniformly sampled and is sampled from the standard Gaussian distribution.
LAControlNet. Although stage-one could remove most degradations, the obtained is often over-smoothed and still far from the distribution of high-quality natural images. We then leverage the pre-trained Stable Diffusion for image reconstruction with our obtained - pairs. First, we utilize the encoder of Stable Diffusion’s pretrained VAE to map into the latent space, and obtain the condition latent . The UNet denoiser performs latent diffusion, which contains an encoder, a middle block, and a decoder. In particular, the decoder receives the features from encoder and fuses them in different scales. Here we create a parallel module (denoted as orange in Figure 2) that contains the same encoder and the middle block as in the UNet denoiser. Then, we concatenate the condition latent with the randomly sampled noisy as the input for the parallel module. Since this concatenation operation will increase the channel number of the first convolutional layer in the parallel module, we initialize the newly added parameters to zero, where all other weights are initialized from the pre-trained UNet denoiser checkpoints. The outputs of the parallel module are added to the original UNet decoder. Moreover, one convolutional layer is applied before the addition operation for each scale. During finetuning, the parallel module and these convolutional layers are optimized simultaneously, where the prompt condition is set to empty. We aim to minimize the following latent diffusion objective:
The obtained result in this stage is denoted as . To summarize, only the skip-connected features in the UNet denoiser are tuned for our specific task. This strategy alleviates overfitting in small training dataset, and could inherit the high-quality generation from Stable Diffusion. More importantly, our conditioning mechanism is more straightforward and effective for image reconstruction task compared to ControlNet , which utilizes an additional condition network trained from scratch for encoding the condition information. In our LAControlNet, the well-trained VAE’s encoder is able to project the condition images into the same representation space as the latent variables. This strategy significantly alleviates the burden on the alignment between the internal knowledge in latent diffusion model and the external condition information. In practice, directly utilizing ControlNet for image reconstruction leads to severe color shifts as shown in the ablation study (see Section 3).
3 Latent Image Guidance for Fidelity-Realness Trade-off
The above guidance could iteratively force spatial alignment and color consistency between latent features, and guide the generated latent to preserve the content of the reference latent. Therefore, one can control how much information (such as structure, layout and color) is maintained from the reference image , thus achieving a transition from generated output to more smooth result. The whole algorithm of our latent image guidance is illustrated in Algorithm 1.
Experiments
Datasets. We train DiffBIR on the ImageNet dataset at resolution for BIR. As for BFR, we use FFHQ dataset and resize it to . To synthesize the LQ images, we utilize the proposed degradation pipeline to process the HQ images during training (please see Appendix A for details). For BSR, we utilize RealSRSet dataset for comparison in a real-world setting. For a more thorough comparison in real-world scenarios, we collect 47 images from the Internet, denoted as Real47. It contains general images of diverse scenes, such as natural outdoor landscapes, old photos, architecture, humans from portraits to dense people crowds, plants, and animals, etc. For BFR task, we evaluate our method on a synthetic dataset CelebA-Test and three real-world datasets: LFW-Test , CelebChild-Test , and WIDER-Test . In particular, CelebA-Test contains 3,000 images selected from the CelebA-HQ dataset, where LQ images are synthesized under the same degradation range as our training settings.
Implementation. The restoration module adopts 8 residual Swin Transformer blocks (RSTB), and each RSTB contains 6 Swin Transformer Layers (STL). The head number is set to 6 and the window size is set to 8. We train the restoration module with a batch size of 96 for 150k iterations. We utilize Stable Diffusion 2.1-base https://github.com/Stability-AI/stablediffusion as the generative prior, and finetune the diffusion model for 25k iterations with a batch size of 192. We use Adam optimizer and set the learning rate to . The training process is conducted on resolution with 8 NVIDIA A100 GPUs. For inference, we adopt spaced DDPM sampling with 50 timesteps. Our DiffBIR is able to handle images with arbitrary sizes larger than . For images with sides , we first upsample them with the short side enlarged to 512, and then resize them back.
Metrics. Regarding the evaluation with ground truth, we adopt the traditional metrics: PSNR, SSIM, and LPIPS . To better evaluate the realness for BIR task, we also include several no-reference image quality assessment (IQA) metrics: MANIQAMANIQA (https://github.com/IIGROUP/MANIQA) won first place in the NTIRE2022 Perceptual Image Quality Assessment Challenge Track 2 No-Reference competition. and NIQE. For BFR, we evaluate the identity preservation - IDS , and employ the widely used perceptual metric FID . We also deploy a user study for a more thorough comparison.
2 Comparisons with State-of-the-Art Methods
For BSR, we compare our DiffBIR with state-of-the-art BSR methods: Real-ESRGAN+ , BSRGAN , SwinIR-GAN , and FeMaSR . The recent state-of-the-art ZIR methods (DDNM and GDP ) are also included DDNM and GDP are selected because they provide an approach to restore images with arbitrary sizes.. Regarding BFR task, we compare with the most recent state-of-the-art methods: DMDNet , GFP-GAN , GPEN , GCFSR , VQFR , CodeFormer , RestoreFormer .
BSR on real-world dataset. We provide the quantitative comparison on real-world datasets in Table 1. It is observed that our DiffBIR obtains the best scores in MANIQA on both the widely used RealSRSet and our collected Real47. While BSRGAN and Real-ESRGAN+ could achieve top-3 results in MANIQA on both two datasets. The visual comparison results are presented in Figure 3. It can be seen that DiffBIR is able to restore text information more naturally, while other methods tend to distort the characters or produce blurry output. On the other hand, our DiffBIR could also generate realistic texture details for natural images, where other methods produce over-smooth results. More visualization results can be found in Figure 11 and Figure 12.
To further compare DiffBIR with other state-of-the-art methods, we conduct a user study on our collected Real47 dataset. This user study compares DiffBIR, SwinIR-GAN, BSRGAN, and RealESRGAN+. For each image, users are asked to rank the results of the four methods and assign 1-4 points to different methods in an ascending order. To be more exact, better result obtains higher score. 31 users are recruited to conduct this user study under detailed instruction. The distribution of scores obtained by each method is shown in Figure 4. It can be observed that DiffBIR achieves the highest median score, and its upper quartile exceeds 3. This indicates that users tend to rank DiffBIR’s results in the first place. The user study results again demonstrate that DiffBIR’s visual results are superior to other methods, which aligns with its highest score on MANIQA.
BFR on both synthetic and real-world datasets. We show the quantitative comparison on both synthetic and real-world datasets in Table 2. For the synthetic dataset CelebA-Test , our DiffBIR achieves the highest FID score. Meanwhile, it is also the top-3 methods regarding PSNR and IDS. This reveals that the proposed DiffBIR can successfully produce results with both high realness and high fidelity. For real-world datasets, DiffBIR obtains the best results on both LFW-Test (mild degradation) and WIDER-Test (heavy degradation) datasets, and comparable results with state-of-the-art methods on CelebChild-Test. Figure 5 depicts a visual comparison of various methods on synthetic dataset. The first example demonstrates that only DiffBIR succeeds in restoring extremely degraded cases while other methods fail. It can be seen from the second example that only DiffBIR can successfully recover the occluded left eye. Figure 6 presents a visual comparison on real-world dataset. It can be observed from the first example that DiffBIR is able to accurately restore the hair, while other methods mistake the hair for a part of the facial area. The second example suggests that our DiffBIR is the only method that can generate realistic details on non-face area (i.e., the decoration in the forehead). More visualization results can be found in Figure 9 and Figure 10.
3 Ablation Studies
The Importance of Restoration Module. In this part, we investigate the significance of our proposed two-stage pipeline. Here, we remove the Restoration Module (RM), and directly finetune the diffusion model with synthesized training pairs. The removal of restoration module leads to a noticeable performance drop in FID/MANIQA across all real-world datasets (see Table 3). The visual comparison is presented in Figure 7(a). As seen from the first example, the one-stage model (w/o RM) regards the degradations as semantic information by mistake. This demonstrates that the restoration module contributes to preserving fidelity. The second example clearly illustrates that solely finetuning the Stable Diffusion without applying the RM cannot fully remove the real-world noise/artifacts. This indicates that the RM is indispensable in degradation removal.
The Necessity of Finetuning Stable Diffusion. Next, we illustrate the necessity of finetuning the latent diffusion model. Zero-shot IR methods [57; 16] provide an effective approach that guides the reverse diffusion process using the degraded image in the image space. Following their methodology, we employ the smoothed result to guide the original Stable Diffusion without finetuning. However, as depicted in Figure 7(b), this guidance strategy tends to generate unrealistic content (i.e., a bird with one leg missing). This demonstrates that the widely used guidance in image space may not effectively generalize to the latent space, thus finetuning Stable Diffusion becomes indispensable for this image reconstruction task.
The Effectiveness of LAControlNet. Then we aim to emphasize the effectiveness of our proposed LAControlNet that encodes to the latent space. Here we compare with ControlNet , which adopts an additional condition network trained from scratch for conditioning the input information. As shown in Figure 7(c), ControlNet tends to output results with color shifts, as there is no explicit regularization on color consistency during training. One might use non-uniform sampling to increase the probability of optimization in the early sampling stage and achieves better color controlling . Nevertheless, our method is much more straightforward and fully exploits the latent diffusion prior.
The Flexibility of Controllable Module. Considering that generative restoration models may produce unexpected details, here we provide a controllable module for users to explore according to their personal preferences. The visualization result is shown in Figure 8. Our experiments suggest that a larger gradient scale tends to produce a high-fidelity smooth result which is close to . As seen from the first row, DiffBIR’s output has some blue artifacts in the dog’s eyes, thus we set to 200 and higher as well for obtaining a better result. Moreover, the background is also changing (tends to be more blurry) as the gradient scale grows.
Conclusion and Limitations
We propose a unified framework for blind image restoration, named DiffBIR, which could achieve realistic restoration results by leveraging the prior knowledge of pre-trained Stable Diffusion. It consists of two stages: the restoration and generation stage, which ensures both fidelity and realness. Extensive experiments have validated the superiority of DiffBIR over existing state-of-the-art methods for both BSR and BFR tasks. Although our proposed DiffBIR has shown promising results, the potential of text-driven image restoration is not explored. Further exploitation in Stable Diffusion for image restoration task is encouraged. On the other hand, our DiffBIR method requires 50 sampling steps to restore a low-quality image, resulting in much higher computational resource consumption and more inference time compared to other image restoration methods.
References
Appendix A Degradation Details
Degradation settings used for training our DiffBIR are introduced in this section. Following , we employ the second-order degradation process to enhance the robustness of the restoration module in real-world scenarios. Specifically, a degradation model in a certain stage consists of three operations: blur, resize, and noise. Blur. We utilize isotropic Gaussian blur or anisotropic Gaussian blur with equal probabilities. The size of the blur kernel follows a uniform distribution ranging from 7 to 21, and the blur sigma is uniformly sampled between 0.2 and 3 for the first degradation process and between 0.2 and 1.5 for the second degradation process. Resize. We consider multiple resize algorithms, including area resize, bilinear interpolation and bicubic resize. The scaling factor for resize follows a uniform distribution ranging from 0.15 to 1.5 for the first degradation process and from 0.3 to 1.2 for the second degradation process. Noise. We incorporate Gaussian noise, Poisson noise, and JPEG compression noise. The scale of Gaussian noise is uniformly sampled between 1 and 30 in the first degradation process and between 1 and 25 in the second degradation process. The scale of Poisson noise is randomly sampled from 0.05 to 3 and 0.05 to 2.5 for the first and second degradation processes, respectively. The quality of JPEG compression follows a uniform distribution ranging from 30 to 95.
Moreover, we combine the degradation settings adopted in blind face restoration. Specifically, we consider a large dowsampling range $[0.1,12]$. In this way, the generation module is trained to remedy the information loss within a wide range.