Learning Degradation Representations for Image Deblurring

Dasong Li, Yi Zhang, Ka Chun Cheung, Xiaogang Wang, Hongwei Qin, Hongsheng Li

Introduction

Image restoration is required to handle various and complicated degradation patterns produced in different degradation processes. The degradation representations act as a crucial component to model the degradation processes and handle complicated degradation patterns, such as different noise levels in image denoising and different combinations of Gaussian blurs and motion blurs in blind super-resolution . However, the degradation representations are less exploited in learning-based deblurring methods and have not been well integrated into state-of-the-art deblurring networks.

The general blurring process can be formulated as

where xx and yy are sharp image and blurry image respectively. F(x,k)F(x,k) is usually modeled as a blurring operator with kernel kk. η\eta represents the Gaussian noise.

A popular paradigm for image deblurring is based on the Maximum A Posterior (MAP) estimate framework,

Based on the limitations of kernel-based blurring modeling, a series of kernel-free approaches are proposed to directly learn the mapping from blurry images to corresponding sharp images. While those methods outperform previous deblurring methods significantly, their performances are still limited in complicated blurry patterns, due to the lack of explicit modeling of the degradation process. This is because, unlike denoising methods, where the noise level might be similar across different images, blurring of different images generally have totally different patterns and cannot be well handled by fixed-weight networks without considering the degradation process. To combine the modeling of degradation and learning-based deblurring, recent works propose to learn explicit degradation representations by using Deep Image Prior (DIP) to reparameterize the kernel kk and the sharp image xx. This inevitably involves the time-consuming iterative inverse optimization and hyperparameter tuning of DIP to adapt to the deblurring process. Moreover, degradation representations have not been taken as a common component in SOTA deblurring methods .

In this paper, we propose to learn explicit degradation representations with a novel joint sharp-to-blurry image reblurring and blurry-to-sharp image deblurring learning framework. Specifically, the degradation representations are learned in the process of sharp-to-blurry image reblurring. The process takes as input a blurry image and learns the degradation representations as a multi-channel spatial latent map to encode the spatially varying blur patterns in replacement of the conventional convolutional blur kernels or the DIP prior. A reblurring generator then takes as input the latent degradation map and the original sharp image and reblurs the sharp image back to its corresponding blurry image.

To effectively integrate the learned degradation representations into the reblurring process, we introduce a multi-scale degradation injection network (MSDI-Net) for achieving conditional image reblurring. The network adopts a U-Net like architecture . The sharp image is fed into the encoder of the U-Net but the degradation is input into the encoder-decoder via the skip-connections to modulate the shortcut encoder feature maps. Specifically, the latent degradation map is gradually upsampled via nearest-neighbor interpolation and a convolution layer to multiple resolutions and are then used to predict spatially varying weighting and bias parameters of the shortcut feature maps at each corresponding resolution. In this way, the learning of the latent degradation representation is supervised by the original blurry image for sharp-to-blurry image reblurring.

To make the learned degradation representations contributing to image deblurring, another blurry-to-sharp image generator network is also introduced, which shares the same MSDI-Net architecture but does not share weights with the reblurring generator. The learned degradation representations are processed similarly to deblur a blurry input image. Specifically, its encoder takes the blurry image as input, while the latent degradation representations are used to modulate the blurry image’s encoder-decoder shortcut feature maps at multiple resolutions. With the help of learned degradation map, the deblurring generator can handle complicated spatially varying blurry patterns, which are adaptively learned from the data to optimize both reblurring and deblurring tasks.

The main contributions of this work are two-fold: 1) We propose a novel joint framework for learning both sharp-to-blurry image reblurring and blurry-to-sharp image deblurring to adaptively encode spatially varying degradations and model the image blurring process, which in turn, benefits the image deblurring performance. 2) The proposed joint reblurring and deblurring framework outperforms state-of-the-art image deblurring methods on the widely used GoPro and RealBlur datasets.

Related Work

In this section, we briefly talk about the related works of image restoration with degradation representations and different image deblurring methods.

Image restoration with degradation representations. Image restoration tasks are usually required to handle different and complicated degradations on real-world applications. The degradation representations have been exploited and taken as one crucial component in several image restoration tasks, such as image denoising and image super-resolution. In image denoising, take the noise variance as one network input to adaptively handle various noise strengths. Several practical denoising methods stabilize the noise variance caused by various ISO and the property of Poisson-Gaussian distribution . Similarly, image blind super-resolution is required to handle various degradations (different Gaussian blurs, motion blurs, and noises) on real-world applications. proposes an unsupervised learning scheme for learning degradation representations based on the assumption that the degradation is the same in an image but can vary for different images. However, this assumption does not hold in image deblurring.

Optimization-based deblurring. A popular approach for image deblurring is based on the Maximum A Posterior (MAP). Most MAP-based methods focus on finding good priors for sharp images and blur kernels. Many priors are designed to model clean images and blur kernels. They include total variation (TV) , hyper-Laplacian prior , l0l_{0}-norm gradient prior and sparse image priors . They all assume the blur kernel is linear and uniform and can be represented as a convolution kernel. However, this assumption does not hold for real-world blurring with the non-uniform kernels. Some non-uniform deblurring methods are proposed based on the assumption that the blur is locally uniform. They are not practical even with high computational costs.

Learning-based deblurring. Many deep deblurring models have been proposed over the past few years. Earlier attempts utilize deep convolution neural networks to facilitate blur kernel estimation. However, there are several limitations in estimating kernels. 1) Simple kernel convolution is not practical on real-world challenging cases. 2) The incorrect kernel estimations, caused by noise and large motions, may cause unpleasing artifacts. 3) Estimating spatially varying kernels requires a huge amount of computation. To avoid the above limitations, a series of kernel-free methods are proposed with much better performances. Nah et al. propose a multi-scale network for image deblurring. Similarly, Tao et al. propose a scale-recurrent structure for image deblurring. Adversarial training is also introduced in image deblurring . Chen et al. introduces a reblur2deblur framework for video deblurring. Zhang et al. proposes a reblurring network to synthesize additional blurry training images. The reblurring network and deblurring network are separate in both two methods. In this work, we combine reblurring network and deblurring network in learning the degradation representations. Recently, multi-stage approaches achieve impressive performance against previous methods. While those deblurring networks outperform the traditional deblurring methods significantly, their performances are still limited due to the lack of explicit degradation modeling.

Learning Blurring Degradation Representations. While the explicit degradation representations have shown convincing improvements on many low-level vision tasks, it is rarely explored in learning-based deblurring methods. SelfDeblur introduces the Deep Image Prior (DIP) to model the clean images and the kernels separately. But, it assumes the blur kernels are linear and uniform, which is not practical in complicated real-world scenes. Tran et al. address the limitation by introducing the explicit representation for the blur kernels and the blur operators and reparameterizing the degradation and the sharp image by using DIP. This inevitably involves alternative optimization, which is time-consuming. The learned representations cannot improve the performance of existing deblurring networks directly and their application is limited to time-consuming DIP-based optimization methods.

Methodology

In this section, we first introduce a joint learning framework for both sharp-to-blurry image reblurring and blurry-to-sharp image deblurring to encode latent spatially varying degradation representations from blurry images. A blur-aware loss is introduced to enhance the image deblurring performance. To more effectively integrate the latent degradation representations for reblurring and deblurring, a multi-scale degradation injection network is proposed for both tasks.

Most existing deblurring methods take the blur kernels as the degradation representations and model the blurring process as a convolution on the input image. However, the simple kernel convolution is not practical on real-world challenging cases and it is usually difficult to estimate the blur kernels in large motions and spatially varying blurring cases. We propose to encode latent degradation representations from blurry images via the joint learning of image reblurring and deblurring. As shown in Fig. 1, we introduce an encoder EE to encode the blurry image yy into the degradation representations E(y)E(y), which is modeled as a multi-channel latent map encoding 2D spatially varying blurring degradation in a latent space. Then an image reblurring generator GrG_{r} and an image deblurring generator GdG_{d} are introduced to generate the reblurred image yˊ\acute{y} and the deblurred image xˊ\acute{x}, respectively. The degradation representations do not only help the reblurring generator GrG_{r} model the degradation process, but also help the deblurring network handle complicated spatially varying degradation patterns. The joint training of reblurring and deblurring strengthens the expressiveness of learned degradation representations.

Sharp-to-blurry image reblurring. Different from modeling the blurring process as a convolution, our generator GrG_{r} models the degradation process via learning to generate the blurry image yˊ\acute{y}, given the sharp image xx and the corresponding degradation representations E(y)E(y). In addition, instead of generating the whole blurry image from scratch, the reblurring generator GrG_{r} learns to predict the residual between the sharp image xx and the blurry image yy. The learning of the blurring degradation process in our framework is therefore formulated as generating a blurry image yˊ\acute{y} from its clean image xx as

where yˊ\acute{y} is the reblurred image conditioned on the sharp image xx and degradation representation E(y)E(y). With the residual learning, the encoder EE is encouraged to neglect the contents of the blurry image and to focus on disentangling the content-independent degradation representation E(y)E(y) from the blurry image yy.

Note that most existing encoder-decoder networks, such as VAE , can also learn implicit image representations via image reconstruction. They aim at encoding the whole-image contents into the latent representations. However, our framework aims at encoding only the degradation information by predicting the residuals between the sharp and blurry images. Such a task actually has a lower difficulty level than reconstructing all contents of an input image and therefore leads to encoding better blurring degradation representations.

We first tried the mainstream L1L_{1} distance as the loss function. But the L1L_{1} loss function merely measures the pixel-wise distance and cannot properly describe the similarity of the blurry patterns of two images. Then the reblurring generator GrG_{r} cannot well generate the blurry images with L1L_{1} loss, which harms the learning of degradation representation. Therefore, we further resort to perceptual loss and adversarial training to distinguish different degradation patterns in training. Adversarial loss is applied on the output of the generator GrG_{r} to distinguish real and fake blurry images so that the decoder can well model the blurring process and improve the expressiveness of degradation representation. We take the hinge loss as the adversarial loss to help the reblurring generator GrG_{r} model the degradation (blurring) process and improve the expressiveness of learned degradation representations. We train the reblurring generator GrG_{r} to generate the blurry images with the multi-scale discriminator DD used in . The training objective for image reblurring is formulated as

where λ1\lambda_{1} balances the L1L_{1} loss and the discriminator loss for image reblurring, and the discriminator DD is a conditional discriminator conditioning on the sharp image xx. Conditioning on the sharp image xx, the conditional discriminator DD can focus on whether the reblurred image yˊ\acute{y} has the same image contents with the corresponding sharp image xx. The reblurring helps extract content-independent degradation information (shown in Fig), which is different with the content-dependent conditional networks .

Blurry-to-sharp image deblurring. To make the learned degradation representations contributing to image deblurring, we also model the image deblurring process with the image reblurring process jointly. The image deblurring generator GdG_{d} follows a similar design to that of GrG_{r}. The deblurring generator GdG_{d} is only required to learn the deblurring residuals between the blurry image yy and its corresponding sharp image xx, and the learned degradation representations E(y)E(y). The blurry-to-sharp image deblurring is modeled as

Thanks to the 2D learnable degradation representations, the image deblurring generator GdG_{d} is aware of the spatially varying blurry patterns and thus can adaptively handle various and complicated degradation patterns. The learning of image deblurring makes the learned degradation representations adapted to the deblurring task. The loss function for image deblurring is formulated as

Discussion of the learned degradation representation. The learned degradation representation has two main advantages against conventional kernel modeling: 1) Our degradation representations can learn non-uniform spatially varying degradations effectively. Figs. 1 and 6 show that the encoder can distinguishe different degradation representations. 2) Interpolating on the latent space of representations can generate blurry images with controllable blurry levels (as shown in Fig. 5). The representations are also content-independent (as shown in Fig. 6), which is different from previous conditional networks . Built on this latent space, the representations show better interpretability and expressiveness.

2 Image Deblurring with Learned Degradation Representations

After obtaining the pre-trained encoder EE, we freeze the pre-trained encoder and re-train the deblurring generator GdG_{d} to illustrate that our improvement is not from the complicated framework but the learned degradation representations. To further demonstrate the generality of learned degradation representations, we also train the deblurring generator GdG_{d} on the RealBlur dataset with the encoder trained on the GoPro dataset .

Following HINet , we select the Peak Signal-to-Noise Ratio (PSNR) loss as the main supervision. We also utilize a blur-aware loss function as extra supervision. The well-trained encoder EE should be quite sensitive to capture various and even subtle blur patterns. We, therefore, define the blur-aware loss as the distance between degradation encoder features of the ground-truth sharp image xx and the estimated sharp image xˊ\acute{x}. The deviation between the encoder feature maps ∥E(x)−E(xˊ)∥1\|E(x)-E(\acute{x})\|_{1} of images xx and xˊ\acute{x} can give more weights on the remaining blurry regions of the network output xˊ\acute{x}. Similar to perceptual loss , the L1L_{1} distances between the encoder feature maps can be calculated at multiple scales.

where E(i)E^{(i)} denotes the ii-th layer of the encoder EE and ∣E(i)∣|E^{(i)}| denotes the number of pixels in feature map E(i)(x)E^{(i)}(x). Using the blur-aware loss makes the deblurring generator GdG_{d} pay more attention to the remaining blurry regions of the network output xˊ\acute{x}. Then the deblurring objective is formulated as

3 Multi-scale Degradation Injection Network

To effectively integrate the degradation E(y)E(y) to predict the reblurring and deblurring residuals, we propose a multi-scale degradation injection network (MSDI-Net) for both reblurring and deblurring. Our reblurring generator GrG_{r} and deblurring generator GdG_{d} consists of two MSDI-Nets, which are stacked as HINet does. for simplicity, the overview of a MSDI-Net is shown in Fig. 2.

where FiF_{i} is the modulated skip-connection features. Since the modulation parameters are predicted from the degradation representations, the learnable modulation of the feature channels makes the deblurring network aware of spatially varying degradations. The degradation-aware feature FiF_{i} is then concatenated with the decoder feature at scale ii and the last decoder layer predicts the image residuals for both image reblurring and deblurring.

Injecting the degradation representations enables the networks to handle various and complicated degradation patterns adaptively. Injection at multiple scales improves the expressiveness of degradation representations by strengthening the connections between the degradation representations and two generators.

Experiments

We train and evaluate our method on the GoPro and RealBlur datasets . The GoPro dataset consists of 2,103 pairs of blurry and sharp images for training and 1,111 pairs for testing. The RealBlur dataset consists of 3,758 pairs for training and 980 pairs for testing. We first train the framework of reblurring and deblurring on GoPro dataset . We apply horizontal flipping and rotation as data augmentation and crop image patch of size 256×256256\times 256 from the dataset for training. λ1\lambda_{1} is set as 30 and λ2\lambda_{2} is set as 10. The networks of whole framework are trained with a batch size of 32 for 200k iterations. Then we freeze the weights of well-trained encoder EE and train the deblurring generator GdG_{d} on the GoPro dataset and the RealBlur dataset respectively. λ3\lambda_{3} is set as 1. The deblurring generator GdG_{d} is trained with a batch size of 64 for 400k iterations. We use the Adam optimizer and the learning rate is set as 3×10−43\times 10^{-4} at the beginning and decreased to 1×10−71\times 10^{-7} following the cosine annealing strategy .

2 Performance comparison

We compare our method with state-of-the-art deblurring methods on the GoPro test dataset . The quantitative results are reported in Table 1. For testing, we slice the whole image into several 256×256256\times 256 patches and test all patches to report the results of HINet , MPRNet-patch256 and our method. Our method achieves 0.45 dB improvement in terms of PSNR over the previous best-performing method HINet .

To evaluate effectiveness and generality of learned degradation representations, we also evaluate our method on the RealBlur dataset. As listed in Table 3, our method achieves the best performance in terms of PSNR and SSIM. Since HINet does not release the model for RealBlur dataset , we train the HINet model on RealBlur dataset based on their released training code. Note that we apply the degradation representations trained on GoPro dataset directly on RealBlur dataset . Our method outperforms the previous SOTA HINet by 0.23 dB PSNR, which demonstrate the generality of learnable degradation representations.

In Table 2, we provide detailed comparisons between our method and MPRNet . We divide the whole gopro dataset into blurriest 10% and sharpest 10% as does. It is observed that the main improvement of our method is the improvement of 0.34dB PSNR in the blurriest 10%, which demonstrates the advantage of proposed degradation learning. What’s more, our method’s computational cost is less than 50% of MPRNet’s .

Fig. 3 and Fig. 4 show example deblurred results from the GoPro and RealBlur test sets by the evaluated approaches. Our method produces sharper images and recovers more details in the regions of texts and moving objects, compared with other methods.

3 Interpolation and decoupleness of degradation representations

4 Ablation study

We evaluate the effectiveness of learning degradation representations and the multi-scale degradation injection network by revising one of the components of our model at a time. Table 4 lists the performances of different settings on the GoPro test set . We first remove the degradation encoder for learning the degradation representations and make GdG_{d} only takes as input the blurry image (denoted as “Ours w/o degradation”). The performance suffers a large drop of 0.47dB PSNR. Then we remove the generator of image reblurring and reserve the encoder to provide additional encoding from the blurry images (denoted as “Our w/o reblurring”), this operation causes a drop of 0.19 dB PSNR, demonstrating that reblurring indeed contributes to the learning of better degradation representations. Then we remove the blur-aware loss function (Eq. (7)). The training objective for deblurring becomes only the PSNR loss (denoted as “Ours w/o blur loss”). The performance drops slightly by about 0.07 PSNR. We further experiment with different ways of integrating degradation representations into the generators. When we remove the injection at multiple scales, the degradation representation is integrated just at the lowest resolution skip connection (denoted as “Our injection w/o multi-scale”). The performance suffers a significant drop of 0.48 PSNR, which means that injecting the degradation representation into a single scale would affect the deblurring performance significantly. Then we replace the modulation and test on concatenating latent degradation map with skip-connection feature maps (denoted as “Our w/ concat injection”). For fair comparisons of computational cost, we add two-layer residual blocks after the feature concatenation. The operation causes a drop of 0.23 PSNR. We also remove the integration totally and concatenate the upsampling degradation maps with the blurry images at the network entrance with the blurry image (denoted as “Ours input w/ concat”), which is the mainstream design in image denoising . Its performance drops by about 0.3 PSNR. All the ablation studies demonstrate the effectiveness of our proposed learnable degradation representations and MSDI-Net for image deblurring.

Conclusions

In this paper, we propose a framework for learning degradation representation with image reblurring and image deblurring. We first utilize an encoder to learn the degradation representations explicitly. Then we propose a multi-scale degradation injection network to effectively integrate the degradation representations for reblurring and deblurring. With the degradation representations, our networks can be aware of and handle spatially varying degradation patterns adaptively. The experimental results demonstrate that our method outperforms other methods with a clear margin.

References