Unsupervised Degradation Representation Learning for Blind Super-Resolution

Longguang Wang, Yingqian Wang, Xiaoyu Dong, Qingyu Xu, Jungang Yang, Wei An, Yulan Guo

Introduction

Single image super-resolution (SR) aims at recovering a high-resolution (HR) image from a low-resolution (LR) observation. Recently, CNN-based methods have dominated the research of SR due to the powerful feature representation capability of deep neural networks. As a typical inverse problem, SR is highly coupled with the degradation model . Most existing CNN-based methods are developed based on an assumption that the degradation is known and fixed (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation differs from their assumption .

To handle various degradations in real-world applications, several methods have been proposed to investigate the non-blind SR problem. Specifically, these methods use a set of degradations (e.g., different combinations of Gaussian blurs, motion blurs and noises) for training and assume the degradation of the test LR image is known at the inference time. These non-blind methods produce promising SR results when the true degradation is known in priori.

To super-resolve real images with unknown degradations, degradation estimation needs to be performed to provide degradation information for non-blind SR networks . However, these non-blind methods are sensitive to degradation estimation. Consequently, the estimation error can further be magnified by the SR network, resulting in obvious artifacts . To address this problem, Gu et al. proposed an iterative kernel correction (IKC) method to correct the estimated degradation by observing previous SR results. By iteratively correcting the degradation, artifact-free results can be gradually produced. Since numerous iterations are required at test time by degradation estimation methods and IKC , these methods are time-consuming.

Unlike the above methods that explicitly estimate the degradation from an LR image, we investigate a different approach by learning a degradation representation to distinguish the latent degradation from other ones. Motivated by recent advances of contrastive learning , a contrastive loss is used to conduct unsupervised degradation representation learning by contrasting positive pairs against negative pairs in the latent space (Fig. 1). The benefits of degradation representation learning are twofold: First, compared to extracting full representations to estimate degradations, it is easier to learn abstract representations to distinguish different degradations. Consequently, we can obtain a discriminative degradation representation to provide accurate degradation information at a single inference. Second, degradation representation learning does not require the supervision from groundtruth degradation. Thus, it can be conducted in an unsupervised manner and is more suitable for real-world applications with unknown degradations.

In this paper, we introduce an unsupervised degradation representation learning scheme for blind SR. Specifically, we assume the degradation is the same in an image but can vary for different images, which is the general case widely used in literature . Consequently, an image patch should be similar to other patches in the same image (i.e., with the same degradation) and dissimilar to patches from other images (i.e., with different degradations) in the degradation representation space, as illustrated in Fig. 1. Moreover, we propose a degradation-aware SR (DASR) network with flexible adaption to different degradations based on the learned representations. Specifically, our DASR incorporates degradation information to perform feature adaption by predicting convolutional kernels and channel-wise modulation coefficients from the degradation representation. Experimental results show that our network can handle various degradations and produce promising results on both synthetic and real-world images under blind settings.

Related Work

In this section, we briefly review several major works for CNN-based single image SR and recent advances of contrastive learning.

SR with Single Degradation. As a pioneer work, a three-layer network is used in SRCNN to learn the LR-HR mapping for single image SR. Since then, CNN-based methods have dominated the research of SR due to their promising performance. Kim et al. proposed a 20-layer network with a residual learning strategy. Lim et al. followed the idea of residual learning and modified residual blocks to build a very deep and wide network, namely EDSR. Zhang et al. then combined residual learning and dense connection to construct a residual dense network (RDN) with over 100 layers. Haris et al. introduced multiple up-sampling/down-sampling layers to provide an error feedback mechanism and used self-corrected features to produce superior results. Recently, channel attention and second-order channel attention are further introduced by RCAN and SAN to exploit feature correlation for improved performance.

SR with Multiple Degradations. Although important advances have been achieved by the above SR methods, they are tailored to a fixed bicubic degradation and suffer severe performance drop when the real degradation differs from the bicubic one . To handle various degradations, several efforts have been made to investigate the non-blind SR problem. Specifically, degradation is first used as an additional input in SRMD to super-resolve LR images under different degradations. Later, dynamic convolutions are further incorporated in UDVD to achieve better performance than SRMD. Recently, Zhang et al. developed an unfolding SR network (USRnet) to handle different degradations by alternately solving a data sub-problem and a prior sub-problem. Hussein et al. introduced a closed-form correction filter to transform an LR image to match the one generated by bicubic degradation. Then, existing networks trained for bicubic degradation can be used to super-resolve the transformed LR image.

Zero-shot methods have also been investigated to achieve SR with multiple degradations. In ZSSR , training is conducted at test time using a degradation and an LR image as its input. Consequently, the network can be adapted to the given degradation. However, ZSSR requires thousands of iterations to converge and is quite time-consuming. To address this limitation, optimization-based meta-learning is used in MZSR to make the network adaptive to a specific degradation within a few iterations during inference.

Since degradation is used as an input for these aforementioned methods, they highly rely on degradation estimation methods for blind SR. Therefore, degradation estimation errors can ultimately introduce undesired artifacts to the SR results . To address this problem, Gu et al. proposed an iterative kernel correction (IKC) method to correct the estimated degradation by observing previous SR results. Luo et al. developed a deep alternating network (DAN) by iteratively estimating the degradation and restoring an SR image.

2 Contrastive Learning

Contrastive learning has demonstrated its effectiveness in unsupervised representation learning. Previous methods usually conduct representation learning by minimizing the difference between the output and a fixed target (e.g., the input itself for auto-encoders). Instead of using a pre-defined and fixed target, contrastive learning maximizes the mutual information in a representation space. Specifically, the representation of a query sample should attract positive counterparts while repelling negative counterparts. The positive counterparts can be transformed versions of the input , multiple views of the input and neighboring patches in the same image . In this paper, image patches generated with the same degradation are considered as positive counterparts and contrastive learning is conducted to obtain content-invariant degradation representations, as shown in Fig. 1.

Methodology

The degradation model of an LR image ILRI^{LR} can be formulated as follows:

where IHRI^{HR} is the HR image, kk is a blur kernel, ⊗\otimes denotes convolution operation, ↓s\downarrow_{s} represents downsampling operation with scale factor ss and nn usually refers to additive white Gaussian noise. Following , we use bicubic downsampler as the downsampling operation. In this paper, we first investigate a noise-free degradation model with isotropic Gaussian kernels and then a more general degradation model with anisotropic Gaussian kernels and noises. Finally, we test our network on real-world degradations.

2 Our Method

Our blind SR framework consists of a degradation encoder and a degradation-aware SR network, as illustrated in Fig. 2. First, the LR image is fed to the degradation encoder (Fig. 2(a)) to obtain a degradation representation. Then, this representation is incorporated in the degradation-aware SR network (Fig. 2(b)) to produce the SR result.

The goal of degradation representation learning is to extract a discriminative representation from the LR image in an unsupervised manner. As shown in Fig. 1, we use a contrastive learning framework for degradation representation learning. Note that, we assume the degradation is the same in each image and varies for different images.

Formulation. Given an image patch (annotated with an orange box in Fig. 1) as the query patch, other patches extracted from the same LR image (e.g., the patch annotated with a red box) can be considered as positive samples. In contrast, patches from other LR images (e.g., patches annotated with blue boxes) can be referred to as negative samples. Then, we encode the query, positive and negative patches into degradation representations using a convolutional network with six layers (Fig. 2(a)). As suggested in SimCLR and MoCo v2 , the resulting representations are further fed to a two-layer multi-layer perceptron (MLP) projection head to obtain x,x+{{x}},{{x}}^{+} and x−{{x}}^{-}. x{{x}} is encouraged to be similar to x+{{x}}^{+} while being dissimilar to x−{{x}}^{-}. Following MoCo , an InfoNCE loss is used to measure the similarity. That is,

where NN is the number of negative samples, τ\tau is a temperature hyper-parameter and ⋅\cdot represents the dot product between two vectors.

where NqueueN_{queue} is the number of samples in the queue and pqueuej{{p}}^{j}_{queue} represents the jthj^{\rm th} negative sample.

Discussion. Existing degradation estimation methods aim at estimating the degradation (usually the blur kernel) at pixel level. That is, these methods learn to extract full representations of the degradation. However, they are time-consuming as numerous iterations are required during inference. For example, KernelGAN conducts network training during test and takes over 60 seconds for a single image . Different from these methods, we aim at learning a “good” abstract representation to distinguish a specific degradation from others rather than explicitly estimating the degradation. It is demonstrated in Sec. 4.2 that our degradation representation learning scheme is effective yet efficient and can obtain discriminative representations at a single inference. Moreover, our scheme does not require the supervision from groundtruth degradation and can be conducted in an unsupervised manner.

2.2 Degradation-Aware SR Network

With degradation representation learning, a degradation-aware SR (DASR) network is proposed to super-resolve the LR image using the resultant representation, as shown in Fig. 2(b).

Network Architecture. Figure 2(b) illustrates the architecture of our DASR network. Degradation-aware block (DA block) is used as the building block and the high-level structure of RCAN is employed. Our DASR network consists of 5 residual groups, with each group comprising of 5 DA blocks.

Discussion. Existing SR networks for multiple degradations commonly concatenate degradation representations with image features and feed them to CNNs to exploit degradation information. However, due to the domain gap between degradation representations and image features, directly processing them as a whole using convolution will introduce interference . Different from these networks, by learning to predict convolutional kernels and modulation coefficients based on the degradation representations, our DASR can well exploit degradation information to adapt to specific degradations. It is demonstrated in Sec. 4.2 that our DASR benefits from DA convolution to achieve flexible adaption to various degradations with better SR performance.

Experiments

We synthesized LR images according to Eq. 1 for training and test. Following , we used 800 training images in DIV2K and 2650 training images in Flickr2K as the training set, and included four benchmark datasets (Set5 , Set14 , B100 and Urban100 ) for evaluation. The size of the Gaussian kernel was fixed to 21×2121\times 21 following . We first trained our network on noise-free degradations with isotropic Gaussian kernels only. The ranges of kernel width σ\sigma were set to [0.2,2.0], [0.2,3.0] and [0.2,4.0] for ×2/3/4\times 2/3/4 SR, respectively. Then, our network was trained on more general degradations with anisotropic Gaussian kernels and noises. Anisotropic Gaussian kernels characterized by a Gaussian probability density function N(0,Σ)N(0,\Sigma) (with zero mean and varying covariance matrix Σ\Sigma) were considered. The covariance matrix Σ\Sigma was determined by two random eigenvalues λ1,λ2∼U(0.2,4)\lambda_{1},\lambda_{2}\sim{U(0.2,4)} and a random rotation angle θ∼U(0,π)\theta\sim{U(0,\pi)}. The range of noise level was set to $$.

During training, 32 HR images were randomly selected and data augmentation was performed through random rotation and flipping. Then, we randomly chose 32 Gaussian kernels from the above ranges to generate LR images. For general degradations, Gaussian noises were also added to the resultant LR images. Next, 64 LR patches of size 48×4848\times 48 (two patches from each LR image as illustrated in Sec. 3.2.1) and their corresponding HR patches were randomly cropped. In our experiments, we set τ\tau and NqueueN_{queue} in Eq. 3 to 0.07 and 8192, respectively. The Adam method with β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999 was used for optimization. We first trained the degradation encoder by optimizing LdegradL_{degrad} for 100 epochs. The initial learning rate was set to 1×10−31\times 10^{-3} and decreased to 1×10−41\times 10^{-4} after 60 epochs. Then, we trained the whole network for 500 epochs. The initial learning rate was set 1×10−41\times 10^{-4} and decreased to half after every 125 epochs. The overall loss function is defined as L=LSR+LdegradL=L_{SR}+L_{degrad}, where LSRL_{SR} is the L1L_{1} loss between SR results and HR images.

2 Experiments on Noise-Free Degradations with Isotropic Gaussian Kernels

We first conduct ablation experiments on noise-free degradations with only isotropic Gaussian kernels. Then, we compare our DASR to several recent SR networks, including RCAN , SRMD , MZSR and IKC . RCAN is a state-of-the-art PSNR-oriented SR method for bicubic degradation. MZSR is a non-blind zero-shot SR method for degradations with isotorpic/anisotropic Gaussian kernels. SRMD is a non-blind SR method for degradations with isotropic/anisotropic Gaussian kernels and noises. IKC is a blind SR method that only considers degradations with isotropic Gaussian kernels. Note that, we do not include DAN , USRnet and correction filter for comparison since their degradation model is different from ours. These methods use ss-fold downsamplerExtract upper-left pixel within each s×ss\times{s} patch. rather than bicubic downsampler as the downsampling operation in Eq. 1. To achieve fair comparison with , we re-trained our DASR using their degradation model and provide the results in the supplemental material.

Degradation Representation Learning. Degradation representation learning is used to produce discriminative representations to provide degradation information. To demonstrate its effectiveness, we introduced a network variant (Model 1) by removing degradation representation learning. Specifically, LdegradL_{degrad} was excluded during training without changing the network. Besides, the separate training of degradation encoder was removed and the whole network was directly trained for 500 epochs.

We first compare the degradation representations learned by models 1 and 4. Specifically, we used B100 to generate LR images with different degradations and fed them to models 1 and 4 to produce degradation representations. Then, these representations are visualized using the T-SNE method . It can be observed in Fig. 3(b) that our degradation representation learning scheme can generate discriminative clusters. Without degradation representation learning, degradations with various kernel widths cannot be well distinguished, as shown in Fig. 3(a). This demonstrates that degradation representation learning facilitates our degradation encoder to learn discriminative representations to provide accurate degradation information. We further compare the SR performance of models 1 and 4 in Table 1. If degradation representation learning is removed, model 1 cannot handle multiple degradations well and produces lower PSNR values, especially for large kernel widths. In contrast, model 4 benefits from accurate degradation information provided by degradation representation learning to achieve better SR performance.

Degradation-Aware Convolutions. With degradation encoder, the extracted degradation representation is incorporated by DA convolutions to achieve flexible adaption to different degradations by predicting convolutional kernels and channel-wise modulation coefficients. To demonstrate the effectiveness of these two key components, we first introduced a variant (Model 2) by replacing DA convolutions with vanilla ones. Specifically, degradation representations are stretched and concatenated with image features as in before being fed to vanilla convolutions. Then, we developed another variant (Model 3) by removing the channel-wise modulation coefficient branch. Note that, we adjust the number of channels in models 2 and 3 to ensure comparable model sizes. From Table 1 we can see that our DASR benefits from both dynamic convolutional kernels and channel-wise modulation coefficients to produce better results for various degradations.

Blind SR vs. Non-Blind SR. We further investigate the upper-bound performance of our DASR network by providing groundtruth degradation. Specifically, we replaced the degradation encoder with 5 FC layers to learn a representation directly from the true degradation (i.e., blur kernel). This network variant (Model 5) was then trained from scratch for 500 epochs. When groundtruth degradation is provided, model 5 achieves improved performance and outperforms SRMDNF by notable margins. Further, SRMDNF is quite sensitive to degradation estimation errors under blind settings, with PSNR values being decreased if degradation is not accurately estimated (e.g., 27.55 vs. 26.66/26.18 for σ ⁣= ⁣3.4\sigma\!=\!3.4). In contrast, our DASR (Model 4) benefits from degradation representation learning to achieve better blind SR performance.

Study of Degradation Representations. Our degradation representations aim at extracting content-invariant degradation information from LR images. To demonstrate this, we conduct experiments to study the effect of different image contents to our degradation representations. Specifically, given an HR image, we first generated an LR image I1I_{1} using a Gaussian kernel kk. Then, we randomly selected another 9 HR images to generate LR images (Ii(i=2,3,...10)I_{i}(i=2,3,...10)) using kk. Next, degradation representations were extracted from Ii(i=1,2,...10)I_{i}(i=1,2,...10) to super-resolve I1I_{1}. Note that, Ii(i=2,3,...10)I_{i}(i=2,3,...10) and I1I_{1} share the same degradation but have different image contents. From Fig. 4 we can see that our network achieves relatively stable performance with degradation representations learned from different image contents. This demonstrates that our degradation representations are robust to image content variations.

Comparison to Previous Networks. We compare DASR to RCAN, SRMD, MZSR and IKC. Pre-trained models of these networks are used for evaluation following their default settings. Quantitative results are shown in Table 2, while visualization results are provided in Fig. 5. Note that, MZSRPre-trained ×4\times 4 model of MZSR is based on ss-fold downsampler./IKC are only tested for ×2/4\times 2/4 SR since their pre-trained models for other scale factors are unavailable. For non-blind SR methods (SRMD and MZSR), we first performed degradation estimation to provide degradation information. Since KernelGAN is quite time-consuming (Table 1), the predictor sub-network in IKC was used to estimate degradations.

It can be observed from Table 2 that RCAN produces the highest PSNR results on bicubic degradation (i.e., kernel width 0) while suffering relatively low performance when the test degradations are different from the bicubic one. Although SRMDNF and MZSR can adapt to the estimated degradations, these methods are sensitive to degradation estimation, as demonstrated in Table 1. Therefore, degradation estimation errors can be magnified by SRMDNF and MZSR, resulting in limited SR performance. Since an iterative correction scheme is used to correct the estimated degradation, IKC outperforms SRMDNF with higher PSNR values being achieved. However, IKC is time-consuming due to its iterations. Compared to IKC, our DASR network achieves better performance for different degradations with shorter running time. That is because, our degradation representation learning scheme can extract “good” representations to distinguish different degradations at a single inference.

Visualization results achieved by different methods are shown in Fig. 5. Since RCAN is trained on the fixed bicubic degradation, it cannot reliably recover missing details when the real degradation differs from the bicubic one. Although SRMDNF can handle multiple degradations, failures can be caused by the degradation estimation error. By iteratively correcting the estimated degradations, IKC achieves better performance than SRMDNF. Compared to other methods, our DASR produces results with much clearer details and higher perceptual quality.

3 Experiments on General Degradations with Anisotropic Gaussian Kernels and Noises

We further conduct experiments on general degradations with anisotropic Gaussian kernels and noises. We first analyze the representations learned from general degradations and then compare the performance of our DASR to RCAN, SRMDNF and IKC under blind settings.

Study of Degradation Representations. Experiments are conducted to investigate the effect of two different components (i.e., blur kernels and noises) to our degradation representations. We first visualize the representations for noise-free degradations with various blur kernels in Fig. 6(a). Then, we randomly select a blur kernel and visualize the representations for degradations with different noise levels in Fig. 6(b). It can be observed that our degradation encoder can easily cluster degradations with different noise levels into discriminative groups and roughly distinguish various blur kernels.

Comparison to Previous Networks. We use 9 typical blur kernels and different noise levels for performance evaluation. To super-resolve noisy LR images using RCAN, SRMDNF and IKC, we first denoise the LR images using DnCNN (a state-of-the-art denoising method) under blind settings. Since the pre-trained model of IKC is trained on isotropic Gaussian kernels only, we further fine-tuned this model on anisotropic Gaussian kernels for fair comparison. The predictor sub-network of the fine-tuned IKC model is used to estimate degradations for SRMDNF.

It can be observed from Table 3 that RCAN produces relatively low performance on complex degradations since it is trained on bicubic degradation only. Since SRMDNF is sensitive to degradation estimation errors, its performance for complex degradations is limited. By iteratively correcting the estimated degradations, IKC performs favorably against SRMDNF. However, IKC is more time-consuming since numerous iterations are required. Different from IKC that focuses on pixel-level degradation estimation, our DASR explores an effective yet efficient approach to learn discriminative representations to distinguish different degradations. Using our degradation representation learning scheme, DASR outperforms IKC in terms of PSNR for various blur kernels and noise levels with running time being reduced by over 7 times. Figure 7 further illustrates the visualization results produced by different methods. Our DASR achieves much better visual quality while other methods suffer obvious blurring artifacts.

4 Experiments on Real Degradations

We further conduct experiments on real degradations to demonstrate the effectiveness of our DASR. Following , DASR trained on isotropic Gaussian kernels is used for evaluation on real images. Visualization results are shown in Fig. 8. It can be observed that our DASR produces visually more promising results with clearer details and fewer blurring artifacts.

Conclusion

In this paper, we proposed an unsupervised degradation representation learning scheme for blind SR to handle various degradations. Instead of explicitly estimating the degradations, we use contrastive learning to extract discriminative representations to distinguish different degradations. Moreover, we introduce a degradation-aware SR (DASR) network with flexible adaption to different degradations based on the learned representations. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experimental results show that our network achieves state-of-the-art performance for blind SR with various degradations.

Acknowledge

The authors would like to thank anonymous reviewers for their insightful suggestions. Xiaoyu Dong is supported by RIKEN Junior Research Associate Program. Part of this work was done when she was a master student at HEU.

References