Unsupervised Image Super-Resolution using Cycle-in-Cycle Generative Adversarial Networks
Yuan Yuan, Siyuan Liu, Jiawei Zhang, Yongbing Zhang, Chao Dong, Liang Lin
Introduction
Recent deep learning based super-resolution (SR) methods have achieved significant improvement either on PSNR values or on visual quality . These methods require supervised learning on high-resolution (HR) and low-resolution (LR) image pairs. However, their common assumption that the downscaling factor is known and the input image is noise-free hinders them from practical usages. In real-world scenarios, the SR problem often have the following properties: 1) HR datasets are unavailable, 2) downscaling method is unknown, 3) input LR images are noisy and blurry. This problem is extremely difficult if the input images suffer from different kinds of degradation. For an easier case, in this study, we assume that input images are degraded with the same processing which is complex and unavailable.
Under the above circumstances, models learned from synthetic data tend to generate similar results as traditional methods or even simple interpolation. In Fig. 1, we show the results of bicubic interpolation and the state-of-the-art deep learning model—EDSR with a noisy input. This is mainly due to the data bias between training and testing images. Detailed survey and analysis of deep learning based methods on real data can be found in .
As an alternative choice, blind SR deal with the real-world data by estimating the down-sampling kernel from internal or external similar patches. However, when the input is noisy, the down-sampling kernel cannot be accurately estimated, and the inverse mapping results are accompanied by amplified noises. There are also works attempting at restoring LR images with addictive Gaussian noises . But real-world noises may neither be addictive nor follow the standard Gaussian distribution, causing noise estimation infeasible. More generally, LR images may suffer from complex noises, blurry and non-uniform down-sampling kernels, which fail almost all existing blind SR methods.
Inspired by the development of unsupervised learning in image-to-image translation, such as CycleGAN or WESPE , we intend to investigate unsupervised strategies to overcome this obstacle. In CycleGAN, images are translated between different domains with unpaired training data. They assume that the input image is of the same size as the output image, with only the difference on styles. However, in SR, output images are several times larger than the inputs, making the direct application of CycleGAN impossible. Further, using a bicubic-upsampled image as the input also could not obtain satisfactory results. SR problem is specific as it requires high quality output but not just a different style.
After exploring several training strategies, we find an effective Cycle-in-Cycle structure, named CinCGAN, which could achieve superior results. The whole pipeline consists of two CycleGANs, while the second GAN covers the first one (See Fig. 2). The first CycleGAN maps the LR image to the clean and bicubic-downsampled LR space. This module ensures that the LR input is fairly denoised/deblurred. We then stack another well-trained deep model with bicubic-downsampling assumption to up-sample the intermediate result to the desired size. Finally, we fine-tune the whole network using adversarial learning in an end-to-end manner. We conduct experiments on the NTIRE2018 Super-Resolution Challengehttps://competitions.codalab.org/competitions/18024 dataset, and show that the proposed Cycle-in-Cycle structure is much stable at training and achieves competitive performance as supervised deep learning methods.
The contributions of this work are three-folds: 1) We study a more general super-resolution problem, where the high-resolution ground truth, down-sampling kernel and degradation function are unavailable. 2) We explore several unsupervised training strategies under the above assumption, and show that super-resolution task is different from conventional image-to-image translation. 3) We propose a Cycle-in-Cycle structure that could achieve comparable results as supervised CNN networks.
Related work
Single image super-resolution (SISR) has been widely studied for decades. Early approaches either rely on natural image statistics or pre-defined models . Later, mapping functions between LR images and HR images are investigated, such as sparse coding based SR methods .
Recently, deep convolution neural networks (CNN) have shown explosive popularity and powerful capability to improve the quality of SR results. Ever since Dong first proposed using CNN for SR and achieved the state-of-the-art performance, plenty of CNN architectures have been studied for SISR. Inspired by the VGG networks used for ImageNet classification, Kim et al. present a very deep network (VDSR) that learns a residual image. For accelerating the speed of SR, FSRCNN and ESPCN extract feature maps at the low-resolution space and up-sample the image at the last layer by transposed convolution and sub-pixel convolution, respectively. All the above mentioned CNN based SR methods aim at minimizing the mean-square error (MSE) between the reconstructed HR image and the ground truth. Based on the observation that minimizing MSE will make the SR results overly smooth, SRGAN combines an adversarial loss and a perceptual loss as the final objective function, and generates visually pleasing images which contain more high frequency details than the MSE-loss based methods. The champion of NTIRE2017 Super-Resolution Challenge , EDSR , employs deeper and wider networks to achieve the state-of-the-art performance by removing the unnecessary modules in SRResNet .
2 Blind Image Super-Resolution
Although a lot of works focus on SR problems with known degradation/downsamping kernels, little works try to solve blind SR—the degradation operation from HR images to LR images are unavailable. Estimating the degradation/blur kernel is an essential step for blind SR. Wang et al. propose a probabilistic framework combined with the image co-occurrence prior to estimate the unknown point spread function (PSF) parameters. According to the property that small image patches will re-appear in natural images, Michaeli and Irani present a method that is able to estimate the optimal blur kernel. Another relevant work introduces a convolution consistency constraint and bi-l0-l2-norm regularization to guide the blur kernel estimation process, achieving state-of-the-art blind SR performance.
In this work, we investigate how deep learning can be beneficial for addressing blind SR problems.
3 Unsupervised Learning
Existing supervised deep learning methods cannot handle blind SR without LR-HR image pairs. In real-world scenarios, where paired data is unavailable, it is essential to find a way to realize unsupervised learning. Recent work on GAN provides a feasible solution, which includes a generator and a discriminator. The generator tries to generate fake images to fool the discriminator, while the discriminator aims at distinguishing the generated results from real data. GAN is widely used to solve the unsupervised learning problems. DualGAN and CycleGAN are two works about image-to-image translation using unsupervised learning, and both of them present an interesting network structure that contains a pair of forward and inverse generators. The forward generator maps domain X to domain Y, while the inverse generator maps the output back to domain X to maintain cycle consistency. Ignatov et al. use the similar architecture to design a weakly supervised photo enhancer (WESPE) that translates ordinary photos to DSLR-quality images.
Different from the proposed method, both DualGAN and CycleGAN deal with input and output images of the same size, while SR requires the output images several times larger than the inputs. Utilizing the property of cycle consistency, we present a Cycle-in-Cycle GAN (CinCGAN) to super-resolve the LR images of which the degradation operators are unknown. Our method achieves a comparable performance with the state-of-the-art supervised CNN based algorithms .
Proposed Method
The conventional formulation of SISR is , where and denote LR and HR image respectively, represents the down-sampling and blurring matrix, and is the addictive noise. Blind SR follow the same assumption, only with unknown . In this work, we study a more general formulation as , where is the down-sampling process, is a degradation function that may introduce complex noises, shift and blur. Here, we assume that , and the paired HR-LR training data are unavailable. Nevertheless, we can obtain a set of LR images that can be used for analysis and unsupervised training.
1) Why applying unsupervised training? As the down-sampling and degradation functions are complex and coupled, it is hard to perform accurate estimation like traditional blind SR methods . The unavailability of HR images in practise also makes supervised training with simulated paired data impractical. This drives us to explore unsupervised learning strategies. 2) What is the difference between SR and image-to-image translation? SR accepts an LR image and outputs a HR image with much larger resolution. Further, SR requires the output to be of high quality, not just a different style. If we directly apply the image-to-image translation methods, we need to up-sample the LR image first by interpolation, which will also enlarge the noisy patterns. Directly applying existing methods like CycleGAN cannot remove such amplified noises, and training becomes very unstable. Experiments (in Sec. 4.4) also show that when the degradation function varies from image to image, it is difficult to deal with all kinds of images in a single forward pass.
Our solution pipeline consists of three steps. First, we learn a mapping from an LR image set to a “clean” LR image set , where images are noise-free and down-sampled from HR images with bicubic kernel. In other words, we deblur and denoise the input images at low resolution. Second, we adopt an existing SR model to super-resolve the intermediate results to the desired resolution. In the end, we combine and fine-tune these two models simultaneously to get the final HR images.
where is the number of training samples. To maintain consistency between input and output , we add a network and let be identical to the input . Hence, we also use a cycle consistency loss as:
In the previous work , the authors introduce an identity loss to preserve color composition between input and output images when they work on painting generation. They claim that the identity loss can help preserve the color of input images. In image SR, we also need to avoid color variation among different iterations, thus we add an identity loss
In addition, we add a total variation (TV) loss to impose spatial smoothness
where and are functions to compute the horizontal and vertical gradient of .
In summary, the final objective loss for the LRclean LR model is a weighted sum of the four losses:
where are the weights of different losses.
2 Jointly Restoration and Super-Resolution
For the identity loss, instead of maintaining the tint consistency between input and output, we consider ensuring the network can generate adequate quality of super-resolved images. We define a new identity loss as:
To sum up, the total loss for fine-tuning the LR to HR networks is
where , for , are weights of each loss.
3 Network Architecture
The architecture of generators and discriminators are shown in Fig. 3. We adapt similar architecture as the work of Zhu et al. , which has shown impressive results for unpaired image-to-image translation. Here, “conv” means convolution layer, where a Leaky ReLU layer with negative slope 0.2 is added right after except for the last convolution layer (we omit it for simplicity). “BN” means a batch normalization layer. The number after symbols and represents kernel size, number of filters and stride size, respectively. For example, k3n64s1 refers to the convolution layer that contains 64 filters, of which the spatial size is 3 and stride is 1.
For the generators and , we use 3 convolution layers at the head and tail, and 6 residual blocks in the middle. The generator shares the same architecture as and , except for the -nd and -rd convolution layers, where the stride is set to 2 to perform down-sampling. As to the discriminator, we use a PatchGAN for . Since we up-sample LR images with a scale of 4, the size of input images is usually less than 70 (we use LR images and HR images for training). Hence, we modify the stride of the first three convolution layers as 1 for discriminator , such that the respective field of is reduced to .
Experiments
In this section, we first introduce the dataset and details we used for training. We then evaluate the performance of the proposed CinCGAN model by comparing with several state-of-the-art SISR methods. Finally, we perform ablation study to validate the advantages of CinCGAN.
We take the track 2 dataset from the NTIRE2018 Super-Resolution Challenge for training. The challenge aims to restore a HR image given a degraded LR image. They provide a high-quality image dataset, DIV2K , which contains 800 training images and 100 validation images. The DIV2K dataset contains almost all kinds of natural scenarios: buildings (indoor and outdoor), forest, lakes, animals, people, etc. The track 2 dataset is degraded from DIV2K dataset, with down-sampling, blurring, pixel shifting and noises. Although the parameters of the degradation operators are fixed for all images, the blur kernels are randomly generated and their resulting pixel shifts vary from image to image. Hence, the degradation kernels of images in the track 2 dataset are unknown and diverse.
Since our purpose is to unsupervised train a network without paired LR-HR data, we take the first 400 images (numbered from 1 to 400) from the training LR set as input images , and the other 400 images (numbered from 401 to 800) from the HR set as demanding HR images . The intermediate clean LR images are directly bicubic down-sampled from . Similar to , we augment data with 90 degree rotation and flipping. Our experiments are performed with a scaling factor of 4. We randomly crop and with size and crop with size . We conduct testing on the provided 100 validation images. Note that, although DIV2K contains paired training dataset, we do not use paired data for supervised training.
2 Training details
We divide our training process into two steps. We first train the model , and for mapping LR images to clean LR images (shown as LRclean LR in Fig. 2). The three parameters in (5) are set to be and , respectively. We train our model with Adam optimizer by setting , and , without weight decay. Learning rate is initialized as and then decreased by a factor of 2 every 40000 iterations. The weights of filters in each layer are initialized using a normal distribution and the batch size is set as 16. We train the model over 400000 iterations, until it converges.
We then jointly fine-tune the LR to HR model (shown as LRHR in Fig. 2). We initialize our SR network by publicly available EDSR modelhttps://github.com/thstkdgus35/EDSR-PyTorch. We set parameters in (10) as and . The optimizer is set almost the same as training the LRclean LR model, except for we initialize learning rate with . As to the weight of identity loss in (5), we set . At each iteration, we update (5) and (10) in turn. We first train and to update the LRclean LR network. We then train , and simultaneously to update the LRHR network.
We implement the proposed networks with PyTorch and train them on a Nvidia Tesla K80 GPU. It takes about 1 day to pre-train the LRclean LR model and about 2 days to jointly fine-tune the LRHR model.
3 Results
We compare the performance of the proposed CinCGAN model with several state-of-the-art SISR methods: FSRCNN , EDSR and SRGAN . We use the publicly available FSRCNN and EDSR models which are trained with paired LR and HR images, where the inputs are clean LR images down-sampled from HR images. To make the results more comparable, we also fine-tune EDSR and SRGAN (labelled as EDSR+ and SRGAN+ respectively) with the paired track 2 dataset. To emphasize the effectiveness of CinCGAN structure, we also try to first denoise the input LR images and then super-resolve the denoised images for comparison. BM3D is one of the state-of-the-art image denoising approach, which is an efficient and powerful denoiser. Hence, we pre-process the test LR images with BM3D first, and then super-resolve it using EDSR (labelled as BM3D+EDSR).
Table 1 shows the average PSNR and SSIM values of the restored test images. It shows that FSRCNN and EDSR cannot work well if the blur and noises are unknown in the training process. After fine-tuning by paired track 2 dataset, EDSR+ and SRGAN+ improve their results and our method can work comparably against SRGAN+ in terms of PSNR and SSIM without paired training data. Although BM3D can remove noise, it also over-smooth the input images. The PSNR and SSIM values of BM3D+EDSR are lower than the proposed method. Several subjective results are illustrated in Fig. 4.
4 Ablation Study
To validate the advantages of the proposed CinCGAN model for the unsupervised SISR problem, we design some other network structures for comparison.
We remove and from the proposed CinCGAN model for our second experiment. We map the input LR images to a set of clean LR images using the same LRclean LR networks shown in Fig. 2; we then super-resolve the converted LR images directly using the network. The whole structure is shown in Fig. 5(b). The corresponding result is illustrated in Fig. 6(b). As we can see, some negligible noise in the resulted clean LR images is magnified and now is visible in the super-resolved images, which affects the visual quality.
Conclusions
We investigate the single image super-resolution problem with a more general assumption: the low-/high-resolution image pairs and the down-sampling process are unavailable. Inspired by the recent successful image-to-image translation applications, we resort to the unsupervised learning methods to solve this problem. Using generative adversarial networks (GAN), the proposed method contains two CycleGANs, where the second GAN covers the first one. The solution pipeline consists of three steps. First, we map the input LR images to the clean and bicubic-downsampled LR space with the first CycleGAN. We then stack another well-trained deep model with bicubic-downsampling assumption to up-sample the intermediate result to the desired size. Finally, we fine-tune the two modules in an end-to-end manner to get the high-resolution out. Experimental results demonstrate that the proposed unsupervised method achieves comparable results as the state-of-the-art supervised models.
Acknowledgement. This work is supported by SenseTime Group Limited and in part by the Projects of National Science Foundations of China (61571254), Guangdong Special Support plan (2015TQ01X16), and Shenzhen Fundamental Research fund (JCYJ20160513103916577).