Image Super-Resolution by Neural Texture Transfer

Zhifei Zhang, Zhaowen Wang, Zhe Lin, Hairong Qi

Introduction

The traditional single image super-resolution (SISR) problem is defined as recovering a high-resolution (HR) image from its low-resolution (LR) observation . As in other fields of computer vision studies, the introduction of convolutional neural networks (CNNs) has greatly advanced the state-of-the-art of SISR. However, due to the ill-posed nature of SISR problems, most existing methods still suffer from blurry results at large upscaling factors, e.g., 4×\times, especially when it comes to the recovery of fine texture present in the original HR image but lost in its LR counterpart. In recent years, perceptual-related constraints, e.g., perception loss and adversarial loss , have been introduced to the SISR problem formulation, leading to major breakthroughs on visual quality under large upscaling factors . However, they tend to hallucinate fake textures and even produce artifacts.

This paper diverts from the traditional SISR and explores the reference-based super-resolution (RefSR). RefSR utilizes rich textures from the HR references (Ref) to compensate for the lost details in the LR images, relaxing the ill-posed issue and producing more detailed and realistic textures with the help of reference images. Note that the Ref images can be obtained from various sources like photo albums, video frames, web image search, etc. There are existing RefSR approaches that adopt internal examples (self-example) or external high-frequency information to enhance textures. However, these approaches assume the reference images possess similar content as that of the LR image and/or with good alignment. Otherwise, their performance would significantly degrade and even become worse than SISR methods. In contrast, the Ref images play a different role in our setting: it does not require well alignment or similar content to the LR image. Instead, we only intend to transfer the semantically relevant texture from Ref images to the output SR image. Ideally, a robust RefSR algorithm should outperform SISR when good Ref images are given, and achieve comparable performance as SISR when Ref images are not provided or do not possess relevant texture at all. Note that content similarity would infer texture similarity but not vice versa.

Inspired by the recent work on image stylization , we propose a new RefSR algorithm, named Super-Resolution by Neural Texture Transfer (SRNTT), which adaptively transfers textures from the Ref images to the SR image. More specifically, SRNTT conducts local texture matching in the feature space and transfers matched textures to the final output through a deep model. The texture transfer model learns the complicated dependency between LR and Ref textures, and leverages similar textures while suppressing dissimilar textures. The example in Fig. 1 illustrates the advantage of the proposed SRNTT compared with two state-of-the-art works, i.e., SRGAN (for SISR) and CrossNet (for RefSR). SRNTT shows significant boost in synthesizing finer texture as compared to the other methods if using a Ref image with similar content (i.e., Fig. 1(a) upper). Even using a Ref image with unrelated content (i.e., Fig. 1(a) lower), SRNTT is still comparable to SRGAN (similar visual quality but less artifacts), demonstrating the adaptiveness/robustness of SRNTT to different Ref images of various levels of content similarity. By contrast, CrossNet would introduce undesired textures from the unrelated Ref image and shows severe performance degradation

In order to facilitate fair comparison and help advance research on the RefSR problem in general, we propose a new dataset, named CUFED5, which provides training and testing sets accompanied with references of different similarity levels in terms of content, texture, color, illumination, view point, etc. The main contributions of this paper are:

We explore a more general RefSR problem, breaking the performance barrier in SISR (i.e., lack of texture detail) and relaxing constraints in existing RefSR (i.e., alignment assumption).

We propose an end-to-end deep model, SRNTT, for the RefSR problem to recover the LR image conditioned on any given references by multi-scale neural texture transfer. We demonstrate the visual improvement, effectiveness, and adaptiveness of the proposed SRNTT by extensive empirical studies.

We build a benchmark dataset, CUFED5, to facilitate the further research and performance evaluation of RefSR methods in handling references with different levels of similarity to the LR input image.

In the rest of this paper, we review the related works in Section 2. The network architecture and training criteria are discussed in Section 3. In Section 4, the proposed dataset CUFED5 is described in detail. The results of both quantitative and qualitative evaluations are presented in Section 5. Finally, Section 6 concludes this paper.

Related Works

In recent years, deep learning based SISR has shown superior performance in terms of either PSNR or visual quality compared to non-deep-learning based methods . The reader could refer to for more comprehensive review. Here we will only focus on deep learning based methods.

A milestone work that introduced CNN into SR was proposed by Dong et al. , where a three-layer fully convolutional network was trained to minimize the mean squared error (MSE) between the SR image and the original HR image. It demonstrated the effectiveness of deep learning in SR and achieved the state-of-the-art performance. Wang et al. combined the strengths of sparse coding and deep network and made considerable improvement over previous models. To speed up the SR process, Dong et al. and Shi et al. extracted features directly from the LR image, that also achieved better performance compared to processing the upscaled LR image through bicubic interpolation. In recent years, the state-of-the-art performance (in PSNR) were all achieved by deep learning based models .

The above mentioned methods, in general, aim at minimizing MSE between the SR and HR images, which might not always be consistent with the human evaluation (i.e., perceptual quality) . Therefore, perceptual-related constraints were incorporated to achieve better visual quality. Johnson et al. demonstrated the effectiveness of adding perception loss using VGG . Ledig et al. introduced adversarial loss from the generative adversarial nets (GANs) to minimize the perceptually relevant distance between the SR and HR images. Sajjadi et al. further incorporated the texture matching loss based on the idea of style transfer to enhance the texture in the SR image. The proposed SRNTT is more closely related to , where perceptual-related constraints (i.e., perceptual loss and adversarial loss) are incorporated to recover more visually plausible SR images.

2 Reference-based Super-Resolution

In contrast to SISR where only a single LR image is used as input, RefSR methods introduce additional images to assist the SR process. In general, the reference images need to possess similar texture and/or content structure with the LR image. The references could be selected from adjacent frames in a video , images from web retrieval , an external database (dictionary) , or images from different view points . There is a batch of SR methods that refer to self patches/neighborhood , which are widely known as self-example based SR. They do not utilize external references, thus more close to SISR problems. These works mostly build the mapping from LR to HR patches and fuse the HR patches at the pixel level or by a shallow model, which is insufficient to model the complicated dependency between the LR image and extracted details from the HR patches. A more generic scenario of utilizing the references was proposed by Yue et al. , which instantly retrieves similar images from web and conducts global registration and local matching. However, they made a strong assumption — the references have to be well aligned to the LR image. In addition, the shallow model for patch blending made its performance highly dependent on how well the references could be aligned. Zheng et al. proposed a deep model based RefSR method and adopted optical flow to align input and reference. However, optical flow is limited in matching long distance correspondences, thus incapable of handling significantly misaligned references. The proposed SRNTT adopts the ideas of local texture (patch) matching which could handle long distance dependency. Like existing RefSR methods, we also “fuse” Ref texture to the final output, but we conduct it in the multi-scale feature space through a deep model, which enables the learning of complicated transfer process from references with scaling, rotation, or even non-rigid deformations.

Approach

The proposed SRNTT aims to estimate the SR image ISRI^{SR} from its LR counterpart ILRI^{LR} and the given reference images IRefI^{Ref}, synthesizing plausible textures conditioned on IRefI^{Ref} while preserving the consistency with ILRI^{LR} in content. An overview of the proposed SRNTT is shown in Fig. 2. The main idea is to search for matching texture from IRefI^{Ref} in the feature space and then transfer matched textures to ISRI^{SR} in a multi-scale fashion, since the features are more robust to the variance of color and illumination. The multi-scale texture transfer simultaneously considers semantic (higher-level) and textual (lower-level) similarity between ILRI^{LR} and IRefI^{Ref}, leading to transferring related textures while suppressing irrelevant textures.

In addition to minimizing the pixel and/or perceptual distance between the output ISRI^{SR} and the original HR image IHRI^{HR} as most existing SR methods do, we further regularize on the texture consistency between ISRI^{SR} and the matched textures from IRefI^{Ref}, enforcing the effectiveness of texture transfer. The final output ISRI^{SR} is synthesized in an end-to-end manner. Texture searching and transfer will be discussed in Sections 3.1 and 3.2, respectively. Section 3.3 will detail the objective function of SRNTT.

We first conduct feature swapping which searches over the entire IRefI^{Ref} for locally similar textures that can be used to replace (or swap) the texture features of ILRI^{LR} for enhanced SR recovery. The feature searching is conducted in HR spatial coordinate to enable direct texture transfer to the final output ISRI^{SR}. Following the self-example matching strategy , we first apply bicubic up-sampling on ILRI^{LR} to get an upscaled LR image ILR↑I^{LR\uparrow} that has the same spatial size as IHRI^{HR}. We also sequentially apply bicubic down-sampling and up-sampling with the same factor on IRefI^{Ref} to obtain a blurry Ref image IRef↓↑I^{Ref\downarrow\uparrow} that matches the frequency band of ILR↑I^{LR\uparrow}. Instead of estimating a global transformation or optical flow, we match the local patches in ILR↑I^{LR\uparrow} and IRef↓↑I^{Ref\downarrow\uparrow} so that there is no constraint on the global structure of the Ref image, which is a key advantage over CrossNet . As LR and Ref patches may also differ in color and illumination, we match their similarity in the neural feature space ϕ(I)\phi(I) to emphasize the structural and textural information. We use inner product to measure the similarity between neural features:

where Pi(⋅)P_{i}(\cdot) denotes sampling the ii-th patch from neural feature map, and si,js_{i,j} is the similarity between the ii-th LR patch and the jj-th Ref patch. The Ref patch feature is normalized for selecting the best match over all jj. The similarity computation can be efficiently implemented as a set of convolution (or correlation) operations over all LR patches with each kernel corresponding to a Ref patch:

where SjS_{j} is the similarity map for the jj-th Ref patch, and ∗\ast denotes the correlation operation. We use Sj(x,y)S_{j}(x,y) to denote the similarity between the LR patch centered at location (x,y)(x,y) and the jj-th Ref patch. Both LR and Ref patches are densely sampled from their images. Based on the similarity score, we can construct a swapped feature map MM to represent texture-enhanced LR image. Each patch in MM centered at (x,y)(x,y) is defined as

where ω(⋅,⋅)\omega(\cdot,\cdot) maps patch center to patch index. Note that while IRef↓↑I^{Ref\downarrow\uparrow} is used for matching (Eq. 2), the raw Ref IRefI^{Ref} is used in swapping (Eq. 3) so that the HR information from the original references is preserved. Due to the dense sampling of LR patches, we take the average of the swapped features Pj∗(ϕ(IRef))P_{j^{*}}(\phi(I^{Ref})) in the regions where they overlap. The resulting swapped feature map MM is used as the basis for the next texture transfer stage.

2 Neural Texture Transfer

Our texture transfer model is designed by merging multiple swapped texture feature maps into a base deep generative network at different feature layers corresponding to various scales, as illustrated in Fig. 2 (blue box). For each scale or neural layer ll, a swapped feature map MlM_{l} is constructed using the method introduced above, with a texture feature encoder ϕl\phi_{l} matching the current scale. The effectiveness of transferring texture across multiple layers is verified by the ablation study in Section 5.3.

We use residual blocks and skip connections to build the base generative network. The network output ψl\psi_{l} at layer ll is defined recursively as

where Res(⋅)\text{Res}(\cdot) denotes the residual blocks, ∥\| denotes channel-wise concatenation, and ↑2×\uparrow_{2\times} denotes 2×2\times upscaling with sub-pixel convolution . The final SR result image is generated after LL layers to reach target HR resolution:

Fig. 3 illustrates the network structure of texture transfer at one scale, where the residual blocks extract related texture from MlM_{l} (i.e., IRefI^{Ref}) conditioned on ψl\psi_{l} (i.e., ILRI^{LR}) and merge it with target content.

Different from traditional SISR methods that only reduce the difference between ISRI^{SR} and the ground truth IHRI^{HR}, our proposed SRNTT method further takes into account the texture difference between ISRI^{SR} and IRefI^{Ref}. That is, we require the texture of ISRI^{SR} to be similar as the swapped feature map MlM_{l} in the feature space of ϕl\phi_{l}. Specifically, we define a texture loss Ltex\mathcal{L}_{tex} as

where Gr(⋅)Gr(\cdot) computes the Gram matrix, and λl\lambda_{l} is a normalization factor corresponding to the feature size of layer ll. Sl∗S^{*}_{l} is a weighting map for all LR patches calculated as the best matching score in Eq. 3. Intuitively, textures dissimilar to ILRI^{LR} will have lower weight, and thus receiving lower penalty in texture transfer. In this way, the texture transfer from IRefI^{Ref} to ISRI^{SR} is adaptively enforced based on the Ref image quality, leading to more robust texture hallucination as demonstrated in Section 5.3.

3 Training Objective

In order to 1) preserve the spatial structure of the LR image, 2) improve the visual quality of the SR image, and 3) take advantage of the rich texture from Ref images, our objective function combines reconstruction loss Lrec\mathcal{L}_{rec}, perceptual loss Lper\mathcal{L}_{per}, adversarial loss Ladv\mathcal{L}_{adv}, and texture loss Ltex\mathcal{L}_{tex}. The reconstruction loss is adopted in most SR methods. The perceptual and adversarial losses improve visual quality. The texture loss already discussed in Eq. 6 is specific to RefSR.

Perceptual loss has been investigated in recent SR works for better visual quality. We adopt the relu5_1 layer of VGG19 ,

where VV and CC indicate the volume and channel number of the feature maps, respectively, and ϕi\phi_{i} denotes the iith channel of the feature maps extracted from the hidden layer of VGG19 model. ∥⋅∥F\|\cdot\|_{F} denotes the Frobenius norm.

4 Implementation Details

We adopt a pre-trained VGG19 model for feature swapping, which is well-known for its power of texture representation . Feature layers relu1_1, relu2_1, and relu3_1 are used as texture encoder ϕl\phi_{l}’s in multiple scales. To speed up the matching process, we only match on the relu3_1 layer and project the correspondence to layers relu2_1 and relu1_1, and use the same correspondence across all layers. The weights for Lrec\mathcal{L}_{rec}, Lper\mathcal{L}_{per}, Ladv\mathcal{L}_{adv}, and Ltex\mathcal{L}_{tex} are 1, 1e-4, 1e-6, and 1e-4, respectively. Adam optimizer is used with the learning rate of 1e-4. The network is pre-trained for 2 epochs, where only Lrec\mathcal{L}_{rec} is applied. Then, all losses are involved to train another 20 epochs.

Our method can be easily extended to handle multiple Ref images. In all our RefSR experiments, we augment each IRefI^{Ref} with its scaled and rotated versions to get more accurate texture matching results.

Dataset

For RefSR problems, the similarity between the LR and Ref images affects SR results significantly. In general, references with various levels of similarity to LR images should be provided for the purpose of both training and evaluating a RefSR algorithm. To the best of our knowledge, there has not been such a dataset available for public usage. We thus construct such a dataset with Ref images at various similarity levels based on the CUFED dataset that contains 1,883 albums capturing diverse events in daily life. The size of each album varies between 30 and 100 images. Within each album, we collect image pairs in different similarity levels based on SIFT feature matching, which characterizes local texture pattern that is in line with the objective of local texture matching.

We define four similarity levels from high to low, i.e., L1, L2, L3, and L4, according to the number of best matches of SIFT features. From each paired images, we randomly crop 160×\times160 patches from one image as the original HR images, and the corresponding references are cropped from the other image. In this way, we collect 13,761 paired patches as the training set. For the testing dataset, each HR image is paired with all four levels of references in order to extensively evaluate the adaptiveness of a reference-based SR method. We use the similar way to collect image pairs as in building the training dataset. In total, the testing set contains 126 groups of samples. Each group consists of one HR image and four references at levels L1, L2, L3, and L4, respectively. Two examples from the testing set are shown in Fig. 4. We refer to the collected training and testing sets as CUFED5, which would largely facilitate the research on RefSR and provide a benchmark for fair comparison.

To evaluate the generalization capacity of the trained model on CUFED5, we test it on Sun80 and Urban100 . The Sun80 dataset has 80 natural images, each of which is accompanied by a series of web-searching references, while the Urban100 dataset contains building images without references.

Experimental Results

In this section, both quantitative and qualitative comparisons are conducted to demonstrate the advantages of the proposed SRNTT in terms of visual quality and texture enrichment. Following standard protocol, we obtain all LR images by bicubic downscaling (4×\times) from the HR images.

We compare the proposed SRNTT with the state-of-the-art SISR and RefSR algorithms Implementation of SR algorithms in comparison: SRCNN: http://mmlab.ie.cuhk.edu.hk/projects/SRCNN.html SelfEx: https://sites.google.com/site/jbhuang0604/publications/struct_sr SCN: http://www.ifp.illinois.edu/~dingliu2/iccv15/ DRCN: http://cv.snu.ac.kr/research/DRCN/ LapSRN: http://vllab.ucmerced.edu/wlai24/LapSRN/ MDSR: https://github.com/LimBee/NTIRE2017 ENet: https://webdav.tue.mpg.de/pixel/enhancenet/ SRGAN: https://github.com/tensorlayer/srgan CrossNet: https://github.com/htzheng/ECCV2018_CrossNet_RefSR as shown in Table 1. The SISR methods in comparison are SRCNN , SelfEx , SCN , DRCN , LapSRN , MDSR , ENet , and SRGAN , among which MDSR has achieved the state-of-the-art performance in PSNR in recent two years, while ENet and SRGAN are considered the state-of-the-art in visual quality. Two RefSR methods are also included in the comparison, i.e., Landmark and the recently proposed CrossNet , which outperforms previous RefSR methods.

Without loss of generality, examples from Sun80 and Urban100 are displayed in Fig. 5. With the help of references, SRNTT outperforms other SR methods on Sun80. On Urban100, however, there is no HR references. We use LR input as the reference and achieve finer texture that could be transferred from the LR image. In general, SRNTT would outperform existing SR methods with the assistance of references, and we could still achieve state-of-the-art SISR performance when there is no HR information from references. Section 5.3 will further demonstrate the adaptiveness of SRNTT by analyzing the performance on references of different similarity levels.

2 Qualitative Evaluation by User Study

To evaluate the visual quality of the SR images, we conduct user study, where SRNTT is compared to SCN , DRCN , MDSR , ENet , SRGAN , Landmark , and CrossNet . We present the users with pair-wise comparisons, i.e., SRNTT vs. other, and ask the users to select the one with higher resolution. For each reference level, 2,400 votes are collected on the testing results from the CUFED5 dataset. Fig. 6 shows the voting results,

where the percentages favoring SRNTT denotes the percentage of users that prefer SRNTT as compared to the algorithms denoted along the horizontal axis. Overall, SRNTT significantly outperforms the other algorithms with over 90% users voting for SRNTT.

3 Ablation Studies

Similarity between LR and Ref images is a key factor to the success of RefSR methods. This section investigates the performance of CrossNet and the proposed SRNTT at different reference levels. Table 2 lists the results at six levels of references, where “HR (warp)” denotes the reference obtained by random translation (quarter to half width/height), rotation (10∼\sim30 degree), and scaling (1.2×\times∼\sim2.0×\times upscaling) from the original HR image. L1, L2, L3, and L4 are the four levels of references from the proposed CUFED5 dataset. “LR” means using the LR input image as the references (there is no external references).

To further investigate the gap between the CrossNet and SRNTT, we conduct an experiment by replacing feature swapping with optical flow (FlowNet2 ) in the SRNTT framework. As shown in Table 2, “SRNTT-flow” shows large degradation even at “HR” level as compared to SRNTT, reflecting the limitation of optical flow in handling large disparity/misalignment. As the reference similarity level decreases, PSNR/SSIM of SRNTT reduces gracefully as well. At “LR” level, SRNTT still achieves comparable performance as the state-of-the-art SISR algorithms (Table 1). We observe that the PSNR of SRNTT-flow is higher than that of SRNTT at the “LR” level because the Ref is identical to the LR input. In this case, optical flow would easily align Ref to LR, while patch matching may have missed some matches.

3.2 Layers for feature swapping

As discussed in Section 3, feature swapping and transfer at multiple scales would increase the performance of SRNTT. Table 3 demonstrates the effectiveness of utilizing multiple scales as compared to using single scale. The relu1/2/3 denotes three layers/scales, i.e., relu1_1, relu2_1, and relu3_1 from VGG19, used in SRNTT for feature swapping. We observe that the performance in PSNR decreases as reducing the number of scales. The relu3 gets the lowest PSNR because relu3_1 is a higher-level layer that carries less high-frequency information, contributing less to texture transfer as compared to relu1_1 and relu2_1. For each reference level, the PSNR follows the similar trend as the number of scales increases. However, it is interesting that relu3 shows decreasing and then increasing trend as the reference similarity decreases. This demonstrates the stronger adaptiveness of relu3 in preserving spacial structure, i.e., low-similarity textures from the references are suppressed, and it tends to focus more on spacial reconstruction instead of textural recovery. Therefore, the multi-scale texture transfer using deep model gains extreme momentum on adaptively learning the complicated transfer process between the content and external texture.

3.3 Effect of texture loss

The weighted texture loss used in the proposed SRNTT is a key difference from most SR methods. Unlike those style transfer works, where the content image is significantly modified to carry the texture from the style image (i.e., the reference), the proposed SRNTT avoids such “stylization” by local matching, adaptive neural transfer, and spatial/perceptual regularization. The local matching ensures spatially consistent texture, neural transfer gains adaptiveness on texture transfer, and spatial/perceptual regularization forces the spacial consistency globally. The effect of texture loss is shown in Fig. 7. The PSNR tested on CUFED5 are 25.25 and 25.61 for SRNTT w/o and with the texture loss, respectively. Without the texture loss, the finer texture from the references cannot be effectively transferred into the output.

Conclusion

This paper exploited the more generic RefSR problem where the references can be arbitrary images. We proposed SRNTT, an end-to-end network structure that performs multi-level adaptive texture transfer from the references to recover more plausible texture in the SR image. Both quantitative and qualitative experiments were conducted to demonstrate the effectiveness and adaptiveness of SRNTT. In addition, a new dataset CUFED5 was constructed to facilitate the evaluation of RefSR methods. It also provides a benchmark for future RefSR research.

References