Cross-Scale Internal Graph Neural Network for Image Super-Resolution

Shangchen Zhou, Jiawei Zhang, Wangmeng Zuo, Chen Change Loy

Introduction

The goal of single image super-resolution (SISR) is to recover the sharp high-resolution (HR) counterpart from its low-resolution (LR) observation. Image SR is an ill-posed problem, since there are multiple HR solutions for a LR input. To solve this inverse problem, many convolutional neural networks (CNNs) have been proposed to capture useful priors by learning mappings between LR and HR images. While immense performance has been achieved, learning from external training data solely still falls short in recovering detailed textures for specific images, especially when the up-scaling factor is large.

Apart from exploiting external paired data, internal image-specific information has also been widely studied in image restoration. Some classical non-local methods have shown the values of capturing correlation among non-local self-similar patches for improving the restoration quality. However, convolutional operations are not able to capture such patterns due to the locality of convolutional kernels. Though the receptive fields are large in the deep networks, some long-range dependencies still cannot be well maintained. Inspired by the classical non-local means method , non-local neural networks are proposed to capture long-range dependencies for video classification. Non-local neural networks are thereafter introduced to image restoration tasks .

These methods, in general, perform self-attention weighting of full connection among positions in the features. Besides non-local neural networks, the neural nearest neighbors network and graph-convolutional denoiser network have been proposed to aggregate kk nearest neighboring patches for image restoration. However, all these methods only exploit correlations of recurrent patches within the same scale, without harvesting any high-resolution information. Different from image denoising, the aggregation of multiple similar patches at the same scale (subpixel misalignments) only improves the performance slightly for SR.

The proposed the cross-scale internal graph neural network (IGNN) is inspired by the traditional self-example based SR methods . Our IGNN is based on cross-scale patch recurrence property verified statistically in that patches in a natural image tend to recur many times across scales. An illustrative example is shown in Figure 1 (a). Given a query patch (yellow square) in the LR image ILI_{L}, many similar patches (solid-marked green squares) can be found in the downsampled image IL↓sI_{L\downarrow s}. Thus the corresponding HR patches (dashed-marked green squares) in the LR image ILI_{L} can also be obtained. Such cross-scale patches provide an indication of what the unknown HR patches of the query patch might look like. The cross-scale patch recurrence property is previously utilized as example-based SR constraints to estimate a HR image or a SR kernel .

In this paper, we model this internal correlations between cross-scale similar patches as a graph, where every patch is a vertex and the edge is similarity-weighted connection of two vertexes from two different scales. Based on this graph structure, we then present our IGNN to process this irregular graph data and explore cross-scale recurrence property effectively. Instead of using this property as constraints , IGNN intrinsically aggregates HR patches using the proposed graph module, which includes two operations: graph construction and patch aggregation. More specifically, as shown in Figure 1 (b)(c), we first dynamically construct a cross-scale graph Gk\mathcal{G}_{k} by searching kk-nearest neighboring patches in the downsampled LR image IL↓sI_{L\downarrow s} for each query patch in the LR image ILI_{L}. After mapping the regions of kk neighbors from IL↓sI_{L\downarrow s} to ILI_{L} scale, the constructed cross-scale graph Gk\mathcal{G}_{k} can provide kk LR/HR patch pairs for each query patch. In Gk\mathcal{G}_{k}, the vertexes are the patches in LR image ILI_{L} and their kk HR neighboring patches and the edges are correlations of these matched LR/HR patches. Inspired by Edge-Conditioned Convolution , we formulate an edge-conditioned patch aggregation operation based on the graph Gk\mathcal{G}_{k}. The operation aggregates kk HR patches conditioned on edge labels (similarity of two matched patches). Different from previous non-local methods that explore and aggregate neighboring patches at the same scale, we search for similar patches at the downsampled LR scale but aggregate HR patches. It allows our network to perform more efficiently and effectively for SISR.

The proposed IGNN obtains kk image-specific LR/HR patch correspondences as helpful complements to the external information learned from a training dataset. Instead of learning a LR-to-HR mapping only from external data as other SR networks do, the proposed IGNN makes full use of kk most likely HR counterparts found from the LR image itself to recover more detailed textures. In this way, the ill-posed issue in SR can be alleviated in IGNN. We thoroughly analyze and discuss the proposed graph module via extensive ablation studies. The proposed IGNN performs favorably against state-of-the-art CNN-based SR baselines and existing non-local neural networks, demonstrating the usefulness of cross-scale graph convolution for image super-resolution.

Methodology

In this section, we start by briefly reviewing the general formulation of some previous non-local methods. We then introduce the proposed cross-scale graph aggregation module (GraphAgg) based on graph message aggregation methods . Built on GraphAgg module, we finally present our cross-scale internal graph neural network (IGNN).

Non-local aggregation strategy has been widely applied in image restoration. Under the assumption that similar patches frequently recur in a natural image, many classical methods, e.g., non-local means and BM3D , have been proposed to aggregate similar patches for image denoising. With the development of deep neural network, the non-local neural networks and some kk-nearest neighbor based networks are proposed for image restoration to explore this non-local self-similarity strategy. For these non-local methods that consider similar patch aggregation, the aggregation process can be generally formulated as:

where Xi\boldsymbol{X}^{i} and Yi\boldsymbol{Y}^{i} are the input and output feature patch (or element) at ii-th location (aggregation center), and Xi\boldsymbol{X}^{i} is also the query item in Eq. (1). Xj\boldsymbol{X}^{j} is the jj-th neighbors included in the neighboring feature patch set Si\mathcal{S}_{i} for ii-th location. The Q(⋅)\boldsymbol{Q}(\cdot) transforms the input X\boldsymbol{X} to the other feature space. As for C(⋅,⋅)\boldsymbol{C(\cdot,\cdot)}, it computes an aggregation weights for transformed neighbors Q(Xj)\boldsymbol{Q}(\boldsymbol{X}^{j}) and the more similar patch relative to Xi\boldsymbol{X}^{i} should have the larger weight. The output is finally normalized by a factor δi(X)\delta_{i}(\boldsymbol{X}), i.e., δi(X)=∑j∈SiC(Xi,Xj)\delta_{i}(\boldsymbol{X})=\sum_{j\in\mathcal{S}_{i}}\boldsymbol{C}(\boldsymbol{X}^{i},\boldsymbol{X}^{j}).

The above aggregation can be treated as a GNN if we treat the feature patches and weighted connections as vertices and edges respectively. The non-local neural networks actually model a fully-connected self-similarity graph. They estimate the aggregation weights between the query item Xi\boldsymbol{X}_{i} and all the spatially nearby patches Xj\boldsymbol{X}_{j} in a d×dd\times d window (or within the whole features). To reduce the memory and computational costs introduced by the above dense connection, some kk-nearest neighbor based networks, e.g., GCDN and N3Net , only consider kk (k≪d2k\ll d^{2}) most similar feature patches for aggregation and treat them as the neighbors in Si\mathcal{S}_{i} for every query Xi\boldsymbol{X}^{i}. For all the above mentioned non-local methods, the aggregated neighboring patches are all in the same scale of the query and no HR information is incorporated, thus leading to a limited performance improvement for SISR. In , Irani et al. notice that patch recurrence also exists across the different scales. They explore these cross-scale recurrent LR/HR pairs as example-based constraints to recover the HR images or to estimate the SR kernels from the LR images.

2 Cross-Scale Graph Aggregation Module

For the aforementioned methods , the patch size of the aggregated feature patches is the same as the query one. Even though it works well for image denoising, it fails to incorporate high-resolution information and only provides limited improvement for SR. Based on the patch recurrency property that similar patches will recur in different scales of nature image, we propose a cross-scale internal graph neural network (IGNN) for SISR. An example of patch aggregation in image domain is shown in Figure 1. For each query patch (yellow square) in ILI_{L}, we search for the kk most similar patches (solid-marked squares) in the downsampled image IL↓sI_{L\downarrow s}. we then aggregate their kk HR corresponding patches (dashed-marked squares) in ILI_{L}.

The connections between cross-scale patches can be well constructed as a graph, where every patch is a vertex and edge is a similarity-weighted connection of two vertices from two different scales. To exploit the information of HR patches for SR, we propose a cross-scale graph aggregation module (GraphAgg) to aggregate HR patches in feature domain. As shown in Figure 2, the GraphAgg includes two operations: Graph Construction and Patch Aggregation.

Graph Construction: We first downsample the input LR image ILI_{L} by a factor of ss using the widely used Bicubic operation. The downsampled image is denoted as IL↓sI_{L\downarrow s}, where the downsampling ratio ss is equal to the desired SR up-scaling factor. Thus the found kk neighboring feature patches in graph Gk\mathcal{G}_{k} are the same size as the desired HR feature patch. However, we find that the downsampling ratio s=2s=2 is better than s=4s=4 for up-scaling ×4\times 4 since it is much more difficult to find accurate ×4\times 4 neighbors directly. Thus, we search for ×2\times 2 neighbors instead and obtain ×2\times 2 HR features FL↑sF_{L\uparrow s} aggregated by GraphAgg. We then concatenate FL↑sF_{L\uparrow s} with features after the first ×2\times 2 upsample PixelShuffle operationTwo ×2\times 2 upsample PixelShuffle operations can achieve ×4\times 4 upsampling ..

To obtain the kk neighboring feature patches, we first extract embedded features ELE_{L} and EL↓sE_{L\downarrow s} by the first three layers of VGG19 from ILI_{L} and IL↓sI_{L\downarrow s}, respectively. Following the notion of block matching in classical non-local methods , for a l×ll\times l query feature patch ELq,lE_{L}^{q,l} in ELE_{L}, we find k l×ll\times l nearest neighboring patches EL↓snr,l,r={1,...,k}E_{L\downarrow s}^{n_{r},l},r=\{1,...,k\} in EL↓sE_{L\downarrow s} according to the Euclidean distance between the query feature patch and neighboring ones. Then, we can get the ls×lsls\times ls HR feature patch ELnr,lsE_{L}^{n_{r},ls} corresponding to EL↓snr,lE_{L\downarrow s}^{n_{r},l} in ELE_{L}. We mark this process with a dashed red line in Figure 2, denoted as Vertex Mapping.

Consequently, a cross-scale kk-nearest neighbor graph Gk(V,E)\mathcal{G}_{k}(\mathcal{V},\mathcal{E}) is constructed. V\mathcal{V} is the patch set (vertices in graph) including a LR patch set Vl\mathcal{V}^{l} and a HR neighboring patch set Vls\mathcal{V}^{ls}, where the size of Vl\mathcal{V}^{l} equals to number of LR patches in ELE_{L}. Set E\mathcal{E} is the correlation set (edges in graph) with size ∣E∣=∣Vl∣×k|\mathcal{E}|=|\mathcal{V}^{l}|\times k, which indicates kk correlations in Vls\mathcal{V}^{ls} for each LR patch in Vl\mathcal{V}^{l}. The two vertices of each edge in this cross-scale graph Gk\mathcal{G}_{k} are LR and HR feature patches, respectively. To measure the similarity of query qq and the rr-th neighbor nrn_{r}, we define the edge label as the difference between the query feature patch ELE_{L} and neighboring patch EL↓snr,lE_{L\downarrow s}^{n_{r},l}, i.e., Dnr→q=ELq,l−EL↓snr,l\mathcal{D}^{n_{r}\to q}=E_{L}^{q,l}-E_{L\downarrow s}^{n_{r},l}. It will be used to estimate aggregation weights in the following Patch Aggregation operation.

We search similar patches from EL↓sE_{L\downarrow s} rather than ELE_{L}, hence our searching space is s2s^{2} times smaller than previous non-local methods. Unlike the fully-connected feature graph in non-local neural networks , we only select kk nearset HR neighbors for aggregation, which further leads to a more efficient network. Following the previous non-local methods , we also design a d×dd\times d searching window in EL↓sE_{L\downarrow s}, which is centered with the position of the query patch in the downsampled scale. As verified statistically in , there are abundant cross-scale recurring patches in the whole single image. Our experiments show that searching for kk HR parches from a window region is sufficient for the network to achieve the desired performance.

Patch Aggregation: Inspired by Edge-Conditioned Convolution (ECC) , we aggregate kk HR neighbors in graph Gk\mathcal{G}_{k} weighted on the edge labels Dnr→q\mathcal{D}^{n_{r}\to q}. Our Patch Aggregation reformulates the general non-local aggregation Eq. (1) as:

where FLnr,lsF_{L}^{n_{r},ls} is the rr-th neighboring ls×lsls\times ls HR feature patch from GraphAgg module input FLF_{L} and FL↑sq,lsF_{L\uparrow s}^{q,ls} is the output HR feature patch at the query location. And the patch2img operator is used to transform the output feature patches into the output feature FL↑sF_{L\uparrow s}. We propose to use an adaptive Edge-Conditioned sub-network (ECN), i.e., Fθ(Dnr→q)\mathcal{F}_{\mathbf{\theta}}\left(\mathcal{D}^{n_{r}\to q}\right), to estimate the aggregation weight for each neighbor according to Dnr→q\mathcal{D}^{n_{r}\to q}, which is the feature difference between the query patch and neighboring patch from the embedded feature ELE_{L}. We use exp(⋅)\text{exp}(\cdot) to denote the exponential function and δq(FL)=∑nr∈Sqexp(Fθ(Dnr→q))\delta_{q}(F_{L})=\sum_{n_{r}\in\mathcal{S}_{q}}\text{exp}\left(\mathcal{F}_{\mathbf{\theta}}\left(\mathcal{D}^{n_{r}\to q}\right)\right) to represent the normalization factor. Therefore, Eq. (1) defines an adaptive edge-conditioned aggregation utilizing the sub-network ECN. By exploiting edge labels (i.e., Dnr→q\mathcal{D}^{n_{r}\to q}), GraphAgg aggregates kk HR feature patches in a robust and flexible manner.

To further utilize FL↑sF_{L\uparrow s}, we use a small Downsampled-Embedding sub-network (DEN) to embed it to a feature with the same resolution as FLF_{L} and then concatenate it with FLF_{L} to get FL′F^{\prime}_{L}. Then FL′F^{\prime}_{L} is used in subsequent layers of the network. Note that the two sub-networks ECN and DEN in Patch Aggregation are both very small networks containing only three convolutional layers, respectively. Please see Figure 2 for more details.

Adaptive Patch Normalization: We observe that the obtained kk HR neighboring patches have some low-frequency discrepancy, e.g., color, brightness, with the query patch. Besides the adaptive weighting by edge-conditioned aggregation, we propose an Adaptive Patch Normalization (AdaPN), which is inspired by Adaptive Instance Normalization (AdaIN) for image style transfer, to align the kk neighboring patches to the query one. Let us denote FLq,l,cF_{L}^{q,l,c} and FLnr,ls,cF_{L}^{n_{r},ls,c} as the cc-th channel of features of the query patch and rr-th HR neighboring patch in FLF_{L}, respectively. The rr-th normalized neighboring patch FLnr,ls,cF_{L}^{n_{r},ls,c} by AdaPN is formulated as:

where σ\sigma and μ\mu are the mean and standard deviation. By aligning the mean and variance of the each neighboring patch features with those of the query patch one, AdaPN transfers the low-frequency information of the query to the neighbors and keep their high-frequency texture information unchanged. By eliminating the discrepancy between query patch and kk neighbor patches, the proposed AdaPN benefits the subsequent feature aggregation.

3 Cross-Scale Internal Graph Neural Network

As shown in Figure 2, we build the IGNN based on GraphAgg. After the GraphAgg module, a final HR feature FL↑sF_{L\uparrow s} is obtained. With a skip connection across different scales, the rich HR information in aggregated HR feature FL↑sF_{L\uparrow s} is passed directly from the middle to the late position in the network. This mechanism allows the HR information to help the network in generating outputs with more details. Besides, the enriched intermediate feature FL′F^{\prime}_{L} is obtained by concatenating the input feature FLF_{L} and the downsampled-embedded feature from FL↑sF_{L\uparrow s} using sub-network DEN. It is then fed into the subsequent layers of the network, enabling the network to explore more cross-scale internal information.

Compared to the previous non-local networks for image restoration that only exploit self-similarity patches with the same LR scale, the proposed IGNN exploits internal recurring patches across different scales. Benefits from the GraphAgg module, IGNN obtains kk internal image-specific LR/HR feature patches as effective HR complements to the external information learned from a training dataset. Instead of learning a LR-to-HR mapping only from external data as other CNN SR networks do, IGNN takes advantage of kk most likely HR counterparts to recover more detailed textures. By kk LR/HR exemplars mining, the ill-posed issue of SR can be mitigated in the IGNN.

To show the effectiveness of our GraphAgg module, we choose the widely used EDSR as our backbone network, which contains 32 residual blocks. The proposed GraphAgg module is only used once in IGNN and it is inserted after the 16th residual block.

In Graph Construction, we use the first three layers of the VGG19 with fixed pre-trained parameters to embed image ILI_{L} and IL↓sI_{L\downarrow s} to ELE_{L} and EL↓sE_{L\downarrow s}, respectively. In Graph Aggregation, both adaptive edge-conditioned network and downsample-embedding network are small network with three convolutional layers. More detailed structures are provided in the supplementary material.

Experiments

Datasets and Evaluation Metrics: Following , we use 800 high-quality (2K resolution) images from DIV2K dataset as training set. We evaluate our models on five standard benchmarks: Set5 , Set14 , BSD100 , Urban100 and Manga109 in three upscaling factors: ×2\times 2, ×3\times 3 and ×4\times 4. All the experiments are conducted with Bicubic (BI) dowmsampling degradation. The estimated high-resolution images are evaluated by PSNR and SSIM on Y channel (i.e., luminance) of the transformed YCbCr space.

In the Graph Aggregation module, we set query patch size ll as 3×33\times 3 and the number of neighbors kk as 5. The size of the searching window dd is 30 within the ss times downsampled LR (i.e., EL↓sE_{L\downarrow s}). Note that our GraphAgg is a plug-in module, and the backbone of our network is based on EDSR. We use the pretrained backbone model to initialize the IGNN in order to improve the training stability and save the training time.

We compare our proposed method with 11 state-of-the-art methods: VDSR , LapSRN , MemNet , EDSR , DBPN , RDN , NLRN , RNAN , SRFBN , OISR , and SAN . Following , we also use self-ensemble strategy to further improve our IGNN and denote the self-ensembled one as IGNN+.

As shown in Table 1, the proposed IGNN outperforms existing CNN-based methods, e.g. VDSR , LapSRN , MemNet , EDSR , DBPN , RDN , SRFBN and OISR , and existing non-local neural networks, e.g. NLRN and RNAN . Similar to OISR , IGNN is also built based on EDSR but has better performance. This demonstrates the effectiveness of the proposed GraphAgg for SISR. In addition, the GraphAgg only has two very small sub-networks (ECN and DEN), each of which only contain three convolutional layers. Thus, the improvement comes from the cross-scale aggregation rather than a larger model size. As to SAN , it performs the best in some cases. However, it uses a very deep network (including 200 residual blocks) which is around seven times deeper than the proposed IGNN.

We also present a qualitative comparison of our IGNN with other state-of-the-art methods, as shown in Figure 3. The IGNN recovers more details with less blurring, especially on small recurring textures. This results demonstrate that IGNN indeed explores the rich texture from cross-scale patch searching and aggregation. Compared with other methods, IGNN obtains image-specific information from the searched kk HR feature patches. Such internal cues complement external information obtained by network learning from the dataset. More visual results are provided in the supplementary material.

2 Analysis and Discussions

In this section, we conduct a number of comparative experiments for further analysis and discussions.

Effectiveness of Graph Aggregation Module: In order to show the effectiveness of the cross-scale aggregation intuitively, we provide a non-learning version, denoted as GraphAgg*, which constructs the cross-scale graph in exactly the same way as to IGNN. Different from IGNNwhose GraphAgg aggregates extracted features in IGNN, GraphAgg* directly aggregates kk neighboring HR patches cropped from the input LR image by simply averaging. As shown in the first row of Figure 4, GraphAgg* is capable of recovering more detailed and sharper result, compared with the Bicubic upsampled input LR image. The results intuitively show the effectiveness of cross-scale aggregation in image SR task. Even though the SR images generated from GraphAgg* are promising, they still contain some artifacts in the second row of Figure 4. The proposed IGNN can remove them and restore better images with finer details by aggregating the features extracted from the network.

To further verify the effectiveness of GraphAgg that aggregation information across different scales, we replace it with the basic non-local block (fully-connected aggregation) and KNN (kk neighbors aggregation) within the same scale. The results in Table 3 show that the basic non-local blocks bring limited improvements of only 0.05 dB in PSNR. Besides, our GraphAgg (cross-scale) outperforms KNN (same-scale) by a considerable margin. In contrast, IGNN shows evident improvements in performance, suggesting the importance of cross-scale aggregation for SISR.

Feature Visualization: To validate the effectiveness of the GraphAgg and AdaPN, we also present the intermediate features from two channels in Figure 5. According to the comparison in red boxes between LR features FLF_{L} (Bicubic) with Bicubic upsampling and HR features FL↑sF_{L\uparrow s} (w AdaPN), It is obvious that FL↑sF_{L\uparrow s} contains more rich and sharp details, suggesting the effectiveness of GraphAgg in obtaining more detailed textures in the feature domain.

From the yellow boxes in the three features (in the same row) shown in Figure 5, HR features FL↑sF_{L\uparrow s} (w/o AdaPN) without Adaptive Patch Normalization exhibit some discrepancy in low-frequency information (e.g., color) with input LR features FLF_{L} (Bicubic). However, this discrepancy is almost eliminated in HR features FL↑sF_{L\uparrow s} (w AdaPN) with normalization by AdaPN. It shows that AdaPN makes the GraphAgg module more accurate and robust in patch aggregation.

Position of Graph Aggregation Module: We compare three positions in the backbone network to integrate GraphAgg, i.e., after the 8th residual block, after the 16th residual block and after the 24th residual block. As summarized in Table 3, performance improvement is observed at all positions. The largest gain is achieved by inserting GraphAgg in the middle, i.e., after the 16th residual block.

Settings for Graph Aggregation Module: We investigate the influence of the searching window size dd and neighborhood number kk in GraphAgg. Table 5 shows the results on Urban100 (×2\times 2) for different size of searching window dd. As expected, the estimated SR image has better quality when dd increases. We also find that d=30d=30 has almost the same performance relative to searching among the whole downsampled features EL↓sE_{L\downarrow s} (d=LR↓sd={LR}_{\downarrow s}). Therefore, we empirically set d=30d=30 (i.e., 30×3030\times 30 window) as a trade-off between the computational complexity and performance.

Table 5 presents the results on Urban100 (×2\times 2) for different number of neighbors kk. In general, more neighbors improve SR results since more HR information can be utilized by GraphAgg. However, the performance does not improve after k=5k=5 since it may be hard to find more than five useful HR neighbors for aggregation.

Effectiveness of Adaptive Patch Normalization and Edge-Conditioned sub-network: The retrieved kk HR neighboring patches are sometimes mismatched with the query patch in low-frequency information, e.g., color, brightness. To solve this problem, we adopt two modules in the proposed GraphAgg, i.e., Adaptive Patch Normalization (AdaPN) and Edge-Conditioned sub-network (ECN), i.e., Fθ(Dnr→q)\mathcal{F}_{\mathbf{\theta}}\left(\mathcal{D}^{n_{r}\to q}\right). To validate the effectiveness of AdaPN and ECN, we compare GraphAgg with three variants: removing AdaPN only (w/o AdaPN), removing ECN only (w/o ECN), and removing both of them (w/o AdaPN and w/o ECN). Table 3.2 shows that the network performs worse when any one of them is removed. Note that we remove ECN by replacing exp(Fθ(Dnr→q))\text{exp}\left(\mathcal{F}_{\mathbf{\theta}}\left(\mathcal{D}^{n_{r}\to q}\right)\right) in Eq. (2) by the metric of weighted Euclidean distance with Gaussian kernel, i.e., exp(−∥Dnr→q∥22/10)\text{exp}\left(-\|\mathcal{D}^{n_{r}\to q}\|_{2}^{2}/10\right) . The above experimental results demonstrate that AdaPN and ECW indeed make the GraphAgg module more robust for the patch aggregation.

Conclusion

We present a novel notion of modelling internal correlations of cross-scale recurring patches as a graph, and then propose a graph network IGNN that explores this internal recurrence property effectively. IGNN obtains rich textures from the kk HR counterparts found from LR features itself to alleviate the ill-posed nature in SISR and recover more detailed textures. The paper has shown the effectiveness of the cross-scale graph aggregation, which passes HR information from HR neighboring patches to LR ones. Extensive results over benchmarks demonstrate the effectiveness of the proposed IGNN against state-of-the-art SISR methods.

Acknowledgments

This research is conducted in collaboration with SenseTime and supported by the Singapore Government through the Industry Alignment Fund-Industry Collaboration Projects Grant. It is also partially supported by Singapore MOE AcRF Tier 1 (2018-T1-002-056) and NTU SUG.

Broader Impact

This paper is an exploratory work on single image super-resolution using graph convolutional network. The main impacts of this work are academia-oriented, it is expected to promote the research progress of image super-resolution and motivate novel methods in the related fields.

As for future societal influence, this work will largely improve the quality of pictures taken by cameras or other mobile devices such as smartphones. In addition, it enhances public safety monitored by vision systems via the higher resolution of the monitoring view. Though there might be some potential risks, e.g., criminals might use this technology for peeping into peoples’ personal privacy. It is worth noticing that the positive social effects of this technology far exceeds the potential problems. We call on people to use this technology and its derivative applications without compromising the personal interest of the general public.

References

Appendix: More Visual Results

In this section, we provide more visual comparisons with seven state-of-the-art SISR networks, i.e., VDSR , EDSR , RDN , RCAN , OISR , SAN , and RNAN , on standard benchmark datasets. As shown in Figure 6 and Figure 7, the proposed IGNN recovers richer and sharper details from the LR images especially in the regions with recurring patterns.