Non-Local Recurrent Network for Image Restoration

Ding Liu, Bihan Wen, Yuchen Fan, Chen Change Loy, Thomas S. Huang

Introduction

Image restoration is an ill-posed inverse problem that aims at estimating the underlying image from its degraded measurements. Depending on the type of degradation, image restoration can be categorized into different sub-problems, e.g., image denoising and image super-resolution (SR). The key to successful restoration typically relies on the design of an effective regularizer based on image priors. Both local and non-local image priors have been extensively exploited in the past. Considering image denoising as an example, local image properties such as Gaussian filtering and total variation based methods are widely used in early studies. Later on, the notion of self-similarity in natural images draws more attention and it has been exploited by non-local-based methods, e.g., non-local means , collaborative filtering , joint sparsity , and low-rank modeling . These non-local methods are shown to be effective in capturing the correlation among non-local patches to improve the restoration quality.

While non-local self-similarity has been extensively studied in the literature, approaches for capturing this intrinsic property with deep networks are little explored. Recent convolutional neural networks (CNNs) for image restoration achieve impressive performance over conventional approaches but do not explicitly use self-similarity properties in images. To rectify this weakness, a few studies apply block matching to patches before feeding them into CNNs. Nevertheless, the block matching step is isolated and thus not jointly trained with image restoration networks.

In this paper, we present the first attempt to incorporate non-local operations in CNN for image restoration, and propose a non-local recurrent network (NLRN) as an efficient yet effective network with non-local module. First, we design a non-local module to produce reliable feature correlation for self-similarity measurement given severely degraded images, which can be flexibly integrated into existing deep networks while embracing the benefit of end-to-end learning. For high parameter efficiency without compromising restoration quality, we deploy a recurrent neural network (RNN) framework similar to such that operations with shared weights are applied recursively. Second, we carefully study the behavior of non-local operation in deep feature space and find that limiting the neighborhood of correlation computation improves its robustness to degraded images. The confined neighborhood helps concentrate the computation on relevant features in the spatial vicinity and disregard noisy features, which is in line with conventional image restoration approaches . In addition, we allow message passing of non-local operations between adjacent recurrent states of RNN. Such inter-state flow of feature correlation facilitates more robust correlation estimation. By combining the non-local operation with typical convolutions, our NLRN can effectively capture and employ both local and non-local image properties for image restoration.

It is noteworthy that recent work has adopted similar ideas on video classification . However, our method significantly differs from it in the following aspects. For each location, we measure the feature correlation of each location only in its neighborhood, rather than throughout the whole image as in . In our experiments, we show that deep features useful for computing non-local priors are more likely to reside in neighboring regions. A larger neighborhood (the whole image as one extreme) can lead to inaccurate correlation estimation over degraded measurements. In addition, our method fully exploits the advantage of RNN architecture - the correlation information is propagated among adjacent recurrent states to increase the robustness of correlation estimation to degradations of various degrees. Moreover, our non-local module is flexible to handle inputs of various sizes, while the module in handles inputs of fixed sizes only.

We introduce NLRN by first relating our proposed model to other classic and existing non-local image restoration approaches in a unified framework. We thoroughly analyze the non-local module and recurrent architecture in our NLRN via extensive ablation studies. We provide a comprehensive comparison with recent competitors, in which our NLRN achieves state-of-the-art performance in image denoising and SR over several benchmark datasets, demonstrating the superiority of the non-local operation with recurrent architecture for image restoration.

Related Work

Image self-similarity as an important image characteristic has been used in a number of non-local-based image restoration approaches. The early works include bilateral filtering and non-local means for image denoising. Recent approaches exploit image self-similarity by imposing sparsity . Alternatively, similar image patches are modeled with low-rankness , or by collaborative Wiener filtering . Neighborhood embedding is a common approach for image SR , in which each image patch is approximated by multiple similar patches in a manifold. Self-example based image SR approaches exploit the local self-similarity assumption, and extract LR-HR exemplar pairs merely from the low-resolution image across different scales to predict the high-resolution image. Similar ideas are adopted for image deblurring .

Deep neural networks have been prevalent for image restoration. The pioneering works include a multilayer perceptron for image denoising and a three-layer CNN for image SR . Deconvolution is adopted to save computation cost and accelerate inference speed . Very deep CNNs are designed to boost SR accuracy in . Dense connections among various residual blocks are included in . Similarly CNN based methods are developed for image denoising in . Block matching as a preprocessing step is cascaded with CNNs for image denoising . Besides CNNs, RNNs have also been applied for image restoration while enjoying the high parameter efficiency .

In addition to image restoration, feature correlations are widely exploited along with neural networks in many other areas, including graphical models , relational reasoning , machine translation and so on. We do not elaborate on them here due to the limitation of space.

Non-Local Operations for Image Restoration

In this section, we first present a unified framework of non-local operations used for image restoration methods, e.g., collaborative filtering , non-local means , and low-rank modeling , and we discuss the relations between them. We then present the proposed non-local operation module.

2 Classic Methods

The proposed framework works with various classic non-local methods for image restoration, including methods based on low-rankness , collaborative filtering , joint sparsity , as well as non-local mean filtering .

Block matching (BM) is a commonly used approach for exploiting non-local image structures in conventional methods . A q×qq\times q spatial neighborhood is set to be centered at each location ii, and Xi\boldsymbol{X}_{i} reduces to the image patch centered at ii. BM selects the KiK_{i} most similar patches (Ki≪q2K_{i}\ll q^{2}) from this neighborhood, which are used jointly to restore Xi\boldsymbol{X}_{i}. Under the proposed non-local framework, these methods can be represented as

Except for the hard block matching, other methods, e.g., the non-local means algorithm , apply soft block matching by calculating the correlation between the reference patch and each patch in the neighborhood. Each element Φ(X)ij\Phi(\boldsymbol{X})_{i}^{j} is determined only by each {Xi,Xj}\{\boldsymbol{X}_{i},\boldsymbol{X}_{j}\} pair, so Φ(X)ij=ϕ(Xi,Xj)\Phi(\boldsymbol{X})_{i}^{j}=\phi(\boldsymbol{X}_{i},\boldsymbol{X}_{j}), where ϕ( ⋅ )\phi(\,\cdot\,) is determined by the distance metric. In , weighted Euclidean distance with Gaussian kernel is applied as the metric, such that ϕ(Xi,Xj)=exp{−∥Xi−Xj∥2,a2/h2}\phi(\boldsymbol{X}_{i},\boldsymbol{X}_{j})=\text{exp}\{-\left\|\boldsymbol{X}_{i}-\boldsymbol{X}_{j}\right\|_{2,a}^{2}/h^{2}\}. Besides, identity mapping is directly used as the embedding in , i.e., G(X)j=Xj\boldsymbol{G}(\boldsymbol{X})_{j}=\boldsymbol{X}_{j}. In this case, the non-local framework in (1) reduces to

The conventional non-local methods suffer from the drawback that parameters are either fixed , or obtained by suboptimal approaches , e.g., the parameters of WNNM are learned based on the low-rankness assumption, which is suboptimal as the ultimate objective is to minimize the image reconstruction error.

3 The Proposed Non-Local Module

Based on the general non-local framework in (1), we propose another soft block matching approach and apply the Euclidean distance with linearly embedded Gaussian kernel as the distance metric. The linear embeddings are defined as follows:

The embedding transforms Wθ\boldsymbol{W}_{\theta}, Wϕ\boldsymbol{W}_{\phi}, and Wg\boldsymbol{W}_{g} are all learnable and have the shape of m×l,  m×l,  m×mm\times l,\;m\times l,\;m\times m, respectively. Thus, the proposed non-local operation can be written as

The proposed non-local operation can be implemented by common differentiable operations, and thus can be jointly learned when incorporated into a neural network. We wrap it as a non-local module by adding a skip connection, as shown in Figure 1, since the skip connection enables us to insert a non-local module into any pre-trained model, while maintaining its initial behavior by initializing Wg\boldsymbol{W}_{g} as zero. Such a module introduces only a limited number of parameters since θ,  ψ\theta,\;\psi and gg are 1×11\times 1 convolutions and m=128,l=64m=128,l=64 in practice. The output of this module on each location only depends on its q×qq\times q neighborhood, so this operation can work on inputs of various sizes.

Relation to Other Methods: Recent works have combined non-local BM and neural networks for image restoration . Lefkimmiatis proposed to first apply BM to noisy image patches. The hard BM results are used to group patch features, and a CNN conducts a trainable collaborative filtering over the matched patches. Qiao et al. combined similar non-local BM with TNRD networks for image denoising. However, as conventional methods , these works conduct hard BM directly over degraded input patches, which may be inaccurate over severely degraded images. In contrast, our proposed non-local operation as soft BM is applied on learned deep feature representations that are more robust to degradation. Furthermore, the matching results in are isolated from the neural network, similar to the conventional approaches, whereas the proposed non-local module is trained jointly with the entire network in an end-to-end manner.

Wang et al. used similar approaches to add non-local operations into neural networks for high-level vision tasks. However, unlike our approach, Wang et al. calculated feature correlations throughout the whole image. which is equivalent to enlarging the neighborhood to the entire image in our approach. We empirically show that increasing the neighborhood size does not always improve image restoration performance, due to the inaccuracy of correlation estimation over degraded input images. Hence it is imperative to choose a neighborhood of a proper size to achieve best performance for image restoration. In addition, the non-local operation in can only handle input images of fixed size, while our module in (6) is flexible to various image sizes. Finally, our non-local module, when incorporated into an RNN framework, allows the flow of correlation information between adjacent states to enhance robustness against inaccurate correlation estimation. This is a new unique formulation to deal with degraded images. More details are provided next.

Non-Local Recurrent Network

In this section, we describe the RNN architecture that incorporates the non-local module to form our NLRN. We adopt the common formulation of an RNN, which consists of a set of states, namely, input state, output state and recurrent state, as well as transition functions among the states. The input, output, and recurrent states are represented as x\boldsymbol{x}, y\boldsymbol{y} and s\boldsymbol{s} respectively. At each time step tt, an RNN receives an input xt\boldsymbol{x}^{t}, and the recurrent state and the output state of the RNN are updated recursively as follows:

where finputf_{\text{input}}, foutputf_{\text{output}}, and frecurrentf_{\text{recurrent}} are reused at every time step. In our NLRN, we set the following:

s0\boldsymbol{s}^{0} is a function of the input image I\boldsymbol{I}.

xt=0,  ∀t∈{1,…,T}\boldsymbol{x}^{t}=0,\;\forall t\in\{1,\dots,T\}, and finput(0)=0f_{\text{input}}(0)=0.

The output state yt\boldsymbol{y}^{t} is calculated only at the time TT as the final output.

We add an identity path from the very first state which helps gradient backpropagation during training , and a residual path of the deep feature correlation between each location and its neighborhood from the previous state. Hence, st={sfeatt,scorrt}\boldsymbol{s}^{t}=\{\boldsymbol{s}_{\text{feat}}^{t},\boldsymbol{s}_{\text{corr}}^{t}\}, and st=frecurrent(st−1,s0),  ∀t∈{1,…,T}\boldsymbol{s}^{t}=f_{\text{recurrent}}(\boldsymbol{s}^{t-1},\boldsymbol{s}^{0}),\;\forall t\in\{1,\dots,T\}, where sfeatt\boldsymbol{s}_{\text{feat}}^{t} denotes the feature map in time tt and scorrt\boldsymbol{s}_{\text{corr}}^{t} is the collection of deep feature correlation. For the transition function frecurrentf_{\text{recurrent}}, a non-local module is first adopted and is followed by two convolutional layers, before the feature s0\boldsymbol{s}^{0} is added from the identity path. The weights in the non-local module are shared across recurrent states just as convolutional layers, so our NLRN still keeps high parameter efficiency as a whole. An illustration is displayed in Figure 3.

Relation to Other RNN Methods: Although RNNs have been adopted for image restoration before, our NLRN is the first to incorporate non-local operations into an RNN framework with correlation propagation. DRCN recursively applies a single convolutional layer to the input feature map multiple times without the identity path from the first state. DRRN applies both the identity path and the residual path in each state, but without non-local operations, and thus there is no correlation information flow across adjacent states. MemNet builds dense connections among several types of memory blocks, and weights are shared in the same type of memory blocks but are different across various types. Compared with MemNet, our NLRN has an efficient yet effective RNN structure with shallower effective depth and fewer parameters, but obtains better restoration performance, which is shown in Section 5 in detail.

Experiments

Dataset: For image denoising, we adopt two different settings to fairly and comprehensively compare with recent deep learning based methods : (1) As in , we choose as the training set the combination of 200 images from the train set and 200 images from the test set in the Berkeley Segmentation Dataset (BSD) , and test on two popular benchmarks: Set12 and Set68 with σ=15,25,50\sigma=15,25,50 following . (2) As in , we use as the training set the combination of 200 images from the train set and 100 images from the val set in BSD, and test on Set14 and the BSD test set of 200 images with σ=30,50,70\sigma=30,50,70 following . In addition, we evaluate our NLRN on the Urban100 dataset , which contains abundant structural patterns and textures, to further demonstrate the capability of using image self-similarity of our NLRN. The training set and test set are strictly disjoint and all the images are converted to gray-scale in each experiment setup. For image SR, we follow and use a training set of 291 images where 91 images are proposed in and other 200 are from the BSD train set. We adopt four benchmark sets: Set5 , Set14 , BSD100 and Urban100 for testing with three upscaling factors: ×2\times 2, ×3\times 3 and ×4\times 4. The low-resolution images are synthesized by bicubic downsampling.

Training Settings: We randomly sample patches whose size equals the neighborhood of non-local operation from images during training. We use flipping, rotation and scaling for augmenting training data. For image denoising, we add independent and identically distributed Gaussian noise with zero mean to the original image as the noisy input during training. We train a different model for each noise level. For image SR, only the luminance channel of images is super-resolved, and the other two color channels are upscaled by bicubic interpolation, following . Moreover, the training images for all three upscaling factors: ×2\times 2, ×3\times 3 and ×4\times 4 are upscaled by bicubic interpolation into the desired spatial size and are combined into one training set. We use this set to train one single model for all these three upscaling factors as in .

We use Adam optimizer to minimize the loss function. We set the initial learning rate as 1e-3 and reduce it by half five times during training. We use Xavier initialization for the weights. We clip the gradient at the norm of 0.50.5 to prevent the gradient explosion which is shown to empirically accelerate training convergence, and we adopt 16 as the minibatch size during training. Training a model takes about 3 days with a Titan Xp GPU. For non-local module, we use circular padding for the neighborhood outside input patches. For convolution, we pad the boundaries of feature maps with zeros to preserve the spatial size of feature maps.

In this section, we analyze our model in the following aspects. First, we conduct the ablation study of using different distance metrics in the non-local module. Table 2 compares instantiations including Euclidean distance, dot product, embedded dot product, Gaussian, symmetric embedded Gaussian and embedded Gaussian when used in NLRN of 12 unfolded steps. Embedded Gaussian achieves the best performance and is adopted in the following experiments.

We compare the NLRN with its variants in terms of PSNR in Table 2. We have a few observations. First, the same model with untied weights performs worse than its weight-sharing counter-part. We speculate that the model with untied weights is prone to model over-fitting and suffers much slower training convergence, both of which undermine its performance. To investigate the function of non-local modules, we implement a baseline RNN with the same parameter number of NLRN, and find it is worse than NLRN by about 0.2 dB, showing the advantage of using non-local image properties for image restoration. Besides, we implement NLRNs where non-local module is used in every other state or every three states, and observe that if the frequency of using non-local modules in NLRN is reduced, the performance decreases accordingly. We show the benefit of propagating correlation information among adjacent states by comparing with the counter-part in terms of restoration accuracy. To further analyze the non-local module, we visualize the feature correlation maps for non-local operations in Figure 5. It can be seen that as the number of recurrent states increases, the locations with similar features progressively show higher correlations in the map, which demonstrates the effectiveness of the non-local module for exploiting image self-similarity.

Figure 5 investigates the influence of the neighborhood size in the non-local module on image denoising results. The performance peaks at q=45q=45. This shows that limiting the neighborhood helps concentrate the correlation calculation on relevant features in the spatial vicinity and enhance correlation estimation. Therefore, it is necessary to choose a proper neighborhood size (rather than the whole image) for image restoration. We select q=45q=45 for the rest of this paper unless stated otherwise.

The unrolling length TT determines the maximum effective depth (i.e., maximum number of convolutional layers) of NLRN. The influence of the unrolling length on image denoising results is shown in Figure 6. The performance increases as the unrolling length rises, but gets saturated after T=12T=12. Given the tradeoff between restoration accuracy and inference time, we adopt T=12T=12 for NLRN in all the experiments.

2 Comparisons with State-of-the-Art Methods

We compare our proposed model with a number of recent competitors for image denoising and image SR, respectively. PSNR and SSIM are adopted for measuring quantitative restoration performance.

Image Denoising: For a fair comparison with other methods based on deep networks, we train our model under two settings: (1) We use the training data as in TNRD , DnCNN and NLNet , and the result is shown in Table 4. We cite the result of NLNet in the original paper , since no public code or model is available. (2) We use the training data as in RED and MemNet , and the result is shown in Table 5. We note that RED uses multi-view testing to boost the restoration accuracy, i.e., RED processes each test image as well as its rotated and flipped versions, and all the outputs are then averaged to form the final denoised image. Accordingly, we perform the same procedure for NLRN and find its performance, termed as NLRN-MV, is consistently improved. In addition, we include recent non-deep-learning based methods: BM3D and WNNM in our comparison. We do not list other methods whose average performances are worse than DnCNN or MemNet. Our NLRN significantly outperforms all the competitors on Urban100 and yields the best results across almost all the noise levels and datasets.

To further show the advantage of the network design of NLRN, we compare different versions of NLRN with several state-of-the-art network models, i.e., DnCNN, RED and MemNet in Table 3. NLRN uses the fewest parameters but outperforms all the competitors. Specifically, NLRN benefits from inherent parameter sharing and uses only less than 1/10 parameters of RED. Compared with the RNN competitor, MemNet, NLRN uses only half of parameters and much shallower depth to obtain better performance, which shows the superiority of our non-local recurrent architecture.

Image Super-Resolution: We compare our model with several recent SISR approaches, including SRCNN , VDSR , DRCN , LapSRN , DRRN and MemNet in Table 6. We crop pixels near image borders before calculating PSNR and SSIM as in . We do not list other methods since their performances are worse than that of DRRN or MemNet. Besides, we do not include SRDenseNet and EDSR in the comparison because the number of parameters in these two network models is over two orders of magnitude larger than that of our NLRN and their training datasets are significantly larger than ours. It can be seen that NLRN yields the best result across all the upscaling factors and datasets. Visual results are provided in Section 7.

Conclusion

We have presented a new and effective recurrent network that incorporates non-local operations for image restoration. The proposed non-local module can be trained end-to-end with the recurrent network. We have studied the importance of computing reliable feature correlations within a confined neighorhood against the whole image, and have shown the benefits of passing feature correlation messages between adjacent recurrent stages. Comprehensive evaluations over benchmarks for image denoising and super-resolution demonstrate the superiority of NLRN over existing methods.

Appendix

Similar to WNNM , BM3D also applies BM first to group similar patches based on their Euclidean distances. The matched patches are then processed via Wiener filtering , and the denoised results of the ii-th group of patches are

2 Visual Results

We show the visual comparison of our NLRN and several competing methods: BM3D , WNNM , and MemNet for image denoising in Figure 7. Our method can recover more details from the noisy measurement. The visual comparison of our NLRN and several recent methods: DRCN , LapSRN , DRRN , and MemNet for image super-resolution is displayed in Figure 8. Our method is able to reconstruct sharper edges and produce fewer artifacts especially in the regions of repetitive patterns.

References