Image Super-Resolution via Dual-State Recurrent Networks
Wei Han, Shiyu Chang, Ding Liu, Mo Yu, Michael Witbrock, Thomas S. Huang
Introduction
In the problem of single-image super-resolution (SR), the aim is to recover a high-resolution (HR) image from a single low-resolution (LR) image. In recent years, SR performance has been significantly improved due to rapid developments in deep neural networks (DNNs). Specifically, convolutional neural networks (CNNs) and residual learning have been widely applied in much recent SR work .
In these approaches, two principles have been consistently observed. The first is that increasing the depth of a CNN model improves SR performance; a deeper model with more parameters can represent a more complex mapping from LR to HR images. In addition, increasing network depth enlarges the size of receptive fields, providing more contextual information that can be exploited to reconstruct missing HR components. The second principle is that adding residual connections (globally , locally or jointly ) prevents the problems of vanishing and exploding gradients, facilitating the training of deep models.
While these recent models have demonstrated promising results, there are also drawbacks. One major issue is that increasing the depth of models by adding new layers introduces more parameters, and thus raises the likelihood of model overfitting. At the same time, larger models demand more storage space, which is a hurdle to deployment in resource-constrained environments (e.g. mobile systems). To resolve this issue, the Deep Recursive Residual Network (DRRN) inspired by the Deeply-Recursive Convolutional Network (DRCN) shares weights across different residual units and achieves state-of-the-art performance with a small number of parameters.
Separate efforts in neural architectural design have recently shown that commonly-used deep structures can be represented more compactly using recurrent neural networks (RNNs). Specifically, Liao and Poggio demonstrated that a weight-sharing Residual Neural Network (ResNet) is equivalent to a shallow RNN. Inspired by their findings, we first explore the connections between the neural architectures of existing SR algorithms and their compact RNN formulations. We note that previous SR models with recursive computation and weight sharing, including DRRN and DRCN, work at a single spatial resolution (bicubic interpolation is first applied to upscale LR images to a desired spatial resolution). This enables their model structures to be represented as a unified single-state RNN. Thus, both DRRN and DRCN can be viewed as a finite unfolding in time of the same RNN structure, but with different transition functions. This is illustrated in Figure 1, and will be discussed in detail in Section 3. It is worth mentioning that we follow the terminology used in , where a “state” can be considered as corresponding to a “layer” in the normal RNN setting.
Based on this compact RNN view of state-of-the-art SR models, in this paper we explore new structures to extend the frontier of SR. The first approach in improving a conventional RNN model is generally to make it multi-layer. We apply this experience in designing the SR architecture in our compact RNN view by adding an additional state, rendering our model a Dual-State Recurrent Network (DSRN), where the two states operate at different spatial resolutions. Specifically, the bottom state captures information at LR, while the top state operates in the HR regime. As with a conventional two-layer stacked RNN, there is a connection from the bottom to the top state via deconvolutional operations. This provides information flow from LR to HR at every single unrolling time. In addition, to allow information flow from previously predicted HR features to LR features, we incorporate a delayed feedback mechanism from the top (HR) state to the bottom one. The overall structure of the proposed DSRN is shown in Figure 2, which not only utilizes parameters efficiently but also allows both LR and HR signals to contribute jointly to learning the mappings.
To demonstrate the effectiveness of the proposed method, we compare DSRN with other recent image SR approaches on four common benchmarks as well as on the DIV2K dataset from the "New Trends in Image Restoration and Enhancement workshop and challenge on image super-resolution (NTIRE SR 2017)" . Extensive experimental results validate that DSRN delivers higher parameter efficiency, low memory consumption and high restoration accuracy.
Related Work
Single image SR has been widely studied in the past few decades and has an extensive literature. In recent years, due to the fast development of deep learning, significant progress has been made in this field. Dong et al. first exploited a fully convolutional neural network, termed SRCNN, to predict the nonlinear LR-HR mapping. It demonstrated superior performance to many other example-based learning paradigms, such as nearest neighbor , sparse representation , neighborhood embedding , random forest , etc. Although all layers of a SRCNN are trained jointly in an end-to-end fashion, conceptually the network is split into three stages: patch representation, non-linear mapping, and reconstruction.
Much of the later work follows a similar network design with more complicated building blocks or advanced optimization techniques . Wang et al. proposed a sparse coding network (SCN) that encodes a sparse representation prior for image SR and can be trained end-to-end, demonstrating the benefit of domain expertise in sparse coding for image SR. Both external and self examples were utilized to synthesize the HR prediction via a neural network in .
Inspired by the success of very deep models on ImageNet challenges , Kim et al. proposed a very deep CNN, VDSR, which stacks 20 convolutional layers with kernels. Both residual learning and adjustable gradient clipping are used to prevent vanishing and exploding gradients. However, as the model gets deeper, the number of parameters increases. To control the size of the model, DRCN introduces 16 recursive layers, each with the same structure and shared parameters. Moreover, DRCN makes use of skip connections and recursive supervision to mitigate the difficulty of training. Tai et al. discovered that many residual SR learning algorithms are based on either global residual learning or local residual learning, which are insufficient for very deep models. Instead, they proposed the DRRN that applies both global and local learning while remaining parameter efficient via recursive learning. More recently, Tong et al. proposed making use of Densely Connected Networks (DenseNet) instead of ResNet as the building block for image SR. They demonstrated that the DenseNet structure is better at combining features at different levels, which boosts SR performance.
Apart from deep models working on bicubic upscaled input images, Shi et al. used a compact network model to conduct convolutions on LR images directly and learned upscaling filters in the last layer, which considerably reduces the computation cost. Similarly, Dong et al. adopted deconvolution layers to accelerate SRCNN in combination with smaller filter sizes and more convolution layers. However, these networks are relatively small and have difficulty capturing complicated mappings owing to limited network capacity. The Laplacian Pyramid Super-Resolution Network (LapSRN) works on LR images directly and progressively predicts sub-band residuals on various scales. Lim et al. proposed the Enhanced Deep Super-Resolution (EDSR) network and a multi-scale variant, which learns different scaled mapping functions in parallel via weight sharing.
Our work is also strongly related to and built upon the idea of viewing a ResNet as an unrolled RNN. It was first proposed in , which aids understanding of a family of deep structures from the perspective of RNNs. Later, Chen et al. unified several different residual functions to provide a better understanding of the design of DNNs with high learning capacity. Recently, the equivalence to RNNs has been further extended to DenseNet. Based on this finding, Dual Path Networks were proposed and showed superior performance to DenseNet and ResNet in a varity of applications.
Single-State Recurrent Networks
In this section, we first revisit the discovery that a ResNet with shared weights can be reformulated as a recurrent system. Then, based on this view, we unite the recent development of SR models with such RNN reformulations to show DRCN and DRRN are structurally equivalent to an unrolled single-state RNN.
To establish the equivalence, we adopt the commonly used definition of a RNN, which is characterized by a set of states and transition functions among the states. A RNN often consists of the input state, output state, and the recurrent states. Depending on the number recurrent states, we describe RNNs as “single-state” (i.e. one recurrent state) or “dual-state” (i.e. two recurrent states). An illustration of a single-state RNN is shown in Figure 1(a). The input, output, and recurrent states are represented as , and respectively. The arrow link indicates the state transition function. The square on the directed cycle indicates that the recurrent function travels one time step forward during the unfolding. Interested readers are referred to for detailed information on this general formulation of a RNN.
Based on Figure 1(a), we unfold along the temporal direction to a fixed length . The unfolded graph is shown in figure 1(b), and the dynamics of a single-state RNN can be characterized by:
where the upper script indicates the -th unrolling. The parameters of , , and are often time-independent, which means these parameters are reused at every unfolding step. This allows us to unify ResNet, DRCN, and DRRN as unrolled networks with the same recurrent structure but with the different realizations of and different rules of parameter sharing.
ResNet: We consider a ResNet in its simplest form without any down-sampling or up-sampling operations. In other words, both of the spatial dimensions and feature dimensions remain the same across all intermediate layers. To render Figure 1(b) equivalent to a ResNet with residual blocks, one possible technique is to make:
be the input image or a function of .
, and . Thus, the state transition becomes .
The recurrent function be the same as a conventional residual block, which contains two convolutional layers with skip connections as shown in Figure 1(c). Differences in color indicate different sets of parameters.
The prediction state be calculated only at the time as the final output.
It is worth mentioning that the only difference between an unrolled RNN following the above definitions and a conventional ResNet is that the parameters in need to be reused among all residual blocks.
DRCN: To realize the DRCN expressible by the same single-state RNN, we define and in the same way as for the ResNet. Since DRCN recursively applies only a single convolutional layer to the input feature map 16 times, with the parameters of the layer reused across the whole network, we could use a single convolutional layer to express . The graph is illustrated in Figure 1(d). Moreover, unlike the ResNet where the output is predicted only at the end of unfolding, DRCN utilizes recursive supervision, which generates an output at every unfolding . The final HR prediction of DRCN is the weighted sum of the outputs at every unfolding .
DRRN: The recurrent structure of DRRN differs only slightly from a ResNet. In a ResNet, the skip connection comes from the previous residual block, whereas in a DRRN the skip connection always comes from the first unrolled state . Figure 1(e) shows the equivalent recurrent function for a DRRN with one recursive block (i.e. ) using the definition in the original paper.
Dual-State Recurrent Networks
Drawing on the connections between state-of-the-art SR models and RNNs, we have investigated new compact RNN architectures for image SR. Specifically, we propose a dual-state design, which adopts two recurrent states enable use of features from both LR and HR spaces. The RNN view of our DSRN is shown in Figure 2(a) and is introduced as follows.
Dual-state design: Unlike single-state models working at the same spatial resolution, DSRN incorporates information from both the LR and HR spaces. Specifically, and in Figure 2(a) indicate the LR state and HR state, respectively. Four colored arrows indicate the transition functions between these two states. The blue (), orange () and yellow () links exist in a conventional two-layer RNN, providing information flow from LR to LR, HR to HR, and LR to HR, respectively. To further enable two-way information flows between and , we add the green link, which is inspired by the delayed feedback mechanism of traditional multi-layer RNNs. Here, it introduces a delayed HR to LR connection. The overall dynamics of our DSRN is given as:
Figure 2(b) demonstrates the same concept via an unfolded graph, where the top row represents HR state while the bottom one is LR. This design choice encourages feature specialization for different resolutions and information sharing across different resolutions.
Transition functions: Our model is characterized by six transition functions. , , , and as illustrated in Figure 2(b). Specifically, we use the standard residual block for both self-transitions. A single convolutional layer is used for the down-sampling transition and a single transposed convolutional (or deconvolutional) layer is used for the up-sampling transition. The strides in both inter-state layers are set to be the same as the SR upscaling factor.
Unfolding details: Similarly to unfolding a single-state RNN to obtain a ResNet, for image SR, we let have no contribution to calculating the state transition. In other words,
for any choice of (e.g. choose ). Furthermore, we set as the output of two convolutional layers with skip connections, which takes the LR input image and transform it into a desired feature space. In addition, is set to zero. Finally, we use deep supervision for the HR prediction, as discussed below.
Deep supervision: The unrolled DSRN is capable of making a prediction at every time step . Denote
as a prediction at the unfolding, where is characterized by a single convolutional layer. Then, instead of taking the prediction only at the final unfolding , we average all the predictions as
Thus, every unrolled layer directly connects to the loss layer to facilitate the training of such a very deep network. Moreover, the model predicts the residual image and minimizes the following mean square error
where is the group-truth image in HR and is the residual map between the ground truth and bicubic upsampled LR image.
Experiments
In this section, we first provide implementation details, including both model hyper-parameters and training data augmentation. Then we analyze a number of design choices and their contributions to final performance. Finally, we compare DSRN to other state-of-the-art methods on several benchmark datasets.
To evaluate the proposed DSRN algorithm, we train our model using 91 images proposed in and test on the following datasets: Set5 , Set14 , B100 and Urban100 . The training data is augmented in a similar way to previous methods , which includes 1) random flipping along the vertical or horizontal axis; 2) random rotation by 90°, 180° or 270°; and 3) random scaling by a factor from [0.5, 0.6, 0.7, 0.8, 0.9, 1]. Tensorflow is used for our full data processing pipeline; the LR training images are generated by the built-in bicubic down-sampling function. We additionally test our algorithm on the DIV2K dataset of the NTIRE SR 2017 challenge , where we use the provided training and validation sets with all of the aforementioned data augmentations except random scaling.
2 Implementation Details
We use our model to super-resolve only the luminance channel of images, and use bicubic interpolation to upscale the other two color channels, following . We train independent models for each scale (2, 3, and 4) with 64 filters on the first input convolutional layer and 128 filters in the rest of the network. All layers use convolution filters. Due to our dual-state design, the feature maps of and in each time step have the same spatial dimensions as the LR and HR images, respectively. We zero-pad the boundaries of feature maps to ensure the spatial size of each feature map is the same as the input size after the convolution is applied.
All the weights in the network are initialized with a uniform distribution using the method proposed in . We use standard stochastic gradient descent (SGD) with momentum 0.95 as our optimizer to minimize the MSE loss function in Equation (6). We search for the best initial learning rate from and reduce it by a factor of 10 three times during the entire training process. This learning rate annealing is driven by observing that the loss on the validation set stops decreasing. Gradient clipping at is adopted during training to prevent the gradient explosion. We sample image patches with a size of and use a mini-batch size of to train our network.
We observe that the recursion defined in Equation (2) may lead to an exponential increase in the scale of feature values, especially when is large. In , the authors proposed the use of unshared batch normalization at every unfolding time to resolve this issue. Batch normalization is not used in our network; we found that normalizing the scale with two scalar parameters was sufficient. Specifically, we use one unshared PReLU activation for each recurrent state after every unrolling step. All other layers have ordinary ReLU as the activation function.
3 Model Analysis
In this section, we analyze our proposed model in the following respects:
Unrolling length: The unrolling length changes the maximum effective depth of the unrolled network. In particular, for a DSRN with times unrolling, the maximum number of convolution layers between input and output of the network is . The multiplier comes from the two layers in a residual block, while the extra 4 is from the auxiliary input and output layers. However, the number of model parameters remains independent of the length of unrolling. Essentially, controls the trade-off between model capacity and computation cost. We study the influence of by training the model with different unrolling lengths. The empirical results are shown in Figure 3. The test performance increases when the number of unfolding steps increases, but the benefit seems to diminish after . Unless otherwise mentioned, we use for all our models. It is worth mentioning that we also experimented with stochastic depth by randomly sampling during training, but we observed no improvement in validation accuracy.
Parameter sharing: We empirically find parameter sharing to be crucial for training a deep recursive model. As shown in Table 2, the same model with untied weights performs much more poorly than its weight-sharing counterpart. Specifically, we observe around 0.2dB performance drop across all three upscaling scales when changing from shared weights to untied weights. We speculate that the model with untied weights suffers a larger risk of model over-fitting and much slower training convergence, both of which diminish the model’s restoration accuracy.
Dual-state and delayed feedback: We compare our DSRN with two baselines under the same unrolling time steps to understand how each module of our model contributes to the final performance: 1) a single-state RNN unrolled ResNet; and 2) a dual-state RNN without delayed feedback connections. The quantitative comparison on the NTIRE SR 2017 challenge is shown in Table 2. Comparing the single-state baseline and the DSRN without feedback, it is clear that considering information from both LR and HR spaces as two separated states provides performance gains. In addition, comparing our models with and without feedback, we realize that incorporating such an information flow from HR space back to LR space consistently improves performance on all three different scales. In all, both the dual-state and delayed feedback designs are beneficial to our model.
State visualization Since DSRN has independent scaling parameters on each unrolled state, the model implicitly learns a weighted-average of all the unrolled states for the final prediction. Empirically we observe that this strategy performs better than output from the last state only. To demonstrate how the network aggregates different unrolled states, we show feature response maps at different unrolling steps in Figure 6, demonstrating that the network distributes slightly different features to each unrolled state.
4 Comparison with the State-of-the-Art
We provide results of evaluation of our model on several public benchmark datasets in Table 1, with three commonly-used evaluation metrics: Peak Signal-to-Noise Ratio (PSNR), Structural SIMilarity (SSIM) and the Information Fidelity Criterion (IFC) . Specifically, we perform a comprehensive comparison between our method and 10 other existing SR algorithms, including both deep learning and non-deep-learning based methods. Note that many recent deep learning based competitors, including VDSR, LapSRN and DRRN, use 291 training samples with the additional 200 from the training set of Berkeley Segmentation Dataset , while our model was trained on only the 91 images. Still, our DSRN method achieves competitive performance across all datasets and scales. It achieves particularly strong performance in the and settings.
In addition, we report quantitative evaluations on the recently developed DIV2K dataset and comparisons with top-ranking algorithms in Table 2. Our method achieves competitive performance with the best algorithm, EDSR+, and outperforms all the other algorithms by a large margin, which demonstrates the effectiveness of our proposed dual-state recurrent structure.
To further analyze the proposed DSRN against other state-of-the-art SR approaches in a qualitative manner, in Figure 4 we present several visual examples of super-resolved images on Set14 with upscaling among different SR approaches. For these competing methods, we use SR results publicly released by the authors. As shown in Figure 4, our method can construct sharp and detailed structures and is less prone to generating spurious artifacts.
Furthermore, the proposed DSRN benefits from inherent parameter sharing and therefore obtains higher parameter efficiency compared to other methods. In Figure 6, we illustrate the parameters-to-PSNR relationship of our model and several state-of-the-art methods, including SRCNN, VDSR, DRCN, DRRN and RED30 . Our method represents a favorable trade-off between model size and SR performance, and has modest inference time. The DSRN takes 0.4s on the x4 task with a 288x288 output image size, on an NVIDIA Titan X GPU.
Conclusion
In this work, we have provided a unique formulation that expresses many state-of-the-art SR models as a finite unfolding of a single-state RNN with various recurrent functions. Based on this, we extend existing methods by considering a dual-state design; the two hidden states of our proposed DSRN operate at different spatial resolutions. One captures the LR information while the other one targets the HR domains. To ensure two-way communication between states, we integrate a delayed feedback mechanism. Thus, the predicted features from both LR and HR states can be exploited jointly for final predictions. Extensive experiments on benchmark datasets have demonstrated that the proposed DSRN performs favorably against state-of-the-art SR models in terms of both efficiency and accuracy. For the future work, we will explore use of our proposed DSRN to capture temporal dependencies for video SR .