Multi-Temporal Recurrent Neural Networks For Progressive Non-Uniform Single Image Deblurring With Incremental Temporal Training

Dongwon Park, Dong Un Kang, Jisoo Kim, Se Young Chun

Introduction

Blind single image deblurring is a challenging ill-posed inverse problem to recover the original sharp image from a given blurred image with or without estimating unknown non-uniform blur kernels and there has been much effort to tackle this problem. One is to simplify the given problem by assuming uniform blur and to recover both the latent ground truth image and the blur kernel . However, uniform blur is often not accurate enough to approximate the actual blur, and thus there has been much research on non-uniform blur by extending the degree of freedom of the blur model from uniform to non-uniform in a limited way compared to the dense matrix . Other non-uniform blur models have been investigated such as additional segmentations within which simple blur models were used or motion estimation based deblurs .

Recently, deep-learning-based approaches for single image deblurring have been proposed with excellent quantitative results and with fast computation time. There are largely two different ways of using deep neural networks (DNNs) for deblurring. One is to use DNNs to explicitly estimate non-uniform blurs and the other is to use DNNs to directly estimate the original sharp image without estimating blurs . Most state-of-the-art methods such as are estimating the original sharp image directly from the given blurred image (see Figure 1). Single-scale (SS) or one-stage approaches (Figure 2 (a)) are frequently used, but many state-of-the-art methods are using multi-scale (MS) approaches (or coarse-to-fine) with down-scaled image(s) in spatial domain.

The MS approaches utilize down-scaled images to restore the latent sharp image progressively over scales as illustrated in Figure 2 (b). This approach makes use of the fact that blurs become relatively smaller as scale of image decreases . Thus, a DNN with MS approach is able to perform deblurring from large blur to small blur progressively. Recently, there have been some works on sharing network parameters of MS structures over scales or efficiently sharing parameters except for feature extraction layers that are independent over scales assuming that blurs are varied with scales . One drawback of typical MS approaches seems to lose much high-frequency information during down-sampling in a sub-optimal way for image deblurring considering the fact that strong edge information is important for reliable deblurring .

In this paper, we investigate an alternative approach, called multi-temporal (MT) approach, to MS approach for single image deblurring. Instead of using down-scaled images, we exploit a typical dataset generation pipeline using high-speed camera to construct a blurred image by averaging multiple frames of images such as the GoPro dataset . We conjecture that recovering sharp latent images from mild blurs is easier than recovering them from severe blurs and propose DNNs to deblur little by little as illustrated in Figure 2 (c). Thus, our MT approach allows to use full image information in the original scale for reliable deblurring and to deblur for mild blurs at each iteration progressively (see Figure 3) for potentially better performance than SS approach. Without any special parameter sharing schemes like , our proposed methods achieved state-of-the-art performance with the smallest number of parameters as shown in Figure 1.

Here is the summary of our contributions: 1) proposing MT approach with incremental temporal training for high-speed camera dataset to divide challenging severe blur into a series of mild blurs and then deblur each mild blur progressively, 2) developing MT-recurrent neural network (RNN) with recurrent feature maps for blind single image deblurring, and 3) achieving state-of-the-art performance on GoPro dataset with the smallest number of parameters among recently proposed blind single image deblurring methods.

Related Works

Conventional approaches to blind single image / video deblurring usually require to explicitly estimate blur kernels. There have been works on estimating uniform blurs using optimization with MS approach , using a model of the spatial randomness of noise and a local smoothness prior , exploiting blurred strong edges to reliably estimate blur kernel , and developing a metric to measure the usefullness of image edges for blur kernel estimation .

There have also been many works on predicting non-uniform blurs assuming spatially linear blur , simplified camera motion , parametrized geometric model in terms of camera rotation velocity during exposure , filter flow framework based blur model , l0l_{0} sparsity for blurs , and dark channel prior . There was also an attempt to exploit multiple images from videos assuming spatially varying blur . There have also been some works to utilize segmentation information by assuming uniform blur on each segmentation area and to segment motion blur using optimization , to simplify motion model as local linear without segmentation using MS approach , and to use bidirectional optical flows for video deblurring .

Recently, many blind single image / video deblurring works employed DNNs for estimating blur kernels and/or original sharp images from given blurred input images. There are several works to predict non-uniform blur kernels explicitly: predicting the probabilistic distribution of motion blur at the patch level , estimating the complex Fourier coefficients of a deconvolution filter , performing blur kernel estimation by division in Fourier space from extracted deep features , and analyzing the spectral content of blurry image patches by reblurring them .

There are also many works to directly estimate the original sharp image from the given blurred input image without explicitly estimating non-uniform blur kernels. For video blind deblurring, there have been some works to exploit temporal information: blending temporal information in spatio-temporal recurrent network for online video deblurring , taking temporal information into account with recurrent deblur network consisting of several deblur blocks , and developing an encoder-decoder network with the input of multiple video frames to accumulate information across frames .

There are a few works for blind single image deblurring without temporal information. Xu et al. proposed a direct estimation of the original sharp image based on optimization to approximate deconvolution by a series of convolution steps using DNNs . Later, Nah et al. proposed a MS network architecture with Gaussian pyramid and MS loss functions and Tao et al. proposed convolution long short-term memory (LSTM)-based MS DNN for single image deblurring . Gao et al. proposed MS parameter sharing and nested skip connections . Zhang et al. proposed a deep multi-patch hierarchical network for different feature levels on the same resolution . Aljadaany et al. proposed a learning both the image prior and data fidelity terms for single image deblurring . Kupyn et al. proposes generative adversarial network (GAN) framework based on feature pyramid network (FPN) and relativistic discriminator with a least-square loss .

Lastly, RNN plays an important role in using sequential data or iterative approach. Zhou proposed spatio-temporal variant RNN for video deblurring. RNN is often introduced to utilize previous frames effectively such as previous features using convolutional LSTM . Similarly, we propose an approach that recurrently makes use of previous feature information for each iteration. However, unlike other RNN based approaches, our proposed methods use incremental temporal training procedure that does not train from the most severe blur to the ground truth, but trains from more blured to less blurred image incrementally.

Multi-Temporal (MT) Approach

The GoPro dataset consists of 15,000 sharp images (frames), captured by GoPro4 Hero Black camera (240 frame per sec), including 22 videos for training and 11 videos for testing . 7-13 frames were averaged to yield blur-sharp image pairs where a middle image among multiple frames was selected as a ground truth as in Figure 4. Temporal level (TL) NN is defined to be a blurred image from NN frames. The GoPro dataset contains TL 7-13.

2 Dataset For Incremental Temporal Training

For the GoPro dataset with TL 1 (ground truth) and TL 7-13 pairs, we further generated data for MT approach and incremental temporal training. For example, for a blurred image with TL 7, we generated intermediate blurred images with TL 1-13 as shown in Figure 4. Thus, our MT approach does not try to estimate TL 1 from TL 7 directly, but tries to estimate from TL 7 to TL 5, TL 5 to TL 3, and finally TL 3 to TL 1, progressively.

We quickly validated our conjecture for MT approach: will it be easier to estimate TL 1 from TL 7 than to estimate TL 1 from TL 5 or TL 3? Table 1 shows the performance of U-Net that was trained only with one TL images for TL 3-13. As TL increases, PSNR clearly decreases. Thus, our conjecture for MT approach seems reasonable.

3 Incremental Temporal Training

Our training method is based on the dataset with more intermediate TL images as explained in Section 3.2. During training, our proposed network is recurrently iterated by 5 or 7. At iteration 1, we train the network with randomly selected temporal blurred images (TL 13 or 11 or 9 or 7) as inputs and desired temporal blurred images (TL 11 or 9 or 7 or 5) as ground truth, respectively. Note that the TL difference between input and ground truth is 2. At the next iteration, the estimated image from iteration 1 is taken as input and desired temporal blurred image (TL 9 or 7 or 5 or 3) as ground truth. Similarly, other iterations are processed sequentially and take the estimated image from previous iteration as input and corresponding desired blurred or sharp image as ground truth. Finally, 1-3 more iterations of training to TL 1 as ground truth are repeated. If the number of iterations is over 7, return to iteration 1 for training the model. Note that model parameters are shared and training is performed independently for each iteration.

4 Progressive Deblurring With MT Approach

The methods of Tao and Nah are based on MS approach for deblurring. The DNN of Tao shares parameters over scales that can be modeled as follows:

We propose MT-RNN with recurrent feature maps using temporal iterations that can be modeled as follows:

5 Proposed MT-RNN With Feature Maps

Our proposed network is based on the network of Tao . Base model is U-Net architecture and consists of encoders and decoders as illustrated in Fig. 5. Each stage has 1 feature extraction layer and residual blocks (Resblocks) that is identical to the Resblock in that consists of 32 channels, 64 channels and 128 channels at the top, middle and bottom encoder-decoders, respectively.

Residual learning Kupyn and Zhou utilize residual learning for deblurring. Both of them produce enhanced image and the network learns a residual image IRI_{R} to correct the blurred image IBI_{B} where Ideblur=IB+IRI_{deblur}=I_{B}+I_{R}. Residual learning is efficient to train the network faster and resulting model generalizes better. In the deblurring problem, input and output are highly correlated. Therefore, the residual learning helps training the network.

We conducted an ablation study for residual learning. In Figure 5, our proposed network takes I0I^{0} and I^i−1\hat{I}^{i-1} as input and residual skip connection is linked with I0I^{0}. Two cases for IBI_{B} was considered: I0I^{0} and non residual skip. PSNR of connection with I0I^{0} is higher than non residual learning by 0.15dB on the GoPro dataset with intermediate TL images.

Recurrent feature maps As shown in Figure 5, recurrent features Fi−1F^{i-1} are from the last ResBlock of each decoder and are concatenated with the feature maps of previous encoder at feature extraction layer:

where fif^{i} is the feature map of previous encoder at the iith iteration. Estimated image I^i−1\hat{I}^{i-1} is concatenated with I0I^{0}:

and then the encoder takes the IcatiI_{cat}^{i} and FencoderiF_{encoder}^{i} as input.

Tao utilized convolutional LSTM for passing intermediate feature maps to the next spatial scale stage. Nah also makes use of hidden state ht−1h_{t-1} in RNN cell. Similarly, our network uses intermediate feature maps Fi−1F^{i-1} from decoder that may include information about blur patterns and intermediate results for IiI^{i}. Thus, Fi−1F^{i-1} is utilized to encode IiI^{i}, having more details of blur patterns and other information for deburring. Using recurrent feature maps Fi−1F^{i-1} improves performance by 0.31dB.

Loss Function We use L1L1 loss function that measures the difference between a restored image and its corresponding latent ground truth normalized by channel, height and width of image. Ground truth images consists of TL 1-11 images.

6 Convergence of MT-RNN over Iterations

Determining the number of iterations for MT-RNN is important for performance. We studied iteration vs. PSNR for the network that was trained only with one type of TL images (e.g., TL 13) for all TL 7, 9, 11, 13. Training was performed until the 7th iteration for all cases. As illustrated in Figure 6, all networks yielded increased PSNR over iterations until 5th or 6th iterations, and then decreased PSNR beyond training iterations. From training procedure, iteration 6 was chosen and it was applied to all experiments for our methods. Note that in all cases with different TL images, our proposed MT-RNN methods outperform state-of-the-art MS methods (Tao ).

Experiments

The GoPro dataset consists of 3214 blurred images with the size of 1280×\times720 that are divided into 2103 training images and 1111 test images. In both validation and test sets, TL 7, 9, 11, 13 images were evenly distributed. We generated more intermediate TL images along with the GoPro dataset so that this new dataset consists of 5500 training, 110 validation and 1200 test images.

2 Implementation Details

We implement the proposed network on pytorch . For fair comparisons, we evaluate our proposed method and state-of-the-art methods on the same machine with NVIDIA Titan V GPU. During training, Adam optimizer was used with learning rate 2×10−42\times 10^{-4}, β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999, and ϵ=10−8\epsilon=10^{-8}. For Tables 2, 3, 4 and 6, total iteration was 92×10392\times 10^{3} with reducing learning rate by half every 46×10346\times 10^{3} iterations and the GoPro dataset with additional intermediate TL images was used. PSNR/SSIM were evaluated on python. Unlike others, for Table 5, total iteration is 46×10446\times 10^{4} with reducing learning rate by half every 46×10346\times 10^{3} iterations and the GoPro dataset was used. PSNR/SSIM were evaluated on MATLAB. Patch size was 256×\times256. Random crop, horizontal flip, and 90∘ rotation were used for data augmentation. Note that since the number of channel is changed by concatenation to replace add operation of skip connection in , 1×\times1 convolution was used.

3 Ablation Studies for Model Architecture

We performed ablation studies from the base model (Tao ) by adding our proposed components such as residual leaning and recurrent feature map. Note that while Tao is spatially iterative due to MS approach, our proposed MT-RNN is temporally iterative due to our proposed MT structure. Table 2 shows PSNR (dB), SSIM and number of parameter(Million) (denoted by Parm) along different components such as Approaches including MS approach or our proposed MT approach, size of kernel (denoted by K), residual learning (denoted by R), recurrent feature map (denoted by F). As a baseline MS network, Tao was used as shown in Table 2 (a). Changing kernel size from 5 to 3 resulted in improved performance by 0.13dB, and substantially decreased parameter size as in Table 2 (b). Using residual learning instead of direct learning also improved performance as in Table 2 (c) due to the effect of specified blurry region on back-propagation. Our proposed MT approach improves performance over conventional MS approach using the same DNN as shown in Table 2 (d). Further improvement was observed when using recurrent feature map as in Table 2 (e). Lastly, it turns out that using MT alone is better than using both MS and MT in performance. Thus, temporal iterative approach helps the network to achieve high performance as in Table 2 (f).

4 Studies on Temporal Steps and Parameters

We studied the effect of temporal steps on performance as shown in Table 3 (g), (h) and (e) with one-stage SS approach (0 temporal step), 2 and 4 temporal steps, respectively. Using MT approach yielded better performance than SS approach and using small temporal step was more advantageous than using larger temporal step in performance even though computation time was increased. Thus, we chose temporal step 2 as the best step size.

Then, we also investigated the effect of parameter size on performance. Table 3 (e), (j), (k) show that the number of parameters is proportional to performance with the cost of increased computation. While twice larger parameters in (k) did not improve performance much over (e), its computation time and memory were substantially increased. Using half the parameter size in (j) did degrade performance substantially while computation speed of (j) is similar to (e).

5 Our MT Approach to Other Deblur DNNs

We investigate the feasibility of applying our proposed MT approach and incremental temporal training to other state-of-the-art deblurring methods such as Kupyn and Zhang where the performances for them are reported in Table 4 (l) and (n), respectively. Table 4 (m) and (o) are performance results when using our proposed MT approach with incremental temporal training in Kupyn and Zhang, respectively. In both cases, our proposed approach successfully increased performance over the baselines.

6 Benchmark Results

We performed studies on the GoPro dataset for benchmarking. Tables 5 presents quantitative results of our proposed methods and other state-of-the-art methods. Our proposed model, MT-RNN (e) were trained with the GoPro dataset along with intermediate TL images and achieve the best result (31.15 dB in PSNR) over other previous state-of-the-art methods on the GoPro test dataset (1111 images). Our MT approach for the network of Zhang (o) also improved performance over the original network of Zhang.

Figure 7 shows qualitative evaluation in the case of four different models which are our proposed method (last row), the work of Nah (2nd row), the work of Tao (3rd row) and the work of Zhang (4th row) for given blurred images (1st row). Qualitative results show that our proposed method outperforms other state-of-the-art methods visually.

Discussion

On Imperfect Ground Truth Videos and images from high speed cameras often have mild blur assuming your subjects or objects move quickly. Thus, obtaining perfect ground truth for single image deblurring problems is quite challenging. Considering imperfect ground truth scenarios for single image deblurring, we perform experiments to observe the behaviors of our proposed MT approaches and conventional MS approaches as illustrated in Figure 8.

During training, we used TL 1, 3, or 5 as ground truth and TL 7 - 13 as input images. Two approaches, MS and MT methods, were applied to these simulations: using TL 3 as ground truth and using TL 5 as ground truth for training.

Table 6 shows that both MT and MS approaches yielded excellent performance assuming known perfect ground truth. However, our proposed MT method yielded about 0.5dB better PSNR than the MS method. When TL 3 images are given as ground truth, both MT and MS methods still yielded good performance. In this case, our MT approach yielded better performance than conventional MS approach, that is also consistent with other performance comparison results in this paper. One of the possible explanations on these results is that TL 3 images are already good enough as ground truth. When TL 5 images are used as ground truth, the performance difference between MS and MT approaches became larger than other cases. Thus, these preliminary results suggest that our proposed MT approaches may be more robust to imperfect ground truth dataset for deblurring than MS approaches.

Decreasing PSNR Beyond Trained Iterations In Figure 6, MT-RNN yielded increasing PSNR during early iterations (usually, before 6 or 7 iterations) and then yielded decreasing PSNR later iterations. To study the reasons for decreasing PSNR after stopping point, we visually investigated deblurred images from our proposed methods. In images, we observed that there are often tiny artifacts appearing near the center of images. Then, as iteration increases, artifacts grows rapidly and they significantly

Computation Time for “Ours-Z” In the Table. 5, Ours-Z takes 2.08 seconds and iterates 6 times, while Zhang takes 0.02 seconds without any iteration. Generally, running time of Ours-Z is expected around 0.02 seconds, but it takes 2.08 seconds in reality. For analyzing this issue, we measure the iteration time with one blurred image. The running time on a example image is 0.015 sec, 0.092 sec, 0.581 sec, 1.073 sec, 1.568 sec and 2.065 sec in the order of iterations. Actually, the first iteration is similar with 0.02 seconds. However, after that, the running time increases exponentially. Further investigation on this issue is necessary such as looking into GPU related issues.

Weight Sharing There are a few works on DNN based MS single image deblurring that share network weights across different scales in MS architecture or that partially share network weights (except for feature extraction layers) so that the number of parameters is reduced significantly while performance is not degraded. Note that our MT approach is similar to weight sharing across temporal iterations. However, partial shared parameters that may be much more efficient were not be investigated in MT structure. Thus, it will be interesting to further investigate partial weights schemes for MT approaches.

Conclusion

In this work, we investigate alternative approach to MS, called multi-temporal (MT) approach, for non-uniform single image deblurring. We propose incremental temporal training with constructed MT level dataset from time-resolved dataset, develop novel MT-RNNs with recurrent feature maps, and investigate progressive single image deblurring over iterations. Our proposed MT methods outperform state-of-the-art MS methods on the GoPro dataset in PSNR with the smallest number of parameters.

Acknowledgments

This work was supported partly by Basic Science Research Program through the National Research Foundation of Korea(NRF) funded by the Ministry of Education(NRF-2017R1D1A1B05035810), the Technology Innovation Program or Industrial Strategic Technology Development Program (10077533, Development of robotic manipulation algorithm for grasping/assembling with the machine learning using visual and tactile sensing information) funded by the Ministry of Trade, Industry & Energy (MOTIE, Korea), and a grant of the Korea Health Technology R&D Project through the Korea Health Industry Development Institute (KHIDI), funded by the Ministry of Health & Welfare, Republic of Korea (grant number: HI18C0316).

References