Rethinking Coarse-to-Fine Approach in Single Image Deblurring
Sung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung, Sung-Jea Ko
Introduction
Single image deblurring aims to recover a latent sharp image from a blurry image . Even with the rapid development of camera modules in the last few decades, blur artifact still exists when camera and/or objects move. Blurry images are not only visually unpleasant but significantly degrade the performance of vision systems including surveillance and autonomous driving systems , necessitating accurate and efficient image deburring techniques.
Owing to the success of deep learning, convolutional neural network (CNN)-based image deblurring methods have been extensively studied and showed promising performance. Early CNN-based image deblurring methods commonly exploit CNN as a blur kernel estimator and construct two-stage image deblurring framework, i.e., CNN-based blur kernel estimation stage and kernel-based deconvoltion stage. On the other hand, recent CNN-based image deblurring methods aim to directly learn the complicated relationship between blurry-sharp image pairs in an end-to-end manner. As a pioneering technique, a deep multi-scale CNN for dynamic scene deblurring (DeepDeblur) is introduced to directly regress a sharp image from a blurry image. DeepDeblur consists of multiple stacked sub-networks to handle multi-scale blur, where each sub-network takes a down-scaled image and gradually recovers a sharp image in a coarse-to-fine manner. Motivated by the success of DeepDeblur, various CNN-based image deblurring methods have been introduced with remarkable performance improvements. Although these methods try to improve the deblurring performance in different aspects, their coarse-to-fine strategies are similar in that multiple sub-networks are stacked. In other words, a coarse-to-fine network design principle has proven to be effective in image deblurring. However, such efficiency comes at the cost of the inevitable increase in the computational complexity and memory usage, making the conventional methods difficult to be used for cost and time-sensitive environments such as mobile devices, vehicles, and robots. Recently, a light-weight CNN is presented for efficient single image deblurring . Specifically, by using optical flow and global motion of blurry images as extra supervision for network training, they design a shallower architecture compared to that of conventional deblurring networks. However, such shallow architecture failed in obtaining deblurring accuracy comparable to state-of-the-art methods.
In this paper, we revisit the coarse-to-fine scheme and present a novel deblurring network called multi-input multi-output UNet (MIMO-UNet) that can handle multi-scale blur with low computational complexity. The proposed MIMO-UNet is a single encoder-decoder-based U-shaped network that has three distinct features.
First, the single decoder of the MIMO-UNet outputs multiple deblurred images, and therefore we name our decoder as multi-output single decoder (MOSD). The MOSD is simple but can mimic conventional network architectures composed of stacked sub-networks and guide the decoder layers to gradually recover latent sharp images in a coarse-to-fine manner. Second, the single encoder of the MIMO-UNet takes multi-scale input images; thus, our encoder is called multi-input single encoder (MISE). Last, asymmetric feature fusion (AFF) is introduced to merge multi-scale features in an efficient manner. The AFF takes features from different scales and merges multi-scale information flow across the encoder and the decoder to improve the deblurring performance. Extensive experiments demonstrate the superiority of the proposed MIMO-UNet compared to the state-of-the-art methods in terms of the PSNR as well as the computational complexity as shown in Figure 1.
Related works
In this section, we review the conventional image deblurring methods that adopt a coarse-to-fine strategy.
As a pioneering work, DeepDeblur directly learns the relation between blurry-sharp image pairs in an end-to-end manner by adopting a coarse-to-fine strategy . Nah et al. also introduced the real-world image deblurring dataset named the GoPro dataset. Specifically, using a sequence of sharp images captured at 240 fps using a GoPro camera, a blurry image, , is obtained by averaging successive sharp images as follow:
where and represent the number of sampled sharp images and the sharp image, respectively. To construct a blurry and sharp image pair for training, the ground-truth sharp image for is chosen by selecting the middle image from the sampled sharp images.
To adopt a coarse-to-fine strategy in CNN for gradual recovery of latent sharp images, DeepDeblur uses multiple stacks of sub-networks as shown in Figure 2(a). Each sub-network consists of a sequence of convolutional layers that maintains the spatial resolution of input feature maps. Different scales of input images are fed into the sub-networks, and the resultant image from a coarser scale sub-network is concatenated with the input of a finer scale sub-network to enable coarse-to-fine information transfer. The reconstruction procedure of DeepDeblur is formulated as follows:
2 PSS-NSC
Inspired by the success of DeepDeblur, Gao et al. presented parameter selective sharing and nested skip connections (PSS-NSC) . As shown in Figure 2(b), the architecture of PSS-NSC is similar to that of DeepDeblur, but has two distinct features. First, each sub-network is structured as an encoder-decoder-based U-Net with symmetric skip connections that directly transfers the feature maps from the encoder to the decoder. Second, since every sub-network commonly aims to recover a sharp image from a blurry image, most network parameters are shared among sub-networks. Therefore, the memory requirement of PSS-NSC is significantly reduced, but the computational complexity is still demanding because the final sharp image is generated after passing through the three sub-networks. The reconstruction procedure of PSS-NSC is formulated as follows:
3 MT-RNN
The network architecture of multi-temporal recurrent neural networks (MT-RNN) is illustrated in Figure 2(c). In MT-RNN, a single U-shaped network is repeated seven times, and the feature maps from the decoder at the previous iteration are transferred to the encoder at the next iteration as green colored arrows. For each iteration, MT-RNN is trained to predict an averaged image obtained using a different number of in Eq. 1, where decreases as the iteration proceeds. Due to the repeated application of a single U-shaped network, MT-RNN has low memory usage but low runtime efficiency. The reconstruction procedure of PSS-NSC is formulated as follows:
Proposed method
We propose MIMO-UNet that fully exploits multi-scale features extracted from an input image. Figure 3 shows the overall architecture of MIMO-UNet. The architecture of MIMO-UNet is based on a single U-Net with significant modifications for efficient multi-scale deblurring. The encoder and decoder of MIMO-UNet are composed of three encoder blocks (EBs) and decoder blocks (DBs). The following subsections detail the three special features of MIMO-UNet, i.e., MISE, MOSD, and AFF.
It has been demonstrated that different levels of blur in images can be better handled from multi-scale images . Various CNN-based deblurring methods have also adopted this idea by taking a blurry image with a different scale as an input of each sub-network .
In our MIMO-UNet, not a sub-network but an EB takes a blurry image with a different scale as an input. In other words, in addition to the downsized feature extracted from the above EB, we extract the feature from the downsampled blurry image and then combine both features. By taking advantage of the complementary information from the downsized feature and the feature obtainable from the downsampled image, our EB is expected to handle diverse image blurs effectively. The use of multi-scale images as an input for a single U-Net has also proven to be effective in other tasks such as depth map super-resolution and object detection .
We first extract the features from the downsampled image using a shallow convolutional module (SCM) as shown in Figure 4(a). Considering efficiency, we use two stacks of and convolutional layers. We concatenate the features from the last layer with the input , and further refine the concatenated features using an additional convolutional layer. The output of the SCM at the level is denoted as , where we use SCM for the second and third levels as shown in Figure 3.
For the fusion of with the output of the level EB, , we apply a convolutional layer with a stride of 2 to , resulting in . The two features and have the same size and thus can be fused. Here, we exploit a feature attention module (FAM) to actively emphasize or suppress the features from the previous scale and learn the spatial/channel importance of the features from SCM. We experimentally demonstrate that this module increases the performance compared to general feature fusion approaches as detailed in Sec. 4.3.
In particular, and are element-wise multiplied with each other, and then the multiplied features are passed through a convolutional layer. The output of the convolutional layer is expected to include complementary information for deblurring, and finally added to to be further refined through following residual blocks, where we used eight modified residual blocks .
2 Multi-output single decoder
In MIMO-UNet, different DBs have feature maps with different sizes. We consider that these multi-scale feature maps can be used to mimic multi-stacked sub-networks. Unlike the intermediate supervision at the sub-network as the conventional coarse-to-fine networks, we apply the intermediate supervision to each DB. The image reconstruction in each level can be formulated as follows:
3 Asymmetric feature fusion
In most conventional coarse-to-fine image deblurring networks, only the features from the coarser-scale sub-network are used for the finer-scale sub-networks, making information flow inflexible. One exceptional method is to cascade the whole network in horizontal or vertical direction, allowing top-to-bottom and bottom-to-top information flow .
Inspired by dense connection between intra-scale features , we present an asymmetric feature fusion (AFF) module as shown in Figure 4(c) to allow information flow from different scales within a single U-Net. Each AFF takes the outputs of all EBs as an input and combines multi-scale features using convolutional layers. The output of the AFF is delivered to its corresponding DB. More specifically, the first-level and second-level AFFs, and , are formulated as follows:
4 Loss function
Likewise with other multi-scale deblurring networks, we use the multi-scale content loss function , where we found that L1 loss produces better results than MSE loss for our network. The content loss is defined as follows:
where is the number of levels. We divide the loss by the number of total elements for normalization.
Recent studies also suggest the auxiliary loss terms in addition to the content loss for the performance improvement . In image enhancement and restoration tasks, auxiliary loss terms that minimize the distance between the input and output in the feature space have been widely used and showed promising results . Since the purpose of deblurring is to restore the lost high-frequency component, it is essential to reduce the difference in the frequency space. To this end, we present multi-scale frequency reconstruction (MSFR) loss function. The MSFR loss measures the L1 distance between multi-scale ground-truth and deblurred images in the frequency domain as follows:
where denotes the fast Fourier transform (FFT) that transfers image signal to the frequency domain. The final loss function for training our network is determined as follows:
where we experimentally set .
Experiments
We used the GoPro and RealBlur training datasets for training our models which consist of 2,103 and 3,758 pairs of blurred and sharp images. The GoPro and Real blur test datasets were used for testing, where the number of image pairs are 1,111 and 980, respectively. For testing on the GoPro test dataset, we trained our model using only the GoPro training dataset.
For every training iteration, we randomly sampled four images and then randomly cropped the sampled images with the size of . For data augmentation, each patch was horizontally flipped with a probability of 0.5. For deblurring of images in the GoPro dataset, we trained our network for 3,000 epochs which were sufficient for convergence. The learning rate was initially set to and decreased by the factor of 0.5 at every 500 epochs. For deblurring of images in the RealBlur dataset, we trained our network for 1,000 epochs, and used the same initial learning rate but decreased it by the factor of 0.5 at every 200 epochs. Our experiments were conducted on Intel i5-8400 and NVIDIA Titan XP.
2 Performance comparison
We compared MIMO-UNet with state-of-the-art deblurring networks . Considering the trade-off between the computational complexity and deblurring accuracy, we evaluated the following three variants of MIMO-UNet: 1) MIMO-UNet employing 8 residual blocks for each EB and DB, 2) MIMO-UNet+ employing 20 residual blocks for each EB and DB, and 3) MIMO-UNet++ estimating the resultant image using MIMO-UNet+ with geometric self-ensemble . The quantitative results on the GoPro test dataset are reported in Table 1. For a fair comparison, the runtime of the models is provided as the runtime measured using the released test code of each model on our PC (left) and the runtime reported in each paper (right).
MIMO-UNet+ and MIMO-UNet++ were slower than MIMO-UNet but still performed deblurring in 0.014s and 0.040s, respectively. The average PSNR of MIMO-UNet++ was obtained as 32.68 dB. MIMO-UNet showed the average processing time of 0.008s and the average PSNR of 31.73 dB. These three models demonstrate the best trade-off between the accuracy and computational complexity as shown in Figure 1.Runtime measured using GPU synchronization mode can be find in on our website. Due to the stacked sub-networks, DeepDeblur, SRN, PSS-NSC, DMPHN, and SAPHN required large computational costs as shown in Table 1. Compared with these methods, MIMO-UNet+ was faster but achieved still higher PSNR scores. Although SRN, PSS-NSC, and MT-RNN employ fewer parameters than the proposed methods, these methods repetitively use parameters in the procedure, and therefore they are slower than the our slowest model MIMO-UNet++. Especially, MIMO-UNet++ was 4.05 times faster and 0.02 dB higher in terms of PSNR compared to MPRNet that is the best method among the conventional methods. The single network-based methods, such as RADN and SVDN, achieved high runtime efficiency compared to the stacked sub-networks. However, MIMO-UNet outperforms SVDM, and MIMO-UNet+ outperforms RADN, in terms of both runtime and PSNR. To validate the effectiveness of the proposed method on the real case scenario, we also evaluated our methods on the recent RealBlur dataset . As listed in Table 2, MIMO-UNet++ recorded the best and the second best performance in terms of PSNR and SSIM, respectively. The several resultant images from the GoPro and RealBlur test datasets are shown in Figure 5 and Figure 6, respectively. For the reproduction of results, we used the author-released network models trained on each dataset, i.e., SRN, PSS-NSC, DMPHN, MT-RNN, and MPRNet were used for the GoPro dataset, and DeblurGAN-v2 and SRN for the RealBlur dataset, respectively. Although the resultant images obtained by the conventional networks exhibit much less blur compared to the input blurry images, local details and structures were not sufficiently deblurred as can be noticed from the magnified image regions, whereas our method produced sharper images.
3 Ablation study
We conducted experiments to analyze the effectiveness of each component of MIMO-UNet on the GoPro test dataset. First, we evaluated the effectiveness of different feature fusion methods in MISE. The proposed FAM was compared with the conventional fusion methods: concatenation and element-wise sum, and achieved the highest performance as listed in Table 3. Second, we tested MIMO-UNet without MOSD, MISE, AFF, and/or MSFR. For comparison, a baseline model was trained without using any of the four components, resulting the average PSNR of 31.16 dB. As shown in Table 4, compared with the baseline model, MOSD improved PSNR by 0.17 dB. The standalone use of MISE showed a marginal effect because multi-scale information is difficult to be used in a simple U-Net. However, when used with MOSD, MISE contributed to the further performance improvement of PSNR by 0.05 dB. AFF improved PSNR by 0.17 dB compared to the baseline model, and the performance gain was further increased to 0.23 dB when AFF was used with MISE. With MISE, MOSD, and AFF, the network achieved 0.30 dB higher PSNR, and finally, the network trained using MSFR achieved 0.57 dB higher PSNR compared to the baseline.
4 Object detection performance evaluation
Single image deblurring can also boost the performance of computer vision tasks when used as a preprocessing technique. Object detection is one of the best examples in which single image deblurring can be used to improve the performance. With the advances in the CNNs, object detection methods have adopted CNNs and achieved significant improvements . However, most of these methods assume blur-free input images, and therefore they often fail to detect objects in blurry images. Figure 7(a) illustrates the failure case of PFPNet , which is one of the state-of-the-art object detectors, in detecting objects from a blurry image, depicting its vulnerability to blurry inputs. When the same PFPNet was applied to the deblurred image obtained using MIMO-UNet++, many of the false negative examples could be successfully detected as shown in Figure 7(b).
Last, we compared the proposed MIMO-UNet++ with the other deblurring techniques in terms of their effectiveness in the object detection task as preprocessing. Similar to the previous experiment, PSS-NSC and DMPHN with the author-provided codes were used for comparison. Although PFPNet was trained using the PASCAL VOC dataset that contains 20 different classes, the blurry images in the GoPro dataset primarily contain only three classes among them, i.e., car, person, and potted plant. Therefore, the average precision (AP) of each object class was measured for the performance evaluation. As shown in Figure 8, the proposed MIMO-UNet++ resulted in the best performance in object detection. Moreover, since the proposed method recorded the fastest execution time, it is most suitable as a preprocessing technique for object detection.
Conclusion
In this paper, we proposed a fast and accurate image deblurring network. Instead of stacking multiple sub-networks for coarse-to-fine deblurring, we presented a single U-Net that has distinct features, enabling much simpler but more effective coarse-to-fine deblurring. The encoder of the network is modified to take multi-scale input images and combine features from different sources. The decoder of the network is also changed to output multi-scale deblurred images during decoding such that coarse-to-fine deblurring can be better performed. A feature fusion method is also introduced to asymmetrically combine multi-scale features for dynamic image deblurring. The experimental results demonstrate that our method outperforms the other conventional methods in regard to the speed and accuracy trade-off.
Acknowlegement
This work was supported by Samsung Electronics Co., Ltd (IO201210-08026-01)