Single Image Defocus Deblurring Using Kernel-Sharing Parallel Atrous Convolutions
Hyeongseok Son, Junyong Lee, Sunghyun Cho, Seungyong Lee
Introduction
Defocus blur of an image occurs when the light ray from a point in the scene forms a circle of confusion (COC) on the camera sensor. The aperture shape and lens design of the camera determine the blur shape, and the blur size varies upon the depth of a scene point and intrinsic camera parameters. In a defocused image, the spatial variance of the blur size is large, while that of the blur shape is relatively small. Single image defocus deblurring remains a challenging problem as it is hard to accurately estimate and remove defocus blur spatially varying in both size and shape.
The conventional two-step approach reduces the complexity of defocus deblurring by assuming an isotropic kernel for the blur shape, such as disc or Gaussian . Based on the assumption, the approach first estimates a defocus map containing the per-pixel blur size of a defocused image, then uses the defocus map to perform non-blind deconvolution on the image. However, real-world defocused images may have more complex kernel shapes than disc or Gaussian, and the discrepancy often hinders accurate defocus map estimation and successful defocus deblurring.
Recently, Abuolaim and Brown proposed DPDNet, the first end-to-end defocus deblurring network, that learns to directly deblur a defocused image without relying on a restrictive blur model. They also presented a defocus deblurring dataset that includes stereo images attainable from a dual-pixel sensor camera. Thanks to the end-to-end learning and the strong supervision provided by the dual-pixel dataset, DPDNet outperforms two-step approaches on deblurring of real-world defocused images. Still, the deblurred results tend to include ringing artifacts or remaining blur, as the conventional encoder-decoder architecture of DPDNet confines its capability in handling spatially variant blur .
In this paper, we propose a novel deep learning approach for single image defocus deblurring based on inverse kernels. It was shown that deconvolution of an image with a given blur kernel can be performed by convolving the image with an inverse kernel , where the inverse kernel could be computed from the given blur kernel using Fourier transform. Xu et al. trained a deep network to learn uniform deconvolution by introducing the property of pseudo inverse kernel into the network. Similarly, we train our network to learn the deconvolution operation by capitalizing the specific characteristics of inverse kernels needed for defocus deblurring. However, due to the spatially varying nature of defocus blur, the inverse kernel required for defocus deblurring also changes per pixel. Training a deep network to learn deconvolution operations of varying inverse kernels would be challenging even with the guidance of defocused and sharp image pairs.
To reduce the complexity, we use the property of defocus blur that the blur shape is similar in a defocused image although the blur size can drastically change. However, instead of assuming any specific blur shape as in the two-step approach, we exploit our observation on inverse kernels; When only the size of a blur changes while keeping the shape, the shape of the corresponding inverse kernel remains the same and the size changes in the same way as the blur (Sec. 3.1). Then, we may constrain our network to simulate inverse kernels with a single shape but with different sizes to cover spatially varying defocus blur. However, it is hard to directly simulate inverse kernels with all possible sizes in practice. Instead, we design our network to carry a few convolutional layers to cover inverse kernels with a discrete set of sizes, and aggregate the outputs of the layers for handling blur with arbitrary size. As a result, compared to the conventional two-step approach and the recent deep learning approach , our method can perform defocus deblurring more accurately by exploiting the properties of defocus blur in the form of inverse kernel.
To implement the network design, we propose a novel kernel-sharing parallel atrous convolutional (KPAC) block. The KPAC block consists of multiple atrous convolution layers with different dilation rates and additional layers for the scale and shape attentions. The multiple atrous convolution layers share the same convolution kernels representing the invariant shape of an inverse kernel, and different dilation rates of the layers correspond to inverse kernels with a discrete set of sizes. To simulate deconvolution using inverse kernels of other sizes, the KPAC block equips a spatial attention mechanism , which we call the scale attention, to aggregate the outputs of the atrous convolution layers. By combining the per-pixel scale attention with multiple atrous convolution layers, the KPAC block can handle the spatially varying size of defocus blur. In addition, the shape of defocus blur may slightly change in a defocused image due to the non-linearity of a camera image pipeline. To handle the variance, we include a channel attention mechanism , which we call the shape attention, in a KPAC block to support the slight shape change of the inverse kernel.
An important benefit of our KPAC block is its small number of parameters, enabled by kernel sharing of the multiple atrous convolution layers. As a result, our defocus deblurring network is lighter-weighted than the previous work , showing the better performance (Sec. 5.2).
Novel deep learning approach for single image defocus deblurring based on inverse kernels,
Novel Kernel-Sharing Parallel Atrous Convolutional (KPAC) block designed upon the properties of spatially varying inverse kernels for defocus deblurring,
Light-weight single image defocus deblurring network, which shows state-of-the-art performance.
Related Work
Conventional approaches perform defocus deblurring in two steps of defocus map estimation and non-blind deconvolution. As they use existing non-blind deconvolution methods for deblurring, they focus on improving the accuracy of defocus map estimation based on parametric blur models such as disc and Gaussian blur. Various methods were proposed for defocus map estimation using hand-crafted features such as edge gradients , sparse coded features , machine learning features , combination of hand-crafted and deep learning features , and an end-to-end deep learning model . These two-step approaches often fail to produce faithful deblurring results due to the restricted blur model as well as the errors of defocus map estimation.
Recently, Abuolaim and Brown proposed the first end-to-end model for deep learning-based defocus deblurring and a dataset for supervised training. The model outperforms conventional two-step approaches, and shows that a dual-pixel image input significantly improves defocus deblurring performance. However, their network structure does not explicitly consider the spatially varying nature of defocus blur with large variance in size but small variance in shape, and the performance has rooms for improvements.
2 Inverse kernel
Convolution of an inverse kernel can be used to perform deblurring on an image . However, a naïve inverse kernel that could be obtained from Wiener deconvolution usually introduces unwanted artifacts such as ringing and amplified noise. To suppress artifacts, previous works took a progressive approach using an image pyramid , regularization using sparse priors , a neural network for postprocessing , and feature space processing . Xu et al. and Ren et al. directly used a CNN to perform non-blind deconvoluton by adapting a separable property of a large inverse kernel, and showed that their approaches are effective for suppressing artifacts. Distinct from that simulate a single inverse kernel for uniform deconvolution, our approach simulates spatially varying inverse kernels needed for defocus deblurring.
Key Idea
In this section, we first present our observation on inverse kernels for deconvolution, and propose inverse kernel-based deconvolution to deal with spatially varying defocus blur (Sec. 3.1). We then present experiments to verify the observation and the proposition (Sec. 3.2).
The inverse kernel-based deconvolution approach is closely related to convolution neural networks (CNN) by nature, as they are based on convolutional operations in the spatial domain. We aim to design a network architecture that takes the benefits of both CNN and inverse kernels for defocus deblurring. To this end, we first introduce the concept and general derivation of pseudo inverse kernel, as in previous works .
We consider a simple blur model defined with convolution operation as
where is a blur kernel, and and are a blurry image and a latent sharp image, respectively. The spatial convolution can be transformed to an element-wise multiplication in the frequency domain as
where denotes the discrete Fourier transform. Then, the latent sharp image can be derived with convolution operation as
where denotes the inverse discrete Fourier transform, and is the spatial pseudo inverse kernel.
We observe that the shape of the corresponding inverse kernel remains the same when the spatial support size of a blur kernel changes, i.e.,
where denotes general upsampling operation with a scale factor in the spatial domain (refer to the supplementary material for more details on kernel upsampling). In Eq. (4), is the kernel with the upsampled resolution, where is used for normalizing the kernel weights. Then, the inverse kernel of is the upsampled version of with the same scale factor .
Based on the observation, to handle defocus blur with varying sizes but with the same shape, we may use inverse kernels sharing a single shape but with different sizes. However, as we aim to utilize a CNN for defocus deblurring, it is not practical to implement a network that carries the inverse kernels of all possible sizes. To reduce the complexity, we may approximate deconvolution of an image by combining the results of applying inverse kernels with a discrete set of sizes to the image. Similarly to a classical approach that handles non-uniform blur using a linear combination of differently deblurred images, spatially varying deconvolution for defocus blur can be approximated as
where is an upsampling factor, and is the per-pixel weight map for the result image obtained using the inverse kernel with an upsampling factor .
2 Experimental validation of the key idea
In this section, we experimentally validate the observed property of an inverse kernel for defocus deblurring (Eq. (4)) and the approach for spatially varying deconvolution based on the property (Eq. (5)). In the validation, we use Wiener deconvolution to compute a spatial inverse kernel and Lanczos upsampling to scale inverse kernels. We use Wiener inverse kernel as it has a finite spatial support analogous to the finite receptive field of a CNN, due to the involvement of signal-to-noise ratio (SNR) as regularization .
Regarding Eq. (4), Fig. 1 shows an example that two kernels on both sides of Eq. (4) produce equivalent deblurring results. We also include the proof of Eq. (4) in the supplementary material. Regarding Eq. (5), Fig. 2 shows that a linear combination of resulting images from the inverse kernels with a discrete set of sizes can approximate deconvolution with an inverse kernel of a different size. Figs. 2b and 2c are the deblurred results obtained by convolving inverse kernels of different sizes. While neither kernel fits the actual blur size, we can still obtain a visually decent result (Fig. 2d) with an approximation accuracyThe accuracy is computed by , where is the mean absolute error, is pixel-wise division, is the deconvolution result using an inverse kernel of a target scale (e.g., Fig. 2e), and is the approximated deconvolution result computed using Eq. (5) (e.g., Fig. 2d). of , by simply applying Eq. (5) with . When we optimize using the method of non-negative least squares, the accuracy further increases up to . These experiments confirm the validity of expanding Eqs. (4) and (5) to a CNN architecture, where the remaining errors could be compensated through deep learning process (Sec. 4). Refer to the supplementary material for more examples.
Network Design
Based on the key idea, we design a network reflecting the property of inverse kernels for defocus deblurring. Instead of directly adopting the linear combination in Eq. (5), we extend the concept of combination with a convolution layer that aggregates the results of multiple inverse kernels, to fully exploit the non-linear nature of deep learning.
Our KPAC block consists of multiple kernel-sharing atrous convolution layers, and modules for the scale and shape attentions (Fig. 3).
In Eq. (6), convolutional kernels {} should represent inverse kernels of the same shape but with different sizes as observed in Eq. (4). We may share the weights of {} to enforce the constraint of the same inverse kernel shape. However, in practice, the weight sharing is not straightforward because of the different sizes of {}.
We resolve this problem with a simple but effective solution based on another important observation; For a blurred region, which is spatially smooth, filtering operations on sparsely sampled pixels (with a dilated kernel) and on densely sampled pixels (with a rescaled kernel) produce similar results. Thus, the upsampling operation for in Eq. (4) can be replaced by a dilation operation, yielding an inverse kernel applied to sparsely sampled pixels. The dilation operation does not change the number of filter weights but only scales the spatial support of the kernel without resampling, enabling direct weight sharing for the convolutional kernels {}.
Fig. 4 shows an example that and produce almost equivalent deblurring results, where denotes the dilation operation. We also experimentally verified that a modified version of Eq. (5) using produces results almost equivalent to Fig. 2. Refer to the supplementary material for the experiment on dilated inverse kernels.
Based on the observation, our KPAC block includes multiple atrous convolution layers with different dilation rates, placed in parallel (Fig. 3). The atrous convolution kernels are composed of the same number of kernel weights regardless of scales and share the same kernel weights, satisfying the constraint of shared inverse kernel shape. That is, we substitute the standard convolution in Eq. (6) by an atrous convolution layer, obtaining
Scale attention
Shape attention
2 Defocus deblurring network
For effective multi-scale processing in the feature space, we adopt an encoder-decoder structure for our defocus deblurring network (Fig. 5). The network consists of three parts: encoder, KPAC blocks, and decoder. Except for the final convolution, all convolution layers include LeakyReLU for the non-linear activation layer.
As our KPAC block is designed upon inverse kernels defined in the linear space, one question naturally follows whether it is proper for the KPAC block to run in the non-linear feature space. We found that the KPAC block still works with non-linear features as CNNs are locally linear . It has also been shown in a recent work that Wiener deconvolution, which is a kind of inverse filters, can be successfully extended to the feature space due to the piecewise linearity of the feature space of CNN.
The proposed KPAC block may not operate exactly the same as conventional inverse kernel-based approaches for deblurring, as it does not explicitly employ inverse kernels. However, the architecture is still constrained with atrous convolutional layers with shared kernels and non-linear aggregation of resulting features, which are designed upon the property of inverse kernels. As a result, the KPAC block would learn restoration kernels that are more robust and effective for defocus deblurring. In addition, while a single KPAC block is designed to model the entire process of deblurring, we can stack multiple KPAC blocks to exploit the iterative nature for removing residual blurs.
For training the defocus deblurring network, we use the mean absolute error (MAE) between a network output and the corresponding ground-truth sharp image as a loss function. We also employ the perceptual loss for restoring more realistic textures. For the perceptual loss, we use the feature map extracted at the ‘conv4_4’ layer in the pre-trained VGG-19 network . When the perceptual loss is used, it is combined with the MAE loss, where the balancing factor is for the perceptual loss.
Experiments
We implemented and evaluated our models using Tensorflow 1.10.0 with NVIDIA Titan Xp GPU. Our final model has two KPAC blocks with a kernel size and the number of atrous convolution layers , as we empirically found it to work well in most cases. We use negative slope coefficient for LeakyReLU layers. We use the Adam optimizer with and to train our models. We train our models for 200k iterations with the fixed learning rate of . We tested a model trained for more iterations with learning rate decay, but its improvement was marginal in PSNR. For evaluation in Sec. 5.2, we train our models with the perceptual loss . For those models, we initialize them with pre-trained models trained with the MAE loss for 200k iterations. Then, we fine-tune the networks with both MAE and perceptual losses for additional 100k iterations with the fixed learning rate of . We use the batch size of 4. Each image in a batch is randomly cropped to .
We use the DPDD dataset for evaluation of our models. The dataset provides 500 image pairs of a real-world defocused image and the corresponding all-in-focus ground-truth image captured by a Canon EOS 5D Mark IV. The dataset consists of training, validation, and testing sets of 350, 74, and 76 pairs of images, respectively. In our experiments, we train and evaluate our models using the training and testing sets, respectively. While the dataset also provides dual-pixel data, we do not use them in our experiments. The dataset provides 16-bit images in the PNG format. We convert them to 8-bit images for our experiments.
1 Analysis
Our KPAC blocks learn spatially varying inverse kernels whose shapes remain the same, but their sizes vary. For effective learning of such inverse kernels, our network shares the convolution weights across multiple atrous convolution layers. In this experiment, we verify the effect of the weight sharing between atrous convolution layers by comparing the performance of models with and without the weight sharing. Both models have two KPAC blocks with kernels. Table 1 shows the deblurring quality and the number of parameters of each model. As shown in the table, our model with the weight sharing not only reduces the number of learning parameters, but also improves the deblurring quality, as its weight sharing structure properly constrains and guides the learning process.
Scale attention
The atrous convolution layers in our KPAC block simulate inverse kernels of different sizes to effectively handle the spatially varying nature of defocus blur. To analyze how they are activated for defocus blur with different sizes, we visualize the scale attention maps of different atrous convolution layers (Fig. 6). The roles of different attention maps may not be strictly distinguished because of the nature of the learning process that implicitly learns the use of different layers. Nevertheless, we can observe a clear tendency that the attention maps of different dilation rates are activated for different blur sizes. For example, the attention map of the dilation rate 1 is activated for pixels with blur of almost any size. On the other hand, the attention map of the dilation rate 5 is activated only for pixels with large blur. This shows that our scheme properly works for handling spatially varying size of defocus blur.
Ablation study
To quantitatively analyze the effect of each component in our KPAC block, we conduct an ablation study (Table 2). We first prepare a baseline model, which uses naïve convolution blocks instead of our KPAC blocks. For the baseline model, we use a conventional residual block that consists of two convolution layers with the filter size of . For a fair comparison, the baseline model includes multiple convolution blocks so that its model size is similar to our model without the weight sharing. We also prepare four variants of the baseline model using two KPAC blocks with the kernel size of , then measure the deblurring performances of the models. Table 2 summarizes the ablation study result. As shown in the table, every component of our proposed approach increases the deblurring quality significantly. Fig. 7 presents a qualitative comparison, which shows that both scale attention and shape attention help our network better handle spatially varying blur and restore fine structures while the models with no attentions suffer from spatially varying blurs.
Atrous convolutions with different dilation rates
We analyze the effect of dilation rates of atrous convolution layers in handling defocus blur with varying scales. In the test, we manually modulated the attention weight maps [] to make our pretrained network use only the feature maps produced by atrous convolution layers of specified scales. With the similar tendency to Fig. 6, we observed that atrous convolutions with dilation rates of 1, 3 and 5 help remove blur of any, medium, and large sizes, respectively (Fig. 8).
Number of KPAC blocks
Our KPAC block can be stacked together so that the network can iteratively remove defocus blur to achieve higher-quality deblurring results. We investigate the performance of different numbers of KPAC blocks. Table 3 shows that even a single KPAC block can effectively remove defocus blur and increase the PNSR by 0.90 dB. As we adopt more KPAC blocks, the PSNR increases although the improvement becomes smaller. After three KPAC blocks, the PSNR starts to decrease possibly due to the increased training complexity. Based on this experiment, we design our final model to have two KPAC blocks as two blocks provide relatively high deblurring quality with small model size.
2 Evaluation
We compare our method with state-of-the-art defocus deblurring methods, including both conventional two-step approaches and the recent end-to-end deep learning-based approach . For all the methods, we produced result images using the source code provided by the authors. For JNB , EBDB and DMENet , we used the non-blind deconvolution method for generating deblurred images using estimated defocus maps. For DPDNet , we used the source code and pre-trained models provided by the author. DPDNet provides two versions of models, each of which takes a single input image and dual-pixel data, which is a pair of sub-aperture images, respectively. We include both of them in our comparison. For evaluation, we measure PSNR and SSIM . We also measure LPIPS for evaluating the perceptual quality as done in .
We include two variants of our model, each of which has a different number of encoding levels, or a different number of downsampling layers in the encoder. By increasing the encoding levels, we can more easily handle large blur with small filters and with a small amount of computations. On the other hand, with fewer encoding levels, it is easier to restore fine-scale details. To inspect the difference between models with different numbers of encoding levels, we include two variants of our model, which have two and three levels, respectively. Both models are trained with both MAE and perceptual loss functions.
Table 4 reports the quantitative comparison. As shown in the table, the classical two-step approaches perform worse than the recent deep-learning based approach . While the DPDNet model with a single input image performs better than the classical approaches, our models outperform both the classical approaches and the DPDNet model with a single input image by a large margin. Moreover, our models outperform the dual-pixel-based DPDNet model even without the strong cue to defocus blur provided by dual-pixel data and with a much small number of parameters. This result clearly proves the effectiveness of our approach. In the supplementary material, we report the performance of the dual-pixel-based variant of our model, which outperforms the dual-pixel-based DPDNet model. Table 4 also shows that our 2- and 3-level models perform similarly. However, we found that the 3-level model tends to better handle extremely large blur as shown in Fig. 9 due to its larger receptive fields.
Fig. 10 shows a qualitative comparison. Our result is produced by the 3-level model. As the figure shows, our method produces sharper results with more details. Even compared to the result of the dual-pixel-based DPDNet , our result has comparably clear details.
Computational cost
We compare the computational cost of our models and DPDNet . Classical two-step approaches rely on computationally heavy non-blind deconvolution algorithms, so we do not include them in this comparison. For the comparison, we measure FLOPs and the average running time per image of size . Table 5 shows that our 3-level model requires small computational cost in FLOPs, which is 10 times smaller than . The table also shows that our 3-level model is slightly faster than the 2-level model as features are more downsampled, even though it has more parameters.
Generalization to other images
Our models are trained using the DPDD dataset , which was generated using one camera. Thus, lastly, we inspect how well our model generalizes to images from other cameras. To this end, we use the CUHK blur detection dataset , which provides 704 defocused images without ground-truth all-in-focus images. The defocused images in the dataset are collected from various sources on the internet. As there are no ground-truth images, we qualitatively inspect the generalization ability. Fig. 11 shows that our method can successfully restore fine details with less visual artifacts compared to the single image-based model of DPDNet . Additional results can be found in the supplementary material.
Conclusion
This paper proposed a single image defocus deblurring framework based on inverse kernels. To effectively simulate spatially varying inverse kernels, we proposed the Kernel-Sharing Parallel Atrous Convolutions (KPAC) block. KPAC provides an effective way to handle spatially varying defocus blur with a small number of atrous convolution layers that share the same convolutional kernel weights but with different dilation rates. KPAC is also equipped with the per-pixel scale attention to further facilitate the handling of spatially varying blur. Thanks to the effective and light-weight structure of KPAC, we can simply stack multiple blocks of KPAC and achieve state-of-the-art deblurring performance. We experimentally validated the effectiveness of KPAC and showed that our method clearly outperforms previous methods with much fewer parameters.
While our method outperforms previous state-of-the-art methods, it may still fail in challenging cases such as large-scale blur, blur with irregular shapes, and bokeh with sharp boundaries (refer to the supplementary material for our deblurring results on images containing such cases). Handling these challenging cases would be an interesting future direction.
Acknowledgements
This work was supported by the Ministry of Science and ICT, Korea, through IITP grants (SW Star Lab, 2015-0-00174; Artificial Intelligence Graduate School Program (POSTECH), 2019-0-01906) and NRF grants (2018R1A5A1060031; 2020R1C1C1014863).
References
Appendix A Detailed Network Architecture
Detailed architectures of overall deblurring network and KPAC block can be found in Tables 6 and 7, respectively.
In this section, we present a formal discussion on Eq. (4) in the main paper. For a 2D image of size , the spatial upsampling can be performed by zero padding to the discrete Fourier transform of \citeSMSmith:2007:FFT\citeSMAshikaga:2014:FFT. Let denote the upsampling operation by a scaling factor , and let be the upsampled result of image by the scaling factor . Then, the discrete Fourier transform of is defined in the range , and can be obtained as:
where is the discrete Fourier transform of , and and are pixel indices in the frequency domain. corresponds to the DC component of . This zero padding-based upsampling is mathematically equivalent to convolution with a sinc kernel \citeSMAshikaga:2014:FFT.
In the remaining of this section, for notational simplicity, we use to indicate the zero-padding operation for upsampling in the frequency domain as well as the upsampling operation in the spatial domain. We also omit the pixel coordinates , e.g., representing as .
We use the Wiener deconvolution to compute the inverse kernel of a blur kernel , i.e.,
where is the discrete Fourier transform of , and is the complex conjugate of . The division operation is done in an element-wise manner. is a noise parameter. Note that when , this inverse kernel is equivalent to the direct inverse kernel .
Regarding the left side of Eq. (10), the upsampling of a blur kernel can be obtained using the zero padding-based upsampling as:
Then, in the frequency domain, from Eq. (8), we have
From Eq. (9), the inverse kernel for the upsampled kernel can be derived by
In the frequency domain, from Eq. (12), we have
Note that the zero padded area in remains zero in .
Regarding the right side of Eq. (10), from Eqs. (11) and (9), the upsampled inverse kernel for kernel can be derived as
Then, in the frequency domain, from Eq. (8), we have
As Eqs. (14) and (16) are equivalent to each other, their spatial domain counterparts and are equivalent too. This proves Eq. (10).
Eq. (4) in the main paper has a scaling factor in both left and right sides. The upsampling operation in the left side of Eq. (10) scales up the total intensity of kernel by times. Similarly, the upsampling operation in the right side of Eq. (10) also scales up the total intensity of inverse kernel by times. Thus, to obtain a properly normalized inverse kernel in the left and right sides, we apply a scaling factor to in the left side, and to in the right side. Then, we obtain Eq. (4) in the main paper.
Appendix C Discussion on Kernel Upsampling
We consider general upsampling operation in Eq. 4 in the main paper for the commutative property between upsampling and inversion of a kernel. For validating the property, in Sec. B of this supplementary material, we used a specific upsampling method using the sinc filter to change the spatial scale of a kernel. Then, there could be a concern whether upsampling of a blur kernel can model actual scale changes of the blur kernel. In our observation, for Gaussian blur kernels with different standard deviations, which is an often-used assumption in existing defocus deblurring approaches, upsampling a kernel is the same as changing the standard deviation of the kernel (Fig. 12).
Nonetheless, the upsampling method may cause a gap in accurately modeling the scale changes of real-world blur when the blur kernel is arbitrary, other than Gaussian. This potential modeling gap in upsampling would be handled by shape and scale attentions in our KPAC blocks together with other convolution layers in our network.
Appendix D Inverse Kernel-based Deconvolution for Spatially Varying Defocus Blur
In Eq. (5) in the main paper, we approximate spatially varying image deconvolution by combining the results obtained from inverse kernels with a discrete set of sizes. In this section, we present additional visual examples to show the validity of our approximated deconvolution. Figs. 13b and 13c are the deblurred results obtained by convolving inverse kernels of different sizes. While neither kernel fits the actual blur size, a linear combination of the deconvolution results still produces a visually pleasing result (Fig. 13d), almost equivalent to the deconvolution result (Fig. 13f) using the inverse kernel of the target scale.
Appendix E Inverse Kernel Sampling for Atrous Convolution
In Sec. 4.1 of the main paper, we claimed that for a blurred region, which is spatially smooth, filtering operations using sparsely sampled pixels (with a dilated kernel) and using densely sampled pixels (with a rescaled kernel) produce similar results. Fig. 14 presents additional visual examples to show that a dense sampling of an inverse kernel can be replaced by its sparse sampling in terms of deblurring performance.
Moreover, we also claimed that a modified version of Eq. (5) in the main paper using the dilated inverse kernels produces results almost equivalent to those from the original Eq. (5). In Fig. 13, we present additional examples to show the validity of the claim, where the modified and original versions of Eq. (5) produce almost same results in Figs. 13e and 13d, respectively.
Appendix F Blur Detection Using a Scale Attention Map
In Sec. 5.1 of the main paper, we showed that the scale attention map for the atrous convolution layer with a dilation rate of 1 in the first KPAC block of our deblurring network captures blur of almost any size. In this section, we evaluate the blur detection performance of the scale attention map using the CUHK dataset . We use 200 test images in the CUHK dataset and measure F-measure and accuracy. Since our attention is computed in the low resolution feature space, small blurs that can be removed in the encoder network would be ignored. Therefore, we upsample input images four times for obtaining attention maps in this test. The result shows that our attention map (F-measure: 0.832 and accuracy: 78.4%) can detect blur comparably to the recent defocus map estimation method (F-measure: 0.839 and accuracy: 76.5%). Fig. 15 shows qualitative examples.
Appendix G Sensitivity to Noise
We investigate the sensitivity of our model to different noise levels, as the shape of a desirable inverse kernel can be affected by noise. Table 8 quantitatively shows the effect of the noise level on our model. For the experiment, we prepare two models. One model is trained with the original dataset of (top row of the table). The other model is trained with defocused images augmented with additive Gaussian noise controlled by within a range (bottom row of the table). Compared to the model trained without the noise augmentation, the model trained with the noise augmentation is more robust to noise, and shows more consistent PSNRs around 25 dB.
Appendix H Handling Irregular Blur
Due to the network design based on kernel weight sharing, our method would be more effective for the case where the majority of blur variation happens in the size. In practice, blur shape can spatially vary as well, e.g., due to lens distortion in a smartphone camera. Still, we observed that our network moderately works on defocused images captured by smartphones in an unseen dataset \citeSMGarg2019ICCV, which usually contain small-sized blurs (Fig. 16). However, as we discussed as limitations in Sec. 6 of the main paper, our network may not properly handle blur with severely irregular shapes or strong highlights (Fig. 17), which are rarely included in the training set .
Appendix I Additional Results
We present additional qualitative results on the DPDD dataset (Figs. 18 and 19) and the CUHK blur detection dataset (Figs. 20 and 21).
Appendix J Our Model Using Dual-Pixel Images
While our model with a single image input shows state-of-the-art performance, the deblurring performance can further boosted by using dual-pixel images as the input. For the experiment, we retrained our model by replacing a single image input as dual-pixel image input with the same training strategy in Sec. 5 of the main paper. Specifically, we concatenate the two images of a dual-pixel image in the channel dimension and use it as the input of the network.
Table 9 shows that dual-pixel input further improves the performance of our model, and our model with dual-pixel input outperforms DPDNet with dual-pixel input by a large margin. Fig. 22 shows that dual-pixel input enables our model to handle fine details better.