Defocus Deblurring Using Dual-Pixel Data

Abdullah Abuolaim, Michael S. Brown

Introduction

This paper addresses the problem of defocus blur. To understand why defocus blur is difficult to avoid, it is important to understand the mechanism governing image exposure. An image’s exposure to light is controlled by adjusting two parameters: shutter speed and aperture size. The shutter speed controls the duration of light falling on the sensor, while the aperture controls the amount of light passing through the lens. The reciprocity between these two parameters allows the same exposure to occur by fixing one parameter and adjusting the other. For example, when a camera is placed in aperture-priority mode, the aperture remains fixed while the shutter speed is adjusted to control how long light is allowed to pass through the lens. The drawback is that a slow shutter speed can result in motion blur if the camera and/or an object in the scene moves while the shutter is open, as shown in Fig. 1. Conversely, in shutter-priority mode, the shutter speed remains fixed while the aperture adjusts its size. The drawback of a variable aperture is that a wide aperture results in a shallow depth of field (DoF), causing defocus blur to occur in scene regions outside the DoF, as shown in Fig. 1. There are many computer vision applications that require a wide aperture but still want an all-in-focus image. An excellent example is cameras on self-driving cars, or cameras on cars that map environments, where the camera must use a fixed shutter speed and the only way to get sufficient light is a wide aperture at the cost of defocus blur.

Our aim is to reduce the unwanted defocus blur. The novelty of our approach lies in the use of data available from dual-pixel (DP) sensors used by modern cameras. DP sensors are designed with two photodiodes at each pixel location on the sensor. The DP design provides the functionality of a simple two-sample light-field camera and was developed to improve how cameras perform autofocus. Specifically, the two-sample light-field provides two sub-aperture views of the scene, denoted in this paper as left and right views. The light rays coming from scene points that are within the camera’s DoF (i.e., points that are in focus) will have no difference in phase between the left and right views. However, light rays coming from scene points outside the camera’s DoF (i.e., points that are out of focus) will exhibit a detectable disparity in the left/right views that is directly correlated to the amount of defocus blur. We refere to it as defocus disparity. Cameras use this phase shift information to determine how to move the lens to focus on a particular location in the scene. After autofocus calculations are performed, the DP information is discarded by the camera’s hardware.

Contribution. We propose a deep neural network (DNN) to perform defocus deblurring that uses the DP images from the sensor available at capture time. In order to train the proposed DNN, a new dataset of 500 carefully captured images exhibiting defocus blur and their corresponding all-in-focus image is collected. This dataset consists of 2000 images – 500 DoF blurred images with their 1000 DP sub-aperture views and 500 corresponding all-in-focus images – all at full-frame resolution (i.e., 6720×44806720\times 4480 pixels). Using this training data, we propose a DNN architecture that is trained in an end-to-end manner to directly estimate a sharp image from the left/right DP views of the defocused input image. Our approach is evaluated against conventional methods that use only a single input image and show that our approach outperforms the existing state-of-the-art approaches in both signal processing and perceptual metrics. Most importantly, the proposed method works by using the DP sensor images that are a free by-product of modern image capture.

Related work

Related work is discussed regarding (1) defocus blur, (2) datasets, and (3) applications exploiting DP sensors.

Defocus deblurring. Related methods in the literature can be categorized into: (1) defocus detection methods or (2) defocus map estimation and deblurring methods . While defocus detection is relevant to our problem, we focus on the latter category as these methods share the goal of ultimately producing a sharp deblurred result.

A common strategy for defocus deblurring is to first compute a defocus map and use that information to guide the deblurring. Defocus map estimation methods estimate the amount of defocus blur per pixel for an image with defocus blur. Representative works include Karaali et al. , which uses image gradients to calculate the blur amount difference between the original image edges and their re-blurred ones. Park et al. introduced a method based on hand-crafted and deep features that were extracted from a pre-trained blur classification network. The combined feature vector was fed to a regression network to estimate the blur amount on edges and then later deblur the image. Shi et al. proposed an effective blur feature using a sparse representation and image decomposition to detect just noticeable blur. Methods that directly deblur the image include Andrès et al.’s approach, which uses regression trees to deblur the image. Recent work by Lee et al. introduced a DNN architecture to estimate an image defocus map using a domain adaptation approach. This approach also introduced the first large-scale dataset for DNN-based training. Our work is inspired by Lee et al.’s success in applying DNNs for the DoF deblurring task. Our distinction from the prior work is the use of the DP sensor information available at capture time.

Defocus blur datasets. There are several datasets available for defocus deblurring. The CUHK and DUT datasets have been used for blur detection and provide real images with their corresponding binary masks of blur/sharp regions. The SYNDOF dataset provided data for defocus map estimation, in which their defocus blur is synthesized based on a given depth map of pinhole image datasets. The datasets of do not provide the corresponding ground truth all-in-focus image. The RTF dataset provided light-field images captured by a Lytro camera for the task of defocus deblurring. In their data, each blurred image has a corresponding all-in-focus image. However, the RTF dataset is small, with only 22 image pairs. While there are other similar and much larger light-field datasets , these datasets were introduced for different tasks (i.e., depth from focus and synthesizing a 4D RGBD light field), which are different from the task of this paper. In general, the images captured by Lytro cameras are not representative of DSLR and smartphone cameras, because they apply synthetic defocus blur, and have a relatively small spatial resolution .

As our approach is to utilize the DP data for defocus deblurring, we found it necessary to capture a new dataset. Our DP defocus blur dataset provides 500 pairs of images of unrepeated scenes; each pair has a defocus blurred image with its corresponding sharp image. The two DP views of the blurred image are also provided, resulting in a total of 2000 images. Details of our dataset capture are provided in Sec. 4. Similar to the patch-wise training approach followed in , we extract a large number of image patches from our dataset to train our DNN.

DP sensor applications. The DP sensor design was developed by Canon for the purpose of optimizing camera autofocus. DP sensors perform what is termed phase difference autofocus (PDAF) , in which the phase difference between the left and right sub-aperture views of the primary lens is calculated to measure the blur amount. Using this phase information, the camera’s lens is adjusted such that the blur is minimized. While intended for autofocus, the DP images have been found useful for other tasks, such as depth map estimation , reflection removal , and synthetic DoF . Our work is inspired by these prior methods and examines the use of DP data for the task of defocus blur removal.

DP image formation

We begin with a brief overview of the DP image formation. As previously mentioned, the DP sensor was designed to improve camera auto-focus technology. Fig. 2 shows an illustrative example of how DP imaging works and how the left/right images are formed. A DP sensor provides a pair of photodiodes for each pixel with a microlens placed at the pixel site, as shown in Fig. 2-A. This DP unit arrangement allows each pair of photodiodes (i.e., dual-pixel) to record the light rays independently. Depending on the sensor’s orientation, this arrangement can be shown as left/right or top/down pair; in this paper, we refer to them as the left/right pair – or L and R. The difference between the two views is related to the defocus amount at that scene point, where out-of-focus scene points will have a difference in phase and be blurred in opposite directions using a point spread function (PSF) and its flipped one . This difference yields noticeable defocus disparity that is correlated to the amount of defocus blur.

The phase-shift process is illustrated in Fig. 2. The person shown in Fig. 2-A is within the camera’s DoF, as highlighted in gray, whereas the textured pyramid is outside the DoF. The light rays from the in-focus object converge at a single DP unit on the imaging sensor, resulting in an in-focus pixel and no disparity between their DP L/R views (Fig. 2-B). The light rays coming from the out-of-focus regions spread across multiple DP units and therefore produce a difference between their DP L/R views, as shown in Fig. 2-C. Intuitively, this information can be exploited by a DNN to learn where regions of the image exhibit blur and the extent of this blur. The final output image is a combination of the L/R views, as shown in Fig. 2-G.

By examining real examples shown in Fig. 3 it becomes apparent how a DNN can leverage these two sub-aperture views as input to deblur the image. In particular, patches containing regions that are out-of-focus will exhibit a notable defocus disparity in the two views that is directly correlated to the amount of defocus blur. By training a DNN with sufficient examples of the L/R views and the corresponding all-in-focus image, the DNN can learn how to detect and correct blurred regions. Animated examples of the difference between the DP views are provided in the supplemental materials.

Dataset collection

Our first task is to collect a dataset with the necessary DP information for training our DNN. While most consumer cameras employ PDAF sensors, we are aware of only two camera manufacturers that provide DP data – Google and Canon. Specifically, Google’s research team has released an application to read DP data from the Google Pixel 3 and 4 smartphones. However, smartphone cameras are currently not suitable for our problem for two reasons. First, smartphone cameras use fixed apertures that cannot be adjusted for data collection. Second, smartphone cameras have narrow aperture and exhibit large DoF; in fact, most cameras go to great lengths to simulate shallow DoF by purposely introducing defocus blur . As a result, our dataset is captured using a Canon EOS 5D Mark IV DSLR camera, which provides the ability to save and extract full-frame DP images.

Dual-pixel defocus deblurring DNN (DPDNet)

Using our captured dataset, we trained a symmetric encoder-decoder CNN architecture with skip connections between the corresponding feature maps . Skip connections are widely used in encoder-decoder CNNs to combine various levels of feature maps. These have been found useful for gradient propagation and convergence acceleration and to allow training of deeper networks as stated in .

We adapt a U-Net-like architecture with the following modifications: an input layer to take a 6-channel input cube (two DP views; each is a 3-channel sRGB image) and an output layer to generate a 3-channel output sRGB image; skip connections of the convolutional feature maps are passed to their mirrored convolutional layers without cropping in order to pass on more feature map detail; and the loss function is changed to be mean squared error (MSE).

where DPDNet is our proposed architecture, and θDPDNet\theta_{\textrm{DPDNet}} is the set of weights and parameters.

Training procedure. The size of input and output layers is set to 512×512×6512\times 512\times 6 and 512×512×3512\times 512\times 3, respectively. This is because we train not on the full-size images but on the extracted image patches. We adopt the weight initialization strategy proposed by He and use the Adam optimizer to train the model. The initial learning rate is set to 2×10−52\times 10^{-5}, which is decreased by half every 60 epochs. We train our model with mini-batches of size 5 using MSE loss between the output and the ground truth as follows:

where nn is the size of the image patch in pixels. During the training phase, we set the dropout rate to 0.40.4. All the models described in the subsequent sections are implemented using Python with the Keras framework on top of TensorFlow and trained with a NVIDIA TITAN X GPU. We set the maximum number of training epochs to 200.

Experimental results

We first describe our data preparation procedure and then evaluation metrics used. This is followed by quantitative and qualitative results to evaluate our proposed method with existing deblurring methods. We also discuss the time analysis and test the robustness of our DP method against different aperture settings.

Data preparation. Our dataset has an equal number of indoor and outdoor scenes. We divide the data into 70%70\% training, 15%15\% validation, and 15%15\% testing sets. Each set has a balanced number of indoor/outdoor scenes. To prepare the data for training, we first downscale our images to be 1680×11201680\times 1120 in size. Next, image patches are extracted by sliding a window of size 512×512512\times 512 with 60%60\% overlap. We empirically found this image size and patch size to work well. An ablation study of different architecture settings is provided in the supplemental materials. We compute the sharpness energy (i.e., by applying Sobel filter) of the in-focus image patches and sort them. We discard 30% of the patches that have the lowest sharpness energy. Such patches represent homogeneous regions, cause an ambiguity associated to the amount of blur, and adversely affect the DNNs training, as found in .

Evaluation metrics. Results are reported on traditional signal processing metrics – namely, PSNR, SSIM , and MAE. We also incorporate the recent learned perceptual image patch similarity (LPIPS) proposed by . The LPIPS metric is correlated with human perceptual similarity judgments as a perceptual metric for low-level vision tasks, such as enhancement and image deblurring.

Qualitative results. In Fig. 6, we present the qualitative results of different defocus deblurring methods. The first row shows the input image with a spatially varying defocus blur; the last row shows the corresponding ground truth sharp image. The rows in between present different methods, including ours. This figure also shows two zoomed-in cropped patches in green and red to further illustrate the difference visually. From the visual comparison with other methods, our DPDNet has the best deblurring ability and is quite similar to the ground truth. EBDB , DMENet , and JNB are not able to handle spatially varying blur with almost unnoticeable difference with the input image. EBDB tends to introduce some artifacts in some cases. Our single image method (i.e., DPDNet-Single) has better deblurring ability compared to other traditional deblurring methods, but it is not at the level of our method that utilizes DP views for deblurring. Our DPDNet method, as shown visually, is effective in handling spatially varying blur. For example, in the second row, the image has a part that is in focus and another is not; our DPDNet method is able to determine the deblurring amount required for each pixel, in which the in-focus part is left untouched. Further qualitative results are provided in our supplemental materials, including results on DP data obtained from a smartphone camera.

Time analysis. We examine evaluating different defocus deblurring methods based on the time required to process a testing image of size 1680×11201680\times 1120 pixels. Our DPDNet directly computes the sharp image in a single pass, whereas other methods use two passes: (1) defocus map estimation and (2) non-blind deblurring based on the estimated defocus map.

Non-learning-based methods (i.e., EBDB and JNB ) do not utilize the GPU and use only the CPU. For the deep-learning method (i.e., DMENet ), it utilizes the GPU for the first pass; however, the deblurring routine is applied on a CPU. This time evaluation is performed using Intel Core i7-6700 CPU and NVIDIA TITAN X GPU. Our DPDNet operates in a single pass and can process the testing image of size 1680×11201680\times 1120 pixels about 1.2×1031.2\times 10^{3} times faster compared to the second-best method (i.e., DMENet), as shown in Table 2.

Robustness to different aperture settings. In our dataset, the image pairs are captured using aperture settings corresponding to f-stops f/22f/22 and f/4f/4. Recall that f/4f/4 results in the greatest DoF and thus most defocus blur. Our DPDNet is trained on diverse images with many different depth values; thus, our training data spans the worst-case blur that would be observed with any aperture settings. To test the ability of our DPDNet in generalizing for scenes with different aperture settings, we capture image pairs with aperture settings f/10f/10 and f/16f/16 for the blurred image and again f/22f/22 for the corresponding ground truth image. Our DPDNet is applied to these less blurred images. Fig. 7 shows the results for four scenes, where each scene’s image has its LPIPS measure compared with the ground truth. For better visual comparison, Fig. 7 provides zoomed-in patches that are cropped from the blurred input (red box) and the deblurred one (green box). These results show that our DPDNet is able to deblur scenes with different aperture settings that have not been used during training.

Applications

Image blur can have a negative impact on some computer vision tasks, as found in . Here we investigate defocus blur effect on two common computer vision tasks – namely, image segmentation and monocular depth estimation.

Conclusion

We have presented a novel approach to reduce the effect of defocus blur present in images captured with a shallow DoF. Our approach leverages the DP data that is available in most modern camera sensors but currently being ignored for other uses. We show that the DP images are highly effective in reducing DoF blur when used in a DNN framework. As part of this effort, we have captured a new image dataset consisting of blurred and sharp image pairs along with their DP images. Experimental results show that leveraging the DP data provides state-of-the-art quantitative results on both signal processing and perceptual metrics. We also demonstrate that our deblurring method can be beneficial for other computer vision tasks. We believe our captured dataset and DP-based method are useful for the research community and will help spur additional ideas about both defocus deblurring and applications that can leverage data from DP sensors. Acknowledgments. This study was funded in part by the Canada First Research Excellence Fund for the Vision: Science to Applications (VISTA) programme and an NSERC Discovery Grant. Dr. Brown contributed to this article in his personal capacity as a professor at York University. The views expressed are his own and do not necessarily represent the views of Samsung Research.

References

S1 Ablation study

In this section, we provide an ablation study of different variations in training our DPDNet with: (1) an extra input image (Sec. S1.1), (2) less E-Blocks and D-Blocks (Sec. S1.2), (3) different input sizes (Sec. S1.3), (4) different ratios of homogeneous region filtering (Sec. S1.4), and (5) different data types (Sec. S1.5). This is related to Sec. 5 and Sec. 6 of the main paper.

S1.2 DPDNet with less blocks

In this section, we train a “lighter” version of our DPDNet with less E-Blocks and D-Blocks. This is done by reducing E-Block 1 and D-Block 4. We refer to this light version as DPDNet-Light. In Table 4, we provide a comparison of DPDNet-Light and our full DPDNet that is proposed in the main paper.

Table 4 shows that our full DPDNet has a better performance compared to the lighter one. Nevertheless, the sacrifice in performance is not too significant, which implies that the DPDNet-Light could be an option for environments with limited computational resources.

S1.3 DPDNet with different input sizes

Our DPDNet is a fully convolutional network. This facilitates training with different input patch sizes with no change required in the network architecture. As such, we consider training with two different patch sizes, namely 256×256256\times 256 pixels and 512×512512\times 512 pixels referred to as DPDNet256 and DPDNet512, respectively.

Table 5 shows that the two different input sizes perform similarly. Particularly, input patch size does not change the performance drastically as long as it is larger than the blur size.

S1.4 DPDNet with different filtering ratios

Homogeneous patches are inherently ambiguous in terms of incurred blur size, and do not provide useful information for network training . As a result, filtering homogeneous patches can be beneficial to the trained network. In this section, different filtering ratios are examined including: 0%0\%, 15%15\%, 30%30\%, and 45%45\%; we refer to them as DPDNet0%, DPDNet15%, DPDNet30%, DPDNet45%, respectively.

In Table 6, we present the results of different filtering ratios. The 30%30\% filtering is a reasonable ratio that has the best quantitative results. Therefore, we filter 30%30\% of the extracted image patches based on the sharpness energy to train our proposed DPDNet as described in Sec. 6 of the main paper.

S1.5 DPDNet with different data types

Our dataset provides high-quality images that are processed to an sRGB encoding with a lossless 16-bit depth per RGB channel. Since we are targeting dual-pixel information which would be obtained directly in the camera’s hardware, in a real hardware implementation we would expect to have such high bit-depth images. However, since most standard encodings still rely on 8-bit image, we provide a comparison of training our DPDNet with 8-bit (DPDNet8-bit) and 16-bit (DPDNet16-bit) input data type.

Based on the numbers in Table 7, DPDNet16-bit has a slightly better performance. In particular, it has a lower LPIPS distance for all categories. As a result, training with 16-bit images is helpful due to the extra information embedded in, and is more representative of the hardware’s data.

S2 Defocus and motion blur discussion

One may be curious if motion blur methods can be used to address the defocus blur problem. While defocus and motion blur both produce a blurring of the underlying latent image, the physical image formation process of these two types of blur are different. Therefore, comparing with methods that solve for motion blur is not expected to give good results. However, for a validity check, we tested the scale recurrent motion deblurring method (SRNet) in using our testing set. This method achieved an average LPIPS of 0.452 and PSNR of 20.12, which is lower than all other existing methods that solve for defocus deblurring. Fig. 9 shows results of applying motion deblurring network SRNet to input image from our dataset.

S3 Use cases

As discussed in Sec. 1 of the main paper, we described how defocus blur is related to the size of the aperture used at capture time. The size of the aperture is often dictated by the desired exposure which is a factor of aperture, shutter speed, and ISO setting. As a result, there is a trade-off between image noise (from ISO gain), motion blur (shutter speed), and defocus blur (aperture). This trade off is referred to as the exposure triangle. In this section, we show some common cases, where defocus deblurring is required.

S4 DPDNet performance for a smartphone DP sensor

In this section, we test our DPDNet on images captured with a smartphone. As we mentioned in Sec. 4 of the main paper, there are two camera manufacturers that provide DP data, namely, Google Pixel 3 and 4 smartphones and Canon EOS 5D Mark IV DSLR. The smartphone camera currently has limitations that make it challenging to train the DPDNet with. First, the Google Pixel smartphone cameras do not have adjustable apertures, so we are unable to capture corresponding “sharp” images using a small aperture as we did with the Canon camera. Second, the data currently available from the Pixel smartphones are not full-frame, but are limited to only one of the Green channels in the raw-Bayer frame. Finally, the smartphone has a very small aperture so most images do not exhibit defocus blur. In fact, many smartphone cameras synthetically apply defocus blur to produce the shallow DoF effect.

As a result, the experiments here are provided to serve as a proof of concept that our method should generalize to other DP sensors. To this end, we examined DP images available in the dataset from to find images exhibiting defocus blur. The L/R views of these images are available in the “animated_dp_examples” directory—located at the same directory as this pdf file.

To use our DPDNet, we replicate the single green channel to be 3-channel image to match our DPDNet input. Fig. 12 shows the deblurring results on images captured by Pixel camera. The image on the left is the input combined image and the image on the right is the deblurred one using our DPDNet. Note that the Pixel android application, used to extract DP data, does not provide the combined image . To obtain it, we average the two views. Fig. 12 visually demonstrates that our DPDNet is able to generalize and deblur for images that are captured by the smartphone camera. Because it is not possible to adjust aperture on the smartphone camera to capture a ground truth image, we cannot report quantitative numbers. The results of two more full images are shown in Fig. 13.

S5 More results

Quantitative results. In Table 8, we provide evaluation of other methods on a single DP view separately using the average LPIPS. Note that a single DP L or R view is formed with a half-disc point spread function in the ideal case. When the two views are combined to form the final output image; the blur kernel would look like a full-disc kernel . Non-blind defocus deblurring methods assume full-disc kernel and the blur kernel of the combined image aligns more with their assumption. More details about DP view formation and modeling DP blur kernels can be found in .

In addition to above, we report in Table 9 the average LPIPS numbers for other methods on the images used to test DPDNet robustness to different aperture settings. Note that the LPIPS numbers here are lower than numbers in Table 1 of the main paper. The reason is that for the robustness test we used f/10 and f/16, which results in less defocus blur compared to the images captured at f/4 (a much wider aperture than f/10 and f/16).