BANet: Blur-aware Attention Networks for Dynamic Scene Deblurring

Fu-Jen Tsai, Yan-Tsung Peng, Yen-Yu Lin, Chung-Chi Tsai, Chia-Wen Lin

I Introduction

Dynamic scene deblurring or blind motion deblurring aims to restore a blurred image with little knowledge about the blur kernel. Scene blur caused by camera shakes, object motions, low shutter speeds, or low frame rates not only degrades the quality of taken images/videos but also results in information loss. Therefore, removing such blurring artifacts to recover image details becomes essential to many downstream vision applications, such as facial detection , text recognition , moving object segmentation , etc., where clean and sharp images are appreciated. Although significant progress has been made in conventional and deep-learning-based approaches , we observe a compromise between accuracy and speed. Owing to this observation, we target to develop an efficient and effective algorithm in this paper for blurred image restoration with its current performance in accuracy and speed shown in Fig. 1.

Deep-learning-based approaches usually reach superior results, given their better feature representation capability toward dynamic scenes. Among the state-of-the-art architectures for deblurring, self-recurrent models have been widely adopted to leverage blurred image repeatability in either multiple scales (MS) , multiple patch levels (MP) , or multiple temporal behaviors (MT) , as shown in Fig. 2(a)–(c). Specifically, the MS models distill multi-scale blur information in a self-recurrent manner and restore blurred images based on the extracted coarse-to-fine features . However, scaling a blurred image to a lower resolution often results in losing edge information . In contrast, the MP models split a blurred input image into multiple patches to estimate and then remove motion blurs of different scales . However, splitting the blurred input and features into equal-sized non-overlapping patches may cause cause contextual information discontinuity, sub-optimal for handling non-uniform blur in dynamic scenes. In , a self-recurrent MT structure was proposed to progressively eliminate non-uniform blurs over multiple iterations. Each iteration would gradually deblur the image until it becomes sharp. However, its inflexible progressive training and inference process may not generalize well for images of varying region-wise blurring degrees. Besides, these existing self-recurrent models, including MS, MP, and MT, cannot achieve high-quality deblurring in real-time (say, 30 HD frames per second).

In addition to model architectures, recent research studies further exploit self-attention to address blur non-uniformity. Suin et al. utilize MP-based processing with self-attention to extract features for areas with global and local motions. However, using a self-recurrent mechanism to generate multi-scale features often leads to a significantly longer inference time. To shorten the latency, Purohit and Rajagopalan selectively aggregate features through learnable pixel-wise attention enabled by deformable convolutions for modeling local blurs in a single forward pass. Despite its effectiveness, self-attention exploring pixel-wise or channel-wise correlations via trainable filters often causes high memory usage, thus only applicable to small-scale features . Furthermore, motion blurs coming from object motions manifest smeared effects and produce directional and local averaging artifacts, which cannot be handled well by inter-pixel/channel correlations.

This paper proposes a Blur-aware Attention Network (BANet) to overcome the above-mentioned issues. BANet is an efficient yet effective single-forward-pass model, as illustrated in Fig. 2(d), which achieves state-of-the-art deblurring performance while working in real-time, as shown in Fig. 1. Specifically, our model stacks multiple layers of the Blur-Aware Module (BAM) for removing motion blurs. BAM separates the deblurring process into two branches, Blur-aware Attention (BA) and Cascaded Parallel Dilated Convolution (CPDC), where BA locates region-wise blur orientations and magnitudes while CPDC adaptively removes blurs based on the attended blurred features. Based on an observation of directional and regional averaging artifacts caused by dynamic blurs, the proposed BA derives region-wise attention by using computationally inexpensive regional averaging to capture blurred patterns of different orientations and magnitudes globally and locally. To derive the orientations and magnitudes of different blurred regions in an image, we reassemble horizontal and vertical blurred responses to catch irregular blur orientations and utilize multi-scale kernels to learn the magnitudes. CDPC leverages two cascaded multi-scale dilated convolutions to deblur image features. As a result, BANet possesses the superior deblurring capability and can support subsequent real-time applications superbly.

In short, our contributions are two-fold. First, BANet is featured with a novel BAM module that exploits region-wise attention to capture blur orientations and magnitudes, making BANet capable of disentangling blur contents of different degrees in dynamic scenes. With the disentangled region-wise blurred patterns, it then utilizes cascaded multi-scale dilated convolution to restore blurred features. Second, our efficient single-forward-pass deep networks perform favorably against state-of-the-art methods with fast inference time.

II Related Work

Dynamic Scene image deblurring is a highly ill-posed problem since blurs stem from various factors in the real world. Conventional image deblurring studies often make different assumptions, such as uniform or non-uniform blurs, and image priors , to model blur characteristics. Namely, these methods impose different constraints on the estimated blur kernels, latent images, or both with handcrafted regularization terms for blur removal. Nevertheless, these methods often attempt to solve a non-convex optimization problem and involve heuristic parameter tuning that is entangled with the camera pipeline; thus, they cannot generalize well to complex real-world examples.

II-B Deblurring via Learning

Learning-based approaches with self-recurrent modules gain great success in single-image deblurring. Particularly, the coarse-to-fine schemes can gradually restore a sharp image on different resolutions (MS) , fields of view (MP) , or temporal characteristics (MT) . Despite the success, self-recurrent models usually lead to longer inference runtime. Recently, non-recurrent methods were proposed for efficient deblurring. For instance, Kupyn et al. suggested using conditional generative adversarial networks to restore blurred images. However, these methods do not well address non-uniform blurs in dynamic scenes, often causing blur artifacts in the deblurred images. To address this issue, Yuan et al. proposed a spatially variant deconvolution network with optical flow estimation to guide deformable convolutions and capture moving objects during model training. Li et al. proposed a depth-guided model for deblurring. However, the optical flow and depth information may not always correlate with blur, which may cause less effective deblurring. Cho et al. proposed an efficient multi-scale deblurring structure with a multi-input multi-output. With multi-scale input, the process adopts a shallow convolution to turn the images into attention masks and multiply them by the same scales’ features. However, its simple feature attention mechanism may not be able to extract blur information comprehensively from an input image, hence limiting its deblurring performance.

II-C Self-attention

Self-attention (SA) has been widely adopted to advance the fields of image processing and computer vision . Recent advances revealed that attention is beneficial for learning inter-pixel correlations to emphasize different local features for removing non-uniform blur. Specifically, Purohit et al. proposed to deblur using SA to explore pixel-wise correlation for non-local feature adaptation. However, since SA requires much memory in O(H2W2)\mathcal{O}(H^{2}W^{2}) space, where HH and WW are the height and width of the input to SA, the method can only apply SA to the smallest-scale features (from a 1280×7201280\times 720 blurred input to 160×90160\times 90 SA’s input), limiting the efficacy of SA. Also, motion blurs cause directional and local averaging artifacts, which merely pixel-wise SA may not address well. Suin et al. proposed an MP architecture with less memory-intensive SA by using global average pooling with space complexity O(dadcHW)\mathcal{O}(d_{a}d_{c}HW), where dad_{a} is the channel dimension of the components query and key in SA, dcd_{c} is the dimension of the component value, and dadc<HWd_{a}d_{c}<HW. Despite the method’s less space complexity, compressing pixel information into the channel domain may lose spatial information, thus degrading deblurring performance. In contrast, we propose an efficient and low memory-cost regional averaging SA to capture non-uniform blur information more accurately. It is with space complexity O(CHW)\mathcal{O}(CHW), where CC is the number of output channels. It can deblur high-resolution input images and achieve superior performance in real-time.

III Proposed Approach

We present the blur-aware attention network (BANet) to address the potential issues in two commonly used techniques for deblurring: self-recurrence and self-attention. Self-recurrent algorithms result in longer inference time due to repeatedly accessing input blurred images. Self-attention based on inter-pixel or inter-channel correlations is memory intensive and cannot explicitly capture regional blurring information. Instead, the proposed BANet is a one-pass residual network consisting of a series of stacked blur-aware modules (BAMs), which serve as the building blocks, to disentangle different patterns of blurriness and remove blurs based on the attended blurred features.

As illustrated in Fig. 3, BANet starts with two convolutional layers, which contain a stride of 22 to downsample the input image to half resolution. BANet employs one transposed convolutional layer to upsample features to the original size. In between, we stack a set of BAMs to correlate regions with similar blur and extract multi-scale content features. A BAM consists of two components, BA and CPDC, where BA distills global and local blur orientations and magnitudes, and CPDC captures multi-scale blurred patterns to eliminate blurs adaptively. Combining BA and CPDC, BAM is a residual-like architecture that derives both global and local multi-scale blurring features in a learnable manner. We detail the two key components, BA and CPDC, in the following.

To accurately restore the motion area displaying directional and averaging artifacts caused by object motions and camera shakes, we propose a region-based self-attention module, called BA, to capture such effects in the global (image) and local (patch) scales. As shown in Fig. 4, BA contains two cascaded parts: multi-kernel strip pooling (MKSP) and attention refinement (AR). MKSP catches multi-scale blurred patterns of different magnitudes and orientations, followed by AR to refine them locally.

Hou et al. presented an SP (strip pooling) method that uses horizontal and vertical one-pixel long kernels to extract long-range band-shape context information for scene parsing. SP averages the input features within a row or a column individually and then fuses the two thin-strip features to discover global cross-region dependencies. Let the input feature maps x=[xi,j,c]∈RH×W×C{\bf x}=[x_{i,j,c}]\in R^{H\times W\times C}, where CC denotes the number of channels. Applying SP to x{\bf x} generates a vertical and a horizontal tensor followed by a 1D convolutional layer with a kernel size of 33. This produces a vertical tensor yv=[yi,cv]∈RH×C{\bf y}^{v}=[y_{i,c}^{v}]\in R^{H\times C} and a horizontal tensor yh=[yj,ch]∈RW×C{\bf y}^{h}=[y_{j,c}^{h}]\in R^{W\times C}, where yi,cv=1W∑j=0W−1xi,j,cy_{i,c}^{v}=\frac{1}{W}\sum^{W-1}_{j=0}x_{i,j,c} and yj,ch=1H∑i=0H−1xi,j,cy_{j,c}^{h}=\frac{1}{H}\sum^{H-1}_{i=0}x_{i,j,c}. The SP operation, after a convolution layer, fuses the two tensors into y=[yi,j,c]∈RH×W×C{\bf y}=[y_{i,j,c}]\in R^{H\times W\times C}, where yi,j,c=yi,cv+yj,chy_{i,j,c}=y_{i,c}^{v}+y_{j,c}^{h}, and then turns the fused tensor into an attention mask Msp{\bf M}_{sp} as

where f1f_{1} is a 1×11\times 1 convolutional layer and σsig(⋅)\sigma_{sig}(\cdot) is the sigmoid function. Although SP has shown its effects on segmenting band-shape objects for scene parsing, it is unsuitable to directly apply SP to an image deblurring task, aiming to locate blurred patterns that tend to involve different orientations and magnitudes, and restore a sharp image.

Motivated by SP, we propose MKSP that adopts strip pooling with different kernel sizes to discover regional and directional averaging artifacts caused by dynamic blurs.

MKSP combines and compares multiple sizes/scales of averaging results followed by concatenation and convolution to catch blurred patterns of different magnitudes and orientations. The idea behind our design is to reassemble different orientations by horizontal and vertical operations on multi-scale results, e.g., the difference between consecutive kernel sizes, and reveal the scales of blurred patterns. We apply convolutional layers to automatically discover these blur-aware operations on the feature level to learn irregular attended features rather than a fixed cropping method on the image level used in MP methods . MKSP averages the input tensors within rows and columns by adaptive average pooling to generate H×n×CH\times n\times C and n×W×Cn\times W\times C long features, where n∈{1,3,5,7}n\in\{1,3,5,7\} represents different scales. Thus, MKSP generates four pairs of tensors, each of which has a vertical and a horizontal tensor followed by a 1D (for n=1n=1) or 2D (for the rest) convolutional layer with the kernel size of 33 or 3×33\times 3, respectively. This produces the vertical tensor yv,n∈RH×n×C{\bf y}^{v,n}\in R^{H\times n\times C} and the horizontal tensor yh,n∈Rn×W×C{\bf y}^{h,n}\in R^{n\times W\times C}, where the vertical tensor is

where the horizontal stride Sh=⌊Wn⌋S_{h}=\lfloor\frac{W}{n}\rfloor and the horizontal-strip kernel size Kh=W−(n−1)ShK_{h}=W-(n-1)S_{h}. Symmetrically, the horizontal tensor is defined by

where the vertical stride Sv=⌊Hn⌋S_{v}=\lfloor\frac{H}{n}\rfloor and the vertical-strip kernel size Kv=H−(n−1)SvK_{v}=H-(n-1)S_{v}.

After determining the horizontal and vertical magnitudes, the orientations of blur patterns are estimated jointly considering the two orthogonal magnitudes. More specifically, MKSP, after a 1D (for n=1n=1) or 2D (for the rest) convolutional layer, fuses each pair of tensors (yv,n{\bf y}^{v,n}, yh,n{\bf y}^{h,n}) into a tensor yn∈RH×W×C{\bf y}^{n}\in R^{H\times W\times C} by

Similar to SP, we concatenate all the fused tensors to yield an attention mask as Mmksp=fout(y1⊕y3⊕y5⊕y7){\bf M}_{mksp}=f_{out}({\bf y}^{1}\oplus{\bf y}^{3}\oplus{\bf y}^{5}\oplus{\bf y}^{7}), where ⊕\oplus stands for the concatenation operation, and fout(⋅)=σsig(Conv(σReLU(Conv(⋅))))f_{out}(\cdot)=\sigma_{sig}(Conv(\sigma_{ReLU}(Conv(\cdot)))) represents a non-linear mapping function consisting of two 3×33\times 3 convolutional layers. The first layer uses the ReLU activation function, and the second uses a sigmoid function. As shown in Fig. 5, the proposed MKSP can generate attention masks that better fit objects or local scenes than those by using SP with only H×1H\times 1 and 1×W1\times W kernels used, which yields rough band-shape masks.

Attention Refinement (AR)

After obtaining the globally attended features by the element-wise multiplication of attention masks Mmksp{\bf M}_{mksp} and input tensor x{\bf x}, we further refine these features locally via a simple attention mechanism using fAR(⋅)f_{AR}(\cdot). The final output of our BA block through the MKSP and AR stages is computed as

The proposed BA facilitates the attention mechanism applied to deblurring since it requires less memory, i.e. O(HWC)\mathcal{O}(HWC), where CC represents the channel dimension, than those adopted in . It disentangles blurred contents with different magnitudes and orientations. Fig. 7 showcases three examples of blur content disentanglement using BA, where we witness that background scenes are differentiated from the foreground scenes because those objects closer to the camera move faster, thus more blurred. Fig. 8 shows more examples of attention maps yielded by BA, which implicitly acts as a gate for propagating relevant blur contents.

III-B Cascaded Parallel Dilated Convolution (CPDC)

Atrous convolution, also called dilated convolution, has been widely applied to computer-vision tasks for enlarging receptive fields and extracting features from objects with different scales without increasing the kernel size. Inspired by this, we design a cascaded parallel dilated convolution (CPDC) block with multiple dilation rates to capture multi-scale blurred objects. Instead of stacking dilated convolutional layers with different rates in parallel, which we call parallel dilated convolution (PDC), our CPDC block cascades two sets of PDC with a single convolutional layer working as a fusion bridge. It can distill patterns more beneficial to deblurring before passing through the second PDC. As an example, Fig. 9(a) shows a PDC block consisting of three 3×33\times 3 dilated convolutional layers with a dilation rate DD (D=1,3,D=1,3, and 55), each of which outputs features with half the number of input channels. After concatenation, the number of the output channels of the PDC block increases by 1.51.5 times. As shown in Fig. 9(b), our CPDC block consists of two PDC blocks bridged by a 3×33\times 3 convolutional layer, which would be more effective in aggregating multi-scale content information for deblurring.

III-C Loss function

In BANet, we utilize the Charbonnier loss as suggested in :

where R\bf{R} and Y\bf{Y} respectively denote the restored image and the ground-truth image, and ε=10−3\varepsilon={10}^{-3} as in . In addition, to enhance the restoration performance, we add an FFT loss to supervise the results in the frequency domain, as adopted in MIMO-UNet+ :

where F\mathcal{F} represents the fast Fourier transform function. At last, we optimize BANet using the total loss LL as

where λ\lambda is set to 0.010.01 empirically.

IV Experiments

This section evaluates the proposed method. In the following, we first describe the experimental setup, then compare our method with the state-of-the-arts, and finally conduct ablation studies to analyze the effectiveness of individual components. The authors from the universities in Taiwan completed the experiments on the datasets.

We evaluate the BANet on three image deblurring benchmark datasets: 1) GoPro that consists of 3,2143,214 pairs of blurred and sharp images of resolution 720×1280720\times 1280, where 2,1032,103 pairs are used for training, and the rest for testing, 2) HIDE that contains 2,0252,025 pairs of HD images, all for testing, and RealBlur that consists of 3,7583,758 pairs for training and 980980 pairs for testing. The RealBlur dataset is further split into two datasets: RealBlur-R collected from raw images and RealBlur-J from JPEG images. We train our model using Adam optimizer with parameters β1=0.9\beta_{1}=0.9 and β2=0.999\beta_{2}=0.999. We set the initial learning rate to 10−410^{-4}, which then decays to 10−710^{-7} based on the cosine annealing strategy. Following , we utilize random cropping, flipping, and rotation for data augmentation. Lastly, we implement our model with PyTorch library on a computer equipped with Intel Xeon Silver 4210 CPU and NVIDIA 2080ti GPU.

IV-B Experimental Results

Quantitative Analysis: We compare our method with 11 latest approaches, including MSCNN , SRN , DSD , DeblurGAN-v2 , DMPHN , EDSD , MTRNN , RADN , SAPHN , MIMO-UNet+ , and MPRNet , which also handle dynamic deblurring on the GoPro test set. For HIDE , we choose nine recent deblurring methods, including DeblurGAN-v2 , SRN , HAdeblur , DSD , DMPHN , MTRNN , SAPHN , MIMO-UNet+ , and MPRNet , according to their availability in released pre-trained weights. For RealBlur , we choose four methods that trained on the RealBlur training set, including DeblurGAN-v2 , SRN , MPRNet , and MIMO-UNet+ .

To better compare with recent approaches, we devise two versions of our model, BANet and BANet+. The only difference between them is the number of channels used in a BAM, and BANet with 128 channels involves 18 million parameters while BANet+ has 40 million parameters with 192 channels. Table I lists the objective scores (PSNR and SSIM), runtime, parameters, and GFLOPs on the GoPro test set for all the compared methods. We observe that the self-recurrent models, MSCNN , SRN , DSD , MTRNN , SAPHN , and MPRNet , consume longer runtime than the non-recurrent ones, i.e., DeblurGAN-v2 , RADN , MIMO-UNet+ , and ours. As reported in Table I, BANet runs faster with fewer parameters and GFLOPs as well as achieves better performance than recurrent-based methods, MSCNN , SRN , DSD , DMPHN , MTRNN , and SAPHN and non-recurrent methods, such as DeblurGAN-v2 and RADN on the GoPro test set. BANet also performs favorably against an efficient multi-scale model, MIMO-UNet+ , with the same runtime and a comparable model size. BANet+ outperforms the best competitor, MPRNet , by 0.37 dB in PSNR with faster runtime (−113-113ms) and lower GFLOPs (−172-172). Table II shows the quantitative results on HIDE . As can be seen, BANet outperforms all the compared methods except for MPRNet with a faster inference time. BANet+ only works comparably to MPRNet since MPRNet seems to perform favorably on HIDE particularly, but our model runs much faster. Table III lists the quantitative comparisons on the RealBlur test set, demonstrating that both BANet and BANet+ outperform the compared methods on the RealBlur-J and RealBlur-R test sets.

Qualitative Analysis: Fig. 10 and Fig. 11 show qualitative comparisons on the GoPro test set and HIDE dataset with previous state-of-the-arts MTRNN , DSD , DMPHN , MIMO-UNet+ , and MPRNet . As observed in Fig. 10, MTRNN , DSD , DMPHN , MIMO-UNet+ , and MPRNet do not well recover regions with texts or severe blurs whereas BANet can restore those regions better. In Fig. 11, MTRNN , DSD , DMPHN , and MIMO-UNet+ do not deblur the striped t-shirt and texts well, while BANet recovers those parts better. Fig. 12 and Fig. 13 demonstrate some deblurred results using DeblurGAN-v2 , SRN , MPRNet , MIMO-UNet+ , and ours, on the RealBlur test set. As can be seen, although all these models can remove blurs, BANet performs favorably on delicate image details.

User Study: We further conduct a user study to evaluate the subjective quality of deblurred results on real blurred images chosen from the RealBlur-J test set. We compare our method (BANet+) against four methods, including MIMO-UNet+ , MPRNet , SRN , and DeblurGAN-v2 . Note that all the methods are trained on the RealBlur-J training set.

In the study, 34 subjects aged from 21 to 40 years participated in the study without any prior knowledge of the experiment. Their vision is either normal or corrected to be normal. We picked 1616 blurred images with varying scenes for the experiment and obtained the deblurred results using all the compared approaches. Since each method is compared against BANet with all the chosen blurred images in the experiment, we have 16×4=6416\times 4=64 image pairs in total. Each subject is shown all the image pairs, one at a time, and asked which one he/she prefers in terms of visual quality. Each image pair is displayed randomly and placed side by side. Subjects are asked to check images carefully before choosing without a time limit.

Table IV shows the subjective evaluation results, where the values represent the percentage that the deblurring results with our method are preferred to the counterparts with the other compared methods for all the votes collected. It indicates that our method obtains over 95% preference votes compared to all the compared methods, which again demonstrates that our approach achieves better subjective visual quality.

IV-C Ablation Study

In the ablation studies, we do all the experiments on the BANet (18M) version.

BAM with Different Components: Table V shows an ablation on different component combinations in our Blur-Aware Module (BAM) tested on the GoPro test set. As can be seen, adding a simple attention refinement (AR) mechanism to PDC (Net1 vs. Net2) can improve PSNR by 0.41 dB, which shows the effectiveness of spatial attention for deblurring. Using MKSP in PDC (Net1 vs. Net4) improves PSNR by 0.72 dB, which has much more performance gain compared to using strip pooling (SP) in Net3 or AR in Net2. Substituting PDC in Net5 with CPDC (Net6), our proposed version of BAM leads to a further performance gain. Thanks to its mechanism for locating blur regions based on both global attention and local convolutions, our BAM attains the best performance while achieving fast inference time.

Numbers of Stacked BAMs: Using more layers to enlarge the receptive field may improve performance for computer vision or image processing tasks. Nevertheless, stacking more layers for deblurring does not guarantee better performance and might consume extra inference time. However, using our residual learning-based BAM design, we can stack multiple layers to expand the effective receptive field for better deblurring. In Table VI, we show performance comparisons with various numbers of BAMs stacked in our model on the GoPro test set. We list four versions: stack-4, stack-8, stack-10, and stack-12, corresponding to 4, 8, 10, and 12 BAMs stacked in BANet. Although the quantitative performance improves with the number of BAMs, the improvement became saturated after 12. Therefore, we choose 10 for its excellent balance between efficiency and visual quality.

Effectiveness of MKSP and CPDC: In Table VII, we investigate the effects of kernel combination of MKSP on the GoPro test set. MKSP with five kernel sizes of 11, 33, 55, 77, and 99 performs a little worse than the first four sizes (11, 33, 55, and 77), indicating that adding the kernel size of 99 would not catch blur features more accurately, thus not helping with the performance. In Table VIII, we verify that CPDC, which uses a single convolution as a fusion bridge, outperforms PDC. For a fair comparison, we also compare CPDC against a PDC variant that stacks two PDCs in a series, called PDC2, with a similar parameter size, and CPDC still performs better.

IV-D Blur-aware Attention vs. Self-Attention

RADN utilizes a similar self-attention (SA) mechanism proposed in for deblurring. It helps connect regions with similar blurs to facilitate global access to relevant features across the entire input feature maps. However, its high memory usage makes applying it to high-resolution images infeasible. Thus, SA is usually employed in network layers on a smaller scale like in RADN , where important blur information would be lost due to down-sampling. In contrast, our proposed region-based attention is more suitable for correlating regions with similar blur characteristics. Moreover, it can process high-resolution images thanks to its low memory consumption. To further demonstrate our BA’s efficacy, we compare the SA with BA using our BANet (stack-4) as a backbone network, as shown in Fig. 14(b). Due to the high memory demand for SA (O(H2W2)\mathcal{O}(H^{2}W^{2})) to process 720×1280720\times 1280 images, we adopt our stack-4 model for training. When testing the networks, we separate the input image into eight sub-images for both SA and BA to deblur, each equipped with a single 2080ti GPU. Since our BA requiring lower memory usage (O(CHW)\mathcal{O}(CHW), where C<<H×WC<<H\times W) can process the image with the full resolution, we also show its result. In Table IX, SA∗ and BA∗ represent deblurring an image with its eight sub-images separately, whereas BA for processing the entire image at once. As can be observed, the proposed BA∗ works much more efficiently than SA∗ with a comparable result. When deblurring the entire image at once, BA undoubtedly performs the best.

V Conclusion

This paper proposes a novel blur-aware attention network (BANet) for single image deblurring. BANet consists of stacked blur-aware modules (BAMs) to disentangle region-wise blur contents of different magnitudes and orientations and aggregate multi-scale content features for more accurate and efficient dynamic scene deblurring. We have investigated and examined our design through demonstrations of attention masks and attended feature maps, as well as extensive ablation studies and performance comparisons. Our extensive experiments demonstrate that the proposed BANet achieves real-time deblurring and performs favorably against state-of-the-art deblurring methods on the GoPro and RealBlur benchmark datasets.

References