Flow-Guided Sparse Transformer for Video Deblurring

Jing Lin, Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Youliang Yan, Xueyi Zou, Henghui Ding, Yulun Zhang, Radu Timofte, Luc Van Gool

Introduction

Video deblurring is a fundamental yet challenging task in low-level computer vision and graphics communities. It aims to restore the latent frames from a blurry video sequence. Serving as a preprocessing technique, video deblurring has wide applications such as video stabilization (Matsushita et al., 2006), tracking (Jin et al., 2005), autonomous driving (Yin et al., 2021), etc. Hand-held devices are more and more popular in capturing videos of dynamic scenes, where prevalent depth variations, abrupt camera shakes, and high-speed object movements lead to undesirable blur in videos. To alleviate the effect of motion blur, researchers have put a lot of efforts into video deblurring.

Conventional methods are mainly based on hand-crafted priors and assumptions, which limits the model capacity. Besides, the assumptions on motion blur and latent frames usually lead to complex energy functions that are difficult to solve. Also, the inaccurately estimated motion blur kernel with hand-crafted priors may easily result in severe artifacts.

In the past decade, video deblurring has witnessed significant progresses with the development of deep learning. Convolutional neural network (CNN) applies a powerful model to learn the mapping from blurry videos to sharp videos under the supervision of a large-scale dataset of blurry-sharp video pairs. CNN-based methods yield impressive performance but show limitations in modeling long-range spatial dependencies and capturing non-local self-similarity.

Recently, the emergence of Transformer provides an alternative to alleviate the constraints of CNN-based methods. Firstly, Transformer excels at modeling long-range spatial dependencies. The contextual information and spatial correlations are critical to restoring the motion blur. Secondly, similar and sharper scene patches from neighboring frames provide crucial cues for video deblurring. Fortunately, the self-attention module in Transformer is dedicated to calculating the correlations among pixels and capturing the self-similarity along the temporal sequence. Thus, Transformer inherently resonates with the goal of learning similar information from spatio-temporal neighborhoods. Nevertheless, directly using existing Transformers for video deblurring has two issues. On one hand, when the standard global Transformer (Dosovitskiy et al., 2021) is utilized, the computational cost is quadratic to the spatio-temporal dimensions. This burden is nontrivial and sometimes unaffordable. Meanwhile, the global Transformer attends to redundant keykey elements, which may easily cause non-convergence issue (Zhu et al., 2020) and over-smoothing results (Li et al., 2019). On the other hand, when the local window-based Transformer (Liu et al., 2021) is used, the self-attention is calculated within position-specific windows, causing limited receptive fields. The model may neglect some keykey elements of similar and sharper scene patches in the spatio-temporal neighborhood when fast motions are present. We summarize the main reason for the above problems, i.e., previous Transformers lack the guidance of motion information, when calculating self-attention. We note that the motion information can be estimated by optical flow.

Exploiting an optical flow estimator to capture motion information and align neighboring frames is a common strategy in video restoration (Makansi et al., 2017; Su et al., 2017; Xue et al., 2019; Pan et al., 2020). Previous flow-based methods mainly adopt the pre-warping strategy. Specifically, they employ an optical flow estimator to produce motion offsets, warp neighboring frames, and align regions corresponding to the same object but misaligned in neighboring image or feature domains. This scheme suffers from the following issues: (i) The interpolating operations in the warping module modify the original image information. As a result, some critical image priors such as self-similarity and sharp textures may be sacrificed. Undesirable artifacts may be introduced to the restored video and the deblurring performance may degrade. (ii) The frame alignment and subsequent representation aggregation are separated. This paradigm is inflexible and does not make full use of optical flow. Besides, the deblurring results are easily affected by the performance of the optical flow estimator. The robustness of this scheme can be further improved.

This work aims to cope with the above problems. We propose a novel method, Flow-Guided Sparse Transformer (FGST), for video deblurring. Firstly, we adopt Transformer instead of CNN as the deblurring model because of its advantages of capturing long-range spatial dependencies and non-local self-similarity. Secondly, to alleviate the limitations of previous Transformers and the pre-warping strategy, we customize Flow-Guided Sparse Multi-head Self-Attention (FGS-MSA) as shown in Fig. 1 (a). For each queryquery element on the reference frame, FGS-MSA guided by an optical flow estimator globally samples spatially sparse keykey elements corresponding to the same scene patch but misaligned in the neighboring frames. These sampled keykey elements provide self-similar and highly related image prior information, which is critical to restoring motion blur. Different from original global and local Transformers, our FGST neither blindly samples redundant keykey elements nor suffers from limited receptive fields. Meanwhile, our alignment scheme is different from the pre-warping operation mainly used by previous flow-based methods. Instead of warping the neighboring frames, our FGST samples keykey elements in consecutive frames to calculate the self-attention. Thus, the original image prior information can be preserved. Thirdly, we promote FGS-MSA to Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA) as shown in Fig. 1 (b). The feature maps are split into non-overlapping windows. Instead of sampling a single keykey element on each neighboring frame for a single queryquery element, FGSW-MSA samples keykey elements assigned by the optical flow corresponding to all the queryquery elements of the window on the reference frame. Thus, FGSW-MSA is more robust to accommodate pixel-level flow offset prediction deviations. Finally, our FGSW-MSA is calculated within a short temporal sequence reducing the computational cost. Hence, the receptive field of FGSW-MSA is spatially global but temporally local. Motivated by RNN-based methods (Nah et al., 2019; Zhong et al., 2020), we propose Recurrent Embedding (RE) to transfer information of past frames and capture long-range temporal dependencies.

Our contributions can be summarized as follows:

We propose a new method, FGST, for video deblurring. To the best of our knowledge, it’s the first attempt to explore the potential of Transformer in this task.

We customize a novel self-attention mechanism, FGS-MSA, and its improved version, FGSW-MSA.

We design an embedding scheme, RE, to transfer frame information and capture temporal dependencies.

Our FGST outperforms SOTA methods on DVD and GOPRO datasets by a large margin and yields more visually pleasing results in real-world video deblurring.

Related Work

In recent years, the deblurring research focus is shifting from single image deblurring (Zoran et al., 2011; Chakrabarti, 2016; Purohit et al., 2020) to the more challenging video deblurring (Cho et al., 2012; Matsushita et al., 2006). Traditional methods (Li et al., 2010; Zhang et al., 2013) are based on hand-crafted image priors and assumptions, which lead to limited generality and representing capacity. With the development of deep learning, recent methods are mainly CNN-based or RNN-based. (Zhang et al., 2018) employ 3D convolutions to model spatio-temporal relations of frames. (Hyun Kim et al., 2017) and (Nah et al., 2019) use RNN-based models to restore the latent frames. However, CNN-based methods show limitations in capturing long-range dependencies while RNN-based methods are not sensitive to patch-level spatial correlation and motion information.

2 Vision Transformer

Transformer is firstly proposed by (Vaswani et al., 2017) for machine translation. Recently, Transformer has been introduced to high-level (Dosovitskiy et al., 2021; Liu et al., 2021; Zhu et al., 2020; Zheng et al., 2021; El-Nouby et al., 2021; Carion et al., 2020; Li et al., 2021b; Ramachandran et al., 2019; Wu et al., 2020; Cai et al., 2020) and low-level vision (Chen et al., 2021; Cai et al., 2021b; Wang et al., 2021; Cao et al., 2021b; Cai et al., 2021a; Hu et al., 2021). (Arnab et al., 2021) factorize the spatial and temporal dimensions of the input video and propose a Transformer model for video classification. (Chen et al., 2021) present a large model IPT pre-trained on large-scale datasets with a multi-task learning scheme. (Cao et al., 2021b) propose VSR-Transformer that uses the self-attention mechanism for better feature fusion in video super-resolution, but image features are still extracted from CNN. (Wang et al., 2021) use Swin Transformer (Liu et al., 2021) blocks to build up a U-shaped structure for single image restoration. In (Vaswani et al., 2021; Cao et al., 2021a; Liu et al., 2021), window-based local self-attention is adopted to replace the global self-attention module of the standard Transformer. However, directly using previous global or local Transformers for video deblurring leads to unaffordable computational cost or limited receptive fields.

3 Flow-based Video Restoration

Optical flow estimators are widely used in video restoration tasks (Gast & Roth, 2019; Xue et al., 2019; Gong et al., 2017; Sun et al., 2015; Makansi et al., 2017; Su et al., 2017; Pan et al., 2020) to align highly related but mis-aligned frames. Previous flow-based video deblurring methods (Xue et al., 2019; Makansi et al., 2017; Su et al., 2017; Pan et al., 2020; Gast & Roth, 2019) mainly adopt the pre-warping strategy, which firstly estimates the optical flow and then warps the neighboring frames. For example, (Su et al., 2017) experiments with pre-warping input images based on classic optical flow methods to register them to the reference frame. Nonetheless, this flow-based pre-warping scheme separates the frame alignment and subsequent information aggregation. The original frame information is sacrificed and the guidance effect of optical flow is not fully explored.

Method

Figure 2 (a) shows the architecture of FGST that adopts the widely used U-shaped structure, consisting of an encoder, a bottleneck, and a decoder. Figure 2 (b) depicts the basic unit of FGST, i.e.i.e., Flow-Guided Attention Block (FGAB).

Subsequently, following the spirit of U-Net (Ronneberger et al., 2015), we customize a symmetrical decoder, which is composed of two FGABs and patch expanding layers. The patch expanding layer is a strided 2×\times2 deconvolution that upsamples the feature maps. To alleviate the information loss caused by downsampling, skip connections are used for feature fusion between the encoder and decoder.

2 Flow-Guided Attention Block

As analyzed in Sec. 1, the standard global Transformer brings quadratic computational complexity with respect to the token number and easily leads to non-convergence issue and over-smoothing results. The previous window-based local Transformers suffer from the limited receptive fields.

To address these problems, we propose to use optical flow as the guidance to sample keykey elements from spatio-temporal neighborhoods when calculating the self-attention. Based on this motivation, we customize the basic unit, FGAB as shown in Fig. 2 (b). FGAB consists of a layer normalization (LN), a Flow-Guided Sparse Window-based Multi-head Self-Attention (FGSW-MSA), a feed-forward network (FFN), and two identity mappings. The FFN is composed of 5 consecutive residual blocks. In this part, we first introduce Flow-Guided Sparse Multi-head Self-Attention (FGS-MSA) and then its improved version, FGSW-MSA.

where rr represents the temporal radius of the neighboring frames. (Δxf,Δyf)({\Delta x_{f}},{\Delta y_{f}}) denotes the value at position (i,ji,j) of the estimated motion offset map, which is predicted from the reference frame vt\boldsymbol{v}_{t} to the neighboring frame vf\boldsymbol{v}_{f}:

where FoF_{o} denotes the mapping function of the optical flow estimator and [⋅\cdot] refers to the rounding operation. Subsequently, FGS-MSA can be formulated as

The standard global MSA leads to quadratic ((THW)2(THW)^{2}) computational complexity while our proposed FGS-MSA contributes to much cheaper linear computational cost with respect to the token number (THW)(THW). Detailed analysis are provided in the supplementary material (SM).

FGSW-MSA. For each neighboring frame, FGS-MSA only samples a single keykey element. When the optical flow estimation is inaccurate, the deblurring performance may be easily affected. To further improve the robustness and reliability of our method, we promote FGS-MSA to FGSW-MSA. As shown in Fig. 1 (b), the feature maps are split into non-overlapping windows. The spatial size of each window is M×MM\times M. Φi,jt\mathbf{\Phi}_{i,j}^{t} denotes the set of queryquery elements in the window centering at position (i,j)(i,j) of the ttht_{th} frame:

For each qm,nt∈Φi,jt\boldsymbol{q}_{m,n}^{t}\in\mathbf{\Phi}_{i,j}^{t}, FGSW-MSA samples not only its corresponding keykey elements in Ωm,nt\mathbf{\Omega}_{m,n}^{t} (Eq. (1)) assigned by the flow offsets but also the keykey elements corresponding to other queryquery elements in Φi,jt\mathbf{\Phi}_{i,j}^{t}. We denote the set of these keykey elements as Ψi,jt\mathbf{\Psi}_{i,j}^{t}, which can be formulated as

Instead of attending to a single keykey element on each neighboring frame for a single queryquery, FGSW-MSA pays attention to the keykey elements from similar and sharper scene patches corresponding to all queryquery elements in Φi,jt\mathbf{\Phi}_{i,j}^{t}. The attending region is enlarged from pixel to window. Thus, FGSW-MSA is more robust to accommodate pixel-level flow prediction deviations. FGSW-MSA can be formulated as

Given the input V\mathbf{V}, the computational complexity is

The computational cost of FGSW-MSA is linear with respect to the number of tokens (THWTHW). Eq. (LABEL:eq:complexity) and (9) reveal the high efficiency and resource economy of our FGST. Please refer to the SM for more detailed analysis.

Discussion. (i) Our FGSW-MSA enjoys much larger receptive fields than W-MSA (Liu et al., 2021). Specifically, according to Eq. (1), (2), (6), and (7), the receptive field of FGSW-MSA can cover the whole input feature map when the estimated flow offset is large enough. In practice, the motion offset predicted by the optical flow estimator between two adjacent frames can reach 40 and 38 pixels on GOPRO and DVD datasets. The input spatial size is 256×\times256. MM is set to 3. Thus, the receptive field of FGSW-MSA can reach 83×\times83 (83 = 40×\times2+3) and 79×\times79 while that of W-MSA is still 3×\times3. (ii) Unlike previous flow-based methods that adopt the pre-warping operation sacrificing the original image information as shown in Fig. 3, our FGST combines motion cues with self-attention calculation. Thus, the original image information can be preserved and the guidance effect of the optical flow can be further explored. In addition, our flow-guided scheme enjoys higher flexibility and robustness because adjacent FGABs sample contents independently. Please refer to the SM for detailed discussions.

3 Recurrent Embedding

Our FGSW-MSA is calculated within a short temporal sequence for the computational complexity consideration (approximately linear to the temporal radius rr in Eq. (9) ). Therefore, the receptive field of FGSW-MSA is temporally local and overlooking the distant frames limits the video deblurring performance. To further capture more robust long-range temporal dependencies, we propose Recurrent Embedding (RE) mechanism. RE is motivated by Recurrent Neural Network (RNN). More specifically, as shown in Fig. 2 (c), we exploit RE in each Transformer layer to transfer information from past frames and establish long-range temporal correlations. With RE, the FGAB is calculated in a recurrent manner for TT time steps. ytl\boldsymbol{y}^{l}_{t}, etl,qtl\boldsymbol{e}^{l}_{t},\boldsymbol{q}^{l}_{t}, ktl\boldsymbol{k}^{l}_{t} respectively denote the output, RE, queryquery elements, and keykey elements of the lthl_{th} FGAB in the ttht_{th} time step. We have

where fw(⋅)f_{w}(\cdot) represents the spatial warping that align the feature map at tt and t−1{t-1} time step, [⋅\cdot,⋅\cdot] is the concatenating operation, fc(⋅)f_{c}(\cdot) denotes 3×\times3 convolution to aggregate the recurrent embedding etl\boldsymbol{e}_{t}^{l} and the output from last FGAB layer ytl−1\boldsymbol{y}_{t}^{l-1}, and ytl=FGAB(qtl,ktl)\boldsymbol{y}_{t}^{l}=\text{FGAB}(\boldsymbol{q}_{t}^{l},\boldsymbol{k}_{t}^{l}) is formulated in details as

where LN denotes the layer normalization and FFN refers to the Feed Forward Network. Our RE sequentially propagates the information from the first frame to the last frame, thus capturing reliable long-range temporal dependencies.

Experiment

DVD. The DVD (Su et al., 2017) dataset consists of 71 videos with 6,708 blurry-sharp image pairs. It is divided into train/test subsets with 61 videos (5,708 image pairs) and 10 videos (1,000 image pairs). DVD is captured with mobile phones and DSLR at a frame rate of 240 fps.

GOPRO. The GOPRO (Nah et al., 2017) benchmark is composed of over 3,300 blurry-sharp image pairs of dynamic scenes. It is obtained by a high-speed camera. The training and testing subsets are split in proportional to 2:1.

Real Blurry Videos. To validate the generality of FGST, we evaluate models on real blurry datasets collected by (Cho et al., 2012). Because the ground truth (GT) is inaccessible, we only compare visual results of FGST and others.

2 Implementation Details

We implement FGST in PyTorch. We adopt a pre-trained SPyNet (Ranjan et al., 2017) as the optical flow estimator. All the modules are trained with the Adam (Kingma & Ba, 2015) optimizer (β1\beta_{1} = 0.9 and β2\beta_{2} = 0.999) for 600 epochs. The initial learning rate is set to 2×\times10-4 and 2.5×\times10-5 respectively for the deblurring model and optical flow estimator. The learning rate is halved every 200 epochs during the training procedure. Patches at the size of 256×\times256 cropped from training frames are fed into the models. The batch size is 8. The temporal radius rr of the neighboring frames is set to 1. The sequence length TT is set to 9 in training and the whole video length in testing. The horizontal and vertical flips are performed for data augmentation. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) (Wang et al., 2004) are adopted as the evaluation metrics. The models are trained with 8 V100 GPUs. L1\mathcal{L}_{1} loss between the restored and GT videos is used for supervision.

3 Quantitative Results

The comparisons between FGST and other SOTA methods are listed in Tabs. 1, 2, and 3c. As can be observed: (i) Our FGST outperforms SOTA methods by a large margin on the two benchmarks. Specifically, as shown in Tab. 1, our FGST surpasses the recent best algorithm ARVo (Li et al., 2021a) by 0.56 dB on DVD. As reported in Tab. 2, our method exceeds Suin et al. (Suin et al., 2021) and TSP (Pan et al., 2020) by 0.80 dB and 1.23 dB respectively on GOPRO. These results demonstrate the effectiveness of our method. (ii) Tab. 3c exhibits efficiency comparisons of different algorithms on GOPRO. The FLOPS is tested at the input size of 1×\times3×\times240×\times240. The running time per frame is tested at the spatial size of 1,280×\times720 on the same RTX 2080 GPU. Our FGST is more cost-effective and achieves a better trade-off between PSNR, Params, FLOPS, and inference speed. For instance, when compared to TSP (Pan et al., 2020), FGST only requires 59.9% (9.70 / 16.19) Params and 36.8% (131.6 / 357.9) FLOPS while achieving even 1.23 dB improvement and 2.34×\times (579.7 / 247.8) speed. This evidence suggests the promising efficiency advantage of our proposed FGST.

4 Qualitative Results

We provide visual comparisons on DVD, GOPRO, and real blurry videos as shown in Figs. 4, 5, and 7. Previous methods are less favorable to restore abrupt motion blur. They either yield over-smoothing images sacrificing fine textural details and structural contents or introduce redundant blotchy texture and chromatic artifacts when fast motions exists. In contrast, our FGST excels at modeling long-range dependencies and exploits motion information to guide the self-attention module to capture non-local self-similarity in spatio-temporal neighborhoods. As a result, FGST is capable of restoring structural contents and textural details while preserving spatial smoothness of the homogeneous regions. Supplementary file provides more visual results.

5 Ablation Study

In this part, we conduct ablation studies on GOPRO dataset. The baseline model is derived by directly removing all the proposed RE and FGSW-MSA modules from our FGST.

Break-down Ablation. We firstly conduct a break-down ablation to investigate the effect of each component toward better performance. The results are reported in Tab. 3a. The baseline model yields 31.18 dB. After applying RE and FGSW-MSA respectively, the deblurring model achieves 1.16 dB and 1.66 dB improvements. While using both RE and FGSW-MSA modules, the model gains by 1.72 dB. The results suggest the effectiveness of RE and FGSW-MSA.

Self-Attention Mechanism. We compare our self-attention mechanisms with other competitors in Tab. 3b. The baseline model yields 31.18 dB while costing 5.15M Params and 43.93G FLOPS. (i) When using global MSA (Dosovitskiy et al., 2021), the feature maps are downsampled into 14\frac{1}{4} size and the channel is increased by 4 times to avoid out of memory and information loss. The deblurring model degrades by 1.98 dB while costing 12.5×\times Params and 3.2×\times FLOPS. This is mainly because global MSA attends to too redundant keykey elements, requiring a large amount of computation and memory resources while leading to ambiguous gradients for input features (Zhu et al., 2020) and thus non-convergence problem. Meanwhile, features from global aggregation tend to over-smooth the predictions of small patterns (Li et al., 2019). (ii) When using local W-MSA (Liu et al., 2021), the model gains by only 0.53 dB while adding 3.11M Params and 64.16G FLOPS. The improvement is limited while the additional burden is nontrivial. That is because W-MSA calculates self-attention within position-specific windows. The receptive field is limited. (iii) Our FGS-MSA exploits the optical flow as the guidance to sample spatially sparse keyskeys of similar and sharper regions in the spatio-temporal neighborhood for each queryquery on the reference frame. Compared to global MSA, the keykey elements of FGST are less but highly related to the selected queryquery. Thus, when using FGS-MSA, the model gains by 1.30 dB while adding 4.54M Params and 81.15G FLOPS. These results show that FGS-MSA costs cheaper resources but achieves better performance than global MSA. When exploiting FGSW-MSA, the model yields an improvement of 1.72 dB while adding 4.55M Params and 87.69G FLOPS. This evidence suggests: (a) FGSW-MSA is more effective than W-MSA in fast motion blur restoration. (b) FGSW-MSA is more reliable than FGS-MSA and achieves better deblurring performance.

In addition, we conduct visual analysis on three adjacent frames by visualizing the last feature map of models with and without (w/o) FGSW-MSA in Fig. 6. Deeper color indicates larger weights. It can be observed that the model without FGSW-MSA responds weakly to similar regions in the neighboring frames. In contrast, the model equipped with FGSW-MSA generates much stronger responses to highly related but misaligned scene patches. Moreover, FGST pays more attention to the regions with fast motion blur. These results demonstrates the effectiveness of FGSW-MSA in capturing non-local self-similarity in dynamic scenes.

Flow-Guided Deformable Convolution. We compare our FGSW-MSA with deformable convolution (DeConv) (Wang et al., 2019) and recent flow-guided deformable convolution (FGDeConv) (Chan et al., 2021) in Tab. 3d. Our proposed FGSW-MSA achieves the most significant improvement. This mainly stems from that FGSW-MSA excels at capturing non-local similarity and long-range dependencies, which are the limitations of CNN-based methods.

Pre-warping Strategy. We compare our FGSW-MSA with the pre-warping strategy mainly adopted by previous methods in Tab. 3e. We start from the baseline model equipped with W-MSA. It can be observed that using FGSW-MSA is 0.30 dB and 0.004 in terms of PSNR and SSIM higher than using pre-warping operation. This performance gap is mainly because the model using our FGSW-MSA can learn from non-corrupted representations of input video and further explore the guidance effect of the optical flow.

Window Size. We change the window size of FGSW-MSA to study its effect. The results are listed in Tab. 3f. We start by setting the window size at 1×\times1 and then gradually increase it. The performance achieves its maximum when the window size is 3×\times3. Thus, the optimal setting is 3×\times3.

Optical Flow Estimator. We adopt three representative optical flow estimators (FlowNet (Dosovitskiy et al., 2015), SPyNet (Ranjan et al., 2017), and PWC-Net (Sun et al., 2018)) to investigate their effects in Tab. 3g. (i) No matter what flow estimator is used, FGST reliably outperforms the baseline model, suggesting the robustness and generality of our method. (ii) The performance of FGST can be further improved by using a better flow estimator. To be specific, when equipped with PWC-Net, FGST is 0.18 dB and 0.13 dB higher than those using FlowNet and SPyNet. These results demonstrate that FGST can directly and conveniently enjoy the benefits of SOTA optical flow estimators.

Conclusion

In this paper, we propose a novel Transformer-based method, FGST, for video deblurring. In FGST, we customize a self-attention mechanism, FGS-MSA, and then promote it to FGSW-MSA. Guided by an optical flow estimator, FGSW-MSA samples spatially sparse but highly related keykey elements corresponding to similar and sharper scene patches in the spatio-temporal neighborhoods. Besides, we present an embedding scheme, RE, to transfer information of past frames and capture long-range temporal dependencies. Comprehensive experiments demonstrate that our FGST significantly surpasses SOTA methods and generates more visually pleasant results in real video deblurring.

Acknowledgements: This work is partially supported by the NSFC fund (61831014), the Shenzhen Science and Technology Project under Grant (CJGJZD20200617102601004, JSGG20210802153150005).

References