KBNet: Kernel Basis Network for Image Restoration

Yi Zhang, Dasong Li, Xiaoyu Shi, Dailan He, Kangning Song, Xiaogang Wang, Hongwei Qin, Hongsheng Li

Introduction

Image restoration is one of the most foundational tasks in computer vision, which aims to remove the unavoidable degradation of the input images and produce clean outputs. Image restoration is highly challenging as it is an ill-posed problem. It does not only play an important role in a wide range of low-level applications (e.g. night sight on smartphones) but also benefits many high-level vision tasks LiuWLWH18 .

Linearly aggregating spatial neighborhood information is a common component and plays a key role in deep neural networks for feature encoding in image restoration. Convolutional neural networks (CNNs) are one of the dominant choices for local information aggregation. They use globally-shared convolution kernels in the convolution operator for aggregating neighboring information, such as using dilated convolutions RIDNet ; chang2020sadnet to increase the receptive fields and adopting multi-stage Zamir2021MPRNet or multi-scale features gu2019self ; zamir2020mirnet for better encoding spatial context. While CNN-based methods show clear performance gains than traditional handcrafted methods BM3D ; NLM ; wavelet ; oldksvd , the convolutions utilized static and spatially invariant kernels for all spatial locations and therefore have limited capacity to handle different types of image structures and textures. While a few works kpn ; xia2020bpn ; mckpn in the low-level task of burst image denoising were proposed to adaptively encode local neighborhoods features, they require heavy computational costs to predict adaptive convolutional kernels of each output pixel location.

Recently, the vision transformers have shown great progress where the attention mechanism uses dot products between pairwise positions to obtain the linear aggregation weights for encoding local information. A few efforts chen2021IPT ; wang2021uformer ; liang2021swinir ; tu2022maxim have been made in image restoration. In IPT chen2021IPT , the global attention layers are adopted to aggregate from all spatial positions, which, however, face challenges on handling high-resolution images. A series of window-based self-attention solutions wang2021uformer ; liang2021swinir ; tu2022maxim have been proposed to alleviate the quadratic computational cost with respect to the input size. While self-attention enables each pixel to adaptively aggerate spatial information from the assigned window, it lacks the inductive biases possessed by convolutions (e.g. locality, translation equivalence, etc.), which is useful in modeling local structures of images. Some work liu2022convnext ; han2021connection ; replknet also reported that using convolutions to aggregate spatial information can produce more satisfying results than self-attention.

In this paper, to tackle the challenge of effectively and efficiently aggregating neighborhood information for image restoration, we propose a kernel basis network (KBNet) with a novel kernel basis attention (KBA) module to adaptively aggregate spatial neighborhood information, which takes advantage of both CNNs and transformers. Since natural images generally share similar patterns across different spatial locations, we first introduce a set of learnable kernel bases to learn various local patterns. The kernel bases specify how the neighborhood information should be aggregated for each pixel, and different learnable kernel bases are trained to capture various image patterns. Intuitively, similar local neighborhoods would use similar kernel bases so that the bases are learned to capture more representative spatial patterns. Given the kernel bases, we adopt a separate lightweight convolution branch to predict the linear combination coefficients of kernel bases for each pixel. The learnable kernel bases are linearly fused by the coefficients to produce adaptive and diverse aggregation weights for each pixel.

Unlike the window-based self-attention in previous works liang2021swinir ; wang2021uformer ; tu2022maxim that uses dot products between pairwise positions to generate spatial aggregation weights, the KBA module generates the weights by adaptively combining the learnable kernel bases. It naturally keeps the inductive biases of convolutions while handling various spatial contexts adaptively. KBA module is also different from existing dynamic convolutions dynamicConvGao ; malleconv or kernel prediction networks mckpn ; kpn ; xia2020bpn ; Wang2019Carafe , which predict all the kernel weights directly. The aggregation weights by the KBA module are predicted by spatially adaptive fusing the shared kernel bases. Thus, our KBA module is more lightweight and easier to optimize. The ablation study validates the proposed design choice.

Based on the KBA module, we design a multi-axis feature fusion (MFF) block to extract and fuse diverse features for image restoration. We combine operators of spatial and channel dimensions to better capture image context. In the MFF block, three branches, including channel attention, spatial-invariant, and pixel-wise adaptive feature extractions are parallelly performed and then fused by point-wise production.

By integrating the proposed KBA module and MFF block into the widely used U-Net architectures with two variants of feed-forward networks, our proposed KBNet achieves state-of-the-art performances on more than ten benchmarks of image denoising, deraining, and deblurring.

The main contributions of this work are summarized as follows:

We propose a novel kernel basis attention (KBA) module to effectively aggregate the spatial information via a series of learnable kernel bases and linearly fusing the kernel bases.

We propose an effective multi-axis feature fusion (MFF) block, which extracts and fuses channel-wise, spatial-invariant, and pixel-wise adaptive features for image restoration.

Our method kernel basis network (KBNet) demonstrates its generalizability and state-of-the-art performances on three image restoration tasks including denoising, deblurring, and deraining.

Related Work

Deep image restoration has been a popular topic in the field of computer vision and has been extensively studied for many years. In this section, we review some of the most relevant works.

Since image restoration is a highly ill-posed problem, many priors or noise models zhang2021rethinkingNoise are adopted to help image restoration. Many properties of natural images have been discussed in traditional methods, like sparsity mairal2007sparse ; SparseDictionary , non-local self-similarity BM3D ; NLM , total variation variationNoiseRemove . Self-similarity is one of the natural image properties used in image restoration. Traditional methods like non-local means NLM , BM3D BM3D leverage the self-similarity to denoise images by averaging the intensities of similar patches in the image. While traditional methods have achieved good performance, they tend to produce blurry results and fail in more challenging cases.

2 CNNs for Image Restoration

Static networks. With the popularity of deep neural networks, learning-based methods become the mainstream of image restoration zhang2020rdn ; zamir2020mirnet ; Zamir2021MPRNet ; chen2022nafnet ; cho2021rethinking_mimo ; dasongBlurRepresentaion ; GroupShift ; ntire2019_denoising ; ntire2019_superresolution ; nah2021ntire . One of the earliest works on deep image denoising using CNNs is the DnCNN DnCNN . Then, lots of CNN-based models have been proposed from different design perspectives: residual learning RIDNet ; RCAN , multi-scale features zamir2020mirnet , multi-stage design Zamir2021MPRNet ; BurstRawVS , non-local information zhang2019residual ; plotz2018N3Net . While they improve the learning capacity of CNN models, they use the normal convolutions as the basic components, which are static and spatially invariant. As a result, plain and textural areas cannot be identified and processed adaptively, which is crucial for image restoration.

Dynamic networks for Image Restoration. Another line of research has focused on leveraging dynamic networks to image restoration tasks. Kernel prediction networks kpn ; mckpn ; xia2020bpn ; malleconv (KPNs) use the kernel-based attention mechanism in a CNN architecture to aggregate spatial information adaptively. But, KPNs predict the kernels directly which have significant memory and computing requirements, as well as their difficulty to optimize.

3 Transformers for Image Restoration

Transformers have shown great progress in natural language and high-level vision tasks. For image restoration, IPT chen2021IPT is the first method to adopt a transformer (both encoder and decoder) into image restoration. But, it leads to heavy computational costs and can only perform on fixed patch size 48×4848\times 48. Most of the following works only utilize the encoder and tend to reduce the computational cost. A seris of window-based self-attention liang2021swinir ; wang2021uformer ; tu2022maxim has been proposed. Each pixel can aggregate the aggregate spatial information through the dot-productions between all pairwise positions. More recently, some work liu2022convnext ; han2021connection ; replknet ; chen2022nafnet indicate that self-attention is not necessary to achieve state-of-the-art results.

Method

In this section, we aim to develop a novel kernel basis network (KBNet) for image restoration. We first describe the kernel basis attention (KBA) module to adaptively aggregate the spatial information. Then, the multi-axis feature fusion (MFF) block is introduced to encode and fuse diverse features for image restoration. Finally, we describe the integration of the MFF block into the U-Net.

How to gather spatial information at each pixel plays an essential role in feature encoding for low-level vision tasks. Most CNN-based methods DnCNN ; cheng2021nbnet ; RIDNet utilize spatial-invariant kernels to encode spatial information of the input image, which cannot adaptively process local spatial context for each pixel. While self-attention tu2022maxim ; wang2021uformer ; liang2021swinir can process spatial context adaptively according to the attention weights from the dot products between pairwise positions, it lacks the inherent inductive biases of convolutions. To tackle the challenges, we propose the kernel basis attention (KBA) module to encode spatial information by fusing learnable kernel bases adaptively.

Here, a natural choice is to normalize the fusion coefficients at each location by the softmax function so that the fusion coefficients sum up to 1. However, we experimentally find that using softmax normalization hinders the final performance since it tends to select the most important kernel basis instead of fusing multiple kernel bases for capturing spatial contexts.

Discussion. Previous kernel prediction methods kpn ; mckpn ; malleconv ; xia2020bpn also predict pixel-wise kernels for convolution. But they adopt a heavy design to predict all kernel weights directly. Specifically, even predicting K×KK\times K depthwise kernels requires producing a (C×K×K)(C\times K\times K) channel feature map, which is quite costly in terms of both computational cost and memory. In contrast, Our method trains a series of learnable kernel bases, which only needs to predict an NN-channel fusing coefficient map (where N≪C×K2N\ll C\times K^{2}). Such a design avoids predicting a large number of kernel weights for each pixel and the representative kernel bases shared across all locations are still efficiently optimized. Compared with window-based self-attention wang2021uformer ; tu2022maxim , our spatial aggregation weights are linearly combined from the shared kernel bases instead of being produced individually through dot products between pairwise positions. The KBA module adopts a set of learnable convolution kernels for modeling different local structures and fuses those kernel weights adaptively for each location. Thus, it benefits from the inductive bias of convolutions while achieving spatially-adaptive processing effectively.

2 Multi-axis Feature Fusion Block

To encode diverse features for image restoration, based on the KBA module, we design a Multi-axis Features Fusion (MFF) block to handle channel and spatial information. As shown in Fig. 2, MFF block first adopts a layer normalization to stabilize the training and then performs spatial information aggregation. A residual shortcut is used to facilitate training convergence. Following the normalization layer, three operators are adopted in parallel. The first operator is a 3×33\times 3 depthwise convolution to capture spatially-invariant features. The second operator is channel attention hu2019squeeze to modulate the feature channels. The third one is our KBA module to adaptively handle spatial features. The three branches output feature maps of the same size. Point-wise multiplication is used to fuse the diverse feature from the three branches directly, which also serves as the non-linear activation chen2022nafnet of the MFF block.

3 Intergration of MFF Block into U-Net

Results

In this section, we first describe the implementation details of our methods. Then, we evaluate our KBNet on popular benchmarks over synthetic denoising, real-world denoising, deraining and deblurring tasks. Finally, we conduct ablation studies on Gaussian denoising to validate important designs of our method and compare KBNet with existing methods.

The default setting of our method is introduced as follows unless otherwise specified. The block numbers for each stage in the encoder and decoder of our KBNet are {2,2,2,2}\{2,2,2,2\} and {4,2,2,2}\{4,2,2,2\}, respectively. For the KBA module, The number of kernel basis is set to 32 by default. For each kernel base, we use a grouped convolution kernel with kernel size 3×33\times 3 and 4 channels for each group. We train our model for 300k iterations for each noise level. The patch size is 256×256256\times 256, batch size is 32, and learning rate is 10−310^{-3} following the training settings of NAFNet chen2022nafnet .

2 Gaussian Denoising Results

The results of Gaussian denoising on color images are shown in Tab. 1. Our method outperforms the previous state-of-the-art Restormer, but only requires half of its computational cost. It is also worth noticing that we do not use progressive patch sampling to improve the training performance as Restormer restormer . Compared with the window-based transformer SwinIR liang2021swinir , we train KBNet for 300k iterations, which is much fewer than 1.6M iterations used in SwinIR liang2021swinir . The performance-efficiency comparisons with previous methods are shown in Fig. 3. Our KBNet achieves state-of-the-art results while requiring half of the computational cost. Some visual results can be found in Fig. 4. Thanks to the pixel-wise adaptive aggregation, KBNet can recover more textures even for some very thin edges.

For Gaussian denoising of gray images, as shown in Tab. 2, KBNet shows consistent performance as the color image denoising. It outperforms the previous state-of-the-art method Restormer restormer slightly, but only uses less than half of its MACs.

3 Raw Image Denoising Results

SIDD: The SIDD dataset is collected on indoor scenes. Five smartphones are used to capture scenes at different noise levels. SIDD contains 320 image pairs for training and 1,2801,280 for validation. As shown in Tab. 3, KBNet achieves state-of-the-art results on SIDD dataset. It outperforms very recent transformer-based methods including Restormer restormer , Uformer wang2021uformer , MAXIM tu2022maxim and CNN-based methods NAFNet chen2022nafnet with fewer MACs. Fig. 6 shows the performance-efficiency comparisons of our method. KBNet achieves the best trade-offs.

SenseNoise: SenseNoise dataset contains 500 diverse scenes, where each image is of high resolution (e.g. 4000 ×\times 3000). It contains both indoor and outdoor scenes with high-quality ground truth. We train the existing methods on the SenseNoise dataset zhang2021IDR under the same training setting of NAFNet chen2022nafnet but use 100k iterations. Since some of the models are too heavy, we scale down their channel numbers for fast training. The performance and MACs are reported in Tab. 4. Our method not only outperforms all other methods but also achieves the best performance-efficiency trade-offs as shown in Fig. 7. Some visualizations are shown in Fig. 5. KBNet produces sharper edges and recovers more vivid colors than previous methods.

4 Deraining and Defocus results

To demonstrate the generalization and effectiveness of our KBNet, we follow the state-of-the-art image restoration method Restormer restormer to conduct experiments on deraining and defocus deblurring. The channels of our MFF module are adjusted to make sure our model uses fewer MACs than Restormer restormer . The training settings are kept the same as the Restormer restormer .

5 Ablation Studies

We conduct extensive ablation studies to validate the effectiveness of components of our method and compare it with existing methods. All ablation studies are conducted on Gaussian denoising with noise level σ=50\sigma=50. We train models for 100k iterations. Other training settings are kept the same as the main experiment on the Gaussian denoising of color images.

Comparison with dynamic spatial aggregation methods. We compare our method with existing dynamic aggregation solutions, including the window-based self-attention wang2021uformer , dynamic convolution AttOverConv20 , and kernel prediction module kpn ; Wang2019Carafe to replace our KBA module in the proposed MFF blocks. As shown in Tab. 7, directly adopting the kernel prediction module in kpn requires heavy computational cost. While the dynamic convolution AttOverConv20 predicts spatial-invariant dynamic kernels, it slightly outperforms the kernel prediction module. This demonstrates that the kernel prediction module is difficult to optimize as it requires predicting a large number of kernel weights directly. Most existing works kpn ; mckpn ; xia2020bpn need to adopt a heavy computational branch to realize the kernel prediction. The shifted window attention requires more MACs while only improving the performance marginally than the dynamic convolution. Our method is more lightweight and improves performance significantly.

Using softmax in KBA module. To linearly combine the kernel bases, a natural choice is to use the softmax function to normalize the kernel fusion coefficients. However, we find that using the softmax would hinder the performance as shown in Tab. 7. Using the softmax function may encourage the fused kernel more focus on a specific kernel basis and reduces the diversity of the fused kernel weights.

The impact of the number of kernel bases. We also validate the influence of the number of kernel bases. As shown in Tab. 8, more kernel bases bring consistent performance improvements since it captures more image patterns and increases the diversity of the spatial information aggregation. In our experiments, we select N=32N=32 for a better performance-efficiency trade-off.

The impact of different branches in MFF block. As shown in Tab. 9, a single 3×33\times 3 depthwise convolution branch produces 29.12 dB on Gaussian denoising. Adding the channel attention branch and KBA module branch successively leads to 0.1dB and 0.25dB respectively. KBA module brings the largest improvement, which indicates the importance of the proposed pixel adaptive spatial information processing.

Visulization of kernel indices. We visualize the kernel indices of different regions in different stages of our KBNet. We project the fusion coefficients to a 3D space by random projection and visualize them as an RGB map. As shown in Fig. 10, similar color indicates that the pixels fuse kernel bases having similar patterns. KBNet can identify different regions and share similar kernel bases for similar textural or plain areas. Different kernels are learned to be responsible for different regions, and they can be optimized jointly during the training.

Conclusion

In this paper, we introduce a kernel basis network (KBNet) for image restoration. The key designs of our KBNet are kernel basis attention (KBA) module and Multi-axis Feature Fusion (MFF) block. KBA module adopts the learnable kernel bases to model the local image patterns and fuse kernel bases linearly to aggregate the spatial information efficiently. Besides, the MFF block aims to fuse diverse features from multiple axes for image denoising, which includes channel-wise, spatial-invariant, and pixel-adaptive processing. In the experiments, KBNet achieves state-of-the-art results on popular synthetic noise datasets and two real-world noise datasets (SIDD and SenseNoise) with less computational cost. It also presents good generalizations and state-of-the-art results on deraining and defocus deblurring.

References