MambaIR: A Simple Baseline for Image Restoration with State-Space Model
Hang Guo, Jinmin Li, Tao Dai, Zhihao Ouyang, Xudong Ren, Shu-Tao Xia
Introduction
Image restoration, aiming to reconstruct a high-quality image from a given low-quality input, is a long-standing problem in computer vision and further has a wide range of sub-problems such as super-resolution, image denoising, etc. Since several pioneering works , Convolutional Neural Networks (CNNs) have gained great success than traditional model-based method , and almost dominate this field in the past few years .
Despite the promising results, these restoration backbones usually suffer from the inherent problem of CNNs. First, the weight-sharing strategy in CNN makes the model remain static for varying inputs, preventing robustness on unseen real-world low-quality images. Second, the receptive field of one convolution layer is constrained by the kernel size, which leads to strong local reductive bias. Although stacking multiple convolutional layers can theoretically enlarge the receptive field, however as shown in Fig. 1(a) the effective receptive field of CNN-based restoration methods is still limited.
Inspired by the success in natural language processing, Transformer has proven superior performance than CNNs in several vision tasks due to weak inductive bias as well as global interaction modeling . However, taming Transformer for image restoration poses a significant challenge since the self-attention mechanism scales quadratically with the number of tokens. Considering that image restoration methods usually set the number of tokens to the image resolution, therefore directly using the attention mechanism will incur a huge computational overhead. Although some works make efforts to divide input into small patches or employ shifted window attention , these strategies harm the long-term dependency as shown in Fig. 1(b), and do not essentially address the quadratic complexity problem.
Recently, structured state-space sequence models (S4), have emerged as an efficient and effective backbone for constructing deep networks . In particular, the improved S4 with selective mechanism and efficient hardware design, also called Mamba, has been proven to surpass Transformer on tasks requiring long-term dependency modeling such as natural language processing . More recently, some variants of Mamba have also presented promising results in computer vision , such as image classification, biomedical image segmentation, etc. However, the power of Mamba in low-level vision is still under-explored, which motivates us to explore the potential of Mamba to enhance the long-range modeling ability for image restoration networks.
In this work, we propose a simple but effective benchmark restoration model, namely MambaIR, to adapt Mamba for image restoration. Following , our MambaIR consists of three stages: shallow feature extraction, deep feature extraction, and high-quality image reconstruction. The shallow feature extraction employs a simple convolution layer to extract the shallow feature. Then the deep feature extraction is performed with several stacked Residual State Space Blocks (RSSBs). As the core component of MambaIR, the RSSB is designed to enhance the capabilities of the original Mamba with image restoration priors. Specifically, to tackle the problem that the original Mamba treats each token equally, ignoring the patch recurrence occurring in the neighborhood, we add additional convolution layers to give local pixels more focus. We further introduce channel attention layer to supplement channel interaction which is beneficial for image restoration but missed in Mamba. In addition, we use the skip connection within each block to offer shortcuts for feature aggregation. Finally, the high-quality image reconstruction stage fuses both shallow and deep features to restore the high-quality output image.
Our main contributions can be summarized as follows:
We are the first to work to adapt state space models for low-level image restoration via extensive experiments to formulate MambaIR, which acts as a simple but effective alternative for CNNs and Transformers.
We propose the Residue State Space Block (RSSB) which can boost the power of the original Mamba with local spatial prior and channel interaction.
Extensive experiments on various tasks demonstrate our MambaIR outperform strong Transformer-based baselines to provide a strong and promising backbone model solution for image restoration.
Related Work
Image restoration has been significantly advanced since the introduction of deep learning by several pioneering works, such as SRCNN for image super-resolution, DnCNN for image denoising and ARCNN for JPEG compression artifact reduction, etc. Early attempts usually elaborate CNNs with techniques such as residual connection , dense connection and others to improve model representation ability. Despite the success, CNN-based restoration methods typically face challenges in effectively modeling global dependencies as well as static weight for varying inputs. As Transformers have been proven to be a strong competitor for CNNs in multiple computer vision tasks , using Transformer for image restoration appears promising. However, it still faces specific challenges from the quadratic computational complexity of the self-attention mechanism . To address this, IPT divides one image into several small patches and processes each patch independently with self-attention. SwinIR further introduces shifted window attention to improve the performance. In addition, promising progress continues to be made in designing efficient and effective attention mechanisms for restoration . Nonetheless, efficient attention design usually comes at the expense of global effective receptive fields, and the issue of quadratic computational complexity is not essentially addressed.
2 State Space Models
State Space Models (SSMs) , stemming from classics control theory , are recently introduced to deep learning as a competitive backbone for state space transforming. The promising property of linearly scaling with sequence length in long-range dependency modeling has attracted great interest from searchers. For example, the Structured State-Space Sequence model (S4) is a pioneer work for the deep state-space model in modeling the long-range dependency. Later, S5 layer is proposed based on S4 and introduces MIMO SSM and efficient parallel scan. Moreover, H3 achieves promising results that nearly fill the performance gap between SSMs and Transformers in natural language. further improve S4 with gating units to obtain the Gated State Space layer to boost the capability. More recently, Mamba , a data-dependent SSM with selective mechanism and efficient hardware design, outperforms Transformers on natural language and enjoys linear scaling with input length. And there are also pioneering works that adopt Mamba to vision tasks such as image classification , video understanding , biomedical image segmentation and others . In this work, we explore the potential of Mamba to image restoration with restoration-specific prior to sever as a simple but effective baseline for future work.
Methodology
This work aims to introduce the recent advanced state space model (SSM), i.e., Mamba, to low-level image restoration. In this section, we begin with a description of the preliminaries of the SSM. Then we give an overview of our simple but effective benchmark model MambaIR, followed by a detailed description of the proposed Residule State-Space Block (RSSB). Finally, we provide a comprehensive discussion of the difference between our MambaIR and previous methods.
After that, the discretization process is typically adopted to integrate Eq. 1 into practical deep learning algorithms. Specifically, denote as the timescale parameter to transform the continuous parameters A, B to discrete parameters , . The commonly used method for discretization is the zero-order hold (ZOH) rule, which is defined as follows:
After the discretization, the discretized version of Eq. 1 with step size can be rewritten as:
However, the formulation in Eq. 3 focuses on the LTI system with parameters remaining static for varying inputs. To address this limitation, recent effort attempt to incorporate the selective scan mechanism in which the matrices , C and are derived from the input data. Inspired by the favorable performance, we also employ the selective scan mechanism to facilitate the awareness of the contextual information embedded in the input.
2 Overall Architecture
3 Residual State-Space Block
The block design in previous Transformer-based image restoration networks mainly follow the Norm Attention Norm MLP flow. Although self-attention and SSM are similar in modeling long-range dependencies, however, we experimentally find that simply replacing Attention with SSM can only obtain sub-optimal results. Therefore, it is promising to tailor a brand-new block structure for Mamba-based restoration networks.
4 Vision State-Space Module
where DWConv represents depth-wise convolution, and denotes the Hadamard product.
5 2D Selective Scan Module
The standard Mamba causally processes the input data, and thus can only capture information within the scanned part of the data. This property is well suited for NLP tasks that involve a sequential nature but poses significant challenges when transferring to non-causal data such as images. To better utilize the 2D spatial information, we follow and introduce the 2D Selective Scan Module (2D-SSM). As shown in Fig. 2(c), the 2D image feature is flattened into a 1D sequence with scanning along four different directions: top-left to bottom-right, bottom-right to top-left, top-right to bottom-left, and bottom-left to top-right. Then the long-range dependency of each sequence is captured according to the discrete state-space equation in Eq. 3. Finally, all sequences are merged using summation followed by the reshape operation to recover the 2D structure.
6 Discussion
Here we would like to discuss the differences between our MambaIR with previous image restoration arts to highlight the contribution of this work. Our MambaIR improves the recent advanced state-space model Mamba with restoration-specific priors, thus allowing the injection of local patch recurrence and channel interaction while maintaining strong long-range capability. Different from Transformer-based methods, the sequential causal property of SSMs exempts us from positional encoding . Meanwhile, the linear complexity with input scale ensures the training and inference efficiency, thus avoiding the shifted window operation which may induce an unsatisfactory effective receptive field.
Experiences
Dataset and Evaluation. Following the setup in previous works , we conduct experiments on distinct image restoration tasks, including image super-resolution (i.e., classic SR and lightweight SR) and real image denoising. For classic and lightweight image super-resolution, we employ DIV2K and Flickr2K as our training datasets, and use Set5 , Set14 , B100 , Urban100 , and Manga109 to evaluate the effectiveness. For real image denoising, we train our model with 320 high-resolution images from SIDD datasets, and use the SIDD test set and DND dataset for testing. Following , we denote the model as MambaIR+ when self-ensemble strategy is used in testing. The performance is evaluated using PSNR and SSIM on the Y channel from the YCbCr color space.
Training Details. In accordance with previous work , we perform data augmentation on the training data by applying horizontal flips and random rotations of , and . Additionally, we crop the original images into patches for image SR. For the classic SR, we use the pre-trained weights from the 2 model to initialize those of 3 and 4, and halve the learning rate and total training iterations to reduce training time . To ensure a fair comparison, we adjust the training batch size to 32 for image SR. We employ the Adam as the optimizer for training our MambaIR with . The initial learning rate is set at and is halved when the training iteration reaches specific milestones. Our MambaIR model is trained with 8 NVIDIA V100 GPUs.
2 Ablation Study
Effects of different designs of RSSB. As the core component, the RSSB can improve Mamba with restoration-specific priors. In this section, we ablate different components of the RSSB. The PSNR results of classic SR, presented in Tab. 1, indicate that (1) Exploiting the local patch recurrence in the image can assist the long-range spatial modeling of Mamaba, without which will cause a performance drop. (2) Without using local convolution and channel attention, i.e., directly employing off-the-shelf Mamba for restoration, can only obtain sub-optimal results due to the domain gap. (3) Replace Conv+ChanelAttention with MLP, whose structure will be similar to Transformer, also leads to unfavorable results, indicating that although both SSMs and Attention have the global modeling ability, different assisting modules should be considered for further improvements.
3 Image Super-Resolution.
Tab. 2 shows the quantitative results between MambaIR and state-of-the-art super-resolution methods. Thanks to the significant global effective receptive filed, our proposed MambaIR achieves the best performance on almost all five benchmark datasets for all scale factors. For example, our Mamba-based baseline outperforms the Transformer-based benchmark model SwinIR by 0.41dB on Manga109 for scale, demonstrating the prospect of Mamba for image restoration. As for the computational complexity(see Fig. 3), our method is far more efficient than the full-attention baseline , and exhibits linear complexity with input resolution which is similar to the efficient attention design such as SwinIR. This observation suggests that out MambaIR has similar scale properties as shifted window attention, and at the same time possesses a global receptive field similar to standard attention. We also give visual comparisons in Fig. 4, and it can be seen that our method can facilitate the reconstruction of sharp edges and natural textures. Moreover, the LAM visualization shown in Fig. 5 demonstrates our MambaIR can utilize more pixels during restoration due to the powerful long-range modeling capability of Mamba.
3.2 Lightweight Image Super-Resolution.
To demonstrate the scalability of our method, we train the Mamba-light model and compare it with state-of-the-art lightweight image SR methods. Following previous works , we also report the number of parameters (#param) and MACs (upscaling a low-resolution image to resolution on all scales). Tab. 3 shows the results. It can be seen that our MambaIR-light outperforms SwinIR-light by up to 0.23dB PSNR on the scale Manga109 dataset with fewer parameters and MACs. The performance results demonstrate the scalability and efficiency of our method.
4 Image Denoising
We further turn to the real image denoising task to evaluate the robustness of our MambaIR when facing real-world degradation. Following , we adopt the progressive training strategy for fair comparison. The results are shown in Tab. 4. As one can see, our method achieves comparable performance with existing state-of-the-art models Restormer and outperforms other methods such as Uformer by 0.12dB PSNR on SIDD dataset. The promising results indicate the ability of our method in real image denoising.
Limitation and Future Work
While MambaIR appears as a competitive restoration backbone against mainstream Transformer-based methods, it can be further improved with restoration-specific designs. For example, in the proposed method, the local similarity and channel interaction priors are employed to obtain better feature representation. However, other restoration prior such as high-frequency enhancement can be also used. As a pure SSM-based method, our MambaIR does not include the Attention module, therefore how to appropriately combine Attention and Mamba remains an interesting research topic. Moreover, although this work has covered multiple image restoration tasks, some other tasks such as image deblurring and deraining can also be explored in the future. Finally, despite the promising results shown, we would like to point out that the Mamba-based image restoration network is still in its early stages. With the increasing interest in Mamba, it will be promising to study the state-space models for low-level vision.
Conclusion
In this work, we explore for the first time the power of the recent advanced state space model i.e., Mamba for image restoration through extensive experiments. By exploiting restoration-specific prior, we propose a brand-new backbone solution, named MambaIR, that can well capture the local patch recurrence and channel interaction while maintaining strong long-range capability and linear complexity, serving as a simple but effective Mamba-based benchmark model for image restoration. Extensive experiments on multiple restoration tasks demonstrate the effectiveness of MambaIR against existing CNN- and Tansformer-based methods.