Uformer: A General U-Shaped Transformer for Image Restoration

Zhendong Wang, Xiaodong Cun, Jianmin Bao, Wengang Zhou, Jianzhuang Liu, Houqiang Li

Introduction

With the rapid development of consumer and industry cameras and smartphones, the requirements of removing undesired degradation (e.g., noise, blur, rain, and so on) in images are constantly growing. Recovering genuine images from their degraded versions, i.e., image restoration, is a classic task in computer vision. Recent state-of-the-art methods are mostly ConvNets-based, which achieve impressive results but show a limitation in capturing long-range dependencies. To address this problem, several recent works start to employ single or few self-attention layers in low resolution feature maps due to the self-attention computational complexity being quadratic to the feature map size.

In this paper, we aim to leverage the capability of self-attention in feature maps at multi-scale resolutions to recover more image details. To this end, we present Uformer, an effective and efficient Transformer-based structure for image restoration. Uformer is built upon an elegant architecture UNet , wher we modify the convolution layers to Transformer blocks while keeping the same overall hierarchical encoder-decoder structure and the skip-connections.

We propose two core designs to make Uformer suitable for image restoration tasks. First, we propose the Locally-enhanced Window (LeWin) Transformer block, which is an efficient and effective basic component. The LeWin Transformer block performs non-overlapping window-based self-attention instead of global self-attention, which significantly reduces the computational complexity on high resolution feature maps. Since we build hierarchical feature maps and keep the window size unchanged, the window-based self-attention at low resolution is able to capture much more global dependencies. On the other hand, local context is essential for image restoration, we further introduce a depth-wise convolutional layer between two fully-connected layers of the feed-forward network in the Transformer block to better capture local context. We also notice that recent works use the similar design for different tasks.

Second, we propose a learnable multi-scale restoration modulator to handle various image degradations. The modulator is formulated as a multi-scale spatial bias to adjust features in multiple layers of the Uformer decoder. Specifically, a learnable window-based tensor is added to features in each LeWin Transformer block to adapt the features for restoring more details. Benefiting from the simple operator and window-based mechanism, it can be flexibly applied for various image restoration tasks in different frameworks.

Based on the above two designs, without bells and whistles, e.g., the multi-stage or multi-scale framework and the advanced loss function , our simple U-shaped Transformer structure achieves state-of-the-art performance on multiple image restoration tasks. For denoising, Uformer outperforms the previous state-of-the-art method (NBNet ) by 0.14 dB and 0.09 dB on the SIDD and DND benchmarks, respectively. For the motion blur removal task, Uformer achieves the best (GoPro , RealBlur-R , and RealBlur-J ) or competitive (HIDE ) performance, displaying its strong capability of deblurring. Uformer also shows the potential on the defocus deblurring task and outperforms the previous best model by 1.04 dB. Also, on the SPAD dataset for deraining, it obtains 47.84 dB on PSNR, an improvement of 3.74 dB over the previous state-of-the-art method . We expect our work will encourage further research to explore Transformer-based architectures for image restoration.

Overall, we summarize the contributions of this paper as follows:

We present Uformer, a general and superior U-shaped Transformer for various image restoration tasks. Uformer is built on the basic LeWin Transformer block that is both efficient and effective.

We present an extra light-weight learnable multi-scale restoration modulator to adjust on multi-scale features. This simple design significantly improves the restoration quality.

Extensive experiments show that Uformer establishes new state-of-the-arts on various datasets for image restoration tasks.

Related Work

Image Restoration Architectures Image restoration aims to restore the clean image from its degraded version. A popular solution is to learn effective models using the U-shaped structures with skip-connection to capture multi-scale information hierarchically for various image restoration tasks, including image denoising , deblurring , and demoireing . Some image restoration methods are inspired by the key insight from the rapid development of image classification . For example, ResNet-based structure has been widely used for general image restoration as well as for specific tasks in image restoration such as super-resolution and image denoising . More CNN-based image restoration architectures can be found in the recent surveys and the NTIRE Challenges .

Until recently, some works start to explore the attention mechanism to boost the performance. For example, squeeze-and-excitation networks and non-local neural networks inspire a branch of methods for different image restoration tasks, such as super-resolution , deraining , and denoising . Our Uformer also applies the hierarchical structure to build multi-scale features while using the newly introduced LeWin Transformer block as the basic building block.

Vision Transformers Transformer shows a significant performance in natural language processing (NLP). Different from the design of CNNs, Transformer-based network structures are naturally good at capturing long-range dependencies in the data by the global self-attention. The success of Transformer in the NLP domain also inspires the computer vision researchers. The pioneering work of ViT directly trains a pure Transformer-based architecture on the medium-size (16×\times16) flattened patches. With large-scale data pre-training (i.e., JFT-300M), ViT gets excellent results compared to state-of-the-art CNNs on image classification.

Since the introduction of ViT, many efforts have been made to reduce the quadratic computational cost of global self-attention for making Transformer more suitable for vision tasks. Some works focus on establishing a pyramid Transformer architecture simlilar to ConvNet-based structure. To overcome the quadratic complexity of original self-attention, self-attention is performed on local windows with the halo operation or window shift to help cross-window interaction, and get promising results. Rather than focusing on image classification, recent works propose a brunch of Transformer-based backbones for more general high-level vision tasks.

Besides high-level discriminative tasks, there are also some Transformer-based works for generative tasks. While there are a lot of explorations in the vision area, introducing Transformer to low-level vision still lacks exploration. Early work makes use of self-attention mechanism to learn texture for super-resolution. As for image restoration tasks, IPT first applies standard Transformer blocks within a multi-task learning framework. However, IPT relies on pretraining on a large-scale synthesized dataset and multi-task learning for good performance. In contrast, we design a general U-shaped Transformer-based structure, which proves to be efficient and effective for image restoration.

Method

In this section, we first describe the overall pipeline and the hierarchical structure of Uformer for image restoration. Then, we provide the details of the LeWin Transformer block which is the basic component of Uformer. After that, we present the multi-scale restoration modulator.

Then, a bottleneck stage with a stack of LeWin Transformer blocks is added at the end of the encoder. In this stage, thanks to the hierarchical structure, the Transformer blocks capture longer (even global when the window size equals the feature map size) dependencies.

where I^\mathbf{\hat{I}} is the ground-truth image, and ϵ=10−3\epsilon=10^{-3} is a constant in all the experiments.

2 LeWin Transformer Block

There are two main challenges to apply Transformer for image restoration. First, the standard Transformer architecture computes self-attention globally between all tokens, which contributes to the quadratic computation cost with respect to the number of tokens. It is unsuitable to apply global self-attention on high-resolution feature maps. Second, the local context information is essential for image restoration tasks since the neighborhood of a degraded pixel can be leveraged to restore its clean version, but previous works suggest that Transformer shows a limitation in capturing local dependencies.

To address the above mentioned two issues, we propose a Locally-enhanced Window (LeWin) Transformer block, as shown in Figure 2(b), which benefits from the self-attention in Transformer to capture long-range dependencies, and also involves the convolution operator into Transformer to capture useful local context. Specifically, given the features at the (l-1l\text{-}1)-th block Xl−1\mathbf{X}_{l-1}, we build the block with two core designs: (1) non-overlapping Window-based Multi-head Self-Attention (W-MSA) and (2) Locally-enhanced Feed-Forward Network (LeFF). The computation of a LeWin Transformer block is represented as:

where Xl′\mathbf{X}_{l}^{\prime} and Xl\mathbf{X}_{l} are the outputs of the W-MSA module and LeFF module, respectively. LN represents the layer normalization . In the following, we elaborate W-MSA and LeFF separately.

Suppose the head number is kk and the head dimension is dk=C/kd_{k}=C/k. Then computing the kk-th head self-attention in the non-overlapping windows can be formulated as follows,

Locally-enhanced Feed-Forward Network (LeFF). As pointed out by previous works , the Feed-Forward Network (FFN) in the standard Transformer suffers limited capability to leverage local context. Actually, neighboring pixels are crucial references for image restoration . To overcome this issue, we add a depth-wise convolutional block to the FFN in our Transformer-based structure following the recent works . As shown in Figure 3, we first apply a linear projection layer to each token to increase its feature dimension. Next, we reshape the tokens to 2D feature maps, and use a 3×33\times 3 depth-wise convolution to capture local information. Then we flatten back the features to tokens and shrink the channels via another linear layer to match the dimension of the input channels. We use GELU as the activation function after each linear/convolution layer.

3 Multi-Scale Restoration Modulator

Different types of image degradation (e.g. blur, noise, rain, etc.) have their own distinctive perturbed patterns to be handled or restored. To further boost the capability of Uformer for approaching various perturbations, we propose a light-weight multi-scale restoration modulator to calibrate the features and encourage more details recovered.

As shown in Figure 2(a) and 2(c), the multi-scale restoration modulator applies multiple modulators in the Uformer decoder. Specially in each LeWin Transformer block, a modulator is formulated as a learnable tensor with a shape of M×M×CM\times M\times C, where MM is the window size and CC is the channel dimension of current feature map. Each modulator is simply served as a shared bias term that is added into all non-overlapping windows before self-attention module. Due to this light-weight addition operation and window-sized shape, the multi-scale restoration modulator introduces marginal extra parameters and computational cost.

We prove the effectiveness of the multi-scale restoration modulator on two typical image restoration tasks: image deblurring and image denoising. The visualization comparisons are presented in Figure 4. We observe that adding the multi-scale restoration modulator makes more motion blur/noising patterns removed and yields a much cleaner image. These results show that our multi-scale restoration modulator truly helps to recover restoration details with little computation cost. One possible explanation is that adding modulators at each stage of the decoder enables a flexible adjustment of the feature maps that boosts the performance for restoring details. This is consistent with the previous work StyleGAN using a multi-scale noise term adding to the convolution features, which realizes stochastic variation for generating photo-realistic images.

Experiments

In this section, we first discuss the experimental setup. After that, we verify the effectiveness and efficiency of Uformer on various image restoration tasks on eight datasets. Finally, we perform comprehensive ablation studies to evaluate each component of our proposed Uformer.

Basic settings. Following the common training strategy of Transformer , we train our framework using the AdamW optimizer with the momentum terms of (0.9,0.999)(0.9,0.999) and the weight decay of 0.02. We randomly augment the training samples using the horizontal flipping and rotate the images by 90∘90^{\circ}, 180∘180^{\circ}, or 270∘270^{\circ}. We use the cosine decay strategy to decrease the learning rate to 1e-61e\text{-}6 with the initial learning rate 2e-42e\text{-}4. We set the window size to 8×\times8 in all LeWin Transformer blocks. The number of Uformer encoder/decoder stages KK equals 4 by default. And the dimension of each head in Transformer block dkd_{k} equals CC. More dataset-specific experimental settings can be found in the supplementary materials.

Architecture variants. For a concise description, we introduce three Uformer variants in our experiments, Uformer-T (Tiny), Uformer-S (Small), and Uformer-B (Base) by setting different Transformer feature channels CC and the numbers of the Transformer blocks in each encoder and decoder stages. The details are listed as follows:

Uformer-T: C=16C=16, depths of Encoder = {2, 2, 2, 2},

Uformer-S: C=32C=32, depths of Encoder = {2, 2, 2, 2},

Uformer-B: C=32C=32, depths of Encoder = {1, 2, 8, 8},

and the depths of Decoder are mirrored depths of Encoder.

Evaluation metrics. We adopt the commonly-used PSNR and SSIM metrics to evaluate the restoration performance. These metrics are calculated in the RGB color space except for deraining where we evaluate the PSNR and SSIM on the Y channel in the YCbCr color space, following the previous work .

2 Real Noise Removal

Table 1 reports the results of real noise removal on the SIDD and DND datasets. We compare Uformer with 8 state-of-the-art denoising methods, including the feature-based BM3D and seven learning-based methods: RIDNet , VDN , CycleISP , NBNet , DANet , MIRNet , and MPRNet . Our Uformer-B achieves 39.89 dB on PSNR, surpassing all the other methods by at least 0.14 dB. As for the DND dataset, we follow the common evaluation strategy and test our model trained on SIDD via the online server testing. Uformer outperforms the previous state-of-the-art method NBNet by 0.09 dB. To verify whether the gains benefit from more computation cost, we present the results of PSNR vs. computational cost in Figure 1. We notice that our Uformer-T can achieve a better performance than most models but with the least computation cost, which demonstrates the efficiency and effectiveness of Uformer. We also show the qualitative results on the SIDD and DND datasets in Figure 5, in which Uformer can not only successfully remove the noise but also keep the texture details.

3 Motion Blur Removal

For motion blur removal, Uformer also shows state-of-the-art performance. We follow the previous method to train Uformer on the GoPro dataset and test it on the four datasets: two synthesized datasets ( HIDE and the test set of GoPro ), and two real-world datasets (RealBlur-R/-J from the RealBlur dataset ). We compare Uformer with ten state-of-the-art methods: Nah et al. , DeblurGAN , Xu et al. , DeblurGAN-v2 , DBGAN , SPAIR , Zhang et al. , SRN , DMPHN , and MPRNet . The results are reported in Table 2. For synthetic deblurring, Uformer gets significant better performance on GoPro than previous state-of-the-art methods and shows a comparable result on the HIDE dataset. As for real-world deblurring, the causes of blur are complicated so the task is usually more challenging. Our Uformer outperforms other methods by at least 0.23 dB and 0.36 dB on RealBlur-R and RealBlur-J, respectively, showing a strong generalization ability. Besides, we show some visual results in Figure 6. Compared with other methods, the images restored by Uformer are more clear and closer to their ground truth.

4 Defocus Blur Removal

We perform defocus blur removal on the DPD dataset . Table 3 and Figure 7 report the quantitative and qualitative results, respectively. Uformer achieves a better performance (1.04 dB, 1.15 dB, 1.44 dB, and 1.87 dB) over previous state-of-the-art methods KPAC , DPDNet , JNB , and DMENet , respectively. From the visualization results, we observe that the images recovered by Uformer are sharper and closer to the ground-truth images.

5 Real Rain Removal

We conduct the deraining experiments on SPAD and compare with 6 deraining methods: GMM , RESCAN , SPANet , JORDER-E , RCDNet , and SPAIR . As shown in Table 4, Uformer presents a significantly better performance, achieving 3.74 dB improvement over the previous best work . This indicates the strong capability of Uformer for deraining on this real derain dataset. We also provide the visual results in Figure 7 where Uformer can remove the rain more successfully while introducing fewer artifacts.

6 Ablation Study

In this section, we analyze the effect of each component of Uformer in detail. The evaluations are conducted on image denoising (SIDD ), deblurring (GoPro , RealBlur ), and deraining (SPAD ) using different variants. The ablation results are reported in Tables 5, 6, and 7.

Transformer vs. convolution. We replace all the LeWin Transformer blocks in Uformer with the convolution-based ResBlocks , resulting in the so-called "UNet", while keeping all others unchanged. Similar to the Uformer variants, we design UNet-T/-S/-B:

UNet-T: C=32C=32, depths of Encoder = {2, 2, 2, 2},

UNet-S: C=48C=48, depths of Encoder = {2, 2, 2, 2},

UNet-B: C=76C=76, depths of Encoder = {2, 2, 2, 2},

and the depths of Decoder are mirrored depths of Encoder.

Table 5 reports the comparison results. We observe that Uformer-T achieves 39.66 dB and outperforms UNet-T by 0.04 dB with fewer parameters and less computation. Uformer-S achieves 39.77 dB and outperforms UNet-S by 0.12 dB with fewer parameters and a slightly higher computation cost. And Uformer-B achieves 39.89 dB which outperforms UNet-B by 0.18 dB. This study indicates the effectiveness of the proposed LeWin Transformer block, compared with the original convolutional block.

Hierarchical structure vs. single scale. We further build a ViT-based architecture which only contains a single scale of the feature maps for image denoising. This architecture employs a head of two convolution layers for extracting features from the input image and also a tail of two convolution layers for the output. 12 standard Transformer blocks are used between the head and the tail. We train the ViT with the hidden dimension of 256 on patch size 16×1616\times 16. The results are presented in Table 5. We observe that the vanilla ViT structure gets an unsatisfactory result compared with UNet, while our Uformer significantly outperforms both the ViT-based and UNet architectures, which demonstrates the effectiveness of hierarchical structure for image restoration.

Where to enhance locality? Table 6 compares the results of no locality enhancement and enhancing locality in the self-attention calculation or the feed-forward network based on Uformer-S and Uformer-B. We observe that introducing locality into the feed-forward network yields 0.03 dB (SIDD), 0.07 dB (RealBlur-R)/0.07 dB (RealBlur-J) over the baseline (no locality enhancement), while introducing locality into the self-attention yields -0.02 dB (SIDD). Further, we combine introducing locality into the feed-forward network and introducing into the self-attention. The results on RealBlur-R/-J also drop from 36.22 dB/29.06 dB to 36.19 dB/28.85 dB, indicating that compared to involving locality into self-attention, introducing locality into the feed-forward network is more suitable for image restoration tasks.

Effect of the multi-scale restoration modulator. In Table 7, to verify the effect of the modulator, we conduct experiments on GoPro for image deblurring, SIDD for image denoising, and SPAD for deraining. For deblurring, we observe that w/ modulator can bring a performance improvement of 0.46 dB, which reveals the effectiveness of the modulator for deblurring. We also compare the results of Uformer-B with/without the modulator on SIDD and SPAD, and the comparisons indicate that the proposed modulator introduces 0.03 dB improvement (SIDD)/0.41 dB improvement (SPAD). In Figure 4, we have provided visual comparisons of Uformer w/ and wo/ the modulator. This study validates the proposed modulator can bring extra ability of restoring more details.

Discussion and Conclusion

In this paper, we have presented an alternative architecture Uformer for image restoration tasks by introducing the Transformer block. In contrast to existing ConvNet-based structures, our Uformer builds upon the main component LeWin Transformer block, which can not only handle local context but also capture long-range dependencies efficiently. To handle various image restoration degradation and enhance restoration quality, we propose a learnable multi-scale restoration modulator inserted into the Uformer decoder. Extensive experiments demonstrate that Uformer achieves state-of-the art performance on several tasks, including denoising, motion deblurring, defocus deblurring, and deraining. Uformer also surpasses the UNet family by a large margin with less computation cost and fewer model parameters.

Limitation and broader impacts. Thanks to the proposed architecture, Uformer achieves the state-of-the-art performance on a variety of image restoration tasks (image denoising, deblurring, and deraining). But we have not evaluated Uformer for more vision tasks such as image-to-image translation, image super-resolution, and so on. We look forward to investigating Uformer for more applications. Meanwhile, we notice that there are several negative impacts caused by abusing image restoration techniques. For example, it may cause human privacy issue with the restored images in surveillance. The techniques may destroy the original patterns for camera identification and multi-media copyright , which hurts the authenticity for image forensics.

References

Appendix A Additional Ablation Study

Table 8 reports the results of whether to use the shifted window design in Uformer. We observe that window shift brings an improvement of 0.01 dB for image denoising. We use the window shift as the default setting in our experiments.

A.2 Variants of Skip-Connections

To investigate how to deliver the learned low-level features from the encoder to the decoder, considering the self-attention computing in Transformer, we present three different skip-connection schemes, including concatenation-based skip-connection, cross-attention as skip-connection, and concatenation-based cross-attention as skip-connection.

Concatenation-based Skip-connection (Concat-Skip). Concat-Skip is based on the widely-used skip-connection in UNet . To build our network, firstly, we concatenate the ll-th stage flattened features El\mathbf{E}_{l} and each encoder stage with the features DK−l+1\mathbf{D}_{K-l+1} from the (K-l+1)(K\text{-}l\text{+}1)-th decoder stage channel-wisely. Here, KK is the number of the encoder/decoder stages. Then, we feed the concatenated features to the W-MSA component of the first LeWin Transformer block in the decoder stage, as shown in Figure 8(a).

Cross-attention as Skip-connection (Cross-Skip). Instead of directly concatenating features from the encoder and the decoder, we design Cross-Skip inspired by the decoder structure in the language Transformer . As shown in Figure 8(b), we first add an additional attention module into the first LeWin Transformer block in each decoder stage. The first self-attention module in this block (the shaded one) is used to seek the self-similarity pixel-wisely from the decoder features DK−l+1\mathbf{D}_{K-l+1}, and the second attention module in this block takes the features El\mathbf{E}_{l} from the encoder as the keys and values, and uses the features from the first module as the queries.

Concatenation-based Cross-attention as Skip-connection (ConcatCross-Skip). Combining above two variants, we also design another skip-connection. As illustrated in Figure 8(c), we concatenate the features El\mathbf{E}_{l} from the encoder and DK−l+1\mathbf{D}_{K-l+1} from the decoder as the keys and values, while the queries are only from the decoder.

Table 9 compares the results of using different skip-connections in our Uformer: concatenating features (Concat), cross-attention (Cross), and concatenating keys and values for cross-attention (ConcatCross). For a fair comparison, we increase the channels in Uformer-S from 32 to 44 in variants Cross and ConcatCross. These three skip-connections achieve similar results, and concatenating features gets slightly better performance. We adopt the feature concatenation as the default setting in Uformer.

Appendix B Additional Experiment for Demoireing

We also conduct an experiment of moire pattern removal on the TIP18 dataset . As shown in Table 10, Uformer outperforms previous methods MopNet , MSNet , CFNet , UNet by 1.53 dB, 2.29 dB, 3.19 dB, and 2.79 dB, respectively. And in Figure 13, we show examples of visual comparisons with other methods. This experiment further demonstrates the superiority of Uformer.

Appendix C Additional Experimental Settings for Different Tasks

Denoising. The training samples are randomly cropped from the original images in SIDD with size 128×128128\times 128, which is also the common training strategy for image denoising in recent works . And the training process lasts for 250 epochs with batch size 32. Then, the trained model is evaluated on the 256×256256\times 256 patches of SIDD and 512×512512\times 512 patches of the DND test images , following . The results on DND are online evaluated.

Motion deblurring. Following previous methods , we train Uformer only on the GoPro dataset , and evaluate it on the test set of GoPro, HIDE , and RealBlur-R/-J . The training patches are randomly cropped from the training set with size 256×256256\times 256. The batch size is set to 32. For validation, we use the central crop with size 256×256256\times 256. The number of training epochs is 3k. For evaluation, the trained model is tested on the full-size test images.

Defocus deblurring. Following the official patch segmentation algorithm of DPD, we crop the training and validation samples to 60% overlapping 512×512512\times 512 patches to train the model. We also discard 30% of the patches that have the lowest sharpness energy (by applying Sobel filter to the patches) as . The whole training process lasts for 160 epochs with batch size 4. For evaluation, the trained model is tested on the full-size test images.

Deraining. We conduct deraining experiments on the SPAD dataset . This dataset contains over 64k 256×256256\times 256 images for training and 1k 512×512512\times 512 images for evaluation. We train Uformer on two GPUs, with mini-batches of size 16 on the 256×256256\times 256 samples. Since this dataset is large enough and the training process converges fast, we just train Uformer for 10 epochs in the experiment. Finally, we evaluate the performance on the test images following the default setting in .

Demoireing. We further validate the effectiveness of Uformer on the TIP18 dataset for demoireing. Since the images in this dataset contain additional borders, following , we crop the central regions with the ratio of [0.15,0.85][0.15,0.85] in all training/validation/testing splits and resize them to 256×256256\times 256 for training and evaluation. Since this task is sensitive to the down-sampling operation, we choose the bilinear interpolation same as the previous work The dataset we used is also downloaded from the Github Page of .. The training epochs are 250.

Appendix D More Visual Comparisons

As shown in Figures 9-13 in this supplementary materials, we give more visual results of our Uformer and others on the five tasks (denoising, motion deblurring, defocus deblurring, deraining, and demoireing) as the supplement of the visualization in the main paper.