Real Image Denoising with Feature Attention

Saeed Anwar, Nick Barnes

Introduction

Image denoising is a low-level vision task that is essential in a number of ways. First of all, during image acquisition, some noise corruption is inevitable and can downgrade the visual quality considerably; therefore, removing noise from the acquired image is a key step for many computer vision and image analysis applications . Secondly, denoising is a unique testing ground for evaluating image prior and optimization methods from a Bayesian perspective . Furthermore, many image restoration tasks can be solved in the unrolled inference through variable splitting methods by a set of denoising subtasks, which further widens the applicability of image denoising .

Generally, denoising algorithms can be categorized as model-based and learning-based. Model-based algorithms include non-local self-similarity (NSS) , sparsity , gradient methods , Markov random field models , and external denoising priors . The model-based algorithms are computationally expensive, time-consuming, unable to suppress the spatially variant noise directly and characterize complex image textures. On the other hand, discriminative learning aims to model the image prior from a set of noisy and ground-truth image sets. One technique is to learn the prior in steps in the context of truncated inference while another approach is to employ brute force learning, for example, MLP and CNN methods . CNN models improved denoising performance, due to their modeling capacity, network training, and design. However, the performance of the current learning models is limited and tailored for a specific level of noise.

A practical denoising algorithm should be efficient, flexible, perform denoising using a single model and handle both spatially variant and invariant noise when the noise standard-deviation is known or unknown. Unfortunately, the current state-of-the-art algorithms are far from achieving all of these aims. We present a CNN model which is efficient and capable of handling synthetic as well as real-noise present in images. We summarize the contributions of this work in the following paragraphs.

Present CNN based approaches for real image denoising employ two-stage models; we present the first model that provides state-of-the-art results using only one stage.

To best of our knowledge, our model is the first to incorporate feature attention in denoising.

Most current models connect the weight layers consecutively; and so increasing the depth will not help improve performance . Also, such networks can suffer from vanishing gradients . We present a modular network, where increasing the number of modules helps improve performance.

We experiment on three synthetic image datasets and four real-image noise datasets to show that our model achieves state-of-the-art results on synthetic and real images quantitatively and qualitatively.

Related Works

In this section, we present and discuss recent trends in the image denoising. Two notable denoising algorithms, NLM and BM3D , use self-similar patches. Due to their success, many variants were proposed, including SADCT , SAPCA , NLB , and INLM which seek self-similar patches in different transform domains. Dictionary-based methods enforce sparsity by employing self-similar patches and learning over-complete dictionaries from clean images. Many algorithms investigated the maximum likelihood algorithm to learn a statistical prior, e.g. the Gaussian Mixture Model of natural patches or patch groups for patch restoration. Furthermore, Levin et al. and Chatterjee et al. , motivated external denoising by showing that an image can be recovered with negligible error by selecting reference patches from a clean external database. However, all of the external algorithms are class-specific.

Recently, Schmidt et al. introduced a cascade of shrinkage fields (CSF) which integrated half-quadratic optimization and random-fields. Shrinkage aims to suppress smaller values (noise values) and learn mappings discriminatively. The CSF assumes the data fidelity term to be quadratic and that it has a discrete Fourier transform based closed-form solution.

Currently, due to the popularity of convolutional neural networks (CNNs), image denoising algorithms have achieved a performance boost. Notable denoising neural networks, DnCNN , and IrCNN predict the residue present in the image instead of the denoised image as the input to the loss function is ground truth noise as compared to the original clean image. Both networks achieved better results despite having a simple architecture where repeated blocks of convolutional, batch normalization and ReLU activations are used. Furthermore, IrCNN and DnCNN are dependent on blindly predicted noise i.e. without taking into account the underlying structures and textures of the noisy image.

Another essential image restoration framework is Trainable Nonlinear Reaction-Diffusion (TRND) which uses a field-of-experts prior into the deep neural network for a specific number of inference steps by extending the non-linear diffusion paradigm into a profoundly trainable parametrized linear filters and the influence functions. Although the results of TRND are favorable, the model requires a significant amount of data to learn the parameters and influence functions as well as overall fine-tuning, hyper-parameter determination, and stage-wise training. Similarly, non-local color net (NLNet) was motivated by non-local self-similar (NSS) priors which employ non-local self-similarity coupled with discriminative learning. NLNet improved upon the traditional methods; but, it lags in performance compared to most of the CNNs due to the adaptaton of NSS priors, as it is unable to find the analogs for all the patches in the image.

Recently, many algorithms focused on blind denoising on real-noisy images . The algorithms benefitted from the modeling capacity of CNNs and have shown the ability to learn a single-blind denoising model; however, the denoising performance is limited, and the results are not satisfactory on real photographs. Generally speaking, real-noisy image denoising is a two-step process: the first involves noise estimation while the second addresses non-blind denoising. Noise clinic (NC) estimates the noise model dependent on signal and frequency followed by denoising the image using non-local Bayes (NLB). In comparison, Zhang et al. proposed a non-blind Gaussian denoising network, termed FFDNet that can produce satisfying results on some of the real noisy images; however, it requires manual intervention to select high noise-level.

Very recently, CBDNet trains a blind denoising model for real photographs. CBDNet is composed of two subnetworks: noise estimation and non-blind denoising. CBDNet also incorporated multiple losses, is engineered to train on real-synthetic noise and real-image noise and enforces a higher noise standard deviation for low noise images. Furthermore, may require manual intervention to improve results. On the other hand, we present an end-to-end architecture that learns the noise and produces results on real noisy images without requiring separate subnets or manual intervention.

CNN Denoiser

Our model is composed of three main modules i.e. feature extraction, feature learning residual on the residual module, and reconstruction, as shown in Figure 2. Let us consider xx is a noisy input image and y^\hat{y} is the denoised output image. Our feature extraction module is composed of only one convolutional layer to extract initial features f0f_{0} from the noisy input:

where Me(⋅)M_{e}(\cdot) performs convolution on the noisy input image. Next, f0f_{0} is passed on to the feature learning residual on the residual module, termed as MflM_{fl},

where frf_{r} are the learned features and Mfl(⋅)M_{fl}(\cdot) is the main feature learning residual on the residual component, composed of enhancement attention modules (EAM) that are cascaded together as shown in Figure 2. Our network has small depth, but provides a wide receptive field through kernel dilation in each EAM initial two branch convolutions. The output features of the final layer are fed to the reconstruction module, which is again composed of one convolutional layer.

where Mr(⋅)M_{r}(\cdot) denotes the reconstruction layer.

where RIDNet(⋅\cdot) is our network and W\mathcal{W} denotes the set of all the network parameters learned. Our feature extraction MeM_{e} and reconstruction module MrM_{r} resemble the previous algorithms . We now focus on the feature learning residual on the residual block, and feature attention.

2 Feature learning Residual on the Residual

In this section, we provide more details on the enhancement attention modules that uses a Residual on the Residual structure with local skip and short skip connections. Each EAM is further composed of DD blocks followed by feature attention. Due to the residual on the residual architecture, very deep networks are now possible that improve denoising performance; however, we restrict our model to four EAM modules only. The first part of EAM covers the full receptive field of input features, followed by learning on the features; then the features are compressed for speed, and finally a feature attention module enhances the weights of important features from the maps. The first part of EAM is realized using a novel merge-and-run unit as shown in Figure 2 second row. The input features branched and are passed through two dilated convolutions, then concatenated and passed through another convolution. Next, the features are learned using a residual block of two convolutions while compression is achieved by an enhanced residual block (ERB) of three convolutional layers. The last layer of ERB flattens the features by applying a 1×11\times 1 kernel. Finally, the output of the feature attention unit is added to the input of EAM.

In image recognition, residual blocks are stacked together to construct a network of more than 1000 layers. Similarly, in image superresolution, EDSR stacked the residual blocks and used long skip connections (LSC) to form a very deep network. However, to date, very deep networks have not been investigated for denoising. Motivated by the success of , we introduce the residual on the residual as a basic module for our network to construct deeper systems. Now consider the m-th module of the EAM is given as

where fmf_{m} is the output of the EAMmEAM_{m} feature learning module, in other words fm=EAMm(fm−1)f_{m}=EAM_{m}(f_{m-1}). The output of each EAM is added to the input of the group as fm=fm+fm−1f_{m}=f_{m}+f_{m-1}. We have observed that simply cascading the residual modules will not achieve better performance, instead we add the input of the feature extractor module to the final output of the stacked modules as

where Ww,b\mathcal{W}_{w,b} are the weights and biases learned in the group. This addition i.e. LSC, eases the flow of information across groups. fgf_{g} is passed to reconstruction layer to output the same number of channels as the input of the network. Furthermore, we use another long skip connection to add the input image to the network output i.e. y^=Mr(fg)+x\hat{y}=M_{r}(f_{g})+x, in order to learn the residual (noise) rather than the denoised image, as this technique helps in faster learning as compared to learning original image due to the sparse representation of the noise.

This section provides information about the feature attention mechanism. Attention has been around for some time; however, it has not been employed in image denoising. Channel features in image denoising methods are treated equally, which is not appropriate for many cases. To exploit and learn the critical content of the image, we focus attention on the relationship between the channel features; hence the name: feature attention (see Figure 3).

An important question here is how to generate attention differently for each channel-wise feature. Images generally can be considered as having low-frequency regions (smooth or flat areas), and high-frequency regions (e.g., lines edges and texture). As convolutional layers exploit local information only and are unable to utilize global contextual information, we first employ global average pooling to express the statistics denoting the whole image, other options for aggregation of the features can also be explored to represent the image descriptor. Let fcf_{c} be the output features of the last convolutional layer having cc feature maps of size h×wh\times w; global average pooling will reduce the size from h×w×ch\times w\times c to 1×1×c1\times 1\times c as:

where fc(i,j)f_{c}(i,j) is the feature value at position (i,j)(i,j) in the feature maps.

Furthermore as investigated in , we propose a self-gating mechanism to capture the channel dependencies from the descriptor retrieved by global average pooling. According to , the mentioned mechanism must learn the nonlinear synergies between channels as well as mutually-exclusive relationships. Here, we employ soft-shrinkage and sigmoid functions to implement the gating mechanism. Let us consider δ\delta, and α\alpha are the soft-shrinkage and sigmoid operators, respectively. Then the gating mechanism is

where HDH_{D} and HUH_{U} are the channel reduction and channel upsampling operators, respectively. The output of the global pooling layer gpg_{p} is convolved with a downsampling Conv layer, activated by the soft-shrinkage function. To differentiate the channel features, the output is then fed into an upsampling Conv layer followed by sigmoid activation. Moreover, to compute the statistics, the output of the sigmoid (rcr_{c}) is adaptively rescaled by the input fcf_{c} of the channel features as

3 Implementation

Our proposed model contains four EAM blocks. The kernel size for each convolutional layer is set to 3×33\times 3, except the last Conv layer in the enhanced residual block and those of the features attention units, where the kernel size is 1×11\times 1. Zero padding is used for 3×33\times 3 to achieve the same size outputs feature maps. The number of channels for each convolutional layer is fixed at 64, except for feature attention downscaling. A factor of 16 reduces these Conv layers; hence having only four feature maps. The final convolutional layer either outputs three or one feature maps depending on the input. As for running time, our method takes about 0.2 second to process a 512×512512\times 512 image.

Experiments

To generate noisy synthetic images, we employ BSD500 , DIV2K , and MIT-Adobe FiveK , resulting in 4k images while for real noisy images, we use cropped patches of 512×512512\times 512 from SSID , Poly , and RENOIR . Data augmentation is performed on training images, which includes random rotations of 90∘, 180∘, 270∘ and flipping horizontally. In each training batch, 32 patches are extracted as inputs with a size of 80×8080\times 80. Adam is used as the optimizer with default parameters. The learning rate is initially set to 10−410^{-4} and then halved after 10510^{5} iterations. The network is implemented in the Pytorch framework and trained with an Nvidia Tesla V100 GPU. Furthermore, we use PSNR as evaluation metric.

2 Ablation Studies

Skip connections play a crucial role in our network. Here, we demonstrate the effectiveness of the skip connections. Our model is composed of three basic types of connections which includes long skip connection (LSC), short skip connections (SSC), and local connections (LC). Table 1 shows the average PSNR for the BSD68 dataset. The highest performance is obtained when all the skip connections are available while the performance is lower when any connection is absent. We also observed that increasing the depth of the network in the absence of skip connections does not benefit performance.

2.2 Feature-attention

Another important aspect of our network is feature attention. Table 1 compares the PSNR values of the networks with and without feature attention. The results support our claim about the benefit of using feature attention. Since the inception of DnCNN , the CNN models have matured, and further performance improvement requires the careful design of blocks and rescaling of the feature maps. The two mentioned characteristics are present in our model in the form of feature-attention and the skip connections.

3 Comparisons

We evaluate our algorithm using the Peak Signal-to-Noise Ratio (PSNR) index as the error metric and compare against many state-of-the-art competitive algorithms which include traditional methods i.e. CBM3D , WNNM , EPLL , CSF and CNN-based denoisers i.e. MLP , TNRD , DnCNN , IrCNN , CNLNet , FFDNet and CBDNet . To be fair in comparison, we use the default setting of the traditional methods provided by the corresponding authors.

In the experiments, we test four noisy real-world datasets i.e. RNI15 , DND , Nam and SSID . Furthermore, we prepare three synthetic noisy datasets from the widely used 12 classical images, BSD68 color and gray 68 images for testing. We corrupt the clean images by additive white Gaussian noise using noise sigma of 15, 25 and 50 standard deviations.

RNI15 provides 15 real-world noisy images. Unfortunately, the clean images are not given for this dataset; therefore, only the qualitative comparison is presented for this dataset.

Nam comprises of 11 static scenes and the corresponding noise-free images obtained by the mean of 500 noisy images of the same scene. The size of the images are enormous; hence, we cropped the images in 512×512512\times 512 patches and randomly selected 110 from those for testing.

DnD is recently proposed by Plotz et al. which originally contains 50 pairs of real-world noisy and noise-free scenes. The scenes are further cropped into patches of size 512×512512\times 512 by the providers of the dataset which resulted in 1000 smaller images. The near noise-free images are not publicly available, and the results (PSNR/SSIM) can only be obtained through the online system introduced by .

SSID (Smartphone Image Denoising Dataset) is recently introduced. The authors have collected 30k real noisy images and their corresponding clean images; however, only 320 images are released for training and 1280 images pairs for validation, as testing images are not released yet. We will use the validation images for testing our algorithm and the competitive methods.

3.2 Grayscale noisy images

In this subsection, we evaluate our model on the noisy grayscale images corrupted by spatially invariant additive white Gaussian noise. We compare against nonlocal self-similarity representative models i.e. BM3D and WNNM , learning based methods i.e. EPLL, TNRD , MLP , DnCNN , IrCNN , and CSF . In Tables 3 and 2, we present the PSNR values on Set12 and BSD68. It is to be remembered here that BSD500 and BSD68 are two disjoint sets. Our method outperforms all the competitive algorithms on both datasets for all noise levels; this may be due to the larger receptive field as well as better modeling capacity.

3.3 Color noisy images

Next, for noisy color image denoising, we keep all the parameters of the network similar to the grayscale model, except the first and last layer are changed to input and output three channels rather than one. Figure 4 presents the visual comparison and Table 4 reports the PSNR numbers between our methods and the alternative algorithms. Our algorithm consistently outperforms all the other techniques published in Table 4 for CBSD68 dataset . Similarly, our network produces the best perceptual quality images as shown in Figure 4. A closer inspection on the vase reveals that our network generates textures closest to the ground-truth with fewer artifacts and more details.

3.4 Real-World noisy images

To further assess the practicality of our model, we employ a real noise dataset. The evaluation is difficult because of the unknown level of noise, the various noise sources such as shot noise, quantization noise etc., imaging pipeline i.e. image resizing, lossy compression etc. Furthermore, the noise is spatially variant (non-Gaussian) and also signal dependent; hence, the assumption that noise is spatially invariant, employed by many algorithms does not hold for real image noise. Therefore, real-noisy images evaluation determines the success of the algorithms in real-world applications.

Next, we visually compare the result of our method with the competing methods on the denoised images provided by the online system of Plotz et al. in Figure 5. The PSNR and SSIM values are also taken from the website. From Figure 5, it is clear that the methods of perform poorly in removing the noise from the star and in some cases the image is over-smoothed, on the other hand, our algorithm can eliminate the noise while preserving the finer details and structures in the star image.

On RNI15 , we provide qualitative images only as the ground-truth images are not available. Figure 6 presents the denoising results on a low noise intensity image. FFDNet and CBDNet are unable to remove the noise in its totality as can been seen near the bottom left of handle and body of the cup image. On the contrary, our method is able to remove the noise without the introduction of any artifacts. We present another example from the RNI15 dataset with high noise in Figure 7. CDnCNN and FFDNet produce results of limited nature as some noisy elements can be seen in the near the eye and gloves of the Dog image. In comparison, our algorithm recovers the actual texture and structures without compromising on the removal of noise from the images.

We present the average PSNR scores of the resultant denoised images in Table 6. Unlike CBDNet , which is trained on Nam to specifically deal with the JPEG compression, we use the same network to denoise the Nam images and achieve favorable PSNR numbers. Our performance in terms of PSNR is higher than any of the current state-of-the-art algorithms. Furthermore, our claim is supported by the visual quality of the images produced by our model as shown in Figure 8. The amount of noise present after denoising by our method is negligible as compared to CDnCNN and other counterparts.

As a last dataset, we employ the SSID real noise dataset which has the highest number of test (validation) images available. The results in terms of PSNR are shown in the second row of Table 6. Again, it is clear that our method outperforms FFDNet and CBDNet by a margin of 9.5dB and 7.93dB, respectively. In Figure 9, we show the denoised results of a challenging image by different algorithms. Our technique recovers the true colors which are closer to the original pixel values while competing methods are unable to restore original colors and in specific regions induce false colors.

Conclusion

In this paper, we present a new CNN denoising model for synthetic noise and real noisy photographs. Unlike previous algorithms, our model is a single-blind denoising network for real noisy images. We propose a novel restoration module to learn the features and to enhance the capability of the network further; we adopt feature attention to rescale the channel-wise features by taking into account the dependencies between the channels. We also use LSC, SSC, and SC to allow low-frequency information to bypass so the network can focus on residual learning. Extensive experiments on three synthetic and four real-noise datasets demonstrate the effectiveness of our proposed model.This work was supported in part by NH&MRC Project grant # 1082358.

References