GRDN:Grouped Residual Dense Network for Real Image Denoising and GAN-based Real-world Noise Modeling

Dong-Wook Kim, Jae Ryun Chung, Seung-Won Jung

Introduction

In the field of image denoising, recent studies show that learning-based methods are more efficient than previous handcrafted methods such as block matching 3D (BM3D) and its variants. It is essential for learning-based methods to have a sufficient amount of dataset with high quality. Because a pair of noisy and noise-free images can be easily constructed by adding synthetic noise to noise-free images, a majority of previous learning-based methods focus on the classic Gaussian denoising task and pay the most attention to the architecture design of networks, especially convolutional neural networks (CNNs). However, due to the gap between synthetically generated noisy images and real-world noisy images, it was found that CNNs trained using synthetic images do not perform well on real-world noisy images and sometimes even inferior to BM3D .

Toward real-world image denoising, there have been two main approaches. The first approach is to find a better statistical model of real-world noise rather than the additive white Gaussian noise . In particular, a combination of Gaussian and Poisson distributions was shown to closely model both signal-dependent and signal-independent noise. The networks trained using these new synthetic noisy images demonstrated the superiority in denoising real-world noisy images. One clear advantage of this approach is that we can have infinitely many training image pairs by simply adding the synthetic noise to noise-free ground-truth images. However, it is still arguable whether the real-world noise can be modeled by statistic models. The second approach is thus in an opposite direction. From real-world noisy images, nearly noise-free ground-truth images can be obtained by inverting an image acquisition procedure . To our knowledge, smartphone image denoising dataset (SIDD) is one of the largest high quality image datasets on the second approach. However, the amount of provided images may not be enough for training a large network and without a sufficient knowhow it is difficult to generate ground-truth images from real-world noisy images. We thus adopt the second approach but applied our own generative adversarial network (GAN)-based data augmentation technique to obtain a larger dataset.

The network architecture is of course the utmost important. In CNN-based image restoration, dense residual blocks (RDBs) have received great attention. In this paper, we propose a new architecture called grouped residual dense network (GRDN). In particular, the proposed architecture adopts the recent residual dense network (RDN) as a component with a minor modification and defines it as grouped residual dense block (GRDB). By cascading the GRDBs with attention modules, we could obtain the state-of-the-art performance in real-world image denoising task . We achieved the best performance in terms of the peak signal-to-noise ratio (PSNR) of 39.93 dB and the structural similarity (SSIM) of 0.9736 in the NTIRE2019 Real Image Denoising Challenge - Track 2:sRGB.

Related Works

Image denoising is one of the most extensively studied topics in image processing. Owing to significant advances in deep learning, CNN-based methods are now dominating in image denoising. However, most previous learning-based image denoising methods have focused on the classic Gaussian denoising task. Toward real-world image denoising, the first approach was to capture a pair of noisy and noise-free images by using different camera settings . It was shown in that earlier learning-based methods were comparable or sometimes even inferior to classic methods such as BM3D. We consider this is mainly because of insufficient quality and quantity of training dataset. Consequently, more abundant and elaborate datasets such as Darmstadt noise dataset (DND) and SIDD were developed, and recent learning-based methods showed their superiority over classic methods for real-world image denoising.

In addition to the efforts in generating high quality datasets, a significant amount of research has been made to find better network architectures for image denoising. From the viewpoint of CNNs, network architectures developed for different image restoration tasks such as image denoising, image deblurring, super-resolution, and compress artifact reduction share similarities. It has been repeatedly demonstrated that one architecture developed for a certain image restoration task also performs well in other restoration tasks . We thus examined many of the architectures developed for different image restoration tasks, especially super-resolution . Among them, RDN and residual channel attention network (RCAN) are most closely related to our network architecture.

In particular, we attempt to take advantage of novel ideas in RDN and RCAN. RCAN introduced residual in residual (RIR) architecture, and the ablation study showed that the performance gain by RIR was the most significant. Thus, we use this RIR principle in our architecture design. In addition, RDN itself is an image restoration network but we use it with modifications as a component of our network and construct a cascaded structure of RDNs as our image denoising network. Recent studies also showed the effectiveness of attention modules. Among many attention modules, convolutional block attention module (CBAM) , an easily implantable module that sequentially estimates channel attention and spatial attention, showed efficacy in general object detection and image classification, and thus we include CBAM into our network.

2 GAN

The amount of training images in publicly available real-world image denoising datasets such as SIDD and DND may not be enough to train a deep and wide neural network. One feasible way of augmenting these datasets is to exploit the capability of GAN . The first GAN-based real-world noise modeling method uses only real-world noisy images for training the noise generator, where the discriminator is trained to distinguish between real and simulated noise signals. The noise generator is then used to add synthetic but realistic noise to noise-free ground-truth images, and the denoising network is finally trained using the generated pairs of ground-truth and noisy images. The real-world image denoising performance was significantly improved by using the dataset generated by GAN.

We improve the previous GAN-based real-world noise simulation technique by including conditioning signals such as the noise-free image patch, ISO, and shutter speed as additional inputs to the generator. The conditioning on the noise-free image patch can help generating more realistic signal-dependent noise and the other camera parameters can increase controllability and variety of simulated noise signals. We also change the discriminator of the previous architecture by using a recent relativistic GAN . Unlike conventional GANs, the discriminator of the relativistic GAN learns to determine which is more realistic between real data and fake data. Our method is different from the conventional relativistic GAN in that both the real and fake data are used as an input to make the discriminator more explicitly compare the two data.

Proposed Methods

Our image denoising network architecture called GRDN is shown in Fig. 1. Our designing principle is to distribute burdens of each layer such that a deeper and wider network can be well trained. To this end, residual connections are applied in four different levels. Down-sampling and up-sampling layers are included to enable a deeper and wider architecture and CBAM is also applied.

Inspired by RDN , we use RDB as shown in Fig. 2(a) as a building module. In RDN, the features from cascaded RDBs are concatenated together and followed by the 1×{\times}1 convolutional layer. We define this feature concatenation part of RDN, as shown in Fig. 2(b), as GRDB and use it as a building module of our GRDN. Note that the original RDN applies convolutional layers before and after GRDB and uses global residual learning for image denoising. However, we consider that RDN imposes a heavy burden to the very last 1×{\times}1 convolutional layer of GRDB. Therefore, we instead cascade GRDBs such that the features from RDBs can be fused in multiple stages. Motivated by many recent image restoration networks including RDN , we also include the global residual connection such that the network can focus on learning the difference between the noisy and ground-truth images. Last, we exploit CBAM as a building module to further improve the denoising performance. The position of the CBAM block was empirically chosen as in-between the upconvolutional layer and the last convolutional layer.

Although GRDN is structurally deeper than RDN , we used the same number of RDBs. Specifically, 16 RDBs were used in the original RDN for image denoising. We use 4 stack of GRDBs and each GRDB consists of 4 RDBs, resulting 16 RDBs in GRDN.

2 GAN-based Real-world Noise Modeling

Motivated by the recent technique , we develop our own generator and discriminator for real-world noise modeling. Likewise with the previous technique , we use residual blocks (ResBlocks) as a building module of the generator. However, we made several modifications to improve the performance of real-world noise modeling. Fig. 3 show the generator architecture. First, we include conditioning signals: the noise-free image patch, ISO, shutter speed, and smartphone model as an additional input to the generator. The conditioning on the noise-free image patch can help generating more realistic signal-dependent noise and the other camera-related parameters can increase controllability and variety of simulated noise signals. To train the generator with these conditioning signals, we used the metadata of SIDD . Second, spectral normalization (SN) is applied before batch normalization in the basic convolutional units like the one used in . Third, our ResBlock includes the residual scaling . SN and residual scaling were empirically found to be useful in training our generator.

Our discriminator architecture as shown in Fig. 4 is also different from the previous GAN-based noise simulation technique . Enhanced super-resolution GAN (ESRGAN) showed that relativistic GAN is effective in generating realistic image textures. Unlike original GAN, the discriminator of relativistic GAN learns to determine which is more realistic between real data and fake data. Let C(x)C(x) denote the non-transformed discriminator output for input image xx. The standard discriminator can then be expressed as D(x)=σ(C(x))D(x)=\sigma(C(x)), σ\sigma is the sigmoid function. The discriminator of relativistic average GAN (RaGAN) adopted in ESGAN is defined as:

The discriminator of the proposed network, defined as conditioned explicit relativistic GAN (cERGAN), is given as

where xcx_{c} denote the conditioning signal. Specifically, we make each conditioning data have the same size as the training patch by replicating values, and thus our xcx_{c} consists of 4 patches: 3 constant patches from smartphone code (e.g. Google Pixel = 0, iPhone 7 = 1, etc), ISO level, and shutter speed, and one noise-free image patch. In addition to xcx_{c}, we also use both xrx_{r} and xfx_{f} as an input of the discriminator. Note that ESGAN uses either xrx_{r} or xfx_{f} as an input of the discriminator.

The loss functions of the generator and discriminator, denoted as LGcERGANL^{cERGAN}_{G} and LDcERGANL^{cERGAN}_{D}, respectively, are finally defined as follows:

In other words, if the second input is xrx_{r} and the third input is xfx_{f}, the discriminator is trained to predict a value close to 1, i.e., xrx_{r} is more realistic than xfx_{f}. If the two inputs are switched, the discriminator is trained to predict a value close to 0, i.e., xfx_{f} is less realistic than xrx_{r}. The generator is trained to fool the discriminator. By requiring the network to explicitly compare between real data and fake data, we could simulate more realistic real-world noise.

Experiments

We implemented all of our models using PyTorch library with Intel i7-8700 @3.20GHz, 32GB of RAM, and NVIDIA Titan XP.

We used the training and validation images of NTIRE 2019 Real Image Denoising Challenge, which is a subset of SIDD dataset . Let ChDB denote the dataset we used for our experiment. Specifically, 320 high-resolution images and 1280 cropped image blocks with the size 256×\times256 were used for training and validation, respectively. The provided images were taken by five smartphone cameras - Apple iPhone 7, Google Pixel, Samsung Galaxy S6 Edge, Motorola Nexus 6, and LG G4. Because the ground-truth images of the test dataset are not publicly available, we report the performance of image denoising models using the validation dataset in this Section. Since we noticed non-marginal degradations around image borders in ground-truth images, we excluded the first and last 8 rows/columns when generating training patches. General data augmentation techniques such as scaling, flipping, and rotation were not applied.

2 Image Denoising

We augmented the provided training dataset by two ways. First, we used the author-provided source code of for adding synthetic noise to the ground-truth images. We also applied our own GAN-based noise simulator described in Sec.3.2 to generate additional synthetic noisy images.

In each training batch, we randomly extracted 16 pairs of ground-truth and noisy image patches. We trained using Adam with β1{\beta_{1}} = 0.9, β2{\beta_{2}} = 0.999. The initial learning rate was set to 10−4{10^{-4}} and then decreased to half at every 2×105{2\times 10^{5}} iteration. We trained the network using L1L_{1} loss. We trained our model for approximately 5 days.

We used 4×{\times}4 filters for up/down-convolutional layer and 1×{\times}1 filters for fusing the features concatenated from RDBs. Otherwise, we used 3×{\times}3 filters. Zero-padding was used and dilation was not used for all convolutional layers. Each RDB has 8 pairs of convolutional layers and ReLU activation layers.

2.2 Comparison to RDN

First, we compared our GRDN model with RDN . The experimental result is shown in Table 1. We re-trained RDN using ChDB. The 1st and 2nd columns in Table 1 correspond to RDN and proposed GRDN. It can be seen that the PSNR of our model is 0.04 dB higher than that of RDN. Note that RDN and GRDN have the same number of RDBs, and thus the number of parameters is similar. Specifically, our basic GRDN model has 22M parameters while RDN has 21.9M parameters.

2.3 Experiments on patch size

Since the original image resolution is very high (more than 12M pixels), the largest possible patch size needs to be used to include sufficient image contents. We thus increased the patch size to 96 ×\times 96, which was the largest possible size in our experimental environment. By comparing the 2nd and 5th columns of Table 1, we can see that the significant performance gain of 0.22dB was obtained by increasing the patch size.

2.4 Experiments on CBAM module

CBAM is a simple but effective module for CNNs. Because it is a lightweight and general module, it can be easily implanted to any CNN architectures without largely increasing the number of parameters. In particular, CBAM can be placed at bottlenecks of the network. Since we have down-sampling and up-sampling layers, we examined different positions and combinations of CBAMs. We concluded that for our model the best position of CBAM is after the up-sampling layer. We believe this indicates that CBAM enhances important features from the up-sampled data. It also helps to construct a final denoised image for the last convolution layer which comes after. The effectiveness of CBAM was found to be dependent on the complexity of the network. Comparing the 2nd and 3rd columns of Table 1, CBAM increased the PSNR by 0.05 dB. However, after increasing the patch size, the gain by CBAM became diluted. Comparing the 5th and 6th columns of Table 1, CBAM even decreased the PSNR by 0.01 dB.

2.5 Hyper parameter adjustment

We compared networks with different numbers of filters and GRDBs. Comparing the 6th and 7th columns of Table 1, a less deeper but more wider network performed 0.02 dB better. Therefore, the model on the 7th column is the best performing model under our hardware constraints.

3 Real-world Noise Modeling

For training the generator and discriminator of cERGAN, we cropped image patches with the size 48×\times48 from real-world noisy images and their ground-truth images from ChDB. We used the batch size of 32 and Adam optimizer with β1=0\beta_{1}=0 and β2=0.9\beta_{2}=0.9. The generator and discriminator were trained for 340k iterations. The initial learning rate was set as 0.0002 for both discriminator and generator, and we linearly decayed the learning rate after 320k iterations such that the learning rate became 0 after the last iteration. Fig. 5 illustrates some of noise image patches generated by the proposed cERGAN. As can be seen in Figs. 5(c) and (d), the proposed cERGAN can generate noise patches close to real-world noise.

The effectiveness of simulated noisy images was evaluated by comparing the proposed image denoising network trained with/without the simulated data. Here, the tested network corresponds to the 4th column of Table 1. We first attempted to train our image denoising network using only the synthesized real-world noisy images obtained by cERGAN. The average PSNR was obtained as 38.63 dB in ChDB validation set, which is inferior to the one we obtained using only the provided ChDB dataset (39.62 dB in Table 1).

Second, we used the author-provided source code of for adding statistically modeled real-world noise to ground-truth images of ChDB. Our image denoising network trained using these dataset only resulted in 36.17 dB, which demonstrates that the proposed GAN-based noise modeling at least performs better than the statistic noise modeling method .

Last, we combined the original ChDB dataset with the synthetic datasets generated by the proposed cERGAN and conventional method . Here, we could test only one configuration: 90% from ChDB, 5% from simulated ChDB using , and 5% from simulated ChDB using cERGAN. Fig. 6 shows that the PSNR obtained using the augmented dataset increases more stably. The resultant PSNR was obtained as 39.64 dB, which is slightly higher than the PSNR obtained using the original dataset (39.62 dB).

NTIRE2019 Image Denoising Challenge

This work is proposed for participating in the NTIRE2019 Real Image Denoising Challenge - Track 2:sRGB. The challenge aims to develop an image denoising system with the highest PSNR and SSIM. The submitted image denoising network corresponds to 7th column of Table 1. One minor change in the submitted model is that we included skip connections for every 2 GRDBs. For training, we used the augmented ChDB using the technique mentioned in Sec. 4.3. Our model ranked 1st place for real image denoising both in terms of PSNR and SSIM. As shown in Table 2, our model outperformed the 2nd rank method by 0.05 dB.

Conclusion

In this paper, we proposed an improved network architecture for real-world image denoising. By using residual connections extensively and hierarchically, our model achieved the state-of-the-art performance. Furthermore, we developed an improved GAN-based real-world noise modeling method.

Although we could evaluate the proposed network only to real-world image denoising, we believe that the proposed network is generally applicable. We thus plan to apply the proposed image denoising network to other image restoration tasks. We also could not fully and quantitatively justify the effectiveness of the proposed real-world noise modeling method. A more elaborate design is clearly necessary for better real-world noise modeling. We believe that our real-world noise modeling method can be extended to other real-world degradations such as blur, aliasing, and haze, which will be demonstrated in our future work.

Acknowledgement

This work was supported by Institute for Information and communications Technology Promotion(IITP) grant funded by the Korea government(MSIP) 2017-0-00072, Development of Audio/Video Coding and Light Field Media Fundamental Technologies for Ultra Realistic Tera-media.

References