Learning to Generate Realistic Noisy Images via Pixel-level Noise-aware Adversarial Training
Yuanhao Cai, Xiaowan Hu, Haoqian Wang, Yulun Zhang, Hanspeter Pfister, Donglai Wei
Introduction
Image denoising is an important yet challenging problem in low-level vision. It aims to restore a clean image from its noisy counterpart. Traditional approaches concentrate on designing a rational maximum a posteriori (MAP) model, containing regularization and fidelity terms, from a Bayesian perspective . Some image priors like low-rankness , sparsity , and non-local similarity are exploited to customize a better rational MAP model. However, these hand-crafted methods are inferior in representing capacity. With the development of deep learning, image denoising has witnessed significant progress. Deep convolutional neural network (CNN) applies a powerful learning model to eliminate noise and has achieved promising performance . These deep CNN denoisers rely on a large-scale dataset of real-world noisy-clean image pairs. Nonetheless, collecting even small datasets is extremely tedious and labor-intensive. The process of acquiring real-world noisy-clean image pairs is to take hundreds of noisy images of the same scene and average them to get the clean image. To get more image pairs, researchers try to synthesize noisy images.
In particular, there are two common settings for synthesizing noisy images. As shown in Fig. 1 (a1), setting1 directly adds the additive white Gaussian noise (AWGN) with the clean RGB image. For a long time, single image denoising is performed with setting1. Nevertheless, fundamentally different from AWGN, real camera noise is generally more sophisticated and signal-dependent. The noise produced by photon sensing is further affected by the in-camera signal processing (ISP) pipeline (e.g., Gama correction, compression, and demosaicing). Models trained with setting1 are easily over-fitted to AWGN and fail in real noise removal. Setting2 is based on ISP-modeling CNN and Poisson-Gaussian noise model that modeling photon sensing with Poisson and remaining stationary disturbances with Gaussian has been adopted in RAW denoising. As shown in Fig. 1 (a2), setting2 adds a Poisson-Gaussian noise with the clean RAW image and then passes the result through a pre-trained RAW2RGB CNN to obtain the RGB noisy counterpart. Notably, when the clean RAW image is unavailable, a pre-trained RGB2RAW CNN is utilized to transform the clean RGB image to its RAW counterpart . However, setting2 has the following drawbacks: (i) The noise is assumed to obey a hand-crafted probability distribution. However, because of the randomness and complexity of real camera noise, it’s difficult to customize a hand-crafted probability distribution to model all the characteristics of real noise. (ii) The ISP pipeline is very sophisticated and hard to be completely modeled. The RAW2RGB branch only learns the mapping from the clean RAW domain to the clean RGB space. However, the mapping from the Poisson-Gaussian noisy RAW domain to the real noisy RGB space can not be ensured. (iii) The ISP pipelines of different devices vary significantly, which results in the poor generality and robustness of ISP modeling CNNs. Thus, whether noisy images are synthesized with setting1 or 2, there still remains a discrepancy between synthetic and real noisy datasets. We notice that GAN utilizes the internal information of the input image and external information from other images when modeling image priors. Hence, we propose to use GAN to adaptively learn the real noise distribution.
GAN is firstly introduced in and has been proven successful in image synthesis and translation . Subsequently, GAN is applied to image restoration and enhancement, e.g., super resolution , style transfer , enlighten , deraining , dehazing , image inpainting , image editing , and mobile photo enhancement . Although GAN is widely applied in low-level vision tasks, few works are dedicated to investigating the realistic noise generation problem . Chen et al. propose a simple GAN that takes Gaussian noise as input to generate noisy patches. However, as in general, this GAN is image-level, i.e., it treats images as samples and attempts to approximate the probability distribution of real-world noisy images. This image-level GAN neglects that each pixel of a real noisy image is a random variable and the real noise is spatio-chromatically correlated, thus results in coarse learning of the real noise distribution.
To alleviate the above problems, this work focuses on learning how to generate realistic noisy images so as to augment the training data for real denoisers. To begin with, we propose a simple yet reasonable noise model that treats each pixel of a real noisy image as a random variable. This noise model splits the noise generation problem into two sub-problems: image domain alignment and noise domain alignment. Subsequently, to tackle these two sub-problems, we propose a novel Pixel-level Noise-aware Generative Adversarial Network (PNGAN). During the training procedure of PNGAN, we employ a pre-trained real denoiser to map the generated and real noisy images into a nearly noise-free solution space to perform image domain alignment. Simultaneously, PNGAN establishes a pixel-level adversarial training that encourages the generator to adaptively simulate the real noise distribution so as to conduct the noise domain alignment. In addition, for better real noise fitting, we present a lightweight yet efficient CNN architecture, Simple Multi-scale Network (SMNet) as the generator. SMNet repeatedly aggregates multi-scale features to capture rich auto-correlation, which provides more sufficient spatial representations for noise simulating. Different from general image-level GAN, our discriminator is pixel-level. The discriminator outputs a score map. Each position on the score map indicates how realistic the corresponding noisy pixel is. With this pixel-level noise-aware adversarial training, the generator is encouraged to create solutions that are highly similar to real noisy images and thus difficult to be distinguished.
In conclusion, our contributions can be summarized into four points:
(1) We formulate a simple yet reasonable noise model. This model treats each noisy pixel as a random variable and then splits the noisy image generation into two parts: image and noise domain alignment.
(2) We propose a novel framework, PNGAN. It establishes an effective pixel-level adversarial training to encourage the generator to favor solutions that reside on the manifold of real noisy images.
(3) We customize an efficient CNN architecture, SMNet learning rich multi-scale auto-correlation for better noise fitting. SMNet serves as the generator in PNGAN costing only 0.8M parameters.
(4) Qualitative validation shows that noise generated by PNGAN is highly similar to real noise in terms of intensity and distribution. Quantitative experiments demonstrate that a series of denoisers finetuned with the generated noisy images achieve SOTA results on four real denoising benchmarks.
Proposed Method
As shown in Fig. 1, the pipeline of using PNGAN to perform data augmentation consists of three phases. (a) is the synthesizing phase. (a1) and (a2) are two common synthetic settings. In this phase, we produce the synthetic noisy image from its clean RGB or RAW counterpart. (b) is the training phase of PNGAN. The generator adopts the synthetic image as input. Which synthetic setting is selected is controlled by the switch. By using a pre-trained real denoiser , PNGAN establishes a pixel-level noise-aware adversarial training between the generator and discriminator so as to simultaneously conduct image and noise domain alignment. is set as RIDNet in this work. (c) is the finetuning phase. Firstly, in (c1), the generator creates extended fake noisy-clean image pairs. Secondly, in (c2), the fake and real data are jointly utilized to finetune a series of real denoisers.
Real camera noise is sophisticated and signal-dependent. Specifically, in the real camera system, the RAW noise produced by photon sensing comes from multiple sources (e.g., short noise, thermal noise, and dark current noise) and is further affected by the ISP pipeline. Besides, illumination changes and camera movement inevitably lead to spatial pixel misalignment and color or brightness deviation. Hence, hand-designed noise models based on mathematical assumptions are difficult to accurately and completely describe the properties of real noise. Different from previous methods, we don’t base our noise model on any mathematical assumptions. Instead, we use CNN to implicitly simulate the characteristics of real noise. We begin by noting that when taking multiple noisy images of the same scene, the noise intensity of the same pixel varies a lot. Simultaneously, affected by the ISP pipeline, the real noise is spatio-chromatically correlated. Thus, the correlation between different pixels of the same real noisy image should be considered. In light of these facts, we treat each pixel of a real noisy image as a random variable and formulate a simple yet reasonable noise model:
2 Pixel-level Noise-aware Adversarial Training
Our goal is to generate realistic noisy images. According to the noise model in Eq. (1), we split this problem into two sub-problems: (i) Image domain alignment aims to align . (ii) Noise domain alignment targets at modeling the distribution of . To handle the sub-problems, PNGAN establishes a novel pixel-level noise-aware adversarial training between and in Fig. 1 (b).
Image Domain Alignment. A very naive strategy to construct both image and noise domain alignment is to directly minimize the distance of and . However, due to the intrinsic randomness, complexity, and irregularity of real noise, directly deploying loss between and is unreasonable and drastically damages the quality of . Besides, as analyzed in Sec. 2.1, each pixel of is a distribution-unknown random variable. This indicates that such a naive strategy challenges the training and may easily cause the non-convergence issue. Therefore, the noise interference should be eliminated while constructing the image domain alignment. To this end, we feed and into to obtain their denoised versions and then perform loss between and :
By using , we can transfer and into a nearly noise-free solution space. The value of is relatively stable. Therefore, minimizing can encourage to favor solutions that after being denoised by converge to . In this way, the image domain alignment is constructed.
Noise Domain Alignment. Becasue of the complexity and variability of real noise, it’s hard to completely seperate from in Eq (1). Fortunately, we note that on the basis of constructing the image domain alignment of , the noise domain alignment of is equivalent to the distribution estimation of . Additionally, as the real noise is signal-dependent, the alignment between and is more beneficial to capture the correlation between noise and scene. We denote the distribution of as , some real noisy pixel samples of as such that , and the distribution of as . Here is the parameter of . Then we formulate the noise domain aligment into a maximum likelihood estimation problem:
where denotes the relativistic discriminator, means the Sigmoid activation, and represents the non-transformed discriminator output. estimates the probability that real data is more realistic than fake data and also directs the generator to create a fake image that is more realistic than real images. The loss functions of and are then defined in a symmetrical form:
During the training procedure, we fix to train and fix to train iteratively. Minimizing and alternately allows us to train a generative model with the goal of fooling the pixel-level discriminator that is trained to distinguish fake noisy images from real noisy images. This pixel-level noise-aware adversarial training scheme encourages to favor perceptually natural solutions that reside on the manifold of real noisy images so as to construct the noise domain alignment.
3 Noisy Image Generating
In Sec. 2.1, we denote the probability distribution of as . Now we customize a light-weight yet efficient CNN architecture, SMNet as to generate . In this section, we firstly introduce the input setting of and subsequently detail the architecture of SMNet.
Input Setting. We aim to generate a realistic noisy image from its clean counterpart. A naive setting is to directly adopt the clean image as the input to generate the noisy image. However, this naive setting is not in line with the fact. When we repeatedly feed the same clean image to a pre-trained , outputs completely the same noisy images. In contrast, when taking multiple pictures in the real world, the real noisy images vary a lot in the intensity of each pixel. This is caused by many factors (e.g., photon sensing noise, ISP pipelines, and illumination conditions). Hence, the naive input setting containing no distribution is unreasonable. We review that the general GANs sample from an initial random distribution (usually Gaussian) to generate a fake image. Hence, the input of should contain a random distribution so as to generate multiple noisy images of the same scene. We note that the two common synthetic settings meet this condition. Therefore, we utilize the two common settings to produce the synthetic image and then adopt the synthetic image as the input of . Subsequently, we propose a light-weight yet efficient architecture, SMNet for better real noise fitting.
where denotes Fast Channel Attention. denotes a conv layer after bilinear interpolation upsampling, 2 is the scale factor. is similarly defined. means Shift-Invariant Downsample , 2 is also the scale factor. is similarly defined. Subsequently, the output feature is derived by:
where represents the last conv layer, denotes the concatenating operation. The architecture of FCA is shown in Fig. 3 (d). We define the input feature as , then FCA can be formulated as:
where represents the Sigmoid activation function, means global average pooling along the spatial wise, denotes 1-Dimension Convolution. In this work, we set = 3, = 2, and = 64.
4 Overall Training Objective
In addition to the aforementioned losses, we employ a perceptual loss function that assesses a solution with respect to perceptually relevant characteristics (e.g., the structural contents and detailed textures):
where denotes the last feature map of VGG16 . Eventually, the training objective is:
where and are two hyper-parameters controlling the importance balance. The proposed PNGAN framework is end-to-end trained by minimizing . Note that the parameters in and VGG16 are fixed. Each mini-batch training procedure is divided into two steps: (i) Fix and train . (ii) Fix and train . This pixel-level adversarial training scheme promotes the ability to distinguish fake noisy images from real noisy images and allows to learn to create the solutions that are highly similar to real camera noisy images and thus difficult to be classified by .
Experiment
Datasets. We first use SIDD train set to train . Then we fix to train on the same set. Subsequently, uses clean images from DIV2K , Flickr2K , BSD68 , Kodak24 , and Urban100 to generate realistic noisy-clean image pairs. We use the generated data and SIDD train set jointly to finetune real denoisers and evaluate them on four real denoising benchmarks: SIDD , DND , PolyU , and Nam . The images in SIDD are collected using five smartphone cameras in 10 static scenes. There are 320 image pairs for training and 1,280 image patch pairs for validation. DND composes 50 noisy-clean image pairs captured by 4 consumer cameras. 1,000 patches at size 512512 are cropped from the collected images. PolyU consists of 40 real camera noisy images. Nam is composed of real noisy images of 11 static scenes.
Implementation Details. We set the hyper-parameter = 610-3, = 810-4. For synthetic setting1, we set the noise intensity, = 50. For synthetic setting2, we directly exploit CycleISP to generate the synthetic noisy input. All the sub-modules (, , and ) are trained with the Adam optimizer ( and ) for 7105 iterations. The initial learning rate is set to 210-4. The cosine annealing strategy is employed to steadily decrease the learning rate from the initial value to 10-6 during the training procedure. Patches at size 128128 cropped from training images are fed into the models. The batch size is set as 8. The horizontal and vertical flips are performed for data augmentation. All the models are trained on RTX8000 GPUs. In the finetuning phase, the learning rate is set to 110-6, other settings remain unchanged.
2 Quantitative Results
Domain Discrepancy Validation. We use the widely applied metric, Maximum Mean Discrepancy (MMD) to measure the domain discrepancy between synthetic and real-world noisy images, PNGAN generating, and real noisy images on four real noisy benchmarks. For DND, we derive a pseudo clean version by denoising the real noisy counterparts with a pre-trained MIRNet . Then we use the pseudo clean version to synthesize noisy images. The results are depicted as a histogram in Fig. 4. For setting1, the domain discrepancy decreases by 74%, 75%, 44%, and 43% on SIDD, DND, PolyU, and Nam when PNGAN is exploited. For setting2, the discrepancy decreases by 64%, 67%, 46%, and 44%. These results demonstrate that PNGAN can narrow the discrepancy between synthetic and real noisy datasets. Please refer to the supplementary for detailed calculation process.
Comparison with SOTA Methods. We use the generated noisy-clean image pairs (setting2) to finetune a series of denoisers. We compare our models with SOTA algorithms on four real denoising datasets: SIDD, DND, PolyU, and Nam. The results are reported in Tab. 1. * denotes denoisers finetuned with image pairs generated by PNGAN. We have the following observations: (i) Our denoisers outperform SOTA methods by a large margin. Specifically, MPRNet* and MIRNet* exceed the recent best method MIRNet by 0.34 and 0.35 dB on SIDD, 0.30 and 0.37 dB on DND. RIDNet*, MPRNet*, and MIRNet* surpass the best performers by 0.36, 1.30, and 1.37 dB on PolyU and 0.01, 1.04, and 1.10 dB on Nam. (ii) Compared with the counterparts that are not finetuned, our models achieve a significant promotion. In particular, RIDNet* is 0.54, 0.29, 0.68, and 0.49 dB higher than RIDNet on SIDD, DND, PolyU, and Nam. MPRNet* achieves 0.35, 0.38, 1.41, and 1.31 dB gain than MPRNet on SIDD, DND, PolyU, and Nam. MIRNet* is improved by 0.35, 0.37, 1.37, and 1.21 dB. This evidence clearly suggests the high similarity between PNGAN generating and real noisy images. Denoisers adapted with our fake image pairs generalize better across different benchmarks.
Train from Scratch. For more strong comparisons, we use the fake noisy images generated from clean SIDD train and DF2K (DIV2K+Flicker2K) respectively to train denoisers from scratch. The PSNR results evaluated on SIDD test are listed in Tab. 2. All models are trained with the same experiment schedule except the training data. It can be observed: (i) On SIDD train, when PNGAN is applied to setting1, denoisers are promoted by 15.65 dB and only 0.84 dB lower than those trained with real data (SIDD train set). While applying PNGAN to setting2 (CycleISP), denoisers are improved by 2.87 dB. Surprisingly, in this case, denoisers achieve almost the same performance as those trained with real data. The relative error is 0.2. (ii) To validate the generality of PNGAN, we also adopt synthetic DF2K noisy-clean image pairs to train denoisers. As shown in the right part of Tab. 2, when PNGAN is applied to setting1, denoisers are promoted by 9.59 dB. While applying PNGAN to setting2, denoisers are improved by 4.35 dB and only 0.75 dB lower than those trained with SIDD real train set. These results convincingly demonstrate: (i) The generated noise is highly similar to the real noise especially when PNGAN is applied to synthetic setting2. (ii) PNGAN can significantly narrow the domain discrepancy between synthetic and real-world noise.
3 Qualitative Results
Visual Examinations of Noisy Images. To intuitively evaluate the generated noisy images, we provide visual comparisons of noisy images on the four real noisy datasets, as shown in Fig. 5. Note that the clean image of DND is pseudo, denoised from its noisy version by a MIRNet. The left part depicts noisy images from SIDD, DND, PolyU, and Nam (top to down). The right part exhibits the patches cropped by the yellow bboxes, from left to right: clean, synthetic setting1, setting2 (CycleISP), PNGAN generating, and real noisy images. As can be seen from the zoom-in patches: (i) Noisy images synthesized by setting1 is signal-independent. The distribution and intensity remain unchanged across diverse scenes, indicating the characteristics of AWGN fundamentally differ from those of the real noise. (ii) Noisy images generated by PNGAN are closer to the real noise than those synthesized by setting2 visually. Noise synthesized by setting2 shows randomness that is obviously inconsistent with the real noise in terms of intensity and distribution. While PNGAN can model spatio-chromatically correlated and non-Gaussian noise more accurately. (iii) Even if passing through the same camera pipeline, different shooting conditions lead to the diversity of real noise. It’s unreasonable for the noise synthesized by CycleISP to show nearly uniform fitting to different input images. In contrast, PNGAN can adaptively simulate more sophisticated and photo-realistic models. This adaptability allows PNGAN to show robust performance across different real noisy datasets.
Visual Comparison of Denoised Images. We compare the visual results of denoisers before and after being finetuned (denoted with *) with the generated data in Fig. 4. We observe that models finetuned with the generated data are more effective in real noise removal. Furthermore, they are capable of preserving the structural content, textural details, and spatial smoothness of the homogeneous regions. In contrast, original models either yield over-smooth images sacrificing fine textural details and structural content or introduce redundant blotchy texture and chroma artifacts.
4 Ablation Study
Break-down Ablations. We perform break-down ablations to evaluate the effects of PNGAN components and SMNet architecture. We select setting1 to synthesize the noisy input from SIDD train set. Then we use the generated data only to train the denoisers from scratch and evaluate them on SIDD test. The PSNR results are reported in Tab. 3. (i) Firstly, is set as SMNet to validate the effects of PNGAN components. We start from Baseline1, no discriminator is used and the loss is directly performed between and in Eq. (2). Denoisers trained with the generated data collapse dramatically, implying the naive strategy mentioned in Sec. 2.2 is unfeasible. When is applied, the denoisers are promoted by 21.81 dB on average. In addition, the PSNR and SSIM between the denoised counterparts of generated and real noisy images are 39.14 dB and 0.928 on average respectively. This evidence indicates that successfully conducts the image domain alignment as mentioned in Sec. 2.2. Subsequently, we use an image-level with stride conv layers to classify whether the whole generated image is real. Nonetheless, the performance of denoisers remains almost unchanged. After deploying , the models are improved by 2.09 dB, suggesting that the pixel-level noise model is more in line with real noise scenes and benefits generating more realistic noisy images. When is used, the denoisers gain a slight improvement by about 0.39 dB, indicating facilitates yielding more vivid results. (ii) Secondly, we only change the architecture of to study the effects of its components. We start from Baseline2 that doesn’t exploit multi-scale feature fusion, SID, and FCA. When we add two different scale branches and use bilinear interpolation to downsample and upsample, denoisers trained with the generated images are promoted by about 1.28 dB. After applying SID and FCA, the denoisers further gain 0.28 and 0.74 dB improvement on average. These results convincingly demonstrate the superiority of the proposed SMNet in real-world noise fitting.
Parameter Analysis. We adopt RIDNet as the baseline to perform parameter analysis. We firstly validate the effects of , in Eq. (13), and the noise intensity of setting1, i.e., . We change the parameters, train , use to generate realistic noisy images from clean images of SIDD train set, train RIDNet with the generated data, and evaluate its performance on SIDD test set. When analyzing one parameter, we fix the others at their optimal values. The PSNR results are shown in Fig. 7. The optimal setting is = 610-3, = 810-4, and = 40 or 50. Secondly, we evaluate the effect of the ratio of finetuning data. We denote the ratio of extended training data (setting2) to SIDD real noisy training data as . We change the value of , finetuned the original RIDNet, and test on three real denoising datasets: SIDD, PolyU, and Nam. The results are listed in Tab. 4. When = 0, all the finetuning data comes from SIDD train set, RIDNet achieves the best performance on SIDD. However, its performance on PolyU and Nam degrades drastically due to the domain discrepancy between different real noisy datasets. We gradually increase the value of to study its effects. The average performance on the three datasets yields the maximum when = 60%.
Conclusion
Too much research focuses on designing a CNN architecture for real noise removal. In contrast, this work investigates how to generate more realistic noisy images so as to boom the denoising performance. We first formulate a noise model that treats each noisy pixel as a random variable. Then we propose a novel framework PNGAN to perform the image and noise domain alignment. For better noise fitting, we customize an efficient architecture, SMNet as the generator. Experiments show that noise generated by PNGAN is highly similar to real noise in terms of intensity and distribution. Denoisers finetuned with the generated data outperform SOTA methods on real denoising datasets.
Acknowledgement
This work is jointly supported by the NSFC fund (61831014), in part by the Shenzhen Science and Technology Project under Grant (ZDYBH201900000002, JCYJ20180508152042002, CJGJZD20200617102601004).