Toward Convolutional Blind Denoising of Real Photographs
Shi Guo, Zifei Yan, Kai Zhang, Wangmeng Zuo, Lei Zhang
Introduction
Image denoising is an essential and fundamental problem in low-level vision and image processing. With decades of studies, numerous promising approaches have been developed and near-optimal performance has been achieved for the removal of additive white Gaussian noise (AWGN). However, in real camera system, image noise comes from multiple sources (e.g., dark current noise, short noise, and thermal noise) and is further affected by in-camera processing (ISP) pipeline (e.g., demosaicing, Gamma correction, and compression). All these make real noise much more different from AWGN, and blind denoising of real-world noisy photographs remains a challenging issue.
In the recent past, Gaussian denoising performance has been significantly advanced by the development of deep CNNs . However, deep denoisers for blind AWGN removal degrades dramatically when applied to real photographs (see Fig. 1(d)). On the other hand, deep denoisers for non-blind AWGN removal would smooth out the details while removing the noise (see Fig. 1(e)). Such an phenomenon may be explained from the characteristic of deep CNNs , where their generalization largely depends on the ability of memorizing large scale training data. In other words, existing CNN denoisers tend to be over-fitted to Gaussian noise and generalize poorly to real-world noisy images with more sophisticated noise.
In this paper, we tackle this issue by developing a convolutional blind denoising network (CBDNet) for real-world photographs. As indicated by , the success of CNN denoisers are significantly dependent on whether the distributions of synthetic and real noises are well matched. Therefore, realistic noise model is the foremost issue for blind denoising of real photographs. According to , Poisson-Gaussian distribution which can be approximated as heteroscedastic Gaussian of a signal-dependent and a stationary noise components has been considered as a more appropriate alternative than AWGN for real raw noise modeling. Moreover, in-camera processing would further makes the noise spatially and chromatically correlated which increases the complexity of noise. As such, we take into account both Poisson-Gaussian model and in-camera processing pipeline (e.g., demosaicing, Gamma correction, and JPEG compression) in our noise model. Experiments show that in-camera processing pipeline plays a pivot role in realistic noise modeling, and achieves notably performance gain (i.e., dB by PSNR) over AWGN on DND .
We further incorporate both synthetic and real noisy images to train CBDNet. On one hand, it is easy to access massive synthetic noisy images. However, the noise in real photographs cannot be fully characterized by our model, thereby giving some leeway for improving denoising performance. On the other hand, several approaches have suggested to get noise-free image by averaging hundreds of noisy images at the same scene. Such solution, however, is expensive in cost, and suffers from the over-smoothing effect of noise-free image. Benefited from the incorporation of synthetic and real noisy images, dB gain on PSNR can be attained by CBDNet on DND .
Our CBDNet is comprised of two subnetworks, i.e., noise estimation and non-blind denoising. With the introduction of noise estimation subnetwork, we adopt an asymmetric loss by imposing more penalty on under-estimation error of noise level, making our CBDNet perform robustly when the noise model is not well matched with real-world noise. Besides, it also allows the user to interactively rectify the denoising result by tuning the estimated noise level map. Extensive experiments are conducted on three real noisy image datasets, i.e., NC12 , DND and Nam . In terms of both quantitative metrics and perceptual quality, our CBDNet performs favorably in comparison to state-of-the-arts. As shown in Fig. 1, both non-blind BM3D and DnCNN for blind AWGN fail to denoise the real-world noisy photograph. In contrast, our CBDNet achieves very pleasing denoising results by retaining most structure and details while removing the sophisticated real-world noise.
To sum up, the contribution of this work is four-fold:
A realistic noise model is presented by considering both heteroscedastic Gaussian noise and in-camera processing pipeline, greatly benefiting the denoising performance.
Synthetic noisy images and real noisy photographs are incorporated for better characterizing real-world image noise and improving denoising performance.
Benefited from the introduction of noise estimation subnetwork, asymmetric loss is suggested to improve the generalization ability to real noise, and interactive denoising is allowed by adjusting the noise level map.
Experiments on three real-world noisy image datasets show that our CBDNet achieves state-of-the-art results in terms of both quantitative metrics and visual quality.
Related Work
The advent of deep neural networks (DNNs) has led to great improvement on Gaussian denoising. Until Burger et al. , most early deep models cannot achieve state-of-the-art denoising performance . Subsequently, CSF and TNRD unroll the optimization algorithms for solving the fields of experts model to learn stage-wise inference procedure. By incorporating residual learning and batch normalization , Zhang et al. suggest a denoising CNN (DnCNN) which can outperform traditional non-CNN based methods. Without using clean data, Noise2Noise also achieves state-of-the-art. Most recently, other CNN methods, such as RED30 , MemNet , BM3D-Net , MWCNN and FFDNet , are also developed with promising denoising performance.
Benefited from the modeling capability of CNNs, the studies show that it is feasible to learn a single model for blind Gaussian denoising. However, these blind models may be over-fitted to AWGN and fail to handle real noise. In contrast, non-blind CNN denoisiers, e.g., FFDNet , can achieve satisfying results on most real noisy images by manually setting proper or relatively higher noise level. To exploit this characteristic, our CBDNet includes a noise estimation subnetwork as well as an asymmetric loss to suppress under-estimation error of noise level.
2 Image Noise Modeling
Most denoising methods are developed for non-blind Gaussian denoising. However, the noise in real images comes from various sources (dark current noise, short noise, thermal noise, etc.), and is much more sophisticated . By modeling photon sensing with Poisson and remaining stationary disturbances with Gaussian, Poisson-Gaussian noise model has been adopted for the raw data of imaging sensors. In , camera response function (CRF) and quantization noise are also considered for more practical noise modeling. Instead of Poisson-Gaussian, Hwang et al. present a Skellam distribution for Poisson photon noise modeling. Moreover, when taking in-camera image processing pipeline into account, the channel-independent noise assumption may not hold true, and several approaches are proposed for cross-channel noise modeling. In this work, we show that realistic noise model plays a pivot role in CNN-based denoising of real photographs, and both Poisson-Gaussian noise and in-camera image processing pipeline benefit denoising performance.
3 Blind Denoising of Real Images
Blind denoising of real noisy images generally is more challenging and can involve two stages, i.e., noise estimation and non-blind denoising. For AWGN, several PCA-based methods have been developed for estimating noise standard deviation (). Rabie models the noisy pixels as outliers and exploits Lorentzian robust estimator for AWGN estimation. For Poisson-Gaussian model, Foi et al. suggest a two-stage scheme, i.e., local estimation of multiple expectation/standard-deviation pairs, and global parametric model fitting.
In most blind denoising methods, noise estimation is closely coupled with non-blind denoising. Portilla adopts a Gaussian scale mixture for modeling wavelet patches of each scale, and utilizes Bayesian least square to estimate clean wavelet patches. Based on the piecewise smooth image model, Liu et al. propose a unified framework for the estimation and removal of color noise. Gong et al. model the data fitting term as the weighted sum of the and norms, and utilize a sparsity regularizer in wavelet domain for handling mixed or unknown noises. Lebrun et al. propose an extension of non-local Bayes approach by modeling the noise of each patch group to be zero-mean correlated Gaussian distributed. Zhu et al. suggest a Bayesian nonparametric technique to remove the noise via the low-rank mixture of Gaussians (LR-MoG) model. Nam et al. model the cross-channel noise as a multivariate Gaussian and perform denoising by the Bayesian nonlocal means filter . Xu et al. suggest a multi-channel weighted nuclear norm minimization (MCWNNM) model to exploit channel redundancy. They further present a trilateral weighted sparse coding (TWSC) method for better modeling noise and image priors . Except noise clinic (NC) , MCWNNM , and TWSC , the codes of most blind denoisers are not available. Our experiments show that they are still limited for removing noise from real images.
Proposed Method
This section presents our CBDNet consisting of a noise estimation subnetwork and a non-blind denoising subnetwork. To begin with, we introduce the noise model to generate synthetic noisy images. Then, the network architecture and asymmetric loss. Finally, we explain the incorporation of synthetic and real noisy images for training CBDNet.
As noted in , the generalization of CNN largely depends on the ability in memorizing training data. Existing CNN denoisers, e.g., DnCNN , generally does not work well on real noisy images, mainly due to that they may be over-fitted to AWGN while the real noise distribution is much different from Gaussian. On the other hand, when trained with a realistic noise model, the memorization ability of CNN will be helpful to make the learned model generalize well to real photographs. Thus, noise model plays a critical role in guaranteeing performance of CNN denoiser.
Different from AWGN, real image noise generally is more sophisticated and signal-dependent . Practically, the noise produced by photon sensing can be modeled as Poisson, while the remaining stationary disturbances can be modeled as Gaussian. Poisson-Gaussian thus provides a reasonable noise model for the raw data of imaging sensors , and can be further approximated with a heteroscedastic Gaussian defined as,
where is the irradiance image of raw pixels. involves two components, i.e., a stationary noise component with noise variance and a signal-dependent noise component with spatially variant noise variance .
Real photographs, however, are usually obtained after in-camera processing (ISP), which further increases the complexity of noise and makes it spatially and chromatically correlated. Thus, we take two main steps of ISP pipeline, i.e., demosaicing and Gamma correction, into consideration, resulting in the realistic noise model as,
where denotes the synthetic noisy image, stands for the camera response function (CRF) uniformly sampled from the 201 CRFs provided in . And is adopted to generate irradiance image from a clean image . represents the function that converts sRGB image to Bayer image and represents the demosaicing function . Note that the interpolation in involves pixels of different channels and spatial locations. The synthetic noise in Eqn. (2) is thus channel and space dependent.
Furthermore, to extend CBDNet for handling compressed image, we can include JPEG compression in generating synthetic noisy image,
For noisy uncompressed image, we adopt the model in Eqn. (2) to generate synthetic noisy images. For noisy compressed image, we exploit the model in Eqn. (3). Specifically, and are uniformly sampled from the ranges of and , respectively. In JPEG compression, the quality factor is sampled from the range $$. We note that the quantization noise is not considered because it is minimal and can be ignored without any obvious effect on denoising result .
2 Network Architecture
As illustrated in Fig. 2, the proposed CBDNet includes a noise estimation subnetwork and a non-blind denosing subnetwork . First, takes a noisy observation to produce the estimated noise level map , where denotes the network parameters of . We let the output of be the noise level map due to that it is of the same size with the input and can be estimated with a fully convolutional network. Then, takes both and as input to obtain the final denoising result , where denotes the network parameters of . Moreover, the introduction of also allows us to adjust the estimated noise level map before putting it to the the non-blind denosing subnetwork . In this work, we present a simple strategy by letting for interactive denoising.
We further explain the network structures of and . adopts a plain five-layer fully convolutional network without pooling and batch normalization operations. In each convolution (Conv) layer, the number of feature channels is set as , and the filter size is . The ReLU nonlinearity is deployed after each Conv layer. As for , we adopt an U-Net architecture which takes both and as input to give a prediction of the noise-free clean image. Following , the residual learning is adopted by first learning the residual mapping and then predicting . The 16-layer U-Net architecture of is also given in Fig. 2, where symmetric skip connections, strided convolutions and transpose convolutions are introduced for exploiting multi-scale information as well as enlarging receptive field. All the filter size is , and the ReLU nonlinearity is applied after every Conv layer except the last one. Moreover, we empirically find that batch normalization helps little for the noise removal of real photographs, partially due to that the real noise distribution is fundamentally different from Gaussian.
Finally, we note that it is also possible to train a single blind CNN denoiser by learning a direct mapping from noisy observation to clean image. However, as noted in , taking both noisy image and noise level map as input is helpful in generalizing the learned model to images beyond the noise model and thus benefits blind denoising. We empirically find that single blind CNN denoiser performs on par with CBDNet for images with lower noise level, and is inferior to CBDNet for images with heavy noise. Furthermore, the introduction of noise estimation subnetwork also makes interactive denoising and asymmetric learning allowable. Therefore, we suggest to include the noise estimation subnetwork in our CBDNet.
3 Asymmetric Loss and Model Objective
Both CNN and traditional non-blind denoisers perform robustly when the input noise is higher than the ground-truth one (i.e., over-estimation error), which encourages us to adopt asymmetric loss for improving generalization ability of CBDNet. As illustrates in FFDNet , BM3D/FFDNet achieve the best result when the input noise and ground-truth noise are matched. When the input noise is lower than the ground-truth one, the results of BM3D/FFDNet contain perceptible noises. When the input noise is higher than the ground-truth one, BM3D/FFDNet can still achieve satisfying results by gradually wiping out some low contrast structure along with the increase of input noise . Thus, non-blind denoisers are sensitive to under-estimation error of noise , but are robust to over-estimation error. With such property, BM3D/FFDNnet can be used to denoise real photographs by setting relatively higher input noise , and this might explain the reasonable performance of BM3D on the DND benchmark in the non-blind setting.
To exploit the asymmetric sensitivity in blind denoising, we present an asymmetric loss on noise estimation to avoid the occurrence of under-estimation error on the noise level map. Given the estimated noise level at pixel and the ground-truth , more penalty should be imposed to their MSE when . Thus, we define the asymmetric loss on the noise estimation subnetwork as,
Furthermore, we introduce a total variation (TV) regularizer to constrain the smoothness of ,
where () denotes the gradient operator along the horizontal (vertical) direction. For the output of non-blind denoising, we define the reconstruction loss as,
To sum up, the overall objective of our CBDNet is,
where and denote the tradeoff parameters for the asymmetric loss and TV regularizer, respectively. In our experiments, the PSNR/SSIM results of CBDNet are reported by minimizing the above objective. As for qualitative evaluation of visual quality, we train CBDNet by further adding perceptual loss on relu3_3 of VGG-16 to the objective in Eqn. (7).
4 Training with Synthetic and Real Noisy Images
The noise model in Sec. 3.1 can be used to synthesize any amount of noisy images. And we can also guarantee the high quality of the clean images. Even though, the noise in real photographs cannot be fully characterized by the noise model. Fortunately, according to , nearly noise-free image can be obtained by averaging hundreds of noisy images from the same scene, and several datasets have been built in literatures. In this case, the scenes are constrained to be static, and it is generally expensive to acquire hundreds of noisy images. Moreover, the nearly noise-free image tends to be over-smoothing due to the averaging effect. Therefore, synthetic and real noisy images can be combined to improve the generalization ability to real photographs.
In this work, we use the noise model in Sec. 3.1 to generate the synthetic noisy images, and use 400 images from BSD500 , 1600 images from Waterloo , and 1600 images from MIT-Adobe FiveK dataset as the training data. Specifically, we use the RGB image x to synthesize clean raw image as a reverse ISP process and use the same to generate noisy image as Eqns. (2) or (3), where is a CRF randomly sampled from those in . As for real noisy images, we utilize the 120 images from the RENOIR dataset . In particular, we alternatingly use the batches of synthetic and real noisy images during training. For a batch of synthetic images, all the losses in Eqn. (7) are minimized to update CBDNet. For a batch of real images, due to the unavailability of ground-truth noise level map, only and are considered in training. We empirically find that such training scheme is effective in improving the visual quality for denoising real photographs.
Experimental Results
Three datasets of real-world noisy images, i.e., NC12 , DND and Nam , are adopted:
includes 12 noisy images. The ground-truth clean images are unavailable, and we only report the denoising results for qualitative evaluation.
DND
contains 50 pairs of real noisy images and the corresponding nearly noise-free images. Analogous to , the nearly noise-free images are obtained by carefully post-processing of the low-ISO images. PSNR/SSIM results are obtained through the online submission system.
Nam
contains 11 static scenes and for each scene the nearly noise-free image is the mean image of 500 JPEG noisy images. We crop these images into patches and randomly select 25 patches for evaluation.
2 Implementation Details
The model parameters in Eqn. (7) are given by , , and . Note that the noisy images from Nam are JPEG compressed, while the noisy images from DND are uncompressed. Thus we adopt the noise model in Eqn. (2) to train CBDNet for DND and NC12, and the model in Eqn. (3) to train CBDNet(JPEG) for Nam.
To train our CBDNet, we adopt the ADAM algorithm with = 0.9. The method in is adopted for model initialization. The size of mini-batch is 32 and the size of each patch is . All the models are trained with 40 epochs, where the learning rate for the first 20 epochs is , and then the learning rate is used to further fine-tune the model. It takes about three days to train our CBDNet with the MatConvNet package on a Nvidia GeForce GTX 1080 Ti GPU.
3 Comparison with State-of-the-arts
We consider four blind denoising approaches, i.e., NC , NI , MCWNNM and TWSC in our comparison. NI is a commercial software and has been included into Photoshop and Corel PaintShop. Besides, we also include a blind Gaussian denoising method (i.e., CDnCNN-B ), and three non-blind denoising methods (i.e., CBM3D , WNNM , FFDNet ). When apply non-blind denoiser to real photographs, we exploit to estimate the noise .
Fig. 3 shows the results of an NC12 images. All the competing methods are limited in removing noise in the dark region. In comparison, CBDNet performs favorably in removing noise while preserving salient image structures.
DND.
Table 1 lists the PSNR/SSIM results released on the DND benchmark website. Undoubtedly, CDnCNN-B cannot be generalized to real noisy photographs and performs very poorly. Although the noise is provided, non-blind Gaussian denoisers, e.g., WNNM , BM3D and FoE , only achieve limited performance, mainly due to that the real noise is much different from AWGN. MCWNNM and TWSC are specially designed for blind denoising of real photographs, and also achieve promising results. Benefited from the realistic noise model and incorporation with real noisy images, our CBDNet achieves the highest PSNR/SSIM results, and slightly better than MCWNNM and TWSC . CBDNet also significantly outperforms another CNN-based denoiser, i.e., CIMM . As for running time, CBDNet takes about 0.4s to process an image. Fig. 4 provides the denoising results of an DND image. BM3D and CDnCNN-B fail to remove most noise from real photograph, NC, NI, MCWNNM and TWSC still cannot remove all noise, and NI also suffers from the over-smoothing effect. In comparison, our CBDNet performs favorably in balancing noise removal and structure preservation.
Nam.
The quantitative and qualitative results are given in Table 2 and Fig. 5. CBDNet(JPEG) performs much better than CBDNet (i.e., dB by PSNR) and achieves the best performance in comparison to state-of-the-arts.
4 Ablation Studies
Instead of AWGN, we consider heterogeneous Gaussian (HG) and in-camera processing (ISP) pipeline for modeling image noise. On DND and Nam, we implement four variants of noise models: (i) Gaussian noise (CBDNet(G)), (ii) heterogeneous Gaussian (CBDNet(HG)), (iii) Gaussian noise and ISP (CBDNet(G+ISP)), and (iv) heterogeneous Gaussian and ISP (CBDNet(HG+ISP), i.e., full CBDNet. For Nam, CBDNet(JPEG) is also included. Table 3 shows the PSNR/SSIM results of different noise models.
G vs HG.
Without ISP, CBDNet(HG) achieves about dB gain over CBDNet(G). When ISP is included, the gain by HG is moderate, i.e., CBDNet(HG+ISP) only outperforms CBDNet(G+ISP) about dB.
w/o ISP.
In comparison, ISP is observed to be more critical for modeling real image noise. In particular, CBDNet(G+ISP) outperforms CBDNet(G) by dB, while CBDNet(HG+ISP) outperforms CBDNet(HG) by dB on DND. For Nam, the inclusion of JPEG compression in ISP further brings a gain of 1.31 dB.
Incorporation of synthetic and real images.
We implement two baselines: (i) CBDNet(Syn) trained only on synthetic images, and (ii) CBDNet(Real) trained only on real images, and rename our full CBDNet as CBDNet(All). Fig. 7 shows the denoising results of these three methods on a NC12 image. Even trained on large scale synthetic image dataset, CBDNet(Syn) still cannot remove all real noise, partially due to that real noise cannot be fully characterized by the noise model. CBDNet(Real) may produce over-smoothing results, partially due to the effect of imperfect noise-free images. In comparison, CBDNet(All) is effective in removing real noise while preserving sharp edges. Also quantitative results of the three models on DND are shown in Table 1. CBDNet(All) obtains better PSNR/SSIM results than CBDNet(Syn) and CBDNet(Real).
Asymmetric loss.
Fig. 8 compares the denoising results of CBDNet with different values, i.e., and . CBDNet imposes equal penalty to under-estimation and over-estimation errors when , and more penalty is imposed on under-estimation error when . It can be seen that smaller (i.e., ) is helpful in improving the generalization ability of CBDNet to unknown real noise.
5 Interactive Image Denoising
Given the estimated noise level map , we introduce a coefficient to interactively modify to . By allowing the user to adjust , the non-blind denoising subnetwork takes and the noisy image as input to obtain denoising result. Fig. 6 presents two real noisy DND images as well as the results obtained using different values. By specifying to the first image and to the second, CBDNet can achieve the results with better visual quality in preserving detailed textures and removing sophisticated noise, respectively. Such interactive scheme can thus provide a convenient means for adjusting the denosing results in practical scenario.
Conclusion
We presented a CBDNet for blind denoising of real-world noisy photographs. The main findings of this work are two-fold. First, realistic noise model, including heterogenous Gaussian and ISP pipeline, is critical in making the learned model from synthetic images be applicable to real-world noisy photographs. Second, the denoising performance of a network can be boosted by incorporating both synthetic and real noisy images in training. Moreover, by introducing a noise estimation subnetwork into CBDNet, we were able to utilize asymmetric loss to improve its generalization ability to real-world noise, and perform interactive denoising conveniently.
Acknowledgements
This work is supported by NSFC (grant no. 61671182, 61872118, 61672446) and HK RGC General Research Fund (PolyU 152216/18E).