Invertible Denoising Network: A Light Solution for Real Noise Removal

Yang Liu, Zhenyue Qin, Saeed Anwar, Pan Ji, Dongwoo Kim, Sabrina Caldwell, Tom Gedeon

Introduction

Image denoising aims to restore clean images from noisy observations. Traditional approaches model denoising as a maximum a posteriori (MAP) optimization problem, with assumptions on the distribution of noise , and natural image priors . Although these algorithms achieve satisfactory performance on removing synthetic noise, their effectiveness on real-world noise is compromised since their assumptions deviate from those in real-world scenarios. Recently, convolutional neural networks (CNNs) have achieved superior denoising performance . These CNNs learn the features of images from a large number of clean and noisy image pairs. However, since real noise is very complex, to achieve better denoising accuracy, CNN denoising models have become increasingly large and complicated . Thus, although some methods can achieve very impressive denoising results, they may not be practical in realistic scenarios such as deploying the model on edge equipment like smartphones and motion sensing devices.

Currently, a substantial amount of research has been devoted to developing neural networks that are invertible . For image denoising, invertible networks are advantageous from the following three aspects: (1) the model is light, as encoding and decoding use the same parameters; (2) they preserve details of the input data since invertible networks are information-lossless ; (3) they save memory during back-propagation because they use a constant amount of memory to compute gradients, regardless of the depth of the network . Hence, invertible models are suitable for small devices like smartphones. We thus study employing invertible networks to address the problem of image denoising. However, applying such networks to remove noise is non-trivial. The original inputs and the reversed results of the traditional invertible models follow the same distribution . In contrast, for image denoising, the input is noisy, and the restored image is clean, following two different distributions. Therefore, invertible denoising networks are required to abandon the noise in the latent space before the reversion. Due to this difficulty, noise removal has not previously been studied and deployed in invertible literature and models.

In this paper, we propose an invertible denoising network, InvDN, to resolve the above difficulties. Unlike previous invertible models, two different latent variables are involved; one incorporates noise and high-frequency clean contents while the other only encodes the clean part. During the forward pass, InvDN transforms the input image to a downscaled latent representation with an increased number of channels. We train InvDN to make the first three channels of the latent representation the same as the low-resolution clean image. Since invertible networks preserve all the information of the input , noisy signals are in the rest of the channels. To remove noise completely, we discard all the channels that contain noise. However, as a side-effect, we also lose some information corresponding to the high-resolution clean image. To reconstruct such missing information, we sample a new latent variable from a prior distribution and combine it with the low-resolution image to restore the clean image.

We are the first to design invertible networks for real image denoising to the best of our knowledge.

The latent variable of traditional invertible networks follows a single distribution. Instead, InvDN has two latent variables following two different distributions. Thus, InvDN can not only restore clean images but also generate new noisy images.

We achieve a new state-of-the-art (SOTA) result on the SIDD test set, using far fewer parameters and less run time than the previous SOTA methods.

InvDN is able to generate new noisy images that are more similar to the original noisy ones.

Related Work

In this section, we summarize and discuss the development and recent trends in image denoising. The widely used denoising methods can be classified into traditional methods and current data-driven deep learning methods.

Traditional Methods. Model-driven denoising methods usually construct a MAP optimization problem with a loss and a regularization term. The assumptions on the noise distribution are needed for most traditional methods to build the model. One assumed distribution is the Mixture of Gaussian, which is used as an approximator for noise on natural patches or patch groups . The regularization term is usually based on the clean image’s prior. Total variation denoising uses the statistical characteristics of images to remove noise. Sparsity is enforced in dictionary learning methods , to learn over-complete dictionaries from clean images. Non-local similarity methods employ non-local patches that share similar patterns. Such a strategy is adopted by the notable BM3D and NLM . However, these models are limited due to the assumptions on the prior of spatially invariant noise or clean images, which are often different from real cases, where the noise is spatially variant.

Data-driven Deep Learning Denoising. Recent years have seen rapid progress in the deep learning methods, boosting the denoising performance to a large extent. The early deep models focus on synthetic noisy image denoising due to a lack of real data. As some large real noise datasets, such as DND and SIDD , have been presented, current research focuses on blind real image denoising. There are two main streams in real image denoising. One is to adapt the methods that work well on the synthetic datasets to the real datasets while considering the gap between these two domains . The current most competitive method along this direction is AINDNet , which applies transfer learning from synthetic to real denoising with the Adaptive Instance Normalization operations.

The other direction is to model real noise with more complicated distributions and design new network architectures . VDN proposed by Yue et al.assumes that noise follows an inverse Gamma distribution, and the clean image we observe is a conjugate Gaussian prior of the unavailable real clean images. They propose a new training objective based on these assumptions and use two parallel branches to learn these two distributions in the same network. Its potential limitation is that the assumptions are not suitable when the noise distribution becomes complicated. Later, DANet abandons the assumptions for noise distributions and employs a GAN framework to train the model. Two parallel branches are also employed in this architecture: one for denoising and the other for noise generation. This design concept is that the three kinds of image pairs (clean and noisy, clean and generated noisy, as well as denoised and noisy) follow the same distribution, so they use a discriminator to train the model. The potential limitation is that GAN-based models’ training is unstable and thus takes longer to converge . Furthermore, both VDN and DANet employ Unet in the parallel branches, making their models very large.

To compress the model size, we explore invertible networks. To the best of our knowledge, few studies apply invertible networks in denoising literature. Noise Flow introduces an invertible architecture to learn real noise distributions to generate real noise as a way of data augmentation. Generating noisy images with Noise Flow requires extra information apart from the sRGB images, including raw-RGB images, ISO, and camera-specific values. They do not propose new denoising backbones. So far, no invertible network for real image denoising has been reported.

Invertible Denoising Network

In this paper, we present a novel denoising architecture consisting of invertible modules, \ie, Invertible Denoising Network (InvDN). For completeness, in this section, we first provide the background of invertible neural networks and then present the details of InvDN.

Invertible networks are originally designed for unsupervised learning of probabilistic models . These networks can transform a distribution to another distribution through a bijective function without losing information . Thus, it can learn the exact density of observations. Using invertible networks, images following a complex distribution can be generated through mapping a given latent variable z\mathbf{z}, which follows a simple distribution pz(z)p_{\mathbf{z}}(\mathbf{z}), to an image instance x∼px(x)\mathbf{x}\sim p_{\mathbf{x}}(\mathbf{x}), \ie, x=f(z)\mathbf{x}=f(\mathbf{z}), where ff is the bijective function learned by the network. Due to the bijective mapping and exact density estimation properties, invertible networks have received increasing attention in recent years and have been applied successfully in applications such as image generation and rescaling .

2 Challenges in Denoising with Invertible Models

Applying invertible models in denoising is different from other applications. The widely used invertible networks employed in image generation and rescaling consider the input and the reverted image to follow the same distribution. Such applications are straightforward candidates for invertible models. Image denoising, however, takes a noisy image as input and reconstructs a clean one, \ie, the input and the reverted outcome follow two different distributions. On the other hand, an invertible transform does not lose any information during the transformation. However, the lossless property of invertibility is not desired for image denoising since the noise information remains while we transform an input image into latent variables. If we can disentangle the noisy and clean signals during an invertible transformation, we may reconstruct a clean image without worrying about losing any important information by abandoning the noisy information. In the following section, we present one way to obtain a clean signal through an invertible transformation.

3 Concept of Design

We denote the original noisy image as y\mathbf{y}, its clean version as x\mathbf{x} and the noise as n\mathbf{n}. We have: p(y)=p(x,n)=p(x)p(n∣x)p(\mathbf{y})=p(\mathbf{x},\mathbf{n})=p(\mathbf{x})p(\mathbf{n}|\mathbf{x}). Using the invertible network, the learned latent representation of observation y\mathbf{y} contains both noise and clean information. It should be noted that it is non-trivial to disentangle them and abandon only the noisy part.

On the other hand, invertible networks utilize different feature extraction approaches comparing with existing deep denoising models. Existing ones usually employ convolutional layers with padding to extract features. However, they are not invertible due to two reasons: Firstly, the padding makes the network non-invertible; secondly, the parameter matrices of convolutions may not be full-rank. Thus, to ensure invertibility, rather than using convolutional layers, it is necessary to utilize invertible feature extraction methods, such as the Squeeze layer and Haar Wavelet Transformation , as presented in Fig. 2. The Squeeze operation reshapes the input into feature maps with more channels according to a checkerboard pattern. Haar Wavelet Transformation extracts the average pooling of the original input as well as the vertical, horizontal, and diagonal derivatives. As a result, the spatial size of the feature maps extracted by the invertible methods is inevitably downscaled.

Therefore, instead of disentangling the clean and noisy signals directly, we aim to separate the low-resolution and high-frequency components of a noisy image. The sampling theory indicates that during the downsampling process, the high-frequency signals are discarded. Since invertible networks are information-lossless , if we make the first three channels of the transformed latent representation to be the same as the downsampled clean image, high-frequency information will be encoded in the remaining channels. Based on the observation that the high-frequency information contains noise as well, we abandon all high-frequency representations before inversion to reconstruct a clean image from low-resolution components. We formally describe the process as follows:

where xLR\mathbf{x}_{\text{LR}} represents the low-resolution clean image. We use xHF\mathbf{x}_{\text{HF}} to represent the high-frequency contents that cannot be obtained by xLR\mathbf{x}_{\text{LR}} when reconstructing the original clean image. Since it is challenging to disentangle xHF\mathbf{x}_{\text{HF}} and n\mathbf{n}, we abandon all the channels representing z∼p(xHF,n∣xLR)\mathbf{z}\sim p(\mathbf{x}_{\text{HF}},\mathbf{n}|\mathbf{x}_{\text{LR}}) to remove spatially variant noise completely. Nevertheless, a side-effect is the loss of xHF\mathbf{x}_{\text{HF}}. To reconstruct xHF\mathbf{x}_{\text{HF}}, we sample zHF∼N(0,I)\mathbf{z}_{\text{HF}}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) and train our invertible network to transform zHF\mathbf{z}_{\text{HF}}, in conjunction with xLR\mathbf{x}_{\text{LR}}, to restore the clean image x\mathbf{x}. In this way, the lost high-frequency clean details xHF\mathbf{x}_{\text{HF}} is embedded in the latent variable zHF\mathbf{z}_{\text{HF}}.

4 Network Architecture

We first employ a supervised approach to guide the network to separate high-frequency and low-resolution components during transformation. After some invertible transformation gg, the noisy image y\mathbf{y} is transformed into its corresponding low-resolution clean image and high-frequency encoding z\mathbf{z}, \ieg(y)=[g(y)LR;z]g(\mathbf{y})=[g(\mathbf{y})_{\text{LR}};\mathbf{z}]. We minimize the following forward objective

where g(y)LRg(\mathbf{y})_{\text{LR}} is the low-frequency components learned by the network, corresponding to three channels of the output representation in the forward pass. MM is the number of pixels. ∣∣⋅∣∣m||\cdot||_{m} is the mm-norm and mm can be either 1 or 2. To obtain the ground truth low-resolution image xLR\mathbf{x}_{\text{LR}}, we down-sample the clean image x\mathbf{x} via bicubic transformation.

To restore the clean image with g(y)LRg(\mathbf{y})_{\text{LR}}, we use inverse transform g−1([g(y)LR;zHF])g^{-1}([g(\mathbf{y})_{\text{LR}};\mathbf{z}_{\text{HF}}]) with random variable zHF\mathbf{z}_{\text{HF}} sampled from normal distribution N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}). The backward objective is written as

where x\mathbf{x} is the clean image. NN is the number of pixels. We train the invertible transformation gg by simultaneously utilizing both forward and backward objectives.

Inspired by , the invertible transform gg we present is of a multi-scale architecture, consisting of several down-scale blocks. Each down-scale block consists of an invertible wavelet transformation followed by a series of invertible blocks. The overall architecture of the InvDN model is demonstrated in Fig. 3.

In this section, we present some empirical performances of InvDN on real denoising tasks.

We evaluate InvDN on three real-world denoising benchmarks, \ie, SIDD , DND , and RNI .

SIDD is taken by five smartphone cameras with small apertures and sensor sizes. We use the medium version of SIDD as the training set, containing 320 clean-noisy pairs for training and 1280 cropped patches from the other 40 pairs for validation. In each iteration of training, we crop the input image into multiple 144×144144\times 144 patches to feed into the network. The reported test results are obtained via an online submission system.

DND is captured by four consumer-grade cameras of differing sensor sizes. It contains 50 pairs of real-world noisy and approximately noise-free images. These images are cropped into 1000 patches of size 512×512512\times 512. Similarly to SIDD, the performance is evaluated by submitting the outputs of the methods to the online system.

RNI15 is composed of 15 real-world noisy images without ground-truths. Therefore, we only provide visual comparisons on this dataset.

2 Training Details.

Our InvDN has two down-scale blocks, each of which is composed of eight invertible blocks. All the models are trained with Adam as the optimizer, with momentum of β1=0.9,β2=0.999\beta_{1}=0.9,\beta_{2}=0.999. The batch size is set as 14, and the initial learning rate is fixed at 2×10−42\times 10^{-4}, which decays by half every 50k iterations. PyTorch is used as the implementation framework, and training is performed on a single 2080-Ti GPU. We augment the data with horizontal and vertical flipping, as well as random rotations of 90×θ90\times\theta where θ=0,1,2,3\theta=0,1,2,3.

3 Experimental Results

We evaluate the methods by commonly used metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), which are also available on the real-noisy DND and SIDD websites.

Quantitative Measure. As mentioned before, we train InvDN using the SIDD medium training set. DND does not provide the training set; therefore, we use the model trained on SIDD. For a fair comparison, the PSNR of other competitive models on the test set is directly taken from the official DND and SIDD leaderboards and verified from the respective articles. Table 2 reports the test results of different denoising models. We can observe that the number of model parameters is directly positively correlated to the model’s performance. From RIDNet to DANet , the number of model parameters increases from 1.49 million (M) to 63.01M with a slight improvement in PSNR.

Nevertheless, InvDN reverses the trend, having only 2.64M parameters. Compared with the most recently proposed SOTA DANet, InvDN only uses less than 4.2% of the number of parameters of DANet. Although InvDN has far fewer parameters, the denoising performance on the test set is better than all the recently proposed denoising models, achieving a new SOTA result for the SIDD datasetWe only compare with the published methods trained on the benchmark SIDD training set.. The performance of InvDN is also comparable with that of the recent competitive models on DND, indicating the generalization ability of our lightweight denoising model.

Qualitative Measure. To further illustrate the better performance of InvDN against other methods, in Fig. 4, we show denoised visual results on images from three different datasets, SIDD, DND and RNI15. The first row of the figure illustrates that InvDN restores well-shaped patterns, while other competing techniques induce blurry textures and artifacts. Furthermore, the second row depicts that InvDN recovers the subtle edges very clearly, whereas other models bring artifacts and over-smoothness. Finally, from the third row, we observe that InvDN reconstructs accurate edges compared to other models that introduce blockiness, fuzziness, and random dots, particularly along the edges.

4 Ablation Study

Squeeze vs Wavelet Transform. We compare the difference of employing the squeeze and Haar wavelet transform in down-scale blocks. We report PSNR on the validation set of the same iteration during training in 4(a). The remaining network components and the parameter settings are the same. We observe that using Haar wavelet converges faster and is more stable than the squeeze operation.

Residual Block vs. Dense Block. Next, we provide comparison between different blocks (ϕi⁡\operatorname{\phi_{i}}) in our network in 4(b). The architecture with residual block achieves higher denoising accuracy than dense block used by . Moreover, the network with residual block has far fewer parameters (2.6M) than that using the dense block (4.3M).

Number of Down-Scale and Invertible Blocks. Now, we study the denoising performance of InvDN with different numbers of down-scale and invertible blocks. We report the PSNR results on the validation set from the same iteration. We discover that increasing the number of invertible blocks boosts denoising accuracy consistently, regardless of the number of down-scale blocks. Moreover, when fixing the number of invertible blocks, we observe that using two down-scale blocks leads to the best denoising effect.

5 Monte Carlo Self Ensemble

To further improve InvDN’s denoising effect without using extra data and training, we introduce Monte Carlo (MC) self-ensemble. We sample the latent variable zHF\mathbf{z}_{\text{HF}} multiple times, resulting in a set of latent variables {zHFi}i=1N\{\mathbf{z}_{\text{HF}}^{i}\}_{i=1}^{N}, where for each zHFi\mathbf{z}_{\text{HF}}^{i}, a corresponding denoised image x^i\hat{x}_{i} is obtained. The final output is the average of {x^i}i=1N\{\hat{x}_{i}\}_{i=1}^{N}. By the law of large numbers , the averaged image is closer to the ground-truth than any individual x^i\hat{x}_{i}. In Fig. 6, we visualize the denoising results before and after using MC self-ensembling by setting the MC size to 16. The contrast between the residual images before and after MC self-ensemble indicates that it further reduces the noise in the denoised image. Quantitatively, 83.35% images witness a performance boost on the SIDD validation set.

6 Analysis of Distributions of 𝐳𝐳\mathbf{z}

For InvDN, there are two types of latent variables: one corresponds to the original high-frequency signal z\mathbf{z} containing noise, and the other sampled zHF\mathbf{z}_{\text{HF}} from N(0,I)\mathcal{N}(\mathbf{0},\mathbf{I}) representing clean details. To analyze, we obtain 500 pairs of z\mathbf{z} and zHF\mathbf{z}_{\text{HF}} from the SIDD validation set. We vectorize each latent variable and plot them in a 3D space with PCA . As 6(a) shows, z\mathbf{z} and zHF\mathbf{z}_{\text{HF}} follow two different distributions.

For the sake of fair comparison, we only train InvDN on the benchmark SIDD training set. However, InvDN also supports generating more noisy images for data augmentation. To generate augmented data, we first conduct sampling z\mathbf{z}. We introduce a tiny disturbance to z\mathbf{z} as z′=z+ϵ⋅v\mathbf{z}^{\prime}=\mathbf{z}+\epsilon\cdot\mathbf{v}, where v∼N(0,I)\mathbf{v}\sim\mathcal{N}(\mathbf{0},\mathbf{I}) and ϵ\epsilon is set as 2×10−42\times 10^{-4}. We expect the reverted image of z′\mathbf{z}^{\prime} to have the same background clean image, yet with a different noise. We visualize the reverted images of z\mathbf{z} and z′\mathbf{z}^{\prime} in 6(b) and 6(c), exhibiting visually different noise corresponding to z\mathbf{z} and z′\mathbf{z}^{\prime}.

To quantitatively evaluate the quality of our generated noisy images, we follow the average KL divergence (AKLD) metric introduced by DANet to measure the similarity between the original and the generated noise, as shown in Table 5. It should be noted here that Noise Flow requires raw-RGB images, ISO, and CAM information to generate noisy images; in other words, it needs the training images as inputs. The AKLD results in Table 5 demonstrate that our generated noisy images are closer to the original noisy images by a large margin.

This paper is the first to study real image denoising with invertible networks. In previous invertible models, the input and the reversed output follow the same distribution. However, for image denoising, the input is noisy, and the restored outcome is clean, following two different distributions. To address this issue, our proposed InvDN transforms the noisy input into a low-resolution clean image as well as a latent representation containing noise. As a result, InvDN can both remove and generate noise. For noise removal, we replace the noisy representation with a new one sampled from a prior distribution to restore clean images; for noise generation, we alter the noisy latent vector to reconstruct new noisy images. Extensive experiments on three real-noise datasets demonstrate the effectiveness of our proposed model in both removing and generating noise.

Acknowledgments: This work was partly supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.2019-0-01906, Artificial Intelligence Graduate School Program(POSTECH)).