Lightweight Pyramid Networks for Image Deraining

Xueyang Fu, Borong Liang, Yue Huang, Xinghao Ding, John Paisley

I Introduction

As a common weather condition, rain impacts not only human visual perception but also computer vision systems, such as self driving vehicles and surveillance systems. Due to the effects of light refraction and scattering, objects in an image are easily blurred and blocked by individual rain streaks. When facing heavy rainy conditions, this problem becomes more severe due to the increased density of rain streaks. Since most existing computer vision algorithms are designed based on the assumption of clear inputs, their performance is easily degraded by rainy weather. Thus, designing effective and efficient algorithms for rain streak removal is a significant problem with many downstream uses. Figure 1 shows an example of our lightweight pyramid network.

Depending on the input data, rain removal algorithms can be categorized into video and single-image based methods.

We first briefly review the rain removal methods in a video, which was the major focus in the early stages of this problem. These methods use both spatial and temporal information from video. The first study on video deraining removed rain from a static background using average intensities from the neighboring frames . Other methods focus on deraining in the Fourier domain , using Gaussian mixture models , low rank approximations and via matrix completions . In , the authors divide rain streaks into sparse ones and dense ones, then a matrix decomposition based algorithm is proposed for deraining. More recently, proposed a patch-based mixture of Gaussians for rain removal in video. Though these methods work well, they require temporal content of video. In this paper we instead focus on the single image deraining problem.

I-A2 Single-image methods

Since information is drastically reduced in individual images, single image deraining is a much more difficult problem. Methods for addressing this problem have employed kernels , low rank approximations and dictionary learning . In , rain streaks are detected and removed by using kernel regression and a non-local mean filtering. In , the authors decompose a rainy image into its low- and high- frequency components. The high-frequency part is processed to extract and remove rain streaks by using sparse-coding based dictionary learning. In , a self learning method is proposed to automatically distinguish rain streaks from the high-frequency part. A discriminative sparse coding method is proposed in . By forcing the coefficient vector of rain layer to be sparse, the objective function is solved to separate background and rain streaks. Other methods have used mixture models and local gradients to model and then remove rain streaks. In , by utilizing Gaussian Mixture Models (GMMs), the authors explore patch-based priors for both the clean and rain layers. The GMM prior for background layers is learned from natural images, while that for rain streaks layers is learned from rainy images. In , three new priors are defined by exploring local image gradients. The priors are used to modeling the objective function which is solved by using alternating direction method of multipliers (ADMM).

Deep learning has also been introduced for this problem. Convolutional neural networks (CNN) have proven useful for a variety of high-level vision tasks as well as various image processing problems . In , a related work based on deep learning was introduced to remove static raindrops and dirt spots from pictures taken through windows. Our previous CNN-based method for removing dynamic rain streaks was introduced by . Here the authors build a relative shallow network with 3 layers to extract features of rain streaks from the high frequency content of a rainy image. Based on the introduction of an effective strategy for training very deep networks , two deeper networks were proposed based on image residuals and multi-scale information . In , the authors utilize the generative adversarial framework to further enhance the textures and improve the visual quality of de-rained results. Recently, in , a density aware multi-stream densely connected CNN is proposed for joint rain density estimation and de-raining. This method can automatically generate rain density label, which is further utilized to guide rain streaks removal.

I-B Our contributions

Though very deep networks achieve excellent performance on single image deraining, a main drawback that potentially limits their application in mobile devices, automatic driving, and other computer vision tasks is their huge number of parameters. As a networks become deeper, more storage space is required . To address this issue, we propose a lightweight pyramid network (LPNet), which contains fewer than 8K parameters, with the single image rain removal problem in mind. Instead of designing a complex network structure, we use problem-specific knowledge to simplify the learning process. Specifically, we first adopt Laplacian pyramids to decompose a degraded/rainy image into different levels. Then we use recursive and residual networks to build a sub-network for each level to reconstruct Gaussian pyramids of derained images. A specific loss function is selected for training each sub-network according to its own physical characteristics and the whole training is performed in a multi-task supervision. The final recovered image is the bottom level of the reconstructed Gaussian pyramid.

The main feature of our LPNet approach is to use the mature Gaussian-Laplacian image pyramid technique to transform one hard problem into several easier sub-problems. In other words, since the Laplacian pyramid contains different levels that can differentiate large scale edges from small scale details, one can design simple and lightweight sub-network to handle each level in a divide-and-conquer way. The contributions of our paper are summarized as follows:

We show how by combining the classical Gaussian-Laplacian pyramid technique with CNN, a simple network structure with few parameters and relative shallow depth is sufficient for excellent performance. To our knowledge the resulting network is far more lightweight (in terms of parameters) among deep networks with comparable performance.

Through multi-scale techniques and recursive and residual deep learning, our proposed network achieves state-of-the-art performances on single image deraining. Although LPNet is trained on synthetic data by necessity, it still generalizes well to real-world images.

We discuss how LPNet can be applied to other fundamental low- and high-level vision tasks in image processing. We also show how LPNet can improve downstream applications such as object recognition.

II Lightweight pyramid network for deraining

In Figure 2, we show our proposed LPNet for single image deraining. To summarize at a high level, we first decompose a rainy image into a Laplacian pyramid and build a sub-network for each pyramid level. Then each sub-network is trained with its own loss function according to the specific physical characteristics of the data at that level. The network outputs a Gaussian pyramid of the derained image. The final derained result is the bottom level of the Gaussian pyramid .

Since rain streaks are blended with object edges and the background scene, it is hard to directly learn the deraining function in the image domain . To simplify the problem, it is natural to train a network on the high-frequency information in images, which primarily contain rain streaks and edges without background interference. Based on this motivation, the authors in use the guided filter to obtain the high-frequency component of an image as the input to a deep network, which is then derained and fused back with the low-resolution information of the same image. However, these two methods fail when very thick rain streaks cannot be extracted by the guided filter. Inspired by this decomposition idea, we instead build a lightweight pyramid of networks to instead simplify the learning processing and reduce the number of necessary parameters as a result.

II-B Stage 1: The Laplacian pyramid

We first decompose a rainy image X\bf{X} into its Laplacian pyramid, which is a set of images LL with NN levels:

where Gn{G_{n}} is the Gaussian pyramid, n=1,...,N−1n=1,...,N-1. The function Gn(X){G_{n}}({\bf{X}}) is computed by downsampling Gn−1(X){G_{n-1}}(\bf{X}) using a Gaussian kernel, with G1(X)=X{G_{1}}(\bf{X})=\bf{X} and LN(X)=GN(X){L_{N}}({\bf{X}})={G_{N}}(\bf{X}).

The reasons we choose the classical Laplacian pyramid to decompose the rainy image are fourfold: 1) The background scene can be fully extracted at the top level of Ln{L_{n}} while the other levels contain rain streaks and details at different spatial scales. Thus, the rain interference is removed and each sub-network only needs to deal with high-frequency components at a single scale. 2) This decomposition strategy will allow the network to take advantage of the sparsity at each level, which motivates many other deraining methods , to simplify the learning problem. However, unlike previous deraining methods that use a single-scale decomposition, LPNet performs a multi-scale decomposition using Laplacian pyramids. 3) As shown in Figure 3, compared with the image domain, deep learning at each pyramid level is more like an identity mapping (e.g., the top row is more similar to the middle row, as evident in the bottom row) which is known to be the situation where residual learning (ResNet) excels . 4) The Laplacian pyramid is a mature algorithm with low computation cost. Most calculations are based on convolutions (Gaussian filtering) which can be easily embedded into existing systems with GPU acceleration.

II-C Stage 2: Sub-network structure

After decomposing X\bf{X} into different pyramid levels, we build a set of sub-networks independently for each level to predict a corresponding clean Gaussian pyramid G(Y)G(\bf{Y}). All the sub-networks have the same network structure with different numbers of kernels. We adopt residual learning for each network structure and recursive blocks to reduce parameters. The sub-network structure can be expressed as follows:

The first layer extracts features from the nnth input level,

where H\bf{H} indexes the feature map, ∗* is the convolution operation, W\bf{W} are weights and b\bf{b} are biases. σ\sigma is an activation function for non-linearity.

To reduce the number of parameters, we build intermediate inference layers in a recursive fashion. The basic idea is to share parameters among recursive blocks. Motivated by our experiments, we adopt three convolutional operations in each recursive block. Calculations in the ttth recursive block are

where F{1,2,3}{\bf{F}}^{\left\{{1,2,3}\right\}} are intermediate features in the recursive block, W{1,2,3}{\bf{W}}^{\left\{{1,2,3}\right\}} and b{1,2,3}{\bf{b}}^{\left\{{1,2,3}\right\}} are shared parameters among TT recursive blocks and t=1,...,Tt=1,...,T.

To help propagate information and back-propagate gradients, the output feature map Hn,t{\bf{H}}_{n,t} of the ttth recursive block is calculated by adding Hn,0{\bf{H}}_{n,0}:

To obtain the output level of the pyramid, the reconstruction layer is expressed as:

After obtaining the output of the Laplacian pyramid L(Y){L}({\bf{Y}}), the corresponding Gaussian pyramid of the derained image can be reconstructed by

where n=1,...,N−1n=1,...,N-1. Since each level of a Gaussian pyramid should equal or lager than 0, we use x=max(0,x)x=max(0,x), which is actually the rectified linear units (ReLU) operation , to simply correct the outputs. The final derained image is the bottom level of the Gaussian pyramid, i.e., G1(Y){G_{1}}(\bf{Y}).

In methods , the authors build similar networks based on the image pyramid, which are the most related to our own work. However, these papers apply similar structures to other tasks such a image generation or super-resolution using different network approaches on the pyramid.

II-D Loss Function

II-E Removing batch normalization

As one of the most effective way to alleviate the internal co-variate shift, batch normalization (BN) is widely adopted before the nonlinearity in each layer in existing deep learning based methods. However, we argue that by introducing image pyramid technology, BN can be removed to improve the flexibility of networks. This is because BN constrains the feature maps to obey a Gaussian distribution. While during our experiments, we found that distributions of lower Laplacian pyramid levels of both clean and rainy images are sparse. To demonstrate this viewpoint, in Figure 4, we show the histogram distributions of each Laplacian pyramid level from 200 clean and light rainy training image pairs from . As can be seen, compared to the image domain in Figure 4(a), distributions of lower pyramid levels, i.e., Figures 4(c) to (f), are more sparse and do not obey Gaussian distribution. This implies that we do not need BN to further constrain the feature maps since the mapping problem already becomes easy to handle. Moreover, removing BN can sufficiently reduce GPU memory usage since the BN layers consume the same amount of memory as the preceding convolutional layers. Based on the above observation and analysis, we remove BN layers from our network to improve flexibility and reduce parameter numbers and computing resource.

II-F Parameter settings

We decompose an RGB image into a 55-level Laplacian pyramid by using a fixed smoothing kernel [0.0625,0.25,0.375,0.25,0.0625][0.0625,0.25,0.375,0.25,0.0625], which is also used to reconstruct the Gaussian pyramid. In our network architecture, each sub-network has the same structure with a different numbers of kernels. The kernel sizes for W{0,1,3,4}{\bf{W}}^{\left\{{0,1,3,4}\right\}} are 3×33\times 3. For W{2}{\bf{W}}^{\left\{{2}\right\}}, the kernel size is 1×11\times 1 to further increase non-linearity and reduce parameters. The number of recursive blocks is T=5T=5 for each sub-network. For the activation function σ\sigma, we use the leaky rectified linear units (LReLUs) with a negative slope of 0.2.

Moreover, as shown in the last row of Figure 3, higher levels are closer to an identity mapping since rain streaks only remain in lower levels. This means for higher levels, fewer parameters are required for learning a good network. Thus, from low to high levels, we set the kernel numbers to 16,8,4,216,8,4,2 and 11, respectively. Since the top level is a tiny and smoothed version of image and rain streaks remain in high-frequency parts, the function of top level sub-network is more like a simple global contrast adjustment. Thus we set the kernel numbers to 11 kernel for the top level. As shown in Figure 2, by connecting the up-sampled version of the output from the higher level, the direct prediction of all sub-networks is actually the clean Laplacian pyramid. We show the intermediate results predicted by each sub-network in Figure 5. It is clear that rain streaks remain in lower levels while higher levels are almost the same. This demonstrates that our diminishing parameter setting is reasonable. As a result, the total number of trainable parameters is only 7,5487,548, far fewer than the hundreds of thousands often encountered in deep learning.

II-G Training details

We use synthetic rainy images from as our training data. This dataset contains 1800 images with heavy rain and 200 images with light rain. We randomly generate one million 80×8080\times 80 clean/rainy patch pairs. We use TensorFlow to train LPNet using the Adam solver with a mini-batch size of 1010. We set the learning rate as 0.0010.001 and finish the training after 33 epochs. The whole network is trained in a end-to-end fashion.

III Experiments

We compare our LPNet with four state-of-the-art deraining methods: the Gaussian Mixture Model (GMM) of , a CNN baseline SRCNN , the deep detail network (DDN) of and joint rain detection and removal (JORDER) , which is also a deep learning method. For fair comparison, all CNN based methods are retrained on the same training dataset.

Three synthetic datasets are chosen for comparison. Two of them are from and each one contains 100100 images. One is synthesized with heavy rain called Rain100H and the other one is with light rain called Rain100L. The third dataset called Rain12 is from which contains 1212 synthetic images. All testing results shown are not included in the training data. Following , for each CNN method we train two models, one for heavy and for light rain datasets. The model trained on the light rainy dataset is used to test Rain12.

Figures 6 to 8 shows visual results from each dataset. As can be seen, GMM fails to remove rain streaks form heavy rainy images. SRCNN and DDN are able to remove the rain streaks while tend to generate obvious artifacts. Our LPNet has comparable visual results with JORDER and outperforms other methods.

We also adopt PSNR and SSIM to perform quantitative evaluations in Table I. Our method has comparable SSIM values with JORDER while outperforming other methods, in agreement with the visual results. Though our result has a lower PSNR value than JORDER method, the visual quality is comparable. This is because PSNR is calculated based on the mean squared error (MSE), which measures global pixel errors without considering local image characters. Moreover, as shown in Table I our LPNet contains far fewer parameters, potentially making LPNet more suitable for storage, e.g., in mobile devices.

III-B Real-world data

In this section, we show that the LPNet learned on synthetic training data still performs well on real-world data. Figure 9 shows five visual results on real-world images. The model trained on the dataset with light rain is used for testing on real-world images. As can be seen, LPNet generates consistently promising derained results on images with different kinds of rain streaks.

Since no ground truth exists, we construct an independent user study to provide realistic feedback and quantify the subjective evaluation. We collect 50 real-world rainy images from the Internet as a new dataset Our code and data will be released soon.. We use the compared five methods to generate de-rained results and randomly order the outputs, as well as the original rainy image, and display them on a screen. We then separately asked 20 participants to rank each image from 1 to 5 subjectively according to quality, with the instructions being that visible rain streaks should decrease the quality and clarity should increase quality (1 represents the worst quality and 5 represents the best quality). We show the average scores in Table II from these 1,000 trials and our LPNet has the best performance. In Figure 10, we show the scatter plot of the rainy inputs vs de-rained user scores. This small-scale experiment gives additional support that our LPnet improves the de-raining on real-world images.

Moreover, when dealing with dense rain, LPNet trained on images with heavy rain has a dehazing effect as shown in Figure 11, which can further improve the visual quality. This is because the highest level sub-network (low-pass component) can adjust image contrast. Although dehazing is not the main focus of this paper, we believe that LPNet can be easily modified for joint deraining and dehazing.

III-C Running time and convergence

To demonstrate the efficiency of LPNet, we show the average running time for a test image in Table III. Three different image sizes are chosen and each one is tested over 100 images. The GMM is implemented on CPUs according to the provided code, while other deep CNN-based methods are tested on both CPU and GPU. All experiments are performed on a server with Intel(R) Xeon(R) CPU E5-2683, 64GB RAM and NVIDIA GTX 1080. The GMM has the slowest running time since complicated inference is required to process each new image. Our method has a comparable and even faster computational time on both CPU and GPU compared with other deep models. This is because LPNet uses relatively shallow networks for each level, so requires fewer convolutions.

We also show the average training loss as a function of training epoch in Figure 12. We observe that LPNet converges quickly on training with both light and heavy rainy datasets. Since heavy rain streaks are harder to handle, as shown in the 1st row of Figure 6, the training error of heavy rain streaks has a vibration.

III-D Parameter settings

In this section, we discuss different parameters setting to study their impact on performance.

We have conducted an experiment on the Rain100H dataset with increased parameters, i.e., 16 feature maps for all convolution layers at each sub-network. The results are shown in Table IV. As can be seen, the SSIM evaluation is better than JORDER and PSNR value is also improved. We believe that the performance can be further improved by using more parameters. However, increasing parameter number requires more storage and computing resources. Figure 13 shows one example by using different parameter numbers. As can be seen, the visual quality is almost the same. Thus, we use our diminishing parameter setting to achieve the balance between effectiveness and efficiency.

III-D2 Skip connections

Though Laplacian pyramid images introduce sparsity in each level to simply the mapping problem, it is still essential to add skip connection in each sub-network. We adopt skip connection for two reasons. First, image information may be lost during feed-forward convolutional operations, using skip connection helps to propagate information flow and improve the deraining performance. Second, using skip connection helps to back-propagate gradient, which can accelerate the training procedure, when updating parameters. In Figure 14 we show the training curves on the heavy rainy dataset with and without all skip connections. As can be seen, using skip connection can bring a faster convergence rate and lower training loss.

III-D3 Loss function

III-E Extensions

Since both Laplacian pyramids and CNNs are fundamental and general image processing technologies, our network design has potential value for other low-level vision tasks. Figure 16 shows the experimental result on image denoising and JPEG artifacts reduction, which shares the property of rainy images in that the desired image is corrupted by high frequency content. This test demonstrates that LPNet can generalize to similar image restoration problems.

III-E2 Pre-processing for high-level vision tasks

Due to the lightweight architecture, our LPNet can potentially be efficiently incorporated into other high-level vision systems. For example, we study the problem of object detection in rainy environments. Since rain steaks can blur and block objects, the performance of object detection will degrade in rainy weather. Figure 17 shows a visual result of object detection by combining with the popular Faster R-CNN model . It is obviously that rain streaks can degrade the performance of Faster R-CNN, i.e., by missing detections and producing low recognition confidence. On the other hand, after deraining by LPNet, the detection performance has a notable improvement over the naive Faster-RCNN.

Additionally, due to the lightweight architecture, using LPNet with Faster R-CNN does not significantly increase the complexity. To process a color image with size of 1024×10241024\times 1024, the running time is 3.7 seconds for Faster R-CNN, and 4.0 seconds for LPNet + Faster R-CNN.

IV Conclusion

In this paper, we have introduced a lightweight deep network that is based on the classical Gaussian-Laplacian pyramid for single image deraining. Our LPNet contains several sub-networks and inputs the Laplacian pyramid to predict the clean Gaussian pyramid. By using the pyramid to simplify the learning problem and adopting recursive blocks to share parameters, LPNet has fewer than 88K parameters while still achieving good performance. Moreover, due to the generality and lightweight architecture, our LPNet has potential values for other low- and high-level vision tasks.

References