Spatially-Adaptive Image Restoration using Distortion-Guided Networks

Kuldeep Purohit, Maitreya Suin, A. N. Rajagopalan, Vishnu Naresh Boddeti

Introduction

Images are often degraded during the data acquisition process, especially under non-ideal imaging conditions. Such degradations can be attributed to the medium and dynamics between the camera, scene elements and the illumination. For instance, as shown in Fig. 1 , (1) precipitation leads to snow/rain streaks occupying the volume between the scene and the camera, (2) presence of rain-drops on the camera lens causes significant degradation in scene visibility, (3) relative motion between the camera or scene elements results in motion blur, and (4) harsh illumination conditions can induce harsh shadows. Despite the disparate source of degradations, they share the same underlying motif that affects the image quality, namely, degradation that is spatially-varying in nature. For example, raindrops and shadows degrade monolithic parts of the image depending on their size and location, motion blur varies with scene depth and degree of motion, and rain streaks effect only sparse regions whose orientation depends on the relative rain direction. Fig. 1 shows representative examples of degraded images and respective distortion-maps. It can be seen that a large number of pixels undergo little or no distortion. Another observation is that the amount of distortion and its spatial distribution is different in every image.

Restoring such images is vital to improve their aesthetic quality as well as the performance of downstream tasks, viz, detection, segmentation, classification, and tracking. Convolutional neural networks (CNNs) currently typify the state-of-the-art for various image restoration tasks. Despite recent progress, existing approaches share several key limitations. Firstly, all layers in their networks are generic CNN layers, which apply the same set of spatially-invariant filters to every degraded image. Such layers are limited in their ability to invert degradations that are highly image-dependent and spatially-varying. Secondly, most network architectures are specifically tailored for individual degradation types as they are based on image formation models. Thirdly, the distortion-localization information embedded in the labeled datasets remains unused or sub-optimally used in all existing solutions.

Static CNN based models trained to directly regress clean intensities from degraded ones, perform poorly when input contains unaffected regions as well as severe intensity distortions in different spatial regions. Conceptually, a stack of fixed learned filters, excelling at restoring pixels degraded with large distortions might not be suitable for reconstructing the texture from unaffected regions. Practically, we observe that such designs often yield poor reconstruction performance (introduce unwanted changes or artifacts on pixels that are not degraded in the input to begin with). The image-dependent nature of spatial distribution and magnitude of distortions only exacerbates the problem faced by static CNNs.

Motivated from the understanding that a restoration network can benefit from adapting to the degradation present in each test image, we propose a distortion-aware model to simultaneously realize the twin goals of restoration and reconstruction. Our spatially-adaptive image restoration architecture (referred to as SPAIR) is suited for any type of degradation which selectively affects parts of the image. It comprises of two components- a distortion-localization network (NetLNet_{L}) and a spatially-guided restoration network (NetRNet_{R}). NetLNet_{L} gathers information from the entire image to estimate a binary mask (localizing high intensity distortions) which steers the processing in NetRNet_{R} to selectively improve only degraded regions.

The proposed NetRNet_{R} comprises of 3 distortion-guided blocks- spatial feature modulator (SFM), sparse convolution module (SC) and a custom sparse non-local module (SNL). SFM utilizes the output mask and intermediate features from NetLNet_{L} to modulate the feature statistics of intermediate features in NetRNet_{R}. SC and SNL improve features in the spatially-sparse degraded regions in an image-dependent manner, without affecting the features in clean regions. SNL locally restores features in distorted regions by adaptively gathering global context from all clean regions. Our key contributions are:

– A two-stage framework to systematically exploit distortion-localization knowledge for directly addressing the challenges associated with diverse spatially-varying degradations in an interpretable manner. It achieves the twin goals of restoration and reconstruction and works across diverse degradation-types.

– Distortion-guided spatially-varying modulation of features statistics in NetRNet_{R} with the help of distortion-mask and features from a pretrained NetLNet_{L}.

– Distortion-guided feature extraction with the help of SC (for local context) and a novel SNL (for global context) modules. These components facilitate spatially-varying restoration while controlling receptive field in an image and location-adaptive manner.

– We demonstrate the versatility of SPAIR by setting new state-of-the-art on 11 synthetic and real-world datasets for various spatially-varying restoration tasks (removing rain-streaks, rain-drops, shadows, and motion blur), outperforming existing approaches designed with task-specific network-engineering. Further, we provide detailed analysis, qualitative results, and generalization tests.

Related Works

Adaptive Inference: Adaptive inference techniques have attracted increasing interest since they enable input-dependent alteration of CNN structure. One class of methods dynamically skip subsets of layers in cascaded CNNs during inference . passes sampled pixels (using a random pattern which is fixed during inference) to CNN layers and fills the remaining locations using simple interpolation. Few approaches exploit sparsity in the input image itself using sub-manifold sparse convolutions, but are unsuitable for non-sparse input data.

However, none of these approaches afford the fine-grained spatial-domain control necessary for spatially-varying image restoration at multiple intermediate layers. For instance, the approaches that skip processing of some layers or prune the network still filter the degraded and other image regions with the same parameters. Methods such as are only applicable to cascade of consecutive residual layers, and do not generalize to encoder-decoder designs (typically used for image restoration) where conditionally altering network depth or channel width is non-trivial. The arbitrary rejection of spatial-domain information proposed in is ill-fitted for general restorations tasks.

Raindrop Removal: Solutions for raindrop removal include both classical as well as CNN based approaches. proposed a clustering and median filtering based restoration, while CNN based approaches include, shallow CNNs but with limited performance, a convolutional-LSTM based model for “joint” learning of rain-map and rain-free image , and a deeper CNN . instead leveraged physical models of raindrop properties (including closedness and roundness) to estimate drop-probability. In contrast to these methods, SPAIR advocates for a pixel selective and adaptive processing to remove raindrops.

Rain-streak Removal: Conventional deraining methods adopt a model-driven methodology utilizing physical properties of rain and prior knowledge of background scenes into an optimization problem. CNN-based approaches include end-to-end negative residual mapping , deeper CNN , multi-stage CNNs with recurrent connections , CNN for predicting density (heavy, medium, light) during deraining , concatenating rain-map for deraining . However, layers in these approaches process all image regions with the same filters (without pixel adaptation). presents model-driven CNN with convolutional dictionary learning. Wang et.al. predicts a rain-map and multiplies it element-wise with feature-maps to enhance them. While SPAIR also utilizes a mask, there are fundamental differences. We estimate a binary mask and utilize it more comprehensively, including for sparse filtering, attention weight calculation and guiding it to non-degraded image regions. SPAIR significantly differs from rain-guided models of in three aspects. (1) They only concatenate the rain-mask at the input. In contrast, we exploit distortion-mask to only perform convolutions and non-local operations on degradation regions. We also transfer feature statistics from clean to degraded regions at multiple intermediate layers using SFM. (2) They lack global context. SPAIR contains SNL module that adaptively gathers all features values within the clean regions of the image. (3) All pixels are passed through same network with spatially rigid processing, which directly contrasts with our work. Ours is the first approach to exploit explicit degradation-guidance to selectively processes degraded pixels and reduce the effect on unaffected regions, for a variety of spatially-sparse degradations.

Shadow Removal: Early works often erased shadows via user interaction or by transferring illumination from non-shadow regions to shadow regions . More robust results have been achieved using CNN based approaches which include using multiple networks , DeshadowNet for illumination estimation in shadow regions , stacked conditional GANs , ARGAN to detect and remove shadow with multiple steps , RIS-GAN to estimate negative residual images and inverse illumination maps for restoration, and finally a cascade of dilated convolutions to jointly estimate shadow-mask and shadow-free image . In contrast to the aforementioned approaches, we propose a two-stage framework wherein the distortion-mask and intermediate learned features of NetLNet_{L} are employed in a principled manner for region-aware and selective restoration.

Motion Blur Removal: Traditional approaches designed priors on image and motion (eg. locally linear blur kernels , planar scenes ) but with limited success in general 3D and dynamic blurred scenes. Recent CNN-based methods directly estimate the latent sharp image , wherein encoder-decoder designs that aggregate features in a coarse-to-fine manner have been proposed . Additionally, explored a design composed of multiple CNNs and RNN and proposed a patch-hierarchical network and stacked its copies along depth to achieve state-of-the-art performance. proposed a recurrent design for efficient deblurring. A limitation shared by all of these methods is the absence of spatially varying adaptive layers. proposed direction-based feature extraction modules suited for efficient motion deblurring. inserts adaptive convolution and attention within the layers of to boost its results. Our distortion-guided sparse architecture performs better than such patch hierachical designs, while generalizing beyond motion-blur and offering consistent gains across other degradations.

Architectures for general Restoration A few solutions have been proposed in literature to address multiple degradation-types. For instance, DuRN make task dependent alterations in their network structure. Similarly, OWAN was proposed to handle multiple degradations present within the same image. However, only addresses simple synthetic degradations that are similar in nature eg. gaussian blur, noise, and jpeg artifacts. SPAIR demonstrates its efficacy on realistic datasets of several physically unrelated degradations which are heavily spatially-varying. In such settings, DuRN and OWAN are quite inferior to our model, as shown in our experiments.

Proposed Network Architecture

An image restoration model needs to solve two equally important tasks: (1) locating the areas to restore in an image, and (2) applying the right filtering mechanism to the corresponding regions. While NetLNet_{L} addresses the former, we realize the latter through a spatially-guided restoration network NetRNet_{R}. A schematic of SPAIR is shown in Fig. 2. The knowledge from intermediate features of pre-trained NetLNet_{L} improves NetRNet_{R}’s training, while the mask itself lends adaptiveness to the restoration process. To realize the twin goals of restoration and reconstruction, distortion-guided filtering of the extracted features in NetRNet_{R} is enabled through SFM (Spatial Feature Modulator), SC (Sparse Convolution), and SNL (Sparse Non Local) modules.

To maximize the generalizability of our approach, we adopt the U-Net topology as our CNN backbone (both for localization and restoration networks). Different versions of this are known to be effective for several restoration tasks such as image deblurring , denoising , and general image-to-image translation . We build a densely connected encoder-decoder structure whose detailed layer-wise description is given in the supplementary. This design delivers competitive performance across all tasks considered and hence acts as a backbone for our NetRNet_{R} (see Sec. 6). NetLNet_{L} is a lightweight version of NetRNet_{R} (with similar structure) since the binary classification (localization) task is simpler than the intensity regression (restoration) task.

Given a degraded image, NetLNet_{L} produces a single channel mask and is trained using binary cross-entropy loss to match the GT binary mask. For datasets with no ground truth mask, we use the absolute difference between degraded image and clean image, and threshold it to obtain a binary mask, classifying pixels into degraded (value 1) or clean (value 0). Empirically, we observed that NetRNet_{R}’s performance improves when NetLNet_{L} is trained to predict only the pixels with severe distortions (as opposed to detecting even minute intensity changes). Note that the distortion-map directly correlates with the difficulty of restoration, and it may differ from the physically occuring degradation-distribution. For instance, when physical rain-steaks are equally distributed throughout the image, the distortion-map would contain more non-zero values in the urban textured regions than the sky regions (since white rain-streaks do not significantly alter the bright intensities in sky).

As depicted in Fig. 2, NetRNet_{R} extracts a features pyramid from input degraded image using a cascade of densely connected layers . These features are fed to decoder that generates the restored image. Although correction of small intensities changes can be learnt by simple convolutional layers (basic building block for all prior works), they struggle for spatially-distributed heavy degradations. For such regions, localization based guidance improves restoration quality. We propose 3 modules to employ trained NetLNet_{L} to convey localization knowledge to NetRNet_{R}.

Since image generation process requires decoder to learn both reconstruction and restoration, each level of the decoder contains an SC and an SNL module. Note that we refrain from using SC or SNL in the encoder layers of NetRNet_{R} since that would completely discard the degraded image intensities (which contain partially-useful information). We employ SFM at multiple levels to perform distortion-guided feature normalization to complement SC and SNL modules.

SFM fuses the features of NetRNet_{R} with intermediate features from layers of the pretrained NetLNet_{L} in an additive manner. We observe that with such feature guidance, early layers of NetRNet_{R} extract more distortion-aware features that correlate strongly to the degradation-variation within the input image. Since both the networks share a similar encoder-decoder structure, the inputs of all strided convolution layers are fused using SFM, as shown in Fig. 2.

In CNNs, feature normalization is known to be important and complementary to feature extraction. The role of SFM is to perform distortion-guided spatial-varying feature normalization. This complements the distortion-guided feature extraction process using local (SC) and global (SNL) context. SFM module performs adaptive shifting of the feature statistics at degraded locations, which aids the restoration process. Studies show that feature mean relates to global semantic information while variance is correlated to local texture. Inspired from this, our SFM modulates features at degraded locations to match the feature statistics (mean and variance) of clean regions.

Given the fused features FF and the predicted mask M\mathcal{M}, we calculate the modulated features FSF^{S} as

2.2 Mask-guided Sparse Convolution (SC)

As discussed earlier, filters of general convolution layers are spatially-invariant and hence are forced to learn the restoration and reconstruction tasks jointly, which impedes the training process and reduces model’s performance. SPAIR harnesses the efficacy of mask-guided sparse convolution that facilitates selective restoration of highly degraded regions, and simplifies the learning process. SC (shown in Fig. 2) contains a densely connected set of 66 guided sparse convolution layers followed by a 1×11\times 1 convolution to reduce the number of channels. Each unit in SC takes the input feature map, FF and the predicted mask M\mathcal{M}. Pixels masked as 11 in M\mathcal{M} are sampled, and passed through a convolution operation, resulting in a sparse feature map FSF^{S} as

2.3 Region-Guided Sparse Non-Local Module (SNL)

Most computer vision tasks are inherently contextual. Frequently used tools for gathering larger context such as dilated convolution , global average or attention pooling , or multi-scale methods etc. can enlarge the receptive field beyond simple convolutional layers. Yet, they are not image-adaptive and still cannot utilize the full feature map effectively. In contrast, a single non-local layer is capable of extending the receptive field to the maximum H×WH\times W size adaptively for each image and each pixel in an image. We claim that such a property is well-suited for a spatially-varying restoration model, where the heavily corrupted regions benefit from the ability to gather relevant features from the whole image. Effectiveness of adaptive global context aggregation has also been explored for recognition/segmentation tasks in .

We hypothesize that within a restoration model, the non-local context aggregation process can benefit immensely from the knowledge of degraded pixel locations. We intuit that propagating heavily degraded information throughout the spatial domain can be counter-productive. Ideally, an adaptive module should learn to completely ignore irrelevant features, but recent vision models (e.g. image captioning ) have shown that this behavior is not practically achieved. They resort to perform additional filtering to remove unnecessary information.

In contrast, the proposed SNL module leverages the distortion mask to control the scope of non-local context aggregation. While restoring degraded pixels, it assigns dynamically estimated non-zero weights to features from only not/less degraded pixel locations, delivering superior performance as it dampens the influence of heavily degraded/corrupted information. Moreover, SNL leaves the features of clean regions unaltered as this operation is performed only on degraded pixel locations.

As illustrated in Fig. 3, SNL comprises of an efficient two-step aggregation approach, with each step comprising of horizontal-vertical scanning of the feature matrix in four fixed directions: left-to-right, top-to-bottom, and vice-versa. Having two steps is important in harvesting full-image contextual information from all pixels. While directional scanning of CNN features has been explored in literature , SPAIR introduces a region-guided and sparse non-local module.

This matrix is then used to weigh the contribution of the features towards the right (Fi,jright\mathbf{F}_{i,j}^{right}) as

Sparse 1×11\times 1 convolution: To perform feature-refinement between two steps of the SNL module, we introduce a sparse 1×11\times 1 convolution. As shown in the subfigure Fig. 3, on the feature locations of interest specified by the binary mask, a point-wise feature representation is extracted. A fully connected layer then accepts and refines the entire stack of these point-wise features. This replaces the 2D convolution on H×WH\times W spatial grid with point-wise 1D convolution on the selected points and facilitates sparse processing.

Datasets and Implementation Details

Rain-Streaks: Using the same experimental setups of the recent approaches on image deraining , we train our model on 1313,712712 clean-rain image pairs gathered from multiple datasets . With this single trained model, we perform evaluation on different test sets, including Rain100H , Rain100L , Test100 , Test2800 , and Test1200 . We also report the error reduction error for each method relative to the best method by translating PSNR to RMSE (RMSE∝10−PSNR/10\textrm{RMSE}\propto\sqrt{10^{-\textrm{PSNR}/10}}) and SSIM to DSSIM (DSSIM=(1−SSIM)/2\textrm{DSSIM}=(1-\textrm{SSIM})/2). We also evaluate SPAIR on SPANet Dataset (real-world rain) containing 2×1052\times 10^{5} training and 10001000 testing images.

Raindrop: We use the AGAN dataset with 861 training and 5858 test samples. Images were generated by placing a raindrop covered glass between the camera and scene.

Shadow: We evaluate our model using a challenging benchmark ISTD containing 13001300 (train) and 540540 (test) images (with real shadows and diverse textured scenes).

Motion Blur: We follow the configuration of and use the GoPro dataset containing 22,103103 image pairs for training and 11,111111 pairs for evaluation. Furthermore, to demonstrate generalizability, we directly evaluate our GoPro trained model on the test set of HIDE and RealBlur datasets. The HIDE dataset is specifically collected for human-aware motion deblurring, containing 22,025025 test images. While the GoPro and HIDE datasets are generated by averaging real videos, the blurred images in RealBlur-J dataset are captured in real-world conditions.

Implementation Details: The NetRNet_{R} for each degradation is trained to minimize l1l_{1} reconstruction loss between the output and the GT clean image. NetLNet_{L} is trained using binary cross entropy loss with respect to the GT binary mask. Each training batch contains randomly cropped RGB patches of size 256×256256\times 256 from degraded images that are randomly flipped horizontally or vertically. The batch-size was 88 for rain-streak, raindrop, and shadow removal and 1616 for deblurring. Both networks use Adam optimizer with initial leaning rate 2×10−42\times 10^{-4}, halved after every 5050 epochs. We use PyTorch library and RTX 2080Ti GPU.

Experimental Evaluation

Rain-streak Removal: Following prior art , we perform quantitative evaluations (PSNR/SSIM scores) on the Y channel (in YCbCr color space). Table 1 reports the results across all five datasets where SPAIR consistently achieves significant gains over the baselines. Compared to the recent algorithm MSPFN , we obtain a performance gain of 2.162.16 dB (averaged across all datasets). Next, for fair comparison with RCDNet , we evaluate SPAIR in their setting (in Table 2) by training and testing on the challenging Rain100H and SPANet (captured in real-world rainy scenes) datasets. While the improvement is modest (0.41 dB) on very heavy rain (Rain100H), it is as large as 3{3} dB on datasets with low rain density, eg. SPANet and Rain100L (since in this case, we selectively process very few pixels without affecting clean pixels), highlighting the advantage of our distortion-adaptive restoration.

Fig. 4 presents qualitative comparisons on challenging images from Rain100H dataset. Our results exhibit significantly higher visual quality than existing methods which fail to recover background textures (1st row) and introduce artifacts (2nd row). SPAIR is robust to changes in scenes and rain densities as it effectively removes rain streaks of different orientations and magnitudes, and generates images that are visually pleasing and faithful to the ground-truth.

Raindrop Removal: Table 3 and Fig. 5 show qualitative and visual comparisons with recent methods . SPAIR outperforms the baselines by a large margin. Our results are visually closer to GT and perceptually better than those of competing methods which often contain artifacts or color distortions.

Shadow Removal: We evaluate our shadow-removal model against traditional and learning based methods including ST-CGAN , DSC , DeShadowNet . Following prior art, results are evaluated in Lab color space using RMSE scores calculated over shadow and non-shadow regions. Fig. 6 and Table 4 show that although CNN-based designs are better than hand-crafted methods, most existing approaches produce shadow boundaries or color inconsistencies. However, SPAIR has minimal artifacts in the shadow boundaries, outperforming the baselines both qualitatively and quantitatively.

Deblurring: We validate our distortion-guided approach for general motion deblurring on 3 benchmarks: GoPro , HIDE , and the real-world blurred images of a recent RealBlur-J . We report the quantitative comparisons with the existing deblurring approaches in Table 5. Overall, SPAIR performs favorably against other algorithms. Note that inspite of training only on the GoPro, it outperforms all methods including on HIDE, without requiring any human bounding box supervision, thereby demonstrating its strong generalization capability.

We evaluate models on RealBlur-J testset under two experimental settings: 1) training on GoPro (to test generalization to real images), and 2) training on RealBlur-J. SPAIR obtains performance gain of 0.390.39 dB over the DMPHN model in setting 11, and 0.440.44 dB over existing best method for setting 22. Our model’s effectiveness is owed to the robustness of the distortion-aware approach.

Visual comparisons on images containing dynamic and 3D scenes are shown in Fig. 7. Often, the results of prior works suffer from incomplete deblurring or artifacts. In contrast, our network demonstrates non-uniform deblurring capability while preserving sharpness. Scene details in the regions containing text, boundaries, and textures are more faithfully restored, making them recognizable.

Network Analysis

This work explores the benefits of distortion-localization guided feature modulation and sparse processing for spatially-varying restoration tasks. Table 6 quantifies the effect of individual design choices on performance of SPAIR on the AGAN (raindrop) and GoPro (motion blur) datasets.

To validate our design choices, we implement the following baselines (reported in Table 6). Net1: Dense encoder-decoder network (CNN backbone of our NetRNet_{R}) with few additional parameters to match NetLNet_{L}. Net2: Net1 guided by NetLNet_{L} using SFM. Net3: Net2 with all densely connected convolutional blocks in decoder replaced with SC modules, Net4: Net3 with non-local (NL) layer introduced in the decoder. Net5: Net4 containing the proposed SNL module instead of NL. Good baseline scores of Net1 for both tasks support our backbone design choice.

Effectivenes of SFM: Net2 introduces SFM blocks (Sec. 3.2.1) which guide restoration network using mask and features of NetLNet_{L} at multiple intermediate levels. The significant improvement in accuracy in comparison to Net1 demonstrates the benefit of degradation guidance.

Effect of SC and SNL modules: Net4 employs the general non-local layer in decoder global context aggregation. Net5 has the same structure as Net4 (sans the NL module), and it feeds the predicted mask as input to the SNL modules. The improvement in behavior and performance is attributed to SNL design which uses explicit distortion-guidance to steer pixel-attention. SNL is more suited than NL for both degraded and clean regions. As explained in Eq. 4, while restoring degraded pixels, SNL assigns dynamically estimated non-zero weights to features originating from only clean pixels in the image. By design, it leaves the features of clean regions unaltered. As reported in Table 6, Net3 vs. Net2 shows the benefits of SC module whereas, Net5 vs. Net3 shows the utility of global context aggregation for restoration. Net5, our final model, shows a significant improvement over CNN baseline (Net1), demonstrating the advantages of our overall solution over static CNNs.

Supplementary Details: We provide additional real results and qualitative comparisons for all four tasks, additional model analysis, and results of multi-task learning in the supplementary document.

Benefit: Many applications (e.g., autonomous vehicles) involve dealing with rain, shadows, blur etc. at different time instances. Designing architectures that are applicable across multiple tasks, without requiring specialized architecture re-engineering is practically very convenient (potentially facilitating customized hardware design). Our versatile design enables this as only the learned weights vary across degradations while the architecture remains the same.

Conclusions

We addressed the single image restoration tasks of removing spatially-varying degradations such as raindrop, rain streak, shadow, and motion blur. We model the restoration task as a combination of degraded-region localization and region guided sparse restoration and propose a guided image restoration framework SPAIR which leverages the features of NetLNet_{L} for spatial modulation of the features in NetRNet_{R} using SFM module. We introduce distortion localization awareness in NetRNet_{R} using sparse convolution module (SC) and sparse non-local attention module (SNL) and show its significant benefits . Extensive evaluation on 1111 datasets across four restoration tasks demonstrates that proposed framework outperforms strong degradation-specific baselines. Ablation analysis and visualizations are shown to validate the effectiveness of its key components.

References