CycleISP: Real Image Restoration via Improved Data Synthesis

Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang, Ling Shao

Introduction

High-level computer vision tasks, such as image classification, object detection and segmentation have witnessed significant progress due to deep CNNs . The major driving force behind the success of CNNs is the availability of large-scale datasets , containing hundreds of thousands of annotated images. However, for low-level vision problems (image denoising, super-resolution, deblurring, etc.), collecting even small datasets is extremely challenging and non-trivial. For instance, the typical procedure to acquire noisy paired data is to take multiple noisy images of the same scene and generate clean ground-truth image by pixel-wise averaging. In practice, spatial pixels misalignment, color and brightness mismatch is inevitable due to changes in lighting conditions and camera/object motion. Moreover, this expensive and cumbersome exercise of acquiring image pairs needs to be repeated with different camera sensors, as they exhibit different noise characteristics.

Consequently, single image denoising is mostly performed in synthetic settings: take a large set of clean sRGB images and add synthetic noise to generate their noisy versions. On synthetic datasets, existing deep learning based denoising models yield impressive results, but they exhibit poor generalization to real camera data as compared to conventional methods . This trend is also demonstrated in recent benchmarks . Such behavior stems from the fact that deep CNNs are trained on synthetic data that is usually generated with the Additive White Gaussian Noise (AWGN) assumption. Real camera noise is fundamentally different from AWGN, thereby causing a major challenge for deep CNNs .

In this paper, we propose a synthetic data generation approach that can produce realistic noisy images both in RAW and sRGB spaces. The main idea is to inject noise in the RAW images obtained with our learned device-agnostic transformation rather than in the sRGB images directly. The key insight behind our framework is that the real noise present in sRGB images is convoluted by the series of steps performed in a regular image signal processing (ISP) pipeline . Therefore, modeling real camera noise in sRGB is an inherently difficult task as compared to RAW sensor data . As an example, noise at the RAW sensor space is signal-dependent; after demosaicking, it becomes spatio-chromatically correlated; and after passing through the rest of the pipeline, its probability distribution not necessarily remains Gaussian . This implies that the camera ISP heavily transforms the sensor noise, and therefore more sophisticated models that take into account the influence of imaging pipeline are needed to synthesize realistic noise than uniform AWGN model .

In order to exploit the abundance and diversity of sRGB photos available on the Internet, the main challenge with the proposed synthesis approach is how to transform them back to RAW measurements. Brooks et al. present a technique that inverts the camera ISP, step-by-step, and thereby allows conversion from sRGB to RAW data. However, this approach requires prior information about the target camera device (e.g., color correction matrices and white balance gains), which makes it specific to a given device and therefore lacks in generalizability. Furthermore, several operations in a camera pipeline are proprietary and such black boxes are very difficult to reverse engineer. To address these challenges, in this paper we propose a CycleISP framework that converts sRGB images to RAW data, and then back to sRGB images, without requiring any knowledge of camera parameters. This property allows us to synthesize any number of clean and realistic noisy image pairs in both RAW and sRGB spaces. Our main contributions are:

Learning a device-agnostic transformation, called CycleISP, that allows us to move back and forth between sRGB and RAW image spaces.

Real image noise synthesizer for generating clean/noisy paired data in RAW and sRGB spaces.

A deep CNN with dual attention mechanism that is effective in a variety of tasks: learning CycleISP, synthesizing realistic noise, and image denoising.

Algorithms to remove noise from RAW and sRGB images, setting new state-of-the-art on real noise benchmarks of DND and SIDD (see Fig. 1). Moreover, our denoising network has much fewer parameters (2.6M) than the previous best model (11.8M) .

CycleISP framework generalizes beyond denoising, we demonstrate this via an additional application i.e., color matching in stereoscopic cinema.

Related Work

The presence of noise in images is inevitable, irrespective of the acquisition method; now more than ever, when majority of images come from smartphone cameras having small sensor size but large resolution. Single-image denoising is a vastly researched problem in the computer vision and image processing community, with early works dating back to 1960’s . Classic methods on denoising are mainly based on the following two principles. (1) Modifying transform coefficients using the DCT , wavelets , etc. (2) Averaging neighborhood values: in all directions using Gaussian kernel, in all directions only if pixels have similar values and along contours .

While these aforementioned methods provide satisfactory results in terms of image fidelity metrics and visual quality, the Non-local Means (NLM) algorithm of Buades et al. makes significant advances in denoising. The NLM method exploits the redundancy, or self-similarity present in natural images. For many years the patch-based methods yielded comparable results, thus prompting studies to investigate whether we reached the theoretical limits of denoising performance. Subsequently, Burger et al. train a simple Multi-Layer Perceptron (MLP) on a large synthetic noise dataset. This method performs well against previous sophisticated algorithms. Several recent methods use deep CNNs and demonstrate promising denoising performance.

Image denoising can be applied to RAW or sRGB data. However, capturing diverse large-scale real noise data is a prohibitively expensive and tedious procedure, consequently leaving us to study denoising in synthetic settings. The most commonly used noise model for developing and evaluating image denoising is AWGN. As such, algorithms that are designed for AWGN cannot effectively remove noise from real images, as reported in recent benchmarks . A more accurate model for real RAW sensor noise contains both the signal-dependent noise component (the shot noise), and the signal-independent additive Gaussian component (the read noise) . The camera ISP transforms RAW sensor noise into a complicated form (spatio-chromatically correlated and not necessarily Gaussian). Therefore, estimating a noise model for denoising in sRGB space requires careful consideration of the influence of ISP. In this paper, we present a framework that is capable of synthesizing realistic noise data for training CNNs to effectively remove noise from RAW as well as sRGB images.

CycleISP

To synthesize realistic noise datasets, we use a two-stage scheme in this work. First, we develop a framework that models the camera ISP both in forward and reverse directions, hence the name CycleISP. Second, using CycleISP, we synthesize realistic noise datasets for the tasks of RAW denoising and sRGB image denoising. In this section, we only describe our CycleISP framework that models the camera ISP as a deep CNN system. Fig. 2 shows the modules of the CycleISP model: (a) RGB2RAW network branch, and (b) RAW2RGB network branch. In addition, we introduce an auxiliary color correction network branch that provides explicit color attention to the RAW2RGB network in order to correctly recover the original sRGB image.

The noise injection module in Fig. 2 is only required when synthesizing noisy data (Section 4), and thus we keep it in the ‘OFF’ state while learning CycleISP. The training process of CycleISP is divided in two steps: the RGB2RAW and RAW2RGB networks are first independently trained, and then joint fine-tuning is performed. Next, we present details of different branches of CycleISP. Note that we use RGB instead of sRGB to avoid notation clutter.

Digital cameras apply a series of operations on RAW sensor data in order to generate the monitor-ready sRGB images . Our RGB2RAW network branch aims to invert the effect of camera ISP. In contrast to the unprocessing technique of , the RGB2RAW branch does not require any camera parameters.

where each RRG contains multiple dual attention blocks, as we shall see in Section 3.3.

The RGB2RAW network is optimized using the L1L_{1} loss in linear and log domains as:

where ϵ\epsilon is a small constant for numerical stability, and Iraw\mathbf{I}_{raw} is the ground-truth RAW image. Similar to , the log loss term is added to enforce approximately equal treatment for all the image values; otherwise the network dedicates more attention to recovering the highlight regions.

2 RAW2RGB Network Branch

While the ultimate goal of RAW2RGB network is to generate synthetic realistic noise data for the sRGB image denoising problem, in this section we first describe how we can map clean RAW images to clean sRGB images (leaving the noise injection module ‘OFF’ in Fig. 2).

Note that Iraw\mathbf{I}_{raw} is the original camera RAW image (not the output of RGB2RAW network) because our objective here is to first learn RAW to sRGB mapping, independently. Color attention unit. To train the CycleISP, we use the MIT-Adobe FiveK dataset that contains images from several different cameras having diverse and complex ISP systems. It is extremely difficult for a CNN to accurately learn a RAW to sRGB mapping function for all different types of cameras (as one RAW image can potentially map to many sRGB images). One solution is to train one network for each camera ISP . However, such solutions are not scalable and the performance may not generalize to other cameras. To address this issue, we propose to include a color attention unit in the RAW2RGB network that provides explicit color attention via a color correction branch.

where ∗\ast denotes convolution, and KK is the Gaussian kernel with standard deviation empirically set to 12. This strong blurring operation ensures that only the color information flows through this branch, whereas the structural content and fine texture comes from the main RAW2RGB network. Using weaker blurring will undermine the effectiveness of the feature tensor Td′T_{d^{\prime}} of Eq. (4). The overall color attention unit process becomes:

where, ⊗\otimes is Hadamard product. To obtain the final sRGB image I^rgb\hat{\mathbf{I}}_{rgb}, the output features TattenT_{atten} from the color attention unit are passed through a RRG module, a convolutional layer M4M_{4} and an upscaling layer MupM_{up} , respectively:

For optimizing RAW2RGB network, we use the L1L_{1} loss:

3 RRG: Recursive Residual Group

Motivated by the advances of recent low-level vision methods based on the residual learning framework , we propose the RRG module, as shown in Fig. 3. The RRG contains PP dual attention blocks (DAB). The goal of each DAB is to suppress the less useful features and only allow the propagation of more informative ones. The DAB performs this feature recalibration by using two attention mechanisms: (1) channel attention (CA) , and (2) spatial attention (SA) . The overall process is:

4 Joint Fine-tuning of CycleISP

Since the RGB2RAW and RAW2RGB networks are initially trained independently, they may not provide the optimal-quality images due to the disconnection between them. Therefore, we perform joint fine-tuning in which the output of RGB2RAW becomes the input of RAW2RGB. The loss function for the joint optimization is:

where β\beta is a positive constant. Note that the RAW2RGB network receives gradients from the RAW2RGB sub-loss (only the second term). Whereas, the RGB2RAW network receives gradients from both sub-losses, thereby effectively contributing to the reconstruction of the final sRGB image.

Synthetic Realistic Noise Data Generation

Capturing perfectly-aligned real noise data pairs is extremely difficult. Consequently, image denoising is mostly studied in artificial settings where Gaussian noise is added to the clean images. While the state-of-the-art image denoising methods have shown promising performance on these synthetic datasets, they do not perform well when applied on real camera images . This is because the synthetic noise data differs fundamentally from real camera data. In this section, we describe the process of synthesizing realistic noise image pairs for denoising both in RAW and sRGB space using the proposed CycleISP method. Data for RAW denoising. The RGB2RAW network branch of the CycleISP method takes as input a clean sRGB image and converts it to a clean RAW image (top branch, Fig. 2). The noise injection module, which we kept off while training CycleISP, is now turned to the ‘ON’ state. The noise injection module adds shot and read noise of different levels to the output of RGB2RAW network. We use the same procedure for sampling shot/read noise factors as in . As such, we can generate clean and its corresponding noisy image pairs {RAWclean\text{RAW}_{clean}, RAWnoisy\text{RAW}_{noisy}} from any sRGB image. Data for sRGB denoising. Given a synthetic RAWnoisy\text{RAW}_{noisy} image as input, the RAW2RGB network maps it to a noisy sRGB image (bottom branch, Fig. 2); hence we are able to generate an image pair {sRGBclean\text{sRGB}_{clean},sRGBnoisy\text{sRGB}_{noisy}} for the sRGB denoising problem. While these synthetic image pairs are already adequate for training the denoising networks, we can further improve their quality with the following procedure. We fine-tune the CycleISP model (Section 3.4) using the SIDD dataset that is captured with real cameras. For each static scene, SIDD contains clean and noisy image pairs in both RAW and sRGB spaces. The fine-tuning process is shown in Fig. 4. Notice that the noise injection module which adds random noise is replaced by (only for fine-tuning) per-pixel noise residue that is obtained by subtracting the real RAWclean\text{RAW}_{clean} image from the real RAWnoisy\text{RAW}_{noisy} image. Once the fine-tuning procedure is complete, we can synthesize realistic noisy images by feeding clean sRGB images to the CycleISP model.

Denoising Architecture

As illustrated in Fig. 5, we propose an image denoising network by employing multiple RRGs. Our aim is to apply the proposed network in two different settings: (1) denoising RAW images, and (2) denoising sRGB data. We use the same network structure under both settings, with the only difference being in the handling of input and output. For denoising in the sRGB space, the input and output of the network are the 3-channel sRGB images. For denoising the RAW images, our network takes as input a 4-channel noisy packed image concatenated with a 4-channel noise level map, and provides us with a 4-channel packed denoised output. The noise level map provides an estimate of the standard deviation of noise present in the input image, based on its shot and read noise parameters .

Experiments

DND . This dataset consists of 5050 pairs of noisy and (nearly) noise-free images captured with four consumer cameras. Since the images are of very high-resolution, the providers extract 2020 crops of size 512×512512\times 512 from each image, thus yielding a total of 10001000 patches. The complete dataset is used for testing because the ground-truth noise-free images are not publicly available. The data is provided for two evaluation tracks: RAW space and sRGB space. Quantitative evaluation in terms of PSNR and SSIM can only be performed through an online server .

SIDD . Due to the small sensor size and high-resolution, smartphone images are much more noisy than those of DSLRs. This dataset is collected using five smartphone cameras. There are 320320 image pairs available for training and 12801280 image pairs for validation. This dataset also provides images both in RAW format and in sRGB space.

2 Implementation Details

All the models presented in this paper are trained with Adam optimizer (β1=0.9\beta_{1}=0.9, and β2=0.999\beta_{2}=0.999) and image crops of 128×128128\times 128. Using the Bayer unification and augmentation technique , we randomly perform horizontal and vertical flips. We set a filter size of 3×33\times 3 for all convolutional layers of the DAB except the last for which we use 1×11\times 1. Initial training of CycleISP. To train the CycleISP model, we use the MIT-Adobe FiveK dataset , which contains 50005000 RAW images. We process these RAW images using the LibRaw library and generate sRGB images. From this dataset, 48504850 images are used for training and 150150 for validation. We use 33 RRGs and 55 DABs for both RGB2RAW and RAW2RGB networks, and 22 RRGs and 33 DABs for the color correction network. The RGB2RAW and RAW2RGB branches of CycleISP are independently trained for 12001200 epochs with a batch size of 4. The initial learning rate is 10−410^{-4}, which is decreased to 10−510^{-5} after 800800 epochs.

Fine-tuning CycleISP. This process is performed twice: first with the procedure presented in Section 3.4, and then with the method of Section 4. In the former case, the output of the CycleISP model is noise-free, and in the latter case, the output is noisy. For each fine-tuning stage, we use 600600 epochs, batch size of 11 and learning rate of 10−510^{-5}. Training denoising networks. We train four networks to perform denoising on: (1) DND RAW data, (2) DND sRGB images, (3) SIDD RAW data, and (4) SIDD sRGB images. For all four networks, we use 44 RRGs and 88 DABs, 6565 epochs, batch size of 16, and initial learning rate of 10−410^{-4} which is decreased by a factor of 10 after every 2525 epochs. We take 11 million images from the MIR flickr extended dataset and split them into a ratio of 90:5:5 for training, validation and testing. All the images are preprocessed with the Gaussian kernel (σ=1\sigma=1) to reduce the effect of noise, and other artifacts. Next, we synthesize clean/noisy paired training data (both for RAW and sRGB denoising) using the procedure described in Section 4.

3 Results for RAW Denoising

In this section, we evaluate the denoising results of the proposed CycleISP model with existing state-of-the-art methods on the RAW data from DND and SIDD benchmarks. Table 1 shows the quantitative results (PSNR/SSIM) of all competing methods on the DND dataset obtained from the website of the evaluation server . Note that there are two super columns in the table listing the values of image quality metrics. The numbers in the sRGB super column are provided by the server after passing the denoised RAW images through the camera imaging pipeline using image metadata. Our model consistently performs better against the learning-based as well as conventional denoising algorithms. Furthermore, the proposed model has ∼\sim5×\times lesser parameters than previous best method . The trend is similar for the SIDD dataset, as shown in Table 2. Our algorithm achieves 6.89 dB improvement in PSNR over the BM3D algorithm.

A visual comparison of our result against the state-of-the-art algorithms is presented in Fig. 1. Our model is very effective in removing real noise, especially the low-frequency chroma noise and defective pixel noise.

4 Results for sRGB Denoising

While it is recommended to apply denoising on RAW data (where noise is uncorrelated and less complex) , denoising is commonly studied in the sRGB domain. We compare the denoising results of different methods on sRGB images from the DND and SIDD datasets. Table 3 and 4 show the scores of image quality metrics. Overall, the proposed model performs favorably against the state-of-the-art. Compared to the recent best algorithm RIDNet , our approach demonstrates the performance gain of 0.33 dB and 0.81 dB on DND and SIDD datasets, respectively.

Fig. 6 and 7 illustrate the sRGB denoising results on DND and SIDD, respectively. To remove noise, most of the evaluated algorithms either produce over-smooth images (and sacrifice image details) or generate images with splotchy texture and chroma artifacts. In contrast, our method generates clean and artifact-free results, while faithfully preserving image details.

5 Generalization Test

To compare the generalization capability of the denoising model trained on the synthetic data generated by our method and that of , we perform the following experiments. We take the (publicly available) denoising model of trained for DND, and directly evaluate it on the RAW images from the SIDD dataset. We repeat the same procedure for our denoising model as well. For a fair comparison, we use the same network architecture (U-Net) and noise model as of . The only difference is data conversion from sRGB to RAW. The results in Table 5 show that the denoising network trained with our method not only performs well on the DND dataset but also shows promising generalization to the SIDD set (a gain of ∼\sim 1 dB over ).

6 Ablations

We study the impact of individual contributions by progressively integrating them to our model. To this end, we use the RAW2RGB network that maps clean RAW image to clean sRGB image. Table 6 shows that the skip connections cause the largest performance drop, followed by the color correction branch. Furthermore, it is evident that the presence of both CA and SA is important, as well as their configuration (see Table 7), for the overall performance.

7 Color Matching For Stereoscopic Cinema

In professional 3D cinema, stereo pairs for each frame are acquired using a stereo camera setup, with two cameras mounted on a rig either side-by-side or (more commonly) in a beam splitter formation . During movie production, meticulous efforts are required to ensure that the twin cameras perform in exactly the same manner. However, oftentimes visible color discrepancies between the two views are inevitable because of the imperfect camera adjustments and impossibility of manufacturing identical lens systems. In movie post-production, color mismatch is corrected by a skilled technician, which is an expensive and highly involved procedure .

With the proposed CycleISP model, we can perform the color matching task, as shown in Fig. 8. Given a stereo pair, we first choose one view as the target and apply morphing to fully register it with the source view. Next, we pass the source RGB image through RGB2RAW model and obtain the source RAW image. Finally, we map back the source RAW image to the sRGB space using the RAW2RGB network, but with the color correction branch providing the color information from the ‘target’ RGB image (rather than the source RGB). Fig. 9 compares our method with three other color matching techniques . The proposed method generates results that are perceptually more faithful to the target views than other competing approaches.

Conclusion

sIn this work, we propose a data-driven CycleISP framework that is capable of converting sRGB images to RAW data and back to sRGB images. The CycleISP model allows us to synthesize realistic clean/noisy paired training data both in RAW and sRGB spaces. By training a novel network for the tasks of denoising the RAW and sRGB images, we achieve state-of-the-art performance on real noise benchmark datasets (DND and SIDD ). Furthermore, we demonstrate that the CycleISP model can be applied to the color matching problem in stereoscopic cinema. Our future work includes exploring and extending the CycleISP model for other low-level vision problems such as super-resolution and deblurring.

Acknowledgments. Ming-Hsuan Yang is supported by the NSF CAREER Grant 149783.

References