NeRF in the Dark: High Dynamic Range View Synthesis from Noisy Raw Images

Ben Mildenhall, Peter Hedman, Ricardo Martin-Brualla, Pratul Srinivasan, Jonathan T. Barron

Introduction

View synthesis methods, such as neural radiance fields (NeRF) , typically use tonemapped low dynamic range (LDR) images as input and directly reconstruct and render new views of a scene in LDR space. This poses no issues for scenes that are well-lit and do not contain large brightness variations, since they can be captured with minimal noise using a single fixed camera exposure setting. However, this precludes many common capture scenarios: images taken at nighttime or in any but the brightest indoor spaces will have poor signal-to-noise ratios, and scenes with regions of both daylight and shadow have extreme contrast ratios that require high dynamic range (HDR) to represent accurately.

Our method, RawNeRF, modifies NeRF to reconstruct the scene in linear HDR color space by supervising directly on noisy raw input images. This bypasses the lossy postprocessing that cameras apply to compress dynamic range and smooth out noise in order to produce visually palatable 8-bit JPEGs. By preserving the full dynamic range of the raw inputs, RawNeRF enables various novel HDR view synthesis tasks. We can modify the exposure level and tonemapping algorithm applied to rendered outputs and even create synthetically refocused images with accurately rendered bokeh effects around out-of-focus light sources.

Beyond these view synthesis applications, we show that training directly on raw data effectively turns RawNeRF into a multi-image denoiser capable of reconstructing scenes captured in near-darkness (Figure 1). The standard camera postprocessing pipeline (e.g., HDR+ ) corrupts the simple noise distribution of raw data, introducing significant bias in order to reduce variance and produce an acceptable output image. Feeding these images into NeRF thus produces a biased reconstruction with incorrect colors, particularly in the darkest regions of the scene (see Figure 2 for an example). We instead exploit NeRF’s ability to reduce variance by aggregating information across frames, demonstrating that it is possible for RawNeRF to produce a clean reconstruction from many noisy raw inputs.

Unlike typical video or burst image denoising methods, RawNeRF assumes a static scene and expects camera poses as input. Provided with these extra constraints, RawNeRF is able to make use of 3D multiview consistency to average information across nearly all of the input frames at once. Since our captured scenes each contain 25-200 input images, this means RawNeRF can remove more noise than feed-forward single or multi-image denoising networks that only make use of 1-5 input images for each output.

In summary, we make the following contributions:

We propose a method for training RawNeRF directly on raw images that can handle high dynamic range scenes as well as noisy inputs captured in the dark.

We show that RawNeRF outperforms NeRF on noisy real and synthetic datasets and is a competitive multi-image denoiser for wide-baseline static scenes.

We showcase novel view synthesis applications made possible by our linear HDR scene representation (varying exposure, tonemapping, and focus).

Related Work

RawNeRF combines concepts from several areas of research. We build upon NeRF as a baseline for high quality view synthesis, bring in ideas from low level image processing to optimize NeRF directly on noisy raw data, and take inspiration from uses of HDR in computer graphics and computational photography to showcase new applications made possible by an HDR scene reconstruction. We briefly cover relevant prior work across each of these areas.

Novel view synthesis is the task of using a set of input images and their camera poses to reconstruct a scene representation capable of rendering novel views. When the input images are densely sampled, it is possible to use direct interpolation in pixel space for view synthesis . A more feasible capture scenario is to capture more widely spaced inputs and use a “proxy” geometry (e.g., a reconstructed triangle mesh) to reproject and combine colors from the input images, using either a heuristic or learned blending function.

Recent work on applying deep learning to view synthesis has focused on volumetric rather than mesh-based scene representations . NeRF directly optimizes a neural volumetric scene representation to match all input images using gradient descent on a rendering loss. Various extensions have improved NeRF’s robustness to varying lighting conditions or added supervision with depth , time-of-flight data , or semantic segmentation labels . As of yet, no approach has extended NeRF to work with high dynamic range color data. Some previous view synthesis methods trained using LDR data jointly solve for per-image scaling factors to account for inconsistent lighting or miscalibration between cameras . ADOP supervises with LDR images and solves for exposure through a differentiable tonemapping step to approximately recover HDR, but does not focus on robustness to noise or supervision with raw data.

2 Denoising

Early neural denoising approaches mostly focused on denoising sRGB images synthetically corrupted with additive white Gaussian noise . In 2017, Plötz and Roth established a real raw image denoising benchmark, which showed that these deep denoisers failed to generalize beyond the synthetic data used during training and were outperformed by standard non-learned methods, such as BM3D . Subsequent work on both single and multi-image denoising demonstrated the benefits of training networks to operate directly on noisy raw input data. Modern cellphone camera pipelines perform a robust averaging of multiple noisy input frames in the raw domain , though they typically cannot afford to employ deep networks due to speed and power limitations.

Another line of research investigated whether denoisers could be trained using only noisy data when no corresponding clean ground truth exists. Noise2Noise demonstrated this was possible given a dataset of pairs of independent noisy observations of the same image, an insight Ehret et al. applied to denoise videos by aligning consecutive noisy frames. Various followups to Noise2Noise proposed modified network architectures allowing supervision with a dataset of single noisy images . Sheth et al. showed that this paradigm could be applied to train a denoiser using a single noisy video, including an application to raw video data. Similarly, RawNeRF is optimized over a single set of images to both denoise and recover the 3D structure of the captured scene.

3 Applications of raw and HDR image data

The value of working directly with raw data has long been noted by digital photographers due to the fact that its preservation of dynamic range allows for maximum postprocessing flexibility, letting users modify exposure, white balance, and tonemapping after the fact. Many works have tried to automate this process by using heuristics or machine learning to map directly from raw data to postprocessed LDR images .

Another line of work focuses on recovering HDR images from LDR inputs. This concept was pioneered by Debevec and Malik , who used a stack of aligned LDR images taken at different exposures to recover and invert the camera’s nonlinear response curve. Current approaches apply machine learning to produce HDR outputs from single or multiple misaligned LDR inputs, either recovering or hallucinating detail in clipped highlights.

Synthetic defocus

Many modern cellphones include a postprocessing option to add synthetic defocus blur after capture . Though it is possible to accurately simulate defocus using a thin-lens model or real multi-element camera lens using ray tracing, most machine learning models use a much faster approximate rendering model, predicting a depth map and applying a depth-varying blur kernel to each discretized depth layer . Performing this blur in HDR space is critical to achieving the correct appearance of defocused bright highlights (known as “bokeh”), as demonstrated by Zhang et al. .

Noisy Raw Input Data

NeRF takes postprocessed low dynamic range (LDR) sRGB color space images as input. This works well when using clean, noise-free images with minimal constrast. However, all real images contain some level of noise, and each step in the camera postprocessing pipeline corrupts this distribution in a certain way. Here we briefly describe the simplified pipeline stages relevant to our method.

When capturing an image, the number of photons hitting a pixel on the camera sensor is converted to an electrical charge, which is recorded as a high bit-depth digital signal (typically 10 to 14 bits). These values are offset by a “black level” to allow for negative measurements due to noise. After black level subtraction, the signal is a noisy measurement yiy_{i} of a quantity xix_{i} proportional to the expected number of photons arriving while the shutter is open. This noise results from both the physical fact that photon arrivals are a Poisson process (“shot” noise) and noise in the readout circuitry that converts the analog electrical signal to a digital value (“read” noise). The combined shot and read noise distribution can be well modeled as a Gaussian whose variance is an affine function of its mean ; importantly, this implies that the distribution of the error yi−xiy_{i}-x_{i} is zero mean.

Color filter demosaicking

Color cameras contain a Bayer color filter array in front of the image sensor such that each pixel’s spectral response curve measures either red, green or blue light. The pixel color values are typically arranged in 2×22\times 2 squares containing two green pixels, one red, and one blue pixel (known as a Bayer pattern), resulting in “mosaicked” data. To generate a full-resolution color image, the missing color channels are interpolated using a demosaicking algorithm . This interpolation correlates noise spatially, and the checkerboard pattern of the mosaic leads to different noise levels in alternating pixels.

Color correction and white balance

The spectral response curves for each color filter element vary between different cameras, and a color correction matrix is used to convert the image from this camera-specific color space to a standardized color space. Additionally, because human perception is robust to the color tint imparted by different light sources, cameras attempt to account for this tint (i.e., make white surfaces appear RGB-neutral white) by scaling each color channel by an estimated white balance coefficient. These two steps are typically combined into a single linear 3×33\times 3 matrix transform, which further correlates the noise between color channels.

Gamma compression and tonemapping

Humans are able to discern smaller relative differences in dark regions compared to bright regions of an image. This fact is exploited by sRGB gamma compression, which optimizes the final image encoding by clipping values outside $$ and applying a nonlinear curve to the signal that dedicates more bits to dark regions at the cost of compressing bright highlights. In addition to gamma compression, tonemapping algorithms can be used to better preserve contrast in high dynamic range scenes (where the bright regions are several orders of magnitude brighter than the darkest) when the image is quantized to 8 bits .

In a slight abuse of terminology, we will refer both of these steps jointly as “tonemapping” in the rest of the paper, indicating the process by which linear HDR values are mapped to nonlinear LDR space for visualization. We will refer to signals before tonemapping as high dynamic range (HDR) and signals after as low dynamic range (LDR). Of all postprocessing operations, tonemapping has the most drastic effect on the noise distribution: clipping completely discards information in the brightest and darkest regions, and after the non-linear tonemapping curve the noise is no longer guaranteed to be Gaussian or even zero mean.

RawNeRF

A neural radiance field (NeRF) is a neural network based scene representation that is optimized to reproduce the appearance of a set of input images with known camera poses. The resulting reconstruction can then be used to render novel views from previously unobserved poses. NeRF’s multilayer perceptron (MLP) network takes 3D position and 2D viewing direction as input and outputs volume density and color. To render each pixel in an output image, NeRF uses volume rendering to combine the colors and densities from many points sampled along the corresponding 3D ray.

Standard NeRF takes clean, low dynamic range (LDR) sRGB color space images with values in the range $$ as input. Converting raw HDR images to LDR images (e.g., using the pipeline described in Section 3) has two significant consequences:

Detail in bright areas is lost when values are clipped from above at one, and detail across the image is compressed by the tonemapping curve and subsequent quantization to 8 bits.

The per-pixel noise distribution becomes biased (no longer zero-mean) after passing through a nonlinear tonemapping curve and being clipped from below at zero.

The goal of RawNeRF is to make use of this information rather than discarding it, optimizing NeRF directly on linear raw input data in HDR color space (Figure 3). In Section 5, we will show that reconstructing NeRF in raw space makes it much more robust to noisy inputs and allows for novel HDR view synthesis applications. First, we detail the changes required to make NeRF work with raw data.

Since the color distribution in an HDR image can span many orders of magnitude, a standard L2 loss applied in HDR space will be completely dominated by error in bright areas and produce an image that has muddy dark regions with low contrast when tonemapped (see Figure 4). Instead, we apply a loss that more strongly penalizes errors in dark regions to align with how human perception compresses dynamic range. One way to achieve this is by passing both the rendered estimate y^\hat{y} and noisy observed intensity yy through a tonemapping curve ψ\psi before the loss is applied:

However, in low-light raw images the observed signal yy is heavily corrupted by zero-mean noise, and a nonlinear tonemap will introduce bias that changes the noisy signal’s expected value (E[ψ(y)]≠ψ(E[y])E[\psi(y)]\neq\psi(E[y])). In order for the network to converge to an unbiased result , we instead use a weighted L2 loss of the form

We can approximate the tonemapped loss (1) in this form by using a linearization of the tone curve ψ\psi around each y^i\hat{y}_{i}:

This corresponds exactly to the relative MSE loss used to achieve unbiased results when training on noisy HDR pathtracing data in Noise2Noise . The curve ψ\psi is proportional to the μ\mu-law function used for range compression in audio processing, and has previously been applied as a tonemapping function when supervising a network to map from a burst of LDR images to an HDR output .

2 Variable exposure training

In scenes with very high dynamic range, even a 10-14 bit raw image may not be sufficient for capturing both bright and dark regions in a single exposure. This is addressed by the “bracketing” mode included in many digital cameras, where multiple images with varying shutter speeds are captured in a burst, then merged to take advantage of the bright highlights preserved in the shorter exposures and the darker regions captured with more detail in the faster exposures.

We can similarly take advantage of variable exposures in RawNeRF (Figure 5). Given a sequence of images IiI_{i} with exposure times tit_{i} (and all other capture parameters held constant), we can “expose” RawNeRF’s linear space color output to match the brightness in image IiI_{i} by scaling it by the recorded shutter speed tit_{i}. In practice, we find that varying exposures cannot be precisely aligned using shutter speed alone due to sensor miscalibration (see supplement). To correct for this, we add a learned per-color-channel scaling factor for each unique shutter speed present in the set of captured images, which we jointly optimize along with the NeRF network. The final RawNeRF “exposure” given a output color y^i\hat{y}_{i} from the network is then min⁡(y^ic⋅ti⋅αtic,1),\min(\hat{y}_{i}^{c}\cdot t_{i}\cdot\alpha_{t_{i}}^{c},1), where cc indexes color channels, and αtic\alpha_{t_{i}}^{c} is the learned scaling factor for shutter speed tit_{i} and channel cc (we constrain αtmaxc=1\alpha_{t_{\textrm{max}}}^{c}=1 for the longest exposure). We clip from above at 1 to account for the fact that pixels saturate in overexposed regions. This scaled and clipped value is passed to the previously described loss (Equation 4).

3 Implementation details

Our implementation is based on the mip-NeRF codebase, which improves upon the positional encoding used in the original NeRF method. Please see that paper for further details on the MLP scene representation and volumetric rendering algorithm. Our only network architecture change is to modify the activation function for the MLP’s output color from a sigmoid to an exponential function to better parameterize linear radiance values. We use the Adam optimizer with batches of 1616k random rays sampled across all training images and a learning rate decaying from 10−310^{-3} to 10−510^{-5} over 500500k steps of optimization.

We find that extremely noisy scenes benefit from a regularization loss on volume density to prevent partially transparent “floater” artifacts. We apply a loss on the variance of the weight distribution used to accumulate color values along the ray during volume rendering; please see the supplement for details.

As our raw input data is mosaicked, it only contains one color value per pixel. We only apply the loss to the active color channel for each pixel, such that optimizing NeRF effectively demosaics the input images. Since any resampling steps will effect the raw noise distribution, we do not undistort or downsample the inputs, and instead train using the full resolution mosaicked images (usually 12MP for our scenes). To achieve this, we use camera intrinsics to account for radial distortion when generating rays. We use full resolution postprocessed JPEG images to calculate camera poses as COLMAP does not support raw images.

Results

We present results exploring two consequences of supervising NeRF with raw HDR data. First, we show that RawNeRF is surprisingly robust to high levels of noise, to the extent that it can act as a competitive multi-image denoiser when applied to wide-baseline images of a static scene. Second, we demonstrate the HDR view synthesis applications enabled by recovering a scene representation that preserves high dynamic range color values.

Recent years have seen an increasing focus on developing deep learning methods for denoising images directly in the raw linear domain . This effort has expanded to include multi-image denoisers that can be applied to burst images or video frames . These multi-image denoisers typically assume that there is a relatively small amount of motion between frames, but that there may be large amounts of object motion within the scene. When nearby frames can be well aligned, these methods merge information from similar image patches (typically across 2-8 neighboring images) to outperform single image denoisers.

By comparison, NeRF (and by extension, RawNeRF) optimizes for a single scene reconstruction that is consistent with all input images. By specializing to wide-baseline static scenes and taking advantage of 3D multiview information, RawNeRF can aggregate observations from much more widely spaced input images than a typical multi-image denoising method.

We collect a real world denoising dataset with 3 different scenes, each consisting of 101 noisy images and a clean reference image merged from stabilized long exposures. The first 100 images are taken handheld across a wide baseline (a standard forward-facing NeRF capture), using a fast shutter speed to accentuate noise. We then capture a stabilized burst of 50-100 longer exposures on a tripod and robustly merge them using HDR+ to create a clean ground truth frame. One additional tripod image taken at the original fast shutter speed serves as a noisy input “base frame” for the deep denoising methods. All images are taken with an iPhone X at 12MP resolution using the wide-angle lens and saved as 12-bit raw DNG files.

Comparisons

In Table 1 and Figure 6, we compare RawNeRF’s joint view synthesis and denoising performance to several recent deep single and multi-image denoising methods. Note that all denoisers require the noisy version of the test image as input, whereas RawNeRF and its ablations only require its camera pose.

We focus our comparison on methods explicitly designed to handle raw input images. Chen et al. (SID) present a single image denoiser that maps from raw inputs to postprocessed LDR images and is trained on a large dataset of noisy raw and clean postprocessed image pairs collected by the authors. Brooks et al. (Unprocess) is a method for training a raw single image denoiser on simulated raw data created from internet image datasets that transfers well to real raw images. RViDeNet trains a raw video denoiser on a combination of Unprocessing-style synthetic data and a new real raw video dataset. Sheth et al. (UDVD) present a “self-supervised” method for training a video denoiser only using noisy data, building on ideas from Noise2Noise and blind-spot networks . UDVD provides network weights specifically trained on the raw video dataset from RViDeNet. For all methods, we use publicly available code and pretrained model weights.

We also compare to two ablations of our method. LDR NeRF represents mip-NeRF trained (as usual) in LDR sRGB space on images postprocessed by a minimal sRGB tonemapping pipeline. “Un+RawNeRF” preprocesses the training images using the single image raw denoiser from Brooks et al. (“Unprocess”) before training RawNeRF.

All compared methods take mosaicked raw images as input. Every deep denoiser uses the noisy “base frame” as input, and the two multi-image denoising networks also receive the nearest images from the wide-baseline capture (based on camera position). We convert the 12-bit raw input to floating point by normalizing with the white and black levels. Since each method was trained on raw data from a different source, they impart different color tints to the output. So this not affect metrics, we calculate a per-color-channel affine transform that best matches each method’s raw output to the ground truth raw image. (The exceptions are SID and LDR NeRF, whose sRGB output we match to the postprocessed sRGB ground truth.) Our basic postprocessing pipeline for visualization and computing sRGB metrics is to apply a bilinear demosaic (when necessary), perform white balance/color correction, rescale white level, clip to $$, and apply the sRGB gamma curve. Please see the supplement for details.

Analysis

Despite simultaneously performing denoising and novel view synthesis, our method is competitive with all compared deep denoisers (Table 1, Figure 6). We suspect that the multi-image denoisers struggle to make use of the additional frames provided from the wide-baseline capture, as the camera movement is larger than in a typical sub-second burst or video clip. By comparison, RawNeRF, despite lacking any explicitly learned image priors, clean training data, or even a “base frame” input image, produces high quality outputs by combining information from across all input images in its reconstruction. Despite the fact that LDR NeRF is directly trained to minimize mean-squared error in sRGB space, RawNeRF achieves significantly better sRGB metrics. We also find that applying a single image denoiser to the inputs before training RawNeRF results in oversmoothed renderings (Un+RawNeRF).

Synthetic noise ablation

In Table 2 and Figure 7, we demonstrate the impact of noise level on RawNeRF image quality. For training, we render 120 linear HDR images using the Lego scene from NeRF , borrowing color correction, white balance, and noise parameters from our iPhone captures’ EXIF metadata to “unprocess” this data into raw space . Since the renderings have a large amount of empty space, we report sRGB PSNR on the object only, by using the provided alpha masks (otherwise error from the background pixels heavily penalizes LDR NeRF). Even in this synthetic setting free from camera miscalibration issues, we can clearly observe the color bias and loss of detail caused by training LDR NeRF on postprocessed noisy data.

2 HDR view synthesis applications

Figures 1, 2, 4, 5, and 8 include examples of varying the exposure level and tonemapping algorithm for images output by RawNeRF, which exist in linear HDR space and can thus be postprocessed like a raw photo from a digital camera. Please see our supplement and video for many more examples.

Synthetic defocus

Given a full 3D model of a scene, physically-based renderers accurately simulate camera lens defocus effects by tracing rays refracted through each lens element , but this process is extremely computationally expensive. A reasonably convincing and much cheaper solution is to apply a varying blur kernel to different depth layers of the scene and composite them together . In Figure 8, we apply this synthetic defocus rendering model to sets of RGBA depth layers precomputed from trained RawNeRF models (similar to a multiplane image ). As shown by Zhang et al. , recovering linear HDR color is critical for achieving the characteristic oversaturated “bokeh balls” around defocused bright light sources.

Discussion

We have demonstrated the benefits of training NeRF directly on linear raw camera images. However, this modification is not without tradeoffs. Most digital cameras can only save raw images at full resolution with minimal compression, resulting in huge storage requirements when capturing tens or hundreds of images per scene. Our method is also dependent on COLMAP’s robustness for computing camera poses, preventing us from capturing scenes below a certain light level. This could potentially be addressed by jointly optimizing RawNeRF and the input camera poses . Finally, despite its robustness to noise, RawNeRF cannot be considered a general purpose denoiser as it cannot handle scene motion and requires orders of magnitude more computation than a feed-forward network.

Despite these shortcomings, we believe that RawNeRF represents a step toward robust, high quality capture of real world environments. Training on raw images with variable exposure allows us to capture scenes with a much wider dynamic range, and robustness to noise makes reconstructing dark nighttime captures possible. Lifting these constraints greatly increases the fraction of the world that can be reconstructed and explored with photorealistic view synthesis.

References

Appendix A Potential negative impact

Training any NeRF model for scene reconstruction has potential negative environmental impact, as current algorithms are very compute-intensive, requiring hours of training per scene even when run on specialized ML accelerators. This also creates an unfair advantage for research groups with access to more computational resources. Future work will likely address this issue, as it blocks the widespread practical adoption of these models.

Any image restoration model could potentially be applied for illicit surveillance purposes. Multi-image denoisers provide the additional capability of potentially revealing details that are not visible in any single image due to noise. ML-based algorithms further complicate this situation by potentially “hallucinating” details in ambiguous regions, either intentionally (as with generative methods) or unintentionally (in the form of reconstruction artifacts). RawNeRF has a minimal ability to hallucinate, as it largely works by simply averaging the input data, but it does occasionally produce high frequency grid-like patterns due to the bias induced by positional encoding.

Appendix B Additional qualitative results

We include additional qualitative results for both dark (Figure 9) and high contrast scenes (Figure 10). We urge the reader to view our supplemental video as the results are more compelling when animated.

Appendix C Training details

We wish to approximate the effect of training with the following loss

while converging to an unbiased result. This can be accomplished by using a locally valid linear approximation for the error term:

Note that we choose to linearize around y^i\hat{y}_{i} because, unlike the noisy observation yiy_{i}, y^i\hat{y}_{i} tends towards the true signal value xi=E⁡[yi]x_{i}=\operatorname{E}[y_{i}] over the course of training.

If we use a weighted L2 loss, then as we train the network we will have y^i→E⁡[yi]=xi\hat{y}_{i}\to\operatorname{E}[y_{i}]=x_{i} in expectation (where xix_{i} is the true signal value). This means that the terms summed in our gradient-weighted loss

will tend towards ψ′(xi)(y^i−yi)\psi^{\prime}(x_{i})(\hat{y}_{i}-y_{i}) over the course of training. Additionally, we note that the gradient of our reweighted loss 7 is a linear approximation of the gradient of the tonemapped loss 5:

In line 10 we substitute the linearization from 6, and in line 11 we exploit the fact that a stop-gradient has no effect for expressions that will not be further differentiated.

C.2 Weight variance regularizer

Our weight variance regularizer is a function of the compositing weights used to calculate the final color for each ray. Given MLP outputs ci,σic_{i},\sigma_{i} for respective ray segments [ti−1,ti)[t_{i-1},t_{i}) with lengths Δi\Delta_{i} (see ), these weights are

If we define a piecewise-constant probability distribution pwp_{w} over the ray segments using these weights, then our variance regularizer is equal to

We will denote this value as t‾\overline{t}. Calculating the regularizer:

We apply a weight between 1×10−21\times 10^{-2} and 1×10−11\times 10^{-1} to Lw\mathcal{L}_{w} (relative to the rendering loss), typically using higher weights in noisier or darker scenes that are more prone to “floater” artifacts. Applying this regularizer with a high weight can result in a minor loss of sharpness, which can be ameliorated by annealing its weight from 0 to 1 over the course of training.

C.3 Findings with alternate loss functions

In practice, we directly scale our loss by the derivative of the desired tone curve:

We also experimented with using a reweighted L1 loss or the negative log-likelihood function of the actual camera noise model (using shot/read noise parameters from the EXIF data) but found that this performed worse than reweighted L2. RawNeRF models supervised with a standard unweighted L2 or L1 loss tended to diverge early in training, particularly in very noisy scenes.

We tried using the unclipped sRGB gamma curve (extended as a linear function below zero and as an exponential function above 1) in our loss, but found that it caused many color artifacts in dark regions. Directly applying our log tone curve (rather than reweighting by its gradient) before the L2 loss caused training to diverge.

C.4 Quality limitations

As briefly mentioned in the main text, our method cannot scale to arbitrary amounts of noise in real world scenes. For our darkest nighttime scenes, we often must run COLMAP multiple times (varying the random seed) or tune its parameters to obtain camera poses. Even when COLMAP reports a successful reconstruction, the results are sometimes poorly aligned at image corners, where the distortion model used for camera intrinsics may not fit well.

RawNeRF itself is prone to reconstruction artifacts in very noisy scenes or scenes captured with few images (under 30), typically in the form of positional encoding grid-like artifacts. These artifacts are often more evident in videos than in still frames. In regions that are essentially pure noise and no signal, RawNeRF sometimes produces a foggy “cloud”, since no multiview information exists to guide its recovery of geometry.

The near and far plane bounds calculated using the point cloud from COLMAP are sometimes wider than the true bounds of the scene. Using these bounds wastes many samples at the front of each ray, which reduces sharpness and can cause additional “floater” artifacts. We therefore sometimes retrain RawNeRF models using tighter depth bounds than those reported by COLMAP.

We found it necessary to use gradient clipping due to the high level of noise in the data we use for supervision. Certain losses (such as standard L2) are prone to producing NaN gradient values and require careful tuning of the clipping values. We found our reweighted loss to be more stable.

Appendix D Data capture and postprocessing details

We captured all images using a 2017 iPhone X with the Halide apphttps://halide.cam/ and a 2020 iPhone SE with the Adobe Lightroom app. We used manual modes in both apps with focus and ISO level fixed for each capture, manually adjusting shutter speed to achieve an exposure with no clipped highlights (except in scenes with varying exposure) and minimal motion blur (at least 1/1001/100s when possible). At night, it was usually necessary to use the maximum ISO level (approximately 2000 on the iPhones) to achieve minimal motion blur. Each capture took around 10-200 seconds, except for the denoising test scenes. All raw images are stored as Adobe DNGhttps://www.adobe.com/content/dam/acom/en/products/photoshop/pdfs/dng_spec_1.4.0.0.pdf files.

We extract the following parameters from the EXIF metadata using exiftool:

We use the standard sRGB gamma curve as a basic tonemap for linear RGB space data:

D.2 Postprocessing pipeline

Our exact postprocessing pipeline for converting raw images to postprocessed sRGB space is detailed below.

Rescale so that the black level is 0 and the white level is 1, preserving values below zero. (The result here is used to train RawNeRF.)

Apply bilinear demosaicking (when necessary).

Apply a color correction matrix (from camera RGB to canonical XYZ) and XYZ-to-RGB matrix, combined into a 3×33\times 3 transformation.

Adjust the exposure to set the white level to the pp-th percentile (p=97p=97 by default).

Apply the sRGB gamma curve to each color channel.

When applying a different tonemapping algorithm, we take the color corrected output from step 6 and pass it through the alternate method, while tuning exposure and other tonemapping parameters manually per scene.

D.3 Camera shutter speed miscalibration

In Section 4.2 of the main text, we discuss our implementation of a learned per-color-channel scaling to account for miscalibration when using variable exposure inputs. Here, we document this miscalibration effect for completeness.

Figure 11 plots data taken from a “sweep” over many shutter speeds. The 2017 iPhone X (used for most data capture in the paper) is held fixed on a tripod, all other parameters (focus, ISO, white balance, etc.) are held fixed, and shutter speeds are sampled roughly logarithmically from 1/100 to 1/10000 seconds. We ensure that no pixels are saturated. To minimize the effect of image noise, we study the average color value yticy_{t_{i}}^{c} for each Bayer filter channel (R, G1, G2, B) over the entire 12MP sensor. Specifically, we plot:

We show an example of the resulting qualitative color shift in Figure 12 using images from one of our three real test scenes. Here the two shutter speeds are 1/1104 and 1/181 seconds, and the relative color shift from the slow to the fast channel is calculated to be (0.89,0.93,0.75)(0.89,0.93,0.75) for red, green, and blue in the raw domain. The effect of undoing this shift before postprocessing is shown in Figure 12b. This miscalibration is another reason for primarily reporting affine-aligned metrics on our real test set, since we cannot rely on perfect color alignment between the input noisy image and the clean ground truth frame.

We do not fully understand the cause of this issue. We speculate that it could be due to the sensor temperature changing over the course of capture, imprecise shutter speed timing for very fast exposures, or any number of other factors related to low level sensor hardware. Given that the effect exists and affects our captures in an unmeasurable manner, it must be accounted for. Using a DSLR or mirrorless camera with a better sensor may avoid this issue.

Appendix E Comparison and ablation details

As mentioned in the main text, we solve for an affine color alignment between each output and the ground truth clean image. For all methods but SID and LDR NeRF, this is done directly in raw Bayer space for each RGGB plane separately. For SID and LDR NeRF (which output images in tonemapped sRGB space), this is done for each RGB plane against the tonemapped sRGB clean image. If the ground truth channel is xx and the channel to be matched is yy, we specifically compute

to get the least-squares fit of an affine transform ax+b≈yax+b\approx y (here z‾\overline{z} indicates the mean over all elements of zz). We then apply the inverse transform as (y−b)/a(y-b)/a to match the estimated yy to xx. In the case where matching happens in the raw domain, we postprocess (y−b)/a(y-b)/a through our standard pipeline (Section D.2) before calculating sRGB-space metrics.

Compared baselines

We provide an overview of each baseline and the pre- and post-processing pipelines used in the main text. Unprocessing is the only method that is a “non-blind” denoiser, and therefore requires a per-pixel noise level as input. We calculate this by using the empirical per-pixel variance from our tripod-aligned fast and clean images to estimate shot and read noise parameters as a best-fit 1D affine transform mapping from clean signal values to empirical variances. Each method required its own relative input rescaling and clipping convention, which we set based on each authors’ source code.

E.2 Synthetic Lego dataset details

In the synthetic Lego dataset, we did not include the effects of remosaicking/demosaicking or quantization when unprocessing/reprocessing the data. We wanted the “infinite” shutter speed case to be perfectly clean, with no degradation resulting from unprocessing and reprocessing in the absence of noise, thus providing an upper bound on possible performance. This example does not particularly test the ability of RawNeRF to encode high dynamic range since the object is diffusely lit, resulting in fairly dim highlights and negligible clipping; instead, it focuses on robustness to noise.

We rendered new randomly sampled images of the scene using the Blender filehttps://drive.google.com/file/d/1RjwxZCUoPlUgEWIUiuCmMmG0AhuV8A2Q/view?usp=sharing provided by the NeRF authors , saving the resulting linear space color data in EXR format. There are 120 images in the training set and 40 images in the test set. Note that metric values on this data are not comparable to metrics on the original scene, since it uses images from different random poses generated using a different postprocessing pipeline.

For completeness, we report the unmasked PSNR values for this experiment in Table 3 (Table 2 in the main text reports masked PSNR), which is heavily skewed by the LDR NeRF’s color bias in the black background regions.

Appendix F Further qualitative ablations

In all LDR NeRF comparisons in the main paper, we use our own simple postprocessing pipeline to generate LDR sRGB inputs from the raw data. However, a standard NeRF implementation would instead use JPEG images directly from the camera, which have a more sophisticated postprocessing pipeline that likely includes noise reduction and a more sophisticated nonlinear tonemap to better compress dynamic range. To satisfy the reader’s potential curiosity, in Figure 13 we provide an example of LDR NeRF trained on iPhone JPEGs versus our LDR images, as well as a RawNeRF result on the same scene.

F.2 Bayer mosaic mask and sensor artifacts

In the main text, we note that we only apply our loss function to the color channel measured by the Bayer filter for each ray. (In practice, we render all three colors for every training ray, then apply a one-hot mask to select the desired output color.) In Figure 14, we show an example of the color noise that emerges when supervising all 3 color channels using bilinearly demosaicked raw images instead of masking the loss. Perhaps surprisingly, we noted that relatively clean regions of the scene seemed to benefit from using all 3 channels of a bilinear demosaicked image as supervision. However, we concluded that the distracting color artifacts induced by demosaicking outweighed this occasional benefit, and opted to use Bayer masking in all scenes.

These artifacts may potentially be caused by broken “hot” pixels that are always fully saturated, in violation of our assumed noise distribution. Bilinear demosaicking would disperse the influence of a hot pixel to many neighboring pixels, potentially increasing its effect on the final trained NeRF. In preliminary experiments, we did not notice any benefit to additionally masking hot pixels when applying a Bayer mask. We did apply a second mask to remove a 4 pixel border from all training images, since many iPhone raw images contained 1 or 2 entire rows or columns of saturated pixels on one side, particularly in bright scenes.

Appendix G Synthetic defocus rendering model

To render defocused images, we use a similar rendering model as prior work that has addressed this task . To avoid prohibitively expensive rendering speeds, we first precompute a multiplane image representation from the trained RawNeRF model. This MPI consists of a series of fronto-parallel RGBA planes (with colors still in linear HDR space), sampled linearly in disparity within a camera frustum at a central camera pose. Given this MPI representation, our rendering algorithm for synthetic defocus (including lateral camera translation) is described in Algorithm 1.

Appendix H Scene index

We provide various details about each scene shown in the paper and video in Table 4.