BARF: Bundle-Adjusting Neural Radiance Fields

Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, Simon Lucey

Introduction

Humans have strong capabilities of reasoning about 3D geometry through our vision from the slightest ego-motion. When watching movies, we can immediately infer the 3D spatial structures of objects and scenes inside the videos. This is because we have an inherent ability of associating spatial correspondences of the same scene across continuous observations, without having to make sense of the relative camera or ego-motion. Through pure visual perception, not only can we recover a mental 3D representation of what we are looking at, but meanwhile we can also recognize where we are looking at the scene from.

Simultaneously solving for the 3D scene representation from RGB images (i.e. reconstruction) and localizing the given camera frames (i.e. registration) is a long-standing chicken-and-egg problem in computer vision — recovering the 3D structure requires observations with known camera poses, while localizing the cameras requires reliable correspondences from the reconstruction. Classical methods such as structure from motion (SfM) or SLAM approach this problem through local registration followed by global geometric bundle adjustment (BA) on both the structure and cameras. SfM and SLAM systems, however, are sensitive to the quality of local registration and easily fall into suboptimal solutions. In addition, the sparse nature of output 3D point clouds (often noisy) limits downstream vision tasks that requires dense geometric reasoning.

Closely related to 3D reconstruction from imagery is the problem of view synthesis. Though not primarily purposed for recovering explicit 3D structures, recent advances on photorealistic view synthesis have opted to recover an intermediate dense 3D-aware representation (e.g. depth , multi-plane images , or volume density ), followed by neural rendering techniques to synthesize the target images. In particular, Neural Radiance Fields (NeRF) have demonstrated its remarkable ability for high-fidelity view synthesis. NeRF encodes 3D scenes with a neural network mapping 3D point locations to color and volume density. This allows the scenes to be represented with compact memory footprint without limiting the resolution of synthesized images. The optimization process of the network is constrained to obey the principles of classical volume rendering , making the learned representation interpretable as a continuous 3D volume density function.

Despite its notable ability for photorealistic view synthesis and 3D scene representation, a hard prerequisite of NeRF (as well as other view synthesis methods) is accurate camera poses of the given images, which is typically obtained through auxiliary off-the-shelf algorithms. One straightforward way to circumvent this limitation is to additionally optimize the pose parameters with the NeRF model via backpropagation. As discussed later in the paper, however, naïve pose optimization with NeRF is sensitive to initialization. It may lead to suboptimal solutions of the 3D scene representation, degrading the quality of view synthesis.

In this paper, we address the problem of training NeRF representations from imperfect camera poses — the joint problem of reconstructing the 3D scene and registering the camera poses (Fig. 1). We draw inspiration from the success of classical image alignment methods and establish a theoretical connection, showing that coarse-to-fine registration is also critical to NeRF. Specifically, we show that positional encoding of input 3D points plays a crucial role — as much as it enables fitting to high-frequency functions , positional encoding is also more susceptible to suboptimal registration results. To this end, we present Bundle-Adjusting NeRF (BARF), a simple yet effective strategy for coarse-to-fine registration on coordinate-based scene representations. BARF can be regarded as a type of photometric BA using view synthesis as the proxy objective. Unlike traditional BA, however, BARF can learn scene representations from scratch (i.e. from randomly initialized network weights), lifting the reliance of local registration subprocedures and allowing for more generic applications.

In summary, we present the following contributions:

We establish a theoretical connection between classical image alignment to joint registration and reconstruction with Neural Radiance Fields (NeRF).

We show that susceptibility to noise from positional encoding affects the basin of attraction for registration, and we present a simple strategy for coarse-to-fine registration on coordinate-based scene representations.

Our proposed BARF can successfully recover scene representations from imperfect camera poses, allowing for applications such as view synthesis and localization of video sequences from unknown poses.

Related Work

Structure from motion (SfM) and SLAM. Given a set of input images, SfM and SLAM systems aim to recover the 3D structure and the sensor poses simultaneously. These can be classified into (a) indirect methods that rely on keypoint detection and matching and (b) direct methods that exploit photometric consistency . Modern pipelines following the indirect route have achieved tremendous success ; however, they often suffer at textureless regions and repetitive patterns, where distinctive keypoints cannot be reliably detected. Researchers have thus sought to use neural networks to learn discriminative features directly from data .

Direct methods, on the other hand, do not rely on such distinctive keypoints — every pixel can contribute to maximizing photometric consistency, leading to improved robustness in sparsely textured environments . They can also be naturally integrated into deep learning frameworks through image reconstruction losses . Our method BARF lies under the broad umbrella of direct methods, as BARF learns 3D scene representations from RGB images while also localizing the respective cameras. However, unlike classical SfM and SLAM that represent 3D structures with explicit geometry (e.g. point clouds), BARF encodes the scenes as coordinate-based representations with neural networks.

View synthesis. Given a set of posed images, view synthesis attempts to simulate how a scene would look like from novel viewpoints . The task has been closely tied to 3D reconstruction since its introduction . Researchers have investigated blending pixel colors based on depth maps or leveraging proxy geometry to warp and composite the synthesized image . However, since the problem is inherently ill-posed, there are still multiple restrictions and assumptions on the synthesized viewpoints.

State-of-the-art methods have capitalized on neural networks to learn both the scene geometry and statistical priors from data. Various representations have been explored in this direction, e.g. depth , layered depth , multi-plane images , volume density , and mesh sheets . Unfortunately, these view synthesis methods still require the camera poses to be known a priori, largely limiting their applications in practice. In contrast, our method BARF is able to effectively learn 3D representations that encodes the underlying scene geometry from imperfect or even unknown camera poses.

Neural Radiance Fields (NeRF). Recently, Mildenhall et al. proposed NeRF to synthesize novel views of static, complex scenes from a set of posed input images. The key idea is to model the continuous radiance field of a scene with a multi-layer perceptron (MLP), followed by differentiable volume rendering to synthesize the images and backpropagate the photometric errors. NeRF has drawn wide attention across the vision community due to its simplicity and extraordinary performance. It has also been extended on many fronts, e.g. reflectance modeling for photorealistic relighting and dynamic scene modeling that integrates the motion of the world . Recent works have also sought to exploit a large corpus of data to pretrain the MLP, enabling the ability to infer the radiance field from a single image .

While impressive results have been achieved by the above NeRF-based models, they have a common drawback — the requirement of posed images. Our proposed BARF allows us to circumvent such requirement. We show that with a simple coarse-to-fine bundle adjustment technique, we can recover from imperfect camera poses (including unknown poses of video sequences) and learn the NeRF representation simultaneously. Concurrent to our work, NeRF– introduced an empirical, two-stage pipeline to estimate unknown camera poses. Our method BARF, in contrast, is motivated by mathematical insights and can recover the camera poses within a single course of optimization, allowing for direct utilities for various NeRF applications and extensions.

Approach

We unfold this paper by motivating with the simpler 2D case of classical image alignment as an example. Then we discuss how the same concept is also applicable to the 3D case, giving inspiration to our proposed BARF.

The steepest descent image J\mathbf{J} can be expanded as

or alternatively, one may choose to solve for warp parameters p1\mathbf{p}_{1} and p2\mathbf{p}_{2} respectively for both images I1\mathcal{I}_{1} and I2\mathcal{I}_{2} through

where M=2M=2 is the number of images. Albeit similar to (1), the image gradients become the analytical Jacobian of the network ∂f(x)∂x\frac{\partial f(\mathbf{x})}{\partial\mathbf{x}} instead of numerical estimation. By manipulating the network ff, this also enables more principled control of the signal smoothness for alignment without having to rely on heuristic blurring on images, making these forms generalizable to 3D scene representations (Sec. 3.2).

2 Neural Radiance Fields (3D)

We discuss the 3D case of recovering the 3D scene representation from Neural Radiance Fields (NeRF) jointly with the camera poses. To signify the analogy to Sec. 3.1, we deliberately overload the notations x\mathbf{x} as 3D points, W\mathcal{W} as camera pose transformations, and ff as the network in NeRF.

Given MM images {Ii}i=1M\{\mathcal{I}_{i}\}_{i=1}^{M}, our goal is to optimize NeRF and the camera poses {pi}i=1M\{\mathbf{p}_{i}\}_{i=1}^{M} over the synthesis-based objective

where I^\hat{\mathcal{I}} also depends on the network parameters Θ\boldsymbol{\Theta}.

One may notice the analogy between the synthesis-based objectives of 2D image alignment (5) and NeRF (8). Similarly, we can also derive the “steepest descent image” as

which is formed via backpropagation in practice. The linearization (9) is also analogous to the 2D case of (3), where the Jacobian of the network ∂y∂x=∂f(x)∂x\frac{\partial\mathbf{y}}{\partial\mathbf{x}}=\frac{\partial f(\mathbf{x})}{\partial\mathbf{x}} linearly relates the change of color c\mathbf{c} and volume density σ\sigma with 3D spatial displacements. To solve for effective camera pose updates Δp\Delta\mathbf{p} through backpropagation, it is also desirable to control the smoothness of ff for predicting coherent geometric displacements from the sampled 3D points {x1,…,xN}\{\mathbf{x}_{1},\dots,\mathbf{x}_{N}\}.

3 On Positional Encoding and Registration

where the kk-th frequency encoding γk(x)\gamma_{k}(\mathbf{x}) is

with the sinusoidal functions operating coordinate-wise. The special case of L=0L=0 makes γ\gamma an identity mapping function. The network ff is thus a composition of f(x)=f′∘γ(x)f(\mathbf{x})=f^{\prime}\circ\gamma(\mathbf{x}), where f′f^{\prime} is the subsequent learnable MLP. Positional encoding allows coordinate-based neural networks, which are typically bandwidth limited, to represent signals of higher frequency with faster convergence behaviors .

The Jacobian of the kk-th positional encoding γk\gamma_{k} is

which amplifies the gradient signals from the MLP f′f^{\prime} by 2kπ2^{k}\pi with its direction changing at the same frequency. This makes it difficult to predict effective updates Δp\Delta\mathbf{p}, since gradient signals from the sampled 3D points are incoherent (in terms of both direction and magnitude) and can easily cancel out each other. Therefore, naïvely applying positional encoding can become a double-edged sword to NeRF for the task of joint registration and reconstruction.

4 Bundle-Adjusting Neural Radiance Fields

We describe our proposed BARF, a simple yet effective strategy for coarse-to-fine registration for NeRF. The key idea is to apply a smooth mask on the encoding at different frequency bands (from low to high) over the course of optimization, which acts like a dynamic low-pass filter. Inspired by recent work of learning coarse-to-fine deformation flow fields , we weigh the kk-th frequency component of γ\gamma as

and α∈[0,L]\alpha\in[0,L] is a controllable parameter proportional to the optimization progress. The Jacobian of γk\gamma_{k} thus becomes

When wk(α)=0w_{k}(\alpha)=0, the contribution to the gradient from the kk-th (and higher) frequency component is nullified.

Starting from the raw 3D input x\mathbf{x} (α=0\alpha=0), we gradually activate the encodings of higher frequency bands until full positional encoding is enabled (α=L\alpha=L), equivalent to the original NeRF model. This allows BARF to discover the correct registration with an initially smooth signal and later shift focus to learning a high-fidelity scene representation.

Experiments

We validate the effectiveness of our proposed BARF with a simple experiment of 2D planar image alignment, and show how the same coarse-to-fine registration strategy can be generalized to NeRF for learning 3D scene representations.

Experimental settings. We investigate how positional encoding impacts this problem by comparing networks with naïve (full) positional encoding and without any encoding. We use a simple ReLU MLP for ff with four 256-dimensional hidden units, and we use the Adam optimizer to optimize both the network weights and the warp parameters for 50005000 iterations with a learning rate of 0.0010.001. For BARF, we linearly adjust α\alpha for the first 20002000 iterations and activate all frequency bands (L=8L=8) for the remaining iterations.

Results. We visualize the registration results in Fig. 4. Alignment with full positional encoding results in suboptimal registration with ghostly artifacts in the recovered image representation. On the other hand, alignment without positional encoding achieves decent registration results, but cannot recover the image with sufficient fidelity. BARF discovers the precise geometric warps with the image representation optimized with high fidelity, quantitatively reflected in Table 1. The image alignment experiment demonstrates the general advantage of BARF for coordinate-based representations.

2 NeRF (3D): Synthetic Objects

We investigate the problem of learning 3D scene representations with Neural Radiance Fields (NeRF) from imperfect camera poses. We experiment with the 8 synthetic object-centric scenes provided by Mildenhall et al. , which consists of M=100M=100 rendered images with ground-truth camera poses for each scene for training.

Experimental settings. We parametrize the camera poses p\mathbf{p} with the se(3)\mathfrak{se}(3) Lie algebra and assume known intrinsics. For each scene, we synthetically perturb the camera poses with additive noise δp∼N(0,0.15I)\delta\mathbf{p}\sim\mathcal{N}(\mathbf{0},0.15\mathbf{I}), which corresponds to a standard deviation of 14.9°14.9\degree in rotation and 0.260.26 in translational magnitude (Fig. 5(a)). We optimize the objective in (8) jointly for the scene representation and the camera poses. We evaluate BARF mainly against the original NeRF model with naïve (full) positional encoding; for completeness, we also compare with the same model without positional encoding.

Implementation details. We follow the architectural settings from the original NeRF with some modifications. We train a single MLP with 128128 hidden units in each layer and without additional hierarchical sampling for simplicity. We resize the images to 400×400400\times 400 pixels and randomly sample 10241024 pixel rays at each optimization step. We choose N=128N=128 sample for numerical integration along each ray, and we use the softplus activation on the volume density output σ\sigma for improved stability. We use the Adam optimizer and train all models for 200200K iterations, with a learning rate of 5 ⁣× ⁣10−45\!\times\!10^{-4} exponentially decaying to 1 ⁣× ⁣10−41\!\times\!10^{-4} for the network ff and 1 ⁣× ⁣10−31\!\times\!10^{-3} decaying to 1 ⁣× ⁣10−51\!\times\!10^{-5} for the poses p\mathbf{p}. For BARF, we linearly adjust α\alpha from iteration 2020K to 100100K and activate all frequency bands (up to L=10L=10) subsequently.

Evaluation criteria. We measure the performance in two aspects: pose error for registration and view synthesis quality for the scene representation. Since both the scene and camera poses are variable up to a 3D similarity transformation, we evaluate the quality of registration by pre-aligning the optimized poses to the ground truth with Procrustes analysis on the camera locations. For evaluating view synthesis, we run an additional step of test-time photometric optimization on the trained models to factor out the pose error that may contaminate the view synthesis quality. We report the average rotation and translation errors for pose and PSNR, SSIM and LPIPS for view synthesis.

Results. We visualize the results in Fig. 6 and report the quantitative results in Table 2. BARF takes the best of both worlds of recovering the neural scene representation with the camera pose successfully registered, while naïve NeRF with full positional encoding finds suboptimal solutions. Fig. 5 shows that BARF can achieve near-perfect registration for the synthetic scenes. Although the NeRF model without positional encoding can also successfully recover alignment, the learned scene representations (and thus the synthesized images) lack the reconstruction fidelity. As a reference, we also compare the view synthesis quality against standard NeRF models trained under ground-truth poses, showing that BARF can achieve comparable view synthesis quality in all metrics, albeit initialized from imperfect camera poses.

3 NeRF (3D): Real-World Scenes

We investigate the challenging problem of learning neural 3D representations with NeRF on real-world scenes, where the camera poses are unknown. We consider the LLFF dataset , which consists of 8 forward-facing scenes with RGB images sequentially captured by hand-held cameras.

Experimental settings. We parametrize the camera poses p\mathbf{p} with se(3)\mathfrak{se}(3) following Sec. 4.2 but initialize all cameras with the identity transformation, i.e. pi=0    ∀i\mathbf{p}_{i}=\mathbf{0}\;\;\forall i. We assume known camera intrinsics (provided by the dataset). We compare against the original NeRF model with naïve positional encoding, and we use the same evaluation criteria described in Sec. 4.2. However, we note that the camera poses provided in LLFF are also estimations from SfM packages ; therefore, the pose evaluation is at most an indication of how well BARF agrees with classical geometric pose estimation.

Implementation details. We follow the same architectural settings from the original NeRF and resize the images to 480×640480\times 640 pixels. We train all models for 200200K iterations and randomly sample 20482048 pixel rays at each optimization step, with a learning rate of 1 ⁣× ⁣10−31\!\times\!10^{-3} for the network ff decaying to 1 ⁣× ⁣10−41\!\times\!10^{-4}, and 3 ⁣× ⁣10−33\!\times\!10^{-3} for the pose p\mathbf{p} decaying to 1 ⁣× ⁣10−51\!\times\!10^{-5}. We linearly adjust α\alpha for BARF from iteration 2020K to 100100K and activate all bands (up to L=10L=10) subsequently.

Results. The quantitative results (Table 3) show that the recovered camera poses from BARF highly agrees with those estimated from off-the-shelf SfM methods (visualized in Fig. 8), demonstrating the ability of BARF to localize from scratch. Furthermore, BARF can successfully recover the 3D scene representation with high fidelity (Fig. 7). In contrast, NeRF with naïve positional encoding diverge to incorrect camera poses, which in turn results in poor view synthesis. This highlights the effectiveness of BARF utilizing a coarse-to-fine strategy for joint registration and reconstruction.

Conclusion

We present Bundle-Adjusting Neural Radiance Fields (BARF), a simple yet effective strategy for training NeRF from imperfect camera poses. By establishing a theoretical connection to classical image alignment, we demonstrate that coarse-to-fine registration is necessary for joint registration and reconstruction with coordinate-based scene representations. Our experiments show that BARF can effectively learn the 3D scene representations from scratch and resolve large camera pose misalignment at the same time.

Despite the intriguing results at the current stage, BARF has similar limitations to the original NeRF formulation (e.g. slow optimization and rendering, rigidity assumption, sensitivity to dense 3D sampling), as well as reliance on heuristic coarse-to-fine scheduling strategies. Nevertheless, since BARF keeps a close formulation to NeRF, many of the latest advances on improving NeRF are potentially transferable to BARF as well. We believe BARF opens up exciting avenues for rethinking visual localization for SfM/SLAM systems and self-supervised dense 3D reconstruction frameworks using view synthesis as a proxy objective.

Acknowledgements. We thank Chaoyang Wang, Mengtian Li, Yen-Chen Lin, Tongzhou Wang, Sivabalan Manivasagam, and Shenlong Wang for helpful discussions and feedback on the paper. This work was supported by the CMU Argo AI Center for Autonomous Vehicle Research.

A Visualizing the Basin of Attraction

We visualize the results in Fig. 9. Naïve positional encoding results in a more nonlinear alignment landscape and a smaller basin of attraction, while not using positional encoding sacrifices the reconstruction quality due to the limited representability of the network ff. In contrast, BARF can widen the basin of attraction while reconstructing the image representation with high fidelity. This also justifies the importance of coarse-to-fine registration for NeRF in the 3D case. Please also refer to the supplementary videos for more visualizations of the basin of attraction.

B Additional NeRF Details & Results

We provide more details and results from our NeRF experiments in this section (for real-world scenes in particular).

As mentioned in the main paper, the optimized solutions of the 3D scenes and camera poses are up to a 3D similarity transformation. Therefore, we evaluate the quality of registration by pre-aligning the optimized poses to the reference poses, which are the ground truth poses for the synthetic objects (Sec. 4.2) and pose estimation computed from SfM packages for the real-world scenes (Sec. 4.3).

We use Procrustes analysis on the camera locations for aligning the coordinate systems. The algorithm details are described in Alg. 1. We write the reference poses {[Ri,ti]}i=1M\{[\mathbf{R}_{i},\mathbf{t}_{i}]\}_{i=1}^{M} and the optimized poses {[R^i,t^i]}i=1M\{[\widehat{\mathbf{R}}_{i},\hat{\mathbf{t}}_{i}]\}_{i=1}^{M} in the form of camera extrinsic matrices, and the aligned poses can be written as {[R^i′,t^i′]}i=1M=\scPreAlign({[Ri,ti]}i=1M,{[R^i,t^i]}i=1M)\{[\widehat{\mathbf{R}}^{\prime}_{i},\hat{\mathbf{t}}^{\prime}_{i}]\}_{i=1}^{M}=\text{\sc PreAlign}(\{[\mathbf{R}_{i},\mathbf{t}_{i}]\}_{i=1}^{M},\{[\widehat{\mathbf{R}}_{i},\hat{\mathbf{t}}_{i}]\}_{i=1}^{M}). After the cameras are Procrustes-aligned, we apply the relative rotation (solved for via the Procrustes analysis process) to account for rotational differences. We measure the rotation error between the SfM poses and the aligned poses from NeRF/BARF by the angular distance as

where ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle is the quaternion inner product. For additional clarity, we provide a more detailed visualization of the optimized camera poses in Fig. 10 (for the LLFF dataset).

To evaluate the quality of novel view synthesis while being minimally affected by camera misalignment, we transform the test views (provided by Mildenhall et al. ) to the coordinate system of the optimized poses by applying the scale/rotation/translation from the Procrustes analysis, as in Alg. 1. The camera trajectories from the baseline NeRF with naïve full positional encoding exhibits large rotational and translational differences compared to SfM poses in general. For this reason, the view synthesis results from the baseline NeRF, whose corresponding test views are also determined using Procrustes analysis, are far from plausible. Unfortunately, there is no other systematic way of determining what the corresponding views held out from the SfM poses would be in the learned coordinate system. Nevertheless, we provide additional qualitative results in Fig. 11, where the novel views are selected from a training view closest to the average pose and sampling translational perturbations. Please also see the supplementary video for more details.

B.2 Real-World Scenes (LLFF Dataset)

Dataset. The LLFF dataset consists of 8 forward-facing scenes with RGB images sequentially captured by hand-held cameras. In the original NeRF paper , the test views were selected by holding out every 8th frame from the video sequence and training with the remaining frames. Unlike Mildenhall et al. , however, we hold out the last 10%10\% of the frames for evaluation and train with the first 90%90\% frames. This train/test split does not assume that the held-out views are interpolations of the training views, which allows a more practical simulation of predicting future viewpoints from previous observations. The statistics of the train/test split for each scene is provided in Table 4.

Full comparison. We provide a more complete evaluation of the LLFF experiment in Table 5, where we also include the baseline without any positional encoding. Note that we consider the same schedule for all scenes in the dataset (adjusting the positional encoding from iterations 2020K to 100100K); due to the per-scene optimization nature, however, the optimal coarse-to-fine scheduling for each scene would actually be data-dependent. Despite this, the coarse-to-fine scheduling considered here already allows BARF to achieve an averaged similar or better performance on real-world scenes. An exhaustive analysis of searching for the best scheduling is currently out of scope of this paper.

In the main LLFF experiments, we sample 3D points along each ray linearly in the inverse depth (disparity) space, where the lower and upper bounds are the image plane and infinity respectively (i.e. 1/znear=11/z_{\text{near}}=1 and 1/zfar=01/z_{\text{far}}=0). To analyze the effect of depth parametrization on the performance of real-world scenes, we run an additional set of the same experiments by sampling the 3D points in the regular (metric) depth space, bounded by znear=1z_{\text{near}}=1 and zfar=20z_{\text{far}}=20.

We report the quantitative results in Table 6. The baseline NeRF with full positional encoding still performs poorly in all metrics. Although the baseline without positional encoding may be slightly better than BARF in this setup, all methods being compared here exhibit better performance when the 3D points are sampled in the inverse depth space. We present empirical results as a supplement and leave a complete analysis of depth parametrization to future work.

References