Block-NeRF: Scalable Large Scene Neural View Synthesis

Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul P. Srinivasan, Jonathan T. Barron, Henrik Kretzschmar

Introduction

Recent advancements in neural rendering such as Neural Radiance Fields have enabled photo-realistic reconstruction and novel view synthesis given a set of posed camera images . Earlier works tended to focus on small-scale and object-centric reconstruction. Though some methods now address scenes the size of a single room or building, these are generally still limited and do not naïvely scale up to city-scale environments. Applying these methods to large environments typically leads to significant artifacts and low visual fidelity due to limited model capacity.

Reconstructing large-scale environments enables several important use-cases in domains such as autonomous driving and aerial surveying . One example is mapping, where a high-fidelity map of the entire operating domain is created to act as a powerful prior for a variety of problems, including robot localization, navigation, and collision avoidance. Furthermore, large-scale scene reconstructions can be used for closed-loop robotic simulations . Autonomous driving systems are commonly evaluated by re-simulating previously encountered scenarios; however, any deviation from the recorded encounter may change the vehicle’s trajectory, requiring high-fidelity novel view renderings along the altered path. Beyond basic view synthesis, scene conditioned NeRFs are also capable of changing environmental lighting conditions such as camera exposure, weather, or time of day, which can be used to further augment simulation scenarios.

Reconstructing such large-scale environments introduces additional challenges, including the presence of transient objects (cars and pedestrians), limitations in model capacity, along with memory and compute constraints. Furthermore, training data for such large environments is highly unlikely to be collected in a single capture under consistent conditions. Rather, data for different parts of the environment may need to be sourced from different data collection efforts, introducing variance in both scene geometry (e.g., construction work and parked cars), as well as appearance (e.g., weather conditions and time of day).

We extend NeRF with appearance embeddings and learned pose refinement to address the environmental changes and pose errors in the collected data. We additionally add exposure conditioning to provide the ability to modify the exposure during inference. We refer to this modified model as a Block-NeRF. Scaling up the network capacity of Block-NeRF enables the ability to represent increasingly large scenes. However this approach comes with a number of limitations; rendering time scales with the size of the network, networks can no longer fit on a single compute device, and updating or expanding the environment requires retraining the entire network.

To address these challenges, we propose dividing up large environments into individually trained Block-NeRFs, which are then rendered and combined dynamically at inference time. Modeling these Block-NeRFs independently allows for maximum flexibility, scales up to arbitrarily large environments and provides the ability to update or introduce new regions in a piecewise manner without retraining the entire environment as demonstrated in Figure 1. To compute a target view, only a subset of the Block-NeRFs are rendered and then composited based on their geographic location compared to the camera. To allow for more seamless compositing, we propose an appearance matching technique which brings different Block-NeRFs into visual alignment by optimizing their appearance embeddings.

Related Work

Researchers have been developing and refining techniques for 3D reconstruction from large image collections for decades , and much current work relies on mature and robust software implementations such as COLMAP to perform this task . Nearly all of these reconstruction methods share a common pipeline: extract 2D image features (such as SIFT ), match these features across different images, and jointly optimize a set of 3D points and camera poses to be consistent with these matches (the well-explored problem of bundle adjustment ). Extending this pipeline to city-scale data is largely a matter of implementing highly robust and parallelized versions of these algorithms, as explored in work such as Photo Tourism and Building Rome in a Day . Core graphics research has also explored breaking up scenes for fast high quality rendering .

These approaches typically output a camera pose for each input image and a sparse 3D point cloud. To get a complete 3D scene model, these outputs must be further processed by a dense multi-view stereo algorithm (e.g., PMVS ) to produce a dense point cloud or triangle mesh. This process presents its own scaling difficulties . The resulting 3D models often contain artifacts or holes in areas with limited texture or specular reflections as they are challenging to triangulate across images. As such, they frequently require further postprocessing to create models that can be used to render convincing imagery . However, this task is mainly the domain of novel view synthesis, and 3D reconstruction techniques primarily focus on geometric accuracy.

In contrast, our approach does not rely on large-scale SfM to produce camera poses, instead performing odometry using various sensors on the vehicle as the images are collected .

2 Novel View Synthesis

Given a set of input images of a given scene and their camera poses, novel view synthesis seeks to render observed scene content from previously unobserved viewpoints, allowing a user to navigate through a recreated environment with high visual fidelity.

Many approaches to view synthesis start by applying traditional 3D reconstruction techniques to build a point cloud or triangle mesh representing the scene. This geometric “proxy” is then used to reproject pixels from the input images into new camera views, where they are blended by heuristic or learning-based methods . This approach has been scaled to long trajectories of first-person video , panoramas collected along a city street , and single landmarks from the Photo Tourism dataset . Methods reliant on geometry proxies are limited by the quality of the initial 3D reconstruction, which hurts their performance in scenes with complex geometry or reflectance effects.

Recent view synthesis work has focused on unifying reconstruction and rendering and learning this pipeline end-to-end, typically using a volumetric scene representation. Methods for rendering small baseline view interpolation often use feed-forward networks to learn a mapping directly from input images to an output volume , while methods such as Neural Volumes that target larger-baseline view synthesis run a global optimization over all input images to reconstruct every new scene, similar to traditional bundle adjustment.

Neural Radiance Fields (NeRF) combines this single-scene optimization setting with a neural scene representation capable of representing complex scenes much more efficiently than a discrete 3D voxel grid; however, its rendering model scales very poorly to large-scale scenes in terms of compute. Followup work has proposed making NeRF more efficient by partitioning space into smaller regions, each containing its own lightweight NeRF network . Unlike our method, these network ensembles must be trained jointly, limiting their flexibility. Another approach is to provide extra capacity in the form of a coarse 3D grid of latent codes . This approach has also been applied to compress detailed 3D shapes into neural signed distance functions and to represent large scenes using occupancy networks .

We build our Block-NeRF implementation on top of mip-NeRF , which improves aliasing issues that hurt NeRF’s performance in scenes where the input images observe the scene from many different distances. We incorporate techniques from NeRF in the Wild (NeRF-W) , which adds a latent code per training image to handle inconsistent scene appearance when applying NeRF to landmarks from the Photo Tourism dataset. NeRF-W creates a separate NeRF for each landmark from thousands of images, whereas our approach combines many NeRFs to reconstruct a coherent large environment from millions of images. Our model also incorporates a learned camera pose refinement which has been explored in previous works .

Some NeRF-based methods use segmentation data to isolate and reconstruct static or moving objects (such as people or cars) across video sequences. As we focus primarily on reconstructing the environment itself, we choose to simply mask out dynamic objects during training.

3 Urban Scene Camera Simulation

Camera simulation has become a popular data source for training and validating autonomous driving systems on interactive platforms . Early works synthesized data from scripted scenarios and manually created 3D assets. These methods suffered from domain mismatch and limited scene-level diversity. Several recent works tackle the simulation-to-reality gaps by minimizing the distribution shifts in the simulation and rendering pipeline. Kar et al. and Devaranjan et al. proposed to minimize the scene-level distribution shift from rendered outputs to real camera sensor data through a learned scenario generation framework. Richter et al. leveraged intermediate rendering buffers in the graphics pipeline to improve photorealism of synthetically generated camera images.

Towards the goal of building photo-realistic and scalable camera simulation, prior methods leverage rich multi-sensor driving data collected during a single drive to reconstruct 3D scenes for object injection and novel view synthesis using modern machine learning techniques, including image GANs for 2D neural rendering. Relying on a sophisticated surfel reconstruction pipeline, SurfelGAN is still susceptible to errors in graphical reconstruction and can suffer from the limited range and vertical field-of-view of LiDAR scans. In contrast to existing efforts, our work tackles the 3D rendering problem and is capable of modeling the real camera data captured from multiple drives under varying environmental conditions, such as weather and time of day, which is a prerequisite for reconstructing large-scale areas.

Background

We build upon NeRF and its extension mip-NeRF . Here, we summarize relevant parts of these methods. For details, please refer to the original papers.

Neural Radiance Fields (NeRF) is a coordinate-based neural scene representation that is optimized through a differentiable rendering loss to reproduce the appearance of a set of input images from known camera poses. After optimization, the NeRF model can be used to render previously unseen viewpoints.

The NeRF scene representation is a pair of multilayer perceptrons (MLPs). The first MLP fσf_{\sigma} takes in a 3D position x\mathbf{x} and outputs volume density σ\sigma and a feature vector. This feature vector is concatenated with a 2D viewing direction d\mathbf{d} and fed into the second MLP fcf_{c}, which outputs an RGB color c\mathbf{c}. This architecture ensures that the output color can vary when observed from different angles, allowing NeRF to represent reflections and glossy materials, but that the underlying geometry represented by σ\sigma is only a function of position.

Each pixel in an image corresponds to a ray r(t)=o+td\mathbf{r}(t)=\mathbf{o}+t\mathbf{d} through 3D space. To calculate the color of r\mathbf{r}, NeRF randomly samples distances {ti}i=0N\{t_{i}\}_{i=0}^{N} along the ray and passes the points r(ti)\mathbf{r}(t_{i}) and direction d\mathbf{d} through its MLPs to calculate σi\sigma_{i} and ci\mathbf{c}_{i}. The resulting output color is

The full implementation of NeRF iteratively resamples the points tit_{i} (by treating the weights wiw_{i} as a probability distribution) in order to better concentrate samples in areas of high density.

To enable the NeRF MLPs to represent higher frequency detail , the inputs x\mathbf{x} and d\mathbf{d} are each preprocessed by a componentwise sinusoidal positional encoding γPE\gamma_{\textrm{PE}}:

where LL is the number of levels of positional encoding.

NeRF’s MLP fσf_{\sigma} takes a single 3D point as input. However, this ignores both the relative footprint of the corresponding image pixel and the length of the interval [ti−1,ti][t_{i-1},t_{i}] along the ray r\mathbf{r} containing the point, resulting in aliasing artifacts when rendering novel camera trajectories. Mip-NeRF remedies this issue by using the projected pixel footprint to sample conical frustums along the ray rather than intervals. To feed these frustums into the MLP, mip-NeRF approximates each of them as Gaussian distributions with parameters μi,Σi\boldsymbol{\mu}_{i},\boldsymbol{\Sigma}_{i} and replaces the positional encoding γPE\gamma_{\textrm{PE}} with its expectation over the input Gaussian

referred to as an integrated positional encoding.

Method

Training a single NeRF does not scale when trying to represent scenes as large as cities. We instead propose splitting the environment into a set of Block-NeRFs that can be independently trained in parallel and composited during inference. This independence enables the ability to expand the environment with additional Block-NeRFs or update blocks without retraining the entire environment (see Figure 1). We dynamically select relevant Block-NeRFs for rendering, which are then composited in a smooth manner when traversing the scene. To aid with this compositing, we optimize the appearances codes to match lighting conditions and use interpolation weights computed based on each Block-NeRF’s distance to the novel view.

The individual Block-NeRFs should be arranged to collectively ensure full coverage of the target environment. We typically place one Block-NeRF at each intersection, covering the intersection itself and any connected street 75% of the way until it converges into the next intersection (see Figure 1). This results in a 50% overlap between any two adjacent blocks on the connecting street segment, making appearance alignment easier between them. Following this procedure means that the block size is variable; where necessary, additional blocks may be introduced as connectors between intersections. We ensure that the training data for each block stays exactly within its intended bounds by applying a geographical filter. This procedure can be automated and only relies on basic map data such as OpenStreetMap .

Note that other placement heuristics are also possible, as long as the entire environment is covered by at least one Block-NeRF. For example, for some of our experiments, we instead place blocks along a single street segment at uniform distances and define the block size as a sphere around the Block-NeRF Origin (see Figure 2).

2 Training Individual Block-NeRFs

Given that different parts of our data may be captured under different environmental conditions, we follow NeRF-W and use Generative Latent Optimization to optimize per-image appearance embedding vectors, as shown in Figure 3. This allows the NeRF to explain away several appearance-changing conditions, such as varying weather and lighting. We can additionally manipulate these appearance embeddings to interpolate between different conditions observed in the training data (such as cloudy versus clear skies, or day and night). Examples of rendering with different appearances can be seen in Figure 4. In § 4.3.3, we use test-time optimization over these embeddings to match the appearance of adjacent Block-NeRFs, which is important when combining multiple renderings.

2.2 Learned Pose Refinement

Although we assume that camera poses are provided, we find it advantageous to learn regularized pose offsets for further alignment. Pose refinement has been explored in previous NeRF based models . These offsets are learned per driving segment and include both a translation and a rotation component. We optimize these offsets jointly with the NeRF itself, significantly regularizing the offsets in the early phase of training to allow the network to first learn a rough structure prior to modifying the poses.

2.3 Exposure Input

2.4 Transient Objects

While our method accounts for variation in appearance using the appearance embeddings, we assume that the scene geometry is consistent across the training data. Any movable objects (e.g. cars, pedestrians) typically violate this assumption. We therefore use a semantic segmentation model to produce masks of common movable objects, and ignore masked areas during training. While this does not account for changes in otherwise static parts of the environment, e.g. construction, it accommodates most common types of geometric inconsistency.

2.5 Visibility Prediction

When merging multiple Block-NeRFs, it can be useful to know whether a specific region of space was visible to a given NeRF during training. We extend our model with an additional small MLP fvf_{v} that is trained to learn an approximation of the visibility of a sampled point (see Figure 3). For each sample along a training ray, fvf_{v} takes in the location and view direction and regresses the corresponding transmittance of the point (TiT_{i} in Equation 2). The model is trained alongside fσf_{\sigma}, which provides supervision. Transmittance represents how visible a point is from a particular input camera: points in free space or on the surface of the first intersected object will have transmittance near 1, and points inside or behind the first visible object will have transmittance near 0. If a point is seen from some viewpoints but not others, the regressed transmittance value will be the average over all training cameras and lie between zero and one, indicating that the point is partially observed. Our visibility prediction is similar to the visibility fields proposed by Srinivasan et al. . However, they used an MLP to predict visibility to environment lighting for the purpose of recovering a relightable NeRF model, while we predict visibility to training rays.

The visibility network is small and can be run independently from the color and density networks. This proves useful when merging multiple NeRFs, since it can help to determine whether a specific NeRF is likely to produce meaningful outputs for a given location, as explained in § 4.3.1. The visibility predictions can also be used to determine locations to perform appearance matching between two NeRFs, as detailed in § 4.3.3.

3 Merging Multiple Block-NeRFs

The environment can be composed of an arbitrary number of Block-NeRFs. For efficiency, we utilize two filtering mechanisms to only render relevant blocks for the given target viewpoint. We only consider Block-NeRFs that are within a set radius of the target viewpoint. Additionally, for each of these candidates, we compute the associated visibility. If the mean visibility is below a threshold, we discard the Block-NeRF. An example of visibility filtering is provided in Figure 2. Visibility can be computed quickly because its network is independent of the color network, and it does not need to be rendered at the target image resolution. After filtering, there are typically one to three Block-NeRFs left to merge.

3.2 Block-NeRF Compositing

We render color images from each of the filtered Block-NeRFs and interpolate between them using inverse distance weighting between the camera origin cc and the centers xix_{i} of each Block-NeRF. Specifically, we calculate the respective weights as wi∝distance⁡(c,xi)−pw_{i}\propto\operatorname{distance}(c,x_{i})^{-p}, where pp influences the rate of blending between Block-NeRF renders. The interpolation is done in 2D image space and produces smooth transitions between Block-NeRFs. We also explore other interpolation methods in § 5.4.

3.3 Appearance Matching

The appearance of our learned models can be controlled by an appearance latent code after the Block-NeRF has been trained. These codes are randomly initialized during training and therefore the same code typically leads to different appearances when fed into different Block-NeRFs. This is undesirable when compositing as it may lead to inconsistencies between views. Given a target appearance in one of the Block-NeRFs, we aim to match its appearance in the remaining blocks. To accomplish this, we first select a 3D matching location between pairs of adjacent Block-NeRFs. The visibility prediction at this location should be high for both Block-NeRFs.

The optimized appearance is iteratively propagated through the scene. Starting from one root Block-NeRF, we optimize the appearance of the neighboring ones and continue the process from there. If multiple blocks surrounding a target Block-NeRF have already been optimized, we consider each of them when computing the loss.

Results and Experiments

In this section we will discuss our datasets and experiments. The architectural and optimization specifics are provided in the supplement. The supplement also provides comparisons to reconstructions from COLMAP , a traditional Structure from Motion approach. This reconstruction is sparse and fails to represent reflective surfaces and the sky.

We perform experiments on datasets that we collect specifically for the task of novel view synthesis of large-scale scenes. Our dataset is collected on public roads using data collection vehicles. While several large-scale driving datasets already exist, they are not designed for the task of view synthesis. For example, some datasets lack sufficient camera coverage (e.g., KITTI , Cityscapes ) or prioritize visual diversity over repeated observations of a target area (e.g., NuScenes , Waymo Open Dataset , Argoverse ). Instead, they are typically designed for tasks such as object detection or tracking, where similar observations across drives can lead to generalization issues.

2 Model Ablations

We ablate our model modifications on a single intersection from the Alamo Square dataset. We report PSNR, SSIM, and LPIPS metrics for the test image reconstructions in Table 1. The test images are split in half vertically, with the appearance embeddings being optimized on one half and tested on the other. We also provide qualitative examples in Figure 7. Mip-NeRF alone fails to properly reconstruct the scene and is prone to adding non-existent geometry and cloudy artifacts to explain the differences in appearance. When our method is not trained with appearance embeddings, these artifacts are still present. If our method is not trained with pose optimization, the resulting scene is blurrier and can contain duplicated objects due to pose misalignment. Finally, the exposure input marginally improves the reconstruction, but more importantly provides us with the ability to change the exposure during inference.

3 Block-NeRF Size and Placement

4 Interpolation Methods

We explore different interpolation methods in Table 3. The simple method of only rendering the nearest Block-NeRF to the camera requires the least amount of compute but results in harsh jumps when transitioning between blocks. These transitions can be smoothed by using inverse distance weighting (IDW) between the camera and Block-NeRF centers, as described in § 4.3.2. We also explored a variant of IDW where the interpolation was performed over projected 3D points predicted by the expected Block-NeRF depth. This method suffers when the depth prediction is incorrect, leading to artifacts and temporal incoherence.

Finally, we experiment with weighing the Block-NeRFs based on per-pixel and per-image predicted visibility. This produces sharper reconstructions of further-away areas but is prone to temporal inconsistency. Therefore, these methods are best used only when rendering still images. Further details are provided in the supplement.

Limitations and Future Work

The proposed method handles transient objects by filtering them out during training via masking using a segmentation algorithm. If objects are not properly masked, they can cause artifacts in the resulting renderings. For example, the shadows of cars often remain, even when the car itself is correctly removed. Vegetation also breaks this assumption as foliage changes seasonally and moves in the wind; this results in blurred representations of trees and plants. Similarly, temporal inconsistencies in the training data, such as construction work, are not automatically handled and require the manual retraining of the affected blocks. Further, the inability to render scenes containing dynamic objects currently limits the applicability of Block-NeRF towards closed-loop simulation tasks in robotics. In the future, these issues could be addressed by learning transient objects during the optimization , or directly modeling dynamic objects . In particular, the scene could be composed of multiple Block-NeRFs of the environment and individual controllable object NeRFs. Separation can be facilitated by the use of segmentation masks or bounding boxes.

In our model, distant objects in the scene are not sampled with the same density as nearby objects which leads to blurrier reconstructions. This is an issue with sampling unbounded volumetric representations. Techniques proposed in NeRF++ and concurrent Mip-NeRF 360 could potentially be used to produce sharper renderings of distant objects.

In many applications, real-time rendering is key, but NeRFs are computationally expensive to render (up to multiple seconds per image). Several NeRF caching techniques or a sparse voxel grid could be used to enable real-time Block-NeRF rendering. Similarly, multiple concurrent works have demonstrated techniques to speed up training of NeRF style representations by multiple orders of magnitude .

Conclusion

In this paper we propose Block-NeRF, a method that reconstructs arbitrarily large environments using NeRFs. We demonstrate the method’s efficacy by building an entire neighborhood in San Francisco from 2.8M images, forming the largest neural scene representation to date. We accomplish this scale by splitting our representation into multiple blocks that can be optimized independently. At such a scale, the data collected will necessarily have transient objects and variations in appearance, which we account for by modifying the underlying NeRF architecture. We hope that this can inspire future work in large-scale scene reconstruction using modern neural rendering methods.

References

Appendix A Model Parameters / Optimization Details

Each Block-NeRF takes between 9 and 24 hours to train (depending on hyperparameters). We train each Block-NeRF on 32 TPU v3 cores available through Google Cloud Compute, which combined offer a total of 16801680 TFLOPS and 512512 GB memory. Rendering an 1200×9001200\times 900px image for a single Block-NeRF takes approximately 5.95.9 seconds. Multiple Block-NeRF can be processed in parallel during inference (typically fewer than 3 Block-NeRFs need to be rendered for a single frame).

Appendix B Block-NeRF Size and Placement

We include qualitative comparisons in Figure 9 on the Mission Bay dataset to complement the quantitative comparisons in (§5.3, Table 2). In this figure, we provide comparisons on two regimes, one where each Block-NeRF contains the same number of weights (left section) and one where the total number of weights across all Block-NeRFs is fixed (right section).

Appendix C Block-NeRF Overlap Comparison

In the main paper, we include experiments on Block-NeRF size and placement (§5.3). For these experiments, we assumed a relative overlap of 50% between each pair of Block-NeRFs, which aids with appearance alignment.

Table 4 is a direct extension of Table 2 in the main paper and shows the effect of varying block overlap in the 8 block scenario. Note that varying the overlap changes the spatial block size. The original setting in the main paper is marked with an asterisk.

The metrics imply that reducing overlap is beneficial for image quality metrics. However, this can likely be attributed to the resulting reduction in block size. In practice, having an overlap between blocks is important to avoid temporal artifacts when interpolating between Block-NeRFs.

Appendix D Block-NeRF Interpolation Details

We experiment with multiple methods to interpolate between Block-NeRFs and find that simple inverse distance weighting (IDW) in image space produces the most appealing videos due to temporal smoothness. We use an IDW power pp of 4 for the Alamo Square renderings and a power of 1 for the Mission Bay renderings. We experiment with 3D inverse distance weighting for each individual pixel by projecting the rendered pixels into 3D space using the expected ray termination depth from the Block-NeRF closest to the target view. The color value of the projected pixel is then determined using inverse distance weighting with the nearest Block-NeRFs. Artifacts occur in the resulting composited renders due to noise in the depth predictions. We also experiment with using the Block-NeRF predicted visibility for interpolation. We consider imagewise visibility where we take the mean visibility of the entire image and pixelwise visibility where were directly utilize the per-pixel visibility predictions. Both of these methods lead to sharper results but come at the cost of temporal inconsistencies. Finally we compare to nearest neighbor interpolation where we only render the Block-NeRF closest to the target view. This results in harsh jumps when transiting between Block-NeRFs.

Appendix E Structure from Motion (COLMAP)

In Table 5, we compare two different rendering options for the densely reconstructed pointcloud. We discard the invisible pixels when computing the PSNR for both methods, making the quantitative results comparable to our Block-NeRF setting.

In Figure 8, we show the qualitative comparisons between two rendering options with PSNR on the corresponding images. This reconstruction is sparse and fails to represent reflective surfaces and the sky.

Appendix F Examples from our Datasets

In Figure 10, we show the camera images from our Mission Bay dataset. In Figure 11, we show both camera images and corresponding segmentation masks from our Alamo Square dataset.

Appendix G Societal Impact

Our method inherits the heavy compute footprint of NeRF models and we propose to apply them at an unprecedented scale. Our method also unlocks new use-cases for neural rendering, such as building detailed maps of the environment (mapping), which could cause more wide-spread use in favor of less computationally involved alternatives. Depending on the scale this work is being applied at, its compute demands can lead to or worsen environmental damage if the energy used for compute leads to increased carbon emissions. As mentioned in the paper, we foresee further work, such as caching methods, that could reduce the compute demands and thus mitigate the environmental damage.

G.2 Application

We apply our method to real city environments. During our own data collection efforts for this paper, we were careful to blur faces and sensitive information, such as license plates, and limited our driving to public roads. Future applications of this work might entail even larger data collection efforts, which raises further privacy concerns. While detailed imagery of public roads can already be found on services like Google Street View, our methodology could promote repeated and more regular scans of the environment. Several companies in the autonomous vehicle space are also known to perform regular area scans using their fleet of vehicles; however some might only utilize LiDAR scans which can be less sensitive than collecting camera imagery.