Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields Reconstruction
Cheng Sun, Min Sun, Hwann-Tzong Chen
Introduction
Achieving free-viewpoint navigation of 3D objects or scenes from only a set of calibrated images as input is a demanding task. For instance, it enables online product showcase to provide an immersive user experience comparing to static image demonstration. Recently, Neural Radiance Fields (NeRFs) have emerged as powerful representations yielding state-of-the-art quality on this task.
Despite its effectiveness in representing scenes, NeRF is known to be hampered by the need of lengthy training time and the inefficiency in rendering new views. This makes NeRF infeasible for many application scenarios. Several follow-up methods have shown significant speedup of FPS in testing phase, some of which even achieve real-time rendering. However, only few methods show training times speedup, and the improvements are not comparable to ours or lead to worse quality . On a single GPU machine, several hours of per scene optimization or a day of pretraining is typically required.
To reconstruct a volumetric scene representation from a set of images, NeRF uses multilayer perceptron (MLP) to implicitly learn the mapping from a queried 3D point (with a viewing direction) to its colors and densities. The queried properties along a camera ray can then be accumulated into a pixel color by volume rendering techniques. Our work takes inspiration from the recent success that uses classic voxel grid to explicitly store the scene properties, which enables real-time rendering and shows good quality. However, their methods can not train from scratch and need a conversion step from the trained implicit model, which causes a bottleneck to the training time.
The key to our speedup is to use a dense voxel grid to directly model the 3D geometry (volume density). Developing an elaborate strategy for view-dependent colors is not in the main scope of this paper, and we simply use a hybrid representation (feature grid with shallow MLP) for colors.
Directly optimizing the density voxel grid leads to super-fast converges but is prone to suboptimal solutions, where our method allocates “cloud” at free space and tries to fit the photometric loss with the cloud instead of searching a geometry with better multi-view consistency. Our solution to this problem is simple and effective. First, we initialize the density voxel grid to yield opacities very close to zero everywhere to avoid the geometry solutions being biased toward the cameras’ near planes. Second, we give a lower learning rate to voxels visible to fewer views, which can avoid redundant voxels that are allocated just for explaining the observations from a small number of views. We show that the proposed solutions can successfully avoid the suboptimal geometry and work well on the five datasets.
Using the voxel grid to model volume density still faces a challenge in scalability. For parsimony, our approach automatically finds a BBox tightly encloses the volume of interest to allocate the voxel grids. Besides, we propose post-activation—applying all the activation functions after trilinearly interpolating the density voxel grid. Previous work either interpolates the voxel grid for the activated opacity or uses nearest-neighbor interpolation, which results in a smooth surface in each grid cell. Conversely, we prove mathematically and empirically that the proposed post-activation can model (beyond) a sharp linear surface within a single grid cell. As a result, we can use fewer voxels to achieve better qualities—our method with dense voxels already outperforms NeRF in most cases.
In summary, we have two main technical contributions. First, we implement two priors to avoid suboptimal geometry in direct voxel density optimization. Second, we propose the post-activated voxel-grid interpolation, which enables sharp boundary modeling in lower grid resolution. The resulting key merits of this work are highlighted as follows:
We achieve visual quality comparable to NeRF at a rendering speed that is about faster.
Our method does not need cross-scene pretraining.
Our grid resolution is about , while the grid resolution in previous work ranges from to to achieve NeRF-comparable quality.
Related work
Images synthesis from novel viewpoints given a set of images capturing the scene is a long-standing task with rich studies. Previous work has presented several scene representations reconstructed from the input images to synthesize the unobserved viewpoints. Lumigraph and light field representation directly synthesize novel views by interpolating the input images but require very dense scene capture. Layered depth images work for sparse input views but rely on depth maps or estimated depth with sacrificed quality. Mesh-based representations can run in real-time but have a hard time with gradient-based optimization without template meshes provided. Recent approaches employ 2D/3D Convolutional Neural Network (CNNs) to estimate multiplane images (MPIs) for forward-facing captures; estimate voxel grid for inward-facing captures. Our method uses gradient-descent to optimize voxel grids directly and does not rely on neural networks to predict the grid values, and we still outperform the previous works with CNNs by a large margin.
Neural radiance fields.
Recently, NeRF stands out to be a prevalent method for novel view synthesis with rapid progress, which takes a moderate number of input images with known camera poses. Unlike traditional explicit and discretized volumetric representations (e.g., voxel grids and MPIs), NeRF uses coordinate-based multilayer perceptrons (MLP) as an implicit and continuous volumetric representation. NeRF achieves appealing quality and has good flexibility with many follow-up extensions to various setups, e.g., relighting , deformation , self-calibration , meta-learning , dynamic scene modeling , and generative modeling . Nevertheless, NeRF has unfavorable limitations of lengthy training progress and slow rendering speed. In this work, we mainly follow NeRF’s original setup, while our method can optimize the volume density explicitly encoded in a voxel grid to speed up both training and testing by a large margin with comparable quality.
Hybrid volumetric representations.
To combine NeRF’s implicit representation and traditional grid representations, the coordinate-based MLP is extended to also conditioning on the local feature in the grid. Recently, hybrid voxels and MPIs representations have shown success in fast rendering speed and result quality. We use hybrid representation to model view-dependent color as well.
Fast NeRF rendering.
NSVF uses octree in its hybrid representation to avoid redundant MLP queries in free space. However, NSVF still needs many training hours due to the deep MLP in its representation. Recent methods further use thousands of tiny MLPs or explicit volumetric representations to achieve real-time rendering. Unfortunately, gradient-based optimization is not directly applicable to their methods due to their topological data structures or the lack of priors. As a result, these methods still need a conversion step from a trained implicit model (e.g., NeRF) to their final representation that supports real-time rendering. Their training time is still burdened by the lengthy implicit model optimization.
Fast NeRF convergence.
Recent works that focus on fewer input views setup also bring faster convergence as a side benefit. These methods rely on generalizable pre-training or external MVS depth information , while ours does not. Further, they still require several per-scene fine-tuning hours or fail to achieve NeRF quality in the full input-view setup . Most recently, NeuRay shows NeRF’s quality with 40 minutes per-scene training time in the lower-resolution setup. Under the same GPU spec, our method achieves NeRF’s quality in 15 minutes per scene on the high-resolution setup and does not require depth guidance and cross-scene pre-training.
Preliminaries
To represent a 3D scene for novel view synthesis, Neural Radiance Fields (NeRFs) employ multilayer perceptron (MLP) networks to map a 3D position and a viewing direction to the corresponding density and view-dependent color emission :
To render the color of a pixel , we cast the ray from the camera center through the pixel; points are then sampled on between the pre-defined near and far planes; the ordered sampled points are then used to query for their densities and colors (MLPs are queried in NeRF). Finally, the queried results are accumulated into a single color with the volume rendering quadrature in accordance with the optical model given by Max :
where is the probability of termination at the point ; is the accumulated transmittance from the near plane to point ; is the distance to the adjacent sampled point, and is a pre-defined background color.
Given the training images with known poses, NeRF model is trained by minimizing the photometric MSE between the observed pixel color and the rendered color :
where is the set of rays in a sampled mini-batch.
Post-activated density voxel grid
A voxel-grid representation models the modalities of interest (e.g., density, color, or feature) explicitly in its grid cells. Such an explicit scene representation is efficient to query for any 3D positions via interpolation:
where is the queried 3D point, is the voxel grid, is the dimension of the modality, and is the total number of voxels. Trilinear interpolation is applied if not specified otherwise.
Density voxel grid for volume rendering.
where the shift is a hyperparameter. Using instead of is crucial to optimize voxel density directly, as it is irreparable when a voxel is falsely set to a negative value with as the density activation. Conversely, allows us to explore density very close to .
Sharp decision boundary via post-activation.
The interpolated voxel density is processed by (Eq. 5) and (Eq. 2b) functions sequentially for volume rendering. We consider three different orderings—pre-activation, in-activation, and post-activation—of plugging in the trilinear interpolation and performing the activation, given a queried 3D point :
The input to the function (Eq. 2b) is omitted for simplicity. We show that the post-activation, i.e., applying all the non-linear activation after the trilinear interpolation, is capable of producing sharp surfaces (decision boundaries) with much fewer grid cells. In Fig. 3, we use a 2D grid cell as an example to show that a grid cell with post-activation can produce a sharp linear boundary, while pre- and in-activation can only produce smooth results and thus require more cells for the surface detail. In Fig. 4, we further use binary image regression as a toy example to compare their capability, which also shows that post-activation can achieve a much better efficiency in grid cell usage.
Fast and direct voxel grid optimization
Typically, a scene is dominated by free space (i.e., unoccupied space). Motivated by this fact, we aim to efficiently find the coarse 3D areas of interest before reconstructing the fine detail and view-dependent effect that require more computation resources. We can thus greatly reduce the number of queried points on each ray in the later fine stage.
Coarse voxels allocation.
We first find a bounding box (BBox) tightly enclosing the camera frustums of training views (See the red BBox in Fig. 2c for an example). Our voxel grids are aligned with the BBox. Let be the lengths of the BBox and be the hyperparameter for the expected total number of voxels in the coarse stage. The voxel size is , so there are voxels on each side of the BBox.
Coarse-stage points sampling.
On a pixel-rendering ray, we sample query points as
where is the camera center, is the ray-casting direction, is the camera near bound, and is a hyperparameter for the step size that can be adaptively chosen according to the voxel size . The query index ranges from to , where is the camera far bound, so the last sampled point stops nearby the far plane.
Prior 1: low-density initialization.
At the start of training, the importance of points far from a camera is down-weighted due to the accumulated transmittance term in Eq. 2c. As a result, the coarse density voxel grid could be accidentally trapped into a suboptimal “cloudy” geometry with higher densities at camera near planes. We thus have to initialize more carefully to ensure that all sampled points on rays are visible to the cameras at the beginning, i.e., the accumulated transmittance rates s in Eq. 2c are close to .
In practice, we initialize all grid values in to and set the bias term in Eq. 5 to
where is a hyperparameter. Thereby, the accumulated transmittance is decayed by for a ray that traces forward a distance of a voxel size . See supplementary material for the derivation and proof.
Prior 2: view-count-based learning rate.
There could be some voxels visible to too few training views in real-world capturing, while we prefer a surface with consistency in many views instead of a surface that can only explain few views. In practice, we set different learning rates for different grid points in . For each grid point indexed by , we count the number of training views to which point is visible, and then scale its base learning rate by , where is the maximum view count over all grid points.
Training objective for coarse representation.
The scene representation is reconstructed by minimizing the mean square error between the rendered and observed colors. To regularize the reconstruction, we mainly use background entropy loss to encourage the accumulated alpha values to concentrate on background or foreground. Please refer to supplementary material for more detail.
2 Fine detail reconstruction
Given the optimized coarse geometry in Sec. 5.1, we now can focus on a smaller subspace to reconstruct the surface details and view-dependent effects. The optimized is frozen in this stage.
Known free space and unknown space.
A query point is in the known free space if the post-activated alpha value from the optimized is less than the threshold . Otherwise, we say the query point is in the unknown space.
Fine voxels allocation.
We densely query to find a BBox tightly enclosing the unknown space, where are the lengths of the BBox. The only hyperparameter is the expected total number of voxels . The voxel size and the grid dimensions can then be derived automatically from as per Sec. 5.1.
Progressive scaling.
Fine-stage points sampling.
The points sampling strategy is similar to Eq. 8 with some modifications. We first filter out rays that do not intersect with the known free space. For each ray, we adjust the near- and far-bound, and , to the two endpoints of the ray-box intersection. We do not adjust if is already inside the BBox.
Free space skipping.
Querying (Eq. 7a) is faster than querying (Eq. 10a); querying for view-dependent colors (Eq. 10b) is the slowest. We improve fine-stage efficiency by free space skipping in both training and testing. First, we skip sampled points that are in the known free space by checking the optimized (Eq. 7a). Second, we further skip sampled points in unknown space with low activated alpha value (threshold at ) by querying (Eq. 10a).
Training objective for fine representation.
We use the same training losses as the coarse stage, but we use a smaller weight for the regularization losses as we find it empirically leads to slightly better quality.
Experiments
We evaluate our approach on five inward-facing datasets. Synthetic-NeRF contains eight objects with realistic images synthesized by NeRF. Synthetic-NSVF contains another eight objects synthesized by NSVF. Strictly following NeRF’s and NSVF’s setups, we set the image resolution to pixels and let each scene have 100 views for training and 200 views for testing. BlendedMVS is a synthetic MVS dataset that has realistic ambient lighting from real image blending. We use a subset of four objects provided by NSVF. The image resolution is pixels, and one-eighth of the images are for testing. Tanks&Temples is a real-world dataset. We use a subset of five scenes provided by NSVF, each containing views captured by an inward-facing camera circling the scene. The image resolution is pixels, and one-eighth of the images are for testing. DeepVoxels dataset contains four simple Lambertian objects. The image resolutions are , and each scene has 479 views for training and 1000 views for testing.
2 Implementation details
We choose the same hyperparameters generally for all scenes. The expected numbers of voxels are set to and in coarse and fine stages if not stated otherwise. The activated alpha values are initialized to be in the coarse stage. We use a higher as the query points are concentrated on the optimized coarse geometry in the fine stage. The points sampling step sizes are set to half of the voxel sizes, i.e., and . The shallow MLP layer comprises two hidden layers with channels. We use the Adam optimizer with a batch size of rays to optimize the coarse and fine scene representations for k and k iterations. The base learning rates are for all voxel grids and for the shallow MLP. The exponential learning rate decay is applied. See supplementary material for detailed hyperparameter setups.
3 Comparisons
We first quantitatively compare the novel view synthesis results in Tab. 1. PSNR, SSIM , and LPIPS are employed as evaluation metrics. Our model with voxels already outperforms the original NeRF and the improved JaxNeRF re-implementation. Besides, our results are also comparable to most of the recent methods, except JaxNeRF+ and Mip-NeRF . Moreover, our per-scene optimization only takes about 15 minutes, while all the methods after NeRF in Tab. 1 need quite a few hours per scene. We also show our model with voxels, which significantly improves our results under all metrics and achieves more comparable results to JaxNeRF+ and Mip-NeRF. We defer detail comparisons on the much simpler DeepVoxels dataset to supplementary material, where we achieve averaged PSNR and outperform NeRF’s and IBRNet’s .
Training time comparisons.
The key merit of our work is the significant improvement in convergence speed with NeRF-comparable quality. In Tab. 2, we show a training time comparison. We also show GPU specifications after each reported time as it is the main factor affecting run-time.
NeRF with a more powerful GPU needs 1–2 days per scene to achieve PSNR, while our method achieves a superior and PSNR in about an minutes per scene respectively. MVSNeRF , IBRNet , and NeuRay also show less per-scene training time than NeRF but with the additional cost to run a generalizable cross-scene pre-training. MVSNeRF , after pre-training, optimizes a scene in 15 minutes as well, but the PSNR is degraded to . IBRNet shows worse PSNR and longer training time than ours. NeuRay originally reports time in lower-resolution (NeuRay-Lo) setup, and we receive the training time of the high-resolution (NeuRay-Hi) setup from the authors. NeuRay-Hi achieves PSNR and requires hours to train, while our method with voxels achieves superior in about minutes. For the early-stopped NeuRay-Hi, unfortunately, only its training time is retained (early-stopped NeuRay-Lo achieves NeRF-similar PSNR). NeuRay-Hi still needs minutes to train with early stopping, while we only need minutes to achieve NeRF-comparable quality and do not rely on generalizable pre-training or external depth information. Mip-NeRF has similar run-time to NeRF but with much better PSNRs, which also signifies using less training time to achieve NeRF’s PSNR. We train early-stopped Mip-NeRFs on our machine and show the averaged PSNR and training time. The early-stopped Mip-NeRF achieves PSNR after hours of training, while we can achieve PSNR in just minutes.
Rendering speed comparisons.
Improving test-time rendering speed is not the main focus of this work, but we still achieve speedups from NeRF— seconds versus seconds per image on our machine.
Qualitative comparison.
Fig. 5 shows our rendering results on the challenging parts and compare them with the results (better than NeRF’s) provided by PlenOctrees .
4 Ablation studies
We mainly validate the effectiveness of the two proposed techniques—post-activation and the imposed priors—that enable voxel grids to model scene geometry with NeRF-comparable quality. We subsample two scenes for each dataset. See supplementary material for more detail and additional ablation studies on the number of voxels, point-sampling step size, progressive scaling, free space skipping, view-dependent colors modeling, and the losses.
We show in Sec. 4 that the proposed post-activated trilinear interpolation enables the discretized grid to model sharper surfaces. In Tab. 3, we compare the effectiveness of post-activation in scene reconstruction for novel view synthesis. Our grid in the fine stage consists of only voxels, where nearest-neighbor interpolation results in worse quality than trilinear interpolation. The proposed post-activation can improve the results further compared to pre- and in-activation. We find that we gain less in the real-world captured BlendedMVS and Tanks and Temples datasets. The intuitive reason is that real-world data introduces more uncertainty (e.g., inconsistent lightning, SfM error), which results in multi-view inconsistent and blurrier surfaces. Thus, the advantage is lessened for scene representations that can model sharper surfaces. We speculate that resolving the uncertainty in future work can increase the gain of the proposed post-activation.
Effectiveness of the imposed priors.
As discussed in Sec. 5.1, it is crucial to initialize the voxel grid with low density to avoid suboptimal geometry. The hyperparameter controls the initial activated alpha values via Eq. 9. In Tab. 4, we compare the quality with different and the view-count-based learning rate. Without the low-density initialization, the quality drops severely for all the scenes. When , we have to train the coarse stage of some scenes for more iterations. The effective range of is scene-dependent. We find generally works well on all the scenes in this work. Finally, using a view-count-based learning rate can further improve the results and allocate noiseless voxels in the coarse stage.
Conclusion
Our method directly optimizes the voxel grid and achieves super-fast convergence in per-scene optimization with NeRF-comparable quality—reducing training time from many hours to 15 minutes. However, we do not deal with the unbounded or forward-facing scenes, while we believe our method can be a stepping stone toward fast convergence in such scenarios. We hope our method can boost the progress of NeRF-based scene reconstruction and its applications.
Acknowledgements: This work was supported in part by the MOST grants 110-2634-F-001-009 and 110-2622-8-007-010-TE2 of Taiwan. We are grateful to National Center for High-performance Computing for providing computational resources and facilities.
In Sec. H, we give more implementation details. In Sec. I, we present additional ablation studies for the important hyperparameters that affect the quality and convergence speed. In Sec. J, we show detailed results of our main ablation studies presented in the main paper. In Sec. K, we report our training time breakdown of the coarse and fine stage on each scene. In Sec. L, we show the detailed quantitative and qualitative results on each scene. Finally, we derive for the low-density initialization in Sec. M, and we prove that the proposed post-activated trilinear interpolation can produce sharp linear surfaces in Sec. N.
H Additional implementation details
We choose the same hyperparameters generally for all scenes in the five datasets.
In the coarse stage, the expected number of voxels is . The activated alpha values are initialized to be .
We set the sampling step sizes to be half of the voxel sizes, i.e., and . During training, we randomly shift the query points on a ray by , where is sampled uniformly in $\bm{d}$ is the viewing direction of the ray.
I Additional ablation experiments
We use the Robot scene to conduct the additional ablation experiments. In all the additional ablation experiments, the fine stage are early stopped at k iterations. The reported training times are measured on our machine with a single RTX2080 Ti GPU.
The hyperparameters and define the numbers of voxels in the coarse stage and the fine stage. In Tab. I1, we show the effect of different setups on the final PSNRs and the training times. We can use more voxels for better quality or use fewer voxels for faster convergence speed.
Effect of the point sampling step size.
In our design, the voxel sizes are automatically derived from the hyperparameters of the number of voxels . We set the step-size-to-voxel-size ratio instead of directly assigning the step size for points sampling on rays, i.e., means that the step size is half a voxel size. We show in Tab. I2 that the step sizes are also important hyperparameters for the speed-quality tradeoff. We can also improve the rendering quality of a trained model by using a finer step size. For the sake of simplicity, we leave future work to explore the effect of different hierarchical point sampling strategies .
Effect of the progressive scaling in the fine stage.
In the fine-stage training, we scale the resolution of our voxel grids progressively. We show in Tab. I3 that such a strategy can improve the training efficiency with a slightly better rendering quality.
Effect of the free space skipping in the fine stage.
There are three models in the fine stage—i) the optimized and frozen coarse density voxel grid , ii) the finer density voxel grid , and iii) the model for the view-dependent color emission. Querying the optimized is more efficient than querying the finer density ; querying the view-dependent color emission is the slowest and computationally expensive. For each query point, we first query and ignore the query point if the activated alpha is below a threshold (i.e., the point is in the known free space). As the training progresses, we can filter out more points that are in the free space defined by the “sculpted” . Specifically, we ignore a query point if the activated alpha from is below a threshold . Thus, we only query for the view-dependent color if the query point passes the two criteria.
In Tab. I4, we show the effect of the free space skipping strategy. In the case of , the coarse voxel grid is only used to determine a tighten BBox enclosing the scene. We finally run out of memory (OOM) at the fine stage if no skipping rule is applied (). When only is applied, we can still sculpt out many irrelevant query points during the progressive scaling and save the memory before our voxel grids are scaled to the finest resolution. The rendering quality is degraded when , which suggests that the imposed priors in the coarse stage are helpful to the final quality (we cannot apply the low-density initialization and the free space skipping at the same time). Finally, we achieve the best rendering quality and convergence speed when both and are applied.
Ablation studies for the view-dependent color modeling.
Our hybrid representation for the view-dependent color emissions in the fine stage comprises a feature voxel grid and a shallow MLP. In Tab. I5, we show different strategies and hyperparameters to model view-dependent colors. Applying only the shallow MLP results in worse outputs as our MLP is much “shallower” than the MLP in NeRF—our two layers with 128 channels versus NeRF’s eight layers with 256 channels. Using only the voxel grid to model view-invariant diffused colors can converge in much less time, but the rendering quality is degraded. In NeRF-based scene reconstruction, spherical harmonic representation , learnable basis , or deferred rendering are also shown to be effective, but we leave the explorations for future work. We also compare the number of dimensions of our feature grid, the number of layers, and the number of channels of the shallow MLP. Using a larger model leads to better rendering quality with the cost of more computation.
Effect of the auxiliary losses.
We show in Tab. I6 that incorporating the per-point rgb loss and the background entropy loss added to the main photometric loss can improve our results. The per-point rgb loss supervises the color emission of each query points directly instead of the accumulated color. The background entropy loss enforces the rendered background probability to concentrate on the background or foreground. We achieve the best results when adding the two auxiliary losses.
J Main ablation studies details
In Tab. J1, we show detailed results for the ablation experiments presented in the main paper. We subsample two scenes for each dataset to conduct the main ablation experiments—Materials and Mic from Synthetic-NeRF dataset; Robot and Lifestyle from Synthetic-NSVF dataset; Character and Statues from BlendedMVS dataset, and Ignatius and Truck from Tanks and Temples dataset.
We show that the proposed post-activation significantly improves our quality. The low-density initialization is shown to be essential to our method and is an important hyperparameter. The view-count-based learning rate can slightly improve the results, and the PSNR without the view-count prior can degrade up to in the worst case.
K Additional training time details
Mip-NeRF requires similar run-time as NeRF but achieves much better quality. As a side benefit, Mip-NeRF can attain NeRF’s PSNR in less time. To compare the convergence speed with Mip-NeRF, we use the code open-sourced by the authors (https://github.com/google/mipnerf). We re-train early-stopped Mip-NeRFs on our machine with only a single RTX2080 Ti GPU. The comparison is shown in Tab. K1. The early-stopped Mip-NeRF that is trained for 6 hours achieves an average PSNR of . On the other hand, our method is only optimized for about minutes and achieves an average PSNR of .
K.2 The detailed training time of our model
L Per-scene analysis
We report the per-scene quantitative comparisons in Tab. L1 for Synthetic-NeRF dataset, Tab. L2 for Synthetic-NSVF dataset, Tab. L3 for BlendedMVS dataset, Tab. L4 for Tanks&Temples dataset, and Tab. L5 for DeepVoxels dataset. Our method achieves comparable results to most of the recent methods, except the Mip-NeRF and JaxNeRF+ . Moreover, all the methods after NeRF on the tables take quite a few hours to train for each scene, while our method only takes about 15 minutes per-scene optimization time as reported in Tab. K2.
L.2 Per-scene qualitative results
We provide more qualitative results in Fig. L2 for Synthetic-NSVF dataset, Fig. L3 for BlendedMVS dataset, and Fig. L5 for DeepVoxels dataset. In Fig. L1, we compare our results with the results provided by JaxNeRF and PlenOctrees on Synthetic-NeRF dataset. Both JaxNeRF and PlenOctrees are stronger baselines than the original NeRF . In Fig. L4, we also compare our results with PlenOctrees on the Tanks&Temples dataset. In the qualitative comparisons, there is no consistent superior method across all the scenes, which matches the results in quantitative comparisons.
M Derivation of low-density initialization
In this section, we derive the bias term in the low-density initialization to make the activated alpha value very close to zero (i.e., ) at the beginning of our coarse stage optimization. More specifically, is a hyperparameter, and means the decay factor of the accumulated transmittance for a ray tracing forward a distance of the voxel size .
Recall that, in practice, we initialize all raw values in to and thus the initial activated densities are all equal to . Then, the decay factor of the accumulated transmittance can be written as
We want to find the such that the decay factor is for passing through a distance of the voxel size . Therefore, we have
which is the equation presented in our main paper to impose the prior of low-density initialization.
N Derivation of post-activation
By expanding the post-activated trilinear interpolation, we have
We want to prove that such a post-activation can produce sharp linear surface (decision boundary) in a single grid cell. Let first prove in the simplest 1D grid and then we generalize it to 2D grid. The case in 3D grid then can be proven easily using the derivations in 1D and 2D grid.
Consider a target function in the form of a shifted unit step function,
where is the position of the target linear surface (decision boundary) in the 1D grid cell. We only derive for the occupancy at right-hand side as the opposite direction can be trivially generalized. Visualization for and with some specific parameters is shown in Fig. N1.
We show that can be made arbitrarily close to . Specifically, given any and satisfying and , we can find the grid values such that
As the function is a monotonically increasing function bounded in $$, the criterion can then be simplified into
By expanding the inequality (14a), we get
By applying the same process to the inequality (14b), we get another inequality:
There are infinite numbers of solutions satisfying the inequalities (15) and (16). Here, we introduce one more constraint to make the derivation and later extensions simpler:
From Eq. 12 with this constraint we can then derive the linear relation between :
Substitute in the inequality (15), we get
Similarly, substitute in the inequality (16), we get
In summary, by adding the extra constraint (17), the inequality (15) and (16) become
which defines the upper bound of such that the conditions on the function in (14a), (14b), and (17) can all be satisfied. The tighter upper bound is
where the is the pre-defined step size in volume rendering and we skip the detailed derivation of Eq. 20 here.
As a verification, we re-use the examples in Fig. 1(b) as the target functions and set the tolerance to and the volume rendering related value to . We can directly find the grid values using the derivations given above. We show the derived numbers and the resulting plot visualization in Fig. N2. It can be seen that the derived can faithfully resemble the target .
N.2 Derivation for a 2D grid cell
We illustrate two situations that a linear boundary crossing a 2D grid cell in Fig. N3, where are the top-left, top-right, bottom-left, and bottom-right values of a 2D grid cell and is the linear boundary parameterized by the vertical position . Let the coordinates of the top-left, top-right, bottom-left, and bottom-right corners be , , , and , so the top and bottom edges of the grid cell are on the horizontal lines and , respectively. Without loss of generality, we assume the linear boundary always intersects the top edge of the 2D grid cell (i.e., ) and the left-hand-side of the boundary is free space. We can generalize to other cases via rotation and flipping, or by negating the grid values. We consider the 2D target function with respect to the decision boundary inside the grid cell as a 2D unit step function, and parameterize it by the vertical coordinate so that on each horizontal scan line the target function can be expressed as
The goal here is to show that, by choosing suitable values for , we are able to approximate the target function closely enough within the required tolerance using our post-activation scheme.
An example is illustrated in Fig. 3(a), where the linear boundary intersects the top edge and the bottom edge of a grid cell. The linear boundary is a linear function of defined by . Based on the results we have derived for 1D, to prove this 2D case we only need to show that Eq. 18 and Eq. 20 are linear function of , so that once we use the two equations to determine for approximating and for approximating , we can readily recover from and on all horizontal scan lines . That is, the approximation criteria in Eq. 14a and Eq. 14b, where the target surface position is , are automatically satisfied by
Eq. 20 is trivially a linear function of given . By substituting in Eq. 18 with from Eq. 20, we get
when . Thus, is also a linear function of given .
Case II: 0<c(0)<10𝑐010<c(0)<1 and 1<c(1)1𝑐11<c(1).
We assume in the target function in Eq. 13, but it finally turns out that the derivations in Sec. N.1 can trivially generalize to . Actually, the derivations also work for as long as we modify the signs in inequality (19) to ‘’, and the upper bound in Eq. 20 becomes lower bound. We verify the results with some examples shown in Fig. N4.
In summary, the derivation in Sec. N.1 also works for extrapolation with minor modifications and the derivation in Case I is directly applicable for Case II.
As a verification, we construct two examples in Fig. N6. The volume rendering step size is set to , the error bound is set to , and we show two results with tolerance and . For each target boundary, we first evaluate and . The values of can then be derived from Eq. 20 and the values of can be derived from Eq. 18.
N.3 Derivation for a 3D grid cell
Similar to the proof for 2D in Sec. N.2, we can assume that the 3D linear surface intersects the top face of a 3D grid cell, as illustrated in Fig. N7. Without loss of generality, we assume the surface intersects the horizontal planes and , and the left-hand-side of the surface,
is free space. The linear surface is . We can use the results in Sec. N.2 to determine the grid values of the top four corners with and the values of the bottom four corners with . The linear boundary at every horizontal slice can be automatically satisfied, as we have shown in Sec. N.2.
N.4 Future extensions
We only prove that the alpha field by post-activated interpolation can be arbitrarily close to a linear surface. We show in Fig. N8 that we can further tune the tolerances at the top edge and bottom edge of a grid to produce sharp and non-linear surfaces. Extending the proof or the capability of post-activation in future work could be helpful to geometry modeling.
Closed form solution when 3D available.
In this work, we only consider the same input setup as NeRF, where only 2D observations and camera poses are available. In cases that the 3D model of the scene is available, an algorithm to convert the 3D model to our post-activated density voxel grid can be helpful. Our representation is directly compatible with gradient-based optimization and volume rendering to support follow-up applications, while 3D in other formats like mesh or point cloud may need more effort. We believe our derivations are useful for future work to develop a closed-form solution to convert the 3D in other formats into our representation.