Volume Rendering of Neural Implicit Surfaces

Lior Yariv, Jiatao Gu, Yoni Kasten, Yaron Lipman

Introduction

Volume rendering is a set of techniques that renders volume density in radiance fields by the so called volume rendering integral. It has recently been shown that representing both the density and radiance fields as neural networks can lead to excellent prediction of novel views by learning only from a sparse set of input images. This neural volume rendering approach, presented in and developed by its follow-ups approximates the integral as alpha-composition in a differentiable way, allowing to learn simultaneously both from input images. Although this coupling indeed leads to good generalization of novel viewing directions, the density part is not as successful in faithfully predicting the scene’s actual geometry, often producing noisy, low fidelity geometry approximation.

We propose VolSDF to devise a different model for the density in neural volume rendering, leading to better approximation of the scene’s geometry while maintaining the quality of view synthesis. The key idea is to represent the density as a function of the signed distance to the scene’s surface, see Figure 1. Such density function enjoys several benefits. First, it guarantees the existence of a well-defined surface that generates the density. This provides a useful inductive bias for disentangling density and radiance fields, which in turn provides a more accurate geometry approximation. Second, we show this density formulation allows bounding the approximation error of the opacity along rays. This bound is used to sample the viewing ray so to provide a faithful coupling of density and radiance field in the volume rendering integral. E.g., without such a bound the computed radiance along a ray (pixel color) can potentially miss or extend surface parts leading to incorrect radiance approximation.

A closely related line of research, often referred to as neural implicit surfaces , have been focusing on representing the scene’s geometry implicitly using a neural network, making the surface rendering process differentiable. The main drawback of these methods is their requirement of masks that separate objects from the background. Also, learning to render surfaces directly tends to grow extraneous parts due to optimization problems, which are avoided by volume rendering. In a sense, our work combines the best of both worlds: volume rendering with neural implicit surfaces.

We demonstrate the efficacy of VolSDF by reconstructing surfaces from the DTU and Blended-MVS datasets. VolSDF produces more accurate surface reconstructions compared to NeRF and NeRF++ , and comparable reconstruction compared to IDR , while avoiding the use of object masks. Furthermore, we show disentanglement results with our method, i.e., switching the density and radiance fields of different scenes, which is shown to fail in NeRF-based models.

Related work

Neural Scene Representation & Rendering Implicit functions are traditionally adopted in modeling 3D scenes . Recent studies have been focusing on model implicit functions with multi-layer perceptron (MLP) due to its expressive representation power and low memory foot-print, including scene (geometry & appearance) representation and free-view rendering . In particular, NeRF has opened up a line of research (see for an overview) combining neural implicit functions together with volume rendering to achieve photo-realistic rendering results. However, it is non-trivial to find a proper threshold to extract surfaces from the predicted density, and the recovered geometry is far from satisfactory. Furthermore, sampling of points along a ray for rendering a pixel is done using an opacity function that is approximated from another network without any guarantee for correct approximation.

Multi-view 3D Reconstruction Image-based 3D surface reconstruction (multi-view stereo) has been a longstanding problem in the past decades. Classical multi-view stereo approaches are generally either depth-based or voxel-based . For instance, in COLMAP (a typical depth-based method) image features are extracted and matched across different views to estimate depth. Then the predicted depth maps are fused to obtain dense point clouds. To obtain the surface, an additional meshing step e.g. Poisson surface reconstruction is applied. However, these methods with complex pipelines may accumulate errors at each stage and usually result in incomplete 3D models, especially for non-Lambertian surfaces as they can not handle view dependent colors. On the contrary, although it produces complete models by directly modeling objects in a volume, voxel-based approaches are limited to low resolution due to high memory consumption. Recently, neural-based approaches such as DVR , IDR , NLR have also been proposed to reconstruct scene geometry from multi-view images. However, these methods require accurate object masks and appropriate weight initialization due to the difficulty of propagating gradients.

Independently from and concurrently with our work here, also use implicit surface representation incorporated into volume rendering. In particular, they replace the local transparency function with an occupancy network . This allows adding surface smoothing term to the loss, improving the quality of the resulting surfaces. Differently from their approach, we use signed distance representation, regularized with an Eikonal loss without any explicit smoothing term. Furthermore, we show that the choice of using signed distance allows bounding the opacity approximation error, facilitating the approximation of the volume rendering integral for the suggested family of densities.

Method

In this section we introduce a novel parameterization for volume density, defined as transformed signed distance function. Then we show how this definition facilitates the volume rendering process. In particular, we derive a bound of the error in the opacity approximation and consequently devise a sampling procedure for approximating the volume rendering integral.

where α,β>0\alpha,\beta>0 are learnable parameters, and Ψβ\Psi_{\beta} is the Cumulative Distribution Function (CDF) of the Laplace distribution with zero mean and β\beta scale (i.e., mean absolute deviation, which is intuitively the L1L_{1} version of the standard deviation),

Figure 1 (center left and right) depicts an example of such a density and SDF. As can be readily checked from this definition, as β\beta approach zero, the density σ\sigma converges to a scaled indicator function of Ω\Omega, that is σ→α1Ω\sigma\rightarrow\alpha\mathbf{1}_{\Omega} for all points x∈Ω∖M{\bm{x}}\in\Omega\setminus{\mathcal{M}}.

Intuitively, the density σ\sigma models a homogeneous object with a constant density α\alpha that smoothly decreases near the object’s boundary, where the smoothing amount is controlled by β\beta. The benefit in defining the density as in equation 2 is two-fold: First, it provides a useful inductive bias for the surface geometry M{\mathcal{M}}, and provides a principled way to reconstruct the surface, i.e., as the zero level-set of dΩd_{\Omega}. This is in contrast to previous work where the reconstruction was chosen as an arbitrary level set of the learned density. Second, the particular form of the density as defined in equation 2 facilitates a bound on the error of the opacity (or, equivalently the transparency) of the rendered volume, a crucial component in the volumetric rendering pipeline. In contrast, such a bound will be hard to devise for a generic MLP densities.

2 Volume rendering of σ𝜎\sigma

In this section we review the volume rendering integral and the numerical integration commonly used to approximate it, requiring a set S{\mathcal{S}} of sample points per ray. In the following section (Section 3.3), we explore the properties of the density σ\sigma and derive a bound on the opacity approximation error along viewing rays. Finally, in Section 3.4 we derive an algorithm for producing a sample S{\mathcal{S}} to be used in the volume rendering numerical integration.

The transparency function of the volume along a ray x{\bm{x}}, denoted TT, indicates, for each t≥0t\geq 0, the probability a light particle succeeds traversing the segment [c,x(t)][{\bm{c}},{\bm{x}}(t)] without bouncing off,

and the opacity OO is the complement probability,

Note that OO is a monotonic increasing function where O(0)=0O(0)=0, and assuming that every ray is eventually occluded O(∞)=1O(\infty)=1. In that sense we can think of OO as a CDF, and

is its Probability Density Function (PDF). The volume rendering equation is the expected light along the ray,

where L(x,n,v)L({\bm{x}},{\bm{n}},{\bm{v}}) is the radiance field, namely the amount of light emanating from point x{\bm{x}} in direction v{\bm{v}}; in our formulation we also allow LL to depend on the level-set’s normal, i.e., n(t)=∇xdΩ(x(t)){\bm{n}}(t)=\nabla_{\bm{x}}d_{\Omega}({\bm{x}}(t)). Adding this dependency is motivated by the fact that BRDFs of common materials are often encoded with respect to the surface normal, facilitating disentanglement as done in surface rendering . We will get back to disentanglement in the experiments section. The integral in equation 7 is approximated using a numerical quadrature, namely the rectangle rule, at some discrete samples S={si}i=1m{\mathcal{S}}=\left\{s_{i}\right\}_{i=1}^{m}, 0=s1<s2<…<sm=M0=s_{1}<s_{2}<\ldots<s_{m}=M, where MM is some large constant:

Sampling. Since the PDF τ\tau is typically extremely concentrated near the object’s boundary (see e.g., Figure 3, right) the choice of the sample points S{\mathcal{S}} has a crucial effect on the approximation quality of equation 8. One solution is to use an adaptive sample, e.g., S{\mathcal{S}} computed with the inverse CDF, i.e., O−1O^{-1}. However, OO depends on the density model σ\sigma and is not given explicitly. In a second, coarse network was trained specifically for the approximation of the opacity OO, and was used for inverse sampling. However, the second network’s density does not necessarily faithfully represents the first network’s density, for which we wish to compute the volume integral. Furthermore, as we show later, one level of sampling could be insufficient to produce an accurate sample S{\mathcal{S}}. Using a naive or crude approximation of OO would lead to a sub-optimal sample set S{\mathcal{S}} that misses, or over extends non-negligible τ\tau values. Consequently, incorrect radiance approximations can occur (i.e., pixel color), potentially harming the learned density-radiance field decomposition. Our solution works with a single density σ\sigma, and the sampling S{\mathcal{S}} is computed by a sampling algorithm based on an error bound for the opacity approximation. Figure 2 compares the NeRF and VolSDF renderings for the same scene. Note the salt and pepper artifacts in the NeRF rendering caused by the random samples; using fixed (uniformly spaced) sampling in NeRF leads to a different type of artifacts shown in the supplementary.

3 Bound on the opacity approximation error

In this section we develop a bound on the opacity approximation error using the rectangle rule. For a set of samples T={ti}i=1n{\mathcal{T}}=\left\{t_{i}\right\}_{i=1}^{n}, 0=t1<t2<⋯<tn=M0=t_{1}<t_{2}<\cdots<t_{n}=M, we let δi=ti+1−ti\delta_{i}=t_{i+1}-t_{i}, and σi=σ(x(ti))\sigma_{i}=\sigma({\bm{x}}(t_{i})). Given some t∈(0,M]t\in(0,M], assume t∈[tk,tk+1]t\in[t_{k},t_{k+1}], and apply the rectangle rule (i.e., left Riemann sum) to get the approximation:

is the rectangle rule approximation, and E(t)E(t) denotes the error in this approximation. The corresponding approximation of the opacity function (equation 5) is

Our goal in this section is to derive a uniform bound over [0,M][0,M] to the approximation O^≈O\widehat{O}\approx O. The key is the following bound on the derivativeAs dΩd_{\Omega} is not differentiable everywhere the bound is on the Lipschitz constant of σ\sigma, see supplementary. of the density σ\sigma inside an interval along the ray x(t){\bm{x}}(t):

The derivative of the density σ\sigma within a segment [ti,ti+1][t_{i},t_{i+1}] satisfies

and Bi={x ∣ ∥x−x(ti)∥<∣di∣}B_{i}=\left\{{\bm{x}}\ |\ \left\|{\bm{x}}-{\bm{x}}(t_{i})\right\|<|d_{i}|\right\}, di=dΩ(x(ti))d_{i}=d_{\Omega}({\bm{x}}(t_{i})).

The proof of this theorem, which is provided in the supplementary, makes a principled use of the signed distance function’s unique properties; the explicit formula for di∗d^{*}_{i} is a bit cumbersome and therefore is deferred to the supplementary as-well. The inset depicts the boundary of the open balls union Bi∪Bi+1B_{i}\cup B_{i+1}, the interval [x(ti),x(ti+1)][{\bm{x}}(t_{i}),{\bm{x}}(t_{i+1})] and the bound is defined in terms of the minimal distance between these two sets, i.e., di∗d^{*}_{i}.

The benefit in Theorem 1 is that it allows to bound the density’s derivative in each interval [ti,ti−1][t_{i},t_{i-1}] based only on the unsigned distance at the interval’s end points, ∣di∣,∣di+1∣|d_{i}|,|d_{i+1}|, and the density parameters α,β\alpha,\beta. This bound can be used to derive an error bound for the rectangle rule’s approximation of the opacity,

Details are in the supplementary. Equation 12 leads to the following opacity error bound, also proved in the supplementary:

For t∈[0,M]t\in[0,M], the error of the approximated opacity O^\hat{O} can be bounded as follows:

Finally, we can bound the opacity error for t∈[tk,tk+1]t\in[t_{k},t_{k+1}] by noting that E^(t)\widehat{E}(t), and consequently also exp⁡(E^(t))\exp(\widehat{E}(t)) are monotonically increasing in tt, while exp⁡(−R^(t))\exp(-\widehat{R}(t)) is monotonically decreasing in tt, and therefore

Taking the maximum over all intervals furnishes a bound BT,βB_{{\mathcal{T}},\beta} as a function of T{\mathcal{T}} and β\beta,

To conclude this section we derive two useful properties, proved in the supplementary. The first, is that sufficiently dense sampling is guaranteed to reduce the error bound BT,ϵB_{{\mathcal{T}},{\epsilon}}:

Fix β>0\beta>0. For any ϵ>0{\epsilon}>0 a sufficient dense sampling T{\mathcal{T}} will provide BT,β<ϵB_{{\mathcal{T}},\beta}<{\epsilon}.

Second, with a fixed number of samples we can set β\beta such that the error bound is below ϵ{\epsilon}:

Fix n>0n>0. For any ϵ>0{\epsilon}>0 a sufficiently large β\beta that satisfies

will provide BT,β≤ϵB_{{\mathcal{T}},\beta}\leq{\epsilon}.

4 Sampling algorithm

In this section we develop an algorithm for computing the sampling S{\mathcal{S}} to be used in equation 8. This is done by first utilizing the bound in equation 15 to find samples T{\mathcal{T}} so that O^\widehat{O} (via equation 10) provides an ϵ\epsilon approximation to the true opacity OO, where ϵ\epsilon is a hyper-parameter, that is BT,β<ϵB_{{\mathcal{T}},\beta}<{\epsilon}. Second, we perform inverse CDF sampling with O^\hat{O}, as described in Section 3.2.

Note that from Lemma 1 it follows that we can simply choose large enough nn to ensure BT,β<ϵB_{{\mathcal{T}},\beta}<{\epsilon}. However, this would lead to prohibitively large number of samples. Instead, we suggest a simple algorithm to reduce the number of required samples in practice and allows working with a limited budget of sample points. In a nutshell, we start with a uniform sampling T=T0{\mathcal{T}}={\mathcal{T}}_{0}, and use Lemma 2 to initially set a β+>β\beta_{+}>\beta that satisfies BT,β+≤ϵB_{{\mathcal{T}},\beta_{+}}\leq{\epsilon}. Then, we repeatedly upsample T{\mathcal{T}} to reduce β+\beta_{+} while maintaining BT,β+≤ϵB_{{\mathcal{T}},\beta_{+}}\leq{\epsilon}. Even though this simple strategy is not guaranteed to converge, we find that β+\beta_{+} usually converges to β\beta (typically 85%85\%, see also Figure 3), and even in cases it does not, the algorithm provides β+\beta_{+} for which the opacity approximation still maintains an ϵ{\epsilon} error. The algorithm is presented below (Algorithm 1).

We initialize T{\mathcal{T}} (Line 1 in Algorithm 1) with uniform sampling T0={ti}i=1n{\mathcal{T}}_{0}=\left\{t_{i}\right\}_{i=1}^{n}, where tk=(k−1)Mn−1t_{k}=(k-1)\frac{M}{n-1}, k∈[n]k\in[n] (we use n=128n=128 in our implementation). Given this sampling we next pick β+>β\beta_{+}>\beta according to Lemma 2 so that the error bound satisfies the required ϵ{\epsilon} bound (Line 2 in Algorithm 1).

In order to reduce β+\beta_{+} while keep BT,β+≤ϵB_{{\mathcal{T}},\beta_{+}}\leq{\epsilon}, nn samples are added to T{\mathcal{T}} (Line 4 in Algorithm 1), where the number of points sampled from each interval is proportional to its current error bound, equation 14.

Assuming T{\mathcal{T}} was sufficiently upsampled and satisfy BT,β+<ϵB_{{\mathcal{T}},\beta_{+}}<{\epsilon}, we decrease β+\beta_{+} towards β\beta. Since the algorithm did not stop we have that BT,β>ϵB_{{\mathcal{T}},\beta}>{\epsilon}. Therefore the Mean Value Theorem implies the existence of β⋆∈(β,β+)\beta_{\star}\in(\beta,\beta_{+}) such that BT,β⋆=ϵB_{{\mathcal{T}},\beta_{\star}}={\epsilon}. We use the bisection method (with maximum of 10 iterations) to efficiently search for β⋆\beta_{\star} and update β+\beta_{+} accordingly (Lines 6 and 7 in Algorithm 1). The algorithm runs iteratively until BT,β≤ϵB_{{\mathcal{T}},\beta}\leq{\epsilon} or a maximal number of 55 iterations is reached. Either way, we use the final T{\mathcal{T}} and β+\beta_{+} (guaranteed to provide BT,β+≤ϵB_{{\mathcal{T}},\beta_{+}}\leq{\epsilon}) to estimate the current opacity O^\widehat{O}, Line 10 in Algorithm 1). Finally we return a fresh set of m=64m=64 samples O^\hat{O} using inverse transform sampling (Line 11 in Algorithm 1). Figure 3 shows qualitative illustration of Algorithm 1, for β=0.001\beta=0.001 and ϵ=0.1{\epsilon}=0.1 (typical values).

5 Training

where LRGB{\mathcal{L}}_{\text{RGB}} is the color loss; ∥⋅∥1\left\|\cdot\right\|_{1} denotes the 11-norm, S{\mathcal{S}} is computed with Algorithm 1, and I^S\hat{I}_{\mathcal{S}} is the numerical approximation to the volume rendering integral in equation 8; here we also incorporate the global feature in the radiance field, i.e., Li=Lψ(x(si),n(si),vp,z(x(si)))L_{i}=L_{\psi}({\bm{x}}(s_{i}),{\bm{n}}(s_{i}),{\bm{v}}_{p},{\bm{z}}({\bm{x}}(s_{i}))). LSDF{\mathcal{L}}_{\text{SDF}} is the Eikonal loss encouraging dd to approximate a signed distance function ; the samples z{\bm{z}} are taken to combine a single random uniform space point and a single point from S{\mathcal{S}} for each pixel pp. We train with batches of size 1024 pixels pp. λ\lambda is a hyper-parameter set to 0.10.1 throughout the the experiments. Further implementation details are provided in the supplementary.

Experiments

We evaluate our method on the challenging task of multiview 3D surface reconstruction. We use two datasets: DTU and BlendedMVS , both containing real objects with different materials that are captured from multiple views. In Section 4.1 we show qualitative and quantitative 3D surface reconstruction results of VolSDF, comparing favorably to relevant baselines. In Section 4.2 we demonstrate that, in contrast to NeRF , our model is able to successfully disentangle the geometry and appearance of the captured objects.

DTU The DTU dataset contains multi-view image (4949 or 6464) of different objects with fixed camera and lighting parameters. We evaluate our method on the 1515 scans that were selected by . We compare our surface accuracy using the Chamfer l1l_{1} loss (measured in mm) to COLMAP0 (which is watertight reconstruction; COLMAP7 is not watertight and provided only for reference) , NeRF and IDR , where for fair comparison with IDR we only evaluate the reconstruction inside the visual hull of the objects (defined by the segmentation masks of ). We further evaluate the PSNR of our rendering compared to . Quantitative results are presented in Table 1. It can be observed that our method is on par with IDR (that uses object masks for all images) and outperforms NeRF and COLMAP in terms of reconstruction accuracy. Our rendering quality is comparable to NeRF’s.

BlendedMVS The BlendedMVS dataset contains a large collection of 113113 scenes captured from multiple views. It supplies high quality ground truth 3D models for evaluation, various camera configurations, and a variety of indoor/outdoor real environments. We selected 99 different scenes and used our method to reconstruct the surface of each object. In contrast to the DTU dataset, BlendedMVS scenes have complex backgrounds. Therefore we use NeRF++ as a baseline for this dataset. In Table 2 we present our results compared to NeRF++. Qualitative comparisons are presented in Fig. 5; since the units are unknown in this case we present relative improvement of

Chamfer distance (in %\%) compared to NeRF. Also in this case, we improve NeRF reconstructions considerably, while being on-par in terms of the rendering quality (PSNR).

Comparison to IDR is the state of the art 3D surface reconstruction method using implicit representation. However, it suffers from two drawbacks: first, it requires object masks for training, which is a strong supervision signal. Second, since it sets the pixel color based only on the single point of intersection of the corresponding viewing ray, it is more pruned to local minima that sometimes appear in the form of extraneous surface parts. Figure 6 compares the same scene trained with IDR with the addition of ground truth masks, and VolSDF trained without masks. Note that IDR introduces some extraneous surface parts (e.g., in marked red), while VolSDF provides a more faithful result in this case.

2 Disentanglement of geometry and appearance

We have tested the disentanglement of scenes to geometry (density) and appearance (radiance field) by switching the radiance fields of two trained scenes. For VolSDF we switched LψL_{\psi}. For NeRF we note that the radiance field is computed as Lψ(z,v)L_{\psi}({\bm{z}},{\bm{v}}), where LψL_{\psi} is a fully connected network with one hidden layer (of width 128 and ReLU activation) and z{\bm{z}} is a feature vector. We tested two versions of NeRF disentanglement: First, by switching the original radiance fields LψL_{\psi} of trained NeRF networks. Second, by switching the radiance fields of trained NeRF models with an identical radiance field model to ours, namely Lψ(x,n,v,z)L_{\psi}({\bm{x}},{\bm{n}},{\bm{v}},{\bm{z}}). As shown in Figure 7 both versions of NeRF fail to produce a correct disentanglement in these scenes, while VolSDF successfully switches the materials of the two objects. We attribute this to the specific inductive bias injected with the use of the density in equation 2.

Conclusions

We introduce VolSDF, a volume rendering framework for implicit neural surfaces. We represent the volume density as a transformed version of the signed distance function to the learned surface geometry. This seemingly simple definition provides a useful inductive bias, allowing disentanglement of geometry (i.e., density) and radiance field, and improves the geometry approximation over previous neural volume rendering techniques. Furthermore, it allows to bound the opacity approximation error leading to high fidelity sampling of the volume rendering integral.

Some limitations of our method present interesting future research opportunities. First, although working well in practice, we do not have a proof of correctness for the sampling algorithm. We believe providing such a proof, or finding a version of this algorithm that has a proof would be a useful contribution. In general, we believe working with bounds in volume rendering could improve learning and disentanglement and push the field forward. Second, representing non-watertight manifolds and/or manifolds with boundaries, such as zero thickness surfaces, is not possible with an SDF. Generalizations such as multiple implicits and unsigned fields could be proven valuable. Third, our current formulation assumes homogeneous density; extending it to more general density models would allow representing a broader class of geometries. Fourth, now that high quality geometries can be learned in an unsupervised manner it will be interesting to learn dynamic geometries and shape spaces directly from collections of images. Lastly, although we don’t see immediate negative societal impact of our work, we do note that accurate geometry reconstruction from images can be used for malice purposes.

Acknowledgments

LY is supported by the European Research Council (ERC Consolidator Grant, "LiftMatch" 771136), the Israel Science Foundation (Grant No. 1830/17), and Carolito Stiftung (WAIC). YK is supported by the U.S.- Israel Binational Science Foundation, grant number 2018680, Carolito Stiftung (WAIC), and by the Kahn foundation.

References

Supplementary Material

Appendix A Additional results

Figure 8 depicts an ablation study that we performed for evaluating the sampling algorithm, by replacing it with other sampling strategies. We compared with the following alternatives: Uniform stands for an uniform sampling of 256256 samples along each ray ; 2-networks denotes a hierarchical sampling using coarse and fine networks as suggested in ; our sampling algorithm as suggested in section 3.4 where the maximal number of iterations is set to 11 or 55 (the choice in the paper). We note that using 11 iteration resembles one level of sampling as in the hierarchical sampling of .

As can be seen in Figure 8 (see also reported Chamfer and PSNR scores) alternative sampling procedures lead to lower accuracy and some artifacts in the geometry (see e.g., head top, and nose areas) and rendering (see e.g., salt and pepper noise and over-smoothed areas).

A.2 Positional encoding ablation

We perform an ablation study on the level of positional encoding used in the geometry network. We note that VolSDF use level 6 positional encoding, while NeRF use level 10. In Figure 9 we show the DTU Bunny scene with positional encoding levels 6 and 10 for both NeRF and VolSDF; we report both PSNR for the rendered images and Chamfer distance to the learned surfaces. Note that higher positional encoding improves specular highlights and details of VolSDF but adds some undesired noise to the reconstructed surface.

A.3 Multi-view 3D reconstruction

Figures 10 and 11 show additional qualitative results for DTU and BlendedMVS datasets, respectively. We further provide a video with qualitative rendering results of the learned geometries and radiance fields from simulated camera paths (including novel views); note that VolSDF rendering alleviates NeRF’s salt and pepper artifacts, while producing higher fidelity geometry approximation.

A.4 Limitations

Figure 12 shows the main failure cases of our method: First, we observe that the geometry for unseen regions is not well defined and can be completed arbitrarily by the algorithm, see e.g., the angle statue head top in (a), and the reverse side of the snow man statue in (b). Second, homogeneous texture-less areas are hard to reconstruct faithfully, see e.g., the white background desk in (c). These limitations can potentially be alleviated with the addition of extra assumptions such as minimal surface reconstruction and/or defining background color or predefined geometry (e.g., a plane).

A.5 Rendering comparison

Figure 2 in the main paper shows a comparison of NeRF and VolSDF rendering for the same scene using the same random sampling strategy of the inverse opacity function. For NeRF, replacing random sampling with regularly-spaced sampling introduces different artifacts as shown in Figure 13. In contrast, VolSDF produces consistent rendering results regardless of the sampling strategy.

A.6 Normal dependency in radiance field

As presented in , incorporating the zero level set normal in surface rendering improves both reconstruction quality and disentanglement of geometry and radiance. Incorporating the level set’s normal has similar effect in the radiance field representation of VolSDF; see Figure 14 where we compare using radiance field with and without normal dependency.

A.7 Disentanglement of geometry and appearance

Figure 15 shows additional results for unsupervised disentanglement of geometry and radiance field (switching the radiance fields of three independently trained scenes). Note how the material and light from one scene is gracefully transferred to the two other scenes. We further provide in the supplementary video the 360 degrees camera path for one of the swaps.

Appendix B Experiments setup

We used the known camera poses to shift the coordinate system, locating the object at the origin. This is done using a least squares solution for the intersection point of all camera principal axes. Let RmaxR_{\textit{max}} be the maximal camera center norm, we further apply a global scale of 3Rmax∗1.1\frac{3}{R_{\textit{max}}*1.1} to place all camera centers inside a sphere of radius 33, centered at the origin.

B.2 Datasets

We used the formal evaluation script to measure the Chamfer l1l_{1} distance between each reconstructed object to its corresponding ground truth point cloud. For fare comparison with , we used their masks (with a dilation of 50 pixels) to remove non visual hull parts from each output of both ours and . Specifically, from an output mesh, we remove a 3D point (and its adjacent triangles) if it is projected to a zero pixel label in any image. Finally, we used the largest connected component mesh for evaluation. The results for COLMAP and IDR are taken from .

BlendedMVS

We used the ground truth meshes supplied by the authors to evaluate the Chamfer l1l_{1} distances from the output surfaces. For each mesh we evaluated the largest connected surface component above the ground plane. To measure the Chamfer l1l_{1} distance we used 100K100K random point samples from each surface.

B.3 Additional implementation details

To capture the high frequencies of the geometry and radiance field, we exploit Positional Encoding (PE) for the position x{\bm{x}} and view direction v{\bm{v}} in the geometry and radiance field. For the position x{\bm{x}} we use 66 PE levels, and for the the view direction v{\bm{v}} we use 44 PE levels, as in . The influence of different levels of frequencies applied to the position is presented in the Section A.2.

Training details

Mesh extraction

We use the Marching Cubes algorithm for extracting each surface mesh from the zero level set of the signed distance function defined by d(x)d({\bm{x}}).

Modeling the background

To satisfy the assumption that all rays, including rays that do not intersect any surface, are eventually occluded (i.e., O(∞)=1O(\infty)=1), we model our SDF as:

where rr denotes a predefined scene bounding sphere (in out experiments r=3r=3, see Section B.1 for more details about camera normalization). During the rendering process, the maximal depth for ray samples, denoted by MM in Sec. 3.2, is set to 2r2r. Intuitively, this modification can be considered as modeling the background using a 360 degrees panorama on the scene boundaries. For modeling more complex backgrounds as in the blendedMVS dataset, we follow the parametrization presented in NeRF++ : the volume outside a radius 3 sphere (the background) is modeled using an additional NeRF network, predicting the background point density and radiance field. Each 3D3D background point (xb,yb,zb)(x_{b},y_{b},z_{b}) is modeled using a 4D4D representation (x′,y′,z′,1r)(x^{\prime},y^{\prime},z^{\prime},\frac{1}{r}) where ∥(x′,y′,z′)∥=1\left\|(x^{\prime},y^{\prime},z^{\prime})\right\|=1 and (xb,yb,zb)=r⋅(x′,y′,z′)(x_{b},y_{b},z_{b})=r\cdot(x^{\prime},y^{\prime},z^{\prime}). For rendering the background component of a pixel, we use 3232 points calculated by sampling 1r\frac{1}{r} uniformly in the range (0,13](0,\frac{1}{3}]. More details regarding this parametrization (referred as the "inverted sphere parametrization") can be found in .

B.4 Baselines

For running NeRF we used a slightly modified version of the pytorch implementation suggested by . Following the official implementation codehttps://github.com/bmild/nerf of NeRF we extracted the 5050 level set of the learned density as NeRF geometry reconstruction.

NeRF++

We used the official codehttps://github.com/Kai-46/nerfplusplus of NeRF++ . All the cameras are normalized inside a unit sphere, and we train two networks (one foreground network in normal coordinates, one background network in inverse sphere coordinates). For foreground, we used 64 uniform samples and 128 hierarchical samples; for background, we used 32 uniform and 64 hierarchical samples. Similar to NeRF experiments, we extracted the 50 level set of the foreground density as reconstruction.

Appendix C Proofs and additional lemmas

As dΩd_{\Omega} is not everywhere differentiable we need to bound the Lipschitz constant of σ\sigma, rather than its derivative. The Lipshcitz constant is defined as a constant Ki>0K_{i}>0 so that ∣σ(x(s))−σ(x(t))∣≤Ki∣s−t∣\left|\sigma({\bm{x}}(s))-\sigma({\bm{x}}(t))\right|\leq K_{i}|s-t|, for all s,t∈[ti,ti+1]s,t\in[t_{i},t_{i+1}].

The Lipshcitz constant of the density σ\sigma within a segment [ti,ti+1][t_{i},t_{i+1}] satisfies

and Bi={x ∣ ∥x−x(ti)∥<∣di∣}B_{i}=\left\{{\bm{x}}\ |\ \left\|{\bm{x}}-{\bm{x}}(t_{i})\right\|<|d_{i}|\right\}, di=dΩ(x(ti))d_{i}=d_{\Omega}({\bm{x}}(t_{i})).

The constant di∗d^{*}_{i} can be computed explicitly using di,di+1,ti,ti+1d_{i},d_{i+1},t_{i},t_{i+1} as described in the next proposition.

The lower distance bound di∗d^{*}_{i} can be computed using the following formulas:

di=dΩ(x(ti))d_{i}=d_{\Omega}({\bm{x}}(t_{i})), hi=2δis(s−δi)(s−∣di∣)(s−∣di+1∣)h_{i}=\frac{2}{\delta_{i}}\sqrt{s(s-\delta_{i})(s-|d_{i}|)(s-|d_{i+1}|)}, s=12(δi+∣di∣+∣di+1∣)s=\frac{1}{2}(\delta_{i}+|d_{i}|+|d_{i+1}|), δi=ti+1−ti\delta_{i}=t_{i+1}-t_{i}.

We denote by Φβ\Phi_{\beta} the Probability Density Function (PDF) of the Laplace distribution (the CDF of which is given in equation 3),

where the first equality is by definition of σ\sigma (equation 2), the first inequality uses the fact that the Lipschitz constant of a continuously differential function is the maximum of (absolute value of) its derivative. The second inequality uses the following Lemma A1 and the fact that Φβ\Phi_{\beta} is symmetric w.r.t. zero, positive, and monotonically decreasing at [0,∞)[0,\infty). The last inequality can be justified by noting that by definition of the distance function M⊂(Bi∪Bi+1)c{\mathcal{M}}\subset(B_{i}\cup B_{i+1})^{c} and therefore min⁡t∈[ti,ti+1]∣dΩ(x(t))∣≥di∗\min_{t\in[t_{i},t_{i+1}]}\left|d_{\Omega}({\bm{x}}(t))\right|\geq d^{*}_{i}. ∎

We wish to find the distance of the two sets: A=[x(ti),x(ti+1)]A=[{\bm{x}}(t_{i}),{\bm{x}}(t_{i+1})] and B=∂((Bi∪Bi+1)c)B=\partial((B_{i}\cup B_{i+1})^{c}). That is di∗=min⁡a∈A,b∈B∥a−b∥d^{*}_{i}=\min_{{\bm{a}}\in A,{\bm{b}}\in B}\left\|{\bm{a}}-{\bm{b}}\right\|. Since these two sets are compact, their cartesian product is compact, and the minimum is achieved. Denote this minimum (a∗,b∗)({\bm{a}}^{*},{\bm{b}}^{*}), where a∗∈A{\bm{a}}^{*}\in A and b∗∈B{\bm{b}}^{*}\in B.

First, if Bi∩Bi+1=∅B_{i}\cap B_{i+1}=\emptyset then di∗=0d^{*}_{i}=0. This case is characterized by ∣di∣+∣di+1∣≤δi|d_{i}|+|d_{i+1}|\leq\delta_{i}. Henceforth we assume ∣di∣+∣di+1∣>δi|d_{i}|+|d_{i+1}|>\delta_{i}.

Now we consider several cases. If a∗=x(ti){\bm{a}}^{*}={\bm{x}}(t_{i}) the minimal distance is ∣di∣|d_{i}|, and similarly the distance is ∣di+1∣|d_{i+1}| if a∗=x(ti+1){\bm{a}}^{*}={\bm{x}}(t_{i+1}). Otherwise, if a∗∈(x(ti),x(ti+1)){\bm{a}}^{*}\in({\bm{x}}(t_{i}),{\bm{x}}(t_{i+1})) Lagrange Multipliers imply that a∗−b∗⊥v{\bm{a}}^{*}-{\bm{b}}^{*}\perp{\bm{v}}, namely is orthogonal to the ray. In this case we have two options to consider: First, b∗∈Bˉi∩Bˉi+1{\bm{b}}^{*}\in\bar{B}_{i}\cap\bar{B}_{i+1}, where Bˉi\bar{B}_{i} is the closure of the ball BiB_{i}. In this case Heron’s formula provides the minimal distance hih_{i}, as the height of the triangle Δ(x(ti),b∗,x(ti+1))\Delta({\bm{x}}(t_{i}),{\bm{b}}^{*},{\bm{x}}(t_{i+1})).

Otherwise b∗∈Bˉi{\bm{b}}^{*}\in\bar{B}_{i} or b∗∈Bˉi+1{\bm{b}}^{*}\in\bar{B}_{i+1}. Lets assume the former (the latter is treated similarly). Lagrange multipliers imply that b∗−a∗∥b∗−x(ti){\bm{b}}^{*}-{\bm{a}}^{*}\|{\bm{b}}^{*}-{\bm{x}}(t_{i}). This necessarily means that a∗=x(ti){\bm{a}}^{*}={\bm{x}}(t_{i}), leading to a contradiction with the assumption that a∗∈(x(ti),x(ti+1)){\bm{a}}^{*}\in({\bm{x}}(t_{i}),{\bm{x}}(t_{i+1})).

To check if a∗∈(x(ti),x(ti+1)){\bm{a}}^{*}\in({\bm{x}}(t_{i}),{\bm{x}}(t_{i+1})) and b∗∈Bˉi∩Bˉi+1{\bm{b}}^{*}\in\bar{B}_{i}\cap\bar{B}_{i+1} it is enough to check that the angles at x(ti){\bm{x}}(t_{i}) and x(ti+1){\bm{x}}(t_{i+1}) in the triangle Δ(x(ti),b∗,x(ti+1))\Delta({\bm{x}}(t_{i}),{\bm{b}}^{*},{\bm{x}}(t_{i+1})) are acute, or equivalently that ∣∣di∣2−∣di+1∣2∣<δi2\left||d_{i}|^{2}-|d_{i+1}|^{2}\right|<\delta^{2}_{i} (the latter can be shown with the Law of Cosines).

The distance function d=dΩd=d_{\Omega} to a compact surface M{\mathcal{M}} is Lipschitz with constant 11, i.e., ∣d(x)−d(y)∣≤∥x−y∥\left|d({\bm{x}})-d({\bm{y}})\right|\leq\left\|{\bm{x}}-{\bm{y}}\right\|, for all x,y{\bm{x}},{\bm{y}}.

Let x∗∈M{\bm{x}}^{*}\in{\mathcal{M}} be the closest point to x{\bm{x}}, and y∗∈M{\bm{y}}^{*}\in{\mathcal{M}} the closets point to y{\bm{y}}. If d(x)d(y)<0d({\bm{x}})d({\bm{y}})<0, then ∣d(x)−d(y)∣=∥x−x∗∥+∥y−y∗∥≤∥y−z∥+∥z−x∥\left|d({\bm{x}})-d({\bm{y}})\right|=\left\|{\bm{x}}-{\bm{x}}^{*}\right\|+\left\|{\bm{y}}-{\bm{y}}^{*}\right\|\leq\left\|{\bm{y}}-{\bm{z}}\right\|+\left\|{\bm{z}}-{\bm{x}}\right\|, for all z∈M{\bm{z}}\in{\mathcal{M}}. However, from mean value theorem we know there exists z∈M{\bm{z}}\in{\mathcal{M}} on the straight line [x,y][{\bm{x}},{\bm{y}}], and therefore ∣d(x)−d(y)∣≤∥x−y∥\left|d({\bm{x}})-d({\bm{y}})\right|\leq\left\|{\bm{x}}-{\bm{y}}\right\|. If d(x)d(y)≥0d({\bm{x}})d({\bm{y}})\geq 0, then ∣d(y)−d(x)∣=∣∣d(y)∣−∣d(x)∣∣\left|d({\bm{y}})-d({\bm{x}})\right|=\left||d({\bm{y}})|-|d({\bm{x}})|\right|. Now,

Therefore ∣d(y)∣−∣d(x)∣≤∥x−y∥|d({\bm{y}})|-|d({\bm{x}})|\leq\left\|{\bm{x}}-{\bm{y}}\right\|. The other direction follows by switching the roles of x,y{\bm{x}},{\bm{y}}. We proved ∣d(x)−d(y)∣≤∥x−y∥\left|d({\bm{x}})-d({\bm{y}})\right|\leq\left\|{\bm{x}}-{\bm{y}}\right\|. ∎

Plugging the Lipschitz bound from Theorem 1 (equation 21) provides that the error in equation 9 for t∈[tk,tk+1]t\in[t_{k},t_{k+1}] can be bounded by

For t∈[0,M]t\in[0,M], the error of the approximated opacity O^\hat{O} can be bounded as follows:

where the last inequality is due to the inequality ∣1−exp⁡(r)∣≤exp⁡(∣r∣)−1\left|1-\exp(r)\right|\leq\exp(|r|)-1 and the bound ∣E(t)∣≤E^(t)|E(t)|\leq\widehat{E}(t) from equation 12. ∎

Fix β>0\beta>0. For any ϵ>0{\epsilon}>0 a sufficient dense sampling T{\mathcal{T}} will provide BT,β<ϵB_{{\mathcal{T}},\beta}<{\epsilon}.

Noting that exp⁡(−R^(t))≤1\exp(-\widehat{R}(t))\leq 1, equation 15 leads to

Lastly we note that the inner sum satisfies ∑iδi2≤Mmax⁡iδi\sum_{i}\delta_{i}^{2}\leq M\max_{i}\delta_{i}, and in turn this implies that dense sampling can achieve arbitrary low error bound. ∎

Fix n>0n>0. For any ϵ>0{\epsilon}>0 a sufficiently large β\beta that satisfies

will provide BT,β≤ϵB_{{\mathcal{T}},\beta}\leq{\epsilon}.

Assuming sufficiently large β\beta that satisfies equation 25, then following the derivation of Lemma 1 and assuming equidistance sampling, i.e. δi=Mn−1\delta_{i}=\frac{M}{n-1}, leads to

Appendix D Numerical integration

Standard numerical quadrature in volume rendering is based on the rectangle rule , which we repeat here for completeness. Quantities with  ^\hat{\ } denote approximated quantities. For a set of discrete samples S={si}i=1m{\mathcal{S}}=\left\{s_{i}\right\}_{i=1}^{m}, 0=s1<s2<…<sm=M0=s_{1}<s_{2}<\ldots<s_{m}=M we let δi=si+1−si\delta_{i}=s_{i+1}-s_{i}, σi=σ(x(si))\sigma_{i}=\sigma({\bm{x}}(s_{i})), and Li=L(x(si),n(si),v)L_{i}=L({\bm{x}}(s_{i}),{\bm{n}}(s_{i}),{\bm{v}}). Applying the rectangle rule (left Riemann sum) to approximate the integral in equation 7 we have

where we assumed that MM is sufficiently large so that the integral over the segement [M,∞)[M,\infty) is negligible, and the integral over [0,M][0,M] is approximated with the rectangle rule, and τ(si)=σiT(si)\tau(s_{i})=\sigma_{i}T(s_{i}) as given in equation 6. The rectangle rule is applied yet again to approximate the transparency TT:

To provide a discrete probability distribution τ^\hat{\tau} that parallels the continuous one τ\tau, defines pi=exp⁡(−σiδi)p_{i}=\exp(-\sigma_{i}\delta_{i}) that can be interpreted as the probability that light passes through the segment [si,si+1][s_{i},s_{i+1}]. Then, using T^(si)=∏j=1i−1pi\hat{T}(s_{i})=\prod_{j=1}^{i-1}p_{i}, and the approximation δiσi≈(1−exp⁡(−δiσi))=1−pi\delta_{i}\sigma_{i}\approx(1-\exp(-\delta_{i}\sigma_{i}))=1-p_{i}, the approximation in equation 26 takes the form

where the discrete probability is given by

establishing that τ^i\hat{\tau}_{i}, i=1,…,mi=1,\ldots,m is indeed a discrete probability. Note that although τ^m\hat{\tau}_{m} is not used in I^\hat{I} above, we added it for implementation ease. Since we use a bounding sphere (see equation 19) and MM is chosen so that sms_{m} is always outside this sphere, τ^m≈0\hat{\tau}_{m}\approx 0.