UNISURF: Unifying Neural Implicit Surfaces and Radiance Fields for Multi-View Reconstruction
Michael Oechsle, Songyou Peng, Andreas Geiger
Introduction
Capturing the geometry and appearance of 3D scenes from a set of images is one of the cornerstone problems in computer vision. Towards this goal, coordinate-based neural models have emerged as a powerful tool for 3D reconstruction of geometry and appearance within the last years.
Many recent methods employ continuous implicit functions parameterized with neural networks as 3D representations of geometry or appearance . These neural 3D representations have shown impressive performance on geometry reconstruction and novel view synthesis from multi-view images. Besides the choice of the 3D representation (e.g., occupancy field, unsigned or signed distance field), one key element for neural implicit multi-view reconstruction is the rendering technique. While some of these works represent the implicit surface as level set and hence render the appearance from surfaces , others integrate densities by drawing samples along the viewing rays .
In existing work, surface rendering techniques have shown impressive performance in 3D reconstruction . However, they require per-pixel object masks as input and an appropriate network initialization since surface rendering techniques only provide gradient information locally where a surface intersects with a ray. Intuitively speaking, optimizing wrt. local gradients can be seen as an iterative deformation procedure applied to an initial neural surface which is often initialized as a sphere. Additional constraints such as mask supervision are necessary for converging to a valid surface, see Fig. 2 for an illustration. Due to their reliance on masks, surface rendering methods are limited to object-level reconstruction and do not scale to larger scenes.
On contrary, volume rendering methods like NeRF have shown impressive results for novel view synthesis, also for larger scenes. However, surfaces extracted as level sets of the underlying volume density are usually non-smooth and contain artifacts due to the flexibility of the radiance field representation which does not sufficiently constrain the 3D geometry in the presence of ambiguities, see Fig. 3.
Contributions: In this paper, we propose UNISURF (UNIfied Neural Implicit SUrface and Radiance Fields) a principled unified framework for implicit surfaces and radiance fields, with the goal of reconstructing solid (i.e., non-transparent) objects from a set of RGB images. Our framework combines the benefits of surface rendering with those of volume rendering, enabling the reconstruction of accurate geometry from multi-view images without masks. By recovering implicit surfaces, we are able to gradually decrease the sampling region for volume rendering during optimization. Starting with large sampling regions enables capturing coarse geometry and resolving ambiguities during early iterations. At a later stage, we draw samples closer to the surface which improves reconstruction accuracy. We show that our approach enables capturing accurate geometry without mask supervision on the DTU MVS dataset , attaining results competitive with state-of-the-art implicit neural reconstruction methods like IDR which use strong mask supervision. Moreover, we also demonstrate our method on scenes from the BlendedMVS dataset as well as synthetic indoor scenes from SceneNet . Code is available at https://github.com/autonomousvision/unisurf.
Related Work
In this section, we first discuss related work from the domain of 3D reconstruction from multi-view images. Next, we provide an overview of recent works on neural implicit representations as well as differentiable rendering.
3D Reconstruction from Multi-View Images: Reconstructing 3D geometry from multiple images has been a longstanding computer vision problem . Before the era of deep learning, classic multi-view stereo (MVS) methods focus on either matching features across views or representing shapes with a voxel grid . The former approaches usually have a complex pipeline requiring additional steps like fusing depth information and meshing , while the latter ones are limited to low resolution due to cubic memory requirements. In contrast, neural implicit representations for 3D reconstruction do not suffer from discretization artifacts as they represent surfaces by the level set of a neural network with continuous outputs.
Recent learning-based MVS methods attempt to replace some parts of the classic MVS pipeline. For instance, some works learn to match 2D features, fuse depth maps , or infer depth maps from multi-view images. Contrary to these learning-based MVS approaches, our method only requires weak 2D supervision during optimization. Moreover, our method yields high-quality 3D geometry and synthesizes photorealistic and consistent novel views.
Neural Implicit Representations: Recently, neural implicit functions have emerged as an effective representation of 3D geometry and appearance as they represent 3D content continuously and without discretization while simultaneously having a small memory footprint. Most of these methods require 3D supervision. However, several recent works demonstrated differentiable rendering for training directly from images . We divide these methods into two groups: surface rendering and volume rendering.
Surface rendering approaches, including DVR and IDR , determine the radiance directly on the surface of an object and provide a differentiable rendering formulation using implicit gradients. This allows for optimizing neural implicit surfaces from multi-view images. Conditioning on the viewing direction allows IDR to capture a high level of detail, even in the presence of non-lambertian surfaces. However, both DVR and IDR require pixel-accurate object masks for all views as input. In contrast, our method leads to similar reconstructions without requiring masks.
NeRF and follow-ups use volume rendering by learning alpha-compositing of a radiance field along rays. This method has shown impressive results on novel view synthesis and does not require mask supervision. However, the recovered 3D geometry is far from satisfactory, see Fig. 3. Several follow-up works (Neural Body D-NeRF and NeRD ) extract meshes using the volume density from NeRF, but none of them considers optimizing surfaces directly. Unlike these works, we aim at capturing accurate geometry and propose a volume rendering formulation that provably approaches surface rendering in the limit.
Background
The two main ingredients for learning neural implicit 3D representations from multi-view images are the 3D representation and the rendering technique linking the 3D representation and the 2D observations. This section provides the relevant background on implicit surface and volumetric radiance representations which we unify in this paper for the case of solid (non-transparent) objects and scenes.
Implicit Surface Models: Occupancy Networks represent surfaces as the decision boundary of a binary occupancy classifier, parameterized by a neural network
where is retrieved by root finding along ray , see for details. The parametersFor convenience, we use the same symbol for all model parameters. of the occupancy field and the color field are determined by optimizing a reconstruction loss via gradient descent as described in .
While surface rendering allows for accurately estimating geometry and appearance, existing approaches strongly rely on supervision with object masks as surface rendering methods are only able to reason about rays that intersect a surface.
Here, is the accumulated transmittance along the ray and is the distance between adjacent samples. As Eq. (3) is differentiable, the parameters of the density field and the color field can be estimated by optimizing a reconstruction loss. We refer to for details.
While NeRF does not require object masks for training due to its volumetric radiance representation, extracting the scene geometry from the volume density requires careful tuning of the density threshold and leads to artifacts due to the ambiguity present in the density field, see Fig. 3.
Method
We now describe our main contribution. In contrast to NeRF which is also applicable to non-solid scenes (e.g., fog, smoke), we restrict our focus to solid objects that can be represented by 3D surfaces and view-dependent surface colors. Our method exploits both, the power of volumetric radiance representations to learn coarse scene structure without mask supervision as well as surface rendering which acts as an inductive bias to represent objects by a set of precise 3D surfaces, leading to accurate reconstructions.
We start by noting that Eq. (3) can be rewritten asWe drop dependencies on the model parameters for clarity.
with alpha values . Assuming solid objects, becomes a discrete occupancy indicator variable which either takes in free space and in occupied space as value:
We recognize this expression as the image formation model for solid objects where the term o(\mathbf{x}_{i})\prod_{j<i}\bigl{(}1-o(\mathbf{x}_{j})\bigr{)} evaluates to for the first occupied sample along ray and to for all other samples. \prod_{j<i}\bigl{(}1-o(\mathbf{x}_{j})\bigr{)} is an indicator for visibility which is if there exists no occupied sample with before sample . Thus, takes the color of the first occupied sample along ray .
To unify implicit surface and volumetric radiance models, we parameterize directly by a continuos occupancy field (1) as opposed to predicting volume density . Following , we condition the color field on the surface normal and a feature vector of the geometry network which empirically induces a useful bias as also observed in for the case of implicit surfaces. Importantly, our unified formulation allows for both volume and surface rendering
where is retrieved by root-finding along ray and , denote the normal and geometry features at , respectively. Note that depends on the occupancy field , but we have dropped this dependency here for clarity. For further details, we refer the reader to the supplementary material.
The advantage of this unified formulation is that it allows for both rendering on the surface directly and rendering throughout the entire volume which enables gradually removing ambiguities during optimization. As evidenced by our experiments, it is indeed critical to combine both for obtaining accurate reconstructions without mask supervision. Being able to quickly recover the surface via root-finding enables more effective volume rendering, successively focusing on and refining the object surfaces as we will describe in Section 4.3. Furthermore, surface rendering enables faster novel view synthesis as illustrated in Fig. 5.
2 Loss Function
We optimize the following regularized loss function
Here, denotes the set of all pixels/rays in the minibatch, is the set of corresponding surface points, is the observed color for pixel/ray and is a small random uniform 3D perturbation. The normal at is given by
which can be computed using double backpropagation .
3 Optimization
The key hypothesis of implicit surface models is that only the region at the first intersection point with the surface contributes to the rendering equation. However, this assumption is not true during early iterations where the surface is not well defined. Consequently, existing methods require strong mask supervision. Conversely, during later iterations, knowledge of the approximate surface is valuable for drawing informative samples when evaluating the volume rendering equation in Eq. (7). Therefore, we utilize a training schedule with a monotonically decreasing sampling interval for drawing samples during volume rendering, as visualized in Fig. 4. In other words, during early iterations, the samples cover the entire optimization volume, effectively bootstrapping the reconstruction process using volume rendering. During later iterations, the samples are drawn closer around the estimated surface. As the surface can be estimated directly from the occupancy field via root-finding , this eliminates the need for hierarchical two-stage sampling as in NeRF. Our experiments demonstrate that this procedure is particularly effective for estimating accurate geometry, while it allows for resolving ambiguities during early iterations.
More formally, let . We obtain samples by drawing depth values using stratified sampling within the interval centered at :
During training, we start with a large interval and gradually decrease for more accurate sampling and optimization of the surface using the following decay schedule
where denotes the iteration number and is a hyperparameter. In fact, it can be shown that for and , volume rendering (7) indeed approaches surface rendering (8): . A formal proof of this limit is provided in the supplementary material.
As evidenced by our experiments, the decay schedule in (14) is critical for capturing detailed geometry as it combines volume rendering of large and uncertain volumes in the beginning of training with surface rendering towards the end of training. To reduce free space artifacts, we combine these samples with points sampled randomly between the camera and the surface. For rays without surface intersection, we use stratified sampling on the entire ray.
4 Implementation Details
Architecture: Similar to Yariv et al. , we use an 8-layer MLP with a Softplus activation function and a hidden dimension of for the occupancy field . We initialize the network such that the decision boundary is a sphere . In contrast, the radiance field is parameterized as a 4-layer ReLU MLP. We encode the 3D location and the viewing direction using Fourier features at octaves. We empirically found for the 3D location and for the viewing direction to work best.
Optimization: In all experiments, we fit our model to multi-view images of a single scene. During optimization of the model parameters, we first randomly sample a view and then pixels/rays from this view based on the camera intrinsics and extrinsics. Next, we render all rays to compute the loss function in Eq. (9). For root-finding, we use uniform sampled points and apply the secant method with steps . For our rendering procedure, we use query points inside the interval and in the free space between the camera and the lower bound of the interval. The interval decay parameters are , and . We use Adam with a learning rate of and optimize for pixels per iteration with two decay steps after k and k iterations. In total, we train our models for k iterations.
Inference: Our method allows for inferring 3D shapes as well as for synthesizing novel view images. For synthesizing images, we can render our representation in two different ways, we can either use volume rendering or surface rendering. In Fig. 5, we show that both rendering approaches lead to similar results. However, we observe that surface rendering is faster than volume rendering.
To extract meshes, we apply the Multiresolution IsoSurface Extraction (MISE) algorithm from . We use as the initial resolution and up-sample the mesh in 3 steps without gradient-based refinement.
Experimental Evaluation
We conduct experiments on multi-view 3D reconstruction to assess our method. First, we provide a qualitative and quantitative comparison of our approach to existing methods (IDR , NeRF , COLMAP ) on the widely used DTU MVS dataset . Second, we show a qualitative comparison for samples from the BlendedMVS dataset and synthetic renderings of scenes from the SceneNet dataset . Third, we analyze our rendering procedure and loss function in an ablation study. In the supplementary, we provide results on the LLFF dataset .
To validate the effectiveness of our method, we compare it to three different baselines.
COLMAP : We consider COLMAP as a classical MVS baseline, as it shows strong performance on multi-view reconstruction and is widely used in related works . We reconstruct a mesh from the output of COLMAP using screened Poisson Surface reconstruction (sPSR) . Following , we show results for the quantitatively best trim parameter and for the setting that results in watertight meshes (trim parameter ).
NeRF : Although NeRF targets novel view synthesis, its volume density admits the extraction of geometry. To extract meshes from NeRF, we define a density threshold of 50. We validate this choice in the supplementary.
IDR : IDR is the state-of-the-art multi-view reconstruction method for neural implicit surfaces. IDR reconstructs surfaces with an impressively high level of detail and handles specular surfaces, but requires input masks. We do not compare to DVR as IDR’s view-dependent modeling has been demonstrated to be superior to DVR, see .
2 Datasets
DTU MVS Dataset : The DTU MVS dataset contains 49 to 64 images at a resolution of as well as extrinsic and intrinsic camera parameters for all views. The dataset consists of objects with different shapes and appearances. Non-lambertian appearance effects make some of the objects particularly challenging. For each scan, ground truth 3D shapes, as well as the official evaluation procedure, are availablehttps://roboimagedata.compute.dtu.dk. Like previous works , we use the ”Surface” method of the evaluation script, and evaluate all methods on meshes cleaned with the respective masks. The official evaluation procedure calculates the Chamfer distance between sampled points of the predicted shapes and ground truth shapes provided in the dataset. For evaluating IDR , we use the pixel-accurate masks for all images which are provided by the authors of IDR.
BlendedMVS Dataset : The BlendedMVS dataset is a large-scale dataset containing multi-view images with respective camera extrinsics and intrinsics. We use examples from the BlendedMVS low-res set with an image resolution of . These examples contain 24 to 64 different views of unmasked images. We define the scene as such that the object is in the center and the closest camera lies near the unit sphere.
SceneNet Dataset : For testing our model on complex indoor scenes, we take two scenes from for evaluation. BlenderProc is applied for rendering images of a part of the scenes containing multiple objects. The first scene is a bedroom scene with a bed, a lamp and a bedside table. The other scene is a living room scene with a sofa, a curtain and a round table. We use 83 and 40 images, respectively.
3 Comparison on DTU
In Table 1, we quantitatively compare our method to the baselines on the DTU MVS dataset. While the COLMAP baseline with trim parameter shows the best performance on the Chamfer distance, it produces non-watertight meshes with incomplete surfaces. Our method performs nearly on par with the state-of-the-art neural implicit model IDR while not relying on strong mask supervision. NeRF and COLMAP () also do not use input masks but exhibit worse performance in terms of Chamfer distance.
In Fig. 6, we show qualitative results for our method and the baselines. While COLMAP provides detailed reconstructions, it leads to incomplete geometry due to trimming. For NeRF, holes and noise artifacts can be observed in the reconstructions. In contrast, our approach and IDR (with input masks) produce accurate surfaces with high-quality details. We remark that our model captures the overall spatial arrangement of the scene accurately while being able to also capture geometric details, e.g., the teeth of the skull and other surface details.
4 Comparison on BlendedMVS and SceneNet
To show our model’s capabilities on more diverse scenes, we use samples of the BlendedMVS dataset and indoor scenes from SceneNet. As there exist no object masks for these scenes, we only consider COLMAP and NeRF as baseline methods. While we have tried running IDR, none of the scenes converged, resulting in degenerate outputs. For scenes with complex backgrounds, we use a background model that learns to capture appearance information outside of our region interest, we refer the reader to the supplementary material for more details.
Our qualitative results in Fig. 8 provide evidence that, unlike existing implicit surface models, our method is able to reconstruct plausible geometry for complex scenes with multiple objects and backgrounds. While COLMAP works well on indoor scenes, it shows artifacts on uniformly colored regions (e.g., round table in the second row). As for the BlendedMVS experiment, NeRF can reason about the overall spatial structure but shows less accurate surfaces with significantly higher levels of noise compared to UNISURF. More results can be found in the supplementary material.
5 Ablation study
We finally investigate the impact of rendering design choices and show ablations of the loss function.
Losses: In Fig. 7, we also show an ablation study of our surface regularization term in (9). Without this regularizer surfaces become less smooth in flat and ambiguous areas, e.g., at the table. This regularization term is particularly useful for regions that are observed less frequently, as it incorporates an inductive bias towards smooth surfaces.
Discussion and Conclusion
This work presents UNISURF, a unified formulation of implicit surfaces and radiance fields for capturing high-quality implicit surface geometry from multi-view images without input masks. We believe that neural implicit surfaces and advanced differentiable rendering procedures play a key role in future 3D reconstruction methods. Our unified formulation shows a path towards optimizing implicit surfaces in a more general setting than possible before.
Limitations: By design, our model is limited to represent solid, non-transparent surfaces. Overexposed and textureless regions are also limiting factors that lead to inaccuraries and non-smooth surfaces. Furthermore, the reconstructions are less accurate at rarely visible regions in the images. Limitations are discussed in more detail in the supplementary.
In future work, for resolving ambiguities from rarely visible and texture-less regions, a prior is necessary for reconstruction. While we incorporate an explicit smoothness prior during optimization, learning a probabilistic neural surface model which captures regularities and uncertainty across objects would help to resolve ambiguities, leading to more accurate reconstructions.
This work was supported by an NVIDIA research gift, the ERC Starting Grant LEGO-3D (850533) and the DFG EXC number 2064/1 - project number 390727645. Songyou Peng is supported by the Max Planck ETH Center for Learning Systems.