Multi-View Mesh Reconstruction with Neural Deferred Shading
Markus Worchel, Rodrigo Diaz, Weiwen Hu, Oliver Schreer, Ingo Feldmann, Peter Eisert
Introduction
The reconstruction of 3D objects based on multiple images is a long standing problem in computer vision. Traditionally, it has been approached by matching pixels between images, often based on photo-consistency constraints or learned features . More recently, analysis-by-synthesis, a technique built around the rendering operation, has re-emerged as a promising direction for reconstructing scenes with complex illumination, materials and geometry . At its core, parameters of a virtual scene are optimized so that its rendered appearance from the input camera views matches the camera images. If the reconstruction focuses on solid objects, these parameters usually include a representation of the object surface.
In gradient descent-based optimizations, analysis-by-synthesis for surfaces is approached differently depending on the differentiable rendering operation at hand. Methods that physically model light transport typically build on prior information such as light and material models . It is common to represent object surfaces with triangle meshes and use differentiable path tracers (e.g., ) to jointly optimize the geometry and parameters like the light position or material diffuse albedo. Due to the inherent priors, these methods do not generalize to arbitrary scenes.
Other methods instead model the rendering operation with neural networks , i.e., the interaction of material, geometry and light is partially or fully encoded in the network weights, without any explicit priors. Surfaces are often represented with implicit functions or more specifically implicit neural representations where the indicator function is modeled by a multi-layer perceptron (MLP) or any other form of neural network and optimized with the rendering networks in an end-to-end fashion.
While fully neural approaches are general, both in terms of geometry and appearance, current methods exhibit excessive runtime, making them impractical for domains that handle a large number of objects or multi-view video (e.g. of human performances ).
We propose Neural Deferred Shading (NDS), a fast analysis-by-synthesis method that combines triangle meshes and neural rendering. The rendering pipeline is inspired by real-time graphics and implements a technique called deferred shading : a triangle mesh is first rasterized and the pixels are then processed by a neural shader that models the interaction of geometry, material, and light. Since the rendering pipeline, including rasterization and shading, is differentiable, we can optimize the neural shader and the surface mesh with gradient descent (Figure 1). The explicit geometry representation enables fast convergence while the neural shader maintains the generality of the modeled appearance. Since triangle meshes are ubiquitously supported, our method can also be readily integrated with existing reconstruction and graphics pipelines. Our technical contributions include:
A fast analysis-by-synthesis pipeline based on triangle meshes and neural shading that handles arbitrary illumination and materials
A runtime decomposition of our method and a state-of-the-art neural approach
An analysis of the neural shader and the influence of its parameters
Related Work
There is a vast body of work on image-based 3D reconstruction for different geometry representations (e.g. voxel grids, point clouds and triangle meshes). Here, we will only focus on methods that output meshes and refer to Seitz et al. for an overview of other approaches.
During the past decades, multi-view methods have primarily exploited photo-consistency across images. Most of these approaches traverse different geometry representations like depth maps or point clouds before extracting (and further refining) a mesh, e.g. . Some methods directly estimate a mesh by deforming or carving an initial mesh (e.g. the visual hull) while minimizing an energy based on cross-image agreement . Recently, learned image features and neural shape priors have been used to drive the mesh deformation process . Our method is similar to full mesh-based approaches in the sense that we do not use intermediate geometry representations. However, we also do not impose strict assumptions on object appearance across images, which enables us to handle non-Lambertian surfaces and varying light conditions.
More than 20 years ago, Rockwood and Winget proposed to deform a mesh so that synthesized images match the input images. Their early analysis-by-synthesis method builds on an objective function with similar terms as ours (and many modern approaches): shading, silhouette, and geometry regularization. Later works propose similar techniques (e.g. ), yet all either assume known material or light parameters or restrict the parameter space with prior information, e.g. by assuming constant material across surfaces. In contrast, we optimize all parameters of the virtual scene and do not assume specific material or light models.
Optimizing many parameters of a complex scene, including geometry, material and light, has only lately become practical, arguably with the advent of differentiable rendering. Differentiable path tracers have been used on top of triangle meshes to recover not only the geometry but also the (spatially varying) reflectance and light , only from images. Related techniques can reconstruct transparent objects . Similarly, we perform analysis-by-synthesis by optimizing a mesh with differentiable rendering. However, we use rasterization and do not simulate light transport. In our framework, the view-dependent appearance is learned by a neural shader, which neither depends on material or light models nor imposes constraints on the acquisition setup (e.g. co-located camera and light).
Besides mesh reconstruction from real world images, analysis-by-synthesis with differentiable rendering has recently been used for image-based geometry processing and appearance-driven mesh simplification . Similar to us, these approaches deform a triangle mesh to reproduce target images, albeit their targets are fully synthetic.
2 Neural Rendering and Reconstruction
In this work, we understand neural rendering as training and using a neural network to synthesize color images from 2D input (e.g. semantic labels or UV coordinates), recently named “2D Neural Rendering” . Neural rendering has been used as integral part of 3D reconstruction methods with neural scene representations.
Introduced by Mildenhall et al. , neural radiance fields are volumetric scene representations used for 3D reconstruction, which are trained to output RGB values and volume densities at points along rays casted from different views. This idea has been adapted by a large number of recent works . Although not strictly based on neural rendering, these methods are related to ours by their analysis-by-synthesis characteristics. While the volumetric representation can handle transparent objects, most methods focus on view synthesis, so extracted surfaces lack geometric accuracy. Lassner et al. propose a volumetric representation based on translucent spheres, which are shaded with neural rendering. Similar to us, they jointly optimize geometry and appearance with a focus on speed, yet most details in their reconstruction are not present in the geometry but “hallucinated” by the neural renderer.
Implicit surfaces encoded in neural networks are another popular geometry representation for 3D reconstruction, most notably occupancy networks and neural signed distance functions . Here, surfaces are implicitly defined by level sets. For 3D reconstruction, these geometry networks are commonly trained end-to-end with a neural renderer to synthesize scenes that reproduce the input images. We also use neural rendering to model the appearance, but represent geometry explicitly with triangle meshes, which can be efficiently optimized and readily integrated into existing graphics workflows.
Similar to us, Thies et al. present a deferred mesh renderer with neural shading. However, their convolutional neural network-based renderer can “hallunicate” colors at image regions that are not covered by geometry, as opposed to our MLP-based shader. Most notably, their method aims at view synthesis, therefore only the renderer weights are optimized, while the mesh vertices remain unchanged.
Method
Given a set of images from calibrated cameras and corresponding masks , we want to estimate the 3D surface of an object shown in the images. To this end, we follow an analysis-by-synthesis approach: we find a surface that reproduces the images when rendered from the camera views. In this work, the surface is represented by a triangle mesh , consisting of vertex positions , a set of edges , and a set of faces . We solve the optimization problem using gradient descent and gradually deform a mesh based on an objective function that compares renderings of the mesh to the input images.
Faithfully reproducing the images via rendering requires an estimate of the surface material and illumination if we simulate light transport, e.g. with a differentiable path tracer . However, because our focus is mainly on the geometry, we do not accurately estimate these quantities and thus also avoid the limitations imposed by material and light models. Instead, we propose a differentiable mesh renderer that implements a deferred shading pipeline and handles arbitrary materials and light settings. At its core, a differentiable rasterizer produces geometry maps per view, which are then processed by a learned shader. See Figure 2 for an overview.
Our differentiable mesh renderer follows the structure of a deferred shading pipeline from real-time graphics: Given a camera , the mesh is rasterized in a first pass, yielding a triangle index and barycentric coordinates per pixel. This information is used to interpolate both vertex positions and vertex normals, creating a geometry buffer (g-buffer) with per-pixel positions and normals. In a second pass, the g-buffer is processed by a learned shader
In addition to a color image, the renderer also produces a mask that indicates if a pixel is covered by the mesh.
2 Objective Function
Finding an estimate of shape and appearance formally corresponds to solving the following minimization problem in our framework
where compares the rendered appearance of the estimated surface to the camera images and regularizes the mesh to avoid undesired vertex configurations.
The appearance objective is composed of two terms
2.2 Geometry Regularization
Naively moving the vertices unconstrained in each iteration quickly leads to undesirable meshes with degenerate triangles and self-intersections. We use a geometry regularization term that favors smooth solutions and is inspired by Luan et al. :
The normal consistency term is defined as
While some prior work (e.g., ) uses El Topo for robust mesh evolution, we found that our geometric regularization sufficiently avoids degenerate vertex configurations. Without El Topo, we are unable to handle topology changes but avoid its impact on runtime performance.
3 Optimization
Our optimization starts from an initial mesh that is computed from the masks and resembles a visual hull . Alternatively, it can start from a custom mesh.
Similar to prior work, we use a coarse-to-fine scheme for the geometry: Starting from a coarsely triangulated mesh, we progressively increase its resolution during optimization. Inspired by Nicolet et al. we remesh the surface with the method of Botsch and Kobbelt , halving the average edge length multiple times at fixed iterations. After mesh upsampling, we also increase the weights of the regularization terms by 4 and decrease the gradient descent step size for the vertices by 25 %, which we empirically found helps to improve the smoothness for highly tesselated meshes.
Since some quantities in geometry regularization (e.g. graph Laplacian) only depend on the connectivity of the mesh, we save time by precomputing them once after upsampling and reusing them in the iterations after.
Experimental Results
We implemented our method on top of the automatic differentiation framework PyTorch and use the Adam optimizer for momentum-based gradient descent. Our differentiable rendering pipeline uses the high-performance primitives by Laine et al. . In our experiments, we run 2000 gradient descent iterations and remesh after 500, 1000, and 1500 iterations. We randomly select one view per iteration to compute the appearance term and shade 75 % of mask pixels. The individual objective terms are weighted with , , , and . All time measurements were taken on a Windows workstation with an Intel Xeon 322.1 GHz CPU, 128 GB of RAM, and an NVIDIA Titan RTX GPU with 24 GB of VRAM.
We demonstrate that our method can be used for multi-view 3D reconstruction. Starting from a coarse visual hull-like mesh, it quickly converges to a reasonable estimate of the object surface. We test our method on the DTU multi-view dataset with the object selection and masks from previous work . We compare our results to two methods: (1) Colmap , a traditional SfM pipeline that serves as a baseline, and (2) IDR , a state-of-the-art analysis-by-synthesis method that uses neural signed distance functions as geometry representation. By default, our Colmap results include trimming (trim7) and we indicate untrimmed results explicitly (trim0).
Figure 4 shows qualitative results for two objects from the DTU dataset and Table 1 contains quantitative results for all objects. We used the official DTU evaluation script to generate the Chamfer-L1 scores and benchmarked all workflows for the time measurements (including data loading times). For IDR and our method we disabled any intermediate visualizations.
In absolute terms (millimeters), the reconstruction accuracy of our method is close to both the traditional baseline and the state-of-the-art neural method. Although Colmap reconstructs many surfaces accurately, only IDR and our method properly handle regions dominated by view-dependent effects (e.g. non-Lambertian materials) and produce watertight surfaces that can remain untrimmed. When Colmap is used without trimming, the reconstruction becomes more complete but it is less accurate than ours for some objects. Our method is limited by the topology and genus of the initial mesh and therefore cannot capture some geometric details that can be recovered with more flexible surface representations. We also observe that our surfaces are often not as smooth as IDR’s and concave regions are not as prominent. The latter is potentially related to the mesh stiffness induced by our geometry regularization.
On the other hand, our method is significantly faster: roughly 10 times faster than Colmap and 80 times faster than IDR in the default configuration. Since the number of iterations is a hyper parameter for IDR and our method (and also has different semantics), we show equal time results for a fair comparison of both (Figure 5). Our method quickly converges to a detailed estimate, while IDR only recovers a rough shape in the same time. Even after 50 minutes, IDR still lacks details present in our result.
Since IDR and our method have similar architectures, i.e., both perform analysis-by-synthesis and use gradient descent to jointly optimize shape and appearance, we can meaningfully compare the runtime in more detail. Table 2 shows the runtime of one gradient descent iteration (see the supplementary material for a full decomposition of our runtime).
An iteration in IDR takes roughly twice the time as in our method, with the majority of time spent for ray marching the implicit function. In contrast, the time taken for rasterizing the triangle mesh is negligible in our method. Most of the ray marching time in IDR can be attributed to evaluating the network, thus switching to a more shallow neural representation could be a way to reduce the runtime. The remaining time is spent on operations like root finding, which could be accelerated by a more optimized implementation.
However, the runtime difference in gradient descent iterations cannot be the sole reason for our fast convergence. Even though iterations in IDR require twice the time, it requires more than twice the total time to show the same level of detail as our method (see Figure 5).
In our method, we noticed that after mesh upsampling finer details quickly appear in the geometry. Thus, our fast convergence time might partially be related to the fact that we can locally increase the geometric freedom with a finer tessellation, while IDR and similar methods have no explicit control over geometric resolution.
2 Mesh Refinement
Many reconstruction workflows based on photo-consistency are very mature, established and deliver high-quality results. Yet, they can fail for challenging materials or different light conditions across images, producing smoothed outputs as a compromise to errors in photo-consistent matching.
Figure 6 shows refinement results for the heads of two persons. We use 1000 iterations and upsample the mesh once. Since the initial mesh is already a good estimate of the true surface but the neural shader is randomly initialized, we rebalance the gradient descent step sizes so that the shader progresses faster than the geometry. We are able to recover details that are lost in the initial reconstruction. Very fine details like the facial hair are still challenging and lead to slight noise in our refined meshes.
3 Analysis of the Neural Shader
Although recovering geometry is the main focus of our work, investigating the neural shader after optimization can potentially give insights into the reconstruction procedure and the information that is encoded in the network.
Figure 7 shows that rendering with a trained shader can be used for basic view synthesis, producing reasonable results for input and novel views. Thus, the shader seems to learn a meaningful representation of appearance, disentangled from the geometry. If desired, the view synthesis quality can be further improved by continuing the optimization of the shader while keeping the mesh fixed.
The neural shader is a composite of two functions (Figure 3)
We investigate the behavior of the view-dependent part by replacing feature vectors in the positional latent space and then extracting colors with the function . More specifically, we replace all features representing one material by a feature representing another material (Figure 8). The results suggest that the function reasonably generalizes to view and normal directions not encountered in combination with the replacement feature, which in this example leads to geometric features of the mesh being still perceivable and thus allows simple material editing.
4 Ablation Studies
We experiment with different encodings and network sizes for the position dependent part of the neural shader. In Figure 9, we show the different results using positional encoding (PE) , Gaussian Fourier features (GFF) , sinusoidal activations (SIREN) and a standard MLP with ReLU activations.
While some of these methods can generate acceptable renderings, they do not necessarily guarantee a sharper geometry. In particular, we observe that although the image rendered with SIREN or GFF has sharp features, the geometric detail of the mesh is less adequate. A possible explanation is that the network might quickly overfit and compensate for geometrical inaccuracies only in the appearance. Conversely, finding the correct direction along which to move the mesh vertices might be more difficult without positional encoding. In our experiments, we have obtained accurate reconstructions using positional encoding with 4 octaves, as opposed to the recommended 10 .
We also examined the effect of different network sizes on the geometry and appearance (see supplementary). While there are no dramatic differences between the results, we observe that configurations with less than 2 layers or with more than 512 units per layer result in fewer geometric details. Additional ablation studies for the initial geometry and the objective function can also be found in the supplementary.
Concluding Remarks
We have presented a fast analysis-by-synthesis pipeline for 3D surface reconstruction from multiple images. Our method jointly optimizes a triangle mesh and a neural shader as representations of geometry and appearance to reproduce the input images, thereby combining the speed of triangle mesh optimization with the generality of neural rendering.
Our approach can match state-of-the-art methods in reconstruction accuracy, yet is significantly faster with average runtimes below 10 minutes.
Using triangle meshes as geometry representation makes our method fully compatible with many traditional reconstruction workflows. Thus, instead of replacing the complete reconstruction architecture, our method can be integrated as a part, e.g., a refinement step.
Finally, a preliminary analysis of the neural shader suggests that it decomposes appearance in a natural way, which could help our understanding of neural rendering and enable simple ways for altering the scene appearance (e.g. for material editing).
While triangle meshes are a simple and fast representation, using them robustly in 3D reconstruction with differentiable rendering is still challenging. We currently avoid undesired mesh configurations (e.g. self-intersections) with carefully weighted geometry regularizers that steer the optimization towards smooth solutions. Finding an appropriate balance between smoothness and rendering terms is not always straightforward and can require fine-tuning for custom data. In this regard, we are excited about integrating very recent work that proposes gradient preconditioning instead of geometry regularization .
Topology changes are a challenge for most mesh-based methods, including ours, and computationally expensive to handle (e.g. with El Topo ). Likewise, adaptive reconstruction-aware subdivision would be preferable over standard remeshing, potentially including learned priors . Instead of moving vertices directly, a (partially pretrained) neural network could drive the deformation , making it more efficient, detail-aware, and less dependent on geometric regularization.
While neural shading is a powerful component of our system, allowing us to handle non-Lambertian surfaces and complex illumination, it is also a major obstruction for interpretability. The effect of changes to the network architecture can only be evaluated with careful experiments and often the black box shader behaves in non-intuitive ways. We experimentally provided preliminary insights but also think that a more thorough analysis is needed. Alternatively, physical light transport models could be combined with more specialized neural components (e.g. irradiance or material) to isolate their effects. In this context, pre-training components to include learned priors also seems like a promising direction.
Although the shader can handle arbitrary material and light in theory, this claim needs more investigation. A possible path is exhaustive experiments with artificial scenes, starting with the simplest cases (how well does it handle a perfectly Lambertian surface?).
Acknowledgements
This work is part of the INVICTUS project that has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 952147.
References
Appendix A Dataset Details
Our evaluation is based on the scripts included in the official DTU MVS dataset release . In our experiments, we use a derived dataset that includes object masks and was assembled in the context of prior work . Our implementation expects the input in a different file structure and we will release the converted dataset with our paper.
We are not aware of any copyright or license attached to the DTU dataset. This dataset is widely used in the scientific community.
A.2 Human Body Dataset
The human body dataset used in the mesh refinement experiments consist of recordings of two persons, taken in a volumetric capture studio . The subjects are hired actors and they consented in writing to be recorded and to their data being used for research purposes. The consent covers showing the subjects’ faces in publications.
Appendix B Implementation Details
Our implementation is built on top of the automatic differentiation framework PyTorch and includes code released in the context of previous publications . For remeshing we use a library with Python bindings ; for differentiable rasterization we use the high-performance primitives by Laine et al. . Additionally, we rely on a variety of other libraries .
In our implementation, we do not assume normalized camera positions. However, we require a rectangular bounding box to normalize the domain to a cube centered at with side length 2.
For 3D reconstruction, we also use the bounding box to construct the initial mesh (Figure 10): we place a grid of points () inside the bounding box volume and project the points into each camera image. If a point lies outside any image or mask region, it is removed. We reconstruct a mesh surface from the remaining points with marching cubes.
Appendix C Experiment Details
In the 3D reconstruction experiments, we set the gradient descent step size to for both the mesh vertices and the neural shader. For mesh refinement, we use a step size of for the vertices and for the shader, progressing the shader faster than the mesh.
For Colmap , we use the official release 3.6 with CUDA support . Similar to prior work , we clean the dense point clouds with masks. We also perform trimming after the screened Poisson surface reconstruction (“trim7”).
For IDR , we use the official implementation and run the reconstruction experiments with camera training, using the dtu_trained_cameras.conf configuration.
In the main work, we compare the runtime of one gradient descent iteration of our method to the one of IDR . The comparison uses data produced by an extensive profiling mode that we implemented for our method and a simpler mode implemented on top of IDR. When profiling, we disable all intermediate outputs (e.g. visualizations). In the case of IDR, we profile calls to RayTracing.forward and inside these calls those to ImplicitNetwork.forward, obtaining measurements for Geometry rendering and SDF evaluation, respectively.
For our method, we record the runtime of most operations during reconstruction. Table 3 shows the full decomposition for one sample from the DTU dataset. Computing the shading term of our objective function and computing the gradients with back propagation are the main bottlenecks of our method.
Appendix D Additional Experimental Results
We investigate how initial meshes with different resolutions, topology, and geometric distance to the target affect the reconstruction (Figure 11). Very coarse initial meshes result in missing details, while fine meshes provide too much geometric freedom, which leads to artifacts. Since we do not support topology changes, holes are only reconstructed if they are present in the initial geometry.
The initial meshes for all DTU objects consist of around 2000 triangles and the output meshes of around 100000 triangles. In terms of scalability, our method handles high-resolution meshes efficiently: in the last 500 iterations (with 2000 iterations in total), the optimization runs on a mesh with the output resolution and still performs fast gradient descent steps.
We investigate the influence of the individual terms of our objective function (Figure 12). Without the Laplacian term, we observe noticeable bumps and crack-like artifacts. The normal term has a small influence but improves smoothness around edges. Without the silhouette term, the object boundaries are not properly reconstructed. Without the shading term, the reconstructed shape resembles a smooth visual hull, missing almost all details.
In Figure 13, we show the results for different architecture configurations. Using very few units per layer leads to sharper geometry while the opposite leads to a smoother surface. In the first case, a sharper geometry does not equate with a better estimation of the surface (e.g., shadows are baked into the geometry). In the latter case, we obtain a smooth geometry but we lose geometrical sharpness. Moreover, with the increase in the network parameters, the optimization time grows accordingly.
We found that using 3 layers with 256 units per layers is a good compromise between network complexity and expressive power and yields the best results.
For SIREN , we noticed that the optimization procedure diverges with our default learning rate. We lowered it to to obtain meaningful results. Despite the lower learning rate, SIREN still converges quickly.
Appendix E Interactive Viewer
Since the renderer we use for optimization has the layout of a standard real-time graphics pipeline, the shader part (i.e. the neural shader) can be readily integrated into other graphics pipelines after training, allowing the interactive synthesis of novel views.
As an example, we implemented a “neural” viewer that uses OpenGL to rasterize positions and normals. Then, we use OpenGL-CUDA interoperability and PyTorch to shade the buffers with a pre-trained neural shader. We can envision an exciting extension where the shader is directly compiled to GPU byte code and used similar to other shaders written in high-level shading languages.
Appendix F Failure Cases
Figure 14 shows some failure cases of our reconstruction method. Since the reconstruction starts from a visual hull-like mesh, deep concavities require large movements in the mesh. This movement is driven by relatively weak shading gradients, which cannot recover these concavities before the mesh becomes overly stiff (e.g. because the gradient descent step size is reduced during remeshing). On the other hand, silhouette gradients are larger than shading gradients, so the mesh easily “grows” to fill the masks.
We also observed that the remeshing operation sometimes produces tangled meshes for some objects. This failure case is unrecoverable and the reconstruction needs to be restarted. We are investigating the issue and will coordinate with the authors of the remeshing library.
Some regions show weak structure in the majority of images for a variety of reasons, for example because they are overexposed. Since we randomly sample one camera view per iteration, there is a high probability to select an image with low structure information for these regions. Using such a view can drive the regions’ vertices in unexpected directions away from the actual surface. If such movement occurs at the wrong time (e.g. right before remeshing), the optimization often fails to correct those vertices.
Appendix G Ethical Considerations
Our method has the same potential for misuse as other multi-view 3D reconstruction pipelines. For example, it could be used to digitally reconstruct dangerous objects (e.g. weapons, weapons parts, ammunition) for reproduction in 3D printing.
Because we also train an appearance model, a malicious actor could potentially modify or edit materials of a reconstructed scene, for example to create fakes by modifying the skin color of a person.
With publicly available code, means for mitigation would be difficult to implement from our side. However, we feel that the overall potential for misuse is low and there are few obvious paths to malicious use or potential damage. Also, our method requires a set of calibrated images, thus is much less accessible than 3D reconstruction or deep fakes based on single images.