Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, Andreas Geiger
Introduction
In recent years, learning-based 3D reconstruction approaches have achieved impressive results . By using rich prior knowledge obtained during the training process, they are able to infer a 3D model from as little as a single image. However, most learning-based methods are restricted to synthetic data, mainly because they require accurate 3D ground truth models as supervision for training.
To overcome this barrier, recent works have investigated approaches that require only 2D supervision in the form of depth maps or multi-view images. Most existing approaches achieve this by modifying the rendering process to make it differentiable . While yielding compelling results, they are restricted to specific 3D representations (e.g. voxels or meshes) that suffer from discretization artifacts and the computational cost limits them to small resolutions or deforming a fixed template mesh. At the same time, implicit representations for shape and texture have been proposed which do not require discretization during training and have a constant memory footprint. However, existing approaches using implicit representations require 3D ground truth for training and it remains unclear how to learn implicit representations from image data alone.
Contribution: In this work, we introduce Differentiable Volumetric Rendering (DVR). Our key insight is that we can derive analytic gradients for the predicted depth map with respect to the network parameters of the implicit shape and texture representation (see Fig. 1). This insight enables us to design a differentiable renderer for implicit shape and texture representations and allows us to learn these representations solely from multi-view images and object masks. Since our method does not have to store volumetric data in the forward pass, its memory footprint is independent of the sampling accuracy of the depth prediction step. We show that our formulation can be used for various tasks such as single- and multi-view reconstruction, and works with synthetic and real data. In contrast to , we do not need to condition the texture representation on the geometry, but learn a single model with shared parameters that represents both geometry and texture. Our code and data are provided at https://github.com/autonomousvision/differentiable˙volumetric˙rendering.
Related Work
3D Representations: Learning-based 3D reconstruction approaches can be categorized wrt. the representation they use as voxel-based , point-based , mesh-based, or implicit representations .
Voxels can be easily processed by standard deep learning architectures, but even when operating on sparse data structures , they are limited to relatively small resolution. While point-based approaches are more memory-efficient, they require intensive post-processing because of missing connectivity information. Most mesh-based methods do not perform post-processing, but they often require a deformable template mesh or represent geometry as a collection of 3D patches which leads to self-intersections and non-watertight meshes.
To mitigate these problems, implicit representations have gained popularity . By describing 3D geometry and texture implicitly, e.g., as the decision boundary of a binary classifier , they do not discretize space and have a fixed memory footprint.
In this work, we show that the volumetric rendering step for implicit representations is inherently differentiable. In contrast to previous works, this allows us to learn implicit 3D shape and texture representations using 2D supervision.
3D Reconstruction: Recovering 3D information which is lost during the image capturing process is one of the long-standing goals of computer vision . Classic multi-view stereo (MVS) methods usually match features between neighboring views or reconstruct the 3D shape in a voxel grid . While the former methods produce depth maps as output which have to be fused in a lossy post-processing step, e.g., using volumetric fusion , the latter approaches are limited by the excessive memory requirements of 3D voxel grids. In contrast to these highly engineered approaches, our generic method directly outputs a consistent representation in 3D space which can be easily converted into a watertight mesh while having a constant memory footprint.
Recently, learning-based approaches have been proposed that either learn to match image features , refine or fuse depth maps , optimize parts of the classical MVS pipeline , or replace the entire MVS pipeline with neural networks that are trained end-to-end . In contrast to these learning-based approaches, our method can be supervised from 2D images alone and outputs a consistent 3D representation.
Differentiable Rendering: We focus on methods that learn 3D geometry via differentiable rendering in contrast to recent neural rendering approaches which synthesize high-quality novel views but do not infer the 3D object. They can again be categorized by the underlying representation of 3D geometry that they use.
Loper et al. propose OpenDR which approximates the backward pass of the traditional mesh-based graphics pipeline and has inspired several follow-up works . Liu et al. replace the rasterization step with a soft version to make it differentiable. While yielding compelling results in reconstruction tasks, these approaches require a deformable template mesh for training, restricting the topology of the output.
Another line of work operates on voxel grids . Paschalidou et al. and Tulsiani et al. propose a probabilistic ray potential formulation. While providing a solid mathematical framework, all intermediate evaluations need to be saved for backpropagation, restricting these approaches to relatively small-resolution voxel grids.
Liu et al. propose to infer implicit representations from multi-view silhouettes by performing max-pooling over the intersections of rays with a sparse number of supporting regions. In contrast, we use texture information enabling us to improve over the visual hull and to reconstruct concave shapes. Sitzmann et al. infer implicit scene representations from RGB images via an LSTM-based differentiable renderer. While producing high-quality renderings, the geometry cannot be extracted directly and intermediate results need to be stored for computing gradients. In contrast, we show that volumetric rendering is inherently differentiable for implicit representations. Thus, no intermediate results need to be saved for the backward pass.
Method
In this section, we describe our Differentiable Volumetric Rendering (DVR) approach. We first define the implicit neural representation which we use for representing 3D shape and texture. Next, we provide a formal description of DVR and all relevant implementation details. An overview of our approach is provided in Fig. 2.
Shape: In contrast to discrete voxel- and point-based representations, we represent the 3D shape of an object implicitly using the occupancy network introduced in :
Texture: Similarly, we can describe the texture of a 3D object using a texture field
Supervision: Recent works have shown that it is possible to learn and with 3D supervision (i.e., ground truth 3D models). However, ground truth 3D data is often very expensive or even impossible to obtain for real-world datasets. In the next section, we introduce DVR, an alternative approach that enables us to learn both and from 2D images alone. For clarity, we drop the condition variable in the following.
2 Differentiable Volumetric Rendering
Our goal is to learn and from 2D image observations. Consider a single image observation. We define a photometric reconstruction loss
Gradients: To obtain gradients of with respect to , we first use the multivariate chain rule:
Here, denotes the Jacobian matrix for a vector-valued function with vector-valued argument and indicates matrix multiplication. By exploiting , we obtain
since both as well as depend on . Because is defined implicitly, calculating is non-trivial. We first exploit that lies on the ray from through . For any pixel , this ray can be described by where is the vector connecting and (see Fig. 3). Since must lie on , there exists a depth value , such that . We call the surface depth. This enables us to rewrite as
For computing the gradient of the surface depth with respect to we exploit implicit differentiation . Differentiating on both sides wrt. , we obtain:
Rearranging (7), we arrive at the following closed form expression for the gradient of the surface depth :
We remark that calculating the gradient of the surface depth wrt. the network parameters only involves calculating the gradient of at wrt. the network parameters and the surface point . Thus, in contrast to voxel-based approaches , we do not have to store intermediate results (e.g., volumetric data) for computing the gradient of the loss wrt. the parameters, resulting in a memory-efficient algorithm. In the next section, we describe our implementation of DVR which makes use of reverse-mode automatic differentiation to compute the full gradient (4).
3 Implementation
To use automatic differentiation, we have to implement the forward and backward pass for the surface depth prediction step . In the following, we describe how both passes are implemented. For more details, we refer the reader to the supplementary material.
Forward Pass: As visualized in Fig. 3, we can determine by finding the first occupancy change on the ray . To detect an occupancy change, we evaluate the occupancy network at equally-spaced samples on the ray . Using a step size of , we can express the coordinates of these point in world-coordinates as
where determines the closest possible surface point. We first find the smallest for which changes from free space () to occupied space ():
We obtain an approximation to the surface depth by applying the iterative secant method to the interval . In practice, we compute the surface depth for a batch of points in parallel. It is important to note that we do not need to unroll the forward pass or store any intermediate results as we exploit implicit differentiation to directly obtain the gradient of wrt. .
Backward Pass: The input to the backward pass is the gradient of the loss wrt. a single surface depth prediction. The output of the backward pass is , which can be computed using (8). In practice, however, we would like to implement the backward pass not only for a single surface depth but for a whole batch of depth values.
We can implement this efficiently by rewriting as
Importantly, the left term in (11) corresponds to a normal backward operation applied to the neural network and the right term in (11) is just an (element-wise) scalar multiplication for all elements in the batch. We can hence conveniently compute the backward pass of the operator by first multiplying the incoming gradient element-wise with a factor and then backpropagating the result through the operator . Both operations can be efficiently parallelized in common deep learning frameworks.
4 Training
During training, we assume that we are given images together with corresponding camera intrinsics, extrinsics, and object masks . As our experiments show, our method works with as little as one image per object. In addition, our method can also incorporate depth information , if available.
For training and , we randomly sample an image and points on the image plane. We distinguish the following three cases: First, let denote the set of points that lie inside the object mask and for which the occupancy network predicts a finite surface depth as described in Section 3.3. For these points we can define a loss directly on the predicted image . Moreover, let denote the points which lie outside the object mask . While we cannot define a photometric loss for these points, we can define a loss that encourages the network to remove spurious geometry along corresponding rays. Finally, let denote the set of points which lie inside the object mask , but for which the occupancy network does not predict a finite surface depth . Again, we cannot use a photometric loss for these points, but we can define a loss that encourages the network to produce a finite surface depth.
RGB Loss: For each point in , we detect the predicted surface depth as described in Section 3.3. We define a photo-consistency loss for the points as
where indicates the ground truth depth value of the sampled image point and denotes the predicted surface depth for pixel .
Freespace Loss: If a point lies outside the object mask but the predicted surface depth is finite, the network falsely predicts surface point . Therefore, we penalize this occupancy with
where BCE is the binary cross entropy. When no surface depth is predicted, we apply the freespace loss to a randomly sampled point on the ray.
Occupancy Loss: If a point lies inside the object mask but the predicted surface depth is infinite, the network falsely predicts no surface points on ray . To encourage predicting occupied space on this ray, we uniformly sample depth values and define
In the single-view reconstruction experiments, we instead use the first point on the ray which lies inside all object masks (depth of the visual hull). If we have additional depth supervision, we use the ground truth depth for the occupancy loss. Intuitively, encourages the network to occupy space along the respective rays which can then be used by in (12) and in (13) to refine the initial occupancy.
Normal Loss: Optionally, our representation allows us to incorporate a smoothness prior by regularizing surface normals. This is useful especially for real-world data as training with 2D or 2.5D supervision includes unconstrained areas where this prior enforces more natural shapes. We define this loss as
where denotes the normal vector, the predicted surface point and a randomly sampled neighbor of .See supplementary for details.
5 Implementation Details
We implement the combined network with fully-connected ResNet blocks and ReLU activation. The output dimension of the last layer is , one dimension for the occupancy probability and three dimensions for the texture. For the single-view reconstruction experiments, we encode the input image with an ResNet-18 encoder network which outputs a -dimensional latent code z. To facilitate training, we start with a ray sampling accuracy of which we iteratively increase to by doubling after , , and thousand iterations. We choose the sampling interval such that it covers the volume of interest for each object. We set for all experiments. We train on a single NVIDIA V100 GPU with a batch size of images with random pixels each. We use the Adam optimizer with learning rate which we decrease by a factor of after and epochs, respectively.
Experiments
We conduct two different types of experiments to validate our approach. First, we investigate how well our approach reconstructs 3D shape and texture from a single RGB image when trained on a large collection of RGB or RGB-D images. Here, we consider both the case where we have access to multi-view supervision and the case where we use only a single RGB-D image per object during training. Next, we apply our approach to the challenging task of multi-view reconstruction, where the goal is to reconstruct complex 3D objects from real-world multi-view imagery.
First, we investigate to which degree our method can infer a 3D shape and texture representation from single-views. We train a single model jointly on all categories.
Datasets: To adhere to community standards , we use the Choy et al. subset ( classes) of the ShapeNet dataset for 2.5D and 3D supervised methods with training, validation, and test splits from . While we use the renderings from Choy et al. as input, we additionally render images of resolution with depth maps and object masks per object which we use for supervision. We randomly sample the viewpoint on the northern hemisphere as well as the distance of the camera to the object to get diverse supervision data. For 2D supervised methods, we adhere to community standards and use the renderings and splits from . Similar to , we train with objects in canonical pose.
Baselines: We compare against the following methods which all produce watertight meshes as output: 3D-R2N2 (voxel-based), Pixel2Mesh (mesh-based), and ONet (implicit representation). We further compare against both the 2D and the 2.5D supervised version of Differentiable Ray Consistency (DRC) (voxel-based) and the 2D supervised Soft Rasterizer (SoftRas) (mesh-based). For 3D-R2N2, we use the pre-trained model from which was shown to produce better results than the original model from . For the other baselines, we use the pre-trained modelsUnfortunately, we cannot show texture results for DRC and SoftRas as texture prediction is not part of the official code repositories. from the authors.
We first consider the case where we have access to multi-view supervision with images and corresponding object masks. In addition, we also investigate the case when ground truth depth maps are given.
Results: We evaluate the results using the Chamfer- distance from . In contrast to previous works , we compare directly wrt. to the ground truth shape models, not the voxelized or watertight versions.
In Table 1 and Fig. 4 we show quantitative and qualitative results for our method and various baselines. We can see that our method is able to infer accurate 3D shape and texture representations from single-view images when only trained on multi-view images and object masks as supervision signal. Quantitatively (Table 1), our method performs best among the approaches with 2D supervision and rivals the quality of methods with full 3D supervision. When trained with depth, our method performs comparably to the methods which use full 3D information. Qualitatively (Fig. 4), we see that in contrast to the mesh-based approaches, our method is not restricted to certain topologies. When trained with the photo-consistency loss , we see that our approach is able to predict accurate texture information in addition to the 3D shape.
1.2 Single-View Supervision
The previous experiment indicates that our model is able to infer accurate shape and texture information without 3D supervision. A natural question to ask is how many images are required during training. To this end, we investigate the case when only a single image with depth and camera information is available. Since we represent the 3D shape in a canonical object coordinate system, the hypothesis is that the model can aggregate the information over multiple training instances, although it sees every object only from one perspective. As the same image is used both as input and supervision signal, we now condition on our renderings instead of the ones provided by Choy et al. .
Results: Surprisingly, Fig. 5 shows that our method can infer appropriate 3D shape and texture when only a single-view is available per object, confirming our hypothesis. Quantitatively, the Chamfer distance of the model trained with and with only a single view () is comparable to the model trained with with views (). The reason for the numbers being worse than in Section 4.1 is that for our renderings, we do not only sample the viewpoint, but also the distance to the object resulting in a much harder task (see Fig. 5).
2 Multi-View Reconstruction
Finally, we investigate if our method is also applicable to multi-view reconstruction in real-world scenarios. We investigate two cases: First, when multi-view images and object masks are given. Second, when additional sparse depth maps are given which can be obtained from classic multi-view stereo algorithms . For this experiment, we do not condition our model and train one model per object.
Dataset: We conduct this experiment on scans , , and from the challenging real-world DTU dataset . The dataset contains or images with camera information for each object and baseline and structured light ground truth data. The presented objects are challenging as their appearance changes in different viewpoints due to specularities. Our sampling-based approach allows us to train on the full image resolution of . We label the object masks ourselves and always remove the same images with profound changes in lighting conditions, e.g., caused by the appearance of scanner parts in the background.
Baselines: We compare against classical approaches that have 3D meshes as output. To this end, we run screened Poisson surface reconstruction (sPSR) on the output of the classical MVS algorithms Campbell et al. , Furukawa et al. , Tola et al. , and Colmap . We find that the results on the DTU benchmark for the baselines are highly sensitive to the trim parameter of sPSR and therefore report results for the trim parameters (watertight output), (good qualitative results) and (good quantitative results). For a fair comparison, we use the object masks to remove all points which lie outside the visual hull from the predictions of the baselines before running sPSR.See supplementary material for details. We use the official DTU evaluation script in “surface mode”.
Conclusion and Future Work
In this work, we have presented Differentiable Volumetric Rendering (DVR). Observing that volumetric rendering is inherently differentiable for implicit representations allows us to formulate an analytic expression for the gradients of the depth with respect to the network parameters. Our experiments show that DVR enables us to learn implicit 3D shape representations from multi-view imagery without 3D supervision, rivaling models that are learned with full 3D supervision. Moreover, we found that our model can also be used for multi-view 3D reconstruction. We believe that DVR is a useful technique that broadens the scope of applications of implicit shape and texture representations.
In the future, we plan to investigate how to circumvent the need for object masks and camera information, e.g., by predicting soft masks and how to estimate not only texture but also more complex material properties.
Acknowledgments
This work was supported by an NVIDIA research gift. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Michael Niemeyer.