Differentiable Volumetric Rendering: Learning Implicit 3D Representations without 3D Supervision

Michael Niemeyer, Lars Mescheder, Michael Oechsle, Andreas Geiger

Introduction

In recent years, learning-based 3D reconstruction approaches have achieved impressive results . By using rich prior knowledge obtained during the training process, they are able to infer a 3D model from as little as a single image. However, most learning-based methods are restricted to synthetic data, mainly because they require accurate 3D ground truth models as supervision for training.

To overcome this barrier, recent works have investigated approaches that require only 2D supervision in the form of depth maps or multi-view images. Most existing approaches achieve this by modifying the rendering process to make it differentiable . While yielding compelling results, they are restricted to specific 3D representations (e.g. voxels or meshes) that suffer from discretization artifacts and the computational cost limits them to small resolutions or deforming a fixed template mesh. At the same time, implicit representations for shape and texture have been proposed which do not require discretization during training and have a constant memory footprint. However, existing approaches using implicit representations require 3D ground truth for training and it remains unclear how to learn implicit representations from image data alone.

Contribution: In this work, we introduce Differentiable Volumetric Rendering (DVR). Our key insight is that we can derive analytic gradients for the predicted depth map with respect to the network parameters of the implicit shape and texture representation (see Fig. 1). This insight enables us to design a differentiable renderer for implicit shape and texture representations and allows us to learn these representations solely from multi-view images and object masks. Since our method does not have to store volumetric data in the forward pass, its memory footprint is independent of the sampling accuracy of the depth prediction step. We show that our formulation can be used for various tasks such as single- and multi-view reconstruction, and works with synthetic and real data. In contrast to , we do not need to condition the texture representation on the geometry, but learn a single model with shared parameters that represents both geometry and texture. Our code and data are provided at https://github.com/autonomousvision/differentiable˙volumetric˙rendering.

Related Work

3D Representations: Learning-based 3D reconstruction approaches can be categorized wrt. the representation they use as voxel-based , point-based , mesh-based, or implicit representations .

Voxels can be easily processed by standard deep learning architectures, but even when operating on sparse data structures , they are limited to relatively small resolution. While point-based approaches are more memory-efficient, they require intensive post-processing because of missing connectivity information. Most mesh-based methods do not perform post-processing, but they often require a deformable template mesh or represent geometry as a collection of 3D patches which leads to self-intersections and non-watertight meshes.

To mitigate these problems, implicit representations have gained popularity . By describing 3D geometry and texture implicitly, e.g., as the decision boundary of a binary classifier , they do not discretize space and have a fixed memory footprint.

In this work, we show that the volumetric rendering step for implicit representations is inherently differentiable. In contrast to previous works, this allows us to learn implicit 3D shape and texture representations using 2D supervision.

3D Reconstruction: Recovering 3D information which is lost during the image capturing process is one of the long-standing goals of computer vision . Classic multi-view stereo (MVS) methods usually match features between neighboring views or reconstruct the 3D shape in a voxel grid . While the former methods produce depth maps as output which have to be fused in a lossy post-processing step, e.g., using volumetric fusion , the latter approaches are limited by the excessive memory requirements of 3D voxel grids. In contrast to these highly engineered approaches, our generic method directly outputs a consistent representation in 3D space which can be easily converted into a watertight mesh while having a constant memory footprint.

Recently, learning-based approaches have been proposed that either learn to match image features , refine or fuse depth maps , optimize parts of the classical MVS pipeline , or replace the entire MVS pipeline with neural networks that are trained end-to-end . In contrast to these learning-based approaches, our method can be supervised from 2D images alone and outputs a consistent 3D representation.

Differentiable Rendering: We focus on methods that learn 3D geometry via differentiable rendering in contrast to recent neural rendering approaches which synthesize high-quality novel views but do not infer the 3D object. They can again be categorized by the underlying representation of 3D geometry that they use.

Loper et al. propose OpenDR which approximates the backward pass of the traditional mesh-based graphics pipeline and has inspired several follow-up works . Liu et al. replace the rasterization step with a soft version to make it differentiable. While yielding compelling results in reconstruction tasks, these approaches require a deformable template mesh for training, restricting the topology of the output.

Another line of work operates on voxel grids . Paschalidou et al. and Tulsiani et al. propose a probabilistic ray potential formulation. While providing a solid mathematical framework, all intermediate evaluations need to be saved for backpropagation, restricting these approaches to relatively small-resolution voxel grids.

Liu et al. propose to infer implicit representations from multi-view silhouettes by performing max-pooling over the intersections of rays with a sparse number of supporting regions. In contrast, we use texture information enabling us to improve over the visual hull and to reconstruct concave shapes. Sitzmann et al. infer implicit scene representations from RGB images via an LSTM-based differentiable renderer. While producing high-quality renderings, the geometry cannot be extracted directly and intermediate results need to be stored for computing gradients. In contrast, we show that volumetric rendering is inherently differentiable for implicit representations. Thus, no intermediate results need to be saved for the backward pass.

Method

In this section, we describe our Differentiable Volumetric Rendering (DVR) approach. We first define the implicit neural representation which we use for representing 3D shape and texture. Next, we provide a formal description of DVR and all relevant implementation details. An overview of our approach is provided in Fig. 2.

Shape: In contrast to discrete voxel- and point-based representations, we represent the 3D shape of an object implicitly using the occupancy network introduced in :

Texture: Similarly, we can describe the texture of a 3D object using a texture field

Supervision: Recent works have shown that it is possible to learn fθf_{\theta} and tθ\mathbf{t}_{\theta} with 3D supervision (i.e., ground truth 3D models). However, ground truth 3D data is often very expensive or even impossible to obtain for real-world datasets. In the next section, we introduce DVR, an alternative approach that enables us to learn both fθf_{\theta} and tθ\mathbf{t}_{\theta} from 2D images alone. For clarity, we drop the condition variable z\mathbf{z} in the following.

2 Differentiable Volumetric Rendering

Our goal is to learn fθf_{\theta} and tθ\mathbf{t}_{\theta} from 2D image observations. Consider a single image observation. We define a photometric reconstruction loss

Gradients: To obtain gradients of L\mathcal{L} with respect to θ\theta, we first use the multivariate chain rule:

Here, ∂g∂x\frac{\partial\mathbf{g}}{\partial\mathbf{x}} denotes the Jacobian matrix for a vector-valued function g\mathbf{g} with vector-valued argument x\mathbf{x} and ⋅\cdot indicates matrix multiplication. By exploiting I^u=tθ(p^)\mathbf{\hat{I}}_{\mathbf{u}}=\mathbf{t}_{\theta}(\mathbf{\hat{p}}), we obtain

since both tθ\mathbf{t}_{\theta} as well as p^\hat{\mathbf{p}} depend on θ\theta. Because p^\mathbf{\hat{p}} is defined implicitly, calculating ∂p^∂θ\tfrac{\partial\hat{\mathbf{p}}}{\partial\theta} is non-trivial. We first exploit that p^\mathbf{\hat{p}} lies on the ray from r0\mathbf{r}_{0} through u\mathbf{u}. For any pixel u\mathbf{u}, this ray can be described by r(d)=r0+dw\mathbf{r}(d)=\mathbf{r}_{0}+d\mathbf{w} where w\mathbf{w} is the vector connecting r0\mathbf{r}_{0} and u\mathbf{u} (see Fig. 3). Since p^\mathbf{\hat{p}} must lie on r\mathbf{r}, there exists a depth value d^\hat{d}, such that p^=r(d^)\mathbf{\hat{p}}=\mathbf{r}(\hat{d}). We call d^\hat{d} the surface depth. This enables us to rewrite ∂p^∂θ\tfrac{\partial\hat{\mathbf{p}}}{\partial\theta} as

For computing the gradient of the surface depth d^\hat{d} with respect to θ\theta we exploit implicit differentiation . Differentiating fθ(p^)=τf_{\theta}(\hat{\mathbf{p}})=\tau on both sides wrt. θ\theta, we obtain:

Rearranging (7), we arrive at the following closed form expression for the gradient of the surface depth d^\hat{d}:

We remark that calculating the gradient of the surface depth d^\hat{d} wrt. the network parameters θ\theta only involves calculating the gradient of fθf_{\theta} at p^\hat{\mathbf{p}} wrt. the network parameters θ\theta and the surface point p^\hat{\mathbf{p}}. Thus, in contrast to voxel-based approaches , we do not have to store intermediate results (e.g., volumetric data) for computing the gradient of the loss wrt. the parameters, resulting in a memory-efficient algorithm. In the next section, we describe our implementation of DVR which makes use of reverse-mode automatic differentiation to compute the full gradient (4).

3 Implementation

To use automatic differentiation, we have to implement the forward and backward pass for the surface depth prediction step θ→d^\theta\to\hat{d}. In the following, we describe how both passes are implemented. For more details, we refer the reader to the supplementary material.

Forward Pass: As visualized in Fig. 3, we can determine d^\hat{d} by finding the first occupancy change on the ray r\mathbf{r}. To detect an occupancy change, we evaluate the occupancy network fθ(⋅)f_{\theta}(\cdot) at nn equally-spaced samples on the ray {pjray}j=1n\{\mathbf{p}_{j}^{\text{ray}}\}_{j=1}^{n}. Using a step size of Δs\Delta s, we can express the coordinates of these point in world-coordinates as

where s0s_{0} determines the closest possible surface point. We first find the smallest jj for which fθf_{\theta} changes from free space (fθ<τf_{\theta}<\tau) to occupied space (fθ≥τf_{\theta}\geq\tau):

We obtain an approximation to the surface depth d^\hat{d} by applying the iterative secant method to the interval [jΔs+s0,(j+1)Δs+s0][j\Delta s+s_{0},(j+1)\Delta s+s_{0}]. In practice, we compute the surface depth for a batch of NpN_{p} points in parallel. It is important to note that we do not need to unroll the forward pass or store any intermediate results as we exploit implicit differentiation to directly obtain the gradient of d^\hat{d} wrt. θ\theta.

Backward Pass: The input to the backward pass is the gradient λ=∂L∂d^\lambda=\tfrac{\partial\mathcal{L}}{\partial\hat{d}} of the loss wrt. a single surface depth prediction. The output of the backward pass is λ∂d^∂θ\lambda\tfrac{\partial\hat{d}}{\partial\theta} , which can be computed using (8). In practice, however, we would like to implement the backward pass not only for a single surface depth d^\hat{d} but for a whole batch of depth values.

We can implement this efficiently by rewriting λ∂d^∂θ\lambda\tfrac{\partial\hat{d}}{\partial\theta} as

Importantly, the left term in (11) corresponds to a normal backward operation applied to the neural network fθf_{\theta} and the right term in (11) is just an (element-wise) scalar multiplication for all elements in the batch. We can hence conveniently compute the backward pass of the operator θ→d^\theta\to\hat{d} by first multiplying the incoming gradient λ\lambda element-wise with a factor and then backpropagating the result through the operator θ→fθ(p^)\theta\to f_{\theta}(\mathbf{\hat{p}}). Both operations can be efficiently parallelized in common deep learning frameworks.

4 Training

During training, we assume that we are given NN images {Ik}k=1N\{\mathbf{I}_{k}\}_{k=1}^{N} together with corresponding camera intrinsics, extrinsics, and object masks {Mk}k=1N\{\mathbf{M}_{k}\}_{k=1}^{N}. As our experiments show, our method works with as little as one image per object. In addition, our method can also incorporate depth information {Dk}k=1N\{\mathbf{D}_{k}\}_{k=1}^{N}, if available.

For training fθf_{\theta} and tθ\mathbf{t}_{\theta}, we randomly sample an image Ik\mathbf{I}_{k} and NpN_{p} points u\mathbf{u} on the image plane. We distinguish the following three cases: First, let P0\mathcal{P}_{0} denote the set of points u\mathbf{u} that lie inside the object mask Mk\mathbf{M}_{k} and for which the occupancy network predicts a finite surface depth d^\hat{d} as described in Section 3.3. For these points we can define a loss Lrgb(θ)\mathcal{L}_{\text{rgb}}(\theta) directly on the predicted image I^k\mathbf{\hat{I}}_{k}. Moreover, let P1\mathcal{P}_{1} denote the points u\mathbf{u} which lie outside the object mask Mk\mathbf{M}_{k}. While we cannot define a photometric loss for these points, we can define a loss Lfreespace(θ)\mathcal{L}_{\text{freespace}}(\theta) that encourages the network to remove spurious geometry along corresponding rays. Finally, let P2\mathcal{P}_{2} denote the set of points u\mathbf{u} which lie inside the object mask Mk\mathbf{M}_{k}, but for which the occupancy network does not predict a finite surface depth d^\hat{d}. Again, we cannot use a photometric loss for these points, but we can define a loss Loccupancy(θ)\mathcal{L}_{\text{occupancy}}(\theta) that encourages the network to produce a finite surface depth.

RGB Loss: For each point in P0\mathcal{P}_{0}, we detect the predicted surface depth d^\hat{d} as described in Section 3.3. We define a photo-consistency loss for the points as

where dd indicates the ground truth depth value of the sampled image point u\mathbf{u} and d^\hat{d} denotes the predicted surface depth for pixel u\mathbf{u}.

Freespace Loss: If a point u\mathbf{u} lies outside the object mask but the predicted surface depth d^\hat{d} is finite, the network falsely predicts surface point p^=r(d^)\hat{\mathbf{p}}=\mathbf{r}(\hat{d}). Therefore, we penalize this occupancy with

where BCE is the binary cross entropy. When no surface depth is predicted, we apply the freespace loss to a randomly sampled point on the ray.

Occupancy Loss: If a point u\mathbf{u} lies inside the object mask but the predicted surface depth d^\hat{d} is infinite, the network falsely predicts no surface points on ray r\mathbf{r}. To encourage predicting occupied space on this ray, we uniformly sample depth values drandomd_{\text{random}} and define

In the single-view reconstruction experiments, we instead use the first point on the ray which lies inside all object masks (depth of the visual hull). If we have additional depth supervision, we use the ground truth depth for the occupancy loss. Intuitively, Loccupancy\mathcal{L}_{\text{occupancy}} encourages the network to occupy space along the respective rays which can then be used by Lrgb\mathcal{L}_{\text{rgb}} in (12) and Ldepth\mathcal{L}_{\text{depth}} in (13) to refine the initial occupancy.

Normal Loss: Optionally, our representation allows us to incorporate a smoothness prior by regularizing surface normals. This is useful especially for real-world data as training with 2D or 2.5D supervision includes unconstrained areas where this prior enforces more natural shapes. We define this loss as

where n(⋅)\mathbf{n}(\cdot) denotes the normal vector, pu^\hat{\mathbf{p}_{\mathbf{u}}} the predicted surface point and qu\mathbf{q}_{\mathbf{u}} a randomly sampled neighbor of pu^\hat{\mathbf{p}_{\mathbf{u}}}.See supplementary for details.

5 Implementation Details

We implement the combined network with 55 fully-connected ResNet blocks and ReLU activation. The output dimension of the last layer is 44, one dimension for the occupancy probability and three dimensions for the texture. For the single-view reconstruction experiments, we encode the input image with an ResNet-18 encoder network gϕ\mathbf{g}_{\phi} which outputs a 256256-dimensional latent code z. To facilitate training, we start with a ray sampling accuracy of n=16n=16 which we iteratively increase to n=128n=128 by doubling nn after 5050, 150150, and 250250 thousand iterations. We choose the sampling interval [s0,nΔs+s0][s_{0},n\Delta s+s_{0}] such that it covers the volume of interest for each object. We set τ=0.5\tau=0.5 for all experiments. We train on a single NVIDIA V100 GPU with a batch size of 6464 images with 10241024 random pixels each. We use the Adam optimizer with learning rate γ=10−4\gamma=10^{-4} which we decrease by a factor of 55 after 750750 and 10001000 epochs, respectively.

Experiments

We conduct two different types of experiments to validate our approach. First, we investigate how well our approach reconstructs 3D shape and texture from a single RGB image when trained on a large collection of RGB or RGB-D images. Here, we consider both the case where we have access to multi-view supervision and the case where we use only a single RGB-D image per object during training. Next, we apply our approach to the challenging task of multi-view reconstruction, where the goal is to reconstruct complex 3D objects from real-world multi-view imagery.

First, we investigate to which degree our method can infer a 3D shape and texture representation from single-views. We train a single model jointly on all categories.

Datasets: To adhere to community standards , we use the Choy et al. subset (1313 classes) of the ShapeNet dataset for 2.5D and 3D supervised methods with training, validation, and test splits from . While we use the renderings from Choy et al. as input, we additionally render 2424 images of resolution 2562256^{2} with depth maps and object masks per object which we use for supervision. We randomly sample the viewpoint on the northern hemisphere as well as the distance of the camera to the object to get diverse supervision data. For 2D supervised methods, we adhere to community standards and use the renderings and splits from . Similar to , we train with objects in canonical pose.

Baselines: We compare against the following methods which all produce watertight meshes as output: 3D-R2N2 (voxel-based), Pixel2Mesh (mesh-based), and ONet (implicit representation). We further compare against both the 2D and the 2.5D supervised version of Differentiable Ray Consistency (DRC) (voxel-based) and the 2D supervised Soft Rasterizer (SoftRas) (mesh-based). For 3D-R2N2, we use the pre-trained model from which was shown to produce better results than the original model from . For the other baselines, we use the pre-trained modelsUnfortunately, we cannot show texture results for DRC and SoftRas as texture prediction is not part of the official code repositories. from the authors.

We first consider the case where we have access to multi-view supervision with N=24N=24 images and corresponding object masks. In addition, we also investigate the case when ground truth depth maps are given.

Results: We evaluate the results using the Chamfer-L1L_{1} distance from . In contrast to previous works , we compare directly wrt. to the ground truth shape models, not the voxelized or watertight versions.

In Table 1 and Fig. 4 we show quantitative and qualitative results for our method and various baselines. We can see that our method is able to infer accurate 3D shape and texture representations from single-view images when only trained on multi-view images and object masks as supervision signal. Quantitatively (Table 1), our method performs best among the approaches with 2D supervision and rivals the quality of methods with full 3D supervision. When trained with depth, our method performs comparably to the methods which use full 3D information. Qualitatively (Fig. 4), we see that in contrast to the mesh-based approaches, our method is not restricted to certain topologies. When trained with the photo-consistency loss LRGB\mathcal{L}_{\text{RGB}}, we see that our approach is able to predict accurate texture information in addition to the 3D shape.

1.2 Single-View Supervision

The previous experiment indicates that our model is able to infer accurate shape and texture information without 3D supervision. A natural question to ask is how many images are required during training. To this end, we investigate the case when only a single image with depth and camera information is available. Since we represent the 3D shape in a canonical object coordinate system, the hypothesis is that the model can aggregate the information over multiple training instances, although it sees every object only from one perspective. As the same image is used both as input and supervision signal, we now condition on our renderings instead of the ones provided by Choy et al. .

Results: Surprisingly, Fig. 5 shows that our method can infer appropriate 3D shape and texture when only a single-view is available per object, confirming our hypothesis. Quantitatively, the Chamfer distance of the model trained with LRGB\mathcal{L}_{\text{RGB}} and LDepth\mathcal{L}_{\text{Depth}} with only a single view (0.4100.410) is comparable to the model trained with LDepth\mathcal{L}_{\text{Depth}} with 2424 views (0.3830.383). The reason for the numbers being worse than in Section 4.1 is that for our renderings, we do not only sample the viewpoint, but also the distance to the object resulting in a much harder task (see Fig. 5).

2 Multi-View Reconstruction

Finally, we investigate if our method is also applicable to multi-view reconstruction in real-world scenarios. We investigate two cases: First, when multi-view images and object masks are given. Second, when additional sparse depth maps are given which can be obtained from classic multi-view stereo algorithms . For this experiment, we do not condition our model and train one model per object.

Dataset: We conduct this experiment on scans 6565, 106106, and 118118 from the challenging real-world DTU dataset . The dataset contains 4949 or 6565 images with camera information for each object and baseline and structured light ground truth data. The presented objects are challenging as their appearance changes in different viewpoints due to specularities. Our sampling-based approach allows us to train on the full image resolution of 1200×16001200\times 1600. We label the object masks ourselves and always remove the same images with profound changes in lighting conditions, e.g., caused by the appearance of scanner parts in the background.

Baselines: We compare against classical approaches that have 3D meshes as output. To this end, we run screened Poisson surface reconstruction (sPSR) on the output of the classical MVS algorithms Campbell et al. , Furukawa et al. , Tola et al. , and Colmap . We find that the results on the DTU benchmark for the baselines are highly sensitive to the trim parameter of sPSR and therefore report results for the trim parameters (watertight output), 55 (good qualitative results) and 77 (good quantitative results). For a fair comparison, we use the object masks to remove all points which lie outside the visual hull from the predictions of the baselines before running sPSR.See supplementary material for details. We use the official DTU evaluation script in “surface mode”.

Conclusion and Future Work

In this work, we have presented Differentiable Volumetric Rendering (DVR). Observing that volumetric rendering is inherently differentiable for implicit representations allows us to formulate an analytic expression for the gradients of the depth with respect to the network parameters. Our experiments show that DVR enables us to learn implicit 3D shape representations from multi-view imagery without 3D supervision, rivaling models that are learned with full 3D supervision. Moreover, we found that our model can also be used for multi-view 3D reconstruction. We believe that DVR is a useful technique that broadens the scope of applications of implicit shape and texture representations.

In the future, we plan to investigate how to circumvent the need for object masks and camera information, e.g., by predicting soft masks and how to estimate not only texture but also more complex material properties.

Acknowledgments

This work was supported by an NVIDIA research gift. The authors thank the International Max Planck Research School for Intelligent Systems (IMPRS-IS) for supporting Michael Niemeyer.

References