DeepVoxels: Learning Persistent 3D Feature Embeddings
Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, Michael Zollhöfer
Introduction
Recent years have seen significant progress in applying generative machine learning methods to the creation of synthetic imagery. Many deep neural networks, for example based on (variational) autoencoders, are able to inpaint, refine, or even generate complete images from scratch HintoS2006; jKingma2014. A very prominent direction is generative adversarial networks GoodfPMXWOCB2014 which achieve impressive results for image generation, even at high resolutions KarraALL2018 or conditional generative tasks IsolaZZE2017. These developments allow us to perform highly-realistic image synthesis in a variety of settings; e.g., purely generative, conditional, etc.
However, while each generated image is of high quality, a major challenge is to generate a series of coherent views of the same scene. Such consistent view generation would require the network to have a latent space representation that fundamentally understands the 3D layout of the scene; e.g., how would the same chair look from a different viewpoint? Unfortunately, this is challenging to learn for existing generative neural network architectures that are based on a series of 2D convolution kernels. Here, spatial layout and transformations of a real, 3D environment would require a tedious learning process which maps 3D operations into 2D convolution kernels Jaderberg2015. In addition, the generator network in these approaches is commonly based on a U-Net architecture with skip connections RonneFB2015. Although skip connections enable efficient propagation of low-level features, the learned 2D-to-2D mappings typically struggle to generalize to large 3D transformations, due to the fact that the skip connections bypass higher-level reasoning.
To tackle similar challenges in the context of learning-based 3D reconstruction and semantic scene understanding, the field of 3D deep learning has seen large and rapid progress over the last few years. Existing approaches are able to predict surface geometry with high accuracy. Many of these techniques are based on explicit 3D representations in the form of occupancy grids Maturana2015; qi2016volumetric, signed distance fields Riegler2017OctNet, point clouds qi2017pointnet; lin2018learning, or meshes Jack2018. While these approaches handle the geometric reconstruction task well, they are not directly applicable to the synthesis of realistic imagery, since it is unclear how to represent color information at a sufficiently high resolution. There also exists a large body of work on learning low-dimensional embeddings of images that can be decoded to novel views tatarchenko2015single; yan2016perspective; dosovitskiy2017learning; eslami2018neural; worrall2017interpretable; rhodin2018unsupervised. Some of these techniques make use of the object’s 3D rotation by explicitly rotating the latent space feature vector worrall2017interpretable; rhodin2018unsupervised. While such 3D techniques are promising, they have thus far not been successful in achieving sufficiently high fidelity for the task of photorealistic image synthesis.
In our work, we aim at overcoming the fundamental limitations of existing 2D generative models by introducing native 3D operations in the neural network architecture. Rather than learning intuitive concepts from 3D vision, such as perspective, we explicitly encode these operations in the network architecture and perform reasoning directly in 3D space. The goal of the DeepVoxels approach is to condense posed input images of a scene into a persistent latent representation without explicitly having to model its geometry (see Fig. 1). This representation can then be applied to the task of novel view synthesis to generate unseen perspectives of a 3D scene without requiring access to the initial set of input images. Our approach is a hybrid 2D/3D one in that it learns to represent a scene in a Cartesian 3D grid of persistent feature embeddings that is projected to the target view’s canonical view volume and processed by a 2D rendering network. This persistent feature volume, which exists in 3D world-space, in combination with a structured, differentiable image formation model, enforces perspective and multi-view geometry in a principled and interpretable manner during training. The proposed approach learns to exploit the underlying 3D scene structure, without requiring supervision in the 3D domain. We demonstrate novel view synthesis with high quality for a variety of scenes based on this new representation. In summary, our approach makes the following technical contributions:
A novel persistent 3D feature representation for image synthesis that makes use of the underlying 3D scene information.
Explicit occlusion reasoning based on learned soft visibility that leads to higher-quality results and better generalization to novel viewpoints.
Differentiable image formation to enforce perspective and multi-view geometry in a principled and interpretable manner during training.
Training without requiring 3D supervision.
In this paper, we present first steps towards 3D-structured neural scene representations. To this end, we limit the scope of our investigation to allow an in-depth discussion of the challenges fundamental to this approach. We assume Lambertian scenes, without specular highlights or other view-dependent effects. While the proposed approach can deal with light specularities, these are not modeled explicitly. Classical approaches will achieve impressive results on the presented scenes. However, these approaches rely on the explicit reconstruction of geometry. Neural scene representations will be essential to develop generative models that can generalize across scenes to solve reconstruction problems where only few observations are available. We thus compare to such baselines exclusively.
Related Work
Our approach lies at the intersection of multiple active research areas, namely generative neural networks, 3D deep learning, deep learning-based view synthesis, and model- as well as image-based rendering.
Deep models for 2D image and video synthesis have recently shown very promising results. Some of these approaches are based on (variational) auto-encoders (VAEs) HintoS2006; jKingma2014 or autoregressive models (AMs), such as PixelCNN Oord:2016. The most promising results so far are based on conditional generative adversarial networks (cGANs) GoodfPMXWOCB2014; RadfoMC2016; MirzaO2014; IsolaZZE2017. In most cases, the generator network has an encoder-decoder architecture HintoS2006, often with skip connections (U-Net) RonneFB2015, which enable efficient propagation of low-level features from the encoder to the decoder. Approaches that convert synthetic images into photo-realistic imagery have been proposed for the special case of human bodies Zhu2018; Chan2018 and faces kim2018DeepVideo. In theory, similar architectures could be used to regress the real-world image corresponding to a given viewpoint, i.e., image-based rendering could be learned from scratch. Unfortunately, these 2D-to-2D translation approaches struggle to generalize to transformations in 3D space, such as rotation and perspective projection, since the underlying 3D scene structure cannot be exploited. We compare to this baseline in Sec. 4 and show that DeepVoxels drastically outperforms it.
D Deep Learning
Recently, deep learning has been successfully applied to many 3D geometric reasoning tasks. Current approaches are able to predict an accurate 3D representation of an object from just a single or multiple views. Many of these techniques make use of classical 3D representations, e.g., occupancy grids Maturana2015; qi2016volumetric, signed distance fields Riegler2017OctNet, 3D point clouds qi2017pointnet; lin2018learning, or meshes Jack2018. While these approaches handle the geometric reconstruction task well, they are not directly applicable to view synthesis, since it is unclear how to represent color information at a sufficiently high resolution. View consistency can be explicitly handled using differentiable ray casting drcTulsiani17. RenderNet RenderNet learns to render in different styles from 3D voxel grid input. Kulkarni et al. Kulkarni2015 learn a disentangled representation of images with respect to various scene properties, such as rotation and illumination. Spatial Transformer Networks Jaderberg2015 can learn spatial transformations of feature maps in the network. Even weakly-supervised yang2015weakly and unsupervised NIPS2016_6600 learning of 3D transformations has been proposed. Our work is also related to CNNs for 3D reconstruction kar2017learning; choy20163d and monocular depth estimation eigen2014depth. A “multi-view stereo machine” kar2017learning can learn 3D reconstruction based on 3D or 2.5D supervision. MapNet henriques2018mapnet performs SLAM based on a scene-specific 2D feature grid representation. In contrast to these approaches, which are focused on geometric reasoning, our goal is to learn an embedding for novel view synthesis. To synthesize multi-view consistent images, we optimize for a persistent, scene-specific 3D embedding over all available 2D observations and enable the network to perform explicit occlusion reasoning. We do not require any 3D ground truth but minimize a 2D photometric reprojection loss exclusively.
Deep Learning for View Synthesis
Recently, a class of deep neural networks has been proposed that directly aim to solve the problem of novel view synthesis. Some techniques predict lookup tables into a set of reference views Park17; zhou2016view or predict weights to blend multi-view images into novel views flynn2016deepstereo. A layered scene representation lsiTulsiani18 can be learned based on a re-rendering loss. A large corpus of work focuses on embedding 2D views of scenes into a learned low-dimensional latent space that is then decoded into a novel view tatarchenko2015single; yan2016perspective; dosovitskiy2017learning; eslami2018neural; worrall2017interpretable; rhodin2018unsupervised; cohen2014transformation. Some of these approaches rely on embedding views into a latent space that does not enforce any geometrical constraints tatarchenko2015single; dosovitskiy2017learning; eslami2018neural, others enforce geometric constraints in varying degrees worrall2017interpretable; rhodin2018unsupervised; cohen2014transformation; falorsi2018explorations, such as learning rotation-equivariant features by explicitly rotating the latent space feature vectors. We focus on optimizing a scene-specific embedding over a training corpus of 2D observations and explicitly account for concepts from 3D vision such as perspective projection and occlusion to constrain the latent space. We demonstrate advantages over weakly structured embeddings in generating high-quality novel views.
Model-Based Rendering
Classic reconstruction approaches such as structure-from-motion exploit multi-view geometry Hartley:2003; szeliski2010computer to build a dense 3D point cloud of the imaged scene schoenberger2016sfm; schoenberger2016mvs; snavely2006photo; agarwal2009building; furukawa2010accurate. A triangular surface representation can be obtained using for example the Poisson Surface Kazhdan:2006 reconstruction technique. However, the reconstructed geometry is often imperfect, coarse, contains holes, and the resulting renderings thus suffer from visible artifacts and are not fully realistic. In contrast, our goal is to learn a representation that efficiently encodes the view-dependent appearance of a 3D scene without having to explicitly reconstruct a geometric model.
Image-Based Rendering
Traditional image-based rendering techniques blend warped versions of the input images to generate new views shum2000review. This idea was first proposed as a computationally efficient alternative to classical rendering lippman1980movie; greene1986environment; chen1993view. Multiple-view geometry can be used to obtain the geometry for warping hedman2016scalable. In other cases, no 3D reconstruction is necessary flynn2016deepstereo; penner2017soft. Some approaches rely on light fields kalantari2016learning. Recently, deep-learning has been used to aid image-based rendering via learning a small subtask, i.e., the computation of the blending weights hedman2018deepblending; flynn2016deepstereo. While this can achieve photorealism, it depends on a dense set of high-resolution photographs to be available at rendering time and requires an error prone reconstruction step to obtain the geometric proxy. Our approach has orthogonal goals: (1) we want to learn an embedding for view synthesis and (2) we want to tackle the problem in a holistic fashion by learning raw pixel output. Thus, our approach is more related to embedding techniques that try to learn a latent space that can be decoded into novel views.
Method
The core of our approach is a novel 3D-structured scene representation called DeepVoxels. DeepVoxels is a viewpoint-invariant, persistent and uniform 3D voxel grid of features. The underlying 3D grid enforces spatial structure on the learned per-voxel code vectors. The final output image is formed based on a 2D network that receives the perspective re-sampled version of this 3D volume, i.e., the canonical view volume of the target view, as input. The 3D part of our approach takes care of spatial reasoning, while the 2D part enables fine-scale feature synthesis. In the following, we first introduce the training corpus and then present our end-to-end approach for finding the scene-specific DeepVoxels representation from a set of multi-view images without explicit 3D supervision.
Our scene-specific training corpus of samples is based on a source view (image and camera pose) and two target views , which are randomly selected from a set of registered multi-view images; see Fig. 1 for an example. We assume that the intrinsic and extrinsic camera parameters are available. These can for example be obtained using sparse bundle adjustment Triggs:1999. For each pair of target views we then randomly select a single source view from the top- nearest neighbors in terms of view direction angle to target view . This sampling heuristic makes it highly likely that points in the source view are visible in the target view . While not essential to training, this ensures meaningful gradient flow for every optimization step, while encouraging multi-view consistency to the random target view . We sample the training corpus dynamically during training.
2 Architecture Overview
Our network architecture is summarized in Fig. 2. On a high level, it can be seen as an encoder-decoder based architecture with the persistent 3D DeepVoxels representation as its latent space. During training, we feed a source view to the encoder and try to predict the target view . We first extract a set of 2D feature maps from the source view using a 2D feature extraction network. To learn a view-independent 3D feature representation, we explicitly lift image features to 3D based on a differentiable lifting layer. The lifted 3D feature volume is fused with our persistent DeepVoxels scene representation using a gated recurrent network architecture. Specifically, the persistent 3D feature volume is the hidden state of a gated recurrent unit (GRU) Cho2014. After feature fusion, the volume is processed by a 3D fully convolutional network. The volume is then mapped to the camera coordinate systems of the two target views via a differentiable reprojection layer, resulting in the canonical view volume. A dedicated, structured occlusion network operates on the canonical view volume to reason about voxel visibility and flattens the view volume to a 2D view feature map (see Fig. 3). Finally, a learned 2D rendering network forms the two final output images. Our network is trained end-to-end, without the need of supervision in the 3D domain, by a 2D re-rendering loss that enforces that the predictions match the target views. In the following, we provide more details.
Here, and specify the position of the voxel center on the screen and is its depth from the camera. Given a pixel and its depth, we can invert this mapping to compute the corresponding 3D point .
Feature Extraction
We extract 2D feature maps from the source view based on a fully convolutional feature extraction network. The image is first downsampled by a series of stride-2 convolutions until a resolution of is reached. A 2D U-Net architecture RFB15a then extracts a feature map that is the input to the subsequent volume lifting.
Lifting 2D Features to 3D Observations
The lifting layer lifts 2D features into a temporary 3D volume, representing a single 3D observation, which is then integrated into the persistent DeepVoxels representation. We position the 3D feature volume in world space such that its center roughly aligns with the scene’s center of gravity, which can be obtained cheaply from the keypoint point cloud obtained from sparse bundle adjustment. The spatial extent is set such that the complete scene is inside the volume. We try to bound the scene as tightly as possible to not lose spatial resolution. Lifting is implemented by a gathering operation. For each voxel, the world space position of its center is projected to the source view’s image space following Eq. 1. We extract a feature vector from the feature map using bilinear sampling and store the result in the code vector associated with the voxel. Note, our approach is based only on a set of registered multi-view images and we do not have access to the scene geometry or depth maps, rather our approach learns automatically to resolve the depth ambiguity based on a gated recurrent network in 3D.
Integrating Lifted Features into DeepVoxels
Lifted observations are integrated into the DeepVoxels representation via an integration network that is based on gated recurrent units (GRUs) Cho2014. In contrast to the standard application of GRUs, the integration network operates on the same volume across the full training procedure, i.e., the hidden state is persistent across all training steps and never reset, leading to a geometrically consistent representation of the whole training corpus. We use a uniform volumetric grid of size voxels, where each voxel has feature channels, i.e., the stored code vector has size . We employ one gated recurrent unit for each voxel, such that at each time step, all the features in a voxel have to be updated jointly. The goal of the gated recurrent units is to incrementally fuse the lifted features and the hidden state during training, such that the best persistent 3D volumetric feature representation is discovered. The gated recurrent units implement the mapping
Here, is the lifted 3D feature volume of the current timestep , the and are trainable 3D convolution weights, and the are trainable tensors of biases. We follow Cho et al. Cho2014 and employ a sigmoid activation to compute the response of the tensor of update gates and reset gates . Based on the previous hidden state , the per-voxel reset values , and the lifted 3D feature volume , the tensor of new feature proposals for the current time step is computed. and are single 3D convolutional layers. The new hidden state , the DeepVoxels representation for the current time step, is computed as a per-voxel linear combination of the old state and the new DeepVoxel proposal . The GRU performs one update step per lifted observation. Afterwards, we apply a 3D inpainting U-Net that learns to fill holes in this feature representation. At test time, only the optimally learned persistent 3D volumetric features, the DeepVoxels, are used to form the image corresponding to a novel target view. The 2D feature extraction, lifting layer and GRU gates are discarded and are not required for inference, see Fig. 2.
Projection Layer
The projection layer implements the inverse of the lifting layer, i.e., it maps the 3D code vectors to the canonical coordinate system of the target view, see Fig. 3 (left). Projection is also implemented based on a gathering operation. For each voxel of the canonical view volume, its corresponding position in the persistent world space voxel grid is computed. An interpolated code vector is then extracted via a trilinear interpolation and stored in the feature channels of the canonical view volume.
Occlusion Module
Occlusion reasoning is essential for correct image formation and generalization to novel viewpoints. To this end, we propose a dedicated occlusion network that computes soft visibility for each voxel. Each pixel in the target view is represented by one column of voxels in the canonical view volume, see Fig. 3 (left). First, this column is concatenated with a feature column encoding the distance of each voxel to the camera, similar as in liu2018intriguing. This allows the occlusion network to reason about voxel order. The feature vector of each voxel in this canonical view volume is then compressed to a low-dimensional feature vector of dimension 4 by a single 3D convolutional layer. This compressed volume is input to a 3D U-Net for occlusion reasoning. For each ray, represented by a single-pixel column, this network predicts a scalar per-voxel visibility weight based on a softmax activation, see Fig. 3 (middle). The canonical view volume is then flattened along the depth dimension with a weighted average, using the predicted visibility values. The softmax weights can further be used to compute a depth map, which provides insight into the occlusion reasoning of the network, see Fig. 3 (right).
Rendering and Loss
Analysis
In this section, we demonstrate that DeepVoxels is a rich and semantically meaningful 3D scene representation that allows high-quality re-rendering from novel views. First, we present qualitative and quantitative results on synthetic renderings of high-quality 3D scans of real-world objects, and compare the performance to strong machine-learning baselines with increasing reliance on geometrically structured latent spaces. Next, we demonstrate that DeepVoxels can also be used to generate novel views on a variety of real captures, even if these scenes may violate the Lambertian assumption. Finally, we demonstrate quantitative and qualitative benefits of explicitly reasoning about voxel visibility via the occlusion module, as well as improved model interpretability. Please see the supplement for further studies on the sensitivity to the number of training images, the size of the voxel volume, as well as noisy camera poses.
We evaluate model performance on synthetic data obtained from rendering high-quality 3D scans (see Fig. 4). We center each scan at the origin and scale it to lie within the unit cube. For the training set, we render the object from 479 poses uniformly distributed on the northern hemisphere. For the test set, we render 1000 views on an Archimedean spiral on the northern hemisphere. All images are rendered in a resolution of and then resized using area averaging to to minimize aliasing. We evaluate reconstruction error in terms of PSNR and SSIM wang2004image.
Implementation
Baselines
We compare to three strong baselines with increasing reliance on geometry-aware latent spaces. The first baseline is a Pix2Pix architecture IsolaZZE2017 that receives as input the per-pixel view direction, i.e., the normalized, world-space vector from camera origin to each pixel, and is trained to translate these images into the corresponding color image. This baseline is representative of recent achievements in 2D image-to-image translation. The second baseline is a deep autoencoder that receives as input one of the top- nearest neighbors of the target view, and the pose of both the target and the input view are concatenated in the deep latent space, as proposed by Tatarchenko et al. tatarchenko2015single. The inputs of this model at training time are thus identical to those of our model. The third baseline learns an interpretable, rotation-equivariant latent space via the method proposed in worrall2017interpretable; cohen2014transformation and used previously in rhodin2018unsupervised, by being fed one of the top- nearest neighbor views and then rotating the latent embedding with the rotation matrix that transforms the input to the output pose. At test time, the previous two baselines receive the top- nearest neighbor to supply the model with the most relevant information. We approximately match the number of parameters of each network, with all baselines having equally or slightly more parameters than our model. We train all baselines to convergence with the same loss function. For the exact baseline architectures and number of parameters, please see the supplement.
Object-specific Novel View Synthesis
We train our network and all baselines on synthetic renders of four high-quality 3D scans. Table 1 compares PSNR and SSIM of the proposed architecture and the baselines. The best-performing baseline is Pix2Pix IsolaZZE2017. This is surprising, since no geometrical constraints are enforced, as opposed to the approach by Worrall et al. worrall2017interpretable. The proposed architecture with strongly structured latent space outperforms all baselines by a wide margin of an average dB. Fig. 4 shows a qualitative comparison as well as further novel views sampled from the proposed model. The proposed model displays robust 3D reasoning that does not break down even in challenging cases. Notably, other models have a tendency to “snap” onto views seen in the training set, while the proposed model smoothly follows the test trajectory. Please see the supplemental video for a demonstration of this behavior. We hypothesize that this improved generalization to unseen views is due to the explicit multi-view constraints enforced by the proposed latent space. The baseline models are not explicitly enforcing projective and epipolar geometry, which may allow them to parameterize latent spaces that are not properly representing the low-dimensional manifold of rotations. Although the resolution of the proposed voxel grid is times smaller than the image resolution, our model succeeds in capturing fine detail much smaller than the size of a single voxel, such as the letters on the sides of the cube or the detail on the vase. This may be due to the use of trilinear interpolation in the lifting and projection steps, which allow for a fine-grained representation to be learned. Please see the video for full sequences, and the supplemental material for two additional synthetic scenes.
Voxel Embedding vs. Rotation-Equivariant Embedding
As reflected in Tab. 1, we outperform worrall2017interpretable by a wide margin both qualitatively and quantitatively. The proposed model is constrained through multi-view geometry, while worrall2017interpretable has more degrees of freedom. Lacking occlusion reasoning, depth maps are not made explicit. The model may thus parameterize latent spaces that do not respect multi-view geometry. This increases the risk of overfitting, which we observe empirically, as the baseline snaps to nearest neighbors seen during training. While the proposed voxel embedding is memory hungry, it is very parameter efficient. The use of 3D convolutions means that the parameter count is independent of the voxel grid size. Giving up spatial structure means Worrell et al. worrall2017interpretable abandon convolutions and use fully connected layers. However, to achieve the same latent space size of features would necessitate more than parameters between just the fully connected layers before and after the feature transformation layer, which is infeasible. In contrast, the proposed 3D inpainting network only has parameters, five orders of magnitude less. To address memory inefficiency, the dense grid may be replaced by a sparse alternative in the future.
Occlusion Reasoning and Interpretability
An essential part of the rendering pipeline is the depth test. Similarly, the rendering network ought to be able to reason about occlusions when regressing the output view. A naive approach might flatten the depth dimension of the canonical camera volume and subsequently reduce the number of features using a series of 2D convolutions. This leads to a drastic increase in the number of network parameters. At training time, this further allows the network to combine features from several depths equally to regress on pixel colors in the target view. At inference time, this results in severe artifacts and occluded parts of the object “shining through” (see Fig. 5). Our occlusion network forces learning to use a softmax-weighted sum of voxels along each ray, which penalizes combining voxels from several depths. As a result, novel views generated by the network with the occlusion module perform much more favorably at test time, as demonstrated in Fig. 5, than networks without the occlusion module. The depth map generated by the occlusion model further demonstrates that the proposed model indeed learns the 3D structure of the scene. We note that the depth map is learned in a fully unsupervised manner and arises out of the pure necessity of picking the most relevant voxel. Please see the supplement for more examples of learned depth maps.
Novel View Synthesis for Real Captures
We train our network on real captures obtained with a DSLR camera. Camera poses, intrinsic camera parameters and keypoint point clouds are obtained via sparse bundle adjustment. The voxel grid origin is set to the respective point cloud’s center of gravity. Voxel grid resolution is set to . Each voxel stores feature channels. Test trajectories are obtained by linearly interpolating two randomly chosen training poses. Scenes depict a drinking fountain, two busts, a globe, and a bag of coffee. See Fig. 6 for example model outputs. The drinking fountain and the globe have noticeable specularities, which are handled gracefully. While the coffee bag is generally represented faithfully, inconsistencies appear on its highly specular surface. Generally, results are of high quality, and only details that are significantly smaller than a single voxel, such as the tiles in the sink of the fountain, show artifacts. Please refer to the supplemental video for detailed results as well as a nearest-neighbor baseline.
Limitations
Although we have demonstrated high-quality view synthesis results for a variety of challenging scenes, the proposed approach still has limitations that can be tackled in the future. By construction, the employed 3D volume is memory inefficient, thus we have to trade local resolution for spatial extent. The proposed model can be trained with a voxel resolution of with feature channels, filling a GPU with GB of memory. Future work on sparse neural networks may replace the dense representation at the core. Please note, compelling results can already be achieved with quite small volume resolutions. Synthesizing images from viewpoints that are significantly different from the training set, i.e., generalization, is challenging for all learning-based approaches. While this is also true for DeepVoxels and detail is lost when viewing scenes from poses far away from training poses, DeepVoxels generally deteriorates gracefully and the 3D structure of the scene is preserved. Please refer to the supplemental material for failure cases as well as examples of pose extrapolation.
Conclusion
We have proposed a novel 3D-structured scene representation, called DeepVoxels, that encodes the view-dependent appearance of a 3D scene using only 2D supervision. Our approach is a first step towards 3D-structured neural scene representations and the goal of overcoming the fundamental limitations of existing 2D generative models by introducing native 3D operations into the network.
Acknowledgements: We thank Robert Konrad, Nitish Padmanaban, and Ludwig Schubert for fruitful discussions, and Robert Konrad for the video voiceover. Vincent Sitzmann was supported by a Stanford Graduate Fellowship. Michael Zollhöfer and Vincent Sitzmann were supported by the Max Planck Center for Visual Computing and Communication (MPC-VCC). Gordon Wetzstein was supported by a National Science Foundation CAREER award (IIS 1553333), by a Sloan Fellowship, and by an Okawa Research Grant. Matthias Nießner and Justus Thies were supported by a Google Research Grant, the ERC Starting Grant Scan2CAD (804724), a TUM-IAS Rudolf Mößbauer Fellowship (Focus Group Visual Computing), and a Google Faculty Award.