NeuralFusion: Online Depth Fusion in Latent Space
Silvan Weder, Johannes L. Schönberger, Marc Pollefeys, Martin R. Oswald
Introduction
Reconstructing the geometry of a scene is a central component of many applications in 3D computer vision. Awareness of the surrounding geometry enables robots to navigate, augmented and mixed reality devices to accurately project the information into the user’s field of view, and serves as the basis for many 3D scene understanding tasks.
In this paper, we consider online surface reconstruction by fusing a stream of depth maps with known camera calibration. The fundamental challenge in this task is that depth maps are typically noisy, incomplete, and contain outliers. Depth map fusion is a key component in many 3D reconstruction methods, like KinectFusion , VoxelHashing , InfiniTAM , and many others. The vast majority of methods builds upon the concept of averaging truncated signed distance functions (TSDFs), as proposed in the pioneering work by Curless and Levoy . This approach is so popular due to its simple, highly parallelizable, and real-time capable way of fusing noisy depth maps into a surface. However, it has difficulties with handling outliers and thin geometry, which can be mainly attributed to the local integration of depth values in the TSDF volume.
To tackle this fundamental limitation, existing methods use various heuristics to filter outliers in decoupled pre- or post-processing steps. Such filtering techniques entail the usual trade-off in terms of balancing accuracy against completeness. Especially in an online fusion system, striking this balance is extremely challenging in the pre-filtering stage, because it is difficult to distinguish between a first surface measurement and an outlier. Consequently, to achieve complete surface reconstructions, one must use conservative pre-filtering, which in turn requires careful post-filtering of outliers by non-local reasoning on the TSDF volume or the final mesh.
This motivates our key idea to use different scene representations for the fusion and final method output, while all prior works perform depth fusion and filtering directly in the output representation. In contrast, we perform the fusion step in a latent scene representation, implicitly learned to encode features like confidence information or local scene information. A final translation step simultaneously filters and decodes this learned scene representation into the final output relevant to downstream applications.
In summary, we make the following contributions:
We propose a novel, end-to-end trainable network architecture, which separates the scene representations for depth fusion and the final output into two different modules.
The proposed latent representation yields more accurate and complete fusion results for larger feature dimensions allowing to balance accuracy against resource demands.
Our network architecture allows for end-to-end learnable outlier filtering within a translation step that significantly improves outlier handling.
Although fully trainable, our approach still only performs very localized updates to the global map which maintains the online capability of the overall approach.
Related Work
Representation Learning for 3D Reconstruction. Many works proposed learning-based algorithms using implicit shape representations. Ladicky et al\onedot learn to regress a mesh from a point cloud using a random forest. The concurrent works OccupancyNetworks , DeepSDF , and IM-NET encode the object surface as the decision boundary of a trained classification or regression network. These methods only use a single feature vector to encode an entire scene or object within a unit cube and are, unlike our approach, not suitable for large-scale scenes. This was improved in which use multiple features to encode parts of the scene. extract image features and learn scene occupancy labels from single or multi-view images. DISN works in a similar fashion but regresses SDF values. Chiyu et al\onedot propose a local implicit grid representation for 3D scenes. However, they encode the scene through direct optimization of the latent representation through a pre-trained neural network. More recently, Scene Representation Networks learn to reconstruct novel views from a single RGB image. Liu et al\onedot learn implicit 3D shapes from 2D images in a self-supervised manner. These works also operate within a unit cube and are difficult to scale to larger scenes. DeepVoxels encodes visual information in a feature grid with a neural network to generate novel, high-quality views onto an observed object. Nevertheless, the work is not directly applicable to a 3D reconstruction task. Our method combines the concept of learned scene representations with data fusion in a learned latent space. Recently, several works proposed a more local scene representation that allows larger scale scenes and multiple objects. Further, learn implicit representations with 2D supervision via differentiable rendering. Overall, none of all mentioned works consider online updates of the shape representation as new information becomes available and adding such functionality is by no means straightforward.
Classic Online Depth Fusion Approaches. The majority of depth map fusion approaches are built upon the seminal “TSDF Fusion” work by Curless and Levoy , which fuses depth maps using an implicit representation by averaging TSDFs on a dense voxel grid. This approach especially became popular with the wide availability of low-cost depth sensors like the Kinect and has led to works like KinectFusion and related works like sparse-sequence fusion , BundleFusion , or variants on sparse grids like VoxelHashing , InfiniTAM , Voxgraph , octree-based approaches , or hierarchical hashing . However, the output of these methods usually contains typical noise artifacts, such as surface thickening and outlier blobs. Another line of works uses a surfel-based representation and directly fuses depth values into a sparse point cloud . This representation perfectly adapts to inhomogeneous sampling densities, requires low memory storage, but also lacks connectivity and topological surface information. An overview of depth fusion approaches is given in .
Classic Global Depth Fusion Approaches. While online approaches only process one depth map at a time, global approaches use all information at once and typically apply additional smoothness priors like total variation , its variants including semantic information , or refine surface details using color information . Consequently, their high compute and memory requirements prevent their application in online scenarios unlike ours.
Learned Global Depth Fusion Approaches. Octnet and its follow-up OctnetFusion fuse depth maps using TSDF fusion into an octree and then post-processes the fused geometry using machine learning. RayNet uses a learned Markov random field and a view-invariant feature representation to model view dependencies. SurfaceNet jointly estimates multi-view stereo depth maps and the fused geometry, but requires to store a volumetric grid for each input depth map. 3DMV optimizes shape and semantics of a pre-fused TSDF scene of given 2D view information. Contrary to these methods, we learn online depth fusion and can process an arbitrary number of input views.
Learned Online Depth Fusion Approaches. In the context of simultaneous localization and mapping, CodeSLAM , SceneCode and DeepFactors learn a 2.5D depth representation and its probabilistic fusion rather than fusing into a full 3D model. The DeepTAM mapping algorithm builds upon traditional cost volume computation with hand-crafted photoconsistency measures, which are fed into a neural network to estimate depth, but full 3D model fusion is not considered. DeFuSR refines depth maps by improving cross-view consistency via reprojection, but it is not real-time capable. Similar to our approach, Weder et al\onedot perform online reconstruction and learn the fusion updates and weights. In contrast to our work, all information is fused into an SDF representation which requires handcrafted pre- or post-filtering to handle outliers, which is not end-to-end trainable. The recent ATLAS method fuses features from RGB input into a voxel grid and then regresses a TSDF volume. While our method learns the fusion of features, they use simple weighted averaging. Their large ResNet50 backbone limits real-time capabilities.
Sequence to Vector Learning. On a high-level, our method processes a variable length sequence of depth maps and learns a 3D shape represented by a fixed length vector. It is thus loosely related to areas like video representation learning or sentiment analysis which processes text, audio or video to estimate a single vectorial value. In contrast to these works, we process 3D data and explicitly model spatial geometric relationships.
Method
Our feature fusion pipeline consists of four key stages depicted as networks in Figure 2. The first stage extracts the current state of the global feature volume into a local, view-aligned feature volume using an affine mapping defined by the given camera parameters. After the extraction, this local feature volume is passed together with the new depth measurement and the ray directions through a feature fusion network. This feature fusion network predicts optimal updates for the local feature volume, given the new measurement and its old state. The updates are integrated back into the global feature volume using the inverse affine mapping defined in the first stage. These three stages form the core of the fusion pipeline and are executed iteratively on the input depth map stream. An additional fourth stage translates the feature volume into an application-specific scene representation, such as a TSDF volume, from which one can finally render a mesh for visualization. We detail our pipeline in the following and we refer to the supplementary material for additional low-level architectural details.
Feature Extraction. The goal of iteratively fusing depth measurements is to (a) fuse information about previously unknown geometry, (b) increase the confidence about already fused geometry, and (c) to correct wrong or erroneous entries in the scene. Towards these goals, the fusion process takes the new measurements to update the previous scene state , which encodes all previously seen geometry. For a fast depth integration, we extract a local view-aligned feature subvolume with one ray per depth measurement centered at the measured depth via nearest neighbor search in the grid positions. Each ray of features in the local feature volume is concatenated with the ray direction and the new depth measurement. This feature volume is then passed to the fusion network.
Feature Fusion. The fusion network fuses the new depth measurements into the existing local feature representation . Therefore, we pass the feature volume through four convolutional blocks to encode neighborhood information from a larger receptive field. Each of these encoding blocks, consists of two convolutional layers with a kernel size of three. These layers are followed by layer normalization and tanh activation function. We found layer normalization to be crucial for training convergence. The output of each block is concatenated with its input, thereby generating a successively larger feature volume with increasing receptive field. The decoder then takes the output of the four encoding blocks to predict feature updates.The decoder consists of four blocks with two convolutional layers and interleaved layer normalization and tanh activation. The output of the final layer is passed through a single linear layer. Finally, the predicted feature updates are normalized and passed as to the feature integration.
Feature Integration. The updated feature state is integrated back into the global feature grid by using the inverse global-local grid correspondences of the extraction mapping. Similar to the extraction, we write the mapped features into the nearest neighbor grid location. Since this mapping is not unique, we aggregate colliding updates using an average pooling operation. Finally, the pooled features are combined with old ones using a per-voxel running average operation, where we use the update counts as weights. This residual update operation ensures stable training and a homogeneous latent space, as compared to direct prediction of the global features. Both the feature extraction and integration steps are inspired by , but they use tri-linear interpolation instead of nearest-neighbor sampling. When extracting and integrating features instead of SDF values, we empirically found that nearest-neighbor interpolation produces better results and leads to more stable convergence during training.
In the final and possibly asynchronous step, we translate the latent scene representation into a representation usable for visualization of the scene (e.g\onedot, signed distance field or occupancy grid). The network architecture in this step is inspired by IM-Net . For efficient and complete translation, we sample a regular grid of world coordinates. Then, at each of these sampled points , the translator aggregates the information stored in the features of the local neighborhood and predicts the TSDF as well as occupancy for this specific grid location. To this end, the translator concatenates the feature vectors of the neighborhood and compresses them into a single feature vector using a linear layer followed by tanh activation. Next, the so combined features are concatenated with the query point feature and passed through the remaining translation network, which consists of four linear layers interleaved with tanh activations and channel-wise dropout preventing the network from overfitting to a single feature channel. According to the desired output ranges, the TSDF head is activated using tanh, while the occupancy head uses a sigmoid activation. After each layer, we concatenate the output with the original query point feature .
where and denote the and norms, and is the binary cross-entropy on the predicted occupancy. The loss is helpful with outliers, whereas the loss improves the reconstruction of fine details. In each step, denotes the number of all updated feature grid locations. When training with outlier contaminated data, we found that setting equal to all visited feature grid locations yields the best results. Therefore, is a crucial hyperparameter when training the pipeline. Moreover, and denote the ground-truth TSDF and occupancy value, respectively. To avoid large deviations for a single feature in the latent space, we regularize the feature grid by penalizing the mean of the channel-wise variance by . We empirically set the loss weights to , , , and .
Experiments
We first discuss implementation details and evaluation metrics before evaluating our method on synthetic and real-world data in comparison to other methods. We further analyze our method for varying numbers of features in an ablation study. We provide additional experiments and results in the supplementary material.
Implementation Details. Our pipeline is implemented in PyTorch and trained on an NVIDIA RTX 2080. All networks were trained using the Adam optimizer with an initial learning rate of , which was adapted using an exponential learning rate scheduler at a rate of . For momentum and beta, we empirically found the default parameters to yield the best results. We trained all networks on synthetic data being augmented with artificial noise and outliers. The batch-size is set to one due to the nature of the sequential fusion process. However, we accumulate the gradients across 8 scene update steps and then update the network parameters. Our un-optimized implementation runs at frames per seconds with a depth map resolution of on an NVIDIA RTX 2080. This demonstrates the real-time applicability of our approach.
Evaluation Metrics. We use the following evaluation metrics to quantify the performance of our approach: Mean Squared Error (MSE), Mean Absolute Distance (MAD), Accuracy (Acc.), Intersection-over-Union (IoU), Mesh Completeness (M.C.) Mesh Accuracy (M.A.), and F1 score. Further details are in the supplementary material.
Datasets. We used the synthetic ShapeNet and ModelNet datasets for performance evaluation. From ShapeNet, we selected 13 classes for training and evaluate on the same test set as consisting of 60 objects from six classes, for which pretrained models are available. For ModelNet , we trained and tested on 10 classes using the train-test split from . We first generated watertight models using the mesh-fusion pipeline used in and computed TSDFs using the mesh-to-sdfhttps://github.com/marian42/mesh_to_sdf library. Additionally, we render depth frames for 100 randomly sampled camera views for each mesh. These depth maps are the input to our pipeline and existing methods. For both datasets, we found that training on one single object per class is sufficient for generalization to any other object and class.
Comparison to Existing Methods. For performance comparisons, we fuse depth maps and augment them with artificial depth-dependent noise as in . We compare to state-of-the-art learned scene representation methods DeepSDF , OccupancyNetworks , and IF-Net , as well as to the online fusion methods TSDF Fusion and RoutedFusion . We further implemented two additional baselines to demonstrate the benefits of a fully learned scene representation for depth map fusion: (1) one baseline performs a learned 2D noise filtering before fusing the frames using TSDF Fusion , and (2) a baseline that post-processes models fused by TSDF Fusion using a simplification of our translation network - the principle is similar to OctnetFusion , but on a dense grid. Further details on these baselines is given in the supplementary material. We compare all baselines on the test set of Weder et al\onedot in Figure 4. For input data augmentation, we used the same scale as in . Figure 4 shows that our method significantly outperforms all existing depth map fusion as well as learned scene representations. We especially emphasize the increase in IoU by more than 10%. This significant increase is due to many fine-grained improvements, where RoutedFusion wrongly predicts the sign, as shown in Figure 5. In all experiments, we set the truncation distance of TSDF Fusion to cm, which is similar to the receptive field of our fusion network.
Higher Input Noise Levels. We also assess our method in fusing depth maps corrupted with higher noise levels on the ModelNet dataset in Figure 6. For this experiment, we augment the input depth maps with three different noise levels. We fuse the corrupted depth maps using standard TSDF Fusion and RoutedFusion . Since Weder et al\onedot showed that their proposed routing network significantly improved the robustness to higher input noise levels, we also tested our method with depth maps pre-processed by a routing network. For these experiments, we use the pre-trained routing network provided by .
Outlier Handling. A main drawback of is its limitation in handling outliers. To this end, we run an experiment, where we augment the input depth maps with random outlier blobs. We create this data by sampling an outlier map from a fixed distribution scaled by a fixed outlier scale. Additionally, we sample three masks with a given probability (outlier fraction) and dilate it once, twice, and three times, respectively. Then, these masks are used to select the outliers from the outlier map. We report the results of this experiment in Figure 7. Note that the results might be better with even higher outlier fractions since we only evaluate on updated grid locations. The consistency in outlier filtering and increase in updated grid locations improves the metrics.
2 Ablation Study
In a series of ablation studies, we discuss several benefits of our pipeline and justify design choices.
Iterative Fusion. Ideally, fusion algorithms should be independent from the number of integrated frames and steadily improve the reconstruction as new information becomes available. In Figure 8, we show that our method is not only better than competing algorithms from the start, but also continuously improves the reconstruction as more data is fused. The metrics are averaged at every fusion step over all scenes in the test set used for all experiments on ShapeNet .
Frame Order Permutation. Our method does not leverage any temporal information from the camera trajectory apart from the previous fusion result. This design choice allows to apply the method also to a broader class of scenarios (e.g\onedotMulti-View Stereo). Ideally, an online fusion method should be invariant to permutations of the fusion frame order. To verify this property, we evaluated the performance of our method in fusing the same set of frames in ten different random frame orders. Figure 9 shows that our method converges to the same result for any frame order and thus seems to be invariant to frame order permutations.
Feature Dimension. An important hyperparameter of our method is the feature dimension . Therefore, we show quantitative results for the reconstruction from noisy and outlier contaminated depth maps using varying in Table 1. We observe that a larger clearly improves the results, but the performance eventually saturates, which justifies our choice of features throughout the paper.
Latent Space Visualization. In order to verify our hypothesis that the translator network mostly filters outliers, since the fusion network can hardly distinguish between first entries and outliers, we visualize the fused latent space in Figure 10. While the translated end result is outlier free, the latent space clearly shows that the fusion network keeps track of most measurements.
3 Real-World Data
We also evaluate on real-world data and large-scale scenes to demonstrate scalability and generalization.
Scene3D Dataset. For real-world data evaluation, we use the lounge and stonewall scenes from the Scene3D dataset . For comparability to , we only fuse every 10th frame from the trajectory using a model solely trained on synthetic ModelNet data augmented with artificial noise and outliers.
In Figure 11, we present qualitative results of reconstructions from real-world depth maps compared to RoutedFusion and TSDF Fusion . We note that our method reconstructs the scene with significantly higher completeness than . This is due to our learned translation from feature to TSDF space, which allows to better handle noise artifacts and outliers without the need for hand-tuned, heuristic post-filtering. Further, we show improved outlier and noise artefact removal compared to TSDF Fusion while being on par with respect to completeness. These results are also quantitatively shown in Table 2.
Tanks and Temples Dataset. In order to demonstrate our methods outlier handling capability, we also run experiments on the Tanks and Temples dataset . We computed stereo depth maps using COLMAP and fused this data using our method, PSR , TSDF Fusion , and RoutedFusion . To demonstrate the easy applicability to new datasets in scenarios with limited ground-truth, we train our method on one single scene (Ignatius) from the Tanks and Temples training set. We reconstructed a dense mesh using Poisson Surface Reconstruction (PSR) , rendered artificial depth maps, and used TSDF Fusion to generate a ground-truth SDF. Then, we used this ground-truth to train the fusion of stereo depth maps. Figure 12 shows the reconstructions of the unseen Caterpillar, Truck, and M60 scene from . Our proposed method significantly reduces the amount of outliers in the scene across all models. While shows comparable results on some scenes, it is heavily dependent on its outlier post-filter, which fails as soon as there are too many outliers in the scene (see also Figure 1). Further, they also pre-process the depth maps using a 2D denoising network while our network uses the raw depth maps.
Limitations. While our pipeline shows excellent generalization capabilities (e.g\onedotgeneralizing from a single MVS training scene), it is biased to the number of observations integrated during training leading to less complete results on some test scene parts with very few observations. However, this issue can be overcome by a more diverse set of training sequences with different number of observations.
Conclusion
We presented a novel approach to online depth map fusion with real-time capability. The key idea is to perform the fusion operation in a learned latent space that allows to encode additional information about undesired outliers and super-resolution complex shape information. The separation of scene representations for fusion and final output allows for an end-to-end trainable post-filtering as a translator network, which takes the latent scene encoding and decodes it into a standard TSDF representation. Our experiments on various synthetic and real-world datasets demonstrate superior reconstruction results, especially in the presence of large amounts of noise and outliers.
Acknowledgments. We thank Akihito Seki from Toshiba Japan for insightful discussions and valuable comments. This research was supported by Toshiba and Innosuisse funding (Grant No. 34475.1 IP-ICT).
Appendix A Evaluation Metrics
In order to compare our method to state-of-the-art learning-based methods and standard TSDF fusion, we compute the following six metrics on reconstructions:
Mean Squared Error (MSE) and Mean Absolute Distance (MAD). The mean squared error measures the reconstruction error on the TSDF field by penalizing large surface deviations and outliers. The mean absolute distance is also computed on the TSDF grid. However, it mainly quantifies the performance on reconstructing fine geometric details.
Accuracy (Acc.), F1 Score, Intersection-over-Union (IoU): The accuracy is computed over the occupancy obtained from the sign of the TSDF grid. We also report the F1 score, which is the harmonic mean of precision and recall. By measuring both, completeness and accuracy, it is a more holistic metric for quantifying the performance of a reconstruction method. Moreover, we measure the IoU on the occupancy grid. The IoU especially quantifies artifacts typically encountered in reconstructions from noisy depth maps, such as surface and corner thickening and the vanishing of fine geometric details.
Mesh Completeness (M.C.) and Accuracy (M.A.) We compute the completeness using the evaluation pipeline from . The completeness describes the distance from points sampled on the ground-truth mesh to the closest point on the reconstructed mesh. Vice-versa, the accuracy computes the distance from points sampled on the reconstructed mesh to the closest point on the ground-truth mesh.
Appendix B Reproducibility
For reproducibility, our source code will be made publicly available upon publication. We further present more details of our fusion pipeline in the following.
Our method consists of four neural network components that are used for (i) sub-volume extraction of the global canonical feature volume, (ii) fusion of previously fused feature with a new depth map , (iii) integration of the fused updates back into the global feature volume, and (iv) translation from the latent feature space to TSDF and occupancy.
(i) Extraction Layer. In the extraction layer, we extract the current state of the global feature volume into a view-dependent canonical feature volume defined by the camera parameters of the current measurement. In a first step, we un-project all depth pixels into the global feature grid using the camera parameters:
where are the coordinates of the un-projected point in world coordinates and are the pixel coordinates with its corresponding depth measurements. In a second step, we sample points around in a window centered at and aligned with the direction of the viewing ray. This procedure is inspired by the extraction step used in . Finally, we convert the coordinates of each sampled point to grid coordinates and extract the current feature state using nearest-neighbor interpolation.
(ii) Feature Fusion Network. The feature fusion network consists of three components: Feature Encoder, Feature Decoder, and Feature Normalization, which are detailed in the following.
Feature Encoder: The feature encoder is built from four network blocks each consisting of the following modules: 1) a 2D convolution having kernel size of 3 and zero padding reducing the number of input channels, 2) a layer normalization, 3) tanh activation, 4) again a 2D convolution having kernel size of 3 and zero padding but without reducing the channels, 5) layer normalization, and 6) tanh activation. The ouput of each block is concatenated with its input and passed to the next block.
Feature Decoder: The feature decoder also consists of four neural blocks, of which each has the same design as the neural blocks in the feature encoder. However, instead of having a kernel-size of three, the convolutional layers in the feature decoder have a kernel-size of one. The motivation behind this choice is that the encoder has already encoded enough neighboring information and, therefore, the decoder predicts the feature updates based on the encoded information for each ray separately.
Feature Normalization: After predicting the feature updates for each position in the local, view-dependent feature volume, we normalize each feature vector. This normalization prevents the feature values from becoming too large and, therefore, it improves the pipeline’s capability to update the scene.
(iii) Integration Layer. In the integration layer, we integrate the predicted feature updates from the fusion network back into the global feature volume. Therefore, we aggregate all updates that are mapped to the same global feature grid location using the correspondence given by the camera parameters. Then, we use an average pooling to combine multiple correspondences to the same grid location. Finally, we update the feature volume by using a running average update similar to .
(iv) Feature Translation Network. The feature translation network renders the output modalities (TSDF and occupancy) from the latent feature representation for a specific query point . It consists of three components: a neighborhood interpolator, a translation MLP, and two network heads predicting the output modalities.
Neighborhood Interpolator: The neighborhood interpolator encodes information from the neighboring feature vector into a single feature vector. Therefore, all neighboring feature vectors are concatenated and passed through a single linear layer followed by tanh activation. The output has the same dimension as one single feature vector.
Translation MLP: The output of the neighborhood interpolator is concatenated with the query point feature vector and passed through the translation MLP. The translation MLP is built from four linear layers interleaved with tanh activations. During training, the output of each layer is further passed through a channel-wise dropout layer with dropout probability . The first layer has 32 output channels, the second layer has 16 output channels, and the third and fourth layer have each 8 output channels. The output of each layer is concatenated with the query point feature vector. The output of the final layer is then passed to the two network output heads.
Network Output Heads: The translation network has two network output heads. Each head takes the output of the MLP as an input and predicts a translation modality. The TSDF head predicts the TSDF using a single linear layer outputting one channel that is followed by a tanh activation. The output of the activation is further scaled by the truncation band of the ground-truth TSDF () to map it into the correct value range. The occupancy head is also passed through a linear layer but activated using a sigmoid activation to map into a unit interval.
B.2 Details on Hyperparameters
Table 3 summarizes the choice of all hyperparameters that we used for all experiments in our work.
Appendix C Qualtitative Results
In this section, we present more qualitative results to demonstrate the performance of our method.
In Figure 14, we show additional qualitative results for the outlier robustness experiment that we presented in the main paper. Our method is consistently better in filtering outliers than existing methods that have no outlier filtering or filter outliers heuristically.
C.2 Real-World Data
In Figure 15 we show more results on the real-world Scene3D dataset. With this experiment, we demonstrate that our method generalizes well to real-world data and is able to fuse and reconstruct measurements of large-scale real-world scenes.
Appendix D Further Evaluation
In order to demonstrate the compactness and generalization performance of our network, we train it only on a single chair object from the ModelNet dataset. We augment the input depth maps with artificial noise of scale . We report the results in Table 4 and show that our method trained on a single object achieves almost the same performance as our method trained on the full training set. Moreover, it outperforms the currently best performing method - RoutedFusion - that is trained on the full training set.
This result indicates the applicability of our method to many real-world scenarios, where the sensor setup might change. In fact, only very little training data is required to retrain our method and achieve state-of-the-art reconstruction results.
D.2 Loss Ablation
We have also run an ablation study to evaluate the importance of the different terms in our loss function. In Figure 13, we show that the combination of all three loss terms yields best results. The binary cross entropy is particularly useful to improve convergence in the beginning of the training as the network learns to predict a coarse shape that is further refined by the losses on the SDF as training progresses.