VolumeFusion: Deep Depth Fusion for 3D Scene Reconstruction
Jaesung Choe, Sunghoon Im, Francois Rameau, Minjun Kang, In So Kweon
Introduction
Multi-view stereo (MVS) is a fundamental research topic that has been extensively investigated over the past decades . The main goal of MVS is to reconstruct a 3D scene from a set of images acquired from different viewpoints. This problem is commonly framed as a correspondence search problem by optimizing photometric or geometric consistency among groups of pixels in different images. Conventional non-learning-based MVS frameworks generally achieve the reconstruction using various 3D representations : depth maps , point cloud , voxels , and meshes .
The recent use of deep neural networks for MVS has proven effective in addressing the limitations of traditional techniques like repetitive patterns, low-texture regions, and reflections. Usually, deep-based MVS methods center on the estimation of the pixel-wise correspondence between a reference image and its surrounding views. While this strategy can be elegantly integrated into deep learning frameworks, they can only work locally where frames largely overlap. For the full 3D reconstruction of the entire scene, these methods require to perform a depth map fusion to merge local reconstructions as post-processing.
More recently, Murez et al. suggest that a direct regression of the Truncated Signed Distance Function (TSDF) volume of the scene is more effective than using intermediate 3D representations, i.e. depth maps. The overall concept of their technique consists in the back-projection of all the extracted image features into a global scene volume from which the network directly regresses the TSDF volume. This end-to-end approach has the advantage of being straightforward and scalable to a large number of images. However, this pioneering work has difficulty in shaping the global structure of complex scenes, such as the corners of rooms or long hallways.
To address this issue, we propose to closely mimic the traditional 3D reconstruction pipeline with two distinct stages: local reconstruction and global fusion. However, unlike the previous studies and concurrent papers , we integrate these two stages in an end-to-end manner. First, our network computes the local geometry, i.e., dense depth maps from neighboring frames. Then, we begin the depth fusion process by merging local depth maps as well as image features in a single volume representation where our network regresses TSDF for the final 3D scene reconstruction. This enables our end-to-end framework to learn a globally consistent volumetric representation in a single forward computation without the need for manually engineered fusion algorithms . To further enhance the robustness of our depth fusion mechanism, we propose the Posed Convolution Layer (Posed-Conv). Compared to the traditional 3D convolution layer that is solely invariant to translation, we propose a more versatile convolution layer that is invariant to both translation and rotation. In short, our Posed Convolution Layer helps to extract pose-invariant feature representation regardless of the orientation of the input image. As a result, our method demonstrates globally consistent shape reconstruction even under wide baselines or large rotations between the views. Our contributions are summarized as follows:
A novel network taking advantage of local MVS and global depth fusion for 3D scene reconstruction.
A new rotation and translation invariant convolution layer, called PosedConv.
Related Works
Multi-View Stereo (MVS) consists in the pixel-wise 3D reconstruction of a scene given a set of unstructured images along with their respective intrinsic and extrinsic parameters. In , the scene representation is used as an axis of taxonomy to categorize MVS into four sub-fields of research: depth maps , point clouds , voxels , and meshes . In particular, The depth map estimation approaches have been widely researched since these strategies are easily scalable to number of multi-view images. These methods usually rely on a small baseline assumption to ensure large overlap with a reference frame. Thus, the photometric matching is achieved via a plane-sweeping algorithm where the most probable depth is estimated for each pixel in the reference frame. Then, a depth map fusion algorithm is required to build a global 3D model from a set of depths.
With the rise of deep neural networks, learning-based MVS methods have achieved promising results. Inspired by stereo matching networks , MVS studies have developed cost volume for unstructured multi-view matching. Relying on basic frameworks, such as DPSNet or MVSNet , follow-up research proposes point-based depth refinement , cascaded depth refinement , and temporal fusion network . After exhaustively estimating a collection of depth maps, depth fusion starts to reconstruct the global 3D scene.
2 Depth Map Fusion
In their seminal work, Curless and Levoy propose a volumetric depth map fusion approach able to deal with noisy depth maps through cumulative weighted signed distance function. Follow-up research, such as KinectFusion or voxel hashing , concentrate on the problem of volumetric representation via depth maps fusion. Recently, with the help of deep learning networks, learning-based volumetric approaches have been proposed. For instance, SurfaceNet and RayNet infer the depth maps from multi-view images and their known camera poses. These methods are close to our strategy, but their networks are trained only by a per-view depth map which is not directly related to depth maps fusion. More recently, RoutedFusion and Neural Fusion introduce a new learning-based depth map fusion using RGB-D sensors. However, these papers concentrate on depth fusion algorithm using noisy and uncertain depth maps from multi-view images, not RGB-D sensors. To the best of our knowledge, our method is the first learning-based depth maps fusion from multi-view images for 3D reconstruction.
Volume Fusion Network
where is an overlapping loss and is a true overlapping mask at pixel . Note that the overlapping mask is an essential element for our depth map fusion stage (see Sec. 3.2).
2 Volumetric Depth Fusion
The traditional fusion process attempts to achieve this reconstruction by using the photometric consistency check of the pixel values that are back-projected using the inferred depth maps . However, these approaches suffer from the following drawbacks: 1) the quality of the final reconstruction depends on the accuracy of the initial depth maps, and 2) the brightness consistency assumption is violated under challenging conditions such as changing lighting conditions and homogeneously textured regions. Alternatively, recent learning-based approaches have been developed, such as RoutedFusion . This approach produces high-quality 3D reconstruction but also requires reliable depth maps from RGB-D sensors. However, it is hardly suited for depth maps obtained via multi-view images which tend to be significantly more noisy.
To overcome these issues, our depth fusion strategy propagates the masked depth maps as well as the feature volumes , which are computed from the image features by back-projecting the image features into the world-referential coordinate system. This strategy allows the network to re-profile the surface of the 3D scene by computing the matching cost of the image features guided by the masked depth maps. To fuse the per-view masked depth maps , we iteratively compute the per-view masked depth map and image feature. First, we declare a 3D volume following and initialize all voxels with . Then, we back-project a masked depth map and compute their voxel location . The value of each voxel occupied by a back-projected depth map is incremented by . We repeat this process for all views, and then average the volume as shown in Fig. 4. This strategy allows to compute the 3D surface probability using all the previously computed depths.
The embedded voxels in the unified scene volume are the initial guidance for shaping the global structure of the target scene through volume fusion. To further enhance the geometric property of each embedded image feature, we introduce the concept of the Posed Convolution Layer (PosedConv) for an accurate reconstruction of the target scene.
3 Posed Convolution Layer
Renowned stereo matching approaches employ a series of 3D convolution layers (3DConv) to find dense correspondences between the left-right image features. Since this strategy proved to be effective, it became the most commonly used technique to build a cost volume. As a result, it has also been widely applied in un-rectified multi-view stereo pipelines . Nonetheless, it appears that this representation is not appropriate for multi-view stereo. To understand this problem, we need to analyze both configurations: calibrated stereo pair and un-calibrated multi-view stereo. For calibrated stereo images, corresponding patches in both views are acquired from the same orientation (aligned optical axis); hence, the translation-invariant convolution operation is suitable to find consistent matches (see Fig. 3-(a)). However, MVS scenarios are more complex since images can be acquired from various orientations; therefore, the 3DConv is not appropriate since it is not rotation invariant (see Fig. 3-(b)). This phenomenon leads to poor matching quality when the orientations of the different viewpoints vary too greatly.
4 Volumetric 3D Reconstruction
where is the ground-truth TSDF volume and is the absolute distance measurements (i.e. L1-loss), and is the voxel location within the TSDF volume.
Finally, our network is trained in an end-to-end manner with the following three different losses:
where , , and are , , and , respectively. The depth loss and the overlapping loss guide the network to perform local multi-view stereo matching for explicit depth estimation. The TSDF loss is intended to transform the explicit geometry depth maps and the per-view image features into an implicit representation through our volume fusion.
Experiment
Following Murez et al. , we conduct the assessment of our method on the ScanNet dataset . The ScanNet dataset consists of 800 indoor scenes. Each scene contains a sequence of RGB images, the corresponding ground-truth depth maps, and the cameras’ parameters. Among them, 700 scenes are used for training while the remaining 100 scenes constitute our testing set. To obtain the ground-truth TSDF volume , we follow the original scheme proposed by a previous study .
2 Comparison with State-of-the-art Methods
To validate the effectiveness of the proposed approach, we compare the reconstruction performance to a variety of traditional geometry-based and deep learning-based methods in both 3D space and the 2.5D depth map domain. Specifically, we use the four common quantitative measures (AbsRel, AbsDiff, SqRel, and RMSE)5 of depth map quality and three common criteria (, Accuracy(Acc), Completeness(Comp), and F-score) We provide a detailed description of these metrics in our supplementary material on 3D reconstruction quality. The quantitative results are reported in Table 1. The evaluation is conducted with the 100 scenes from the ScanNet test set following the evaluation pipeline described in Murez et al. . As our experimental results show, the proposed method outperforms all competitors for all evaluation metrics of depth map by a large margin. We conjecture that the significant performance gap results from our depth map fusion method, which effectively matches multiple images even at a large viewpoint difference and increases the number of observations for matching.
In the evaluation of 3D reconstruction in Table 1, our approach outperforms the others especially on the and Comp. It suggests that our method is outstanding for shaping the global structure of the scene. For the Acc metric, our method demonstrates the second-best results after GPMVS , which shows that temporal fusion is a potential method to improve the quality of the multi-view depth map. However, for the 3D reconstruction of the overall scene (, Comp), our method largely outperforms the temporal fusion method . The best F-score is obtained by COLMAP . As described in DPSNet , COLMAP has the strength to accurately reconstruct the edges or corners where distinctive features are extracted.
Additionally, we show qualitative results of the depth fusion results in Fig. 6, and the 3D scene reconstruction in Fig. 5. Compared to the most recent competitive technique , our method better preserves the global structure of the scene, especially for the shapes of complex rooms – e.g. with corridors. Moreover, concerning the quality of depth estimation, our method shows improved depth accuracy after fusion through volume fusion. We attribute the performance improvement to our two-stage approach that exploits the pose-invariant features that re-profile the surface of the scene with the robust matching in a world-referential coordinate, i.e., the unified scene volume ).
3 Ablation Study
In this section, we propose to evaluate the contribution of each proposed component (PosedConv, depth map fusion, and overlapping mask), through an extensive ablation study. The obtained results are shown in Tables 2, 3, and 4.
In Table 2, we validate the PosedConv by comparing with the origin 3D Conv as in Fig. 3-(a)). This result highlights the relevance of PosedConv since it consistently improves the depth (AbsRel and RMSE) and reconstruction ( and F-score) accuracy. We attribute these results to the rotation-invariant and translation-invariant features extracted from our PosedConv, which helps to re-profile the 3D surface in the depth map fusion stage (Fig. 2).
In Table 3, we conduct an ablation study regarding the two-stage strategy of our method. Single-stage represents the direct TSDF regression as proposed in Murez et al. , and two-stage means our volume fusion network equipped with the PosedConv and depth map fusion. To verify the necessity of the depth fusion (i.e., depth maps embedding), we also include one additional method (w/o depth embedding) whose unified scene volume only contains pose-invariant feature volume . The results demonstrate that our two-stage approach outperforms the single-stage strategy . Moreover, when we embed masked depth maps , the performance gap to the previous work increases. Thanks to our volume fusion network with PosedConv and differentiable depth map fusion scheme, we obtain the best 3D reconstruction results.
Lastly, we validate the overlapping mask in Table 4. This mask is used to filter out the depth values at the non-overlapping region between local multi-view images. It shows that applying overlapping masks consistently improves the quality of the depth maps and 3D reconstruction. Based on these results, we confirm that using overlapping masks under varying camera motion is effective for depth map fusion.
Conclusion
In this work, we presented an end-to-end volume fusion network for 3D scene reconstruction using a set of images with known camera poses. The specificity of our strategy is its two-stage structure that mimics traditional techniques: local multi-view stereo and Volumetric Depth Fusion. For the local depth maps estimation, we designed a novel multi-view stereo method that also estimates an overlapping mask. This additional output enables filtering out of the depth measurements that do not have the pixel correspondence among the neighbor views, which in turn, improves the overall reconstruction of the scene. In our depth fusion, we fuse the masked depth maps as well as the image features to re-profile the 3D surface by computing the matching cost. To improve the robustness of matching between images with different orientations – in the world-referential coordinate (i.e., unified scene volume), we introduce the Posed Convolution Layer that extracts pose-invariant features. Finally, our network infers a TSDF volume that describes the global structure of the target scene. Despite the consistency and accuracy of our 3D reconstructions, volumetric representation requires huge computation power and memory that restricts the resolution of the resulting TSDF. To cope with this problem, future works for volume-free fusion via local multi-view information should be explored. In this context, our method represents a solid base for the development of future research in multi-view stereo and depth map fusion for 3D scene reconstruction.
ACKNOWLEDGMENT
This work was supported by NAVER LABS Corporation [SSIM: Semantic and scalable indoor mapping]
References
A Overview
This document provides additional information that we did not fully cover in the main paper, VolumeFusion: Deep Depth Fusion for 3D Scene Reconstruction. All references are consistent with the main paper.
B Evaluation Metrics
In this section, we describe the metrics for depth and 3D reconstruction evaluation. Depth evaluation involves four different metrics: Abs Rel, Abs Diff, Sq Rel, and RMSE. Each of these metrics is calculated as:
Regarding the 3D geometry evaluation, we propose four different metrics: , accuracy (Acc), completeness (Comp), and F-score. is the absolute difference between the ground-truth TSDF and the inferred TSDF, the accuracy is the distance from predicted points to the ground-truth points, completeness is the distance from the ground-truth points to the predicted points, and F-score is the harmonic mean of the precision and recall. Each of these metrics is calculated as:
C Implementation and Training Scheme
D Discrete Kernel Rotation
This section describes the details of PosedConv. The idea of the PosedConv is to rotate the reservoir kernel by using the known camera poses to extract rotation-invariant features (Sec. 3.3 of the manuscript).
The rotated kernel can be simply computed by rotation followed by linear interpolation, called Rotation-by-Interpolation, as shown in Fig. 7-(a). However, the naively rotated kernels often fail to interpolate properly rotated reservoir kernel values because of the different radius from the center to the boundary. Thus, we design a discrete kernel rotation performed on a unit sphere to minimize the loss of boundary kernel information as shown in Fig. 7-(b), called Discrete Kernel Rotation.
In detail, we first transform the 3D cube into a unit sphere by normalizing its distance from the center of the kernel. Within this unit sphere, the rotated sphere also becomes a unit sphere so that we can alleviate the information drops during the rotation. After rotating the unit sphere, we denormalize the radius of the sphere to form the 3D cube.
This process rotates the reservoir kernel of the n-th camera view by using the rotation matrix , as depicted in Fig. 7 of the main paper. Concretely, we iteratively apply the discrete kernel rotation for each n-th feature volume , as in Fig. 2 of the main paper. Given the reservoir kernel and the rotation matrix , the algorithm infers the rotated kernel as in Alg. 1. To densely fill the rotated kernel with the reservoir kernel , we need to repetitively apply the reverse warping process.
Each voxel at the voxel coordinate , is transformed into its position on the unit sphere through the NORM function (Line 4 in Alg. 1). The details of the NORM function is described in Alg. 2 and this is the fundamental difference between our Discrete Kernel Rotation (Fig. 7-(b) of the main paper and Alg. 1) and rotation-by-interpolation (Fig. 7-(a) of the main paper and Alg. 4). Since this is a reverse warping process, we use the given rotation matrix (Line 5 in Alg. 1) that is the inverse of the rotation matrix as depicted in Fig. 7 of the main paper. To do so, we obtain the rotated location lying on the unit sphere (Line 5 in Alg. 1).
To find the corresponding location in the reservoir kernel, we denormalize the rotated location and obtain the rotated voxel coordinate at the reservoir kernel (Line 6 in Alg. 1) where DeNorm function is precisely described in Alg. 3. Since the rotated voxel coordinate is not located outside of the boundary of the reservoir kernel , we directly apply bilinear interpolation to extract the kernel response at at the reservoir kernel (Line 7 in Alg. 1). Contrary to our discrete kernel rotation, naive rotation-by-interpolation must check whether the rotated voxel coordinate is out of the boundary of the reservoir kernel , or it sometimes has difficulty in interpolating the reservoir kernel into the rotated kernel (Lines 8-10 in Alg. 4).
Finally, we compute the rotated kernel for the n-th camera view. As described in Sec. 3.3 of the main paper, we identically apply the conventional 3D convolution operation that is equipped with our rotated kernels .
To validate the necessity of our discrete kernel rotation, we conduct an ablation study as in Table 5. Alongside the results in Table 2 of the manuscript, we additionally report the accuracy when we use naive Rotation-by-Interpolation (i.e., Rot interp). It shows that the rotating the kernel in both ways (Rot interp and discrete kernel rotation) improves the quality of depth and 3D geometry, but our PosedConv further achieves the higher accuracy than that of Rotation-by-Interpolation.
E Overlapping Mask
In this section, we present an example figure of an overlapping mask as in Fig. 8. In the first stage of our network, multi-view stereo, we utilize three neighbor views to infer a depth map and an overlapping mask. This overlapping mask is used to filter out the uncertain depth values at the specific pixels that have no corresponding pixels in the neighbor views. As shown in Fig. 8, our network properly infers the overlapping mask in a referential camera viewpoint. Note that the referential image is in between two adjacent views, the overlapping region is usually located at the center of the referential images.
F Combined Ablation Results.
To clearly show influences from our contributions, we merge ablation results in the manuscripts as in Table 6. The quality of the reconstruction is largely improved with both PosedConv and overlapping masks, simultaneously. This is because overlapping masks are designed to filter out the uncertain depth values that can hurt the quality of 3D reconstruction done by PosedConv. In conclusion, our two novel contributions, PosedConv and overlapping masks, are complementary and lead to precise 3D scene reconstruction.
G Additional Results
We conduct quantitative evaluations on the 7-Scenes dataset (Shotton et al., CVPR’13) in Table 7. Similarly, our network achieves state-of-the-art performance against the recent approaches. For a fair comparison, we strictly follow the assessment pipelines of these concurrent approaches ( and Long et al. in ECCV’20).