PatchmatchNet: Learned Multi-View Patchmatch Stereo

Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Pablo Speciale, Marc Pollefeys

Introduction

Given a collection of images with known camera parameters, multi-view stereo (MVS) describes the task of reconstructing the dense geometry of the observed scene. Despite being a fundamental problem of geometric computer vision that has been studied for several decades, MVS is still a challenge. This is due to a variety of de-facto unsolved problems occurring in practice such as occlusion, illumination changes, untextured areas and non-Lambertian surfaces .

The success of Convolutional Neural Networks (CNN) in almost any field of computer vision ignites the hope that data driven models can solve some of these issues that classical MVS models struggle with. Indeed, many learning-based methods appear to fulfill such promise and outperform some traditional methods on MVS benchmarks . While being successful at the benchmark level, most of them do only pay limited attention to scalability, memory and run-time. Currently, most learning-based MVS methods construct a 3D cost volume, regularize it with a 3D CNN and regress the depth. As 3D CNNs are usually time and memory consuming, some methods down-sample the input during feature extraction and compute both, the cost volume and the depth map at low-resolution. Yet, according to Fig. 1, delivering depth maps at low resolution can harm accuracy. Methods that do not scale up well to realistic image sizes of several mega-pixel cannot exploit the full resolution due to memory limitations. Evidently, low memory and time consumption are key to enable processing on memory and computational restricted devices such as phones or mixed reality headsets, as well as in time critical applications. Recently, researchers tried to alleviate these limitations. For example, R-MVSNet decouples the memory requirements from the depth range and sequentially processes the cost volume at the cost of an additional runtime penalty. include cascade 3D cost volumes to predict high-resolution depth map from coarse to fine with high efficiency in time and memory.

Several traditional MVS methods abandon the idea of holding a structured cost volume completely and instead are based on the seminal Patchmatch algorithm. Patchmatch adopts a randomized, iterative algorithm for approximate nearest neighbor field computation . In particular, the inherent spatial coherence of depth maps is exploited to quickly find a good solution without the need to look through all possibilities. Low memory requirements – independent of the disparity range – and an implicit smoothing effect make this method very attractive for our deep learning based MVS setup.

In this work, we propose PatchmatchNet, a novel cascade formulation of learning-based Patchmatch, which aims at decreasing memory consumption and run-time for high-resolution multi-view stereo. It inherits the advantages in efficiency from classical Patchmatch, but also aims to improve the performance with the power of deep learning.

Contributions: (i) We introduce the Patchmatch idea into an end-to-end trainable deep learning based MVS framework. Going one step further, we embed the model into a coarse-to-fine framework to speed up computation. (ii) We augment the traditional propagation and cost evaluation steps of Patchmatch with learnable, adaptive modules that improve accuracy and base both steps on deep features. We estimate visibility information during cost aggregation for the source views. Moreover, we propose a robust training strategy to introduce randomness into training for improved robustness in visibility estimation and generalization. (iii) We verify the effectiveness of our method on various MVS datasets, \egDTU , Tanks & Temples and ETH3D . The results demonstrate that our PatchmatchNet achieves competitive performance, while reducing memory consumption and run-time compared to most learning-based methods.

Related Work

Traditional MVS. Traditional MVS methods can be divided into four categories: voxel-based , surface evolution based , patch-based and depth map based . Comparatively, depth map based methods are more concise and flexible. Here, we discuss Patchmatch Stereo methods in this category. Galliani et al. present Gipuma, a massively parallel multi-view extension of Patchmatch stereo. It uses a red-black checkerboard pattern to parallelize message-passing during propagation. Schönberger et al. present COLMAP, which jointly estimates pixel-wise view selection, depth map and surface normal. ACMM adopts adaptive checkerboard sampling, multi-hypothesis joint view selection and multi-scale geometric consistency guidance. Based on the idea of Patchmatch, we propose our learning-based Patchmatch, which inherits the efficiency from classical Patchmatch, but also improves the performance leveraging deep learning.

Learning-based stereo. GCNet introduces 3D cost volume regularization for stereo estimation and regresses the final disparity map with a soft argmin operation. PSMNet adds spatial pyramid pooling (SPP) and applies a 3D hour-glass network for regularization. DeepPruner develops a differentiable Patchmatch module, without learnable parameters, to discard most disparities and then builds a lightweight cost volume, which is regularized by a 3D CNN. In contrast, we do not apply any cost volume regularization but extend the original Patchmatch idea into the deep learning era. Xu et al. propose a sparse point based intra-scale cost aggregation method with deformable convolution . Likewise, we propose a strategy to adaptively sample points for spatial cost aggregation.

Learning-based MVS. Voxel-based methods are restricted to small-scale reconstructions, due to the drawbacks of a volumetric representation. In contrast, based on plane-sweep stereo , many recent works use depth maps to reconstruct the scene. They build cost volumes with warped features from multiple views, regularize them with 3D CNNs and regress the depth. As 3D CNNs are time and memory consuming, they usually use down-sampled cost volumes. To reduce memory, R-MVSNet sequentially regularizes 2D cost maps with GRU but sacrifices run-time. Current research targets to improve efficiency and also estimate high-resolution depth maps. CasMVSNet proposes cascade cost volumes based on a feature pyramid and estimates the depth map in a coarse-to-fine manner. UCS-Net proposes cascade adaptive thin volumes, which use variance-based uncertainty estimates for an adaptive construction. CVP-MVSNet forms an image pyramid and also constructs a cost volume pyramid. To accelerate propagation in Patchmatch, we likewise employ a cascade formulation. In addition to cascade cost volumes, PVSNet learns to predict visibility for each source image. An anti-noise training strategy is used to introduce disturbing views. We also learn a strategy to adaptively combine the information of multiple views based on visibility information. Moreover, we propose a robust training strategy to include randomness into the training to improve robustness in visibility estimation and generalization. Fast-MVSNet constructs a sparse cost volume to learn a sparse depth map and then use high-resolution RGB image and 2D CNN to densify it. We build a refinement module and use the RGB image to guide the up-sampling of the depth map based on MSG-Net .

Method

In this section, we introduce the structure of PatchmatchNet, illustrated in Fig. 2. It consists of multi-scale feature extraction, learning-based Patchmatch included iteratively in a coarse-to-fine framework, and a spatial refinement module.

Given NN input images of size W ⁣× ⁣HW\!\times\!H, we use I0\mathbf{I}_{0} and {Ii}i=1N−1{\left\{\mathbf{I}_{i}\right\}}_{i=1}^{N-1} to denote reference and source images respectively. Before we apply our learning-based Patchmatch algorithm, we extract pixel-wise features from our inputs, similar to Feature Pyramid Network (FPN) . Features are extracted hierarchically at multiple resolutions, which allows us to advance our depth map estimation in a coarse-to-fine manner.

2 Learning-based Patchmatch

Following traditional Patchmatch and subsequent adaptations to depth map estimation , our learnable Patchmatch consists of the following three main steps:

Initialization: generate random hypotheses;

Propagation: propagate hypotheses to neighbors;

Evaluation: compute the matching costs for all the hypotheses and choose best solutions.

After initialization, the approach iterates between propagation and evaluation until convergence. Leveraging deep learning, we propose an adaptive version of the propagation (Sec. 3.2.2) and evaluation (Sec. 3.2.3) module and also adjust the initialization (Sec. 3.2.1). The detailed structure of our Patchmatch pipeline is illustrated in Fig. 3. In a nutshell, the propagation module adaptively samples the points for propagation based on the extracted deep features. Our adaptive evaluation learns to estimate visibility information for cost computation and adaptively samples the spatial neighbors to aggregate the costs again based on deep features. Unlike , we refrain from parameterizing the per-pixel hypothesis as a slanted plane, due to heavy memory penalties. Instead, we rely on our learned adaptive evaluation to organize the spatial pattern within the window over which matching costs are computed.

In the very first iteration of Patchmatch, the initialization is performed in a random manner to promote diversification. Based on a pre-defined depth range [dmin,dmax][d_{min},d_{max}], we sample per pixel DfD_{f} depth hypotheses in the inverse depth range, corresponding to uniform sampling in image space. This helps our model be applicable to complex and large-scale scenes . To ensure we cover the depth range evenly, we divide the (inverse) range into DfD_{f} intervals and ensure that each interval is covered by one hypothesis.

For subsequent iterations on stage kk, we perform local perturbation by generating per pixel NkN_{k} hypotheses uniformly in the normalized inverse depth range RkR_{k} and gradually decrease RkR_{k} for finer stages. To define the center of RkR_{k}, we utilize the estimation from previous iteration, possibly up-sampled from a coarser stage. This delivers a more diverse set of hypotheses than just using propagation. Sampling around the previous estimation can refine the result locally and correct wrong estimates (see supplementary).

2.2 Adaptive Propagation

Spatial coherence of depth values does in general only exist for pixel from the same physical surface. Hence, instead of propagating depth hypotheses naively from a static set of neighbors as done for Gipuma and DeepPruner , we want to perform the propagation in an adaptive manner, which gathers hypotheses from the same surface. This helps Patchmatch converge faster and deliver more accurate depth maps. Fig. 4 illustrates the idea and functionality of our strategy. Our adaptive scheme tends to gather hypotheses from pixels of the same surface – for both the textured object and the textureless region – enabling us to effectively collect more promising depth hypotheses compared to using just a static pattern.

We base our implementation of the adaptive propagation on Deformable Convolution Networks . As the approach is identical for each resolution stage, we omit subindices denoting the stage. To gather KpK_{p} depth hypotheses for pixel p\mathbf{p} in the reference image, our model learns additional 2D offsets {Δoi(p)}i=1Kp{\left\{\Delta\mathbf{o}_{i}(\mathbf{p})\right\}}_{i=1}^{K_{p}} that are applied on top of fixed 2D offsets {oi}i=1Kp{\left\{\mathbf{o}_{i}\right\}}_{i=1}^{K_{p}}, organized as a grid. We apply a 2D CNN on the reference feature map F0\mathbf{F}_{0} to learn additional 2D offsets for each pixel p\mathbf{p} and get the depth hypotheses Dp(p)\mathbf{D}_{p}(\mathbf{p}) via bilinear interpolation as follows:

where D\mathbf{D} is the depth map from previous iteration, possibly up-sampled from a coarser stage.

2.3 Adaptive Evaluation

The adaptive evaluation module performs the following steps: differentiable warping, matching cost computation, adaptive spatial cost aggregation and depth regression. As the approach is identical at each resolution stage, we omit subindices to ease notation.

Differentiable Warping. Following plane sweep stereo , most learning-based MVS methods establish front-to-parallel planes at sampled depth hypotheses and warp the feature maps of source images into them. Equipped with intrinsic matrices {Ki}i=0K\left\{K_{i}\right\}_{i=0}^{K} and relative transformations {[R0,i∣t0,i]}i=1K\left\{\left[\mathbf{R}_{0,i}|\mathbf{t}_{0,i}\right]\right\}_{i=1}^{K} of reference view and source view ii, we compute the corresponding pixel pi,j:=pi(dj)\mathbf{p}_{i,j}:=\mathbf{p}_{i}(d_{j}) in the source for a pixel p\mathbf{p} in the reference, given in homogeneous coordinates, and depth hypothesis dj:=dj(p)d_{j}:=d_{j}(\mathbf{p}) as follows:

We obtain the warped source feature maps of view ii and the jj-th set of (per pixel different) depth hypotheses, Fi(pi,j)\mathbf{F}_{i}(\mathbf{p}_{i,j}), via differentiable bilinear interpolation.

Matching Cost Computation. For multi-view stereo, this step has to integrate information from an arbitrary number of source views into a single cost per pixel p\mathbf{p} and depth hypothesis djd_{j}. To that end, we compute the cost per hypothesis via group-wise correlation and aggregate over the views with a pixel-wise view weight . In that manner we can employ visibility information during cost aggregation and gain robustness. Finally, the per group costs are projected into a single number, per reference pixel and hypothesis, by a small network.

To find pixel-wise view weights, {wi(p)}i=1N−1{\left\{\mathbf{w}_{i}(\mathbf{p})\right\}}_{i=1}^{N-1}, we exploit the diversity of our initial set of depth hypotheses in the first iteration on stage 3 (Sec. 3.2.1). We intend wi(p)\mathbf{w}_{i}(\mathbf{p}) to represent the visibility information of pixel p\mathbf{p} in the source image Ii\mathbf{I}_{i}. The weights are computed once and kept fixed and up-sampled for finer stages.

where Pi(p,j)\mathbf{P}_{i}(\mathbf{p},j) intuitively represents confidence of visibility for the range covered by the jj-th depth hypothesis at p\mathbf{p}.

The final per group similarities Sˉ(p,j)\mathbf{\bar{S}}(\mathbf{p},j) for pixel p\mathbf{p} and the jj-th hypothesis are the weighted sum of Si(p,j)\mathbf{S}_{i}(\mathbf{p},j) and the view weight wi(p)\mathbf{w}_{i}(\mathbf{p}):

where wkw_{k} and dkd_{k} weight the cost C\mathbf{C} based on feature and depth similarity (details in supplementary). Similar to adaptive propagation, the per pixel sets of displacements {Δpk}k=1Ke\left\{\Delta\mathbf{p}_{k}\right\}_{k=1}^{K_{e}} are found by applying a 2D CNN on the reference feature map F0\mathbf{F}_{0}. Fig. 5 exemplifies the learned adaptive aggregation window. Sampled locations stay within object boundaries, while for the textureless region, the sampling points aggregate over a larger spatial context, which can potentially reduce the ambiguity of estimation.

3 Depth Map Refinement

Instead of using Patchmatch also on the finest resolution level (stage 0), we find it sufficient to directly up-sample (from resolution W2 ⁣× ⁣H2\frac{W}{2}\!\times\!\frac{H}{2} to W ⁣× ⁣HW\!\times\!H) and refine our estimation with the RGB image. Based on MSG-Net , we design a depth residual network. To avoid being biased for a certain depth scale, we pre-scale the input depth map into the range $andconvertitbackafterrefinement.Ourrefinementnetworklearnstooutputaresidualthatisaddedtothe(up−sampled)estimationfromPatchmatch,and convert it back after refinement. Our refinement network learns to output a residual that is added to the (up-sampled) estimation from Patchmatch,\mathbf{D},togettherefineddepthmap, to get the refined depth map\mathbf{D}_{ref}.Thisnetworkindependentlyextractsfeaturemaps. This network independently extracts feature maps\mathbf{F}_{D}andand\mathbf{F}_{I}fromfrom\mathbf{D}andand\mathbf{I}_{0}andappliesdeconvolutiononand applies deconvolution on\mathbf{F}_{D}$ to up-sample the feature map to the image size. Multiple 2D convolution layers are applied on top of the concatenation of both feature maps – depth map and image – to deliver the depth residual.

4 Loss Function

Loss function LtotalL_{total} considers the losses among all the depth estimation and rendered ground truth with same resolution as a sum:

We adopt the smooth L1L1 loss for LikL_{i}^{k}, the loss of the ii-th iteration of Patchmatch on stage kk (k=1,2,3k=1,2,3) and Lref0L_{ref}^{0}, the loss for final refined depth map.

Experiments

We evaluate our work on multiple datasets, such as DTU , Tanks & Temples and ETH3D and analyze each new component with an ablation study.

The DTU dataset is an indoor multi-view stereo dataset with 124 different scenes where all scenes share the same camera trajectory. We use the training, testing and validation split introduced in . The Tanks & Temples dataset is provided as a set of video sequences in realistic environments. It is divided into intermediate and advanced datasets. ETH3D benchmark consists of calibrated high-resolution images of scenes with strong viewpoint variations. It is divided into training and test datasets.

2 Robust Training Strategy

Many learning-based methods select two best source views based on view selection scores to train models on DTU . However, the selected source views have a strong visibility correlation with the reference view, which may affect the training of the pixel-wise view weight network. Instead, we propose a robust training strategy based on PVSNet . For each reference view, we randomly choose four from the ten best source views for training. This strategy increases the diversity at training time and augments the dataset on the fly, which improves the generalization performance. In addition, training on those random source views with weak visibility correlation generates further robustness for our visibility estimation.

3 Implementation Details

We implement the model with PyTorch and train it on DTU’s training set . We set the image resolution to 640×512640\times 512 and the number of input images to N=5N=5. The selection of source images is based on the proposed robust training strategy. We set the iteration number of Patchmatch on stages 3,2,13,2,1 as 2,2,12,2,1. For initialization, we set Df ⁣= ⁣48D_{f}\!=\!48. For local perturbation, we set R3 ⁣= ⁣0.38,R2 ⁣= ⁣0.09,R1 ⁣= ⁣0.04R_{3}\!=\!0.38,R_{2}\!=\!0.09,R_{1}\!=\!0.04 (see supplementary), N3 ⁣= ⁣16,N2 ⁣= ⁣8,N1 ⁣= ⁣8N_{3}\!=\!16,N_{2}\!=\!8,N_{1}\!=\!8. In the adaptive propagation, we set KpK_{p} on stages 3,2,13,2,1 to 16, 8, 0 (no propagation for last iteration on stage 1, see supplementary). For the adaptive evaluation, we use Ke=9K_{e}=9 on all stages. We train our model with Adam (β1 ⁣= ⁣0.9,β2 ⁣= ⁣0.999\beta_{1}\!=\!0.9,\beta_{2}\!=\!0.999) for 8 epochs with a learning rate of 0.001. Here, we use a batch size of 4 and train on 2 Nvidia GTX 1080Ti GPUs. After depth estimation, we reconstruct point clouds similar to MVSNet .

4 Benchmark Performance

Memory and Run-time Comparison. We compare the memory consumption and run-time with several state-of-the-art learning-based methods that achieve competing performance with low memory consumption and run-time: CasMVSNet , UCS-Net and CVP-MVSNet . These methods propose a cascade formulation of 3D cost volumes and output depth maps at the same resolution as the input images. As shown in Fig. 7, memory and run-time of all the methods increase almost linearly w.r.t. the resolution as the number of depth hypotheses is fixed (notably this will lead to under-sampling of the enlarged image space for methods using a naive cost volume approach). Note that at higher resolutions other methods could not fit into the memory of the GPU used for evaluation. We further observe that memory consumption and run-time increase much slower for PatchmatchNet than for other methods. For example, at a resolution of 1152 ⁣× ⁣8641152\!\times\!864 (51.8%), memory consumption and run-time are reduced by 67.1% and 66.9% compared to CasMVSNet, by 55.8% and 63.9% compared to UCS-Net and by 68.5% and 83.4% compared to CVP-MVSNet. Combining the results in Table 4.4, we conclude that our method is much more efficient in memory consumption and run-time than most state-of-the-art learning-based methods, at a very competitive performance.

Evaluation on Tanks & Temples Dataset. We use the model trained on DTU without any fine-tuning. For evaluation, we set the input image size to 1920×10561920\times 1056 and the number of views NN to 7. The camera parameters and sparse point cloud are recovered with OpenMVG . During evaluation, the GPU memory and run-time for each depth map are 2887 MB and 0.505 s respectively. As shown in Table 4.4, the performance of our method on the intermediate dataset is comparable to CasMVSNet , which has the highest score. For the more complex advanced dataset, our method performs best among all the methods. Overall, due to its simple, scalable structure, our PatchmatchNet demonstrates competitive generalization performance, low memory consumption and low run-time compared to state-of-the-art learning-based methods that commonly use 3D cost volume regularization.

We visualize the pixel-wise view weight in Fig. 8. Brighter colors indicate co-visible areas, \ieregions in the reference image that are also (well) visible in the source images. Conversely, areas that are not visible in the source images receive a dark color, corresponding to a low weight. Also pixel near depth discontinuities appear slightly darker than surrounding areas. By inspection, we conclude that our pixel-wise view weight is indeed capable to determine co-visible areas between the reference and source views.

Pixel-wise View Weight (VW) & Robust Training Strategy (RT). In this experiment, we resign from pixel-wise view weighting (w/o VW) and do not follow our strategy but choose four best source views for training (w/o RT). To investigate the generalization performance, we further test on Tanks & Temples and ETH3D . Table 4.5 shows a similar performance on DTU for all the models, yet, we observe a drop in performance on the other datasets without pixel-wise view weight or the robust training strategy. This proves that these two modules lead to improved robustness and a better generalization performance.

Similar to MVSNet , the point cloud reconstruction mainly consists of photometric consistency filtering, geometric consistency filtering and depth fusion. Photometric consistency filtering is used to filter out those depth hypotheses that have low confidence. Based on MVSNet , we define the confidence as the probability sum of the depth hypotheses that fall in a small range near the estimation. We use the probability P\mathbf{P} (\cfEq. 7 from the paper) from the last iteration of Patchmatch on stage 1 for filtering. In this iteration, we only perform local perturbation, without adaptive propagation. At stage 1, operating at a quarter the image resolution and with the algorithm almost converged, the hypotheses obtained via propagation from spatial neighbors are usually very similar to the current solution. Such irregular sampling of the probability space causes bias in the regression (\cfEq. 7 from the paper) and the estimate becomes over-confident at the current solution, where most propagated samples are located. In contrast, by performing only the local perturbation, the depth hypotheses are uniformly distributed in the inverse depth range. Contrary to previous iterations, we compute the estimated depth at pixel p\mathbf{p}, D(p)\mathbf{D}(\mathbf{p}), by utilizing the inverse depth regression , which is based on the soft argmin operation :

where P(p,j)\mathbf{P}(\mathbf{p},j) is the probability for pixel p\mathbf{p} at the jj-th depth hypothesis. Then we compute the probability sum of four depth hypotheses that are nearest to the estimation to measure the confidence .

Recall that in Eq. 6 of the paper we utilize two weights to aggregate our spatial costs, {wk}k=1Ke\left\{w_{k}\right\}_{k=1}^{K_{e}} based on spatial feature similarity and {dk}k=1Ke\left\{d_{k}\right\}_{k=1}^{K_{e}} based on the similarity of depth hypotheses. The feature weights {wk}k=1Ke\left\{w_{k}\right\}_{k=1}^{K_{e}} at a pixel p\mathbf{p} are based on the feature similarity at the sampling locations around p\mathbf{p}, measured in the reference feature map F0\mathbf{F}_{0}. Given the sampling positions {p+pk+Δpk}k=1Ke\left\{\mathbf{p}+\mathbf{p}_{k}+\Delta\mathbf{p}_{k}\right\}_{k=1}^{K_{e}}, we extract the corresponding features from F0\mathbf{F}_{0} via bilinear interpolation. Then we apply group-wise correlation between the features at each sampling location and p\mathbf{p}. The results are concatenated into a volume on which we apply 3D convolution layers with 1 ⁣× ⁣1 ⁣× ⁣11\!\times\!1\!\times\!1 kernels and sigmoid non-linearities to output normalized weights that describe the similarity between each sampling point and p\mathbf{p}.

As discussed in Sec. 1, neighboring pixels will be assigned different depth values throughout the estimation process. For pixel p\mathbf{p} and the jj-th depth hypothesis, our depth weights {dk}k=1Ke\left\{d_{k}\right\}_{k=1}^{K_{e}} take this into account and downweight the influence of samples with large depth difference, especially when located across depth discontinuities. To that end, we collect the absolute difference in inverse depth between each sampling point and pixel p\mathbf{p} with their jj-th hypotheses, and obtain the weights by applying a sigmoid function on the, again, inverted differences for normalization.

We use multiple stages to estimate the depth map in a coarse-to-fine manner. Here, we analyze the effectiveness of our multi-stage framework. We upsample the estimated depth maps on stages 3, 2 and 1, to the same resolution as the input and then reconstruct the point clouds. As shown in Table 5, the reconstruction quality gradually increases from coarser stages to finer ones. This shows that our multi-stage framework can reconstruct the scene geometry with increasing accuracy and completeness.

We visualize the sampling locations in two typical situations, at an object boundary and a textureless region. As shown in Fig. 11, for the pixel p\mathbf{p} at the object boundary, all sampling points tend to be located on the same surface as p\mathbf{p}. In contrast, for the pixel q\mathbf{q} in the textureless region, the sampling points are spread over a larger region. By sampling from a large region, a more diverse set of depth hypotheses can be propagated to q\mathbf{q} and reduce the local ambiguity for depth estimation in the textureless area. The visualization shows two examples how the adaptive propagation successfully adapts the sampling to different challenging situations.

Here, we again visualize the sampling locations for two situations, at an object boundary and a textureless region. Fig.12 demonstrates that for the pixel p\mathbf{p} at the object boundary, sampling points tend to stay within the boundaries of the object, such that they focus on similar depth regions. For the pixel q\mathbf{q} in the textureless region, the points are distributed sparsely to sample from a large context, which helps to obtain reliable matching and to reduce the ambiguity. Again, the visualization demonstrates how our adaptive evaluation adapts the sampling for the spatial cost aggregation to different situations.

We visualize reconstructed point clouds from DTU’s evaluation set , Tanks & Temples dataset and ETH3D benchmark in Fig. 13, 14, 15.