IterMVS: Iterative Probability Estimation for Efficient Multi-View Stereo

Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Marc Pollefeys

Introduction

Multi-view stereo (MVS) describes the technology to reconstruct dense 3D models of observed scenes from a set of calibrated images. MVS is a fundamental problem of geometric computer vision and a core technique for applications like augmented/virtual reality, autonomous driving and robotics. Albeit being studied extensively for decades, the conditions that occur in real-world application scenarios pose problems such as occlusion, illumination changes, low-textured areas and non-Lambertian surfaces that remain unsolved up to now.

Traditional methods suffer from hand-crafted modeling and matching metrics and struggle under those challenging conditions. In comparison, recent data-driven approaches based on Convolutional Neural Networks (CNN) demonstrate a significantly improved performance on various MVS benchmarks .

A popular representative, MVSNet , constructs a 3D cost volume that is regularized with a 3D CNN and regresses the depth map from the probability volume. While this methodology achieves impressive performance on benchmarks, it does not scale up well to high-resolution images or large-scale scenes, because the 3D CNN is memory and run-time consuming. However, low run-time and power consumption are the key in most industrial applications and resource friendly methodologies become more important. Aiming to improve the efficiency, recent variants of MVSNet are proposed, which can be mainly divided into two categories: recurrent methods and multi-stage methods . Several recurrent methods can relax the memory consumption by regularizing the cost volume sequentially with GRU or convolutional LSTM , however, at the cost of an increase in run-time. In contrast, multi-stage methods utilize cascade cost volumes and estimate the depth map from coarse to fine. While this approach can lead to high efficiency in both memory and run-time, a reduced search range at finer stages implies a limitation in recovering from errors induced at coarse resolutions .

A current method that unites a competitive performance with highest efficiency in memory and run-time among all learning-based methods is PatchmatchNet . Based on traditional PatchMatch , PatchmatchNet combines learned adaptive propagation and evaluation modules with a cascaded structure. While sharing the common limitation of coarse-to-fine methods, the generalization ability of PatchmatchNet appears limited when compared to other multi-stage methods .

In this work, we propose IterMVS, a novel GRU-based iterative method aimed at further improving efficiency as well as performance for high-resolution MVS.

Contributions: (i) We propose a novel and lightweight GRU-based probability estimator that encodes the per-pixel probability distribution of depth in its hidden state. This compressed representation does not require to keep the probability volume in memory the whole time. In each iteration, multi-scale matching information is injected to update the pixel-wise depth distribution. Compared to coarse-to-fine methods, the GRU-based probability estimator always operates at the same resolution, utilizes a large search range and keeps track of the distribution over the full depth range. (ii) We propose a simple, yet effective depth estimation strategy that combines both classification and regression, which is robust to multi-modal distributions but also achieves sub-pixel precision. (iii) We verify the effectiveness of our method on various MVS datasets, e.g. DTU , Tanks & Temples and ETH3D . The results demonstrate that IterMVS achieves very competitive performance, while showing highest efficiency in both memory and run-time among all the learning-based methods, Fig. 1. Compared with PatchmatchNet , IterMVS is more efficient in both memory and run-time, achieves comparable performance on DTU and demonstrates much better generalization ability on Tanks & Temples and ETH3D .

Related Work

Traditional MVS. Based on the scene representations, traditional MVS methods can be divided into three main categories: volumetric, point cloud based and depth map based. Volumetric methods discretize 3D space into voxels and label each as inside or outside of the true surface. Operating in scene space usually comes at the price of large memory and run-time consumption, limiting applications to scenes of smaller scale. Point cloud based methods operate directly on 3D points and often employ propagation to gradually densify the reconstruction. By decoupling the problem into depth map estimation and fusion, depth map based methods are more concise and flexible. Galliani et al. propose Gipuma, a multi-view extension of Patchmatch stereo, which uses a red-black checkerboard pattern to parallelize propagation. In COLMAP, Schönberger et al. jointly estimate pixel-wise view selection, depth map and surface normal. Although traditional depth map based methods can achieve impressive results, hand-crafted models and features limit the performance under challenging conditions.

Data-driven MVS. Recently, data-driven methods dominate the research for MVS. Several volumetric methods first compute a cost volume from multiple images and infer surface voxels after cost volume regularization with a 3D CNN. However, similar to traditional volumetric methods, they are restricted to small-scale reconstructions. More common are depth map based methods that often operate in similar fashion. MVSNet can be seen as a blueprint. It computes an initial cost volume from the features that is regularized with a 3D CNN and regresses the depth map from the probability volume. The high memory consumption of 3D CNNs often limits these methods to down-sampled cost volumes and depth maps. Recently, several variants based on MVSNet were proposed that aim to reduce memory and run-time consumption. The two main ideas involve recurrent and multi-stage methods . R-MVSNet sequentially regularizes 2D slices of the cost volume with a GRU . D2D^{2}HC-RMVSNet augments R-MVSNet with a complex convolutional LSTM . The main drawback of these recurrent methods is the high run-time. In contrast, multi-stage methods achieve efficiency in both memory and run-time. They operate on cascade cost volumes and estimate the depth map in a coarse-to-fine manner. First, a low resolution depth map is computed utilizing a large but coarse sampling interval. After upsampling, the estimation is refined at a higher sampling rate but at a smaller interval and search range. PatchmatchNet further proposes an adaptive procedure based on PatchMatch that achieves superior efficiency among all learning-based methods. Despite their impressive performance, coarse-to-fine methods have difficulty to recover from errors introduced at coarse resolutions , where the sampling interval is large but the sampling frequency low. In contrast, we estimate the depth map at a relatively high resolution and generate hypotheses in a fixed large search range in each GRU iteration. Further, we let the hidden state of our GRU encode the probability distribution for the whole depth range.

Iterative Update. Recently, RAFT proposes to estimate optical flow by iteratively updating a motion field through a GRU, which emulates first-order optimization. The idea is further adopted in stereo , scene flow and SfM . In our work, we let the GRU model a probability distribution per pixel from which we predict the depth map. The hidden state is updated in each iteration to more accurately model the pixel-wise probability distribution.

Method

In this section, we introduce the detailed structure of IterMVS, illustrated in Fig. 2.

It consists of a multi-scale feature extractor, an iterative GRU-based probability estimator that models a probability distribution of the depth at each pixel, and a spatial upsampling module.

Given NN input images of size W ⁣× ⁣HW\!\times\!H, we use I0\mathbf{I}_{0} and {Ii}i=1N−1{\left\{\mathbf{I}_{i}\right\}}_{i=1}^{N-1} to denote reference and source images respectively. Similar to , we extract multi-scale features from the images with a Feature Pyramid Network (FPN) . We attain features at 3 scale levels, and denote the feature of the ii-th image at level ll by Fi,l\mathbf{F}_{i,l}. The features Fi,l\mathbf{F}_{i,l} are stored at 1/2l1/2^{l} resolution and possess C ⁣= ⁣16,32,64C\!=\!16,32,64 channels at level l ⁣= ⁣1,2,3l\!=\!1,2,3 respectively. Although we refrain from using an explicit coarse-to-fine structure, e.g. , we can include multi-scale context information by using our multi-scale features for matching similarity computation in each iteration of the GRU. This improves performance as shown in Table 4.

2 GRU-based Probability Estimator

Differentiable Warping. Following most learning-based MVS methods , we warp the source features into front-to-parallel planes w.r.t. the reference view at the given depth hypotheses. Specifically, for a pixel p\mathbf{p} in the reference view and the jj-th depth hypothesis dj:=dj(p)d_{j}:=d_{j}(\mathbf{p}), with known intrinsic {Ki}i=0N−1\left\{\mathbf{K}_{i}\right\}_{i=0}^{N-1} and relative transformations {[R0,i∣t0,i]}i=1N−1\left\{\left[\mathbf{R}_{0,i}|\mathbf{t}_{0,i}\right]\right\}_{i=1}^{N-1} between reference view and source view ii, we can compute the corresponding pixel pi,j:=pi(dj)\mathbf{p}_{i,j}:=\mathbf{p}_{i}(d_{j}) in the source view as:

After de-homogenization, we obtain the warped feature Fi(pi,j)\mathbf{F}_{i}(\mathbf{p}_{i,j}) via bilinear interpolation.

Finally, the integrated matching similarity Sinitial(p,j)\mathbf{S}_{\textrm{initial}}(\mathbf{p},j) for pixel p\mathbf{p} and depth hypothesis djd_{j} is given by:

where σ(⋅)\sigma(\cdot) denotes the sigmoid nonlinearity, ⊙\odot denotes the Hadamard Product. With the multi-scale matching information injected in each iteration, the hidden state can encode the pixel-wise depth distribution more accurately (Table 5).

Usual strategies to predict a depth value from such sampled distributions are to take the argmax or the soft argmax . The former corresponds to measuring the Kullback-Leibler divergence between a one-hot encoding of the ground truth and P\mathbf{P}, but cannot deliver solutions beyond the discretization level (e.g. ‘sub-pixel’ solutions). The latter corresponds to measuring the distance of the expectation of P\mathbf{P} to the ground truth depth. While the expectation can take any continuous value, the measure cannot handle multiple modes in P\mathbf{P} and strictly prefers unimodal distributions. At this point we propose a new hybrid strategy that combines classification and regression. It is robust to multi-modal distributions but also achieves ‘sub-pixel’ precision that is not limited by the sampling resolution. Specifically, we find the index X(p)\mathbf{X}(\mathbf{p}) with the highest probability for pixel p\mathbf{p} from probability P\mathbf{P}:

With a radius rr, we take the expectation in the local inverse range to compute the depth estimate D(p)\mathbf{D}(\mathbf{p}):

3 Spatial Upsampling

4 Loss Function

where the probability Pinitial\mathbf{P}_{\textrm{initial}} is generated from Sˉinitial\mathbf{\bar{S}}_{\textrm{initial}} by applying softmax along the depth dimension. Then we bilinearly upsample Dinitial\mathbf{D}_{\textrm{initial}} to 1/41/4 resolution. The losses are based on D′\mathbf{D}^{\prime} that is converted from a depth map D\mathbf{D}:

We can summarize our loss function as follows:

and only consider NvalidN_{\textrm{valid}} pixels with valid ground truth depth. The loss function is composed of five component-level losses: classification loss Lclass,kL_{\textrm{class},k}, regression losses LinitialL_{\textrm{initial}}, Lregress,kL_{\textrm{regress},k}, LupsampleL_{\textrm{upsample}} and confidence loss Lconf,kL_{\textrm{conf},k},

where γ\gamma is empirically set to 0.0020.002. To train the classification, we define Q\mathbf{Q} as the one-hot encoding generated by the ground truth depth Dgt,2\mathbf{D}_{\textrm{gt},2}, peaking at the nearest discrete location and use cross-entropy loss for LclassL_{\textrm{class}}. The subsequent regression can only change the initial estimate of the classification by rr in any direction. Hence, the loss LregressL_{\textrm{regress}} only considers pixels whose ground truth classification index, Xgt,2(p)\mathbf{X}_{\textrm{gt},2}(\mathbf{p}), falls within a radius rr of the estimated index Xk(p)\mathbf{X}_{k}(\mathbf{p}). For the 1st1^{st} epoch of the training, we exclude LregressL_{\textrm{regress}} and LconfL_{\textrm{conf}} from the loss to warm up the classification.

Experiments

The DTU dataset is an indoor multi-view stereo dataset with 124 different scenes and 7 different lighting conditions. We use the training, testing and validation split introduced in . BlendedMVS is a large-scale dataset, which provides over 17k high-quality training samples of various scenes. Tanks & Temples is a large-scale outdoor dataset consisting of complex environments that are captured under real-world conditions. The ETH3D benchmark consists of calibrated high-resolution images of real-world scenes with strong viewpoint variations.

2 Implementation Details

Implemented with PyTorch, two models are trained on DTU (Ours) and Large-Scale BlendedMVS (Ours-LS) respectively. For DTU, we use an image resolution of 640×512640\times 512 and the number of input images to N=5N=5. Observing that a fixed depth range (e.g. MVSNet ) filters out some of the ground truth signals on DTU, we propose to estimate a depth range per-view from the sparse ground truth point clouds for both training and evaluation. This is comparable to the treatment of real-world scenes in MVS, where the per-view depth range is usually estimated from a sparse SfM reconstruction, e.g. . For BlendedMVS, we use an image resolution of 768×576768\times 576 and N=5N=5 input images. Here, the depth range is provided by the dataset. To improve the robustness during training, we follow to randomly select the source views and also scale the scene within the range of [0.8,1.25][0.8,1.25].

We use D1 ⁣= ⁣32D_{1}\!=\!32 during initialization and perform K=4K=4 update iterations. Here, we set R1=2−7,R2=2−5,R3=2−3R_{1}=2^{-7},R_{2}=2^{-5},R_{3}=2^{-3} and N1=4,N2=4,N3=2N_{1}=4,N_{2}=4,N_{3}=2. When predicting the depth, we use D2=256D_{2}=256 and r=4r=4. We train for 16 epochs using Adam (β1 ⁣= ⁣0.9,β2 ⁣= ⁣0.999\beta_{1}\!=\!0.9,\beta_{2}\!=\!0.999). The learning rate is initially set to 0.001, and halved after 4, 8 and 12 epochs. We set the batch size to 4 and 2 for the models trained on DTU and BlendedMVS. The models are trained on a single Nvidia RTX 2080Ti GPU. After depth estimation, we perform confidence filtering (threshold τ\tau is set to 0.30.3) and geometric consistency filtering to remove outliers and then reconstruct the point clouds (see supplementary).

3 Benchmark Performance

Evaluation on Tanks & Temples. We set the input image size to 1920×10241920\times 1024, the number of views, NN, to 7 and the number of iterations, KK, to 4. The camera parameters and depth ranges are estimated with OpenMVG . During evaluation, the GPU memory consumption and run-time for estimating each depth map are 2426 MB and 0.405 s respectively. Among the methods trained on DTU , as shown in Table 2, the F-score of our method on the intermediate dataset is similar to CasMVSNet and much better than many other multi-stage methods . On the complex advanced dataset, our method performs best in Recall and F-score among all the methods. Compared with PatchmatchNet , our method performs much better on all the metrics of both datasets. Training on the large-scale BlendedMVS dataset appears to improve the generalization ability of our method on both datasets compared with the model trained on DTU only. On the advanced dataset, our method performs better in F-score than Vis-MVSNet , which is a state-of-the-art multi-stage method. Overall, operating at a notably low memory and run-time, our method demonstrates very competitive generalization performance compared to most learning-based methods.

Evaluation on ETH3D Benchmark. We set the input image size to 1920 ⁣× ⁣12801920\!\times\!1280, the number of views NN to 7 and the number of iterations KK to 4. The camera parameters and depth ranges are estimated with COLMAP . During evaluation, the GPU memory consumption and run-time for estimating each depth map are 2898 MB and 0.410 s respectively. The results are summarized in Table 3. When trained on DTU , our method performs much better in accuracy than PVSNet and PatchmatchNet . On the training dataset, our method achieves comparable performance in F1F_{1}-score as PVSNet and COLMAP . On the test dataset, the performance of our method in F1F_{1}-score is better than PatchmatchNet , PVSNet and COLMAP , which is competitive to ACMH . When trained on BlendedMVS , the performance of our method improves significantly in all metrics on both datasets compared with the model solely trained on DTU. Our method performs best in F1F_{1}-score on both datasets among all competitors. This analysis further demonstrates the effectiveness, efficiency and generalization capabilities of our method.

We conduct an ablation study of our model trained on DTU to analyze the effectiveness of its components. Results of the first six experiments are summarized in Table 4.

Depth Prediction. For the depth estimated from GRU, instead of using both classification and regression, we also try both argmax (classification only) and soft argmax (regression only), i.e. taking the expectation on all the D2D_{2} depth samples. Our hybrid strategy performs better on both DTU and ETH3D . Apparently, the gap on the unseen ETH3D data widens when (also) utilizing the classification loss, whereas classification only leads to inferior accuracy on DTU. We conjecture that soft argmax (regression) only training could be more prone to overfitting, implied by the tendency to force single peaked distributions.

Confidence. We compare our idea of explicitly learning the confidence from the hidden state to a commonly used strategy that defines the confidence as the sum of 4 probability samples near the estimate from the probability volume P\mathbf{P}. Following MVSNet , we use a threshold of τ ⁣= ⁣0.8\tau\!=\!0.8 for 3D reconstruction. Our strategy performs slightly better on DTU and ETH3D . A visualization in Fig. 4 shows that our learned confidence can deliver more accurate predictions at object boundaries. When taking confidence as a sum of probabilities, the un-matchable uniformly colored background pixel near the boundary receives a confidence as high as foreground pixels, due to an over-smoothing effect induced by averaging probabilities.

Scale of Feature. Recall that we utilize multi-scale features for similarity computation throughout the Iterative Update. Here, we also try to only use level 2 features (we still employ level 3 features in the Initialization phase to reduce computation). Our analysis shows that our model benefits from using multi-scale context information on both DTU and ETH3D .

Depth Upsampling. We compare learned upsampling to simple bilinear upsampling. Clearly, the learned approach delivers better results on both DTU and ETH3D .

Inverse Depth Loss. By default, we compute the losses LinitialL_{\textrm{initial}}, LregressL_{\textrm{regress}}, LupsampleL_{\textrm{upsample}} w.r.t. the inverse depth. We also try computing the losses without converting to the inverse depth range. While the result on DTU is similar, the performance on the large-scale ETH3D dataset, is explicitly improved when using inverse depth.

Pixel-wise View Weight. By default, we estimate the pixel-wise view weight to robustly integrate the matching information from the source views. Here, we compare to the simpler approach of taking the average over the views. The large performance difference on ETH3D can be explained by strong viewpoint variations within the dataset, such that our robust integration of the matching information across the views becomes very impactful.

Number of Iterations. Table 5 relates the performance in accuracy, completeness and overall quality to the number of iterations on the GRU. Running more iterations allows to inject more matching information into the hidden state, leading to more accurate depth probability distributions per pixel. Although we limit our model to K ⁣= ⁣4K\!=\!4 iterations during training, the performance keeps improving beyond that in all metrics. Comparing the result with Table 4.3, we can even observe that our method achieves state-of-the-art performance at K ⁣= ⁣16K\!=\!16 on DTU. Despite the increased run-time, the inference is still faster than most multi-stage methods . We conclude by noting that the flexibility to run our model for an arbitrary number of iterations enables the user to trade-off time efficiency for performance in dependence of the application.

Number of Views. We vary the number of views NN and summarize the results in Table 6. Multi-view information can help to alleviate problems such as occlusions and the reconstruction quality improves at a higher value, saturating at around 6 views.

While our network structure allows to trade-off speed and accuracy by adjusting the number of iterations during inference, we need to fix the number of samples that our probability distributions consist of. This number, D2D_{2}, is afterwards determined by the network structure and cannot be adjusted for different scenes. Likewise predetermined is the range of the data term samples placed around the current solution that are ingested into the CNN. In our model those samples cover 2R3 ⁣= ⁣1/4th2R_{3}\!=\!1/4^{\textrm{th}} of the total inverse depth range.

We present IterMVS, a novel learning-based MVS method combining highest efficiency and competitive reconstruction quality. We propose to explicitly encode a pixel-wise probability distribution of depth in the hidden state of a GRU-based estimator. In each iteration, we inject multi-scale matching information and extract the – in the inverse depth range – uniformly sampled depth distribution to estimate depth map and confidence. Extensive experiments on DTU, Tanks & Temples and ETH3D show highest efficiency in both memory and run-time, and a better generalization ability than many state-of-the-art learning-based methods.

Following RAFT , we upsample the depth map from 1/41/4 to full resolution. Specifically, the depth of each pixel in the high resolution depth map is a convex combination of its 9 neighbors at the coarse resolution. The weights are learned from the reference feature map. Fig. 5 illustrates the upsampling process.

Before fusing the depth maps, we filter out unreliable depth estimates, following MVSNet . There are two filtering steps: geometric consistency filtering and confidence filtering.

Geometric Consistency Filtering. Following MVSNet , we apply a geometric constraint to measure the consistency of depth estimates among multiple views. For each pixel p\mathbf{p} in the reference view, we project it, using its depth d0d_{0}, to a pixel pi\mathbf{p}_{i} in the ii-th source view. After looking up its depth did_{i} in the source view, we reproject pi\mathbf{p}_{i} into the reference view, and look up the depth dreprojd_{\textrm{reproj}} at this location, preproj\mathbf{p}_{\textrm{reproj}}. We consider pixel p\mathbf{p} and its depth d0d_{0} as consistent to the ii-th source view, if the distances, in image space and depth, between the original estimate and its reprojection satisfy:

where δ=1\delta=1 and ε=0.01\varepsilon=0.01 are two thresholds. Finally, we accept the estimations as reliable, if they are consistent in at least NgeoN_{\textrm{geo}} source views.

Confidence Filtering. Since our learned confidence indicates how close the estimation is to the ground truth depth, we use it to filter out estimations with high uncertainty. Specifically, we use a confidence threshold τ=0.3\tau=0.3 throughout the experiments to filter out all the pixels with confidence lower than it.

Our GRU-based probability estimator encodes the per-pixel probability distribution of depth with the hidden state. A 2D CNN is applied on the hidden state to estimate the probability of D2D_{2} samples that are uniformly distributed in the inverse depth range for each pixel. We visualize this probability distribution for various scenes in Fig. 6. For pixels with distinct features, the probability has a single dominant peak and the estimation is precise. For some challenging situations, where distributions are non-peaky or multi-modal, e.g. in textureless areas, our hybrid depth estimation strategy can still robustly produce estimations as accurate as possible. We also visualize the update process of probability distribution in Fig. 7. We observe that the probability distribution becomes more focused, several local maxima get suppressed and the estimation becomes more precise with more GRU iterations. In each iteration, multi-scale matching information is injected into the hidden state. This allows the hidden state to more accurately model the per-pixel probability distribution of depth with each iteration.

Several examples for our estimated pixel-wise view weights are depicted in Fig. 8. Comparing the view weights with the visible areas in the reference validates that the all visible areas receive higher weights, while occluded and invisible parts have very low weights. Interestingly, in the first two images, pixels on the windows have low weights. Here, especially the upper row of windows mirror the surrounding buildings and cannot provide reasonable matching information. The other images have low weights in visible regions at areas with strong perspective and specular reflections as well as occlusions, while fronto-parallel and textured regions achieve higher weights. We conclude that our pixel-wise view weight is capable to determine co-visible areas between the reference and source images.

We visualize the reconstructed point clouds from DTU’s evaluation set , Tanks & Temples dataset and ETH3D benchmark in Fig. 9, 10 and 11.

Currently, the learned confidence is only used to filter out unreliable estimates before depth fusion. However, we believe it will be a promising direction to further exploit the confidence in each GRU iteration. For example, one can refine the depth of unconfident areas with the information propagated from those confident areas . Another idea would be to focus more effort on unconfident areas only, while leaving the confident areas unchanged, which further should improve efficiency.