Pixel-Perfect Structure-from-Motion with Featuremetric Refinement
Philipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc Pollefeys
Introduction
Mapping the world is an important requirement for spatial intelligence applications in augmented reality or robotics. Tasks like visual localization or path planning can benefit from accurate sparse or dense 3D reconstructions of the environment. These can be built from images using Structure-from-Motion (SfM), which associates observations across views to estimate camera parameters and 3D scene geometry. Sparse reconstruction based on matching local image features is the most common due to its scalability and its robustness to appearance changes introduced by varying devices, viewpoints, and temporal conditions found in crowdsourced scenarios .
SfM assumes that sparse interest points can be reliably detected across views. It typically selects such points for each image independently and relies on these initial detections for the remainder of the reconstruction process. However, detecting keypoints from a single view is inherently inaccurate due to appearance changes and discrete image sampling . The advent of convolutional neural network (CNNs) for detection has magnified this issue, as they generally do not retain local image information and instead favor global context.
Multi-view geometric optimization with bundle adjustment is commonly used to refine cameras and points using reprojection errors. Dusmanu et al. proposed to refine keypoint locations prior to SfM via an analogous geometric cost constrained with local optical flow. This can improve SfM, but has limited accuracy and scalability.
In this work, we argue that local image information is valuable throughout the SfM process to improve its accuracy. We adjust both keypoints and bundles, before and after reconstruction, by direct image alignment in a learned feature space. Exploiting this locally-dense information is significantly more accurate than geometric optimization, while deep, high-dimensional features extracted by a CNN ensure wider convergence in challenging conditions. This formulation elegantly combines globally-discriminative sparse matching with locally-accurate dense details. It is applicable to both incremental and global SfM irrespective of the types of sparse or dense features.
We validate our approach in experiments evaluating the accuracy of both 3D structure and camera poses in various conditions. We demonstrate drastic improvements for multiple hand-crafted and learned local features using off-the-shelf CNNs. The resulting system produces accurate reconstructions and scales well to large scenes with thousands of images. In the context of visual localization, it can, in addition to providing a more accurate map, also refine poses of single query images with minimal overhead.
For the benefit of the research community, we will release our code as an extension to COLMAP and to the popular localization toolbox hloc . We believe that our featuremetric refinement can significantly improve the accuracy of existing datasets and push the community towards sub-pixel accurate localization at large scale.
Related work
Image matching is at the core of SfM and visual SLAM, which typically rely on sparse local features for their efficiency and robustness. The process i) detects a small number of interest points, ii) computes their visual descriptors, iii) matches them with a nearest neighbor search, and iv) verifies the matches with two-view epipolar estimation and RANSAC. The correspondences then serve for relative or absolute pose estimation and 3D triangulation. As keypoints are sparse, small inaccuracies in their locations can result in large errors for the estimated geometric quantities.
Differently, dense matching considers all pixels in each image, resulting in denser and more accurate correspondences. It has been successful for constrained settings like optical flow or stereo depth estimation , but is not suitable for large-scale SfM due to its high computational cost due to many redundant correspondences. Several recent works improve the matching efficiency by first matching coarsely and subsequently refining correspondences using a local search. This is however limited to image pairs and thus cannot create point tracks required by SfM.
Our work combines the best of both paradigms by leveraging dense local information to refine sparse observations. It is inherently amenable to SfM as it can optimize all locations over multiple views in a track simultaneously.
Subpixel estimation is a well-studied problem in correspondence search. Common approaches either upsample the input images or fit polynomials or Gaussian distributions to local image neighborhoods . With the widespread interest in CNNs for local features, solutions tailored to 2D heatmaps have been recently developed, such as learning fine local sub-heatmaps or estimating subpixel corrections with regression or the soft-argmax . Cleaner heatmaps can also arise from aggregating predictions over multiple virtual views using data augmentation .
Detections or local affine frames can be combined across multiple views with known poses in a least-squares geometric optimization . Dusmanu et al. instead refine keypoints solely based on tentative matches, without assuming known geometry. This geometric formulation exhibits remarkable robustness, but is based on a local optical flow whose estimation for each correspondence is expensive and approximate. We unify both keypoint and bundle optimizations into a joint framework that optimizes a featuremetric cost, resulting in more accurate geometries and a more efficient keypoint refinement.
Direct alignment optimizes differences in pixel intensities by implicitly defining correspondences through the motion and geometry. It therefore does not suffer from geometric noise and is naturally subpixel accurate via image interpolation. Direct photometric optimization has been successfully applied to optical flow , visual odometry , SLAM , multi-view stereo (MVS) , and pose refinement . It generally fails for moderate displacements or appearances changes, and is thus not suitable for large-baseline SfM. One notable work by Woodford & Rosten refines dense SfM+MVS models with a robust image normalization. It focuses on dense mapping with accurate initial poses and moderate appearance changes. Georgel et al. instead estimate more accurate relative poses by elegantly combining photometric and geometric costs. They show that dense information can improve sparse estimation but their approach ignores appearance changes. Differently, our work improves the entire SfM pipeline starting with tentative matches and addresses larger, challenging changes.
To improve on the weaknesses of photometric optimization, numerous recent works align multi-dimensional image representations. Examples of this featuremetric optimization include frame tracking with handcrafted or learned descriptors , optical flow , MVS , and dense SfM in small scenes . Closer to our work, PixLoc learns deep features with a large basin of convergence for wide-baseline pose refinement. It improves the accuracy of sparse matching but is designed for single images and disregards the scalability to multiple images or large scenes. Here we extend this paradigm to other steps of SfM and propose an efficient algorithm that scales to thousands of images. We show that learning task-specific wide-context features is not necessary and demonstrate highly accurate refinements with off-the-shelf features.
In conclusion, our work is the first to apply robust featuremetric optimization to a large-scale sparse reconstruction problem and show significant benefits for visual localization.
Background
Global refinement: Bundle adjustment is the gold standard for refining structure and poses given initial estimates. It minimizes the total geometric error
where is the set the images and keypoints in track , projects to the image plane, and is a robust norm . This formulation implicitly refines the keypoints while ensuring their geometric consistency. It however ignores the uncertainty of the initial detections and thus requires many observations to reduce the geometric noise. Operating on an existing reconstruction, it cannot recover observations arising from noisy keypoints that are matched correctly but discarded by the geometric verification.
where denotes the set of matches that forms the track and is a lookup with subpixel interpolation. A deep neural network is trained to regress the flow of a single point from two input patches and the flow field is interpolated from a sparse grid. This dramatically improves the keypoint accuracy, but some errors remain as the regression and interpolation are only approximate.
Both bundle and keypoint adjustments are based on geometric observations, namely keypoint locations and flow, but do not account for their respective uncertainties. They thus require a large number of observations to average out the geometric noise and their accuracy is in practice limited.
Approach
Summarizing dense image information into sparse points is necessary to perform global data association and optimization at scale. However, refining geometry is an inherently local operation, which, we show, can efficiently benefit from locally-dense pixels. Given constraints provided by coarse but global correspondences or initial 3D geometry, the dense information only needs to be locally accurate and invariant but not globally discriminative. While SfM typically discards image information as early as possible, we instead exploit it in several steps of the process thanks to direct alignment. Leveraging the power of deep features, this translates into featuremetric keypoint and bundle adjustments that elegantly integrate into any SfM pipeline by replacing their geometric counterparts. Figure 2 shows an overview.
We first introduce the featuremetric optimization in Section 4.1. We then describe our formulations of keypoint adjustment, in Section 4.2, and bundle adjustment, in Section 4.3, and analyze their efficiency.
This flow can be efficiently computed at any location in a neighborhood around , without approximate interpolation nor descriptor matching. It naturally emerges from the direct optimization of the photometric error, which can be minimized with second-order methods in the same way as the aforementioned geometric costs. Unlike the flow regressed from a black-box neural network , this flow can be made consistent across multiple view by jointly optimizing the cost over all pairs of observations.
2 Keypoint adjustment
Once local features are detected, described, and matched, we refine the keypoint locations before geometrically verifying the tentative matches.
Track separation: Connected components in the matching graph define tentative tracks – sets of keypoints that are likely to observe the same 3D point, but whose observations have not yet been geometrically verified. Because a 3D point has a single projection on a given image plane, valid tracks cannot contain multiple keypoints detected in the same image. We can leverage this property to efficiently prune out most incorrect matches using the track separation algorithm introduced in . This speeds up the subsequent optimization and reduces the noise in the estimation.
Objective: We then adjust the locations of 2D keypoints belonging to the same track by optimizing its featuremetric consistency along tentative matches with the cost
Efficiency : This direct formulation simply compares pre-computed features on sparse points and is thus much more scalable than patch flow regression (Eq. 2), which performs a dense local correlation for each correspondence. All tracks are optimized independently, which is very fast in practice despite the sheer number of tentative matches.
Once all tracks are refined, the geometric estimation proceeds, typically using two-view epipolar geometric verification followed by incremental or global SfM.
3 Bundle adjustment
The estimated structure and motion can then be refined with a similar featuremetric cost. Here keypoints are implicitly defined by the projections of the 3D points into the 2D image planes, and only poses and 3D points are optimized.
The reference is selected at the beginning of the optimization and kept fixed from then on. This reduces the drift of the points significantly, as also noted in , but is more flexible than the common ray-based parametrization .
This ensures robustness to outlier observations and accounts for the unknown topology of the feature space.
Efficiency: Compared to the keypoint adjustment (Eq. 4), using a reference feature reduces the number of residuals from to . On the other hand, all tracks need to be updated simultaneously because of the interdependency caused by the camera poses. To accelerate the convergence, we form a reduced camera system based on the Schur complement and use embedded point iterations . The refinement generally converges within a few camera updates.
4 Implementation
Dense extractor: Our refinement can work with any off-the-shelf CNN that produces feature maps that are locally discriminative. These should be of the same resolution as the input (stride 1) to enable subpixel accuracy. The radius of convergence, or context, of such features depends on the amount of noise in the keypoints. Most detectors like SIFT have at most a few pixels of error, while others like D2-Net exhibit a much larger detection noise. In our experiments, we use S2DNet for dense feature extraction, as it computes fine features very efficiently in only 4 convolutions, but also produce, if required, deeper features with a larger context. These can then be combined into a multi-level optimization scheme that sequentially refines based on coarse to fine features. The convergence can thus be adjusted depending on the detector and on the image resolution. We show in Section 5.4 that other dense features work well too.
Optimization: The optimization problems of both keypoint and bundle adjustments are solved with the Levenberg-Marquardt algorithm implemented using Ceres . Feature maps are stored as collections of 1616 patches centered around the initial keypoint detections. We thus constrain points to move at most 8 pixels. The feature lookup is implemented as bicubic interpolation. We use the Cauchy loss with a scale of 0.25. The robust mean in Eq. 7 is computed with iteratively reweighted least squares .
Run time and memory: S2DNet can extract 3-5 dense feature maps per second and both featuremetric adjustments run in less than 5 minutes for 100 images. As these features are 128-dimensional, the memory consumption can be a bottleneck. We believe that much fewer dimensions are actually required for refinement, and retraining a compact feature extractor would improve the efficiency of the optimization.
Experiments
We evaluate our featuremetric refinement on various SfM tasks with several handcrafted and learned local features and show substantial improvements for all of them. We first evaluate its accuracy on the tasks of triangulation and camera pose estimation in Sections 5.1 and 5.2, respectively. We then assess in Section 5.3 the impact of the refinement on two-view and multi-view pose estimation for end-to-end reconstruction in challenging conditions. Lastly, Section 5.4 analyzes the validity and scalability of our design decisions through an ablation study.
We first evaluate the accuracy of the refined 3D structure given known camera poses and intrinsics.
Evaluation: We use the ETH3D benchmark , which is composed of 13 indoor and outdoor scenes and provides images with millimeter-accurate camera poses and highly-accurate ground truth dense reconstructions obtained with a laser scanner. We follow the protocol introduced in , in which a sparse 3D model is triangulated for each scene using COLMAP with fixed camera poses and intrinsics. Following the original benchmark setup, we report the accuracy and completeness of the reconstruction, in %, as the ratio of triangulated and ground-truth dense points that are within a given distance of each other.
Baselines: We evaluate our featuremetric refinement with the hand-crafted local features SIFT and the learned ones SuperPoint , D2-Net , and R2D2 , using the associated publicly available code repositories. We compare our approach to the geometric optimization of , referred here as Patch Flow. We re-compute the numbers provided in the original paper using the code provided by the authors.
Results: Table 1 shows that our approach results in significantly more accurate and complete 3D reconstructions compared to the traditional geometric SfM. It is more accurate than Patch Flow, especially at the strict threshold of 1cm, and exhibits similar completeness. The improvements are consistent across all local features, both indoors and outdoors. The gap with Patch Flow is especially large for SIFT, which already detects well-localized keypoints. This confirms that our featuremetric optimization better captures low-level image information and yields a finer alignment. Patch Flow is more complete for larger thresholds as it partly solves a different problem by increasing the keypoint repeatability with its large receptive field, while we focus on their localization.
2 Camera pose estimation
We now evaluate the impact of our refinement on the task of camera pose estimation from a single image.
Evaluation: We again follow the setup of based on the ETH3D benchmark. For each scene, 10 images are randomly selected as queries. For each of them, the remaining images, excluding the 2 most covisible ones, are used to triangulate a sparse 3D partial model. Each query is then matched against its corresponding partial model and the resulting 2D-3D matches serve to estimate its absolute pose using LO-RANSAC+PnP followed by geometric refinement. We compare the 130 estimated query poses to their ground truth and report the area under the cumulative translation error curve (AUC) up to 1mm, 1cm, and 10cm.
Baselines: Patch Flow performs multi-view optimization over each partial model independently as well as over the matches between each query and its partial model. Similarly, we first refine each partial model as in Section 5.1. We then adjust the query keypoints using its tentative matches, estimate an initial pose, and refine it with featuremetric BA.
Results: The AUC and its cumulative plot are shown in Table 2. Our refinement substantially improves the localization accuracy for all local features, including SIFT, for which Patch Flow does not show any benefit. At all error thresholds, featuremetric optimization is consistently more accurate than its geometric counterparts. The accuracy of SuperPoint is raised far higher than other detectors, despite the high sparsity of the 3D models that it produces. This shows how more accurate keypoint detections can result in much more accurate visual localization.
3 End-to-end Structure-from-Motion
While the previous experiments precisely quantify the accuracy of the refinement, they do not contain any variations of appearance or camera models. We thus turn to crowd-sourced imagery and evaluate the benefits of our featuremetric optimization in an end-to-end reconstruction pipeline.
Evaluation: We use the data, protocol, and code of the 2020 Image Matching Challenge . It is based on large collections of crowd-sourced images depicting popular landmarks around the world. Pseudo ground truth poses are obtained with SfM and used for two tasks. The stereo task evaluates relative poses estimated from image pairs by decomposing their epipolar geometry. This is a critical step of global SfM as it initializes its global optimization. The multiview task runs incremental SfM for small subsets of images, making the SfM problem much harder, and evaluates the final relative poses within each subset. For each task, we report the AUC of the pose error at the threshold of 5°, where the pose error is the maximum of the angular errors in rotation and translation. As the evaluation server accepts at most correspondences, we cannot evaluate our method using the test data. We instead test on a subset of the publicly available validation scenes, and tune the RANSAC and matching parameters on the remaining scenes. More details on this setup are provided in the Appendix.
Baselines: We evaluate our refinement in combination with SIFT , D2-Net , and SuperPoint+SuperGlue . We limit the number of detected keypoints to 2k for computational reasons, but increase this number to 4k for D2-Net as it otherwise performs poorly. In the stereo task, we adjust the keypoints using the entire exhaustive tentative match graph (4950 pairs per scene). We use LO-DEGENSAC for match verification, the ratio test for SIFT, and the mutual check for SIFT and D2-Net. In the multiview task, we adjust keypoints for each subset independently, considering only the matches between images in the subset, and run our bundle adjustment after SfM.
Results: Table 3 summarizes the results. For stereo, our featuremetric keypoint adjustment significantly improves the accuracy of the two-view epipolar geometries across all local features and despite the challenging conditions. In multiview setting, it also improves the accuracy of the SfM poses, especially for small sets of images. Featuremetric optimization is particularly effective in this situation, as geometric optimization cannot fully suppress the detection noise due to the small number of observations. We visualize tracks of a 5-image reconstruction in Figure 4 and highlight the accuracy of the refined SfM model.
4 Additional insights
Ablation study: Table 4 shows the performance of several variants of our featuremetric optimization on ETH3D in terms of triangulation (scene Facade only) and localization (all scenes). We compare both types of adjustments, minor tweaks, and different image representations, including NCC-normalized intensity patches with fronto-parallel warping. Our final configuration, based on on the dense features of S2DNet , performs best across all metrics. We will now show that it is also fairly efficient.
Scalability: We run SfM on subsets of images of the Aachen Day-Night dataset . Figure 3 shows the run times of the refinement for subsets of 10, 100 and 1000 images. The featuremetric refinement is an order of magnitude faster than Patch-Flow . Precomputing distance maps reduces the peak memory requirement of the bundle adjustment from 80 GB to less than 10GB for 1000 images. As storing feature maps only requires 50 GB of disk space, this refinement can easily run on a desktop PC. We thus refined the entire Aachen Day-Night v1.1 model, composed of 7k images, in less than 2 hours. Scene partitioning could further reduce the peak memory. See Appendix D for more details.
Conclusion
In this paper we argue that the recipe for accurate large-scale Structure-from-Motion is to perform an initial coarse estimation using sparse local features, which are by necessity globally-discriminative, followed by a refinement using locally-accurate dense features. Since the dense feature only need to be locally-discriminative, they can afford to capture much lower-level texture, leading to more accurate correspondences. Through extensive experiments we show that this results in more accurate camera poses and structure; in challenging conditions and for different local features.
While we optimize against dense feature maps, we keep the sparse scene representation of SfM. This ensures not only that the approach is scalable but also that the resulting 3D model is compatible with downstream applications, e.g. mapping for visual localization. Since our refinement works well even with few observations, as it does not need to average out the keypoint detection noise, it has the potential to achieve more accurate results using fewer images.
We thus believe that our approach can have a large impact in the localization community as it can improve the accuracy of the ground truth poses of standard benchmark datasets, of which many are currently saturated. Since this refinement is less sensitive to under-sampling, it enables benchmarking for crowd-sourced scenarios beyond densely-photographed tourism landmarks.
Acknowledgements: The authors thank Mihai Dusmanu, Rémi Pautrat, Marcel Geppert, and the anonymous reviewers for their thoughtful comments. Paul-Edouard Sarlin was supported by gift funding from Huawei, and Viktor Larsson by an ETH Zurich Postdoctoral Fellowship.
Appendix
Appendix A Additional results on ETH3D
We refine the triangulation of SuperPoint keypoints for the ETH3D Courtyard scene and show in Figure 5 the distribution of triangulation errors for points observed by different numbers of images (track length). Our featuremetric refinement provides the largest improvement for points with low track length, for which the estimates of the traditional geometric BA are dominated by the noise of the keypoint detection. For larger track lengths, the refined point cloud has an accuracy close to the Faro Focus X 330 laser scanner from which the ground truth is computed.
We show in Figure 10 the raw and refined point clouds for SuperPoint and D2-Net. The benefits of our refinement are easily visible in 3D. Planar walls exhibit fewer noisy keypoints and the refined point clouds are more complete.
A.2 Camera pose estimation
We analyze in Table 5 how the different kinds of adjustments impact the accuracy of camera localization. The full method presented in the main paper first refines the 3D SfM model with featuremetric keypoint and bundle adjustments. It then refines each keypoint in the query image using its tentative 2D-3D correspondences by minimizing the featuremetric error between its observation in the query and the most similar observation of the respective 3D points. Refining the query keypoints before RANSAC increases the number of inlier matches and stabilizes the pose estimation in challenging scenarios where few 3D points are matches.
Once an initial pose is estimated with PnP+RANSAC, we refine it via a small featuremetric bundle adjustment over the inlier correspondences. This optimizes each query keypoint against the closest descriptor within the matched track. As opposed to refining each query keypoint against all observations in the matched track, this has the benefit of scaling linearly in the number of query keypoints and yields a similar accuracy.
Appendix B Impact of various parameters
Figure 6 shows how much our refinement displaces the detected keypoints during the triangulation of SuperPoint on Courtyard using dense features extracted from 1600x1066-pixel images. When using full feature maps without any constraints in keypoint adjustment, most points are moved by more than 1 pixel, but most often by less than 8 pixels. This confirms that storing the feature maps as 1616 patches is sufficient and rather conservative.
We show in Figure 7 the accuracy of the triangulation for various patch sizes. Smaller 1010 patches achieve sufficient accuracys and require significantly less memory.
B.2 Image resolution
The image resolution at which the dense features are extracted has a large impact on the accuracy of the refinement. In Figure 8 we quantify in the impact on both triangulation accuracy and run time for the ETH3D Courtyard scene (38 images). The accuracy drops significantly when the resolution is smaller than 16001066px, which amounts to 25% of the full image resolution. Doubling the resolution to 32002132px yields noticeable improvements, albeit significantly increases the extraction time and the consumption of GPU VRAM. As a reference, extracting only fine-level S2DNet features (4 convolutions) from 32002132px images requires around 10GB of GPU VRAM.
B.3 Reference selection for keypoint adjustment
Selecting some observations as references is necessary to avoid the drift. In a given track, the keypoint adjustment selects the point that is the most connected (topological center), while the bundle adjustment selects the point closest to the robust mean in feature space (feature center). Could we use the feature center for selecting the reference of the keypoint adjustment? By minimizing the feature distance to this unique reference, we could reduce the number of residuals from quadratic (pairwise constraints) to linear (unary constraints) and thus accelerate the optimization.
Retaining pairwise constraints however allows the optimization to separate tracks that were incorrectly merged by the track separation algorithm. This is not necessary in the bundle adjustment, as tracks are already filtered by the robust geometric estimation and can thus be assumed to be correct, but is common for unverified track. We evaluate the impact of the reference selection in the keypoint adjustment and report the results in Table 6. For both SuperPoint and D2-Net, using the feature center results in lower completeness and accuracy than the topological center. It also results in a lower track length, which confirms that the topological reference allows to retain incorrectly-merged tracks. Since the feature center still performs relatively well, it could be considered in case of tighter computational constraints.
Furthermore, Table 6 highlights the importance of the featuremetric keypoint adjustment. The benefits are larger for D2-Net, which detects very noisy keypoints. As a consequence, many correct albeit noisy matches are rejected by the geometric verification. Our keypoint adjustment not only allows more points to be triangulated, thus increasing the completeness of the model, but also increases the accuracy of the triangulated points.
B.4 Number of feature levels
Using multiple feature levels enlarges the basin of convergence but increases the computational requirements. The radius of convergence that is required depends on the noise of the keypoint detector and on the resolution of the image from which keypoints are detected. When performing detection and refinement at identical image resolutions, the optimal displacement is at most a few pixels for most keypoint detectors. In this case, the fine level of S2DNet feature maps is sufficient. We empirically measured that its radius of convergence is approximately 3 pixels, although the multiview constraints enable to refine over much larger distances.
We thus use a single feature level for all experiments involving SIFT, SuperPoint, and R2D2. D2-Net require a different treatment, as its detection noise is significantly larger. This is partly due to the aggressive downsampling of its CNN backbone and to the low resolution of its output heatmap. As a consequence, we employ both fine and medium feature levels for D2-Net. Both keypoint and bundle adjustments run the optimization successively at the coarser and finer levels.
B.5 Dimensionality of the features
Throughout this paper, we used 128-dimensional dense features extracted by S2DNet . Relying on compact features would easily reduce the memory footprint and the run time of the refinement. To demonstrate these benefits, we show in Figure 9 the relationship between the dimension, the run time of the BA, and the triangulation accuracy when retaining only the first channels of the S2DNet features. Features with fewer dimensions yield a faster refinement. The accuracy drops moderately but we expect a smaller reduction with features explicitly trained for smaller dimensions.
Appendix C Cost map approximation
We mention in Section 4.4 that the memory efficiency of the bundle adjustment can be improved by precomputing the featuremetric cost. We provide here more details.
This error is zero at points on the discrete grid and increases with the roughness of the feature space. This approximation thus displaces the local minimum of the cost by at most 1 pixel but most often by much less.
Improvement: This approximation however degrades the correctness of the approximate Hessian matrix that the Levenberg-Marquardt algorithm relies on for fast convergence. We found that also optimizing the squared spatial derivatives of this cost significantly improves the convergence. This simply amounts to augmenting the scalar residual map with dense derivative maps:
This improvement results in three-dimensional residuals, which is still smaller than when . Using the spatial derivatives, we can also compute an exact, more accurate bicubic spline interpolation of the cost landscape.
Evaluation: We now show experimentally that this approximation often does not, or only minimally, impairs the accuracy of the refinement. Table C reports the results of the triangulation of SuperPoint features on the ETH3D dataset. The approximation reduces the accuracy by less than 1% and does not alter the completeness. It however significantly reduces the memory consumption of the bundle adjustment, allowing it to scale to thousands of images. Note that all experiments in Sections 5.1, 5.2, and 5.3 do not use the cost map approximation as the corresponding scenes are relatively small.
For the experiments on ETH3D, we use the evaluation code provided by Dusmanu et al. . We use the original implementations of SuperPoint , D2-Net , and R2D2 , and extract root-normalized SIFT features using COLMAP . For both sparse and dense feature extraction, the images are resized so that their longest dimension is equal to 1600 pixels. The tentative matches are filtered according to the recipe described in .
D.2 Structure-from-Motion - Section 5.3
We tune the hyperparameters on the training scenes Temple Nara Japan, Trevi Fountain, and Brandenburg Gate. The results in the main paper are computed on the test scenes Sacre Coeur, Saint Peter’s Square, and Reichstag, using the data and code provided by the challenge organizers.
For SIFT , we use the mutual check, a ratio test with threshold 0.85 for the multi-view and 0.9 for the stereo tasks, and DEGENSAC with an inlier threshold of 0.5px. For D2-Net , we use the mutual check and inlier thresholds of 2px and 0.5px for raw and refined keypoints, respectively. For SuperPoint+SuperGlue , we do not use additional match filtering and we select an inlier thresholds of 1.1px and 0.5px for raw and refined keypoints, respectively. All sparse local and dense features are extracted at full image resolution, which is generally not larger than 1024px.
D.3 Ablation study - Section 5.4
The triangulation metrics are reported for the ETH3D scene Facade, which is the largest with 76 images. We use SuperPoint local features as they perform best in all earlier experiments and we store dense feature maps in every experiment. The localization AUC is measured over all 13 scenes in ETH3D with 10 holdout images per scene. We now detail the different baselines.
Localization is achieved in “F-KA” by first refining the keypoints, triangulating the map and finally performing query keypoint adjustment as described in section A.2. For localization with “F-BA”, we refined the triangulated model using featuremetric bundle adjustment and then refined the pose from PnP+RANSAC using qBA.
In the entry “w/ F-BA drift”, we use the robust reference (Eq. 7) to select the observation in each track which is most similar to the robust reference as the source frame. The optimizer then minimizes the error between each other observation and the current, moving reference of the source frame. Since only the index of the source frame is fixed during the optimization, this method does not account for drift, which appears to yield higher accuracy but suffers from repeatability problems during localization.
The baseline “PatchFlow + F-BA” uses the keypoint refinement from Dusmanu et al. as initialization, and runs our featuremetric bundle adjustment on top of it. We used the exact same parameters for PatchFlow as presented in .
The entry “higher resolution” corresponds to input images at double the resolution than all the other experiments, i.e. 3200 pixels in the longest dimension.
For the “photometric” baseline, we use RGB images (while Woodford et al. use grayscale images), we warp patches of 44 pixels at the featuremap resolution (1600 pixels in the longest dimension) with fronto-parallel assumption, and apply normalized cross correlation (NCC). Identically to our featuremetric BA and to LSPBA , the source frame is selected as the observation closest to the robust mean.
We report results for dense features extracted from a VGG-16 CNN, trained on ImageNet , at the layer conv1_2 (64 channels) and for the fine feature map predicted by PixLoc (32 channels). The model of PixLoc, trained on MegaDepth , was kindly provided by its authors. In DSIFT (128 channels), we apply a bin size of 4 and a step size of 1 and refer to the VLFeat implementation for more details.
D.4 Scalability
All experiments were conducted on 8 CPU cores (Intel Xeon E5-2630v4) and one NVIDIA RTX 1080 Ti. The subsets from the Aachen Day-Night v1.1 model were selected as the images with the largest visibility overlap, in descending order. To accelerate the feature matching, each image was matched only to its top 20 most covisible reference images in the original Aachen SfM model. We use SuperPoint features and match image pairs with the mutual check and distance thresholding at 0.7. During BA, we apply the sparse Schur solver from Ceres for each linear system in LM, while we use sparse Cholesky in KA, similar to . Featuremetric bundle adjustment is stopped after 30 iterations while KA runs for at most 100 iterations and stops when parameters change by less than .
To refine the full Aachen Day-Night model, we use SuperPoint features matched with SuperGlue from the Hierarchical Localization toolbox . We refine the keypoints with KA, then triangulate the points with fixed poses from the reference model. Finally, we run a full bundle adjustment of the model with the proposed approximation by cost maps.