DUSt3R: Geometric 3D Vision Made Easy
Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, Jerome Revaud
Introduction
Unconstrained image-based dense 3D reconstruction from multiple views is one of a few long-researched end-goals of computer vision . In a nutshell, the task aims at estimating the 3D geometry and camera parameters of a particular scene, given a set of photographs of this scene. Not only does it have numerous applications like mapping , navigation , archaeology , cultural heritage preservation , robotics , but perhaps more importantly, it holds a fundamentally special place among all 3D vision tasks. Indeed, it subsumes nearly all of the other geometric 3D vision tasks. Thus, modern approaches for 3D reconstruction consists in assembling the fruits of decades of advances in various sub-fields such as keypoint detection and matching , robust estimation , Structure-from-Motion (SfM) and Bundle Adjustment (BA) , dense Multi-View Stereo (MVS) , etc.
In the end, modern SfM and MVS pipelines boil down to solving a series of minimal problems: matching points, finding essential matrices, triangulating points, sparsely reconstructing the scene, estimating cameras and finally performing dense reconstruction. Considering recent advances, this rather complex chain is of course a viable solution in some settings , yet we argue it is quite unsatisfactory: each sub-problem is not solved perfectly and adds noise to the next step, increasing the complexity and the engineering effort required for the pipeline to work as a whole. In this regard, the absence of communication between each sub-problem is quite telling: it would seem more reasonable if they helped each other, i.e. dense reconstruction should naturally benefit from the sparse scene that was built to recover camera poses, and vice-versa. On top of that, key steps in this pipeline are brittle and prone to break in many cases . For instance, the crucial stage of SfM that serves to estimate all camera parameters, is typically known to fail in many common situations, e.g. when the number of scene views is low , for objects with non-Lambertian surfaces , in case of insufficient camera motion , etc. This is concerning, because in the end, “an MVS algorithm is only as good as the quality of the input images and camera parameters” .
In this paper, we present DUSt3R, a radically novel approach for Dense Unconstrained Stereo 3D Reconstruction from un-calibrated and un-posed cameras. The main component is a network that can regress a dense and accurate scene representation solely from a pair of images, without prior information regarding the scene nor the cameras (not even the intrinsic parameters). The resulting scene representation is based on 3D pointmaps with rich properties: they simultaneously encapsulate (a) the scene geometry, (b) the relation between pixels and scene points and (c) the relation between the two viewpoints. From this output alone, practically all scene parameters (i.e. cameras and scene geometry) can be straightforwardly extracted. This is possible because our network jointly processes the input images and the resulting 3D pointmaps, thus learning to associate 2D structures with 3D shapes, and having the opportunities of solving multiple minimal problems simultaneously, enabling internal ‘collaboration’ between them.
Our model is trained in a fully-supervised manner using a simple regression loss, leveraging large public datasets for which ground-truth annotations are either synthetically generated , reconstructed from SfM softwares or captured using dedicated sensors . We drift away from the trend of integrating task-specific modules , and instead adopt a fully data-driven strategy based on a generic transformer architecture, not enforcing any geometric constraints at inference, but being able to benefit from powerful pretraining schemes. The network learns strong geometric and shape priors, which are reminiscent of those commonly leveraged in MVS, like shape from texture, shading or contours .
To fuse predictions from multiple images pairs, we revisit bundle adjustment (BA) for the case of pointmaps, hereby achieving full-scale MVS. We introduce a global alignment procedure that, contrary to BA, does not involve minimizing reprojection errors. Instead, we optimize the camera pose and geometry alignment directly in 3D space, which is fast and shows excellent convergence in practice. Our experiments show that the reconstructions are accurate and consistent between views in real-life scenarios with various unknown sensors. We further demonstrate that the same architecture can handle real-life monocular and multi-view reconstruction scenarios seamlessly. Examples of reconstructions are shown in Fig. 1 and in the accompanying video.
In summary, our contributions are fourfold. First, we present the first holistic end-to-end 3D reconstruction pipeline from un-calibrated and un-posed images, that unifies monocular and binocular 3D reconstruction. Second, we introduce the pointmap representation for MVS applications, that enables the network to predict the 3D shape in a canonical frame, while preserving the implicit relationship between pixels and the scene. This effectively drops many constraints of the usual perspective camera formulation. Third, we introduce an optimization procedure to globally align pointmaps in the context of multi-view 3D reconstruction. Our procedure can extract effortlessly all usual intermediary outputs of the classical SfM and MVS pipelines. In a sense, our approach unifies all 3D vision tasks and considerably simplifies over the traditional reconstruction pipeline, making DUSt3R seem simple and easy in comparison. Fourth, we demonstrate promising performance on a range of 3D vision tasks In particular, our all-in-one model achieves state-of-the-art results on monocular and multi-view depth benchmarks, as well as multi-view camera pose estimation.
Related Work
For the sake of space, we summarize here the most related works in 3D vision, and refer the reader to the appendix in Sec. for a more comprehensive review.
Structure-from-Motion (SfM) aims at reconstructing sparse 3D maps while jointly determining camera parameters from a set of images. The traditional pipeline starts from pixel correspondences obtained from keypoint matching between multiple images to determine geometric relationships, followed by bundle adjustment to optimize 3D coordinates and camera parameters jointly. Recently, the SfM pipeline has undergone substantial enhancements, particularly with the incorporation of learning-based techniques into its subprocesses. These improvements encompass advanced feature description , more accurate image matching , featuremetric refinement , and neural bundle adjustment . Despite these advancements, the sequential structure of the SfM pipeline persists, making it vulnerable to noise and errors in each individual component.
MultiView Stereo (MVS) is the task of densely reconstructing visible surfaces, which is achieved via triangulation between multiple viewpoints. In the classical formulation of MVS, all camera parameters are supposed to be provided as inputs. The fully handcrafted , the more recent scene optimization based , or learning based approaches all depend on camera parameter estimates obtained via complex calibration procedures, either during the data acquisition or using Structure-from-Motion approaches for in-the-wild reconstructions. Yet, in real-life scenarios, the inaccuracy of pre-estimated camera parameters can be detrimental for these algorithms to work properly. In this work, we propose instead to directly predict the geometry of visible surfaces without any explicit knowledge of the camera parameters.
Direct RGB-to-3D. Recently, some approaches aiming at directly predicting 3D geometry from a single RGB image have been proposed. Since the problem is by nature ill-posed without introducing additional assumptions, these methods leverage neural networks that learn strong 3D priors from large datasets to solve ambiguities. These methods can be classified into two groups. The first group leverages class-level object priors. For instance, Pavllo et al. propose to learn a model that can fully recover shape, pose, and appearance from a single image, given a large collection of 2D images. While this type of approach is powerful, it does not allow to infer shape on objects from unseen categories. A second group of work, closest to our method, focuses instead on general scenes. These methods systematically build on or re-use existing monocular depth estimation (MDE) networks . Depth maps indeed encode a form of 3D information and, combined with camera intrinsics, can straightforwardly yield pixel-aligned 3D point-clouds. SynSin , for example, performs new viewpoint synthesis from a single image by rendering feature-augmented depthmaps knowing all camera parameters. Without camera intrinsics, one solution is to infer them by exploiting temporal consistency in video frames, either by enforcing a global alignment et al. or by leveraging differentiable rendering with a photometric reconstruction loss . Another way is to explicitly learn to predict camera intrinsics, which enables to perform metric 3D reconstruction from a single image when combined with MDE . All these methods are, however, intrinsically limited by the quality of depth estimates, which arguably is ill-posed for monocular settings.
In contrast, our network processes two viewpoints simultaneously in order to output depthmaps, or rather, pointmaps. In theory, at least, this makes triangulation between rays from different viewpoint possible. Multi-view networks for 3D reconstruction have been proposed in the past. They are essentially based on the idea of building a differentiable SfM pipeline, replicating the traditional pipeline but training it end-to-end . For that, however, ground-truth camera intrinsics are required as input, and the output is generally a depthmap and a relative camera pose . In contrast, our network has a generic architecture and outputs pointmaps, i.e. dense 2D field of 3D points, which handle camera poses implicitly and makes the regression problem much better posed.
Pointmaps. Using a collection of pointmaps as shape representation is quite counter-intuitive for MVS, but its usage is widespread for Visual Localization tasks, either in scene-dependent optimization approaches or scene-agnostic inference methods . Similarly, view-wise modeling is a common theme in monocular 3D reconstruction works and in view synthesis works . The idea being to store the canonical 3D shape in multiple canonical views to work in image space. These approaches usually leverage explicit perspective camera geometry, via rendering of the canonical representation.
Method
Before delving into the details of our method, we introduce below the essential concept of pointmaps.
Network architecture. The architecture of our network is inspired by CroCo , making it straightforward to heavily benefit from CroCo pretraining . As shown in Fig. 2, it is composed of two identical branches (one for each image) comprising each an image encoder, a decoder and a regression head. The two input images are first encoded in a Siamese manner by the same weight-sharing ViT encoder , yielding two token representations and :
The network then reasons over both of them jointly in the decoder. Similarly to CroCo , the decoder is a generic transformer network equipped with cross attention. Each decoder block thus sequentially performs self-attention (each token of a view attends to tokens of the same view), then cross-attention (each token of a view attends to all other tokens of the other view), and finally feeds tokens to a MLP. Importantly, information is constantly shared between the two branches during the decoder pass. This is crucial in order to output properly aligned pointmaps. Namely, each decoder block attends to tokens from the other branch:
for for a decoder with blocks and initialized with encoder tokens and . Here, denotes the -th block in branch , and are the input tokens, with the tokens from the other branch. Finally, in each branch a separate regression head takes the set of decoder tokens and outputs a pointmap and an associated confidence map:
Discussion. The output pointmaps and are regressed up to an unknown scale factor. Also, it should be noted that our generic architecture never explicitly enforces any geometrical constraints. Hence, pointmaps do not necessarily correspond to any physically plausible camera model. Rather, we let the network learn all relevant priors present from the train set, which only contains geometrically consistent pointmaps. Using a generic architecture allows to leverage strong pretraining technique, ultimately surpassing what existing task-specific architectures can achieve. We detail the learning process in the next section.
2 Training Objective
3D Regression loss. Our sole training objective is based on regression in the 3D space. Let us denote the ground-truth pointmaps as and , obtained from Eq. 1 along with two corresponding sets of valid pixels on which the ground-truth is defined. The regression loss for a valid pixel in view is simply defined as the Euclidean distance:
To handle the scale ambiguity between prediction and ground-truth, we normalize the predicted and ground-truth pointmaps by scaling factors and , respectively, which simply represent the average distance of all valid points to the origin:
Confidence-aware loss. In reality, and contrary to our assumption, there are ill-defined 3D points, e.g. in the sky or on translucent objects. More generally, some parts in the image are typically harder to predict that others. We thus jointly learn to predict a score for each pixel which represents the confidence that the network has about this particular pixel. The final training objective is the confidence-weighted regression loss from Eq. 2 over all valid pixels:
where is the confidence score for pixel , and is a hyper-parameter controlling the regularization term . To ensure a strictly positive confidence, we typically define . This has the effect of forcing the network to extrapolate in harder areas, e.g. like those ones covered by a single view. Training network with this objective allows to estimate confidence scores without an explicit supervision. Examples of input image pairs with their corresponding outputs are shown in Fig. 3 and in the appendix in Figs. 4, 5 and 8.
3 Downstream Applications
The rich properties of the output pointmaps allows us to perform various convenient operations with relative ease.
Point matching. Establishing correspondences between pixels of two images can be trivially achieved by nearest neighbor (NN) search in the 3D pointmap space. To minimize errors, we typically retain reciprocal (mutual) correspondences between images and , i.e. we have:
Recovering intrinsics. By definition, the pointmap is expressed in ’s coordinate frame. It is therefore possible to estimate the camera intrinsic parameters by solving a simple optimization problem. In this work, we assume that the principal point is approximately centered and pixel are squares, hence only the focal remains to be estimated:
with and . Fast iterative solvers, e.g. based on the Weiszfeld algorithm , can find the optimal in a few iterations. For the focal of the second camera, the simplest option is to perform the inference for the pair and use above formula with instead of .
Relative pose estimation can be achieved in several fashions. One way is to perform 2D matching and recover intrinsics as described above, then estimate the Epipolar matrix and recover the relative pose . Another, more direct, way is to compare the pointmaps (or, equivalently, ) using Procrustes alignment to get the relative pose :
which can be achieved in closed-form. Procrustes alignment is, unfortunately, sensitive to noise and outliers. A more robust solution is finally to rely on RANSAC with PnP .
Absolute pose estimation, also termed visual localization, can likewise be achieved in several different ways. Let denote the query image and the reference image for which 2D-3D correspondences are available. First, intrinsics for can be estimated from . One possibility consists of obtaining 2D correspondences between and , which in turn yields 2D-3D correspondences for , and then running PnP-RANSAC . Another solution is to get the relative pose between and as described previously. Then, we convert this pose to world coordinate by scaling it appropriately, according to the scale between and the ground-truth pointmap for .
4 Global Alignment
The network presented so far can only handle a pair of images. We now present a fast and simple post-processing optimization for entire scenes that enables the alignment of pointmaps predicted from multiple images into a joint 3D space. This is possible thanks to the rich content of our pointmaps, which encompasses by design two aligned point-clouds and their corresponding pixel-to-3D mapping.
Pairwise graph. Given a set of images for a given scene, we first construct a connectivity graph where images form vertices and each edge indicates that images and shares some visual content. To that aim, we either use existing off-the-shelf image retrieval methods, or we pass all pairs through network (inference takes 40ms on a H100 GPU) and measure their overlap based on the average confidence in both pairs, then we filter out low-confidence pairs.
Here, we abuse notation and write for if . The idea is that, for a given pair , the same rigid transformation should align both pointmaps and with the world-coordinate pointmaps and , since and are by definition both expressed in the same coordinate frame. To avoid the trivial optimum where , we enforce that .
Recovering camera parameters. A straightforward extension to this framework enables to recover all cameras parameters. By simply replacing (i.e. enforcing a standard camera pinhole model as in Eq. 1), we can thus estimate all camera poses , associated intrinsics and depthmaps for .
Discussion. We point out that, contrary to traditional bundle adjustment, this global optimization is fast and simple to perform in practice. Indeed, we are not minimizing 2D reprojection errors, as bundle adjustment normally does, but 3D projection errors. The optimization is carried out using standard gradient descent and typically converges after a few hundred steps, requiring mere seconds on a standard GPU.
Experiments with DUSt3R
Training data. We train our network with a mixture of eight datasets: Habitat , MegaDepth , ARKitScenes , MegaDepth , Static Scenes 3D , Blended MVS , ScanNet++ , CO3D-v2 and Waymo . These datasets feature diverse scenes types: indoor, outdoor, synthetic, real-world, object-centric, etc. When image pairs are not directly provided with the dataset, we extract them based on the method described in . Specifically, we utilize off-the-shelf image retrieval and point matching algorithms to match and verify image pairs. All in all, we extract 8.5M pairs in total.
Training details. During each epoch, we randomly sample an equal number of pairs from each dataset to equalize disparities in dataset sizes. We wish to feed relatively high-resolution images to our network, say 512 pixels in the largest dimension. To mitigate the high cost associated with such input, we train our network sequentially, first on 224224 images and then on larger 512-pixel images. We randomly select the image aspect ratios for each batch (e.g. 16/9, 4/3, etc), so that at test time our network is familiar with different image shapes. We simply crop images to the desired aspect-ratio, and resize so that the largest dimension is 512 pixels.
We use standard data augmentation techniques and training set-up overall. Our network architecture comprises a Vit-Large for the encoder , a ViT-Base for the decoder and a DPT head . We refer to the appendix in Sec. for more details on the training and architecture. Before training, we initialize our network with the weights of an off-the-shelf CroCo pretrained model . Cross-View completion (CroCo) is a recently proposed pretraining paradigm inspired by MAE that has been shown to excel on various downstream 3D vision tasks, and is thus particularly suited to our framework. We ablate in Sec. 4.6 the impact of CroCo pretraining and increase in image resolution.
Evaluation. In the remainder of this section, we benchmark DUSt3R on a representative set of classical 3D vision tasks, each time specifying datasets, metrics and comparing performance with existing state-of-the-art approaches. We emphasize that all results are obtained with the same DUSt3R model (our default model is denoted as ‘DUSt3R 512’, other DUSt3R models serves for the ablations in Section Sec. 4.6), i.e. we never finetune our model on a particular downstream task. During test, all test images are rescaled to 512px while preserving their aspect ratio. Since there may exist different ‘routes’ to extract task-specific outputs from DUSt3R, as described in Sec. 3.3 and Sec. 3.4, we precise each time the employed method.
Qualitative results. DUSt3R yields high-quality dense 3D reconstructions even in challenging situations. We refer the reader to the appendix in Sec. for non-cherrypicked visualizations of pairwise and multi-view reconstructions.
Datasets and metrics. We first evaluate DUSt3R for the task of absolute pose estimation on the 7Scenes and Cambridge Landmarks datasets . 7Scenes contains 7 indoor scenes with RGB-D images from videos and their 6-DOF camera poses. Cambridge-Landmarks contains 6 outdoor scenes with RGB images and their associated camera poses, which are obtained via SfM. We report the median translation and rotation errors in (cm/∘), respectively.
Protocol and results. To compute camera poses in world coordinates, we use DUSt3R as a 2D-2D pixel matcher (see Section 3.3) between a query and the most relevant database images obtained using off-the-shelf image retrieval AP-GeM . In other words, we simply use the raw pointmaps output from without any refinement, where is the query image and is a database image. We use the top 20 retrieved images for Cambridge-Landmarks and top 1 for 7Scenes and leverage the known query intrinsics. For results obtained without using ground-truth intrinsics parameters, refer to the appendix in Sec. .
We compare our results against the state of the art in Table 1 for each scene of the two datasets. Our method obtains comparable accuracy compared to existing approaches, being feature-matching ones or end-to-end learning-based methods , even managing to outperform strong baselines like HLoc in some cases. We believe this to be significant for two reasons. First, DUSt3R was never trained for visual localisation in any way. Second, neither query image nor database images were seen during DUSt3R’s training.
2 Multi-view Pose Estimation
We now evaluate DUSt3R on multi-view relative pose estimation after the global alignment from Sec. 3.4.
Datasets. Following , we use two multi-view datasets, CO3Dv2 and RealEstate10k for the evaluation. CO3Dv2 contains 6 million frames extracted from approximately 37k videos, covering 51 MS-COCO categories. The ground-truth camera poses are annotated using COLMAP from 200 frames in each video. RealEstate10k is an indoor/outdoor dataset with 10 million frames from about 80K video clips on YouTube, the camera poses being obtained by SLAM with bundle adjustment. We follow the protocol introduced in to evaluate DUSt3R on 41 categories from CO3Dv2 and 1.8K video clips from the test set of RealEstate10k. For each sequence, we random select 10 frames and feed all possible 45 pairs to DUSt3R.
Baselines and metrics. We compare DUSt3R pose estimation results, obtained either from PnP-RANSAC or global alignment, against the learning-based RelPose , PoseReg and PoseDiffusion , and structure-based PixSFM , COLMAP+SPSG (COLMAP extended with SuperPoint and SuperGlue ). Similar to , we report the Relative Rotation Accuracy (RRA) and Relative Translation Accuracy (RTA) for each image pair to evaluate the relative pose error and select a threshold to report RTA@ and RRA. Additionally, we calculate the mean Average Accuracy (mAA), defined as the area under the curve accuracy of the angular differences at .
Results. As shown in Table 2, DUSt3R with global alignment achieves the best overall performance on the two datasets and significantly surpasses the state-of-the-art PoseDiffusion . Moreover, DUSt3R with PnP also demonstrates superior performance over both learning and structure-based existing methods. It is worth noting that RealEstate10K results reported for PoseDiffusion are from the model trained on CO3Dv2. Nevertheless, we assert that our comparison is justified considering that RealEstate10K is not used either during DUSt3R’s training. We also report performance with less input views (between 3 and 10) in the appendix (Sec. ), in which case DUSt3R also yields excellent performance on both benchmarks.
3 Monocular Depth
For this monocular task, we simply feed the same input image to the network as . By design, depth prediction is simply the coordinate in the predicted 3D pointmap.
Datasets and metrics. We benchmark DUSt3R on two outdoor (DDAD , KITTI ) and three indoor (NYUv2 , BONN , TUM ) datasets. We compare DUSt3R ’s performance to state-in-the-art methods categorized in supervised, self-supervised and zero-shot settings, this last category corresponding to DUSt3R. We use two metrics commonly used in the monocular depth evaluations : the absolute relative error between target and prediction , , and the prediction threshold accuracy, .
Results. In zero-shot setting, the state of the art is represented by the recent SlowTv . This approach collected a large mixture of curated datasets with urban, natural, synthetic and indoor scenes, and trained one common model. For every dataset in the mixture, camera parameters are known or estimated with COLMAP. As Table 2 shows, DUSt3R adapts well to outdoor and indoor environments. It outperforms the self-supervised baselines and performs on-par with state-of-the-art supervised baselines .
4 Multi-view Depth
We evaluate DUSt3R for the task of multi-view stereo depth estimation. Likewise, we extract depthmaps as the -coordinate of predicted pointmaps. In the case where multiple depthmaps are available for the same image, we rescale all predictions to align them together and aggregate all predictions via a simple averaging weighted by the confidence.
Datasets and metrics. Following , we evaluate it on the DTU , ETH3D , Tanks and Temples , and ScanNet datasets. We report the Absolute Relative Error (rel) and Inlier Ratio with a threshold of 1.03 on each test set and the averages across all test sets. Note that we do not leverage the ground-truth camera parameters and poses nor the ground-truth depth ranges, so our predictions are only valid up to a scale factor. In order to perform quantitative measurements, we thus normalize predictions using the medians of the predicted depths and the ground-truth ones, as advocated in .
Results. We observe in Table 3 that DUSt3R achieves state-of-the-art accuracy on ETH-3D and outperforms most recent state-of-the-art methods overall, even those using ground-truth camera poses. Timewise, our approach is also much faster than the traditional COLMAP pipeline . This showcases the applicability of our method on a large variety of domains, either indoors, outdoors, small scale or large scale scenes, while not having been trained on the test domains, except for the ScanNet test set, since the train split is part of the Habitat dataset.
5 3D Reconstruction
Finally, we measure the quality of our full reconstructions obtained after the global alignment procedure described in Sec. 3.4. We again emphasize that our method is the first one to enable global unconstrained MVS, in the sense that we have no prior knowledge regarding the camera intrinsic and extrinsic parameters. In order to quantify the quality of our reconstructions, we simply align the predictions to the ground-truth coordinate system. This is done by fixing the parameters as constants in Eq. 5. This leads to consistent 3D reconstructions expressed in the coordinate system of the ground-truth.
Datasets and metrics. We evaluate our predictions on the DTU dataset. We apply our network in a zero-shot setting, i.e. we do not finetune on the DTU train set and apply our model as is. In Tab. 4 we report the averaged accuracy, averaged completeness and overall averaged error metrics as provided by the authors of the benchmarks. The accuracy for a point of the reconstructed shape is defined as the smallest Euclidean distance to the ground-truth, and the completeness of a point of the ground-truth as the smallest Euclidean distance to the reconstructed shape. The overall is simply the mean of both previous metrics.
Results. Our method does not reach the accuracy levels of the best methods. In our defense, these methods all leverage GT poses and train specifically on the DTU train set whenever applicable. Furthermore, best results on this task are usually obtained via sub-pixel accurate triangulation, requiring the use of explicit camera parameters, whereas our approach relies on regression, which is known to be less accurate. Yet, without prior knowledge about the cameras, we reach an average accuracy of , with a completeness of , for an overall average distance of . We believe this level of accuracy to be of great use in practice, considering the plug-and-play nature of our approach.
6 Ablations
We ablate the impact of the CroCo pretraining and image resolution on DUSt3R’s performance. We report results in tables Tab. 1, Tab. 2, Tab. 3 for the tasks mentioned above. Overall, the observed consistent improvements suggest the crucial role of pretraining and high resolution in modern data-driven approaches, as also noted by .
Conclusion
We presented a novel paradigm to solve not only 3D reconstruction in-the-wild without prior information about scene nor cameras, but a whole variety of 3D vision tasks as well.
Appendix
This appendix provides additional details and qualitative results of DUSt3R. We first present in Sec. qualitative pairwise predictions of the presented architecture on challenging real-life datasets. This section also contains the description of the video accompanying this material. We then propose an extended related works in Sec. , encompassing a wider range of methodological families and geometric vision tasks. Sec. provides auxiliary ablative results on multi-view pose estimation, that did not fit in the main paper. We then report in Sec. results on an experimental visual localization task, where the camera intrinsics are unknown. Finally, we details the training and data augmentation procedures in Sec. .
Qualitative results
We present some visualization of DUSt3R’s pairwise results in Figs. 4, 5, 6, 7 and 8. Note these scenes were never seen during training and were not cherry-picked. Also, we did not post-process these results, except for filtering out low-confidence points (based on the output confidence) and removing sky regions for the sake of visualization, i.e. these figures accurately represent the raw output of DUSt3R. Overall, the proposed network is able to perform highly accurate 3D reconstruction from just two images. In Fig. 9, we show the output of DUSt3R after the global alignment stage. In this case, the network has processed all pairs of the 4 input images, and outputs 4 spatially consistent pointmaps along with the corresponding camera parameters.
Note that, for the case of image sequences captured with the same camera, we never enforce the fact that camera intrinsics must be identical for every frame, i.e. all intrinsic parameters are optimized independently. This remains true for all results reported in this appendix and in the main paper, e.g. on multi-view pose estimation with the CO3Dv2 and RealEstate10K datasets.
Supplementary Video. We attach to this appendix a video showcasing the different steps of DUSt3R. In the video, we demonstrate dense 3D reconstruction from a small set of raw RGB images, without using any ground-truth camera parameters (i.e. unknown intrinsic and extrinsic parameters). We show that our method can seamlessly handle monocular predictions, and is able to perform reconstruction and camera pose estimation in extreme binocular cases, where the cameras are facing each other. In addition, we show some qualitative reconstructions of rather large scale scenes from the ETH3D dataset .
Extended Related Work
For the sake of exposition, Section 2 of the main paper covered only some (but not all) of the most related works. Because this work covers a large variety of geometric tasks, we complete it in this section with a few equally important topics.
Implicit Camera Models. In our work, we do not explicitly output camera parameters. Likewise, there are several works aiming to express 3D shapes in a canonical space that is not directly related to the input viewpoint. Shapes can be stored as occupancy in regular grids , octree structures , collections of parametric surface elements , point clouds encoders , free-form deformation of template meshes or per-view depthmaps . While these approaches arguably perform classification and not actual 3D reconstruction , all-in-all, they work only in very constrained setups, usually on ShapeNet and have trouble generalizing to natural scenes with non object-centric views . The question of how to express a complex scene with several object instances in a single canonical frame had yet to be answered: in this work, we also express the reconstruction in a canonical reference frame, but thanks to our scene representation (pointmaps), we still preserve a relationship between image pixels and the 3D space, and we are thus able to perform 3D reconstruction consistently.
Dense Visual SLAM. In visual SLAM, early works on dense 3D reconstruction and ego-motion estimation utilized active depth sensors . Recent works on dense visual SLAM from RGB video stream are able to produce high-quality depth maps and camera trajectories , but they inherit the traditional limitations of SLAM, e.g. noisy predictions, drifts and outliers in the pixel correspondences. To make the 3D reconstruction more robust, R3D3 jointly leverages jointly multi-camera constraints and monocular depth cues. Most recently, GO-SLAM proposed real-time global pose optimization by considering the complete history of input frames and continuously aligning all poses that enables instantaneous loop closures and correction of global structure. Still, all SLAM methods assume that the input consists of a sequence of closely related images, e.g. with identical intrinsics, nearby camera poses and small illumination variations. In comparison, our approach handles completely unconstrained image collections.
3D reconstruction from implicit models has undergone significant advancements, largely fueled by the integration of neural networks . Earlier approaches utilize Multi-Layer Perceptron (MLP) to generate continuous surface outputs with only posed RGB images. Innovations like Nerf and its follow-ups have pioneered density-based volume rendering to represent scenes as continuous 5D functions for both occupancy and color, showing exceptional ability in synthesizing novel views of complex scenes. To handle large-scale scenes, recent approaches introduce geometry priors to the implicit model, leading to much more detailed reconstructions. In contrast to the implicit 3D reconstruction, our work focuses on the explicit 3D reconstruction and showcases that the proposed DUSt3R can not only have detailed 3D reconstruction but also provide rich geometry for multiple downstream 3D tasks.
RGB-pairs-to-3D takes its roots in two-view geometry and is considered as a stand-alone task or an intermediate step towards the multi-view reconstruction. This process typically involves estimating a dense depth map and determining the relative camera pose from two different views. Recent learning-based approaches formulate this problem either as pose and monocular depth regression or pose and stereo matching . The ultimate goal is to achieve 3D reconstruction from the predicted geometry . In addition to reconstruction tasks, learning from two views also gives an advance in unsupervised pretraining; the recently proposed CroCo introduces a pretext task of cross-view completion from a large set of image pair to learn 3D geometry from unlabeled data and to apply this learned implicit representation to various downstream 3D vision tasks. Our method draws inspiration from the CroCo pipeline, but diverges in its application. Instead of focusing on model pretraining, our approach leverages this pipeline to directly generate 3D pointmaps from the image pair. In this context, the depth map and camera poses are only by-products in our pipeline.
Multi-view Pose Estimation
We include additional results for the multi-view pose estimation task from the main paper (in Sec. 4.2). Namely, we compute the pose accuracy for a smaller number of input images (they are randomly selected from the entire test sequences). Tab. 5 reports our performance and compares with the state of the art. Numbers for state-of-the-art methods are borrowed from the recent PoseDiffusion paper’s tables and plots, hence some numbers are only approximate. Our method consistently outperforms all other methods on the CO3Dv2 dataset by a large margin, even for small number of frames. As can be observed in Fig. 8 and in the attached video, DUSt3R handles opposite viewpoints (i.e. nearly 180∘ apart) seemingly without much troubles. In the end, DUSt3R obtains relatively stable performance, regardless of the number of input views. When comparing with PoseDiffusion on RealEstate10K, we report performances with and without training on the same dataset. Note that DUSt3R’s training data include a small subset of CO3Dv2 (we used 50 sequences for each category, i.e. less than 7% of the full training set) but no data from RealEstate10K whatsoever.
An example of reconstruction on RealEstate10K is shown in Fig. 9. Our network outputs a consistent pointcloud despite wide baseline viewpoint changes between the first and last pairs of frames.
Visual localization
We include additional results of visual localization on the 7-scenes and Cambridge-Landmarks datasets . Namely, we experiment with a scenario where the focal parameter of the querying camera is unknown. In this case, we feed the query image and a database image into DUSt3R, and get an un-scaled 3D reconstruction. We then scale the resulting pointmap according to the ground-truth pointmap of the database image, and extract the pose as described in Sec. 3.3 of the main paper. Tab. 6 shows that this method performs reasonably well on the 7-scenes dataset, where the median translation error is on the order of a few centimeters. On the Cambridge-Landmarks dataset, however, we obtain considerably larger errors. After inspection, we find that the ground-truth database pointmaps are sparse, which prevents any reliable scaling of our reconstruction. On the contrary, 7-scenes provides dense ground-truth pointmaps. We conclude that further work is necessary for ”in-the-wild” visual-localization with unknown intrinsics.
Training details
Relation between depthmaps and pointmaps. As a result, the depth value at pixel in image can be recovered as
Therefore, all depthmaps displayed in the main paper and this appendix are straightforwardly extracted from DUSt3R’s output as and for images and , respectively.
Dataset mixture. DUSt3R is trained with a mixture of eight datasets: Habitat , ARKitScenes , MegaDepth , Static Scenes 3D , Blended MVS , ScanNet++ , CO3Dv2 and Waymo . These datasets feature diverse scene types: indoor, outdoor, synthetic, real-world, object-centric, etc. Table 8 shows the number of extracted pairs in each datasets, which amounts to 8.5M in total.
Data augmentation. We use standard data augmentation techniques, namely random color jittering and random center crops, the latter being a form of focal augmentation. Indeed, some datasets are captured using a single or a small number of camera devices, hence many images have practically the same intrinsic parameters. Centered random cropping thus helps in generating more focals. Crops are centered so that the principal point is always centered in the training pairs. At test time, we observe little impact on the results when the principal point is not exactly centered. During training, we also systematically feed each training pair as well as its inversion to help generalization. Naturally, tokens from these two pairs do not interact.
.2 Training hyperparameters
We report the detailed hyperparameter settings we use for training DUSt3R in Table 7.