UniDepth: Universal Monocular Metric Depth Estimation

Luigi Piccinelli, Yung-Hsu Yang, Christos Sakaridis, Mattia Segu, Siyuan Li, Luc Van Gool, Fisher Yu

Introduction

The precise pixel-wise depth estimation is crucial to understanding the geometric scene structure, with applications in 3D modeling , robotics , and autonomous vehicles . However, delivering reliable metric scaled depth outputs is necessary to perform 3D reconstruction effectively, thus motivating the challenging and inherently ill-posed task of Monocular Metric Depth Estimation (MMDE).

While existing MMDE methods have demonstrated remarkable accuracy across different benchmarks, they require training and testing on datasets with similar camera intrinsics and scene scales. Moreover, the training datasets typically have a limited size and contain little diversity in scenes and cameras. These characteristics result in poor generalization to real-world inference scenarios , where images are captured in uncontrolled, arbitrarily structured environments and cameras with arbitrary intrinsics.

Only a few methods have addressed the challenging task of generalizable MMDE. However, these methods assume controlled setups at test time, including camera intrinsics. While this assumption simplifies the task, it has two notable drawbacks. Firstly, it does not address the full application spectrum, e.g. in-the-wild video processing and crowd-sourced image analysis. Secondly, the inherent camera parameter noise is directly injected into the model, leading to large inaccuracies in the high-noise case.

In this work, we address the more demanding task of generalizable MMDE without any reliance on additional external information, such as camera parameters, thus defining the universal MMDE task. Our approach, named UniDepth, is the first that attempts to solve this challenging task without restrictions on scene composition and setup and distinguishes itself through its general and adaptable nature. Unlike existing methods, UniDepth delivers metric 3D predictions for any scene solely from a single image, waiving the need for extra information about scene or camera. Furthermore, UniDepth flexibly allows for the incorporation of additional camera information at test time.

Our design introduces a camera module that outputs a non-parametric, i.e. dense camera representation, serving as the prompt to the depth module. However, relying only on this single additional module clearly results in challenges related to training stability and scale ambiguity. We propose an effective pseudo-spherical representation of the output space to disentangle the camera and depth dimensions of this space. This representation employs azimuth and elevation angle components for the camera and a radial component for the depth, forming a perfect orthogonal space between the camera plane and the depth axis. Moreover, the camera components are embedded through Laplace spherical harmonic encoding. Figure 1 depicts our camera self-prompting mechanism and the output space. Additionally, we introduce a geometric invariance loss to enhance the robustness of depth estimation. The underlying idea is that the camera-conditioned depth features from two views of the same image should exhibit reciprocal consistency. In particular, we sample two geometric augmentations, creating a pair of different views for each training image, thus simulating different apparent cameras for the original scene.

Our overall contribution is the first universal MMDE method, UniDepth, that predicts a point in metric 3D space for each pixel without any input other than a single image. In particular, first, we design a promptable camera module, an architectural component that learns a dense camera representation and allows for non-parametric camera conditioning. Second, we propose a pseudo-spherical representation of the output space, thus solving the intertwined nature of camera and depth prediction. In addition, we introduce a geometric invariance loss to disentangle the camera information from the underlying 3D geometry of the scene. Moreover, we extensively test UniDepth and re-evaluate seven MMDE State-of-the-Art (SotA) methods on ten different datasets in a fair and comparable zero-shot setup to lay the ground for the generalized MMDE task. Owing to its design, UniDepth consistently sets the new state of the art even compared with non-zero-shot methods, ranking first in the competitive official KITTI Depth Prediction Benchmark.

Related Work

Metric and Scale-agnostic Depth Estimation. It is crucial to distinguish Monocular Metric Depth Estimation (MMDE) from scale-agnostic, namely up-to-a-scale, monocular depth estimation. MMDE SotA approaches typically confine training and testing to the same domain. However, challenges arise, such as overfitting to the training scenario leading to considerable performance drops in the presence of minor domain gaps, often overlooked in benchmarks like NYU-Depthv2 (NYU) and KITTI . On the other hand, scale-agnostic depth methods, including MiDaS , OmniData , and LeReS , show robust generalization by training on extensive datasets. Their limitation lies in the absence of a metric output, hindering practical usage in downstream applications.

General Monocular Metric Depth Estimation. Recent efforts focus on developing MMDE models for general depth prediction across diverse domains. These models often leverage camera awareness, either by directly incorporating external camera parameters into computations or by normalizing the shape or output depth based on intrinsic properties, as seen in .

However, these generalizable MMDE methods often adopt specific strategies to enhance performance, e.g. geometric pretraining or dataset-specific prior like reshaping . In addition, these methods assume access to noiseless camera intrinsics both at training and test time, also limiting their applicability to pinhole camera models. Additionally, SotA methods depend on a predefined backprojection operation, blurring the distinction between learning depth and the 3D scene. In contrast, our approach aims to overcome these limitations, presenting a more demanding perspective, e.g. universal MMDE. Universal MMDE involves directly predicting the 3D scene from the input image without any additional information other than the latter. Notably, we do not require any additional prior information at test time, such as access to camera information.

UniDepth

MMDE SotA methods typically assume access to the camera intrinsics, thus blurring the line between pure depth estimation and actual 3D estimation. In contrast, UniDepth aims to create a universal MMDE model deployable in diverse scenarios without relying on any other external information, such as camera intrinsic, thus leading to 3D space estimation by design. However, attempting to directly predict 3D points from a single image without a proper internal representation neglects geometric prior knowledge, i.e. perspective geometry, burdening the learning process with re-learning laws of perspective projection from data.

Sec. 3.1 introduces a pseudo-spherical representation of the output space to inherently disentangle camera rays’ angles from depth. In addition, our preliminary studies indicate that depth prediction clearly benefits from prior information on the acquisition sensor, leading to the introduction of a self-prompting camera operation in Sec. 3.2. Further disentanglement at the level of internal depth features is achieved through a geometric invariance loss, outlined in Sec. 3.3. This loss ensures depth features remain invariant when conditioned on the bootstrapped camera predictions, promoting robust camera-aware depth predictions. The overall architecture and the resulting optimization induced by the combination of design choices are detailed in Sec. 3.4.

The general purpose nature of our MMDE method requires inferring both depth and camera intrinsics to make 3D predictions based only on imagery observations. We design the 3D output space presenting a natural disentanglement of the two sub-tasks, namely depth estimation and camera calibration. In particular, we exploit the pseudo-spherical representation where the basis is defined by azimuth, elevation, and log-depth, i.e. (θ\theta,ϕ\phi,zlog⁡z_{\log}), in contrast to the Cartesian representation (xx,yy,zz). The strength of the proposed pseudo-spherical representation lies in the decoupling of camera (θ\theta,ϕ\phi) and depth (zlog⁡z_{\log}) components, ensuring their orthogonality by design, in contrast to the entanglement present in Cartesian representation.

where Pml\mathcal{P}^{l}_{m} is the associated Legendre polynomial of degree ll and order mm, and αml\alpha^{l}_{m} is a normalizing constant. In particular, the spherical harmonics on the unit sphere form an orthogonal basis of the spherical manifold and preserve inner products. The total number of harmonics utilized is 8181, resulting from capping the degree ll to 8. SHE is utilized as a mathematic sounder choice compared to, e.g. the Fourier Transform, to produce the camera embeddings.

2 Self-Promptable Camera

The camera module plays a crucial role in the final 3D predictions since its angular dense output accounts for two dimensions of the output space, namely azimuth and elevation. Most importantly, these embeddings prompt the depth module to ensure a bootstrapped prior knowledge of the input scene’s global depth scale. The prompting is fundamental to avoid mode collapse in the scene scale and to alleviate the depth module from the burden of predicting depth from scratch as the scale is already modeled by camera output.

Nonetheless, the internal representation of the camera module is based on a pinhole parameterization, namely via focal length (fxf_{x}, fyf_{y}) and principal point (cxc_{x}, cyc_{y}). The four tokens conceptually corresponding to the intrinsics are then projected to scalar values, i.e., Δfx\Delta f_{x}, Δfy\Delta f_{y}, Δcx\Delta c_{x}, Δcy\Delta c_{y}. However, they do not directly represent the camera parameters, but the multiplicative residuals to a pinhole camera initialization, namely H2\frac{H}{2} for y-components and W2\frac{W}{2} for x-components, leading to fx=ΔfxW2f_{x}=\frac{\Delta f_{x}W}{2}, fy=ΔfyH2f_{y}=\frac{\Delta f_{y}H}{2}, cx=ΔcxW2c_{x}=\frac{\Delta c_{x}W}{2}, cy=ΔcyH2c_{y}=\frac{\Delta c_{y}H}{2}, leading to invariance towards input image sizes.

Figure 3 illustrates one of the main benefits of our camera module. In particular, in high-noise intrinsics or camera-agnostic scenarios, UniDepth can bootstrap the camera prediction, thus displaying total noise insensitivity. However, we can substitute the camera module output to improve 3D reconstruction peak performance if any external dense camera representation is provided. This adaptability enhances the model’s versatility, allowing it to operate seamlessly in diverse setups. Moreover, Figure 3 suggests that training with noisy self-prompts enhances the robustness of UniDepth to noisier external intrinsics if given at test time.

3 Geometric Invariance Loss

The spatial locations from the same scene captured by different cameras should correspond when the depth module is conditioned on the specific camera. To this end, we propose a geometric invariance loss to enforce the consistency of camera-prompted depth features of the same scene from different acquisition sensors. In particular, consistency is enforced on features extracted from identical 3D locations.

For each image, we perform NN distinct geometrical augmentations, denoted as {Ti}i=1N\{\mathcal{T}_{i}\}_{i=1}^{N}, with N=2N=2 in our experiments. This operation involves involves sampling a rescaling factor r∼2Ur\sim 2^{\mathcal{U}_{}} and a relative translation on the xx-axis t∼U[−0.1,0.1]t\sim\mathcal{U}_{[-0.1,0.1]}, then cropping it to the network’s input shape. This is analogous to sampling a pair of images from the same scene and extrinsic parameters but captured by different cameras. Let Ci\mathbf{C}_{i} and Di∣Ei\mathbf{D_{i}|E_{i}} describe the predicted camera representation and camera-prompted depth features, respectively, corresponding to augmentation Ti\mathcal{T}_{i}. It is evident that the camera representations differ when two diverse geometric augmentations are applied, i.e., Ci≠Cj\mathbf{C}_{i}\neq\mathbf{C}_{j} if Ti≠Tj\mathcal{T}_{i}\neq\mathcal{T}_{j}. Therefore, the geometric invariance loss can be expressed as

4 Network Design

Optimization. The optimization process is guided by a re-formulation of the Mean Squared Error (MSE) loss in the final 3D output space (θ\theta,ϕ\phi,zlog⁡z_{\log}) from Sec. 3.1 as:

The loss defined here serves as a motivation for the designed output representation. Specifically, employing a Cartesian representation and applying the loss directly to the output space would result in backpropagation through (xx, yy), and zlog⁡z_{\log} errors. However, xx and yy components are derived as rx⋅zr_{x}\cdot z and ry⋅zr_{y}\cdot z as detailed in Sec. 3.1. Consequently, the gradients of camera components, expressed by (rxr_{x}, ryr_{y}), and of depth become intertwined, leading to suboptimal optimization as discussed in Sec. 4.3.

Experiments

In-domain training datasets. The training dataset utilized is the ensemble of Argoverse2 , Waymo , DrivingStereo , Cityscapes , BDD100K , Mapillary-PSD , A2D2 , ScanNet , and Taskonomy . The resulting dataset amounts roughly to 3M real-world images with different cameras and domains, compared to, e.g. Metric3D and ZeroDepth which exploit 8M and 17M training images, respectively.

Zero-shot testing datasets. We evaluate the generalizability of the compared models by testing them on ten datasets not seen during training. More precisely, each method is tested on validation splits from SUN-RGBD without NYU split, Diode Indoor , IBims-1 , VOID HAMMER , ETH-3D , nuScenes , and DDAD with split proposed in and evaluated with official masks. Also, UniDepth and the models from are zero-shot-tested on NYU-Depth V2 and KITTI . In particular, KITTI testing is performed on the corrected Eigen-split test set with the Garg evaluation mask , while NYU testing uses the evaluation mask from .

Implementation Details. UniDepth is implemented in PyTorch and CUDA . For training, we use the AdamW optimizer (β1=0.9\beta_{1}=0.9, β2=0.999\beta_{2}=0.999) with an initial learning rate of 0.00010.0001. The learning rate is divided by a factor of 10 for the backbone weights for every experiment and weight decay is set to 0.10.1. As the learning rate scheduler, we exploit Cosine Annealing to one-tenth starting from 30% of the training. We run 1M optimization iterations with a batch size of 128, each training dataset is uniformly represented in each batch. In particular, we sample 64 images and then we sample two different augmented views of the same image for consistency loss. The augmentations include both geometric and appearance (random brightness, gamma, saturation, hue shift, and grayscale) augmentations. ViT-L backbone is initialized with weights from DINO-pre-trained models, and ConvNext-L is ImageNet -pre-trained. The required training time amounts to roughly 12 days on 8 NVIDIA A100. Ablations are conducted with three different seeds and for 100k training iterations, using a randomly sampled subset with a size equal to 20% of the original training set.

2 Comparison with the State of the Art

For the sake of fair comparison, we provide in Table 4 a comparison between Metric3D, iDisc, and UniDepth where the latter two are retrained on a strict subset of Metric3D’s data, namely accounting for one-quarter of the original Metric3D dataset, with same framework detailed in Sec. 4.1. The results are two-fold: they demonstrate how UniDepth still surpasses Metric3D with a subsplit of the training set, and how MMDE SotA methods designed for single-domain can not fully exploit the training diversity. Qualitative results in Fig. 4 emphasize how the method excels in capturing the overall scale and scene complexity in a zero-shot setup.

3 Ablation Study

The importance of each component introduced in UniDepth in Sec. 3 is evaluated by ablating the method in Table 5. All ablations exploit the predicted camera representation, if not stated otherwise. The first distinction involves the Oracle model, which operates under ideal conditions with known camera information during training and testing, addressing a task similar to . On the other hand, Baseline is a straightforward encoder-decoder implementation with a (xx,yy,zz) output, as outlined at the beginning of Sec. 3, while Baseline++ exploits the proposed pseudo-spherical representation. Modules’ architectures are consistent across experiments. The In-Domain column reflects testing on validation splits of training domains, while Out-of-Domain corresponds to zero-shot testing, as detailed in Sec. 4.1. Notably, In-Domain results exhibit a higher degree of homogeneity compared to Out-of-Domain, which is noisier yet more informative for gauging expected performances in downstream applications and in-the-wild deployment.

Architecture. The Oracle model demonstrates more robust scale-dependent performance during zero-shot testing compared to the Full model, highlighting how the proposed task is inherently more demanding. The Baseline model illustrates an approach to the problem without utilizing external information and lacking a proper design for both internal and output space. This approach yields markedly inferior results for both In-Domain and Out-of-Domain scenarios in terms of depth and 3D reconstruction metrics.

Optimization and Output Representation. All ablations employ the same loss LλMSE\mathcal{L}_{\lambda MSE}, but across different output spaces. In row 4, a Cartesian output space is used instead of a pseudo-spherical from Sec. 3.1, which results in substantially inferior performance due to the respective intertwined formulation of camera and depth output spaces. The Baseline (row 8) also employs a Cartesian representation, but the negative impact of this choice is less pronounced in this model because of the absence of a camera module. More specifically, the decoder of Baseline is not conditioned on inaccurate prior camera and scale information as in row 4. Moreover, row 9 corresponds to Baseline with pseudo-spherical representation. Comparison between row 8 and row 9 shows that when predicting directly the 3D outputs, the choice of the output representation is still relevant in defining a better internal representation and optimization. Row 5 demonstrates the positive impact of the geometric invariance loss. This loss contributes to enhanced in-domain and out-of-domain performance by promoting the invariance of depth features to appearance variations owing to different camera intrinsics. Furthermore, stopping the gradient from propagating from the Camera Module to the Encoder (row 7), as described in Sec. 3.4, proves particularly beneficial in avoiding scale and camera overfitting in zero-shot testing, and stabilizes the training. The more stable training is obtained by limiting the dominant effect that camera supervision has on the gradient of the Encoder weights compared to depth supervision.

Conclusion

In this work, we propose UniDepth to predict metric 3D points in diverse scenes relying solely on a single input image. Through meticulous ablation studies, we systematically address the challenges inherent in universal MMDE tasks, underscoring the pivotal contributions of our work. The designed self-prompting camera allows camera-free test time application and renders the model more robust against camera noise. The introduced pseudo-spherical output space representation adequately disentangles the camera and depth of the optimization process. Furthermore, the proposed geometric invariance loss effectively ensures camera-aware depth consistency. Extensive validations unequivocally exhibit how UniDepth sets the new state of the art across multiple benchmarks in a zero-shot regime, even surpassing in-domain trained methods. This attests to the robustness and efficacy of our model and, most importantly, outlines its potential to propel the field of MMDE to new frontiers.

Acknowledgment. This work is funded by Toyota Motor Europe via the research project TRACE-Zürich.

References

A Results

KITTI benchmark . Table 6 clearly shows the compelling performance of UniDepth on the official KITTI private test set. Results of the latest published methods are reported. The table is fetched from the official KITTI leaderboard for depth prediction. In particular, UniDepth ranks first in the KITTI benchmark at the time of submission among all methods, published and not.

KITTI Eigen-split and NYUv2-Depth. For the sake of completeness, we report the “standard” metrics results in Table 7 and Table 8 on KITTI Eigen-split and NYU validation set, respectively. It is worth noting that the typical metrics δ2\delta_{2} and, especially, δ3\delta_{3} are saturated, thus not informative. Therefore, we advocate our choice of not reporting them in the main paper and prefer to report δ0.5\delta_{0.5}. Moreover, we suggest in future works the use of the area under the curve of the δ\delta metrics as a more informative and comprehensive metric, instead of the values at fixed thresholds, i.e. {1.25i}i=13\{1.25^{i}\}^{3}_{i=1}.

B Ablations

B.2 Alternative pseudo-spherical representation

Furthermore, we ablate our camera prompting with respect to CAMConvs in Table 11

C Datasets

Details of training and testing datasets are presented in Table 12. The training datasets are processed in a way that the interval between two consecutive RGB and GT depth frames is not smaller than one second. We do not apply any post-processing apart from the aforementioned subsampling. The total amount of training samples accounts for 3’743’000 samples. SUN-RGBD validation set involves also NYU test set. Therefore, we removed the samples corresponding to NYU test set to avoid any overlap between test sets. As per standard practice, KITTI Eigen-split corresponds to the corrected and accumulated GT depth maps with 45 images with inaccurate GT discarded from the original 697 images.

C.2 Diode Indoor ground-truth correction

Diode ground-truth depth is not perfectly accurate on boundaries, in particular, a simple inspection shows how depth in boundaries presents low values, but greater than zero. These artifacts present in the GT affect the validation pipeline and results. Therefore, we design a simple image processing algorithm, outlined in Algorithm 1, that, first, detects the aforementioned boundary artifacts and, second, masks the depth in the corresponding neighborhoods. Thanks to masking those boundaries, the corresponding regions are ignored during validation.

D Model Complexity

Table 13 displays the parameters and inference complexity of UniDepth and other SotA methods. UniDepth with ViT-L backbone is comparable to ZoeDepth in terms of efficiency and model parameters; however UniDepth surpasses it in terms of performance as stated in Sec. 4. Metric3D displays an improved efficiency due to the fully convolutional and relatively low dimensionality designed in the decoder. It is worth highlighting how ZeroDepth presents a low efficiency although based on ResNet-18, we argue that this is due to the expensive full-resolution cross-attention in the decoder. The last two rows in Table 13 analyze separately the complexity of the single Camera and Depth Module. The Camera Module is a lightweight component accounting for 13.4M parameters. On the other hand, the Depth Module amounts to more than half of the total latency, despite the limited memory consumption. The Depth Module’s high latency is due to the several (6) self-attention layers in the decoder.

E Network Architecture

Depth Module. The depth latents are initialized as the average of the features F\mathbf{F} along the BB dimension. Then, the latents are conditioned on the original feature tensor F\mathbf{F} via one cross-attention layer where two projections of F\mathbf{F} account for keys and values and L\mathbf{L} as queries. In addition, one MLP is applied, seamlessly as in the Camera Module. Furthermore, the depth features are conditioned on the camera prompts E\mathbf{E} with one additional cross-attention layer, where keys and values are two projections of camera embeddings E\mathbf{E}, and one MLP as above. The features are decoded in three consecutive stages. The first stage applies three self-attention layers with E\mathbf{E} as positional encoding. The features are then processed with one ConvNext layer, upsampled by a factor of two, and the channels are halved. The second and third stages are similar, although the second stage presents two self-attention layers and the third only one. In the second and third stages, MLP’s hidden channel dimension is sequentially halved, too, from the initial aforementioned value of 2048. Each stage’s output is projected to a dimension one. Therefore, the three output maps are interpolated to a common shape, i.e. (H2\frac{H}{2}, W2\frac{W}{2}), and pixel-wise averaged. The final log-depth output Zlog⁡\mathbf{Z}_{\log} is obtained by upsampling the obtained tensor to the input shape (HH, WW). The final depth is element-wise exponentiation of Zlog⁡\mathbf{Z}_{\log}.

F Visualization

We provide here twenty more qualitative comparisons, two for each zero-shot test set: KITTI, NYU, Diode, ETH3D in Fig. 5, DDAD, NuScenes, SUN-RGBD, IBims-1 in Fig. 6, and Fig. 7 displays VOID and HAMMER. The error maps are shown after applying median-based rescaling. The rescaling was deemed necessary to avoid some of the error maps being completely red and not informative. Due to sparsity, DDAD and Nuscenes GT and error maps are dilated by a factor of 5, leading to visible GT depth and error maps.