Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao

Introduction

The field of computer vision and natural language processing is currently experiencing a revolution with the emergence of “foundation models” that demonstrate strong zero-/few-shot performance in various downstream scenarios . These successes primarily rely on large-scale training data that can effectively cover the data distribution. Monocular Depth Estimation (MDE), which is a fundamental problem with broad applications in robotics , autonomous driving , virtual reality , etc., also requires a foundation model to estimate depth information from a single image. However, this has been underexplored due to the difficulty of building datasets with tens of millions of depth labels. MiDaS made a pioneering study along this direction by training an MDE model on a collection of mixed labeled datasets. Despite demonstrating a certain level of zero-shot ability, MiDaS is limited by its data coverage, thus suffering disastrous performance in some scenarios.

In this work, our goal is to build a foundation model for MDE capable of producing high-quality depth information for any images under any circumstances. We approach this target from the perspective of dataset scaling-up. Traditionally, depth datasets are created mainly by acquiring depth data from sensors , stereo matching , or SfM , which is costly, time-consuming, or even intractable in particular situations. We instead, for the first time, pay attention to large-scale unlabeled data. Compared with stereo images or labeled images from depth sensors, our used monocular unlabeled images exhibit three advantages: (i) (simple and cheap to acquire) Monocular images exist almost everywhere, thus they are easy to collect, without requiring specialized devices. (ii) (diverse) Monocular images can cover a broader range of scenes, which are critical to the model generalization ability and scalability. (iii) (easy to annotate) We can simply use a pre-trained MDE model to assign depth labels for unlabeled images, which only takes a feedforward step. More than efficient, this also produces denser depth maps than LiDAR and omits the computationally intensive stereo matching process.

We design a data engine to automatically generate depth annotations for unlabeled images, enabling data scaling-up to arbitrary scale. It collects 62M diverse and informative images from eight public large-scale datasets, e.g., SA-1B , Open Images , and BDD100K . We use their raw unlabeled images without any forms of labels. Then, in order to provide a reliable annotation tool for our unlabeled images, we collect 1.5M labeled images from six public datasets to train an initial MDE model. The unlabeled images are then automatically annotated and jointly learned with labeled images in a self-training manner .

Despite all the aforementioned advantages of monocular unlabeled images, it is indeed not trivial to make positive use of such large-scale unlabeled images , especially in the case of sufficient labeled images and strong pre-training models. In our preliminary attempts, directly combining labeled and pseudo labeled images failed to improve the baseline of solely using labeled images. We conjecture that, the additional knowledge acquired in such a naive self-teaching manner is rather limited. To address the dilemma, we propose to challenge the student model with a more difficult optimization target when learning the pseudo labels. The student model is enforced to seek extra visual knowledge and learn robust representations under various strong perturbations to better handle unseen images.

Furthermore, there have been some works demonstrating the benefit of an auxiliary semantic segmentation task for MDE. We also follow this research line, aiming to equip our model with better high-level scene understanding capability. However, we observed when an MDE model is already powerful enough, it is hard for such an auxiliary task to bring further gains. We speculate that it is due to severe loss in semantic information when decoding an image into a discrete class space. Therefore, considering the excellent performance of DINOv2 in semantic-related tasks, we propose to maintain the rich semantic priors from it with a simple feature alignment loss. This not only enhances the MDE performance, but also yields a multi-task encoder for both middle-level and high-level perception tasks.

Our contributions are summarized as follows:

We highlight the value of data scaling-up of massive, cheap, and diverse unlabeled images for MDE.

We point out a key practice in jointly training large-scale labeled and unlabeled images. Instead of learning raw unlabeled images directly, we challenge the model with a harder optimization target for extra knowledge.

We propose to inherit rich semantic priors from pre-trained encoders for better scene understanding, rather than using an auxiliary semantic segmentation task.

Our model exhibits stronger zero-shot capability than MiDaS-BEiTL-512{}_{\textrm{L-512}} . Further, fine-tuned with metric depth, it outperforms ZoeDepth significantly.

Related Work

Monocular depth estimation (MDE). Early works primarily relied on handcrafted features and traditional computer vision techniques. They were limited by their reliance on explicit depth cues and struggled to handle complex scenes with occlusions and textureless regions.

Deep learning-based methods have revolutionized monocular depth estimation by effectively learning depth representations from delicately annotated datasets . Eigen et al. first proposed a multi-scale fusion network to regress the depth. Following this, many works consistently improve the depth estimation accuracy by carefully designing the regression task as a classification task , introducing more priors , and better objective functions , etc. Despite the promising performance, they are hard to generalize to unseen domains.

Zero-shot depth estimation. Our work belongs to this research line. We aim to train an MDE model with a diverse training set and thus can predict the depth for any given image. Some pioneering works explored this direction by collecting more training images, but their supervision is very sparse and is only enforced on limited pairs of points.

To enable effective multi-dataset joint training, a milestone work MiDaS utilizes an affine-invariant loss to ignore the potentially different depth scales and shifts across varying datasets. Thus, MiDaS provides relative depth information. Recently, some works take a step further to estimate the metric depth. However, in our practice, we observe such methods exhibit poorer generalization ability than MiDaS, especially its latest version . Besides, as demonstrated by ZoeDepth , a strong relative depth estimation model can also work well in generalizable metric depth estimation by fine-tuning with metric depth information. Therefore, we still follow MiDaS in relative depth estimation, but further strengthen it by highlighting the value of large-scale monocular unlabeled images.

Leveraging unlabeled data. This belongs to the research area of semi-supervised learning , which is popular with various applications . However, existing works typically assume only limited images are available. They rarely consider the challenging but realistic scenario where there are already sufficient labeled images but also larger-scale unlabeled images. We take this challenging direction for zero-shot MDE. We demonstrate that unlabeled images can significantly enhance the data coverage and thus improve model generalization and robustness.

Depth Anything

Our work utilizes both labeled and unlabeled images to facilitate better monocular depth estimation (MDE). Formally, the labeled and unlabeled sets are denoted as Dl={(xi,di)}i=1M\mathcal{D}^{l}=\{(x_{i},d_{i})\}_{i=1}^{M} and Du={ui}i=1N\mathcal{D}^{u}=\{u_{i}\}_{i=1}^{N} respectively. We aim to learn a teacher model TT from Dl\mathcal{D}^{l}. Then, we utilize TT to assign pseudo depth labels for Du\mathcal{D}^{u}. Finally, we train a student model SS on the combination of labeled set and pseudo labeled set. A brief illustration is provided in Figure 2.

This process is similar to the training of MiDaS . However, since MiDaS did not release its code, we first reproduced it. Concretely, the depth value is first transformed into the disparity space by d=1/td=1/t and then normalized to 0∼\sim1 on each depth map. To enable multi-dataset joint training, we adopt the affine-invariant loss to ignore the unknown scale and shift of each sample:

where di∗d^{*}_{i} and did_{i} are the prediction and ground truth, respectively. And ρ\rho is the affine-invariant mean absolute error loss: ρ(di∗,di)=∣d^i∗−d^i∣\rho(d^{*}_{i},d_{i})=|\hat{d}^{*}_{i}-\hat{d}_{i}|, where d^i∗\hat{d}^{*}_{i} and d^i\hat{d}_{i} are the scaled and shifted versions of the prediction di∗d^{*}_{i} and ground truth did_{i}:

where t(d)t(d) and s(d)s(d) are used to align the prediction and ground truth to have zero translation and unit scale:

To obtain a robust monocular depth estimation model, we collect 1.5M labeled images from 6 public datasets. Details of these datasets are listed in Table 1. We use fewer labeled datasets than MiDaS v3.1 (12 training datasets), because 1) we do not use NYUv2 and KITTI datasets to ensure zero-shot evaluation on them, 2) some datasets are not available (anymore), e.g., Movies and WSVD , and 3) some datasets exhibit poor quality, e.g., RedWeb (also low resolution) . Despite using fewer labeled images, our easy-to-acquire and diverse unlabeled images will comprehend the data coverage and greatly enhance the model generalization ability and robustness.

Furthermore, to strengthen the teacher model TT learned from these labeled images, we adopt the DINOv2 pre-trained weights to initialize our encoder. In practice, we apply a pre-trained semantic segmentation model to detect the sky region, and set its disparity value as 0 (farthest).

2 Unleashing the Power of Unlabeled Images

This is the main point of our work. Distinguished from prior works that laboriously construct diverse labeled datasets, we highlight the value of unlabeled images in enhancing the data coverage. Nowadays, we can practically build a diverse and large-scale unlabeled set from the Internet or public datasets of various tasks. Also, we can effortlessly obtain the dense depth map of monocular unlabeled images simply by forwarding them to a pre-trained well-performed MDE model. This is much more convenient and efficient than performing stereo matching or SfM reconstruction for stereo images or videos. We select eight large-scale public datasets as our unlabeled sources for their diverse scenes. They contain more than 62M images in total. The details are provided in the bottom half of Table 1.

Technically, given the previously obtained MDE teacher model TT, we make predictions on the unlabeled set Du\mathcal{D}^{u} to obtain a pseudo labeled set D^u\hat{\mathcal{D}}^{u}:

With the combination set Dl∪Du^\mathcal{D}^{l}\cup\hat{\mathcal{D}^{u}} of labeled images and pseudo labeled images, we train a student model SS on it. Following prior works , instead of fine-tuning SS from TT, we re-initialize SS for better performance.

Unfortunately, in our pilot studies, we failed to gain improvements with such a self-training pipeline, which indeed contradicts the observations when there are only a few labeled images . We conjecture that, with already sufficient labeled images in our case, the extra knowledge acquired from additional unlabeled images is rather limited. Especially considering the teacher and student share the same pre-training and architecture, they tend to make similar correct or false predictions on the unlabeled set Du\mathcal{D}^{u}, even without the explicit self-training procedure.

To address the dilemma, we propose to challenge the student with a more difficult optimization target for additional visual knowledge on unlabeled images. We inject strong perturbations to unlabeled images during training. It compels our student model to actively seek extra visual knowledge and acquire invariant representations from these unlabeled images. These advantages help our model deal with the open world more robustly. We introduce two forms of perturbations: one is strong color distortions, including color jittering and Gaussian blurring, and the other is strong spatial distortion, which is CutMix . Despite the simplicity, the two modifications make our large-scale unlabeled images significantly improve the baseline of labeled images.

We provide more details about CutMix. It was originally proposed for image classification, and is rarely explored in monocular depth estimation. We first interpolate a random pair of unlabeled images uau_{a} and ubu_{b} spatially:

where MM is a binary mask with a rectangle region set as 1.

The unlabeled loss Lu\mathcal{L}_{u} is obtained by first computing affine-invariant losses in valid regions defined by MM and 1−M1-M, respectively:

We use CutMix with 50% probability. The unlabeled images for CutMix are already strongly distorted in color, but the unlabeled images fed into the teacher model TT for pseudo labeling are clean, without any distortions.

3 Semantic-Assisted Perception

There exist some works improving depth estimation with an auxiliary semantic segmentation task. We believe that arming our depth estimation model with such high-level semantic-related information is beneficial. Besides, in our specific context of leveraging unlabeled images, these auxiliary supervision signals from other tasks can also combat the potential noise in our pseudo depth label.

Therefore, we made an initial attempt by carefully assigning semantic segmentation labels to our unlabeled images with a combination of RAM + GroundingDINO + HQ-SAM models. After post-processing, this yields a class space containing 4K classes. In the joint-training stage, the model is enforced to produce both depth and segmentation predictions with a shared encoder and two individual decoders. Unfortunately, after trial and error, we still could not boost the performance of the original MDE model. We speculated that, decoding an image into a discrete class space indeed loses too much semantic information. The limited information in these semantic masks is hard to further boost our depth model, especially when our depth model has established very competitive results.

Therefore, we aim to seek more informative semantic signals to serve as auxiliary supervision for our depth estimation task. We are greatly astonished by the strong performance of DINOv2 models in semantic-related tasks, e.g., image retrieval and semantic segmentation, even with frozen weights without any fine-tuning. Motivated by these clues, we propose to transfer its strong semantic capability to our depth model with an auxiliary feature alignment loss. The feature space is high-dimensional and continuous, thus containing richer semantic information than discrete masks. The feature alignment loss is formulated as:

where cos⁡(⋅,⋅)\cos(\cdot,\cdot) measures the cosine similarity between two feature vectors. ff is the feature extracted by the depth model SS, while f′f^{\prime} is the feature from a frozen DINOv2 encoder. We do not follow some works to project the online feature ff into a new space for alignment, because a randomly initialized projector makes the large alignment loss dominate the overall loss in the early stage.

Another key point in feature alignment is that, semantic encoders like DINOv2 tend to produce similar features for different parts of an object, e.g., car front and rear. In depth estimation, however, different parts or even pixels within the same part, can be of varying depth. Thus, it is not beneficial to exhaustively enforce our depth model to produce exactly the same features as the frozen encoder.

To solve this issue, we set a tolerance margin α\alpha for the feature alignment. If the cosine similarity of fif_{i} and fi′f^{\prime}_{i} has surpassed α\alpha, this pixel will not be considered in our Lfeat\mathcal{L}_{feat}. This allows our method to enjoy both the semantic-aware representation from DINOv2 and the part-level discriminative representation from depth supervision. As a side effect, our produced encoder not only performs well in downstream MDE datasets, but also achieves strong results in the semantic segmentation task. It also indicates the potential of our encoder to serve as a universal multi-task encoder for both middle-level and high-level perception tasks.

Finally, our overall loss is an average combination of the three losses Ll\mathcal{L}_{l}, Lu\mathcal{L}_{u}, and Lfeat\mathcal{L}_{feat}.

Experiment

We adopt the DINOv2 encoder for feature extraction. Following MiDaS , we use the DPT decoder for depth regression. All labeled datasets are simply combined together without re-sampling. In the first stage, we train a teacher model on labeled images for 20 epochs. In the second stage of joint training, we train a student model to sweep across all unlabeled images for one time. The unlabeled images are annotated by a best-performed teacher model with a ViT-L encoder. The ratio of labeled and unlabeled images is set as 1:2 in each batch. In both stages, the base learning rate of the pre-trained encoder is set as 5e-6, while the randomly initialized decoder uses a 10×\times larger learning rate. We use the AdamW optimizer and decay the learning rate with a linear schedule. We only apply horizontal flipping as our data augmentation for labeled images. The tolerance margin α\alpha for feature alignment loss is set as 0.15. For more details, please refer to our appendix.

2 Zero-Shot Relative Depth Estimation

As aforementioned, this work aims to provide accurate depth estimation for any image. Therefore, we comprehensively validate the zero-shot depth estimation capability of our Depth Anything model on six representative unseen datasets: KITTI , NYUv2 , Sintel , DDAD , ETH3D , and DIODE . We compare with the best DPT-BEiTL-512{}_{\textrm{L-512}} model from the latest MiDaS v3.1 , which uses more labeled images than us. As shown in Table 2, both with a ViT-L encoder, our Depth Anything surpasses the strongest MiDaS model tremendously across extensive scenes in terms of both the AbsRel (absolute relative error: ∣d∗−d∣/d|d^{*}-d|/d) and δ1\delta_{1} (percentage of max⁡(d∗/d,d/d∗)<1.25\max(d^{*}/d,d/d^{*})<1.25) metrics. For example, when tested on the well-known autonomous driving dataset DDAD , we improve the AbsRel (↓\downarrow) from 0.251 →\rightarrow 0.230 and improve the δ1\delta_{1} (↑\uparrow) from 0.766 →\rightarrow 0.789.

Besides, our ViT-B model is already clearly superior to the MiDaS based on a much larger ViT-L. Moreover, our ViT-S model, whose scale is less than 1/10 of the MiDaS model, even outperforms MiDaS on several unseen datasets, including Sintel, DDAD, and ETH3D. The performance advantage of these small-scale models demonstrates their great potential in computationally-constrained scenarios.

It is also worth noting that, on the most widely used MDE benchmarks KITTI and NYUv2, although MiDaS v3.1 uses the corresponding training images (not zero-shot anymore), our Depth Anything is still evidently superior to it without training with any KITTI or NYUv2 images, e.g., 0.127 vs. 0.076 in AbsRel and 0.850 vs. 0.947 in δ1\delta_{1} on KITTI.

3 Fine-tuned to Metric Depth Estimation

Apart from the impressive performance in zero-shot relative depth estimation, we further examine our Depth Anything model as a promising weight initialization for downstream metric depth estimation. We initialize the encoder of downstream MDE models with our pre-trained encoder parameters and leave the decoder randomly initialized. The model is fine-tuned with correponding metric depth information. In this part, we use our ViT-L encoder for fine-tuning.

We examine two representative scenarios: 1) in-domain metric depth estimation, where the model is trained and evaluated on the same domain (Section 4.3.1), and 2) zero-shot metric depth estimation, where the model is trained on one domain, e.g., NYUv2 , but evaluated in different domains, e.g., SUN RGB-D (Section 4.3.2).

As shown in Table 3 of NYUv2 , our model outperforms the previous best method VPD remarkably, improving the δ1\delta_{1} (↑\uparrow) from 0.964 →\rightarrow 0.984 and AbsRel (↓\downarrow) from 0.069 to 0.056. Similar improvements can be observed in Table 4 of the KITTI dataset . We improve the δ1\delta_{1} (↑\uparrow) on KITTI from 0.978 →\rightarrow 0.982. It is worth noting that we adopt the ZoeDepth framework for this scenario with a relatively basic depth model, and we believe our results can be further enhanced if equipped with more advanced architectures.

3.2 Zero-Shot Metric Depth Estimation

We follow ZoeDepth to conduct zero-shot metric depth estimation. ZoeDepth fine-tunes the MiDaS pre-trained encoder with metric depth information from NYUv2 (for indoor scenes) or KITTI (for outdoor scenes). Therefore, we simply replace the MiDaS encoder with our better Depth Anything encoder, leaving other components unchanged. As shown in Table 5, across a wide range of unseen datasets of indoor and outdoor scenes, our Depth Anything results in a better metric depth estimation model than the original ZoeDepth based on MiDaS.

4 Fine-tuned to Semantic Segmentation

In our method, we design our MDE model to inherit the rich semantic priors from a pre-trained encoder via a simple feature alignment constraint. Here, we examine the semantic capability of our MDE encoder. Specifically, we fine-tune our MDE encoder to downstream semantic segmentation datasets. As exhibited in Table 7 of the Cityscapes dataset , our encoder from large-scale MDE training (86.2 mIoU) is superior to existing encoders from large-scale ImageNet-21K pre-training, e.g., Swin-L (84.3) and ConvNeXt-XL (84.6). Similar observations hold on the ADE20K dataset in Table 8. We improve the previous best result from 58.3 →\rightarrow 59.4.

We hope to highlight that, witnessing the superiority of our pre-trained encoder on both monocular depth estimation and semantic segmentation tasks, we believe it has great potential to serve as a generic multi-task encoder for both middle-level and high-level visual perception systems.

5 Ablation Studies

Unless otherwise specified, we use the ViT-L encoder for our ablation studies here.

Zero-shot transferring of each training dataset. In Table 6, we provide the zero-shot transferring performance of each training dataset, which means that we train a relative MDE model on one training set and evaluate it on the six unseen datasets. With these results, we hope to offer more insights for future works that similarly aim to build a general monocular depth estimation system. Among the six training datasets, HRWSI fuels our model with the strongest generalization ability, even though it only contains 20K images. This indicates the data diversity counts a lot, which is well aligned with our motivation to utilize unlabeled images. Some labeled datasets may not perform very well, e.g., MegaDepth , however, it has its own preferences that are not reflected in these six test datasets. For example, we find models trained with MegaDepth data are specialized at estimating the distance of ultra-remote buildings (Figure 1), which will be very beneficial for aerial vehicles.

Effectiveness of 1) challenging the student model when learning unlabeled images, and 2) semantic constraint. As shown in Table 9, simply adding unlabeled images with pseudo labels does not necessarily bring gains to our model, since the labeled images are already sufficient. However, with strong perturbations (S\mathcal{S}) applied to unlabeled images during re-training, the student model is challenged to seek additional visual knowledge and learn more robust representations. Consequently, the large-scale unlabeled images enhance the model generalization ability significantly.

Moreover, with our used semantic constraint Lfeat\mathcal{L}_{feat}, the power of unlabeled images can be further amplified for the depth estimation task. More importantly, as emphasized in Section 4.4, this auxiliary constraint also enables our trained encoder to serve as a key component in a multi-task visual system for both middle-level and high-level perception.

Comparison with MiDaS trained encoder in downstream tasks. Our Depth Anything model has exhibited stronger zero-shot capability than MiDaS . Here, we further compare our trained encoder with MiDaS v3.1 trained encoder in terms of the downstream fine-tuning performance. As demonstrated in Table 10, on both the downstream depth estimation task and semantic segmentation task, our produced encoder outperforms the MiDaS encoder remarkably, e.g., 0.951 vs. 0.984 in the δ1\delta_{1} metric on NYUv2, and 52.4 vs. 59.4 in the mIoU metric on ADE20K.

Comparison with DINOv2 in downstream tasks. We have demonstrated the superiority of our trained encoder when fine-tuned to downstream tasks. Since our finally produced encoder (from large-scale MDE training) is fine-tuned from DINOv2 , we compare our encoder with the original DINOv2 encoder in Table 11. It can be observed that our encoder performs better than the original DINOv2 encoder in both the downstream metric depth estimation task and semantic segmentation task. Although the DINOv2 weight has provided a very strong initialization (also much better than the MiDaS encoder as reported in Table 10), our large-scale and high-quality MDE training can further enhance it impressively in downstream transferring performance.

6 Qualitative Results

We visualize our model predictions on the six unseen datasets in Figure 3. Our model is robust to test images from various domains. In addition, we compare our model with MiDaS in Figure 4. We also attempt to synthesis new images conditioned on the predicted depth maps with ControlNet . Our model produces more accurate depth estimation than MiDaS, as well as better synthesis results, although the ControlNet is trained with MiDaS depth. For more accurate synthesis, we have also re-trained a better depth-conditioned ControlNet based on our Depth Anything, aiming to provide better control signals for image synthesis and video editing. Please refer to our project page or the following supplementary material for more qualitative results,

Conclusion

In this work, we present Depth Anything, a highly practical solution to robust monocular depth estimation. Different from prior arts, we especially highlight the value of cheap and diverse unlabeled images. We design two simple yet highly effective strategies to fully exploit their value: 1) posing a more challenging optimization target when learning unlabeled images, and 2) preserving rich semantic priors from pre-trained models. As a result, our Depth Anything model exhibits excellent zero-shot depth estimation ability, and also serves as a promising initialization for downstream metric depth estimation and semantic segmentation tasks.

More Implementation Details

We resize the shorter side of all images to 518 and keep the original aspect ratio. All images are cropped to 518×\times518 during training. During inference, we do not crop images and only ensure both sides are multipliers of 14, since the pre-defined patch size of DINOv2 encoders is 14. Evaluation is performed at the original resolution by interpolating the prediction. Following MiDaS , in zero-shot evaluation, the scale and shift of our prediction are manually aligned with the ground truth.

When fine-tuning our pre-trained encoder to metric depth estimation, we adopt the ZoeDepth codebase . We merely replace the original MiDaS-based encoder with our stronger Depth Anything encoder, with a few hyper-parameters modified. Concretely, the training resolution is 392×\times518 on NYUv2 and 384×\times768 on KITTI to match the patch size of our encoder. The encoder learning rate is set as 1/50 of the learning rate of the randomly initialized decoder, which is much smaller than the 1/10 adopted for MiDaS encoder, due to our strong initialization. The batch size is 16 and the model is trained for 5 epochs.

When fine-tuning our pre-trained encoder to semantic segmentation, we use the MMSegmentation codebase . The training resolution is set as 896×\times896 on both ADE20K and Cityscapes . The encoder learning rate is set as 3e-6 and the decoder learning rate is 10×\times larger. We use Mask2Former as our semantic segmentation model. The model is trained for 160K iterations on ADE20K and 80K iterations on Cityscapes both with batch size 16, without any COCO or Mapillary pre-training. Other training configurations are the same as the original codebase.

More Ablation Studies

All ablation studies here are conducted on the ViT-S model.

The necessity of tolerance margin for feature alignment. As shown in Table 12, the gap between the tolerance margin of 0 and 0.15 or 0.30 clearly demonstrates the necessity of this design (mean AbsRel: 0.188 vs. 0.175).

Applying feature alignment to labeled data. Previously, we enforce the feature alignment loss Lfeat\mathcal{L}_{feat} on unlabeled data. Indeed, it is technically feasible to also apply this constraint to labeled data. In Table 13, apart from applying Lfeat\mathcal{L}_{feat} on unlabeled data, we explore to apply it to labeled data. We find that adding this auxiliary optimization target to labeled data is not beneficial to our baseline that does not involve any feature alignment (their mean AbsRel values are almost the same: 0.180 vs. 0.179). We conjecture that this is because the labeled data has relatively higher-quality depth annotations. The involvement of semantic loss may interfere with the learning of these informative manual labels. In comparison, our pseudo labels are noisier and less informative. Therefore, introducing the auxiliary constraint to unlabeled data can combat the noise in pseudo depth labels, as well as arm our model with semantic capability.

Limitations and Future Works

Currently, the largest model size is only constrained to ViT-Large . Therefore, in the future, we plan to further scale up the model size from ViT-Large to ViT-Giant, which is also well pre-trained by DINOv2 . We can train a more powerful teacher model with the larger model, producing more accurate pseudo labels for smaller models to learn, e.g., ViT-L and ViT-B. Furthermore, to facilitate real-world applications, we believe the widely adopted 512×\times512 training resolution is not enough. We plan to re-train our model on a larger resolution of 700+ or even 1000+.

More Qualitative Results

Please refer to the following pages for comprehensive qualitative results on six unseen test sets (Figure 5 for KITTI , Figure 6 for NYUv2 , Figure 7 for Sintel , Figure 8 for DDAD , Figure 9 for ETH3D , and Figure 10 for DIODE ). We compare our model with the strongest MiDaS model , i.e., DPT-BEiTL-512{}_{\textrm{L-512}}. Our model exhibits higher depth estimation accuracy and stronger robustness.

References