Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
Bingxin Ke, Anton Obukhov, Shengyu Huang, Nando Metzger, Rodrigo Caye Daudt, Konrad Schindler
Introduction
Monocular depth estimation aims to transform a photographic image into a depth map, i.e., regress a range value for every pixel. The task arises whenever the 3D scene structure is needed, and no direct range or stereo measurements are available. Clearly, undoing the projection from the 3D world to a 2D image is a geometrically ill-posed problem and can only be solved with the help of prior knowledge, such as typical object shapes and sizes, likely scene layouts, occlusion patterns, etc. In other words, monocular depth implicitly requires scene understanding, and it is no coincidence that the advent of deep learning brought about a leap in performance. Following the seminal work of Eigen et al. , depth estimation is nowadays cast as neural image-to-image translation and learned in a supervised fashion using collections of paired and well-aligned RGB and depth images. Early methods of this type were limited to a narrow domain defined by their training data, often indoor or driving scenes. More recently, there has been a quest to train generic depth estimators that can be either used off-the-shelf across a broad range of scenes or fine-tuned to a specific application scenario with a small amount of data. These models generally follow the strategy first employed by MiDAS to achieve generality, namely to train a high-capacity model with data sampled from many different RGB-D datasets (respectively, domains). The latest developments include moving from convolutional encoder-decoder networks to increasingly large and powerful vision transformers , and training on more and more data and with additional surrogate tasks to amass even more knowledge about the visual world, and consequently to produce better depth maps. Importantly, visual cues for depth depend not only on the scene content but also on the (generally unknown) camera intrinsics . For general in-the-wild depth estimation, it is often preferred to estimate affine-invariant depth (i.e., depth values up to a global offset and scale), which can also be determined without objects with known sizes that could serve as “scale bars”.
The intuition behind our work is the following: Modern image diffusion models have been trained on internet-scale image collections specifically to generate high-quality images across a wide array of domains . If the cornerstone of monocular depth estimation is indeed a comprehensive, encyclopedic representation of the visual world, then it should be possible to derive a broadly applicable depth estimator from a pretrained diffusion model. In this paper, we set out to explore this option and develop Marigold, a latent diffusion model (LDM) based on Stable Diffusion , along with a fine-tuning protocol to adapt the model for depth estimation. The key to unlocking the potential of a pretrained diffusion model is to keep its latent space intact. We find this can be done efficiently by modifying and fine-tuning only the denoising U-Net. Turning Stable Diffusion into Marigold requires only synthetic RGB-D data (in our case, the Hypersim and Virtual KITTI datasets) and a few GPU days on a single consumer graphics card. Empowered by the underlying diffusion prior of natural images, Marigold exhibits excellent zero-shot generalization: Without ever having seen real depth maps, it attains state-of-the-art performance on several real datasets. To summarize, our contributions are:
A simple and resource-efficient fine-tuning protocol to convert a pretrained LDM image generator into an image-conditional depth estimator;
Marigold, a state-of-the-art, versatile monocular depth estimation module that offers excellent performance across a wide variety of natural images.
Related Work
At the technical level, monocular depth estimation is a dense, structured regression task. The pioneering work of Eigen et al. introduced a multi-scale network and showed that the result can be converted to metric depth for a dataset recorded with a single sensor. Successive improvements have come from various fronts, including ordinal regression , planar guidance maps , neural conditional random fields , vision transformers , a piecewise planarity prior , first-order variational constraints and variational autoencoders . Some authors treat depth estimation as a combined regression-classification task, using various binning strategies like AdaBins or BinsFormer to discretize depth range. A notable recent trend involves training generative models, especially diffusion models for monocular depth estimation . Recently, a few works like Metric3D and ZeroDepth revisited the depth estimation by explicitly feeding camera intrinsics as additional input.
Estimating depth “in the wild” refers to methods that are successful across a wide range of (possibly unfamiliar) settings, a particularly challenging task. It has been addressed by constructing large and diverse depth datasets and designing algorithms to handle that diversity. Depth-in-the-wild (DIW) was perhaps the earliest work to introduce an uncontrolled dataset and to predict relative (ordinal) depth for it. OASIS introduced relative depth and normals to better generalize across scenes. However, relative depth predictions (depth ordering) are of limited use for many downstream tasks, which has led several authors to consider affine-invariant depth. In that setting, depth is estimated up to an unknown (global) offset and scale. It offers a viable compromise between the ordinal and metric cases: on the one hand, it can handle general scenes consisting of unfamiliar objects; on the other hand, depth differences between different objects or scene parts are still geometrically meaningful relative to each other. MegaDepth and DiverseDepth utilize large internet photo collections to train models that can adapt to unseen data, while MiDaS achieves generality by training on a mixture of multiple datasets. The step from CNNs to vision transformers has further boosted performance, as evidenced by DPT (MiDaS v3) and Omnidata . LeReS proposed a two-stage framework that first predicts affine-invariant depth, then upgrades it to metric depth by estimating the shift and focal length (for which it uses a separate training set with 350k samples). HDN introduced multi-scale depth normalization to improve the prediction details and smoothness further. While this enables the depth estimator to handle images captured with different known cameras, it does not include the true in-the-wild setting, where the camera intrinsics of the test images are unknown. Our method addresses affine-invariant depth estimation but doesn’t focus on compiling an extensive, annotated training dataset. Rather, we utilize the broader image priors in image LDMs and apply fine-tuning.
2 Diffusion Models
Denoising Diffusion Probabilistic Models (DDPMs) emerged as a powerful class of generative models. They learn to reverse a diffusion process that progressively degrades images with Gaussian noise so that they can draw samples from the data distribution by applying the reverse process to random noise. This idea was extended to DDIMs , which provide a non-Markovian shortcut for the diffusion process. Conditional diffusion models are an extension of DDPMs that ingest additional information on which the output is then conditioned, similar to cGAN and cVAE . Conditioning can take various forms, including text , other images , or semantic maps .
In the realm of text-based image generation, Rombach et al. have trained a diffusion model on the large-scale image and text dataset LAION-5B and demonstrated image synthesis with previously unattainable quality. The cornerstone of their approach is a latent diffusion model (LDM), where the denoising process is run in an efficient latent space, drastically reducing the complexity of the learned mapping. Such models distill internet-scale image sets into model weights, thereby developing a rich scene understanding prior, which we harness for monocular depth estimation.
3 Diffusion for Monocular Depth Estimation
Several methods have tried to use DDPMs for metric depth estimation. The DDP approach proposes an architecture to encode the image but decode a depth map and has obtained state-of-the-art results on the KITTI dataset. DiffusionDepth performs diffusion in the latent space, conditioned on image features extracted with a SwinTransformer. DepthGen extends a multi-task diffusion model to metric depth prediction, including handling noisy ground truth. Its successor DDVM emphasizes pretraining on synthetic and real data for enhanced depth estimation. Finally, VPD employs a pretrained Stable Diffusion model as an image feature extractor with additional text input.
Our approach advances beyond these methods, which perform well but only in their specific training domains. We explore the potential of pretrained LDMs for single-image depth estimation across diverse, real-world settings.
4 Foundation Models
Vision foundation models (VFMs) are large neural networks trained on internet-scale data. The extreme scaling leads to the emergence of high-level visual understanding, such that the model can then be used as is or fine-tuned to a wide range of downstream tasks with minimal effort . Prompt tuning methods can efficiently adapt VFMs towards dedicated scenarios by designing suitable prompts. Feature adaptation methods can further pivot VFMs towards different tasks. For example, VPD shows the potential to extract features from a pre-trained text-to-image model for (domain-specific) depth estimation. Direct tuning enables more flexible adaptation, not only for few-shot customization scenarios like DreamBooth but also for object detection, as in 3DiffTection .
As we show in this paper, Marigold can be interpreted as an instance of this type of tuning, where StableDiffusion plays the role of the foundation model. With as few as 74k synthetic depth samples, we obtain state-of-the-art on multiple datasets of real images. Beyond those datasets, our model exhibits strong in-the-wild performance (cf. Fig. 1).
Method
In the forward process, which starts at from the conditional distribution, Gaussian noise is gradually added at levels to obtain noisy samples as
where , and is the variance schedule of a process with steps. In the reverse process, the conditional denoising model parameterized with learned parameters gradually removes noise from to obtain .
At training time, parameters are updated by taking a data pair from the training set, noising with sampled noise at a random timestep , computing the noise estimate and minimizing one of the denoising diffusion objective functions. The canonical standard noise objective is given as follows :
At inference time, is reconstructed starting from a normally-distributed variable , by iteratively applying the learned denoiser .
Unlike diffusion models that work directly on the data, latent diffusion models perform diffusion steps in a low-dimensional latent space, offering computational efficiency and suitability for high-resolution image generation . The latent space is constructed in the bottleneck of a variational autoencoder (VAE) trained independently of the denoiser to enable latent space compression and perceptual alignment with the data space. To translate our formulation into the latent space, for a given depth map , the corresponding latent code is given by the encoder : . Given a depth latent code, a depth map can be recovered with the decoder : . The conditioning image is also naturally translated into the latent space as . The denoiser is henceforth trained in the latent space: . The adapted inference procedure involves one extra step – the decoder reconstructing the data from the estimated clean latent : .
2 Network Architecture
One of our main objectives is training efficiency since diffusion models are often extremely resource-intensive to train. Therefore, we base our model on a pretrained text-to-image LDM (Stable Diffusion v2 ), which has learned very good image priors from LAION-5B . With minimal changes to the model components, we turn it into an image-conditioned depth estimator. Fig. 2 contains an overview of the proposed fine-tuning procedure.
Depth encoder and decoder. We take the frozen VAE to encode both the image and its corresponding depth map into a latent space for training our conditional denoiser. Given that the encoder, which is designed for 3-channel (RGB) inputs, receives a single-channel depth map, we replicate the depth map into three channels to simulate an RGB image. At this point, the data range of the depth data plays a significant role in enabling affine-invariance. We discuss our normalization approach in Sec. 3.3. We verified that without any modification of the VAE or the latent space structure, the depth map can be reconstructed from the encoded latent code with a negligible error, i.e., . At inference time, the depth latent code is decoded once at the end of diffusion, and the average of three channels is taken as the predicted depth map.
Adapted denoising U-Net. To implement the conditioning of the latent denoiser on input image , we concatenate the image and depth latent codes into a single input along the feature dimension. The input channels of the latent denoiser are then doubled to accommodate the expanded input . To prevent inflation of activations magnitude of the first layer and keep the pre-trained structure as faithfully as possible, we duplicate the weight tensor of the input layer and divide its values by two.
3 Fine-Tuning Protocol
Affine-invariant depth normalization. For the ground truth depth maps , we implement a linear normalization such that the depth primarily falls in the value range $$, to match the designed input value range of the VAE. Such normalization serves two purposes. First, it is the convention for working with the original Stable Diffusion VAE. Second, it enforces a canonical affine-invariant depth representation independent of the data statistics – any scene must be bounded by near and far planes with extreme depth values. The normalization is achieved through an affine transformation computed as
where and correspond to the and percentiles of individual depth maps. This normalization allows Marigold to focus on pure affine-invariant depth estimation.
Training on synthetic data. Real depth datasets suffer from missing depth values caused by the physical constraints of the capture rig and the physical properties of the sensors. Specifically, the disparity between cameras and reflective surfaces diverting LiDAR laser beams are inevitable sources of ground truth noise and missing pixels . In contrast to prior work that utilized diverse real datasets to achieve generalization , we train exclusively with synthetic depth datasets. As with the depth normalization rationale, this decision has two objective reasons. First, synthetic depth is inherently dense and complete, meaning that every pixel has a valid ground truth depth value, allowing us to feed such data into the VAE, which can not handle data with invalid pixels. Second, synthetic depth is the cleanest possible form of depth, which is guaranteed by the rendering pipeline. If our assumption about the possibility of fine-tuning a generalizable depth estimation from a text-to-image LDM is correct, then synthetic depth gives the cleanest set of examples and reduces noise in gradient updates during the short fine-tuning protocol. Thus, the remaining concern is the sufficient diversity or domain gaps between synthetic and real data, which sometimes limits generalization ability. As demonstrated in our experiments, our choice of synthetic datasets leads to impressive zero-shot transfer.
Previous works have explored deviations from the original DDPM formulations, such as non-Gaussian noise or non-Markovian schedule shortcuts . Our proposed setting and the fine-tuning protocol outlined above are permissive to changes to the noise schedule at the fine-tuning stage. We identified a combination of multi-resolution noise and an annealed schedule to converge faster and substantially improve performance over the standard DDPM formulation. The multi-resolution noise is composed by superimposing several random Gaussian noise images of different scales, all upsampled to the U-Net input resolution. The proposed annealed schedule interpolates between the multi-resolution noise at and standard Gaussian noise at .
4 Inference
Latent diffusion denoising. The overall inference pipeline is presented in Fig. 3. We encode the input image into the latent space, initialize depth latent as standard Gaussian noise, and progressively denoise it with the same schedule as during fine-tuning. We empirically find that initializing with standard Gaussian noise gives better results than with multi-resolution noise, although the model is trained on the latter. We follow DDIM’s approach to perform non-Markovian sampling with re-spaced steps for accelerated inference. The final depth map is decoded from the latent code using the VAE decoder and postprocessed by averaging channels.
Test-time ensembling. The stochastic nature of the inference pipeline leads to varying predictions depending on the initialization noise in . Capitalizing on that, we propose the following test-time ensembling scheme, capable of combining multiple inference passes over the same input. For each input sample, we can run inference times. To aggregate these affine-invariant depth predictions , we jointly estimate the corresponding scale and shift , relative to some canonical scale and range, in an iterative manner. The proposed objective minimizes the distances between each pair of scaled and shifted predictions , where . In each optimization step, we calculate the merged depth map by the taking pixel-wise median . An extra regularization term , is added to prevent collapse to the trivial solution and enforce the unit scale of . Thus, the objective function can be written as follows:
where the binominal coefficient represents the number of possible combinations of image pairs from images. After the iterative optimization for spatial alignment, the merged depth is taken as our ensembled prediction. Note that this ensembling step requires no ground truth for aligning independent predictions. This scheme enables a flexible trade-off between computation efficiency and prediction quality by choosing accordingly.
Experiments
We implement Marigold using PyTorch and utilize Stable Diffusion v2 as our backbone, following the original pre-training setup with a v-objective . We disable text conditioning and perform steps outlined in Sec. 3.2. During training, we apply the DDPM noise scheduler with 1000 diffusion steps. At inference time, we apply DDIM scheduler and only sample 50 steps. For the final prediction, we aggregate results from 10 inference runs with varying starting noise. Training our method takes 18K iterations using a batch size of 32. To fit one GPU, we accumulate gradients for 16 steps. We use the Adam optimizer with a learning rate of . Additionally, we apply random horizontal flipping augmentation to the training data. Training our method to convergence takes approximately 2.5 days on a single Nvidia RTX 4090 GPU card.
2 Evaluation
Training datasets. We train Marigold on two synthetic datasets covering both indoor and outdoor scenes. Hypersim is a photorealistic dataset with 461 indoor scenes. We use the official split with around 54K samples from 365 scenes for training. Incomplete samples are filtered out. RGB images and depth maps are resized to size. Depth is normalized with the dataset statistics. Additionally, we transform the original distances relative to the focal point into conventional depth values relative to the focal plane. The second dataset, Virtual KITTI is a synthetic street-scene dataset featuring 5 scenes under varying conditions like weather or camera perspectives. Four scenes containing a total of around 20K samples are used for training. We crop the images to the KITTI benchmark resolution and set the far plane to 80 meters.
Evaluation datasets. We evaluate Marigold on 5 real datasets that are not seen during training. NYUv2 and ScanNet are both indoor scene datasets captured with an RGB-D Kinect sensor. For NYUv2, we utilize the designated test split, comprising a total of 654 images. In the case of the ScanNet dataset, we randomly sampled 800 images from the 312 official validation scenes for testing. KITTI is a street-scene dataset with sparse metric depth captured by a LiDAR sensor. We employ the Eigen test split made of 652 images. ETH3D and DIODE are two high-resolution datasets, both featuring depth maps derived from LiDAR sensor measurements. For ETH3D, we incorporate all 454 samples with available ground truth depth maps. For DIODE, we use the entire validation split, which encompasses 325 indoor samples and 446 outdoor samples.
Evaluation protocol. Following the protocol of affine-invariant depth evaluation , we first align the estimated merged prediction to the ground truth with the least squares fitting. This step gives us the absolute aligned depth map in the same units as the ground truth. Next, we apply two widely recognized metrics for assessing quality of depth estimation. The first is Absolute Mean Relative Error (AbsRel), calculated as: , where is the total number of pixels. The second metric, accuracy, measures the proportion of pixels satisfying .
Comparison with other methods. We compare Marigold to six baselines, each claiming zero-shot generalization. DiverseDepth , LeReS and HDN estimate affine-invariant depth maps, while MiDaS , DPT , and Omnidata produce affine-invariant disparities. As shown in Tab. 1, Marigold outperforms prior art in most cases and secures the highest overall ranking. Despite being trained solely on synthetic depth datasets, the model can well generalize to a wide range of real scenes. This successful adaptation of diffusion-based image generation models toward depth estimation confirms our initial hypothesis that a comprehensive representation of the visual world is the cornerstone of monocular depth estimation. It also shows that our fine-tuning protocol was successful in adapting Stable Diffusion for this task without unlearning such visual priors.
For a visual assessment, we present qualitative comparison in Fig. 4. Additionally, in Fig. 5, we provide 3D visualizations of reconstructed surface normals. Marigold not only correctly captures the scene layout, such as the spatial relationships between walls and furniture in the first example in Fig. 5, but also captures fine-grained details, as indicated by the arrows in Fig. 4. Moreover, the reconstruction of flat surfaces, especially walls, is significantly better (see Fig. 4). Furthermore, our method effectively models common shapes and their layouts, once again aligning with our expectations regarding the generative prior.
3 Ablation Studies
Two zero-shot validation sets are selected for the ablation studies – the official training split of NYUv2 , consisting of 785 samples, and a randomly selected subset of 800 images from the KITTI Eigen training split. Refer to supplementary sections for extra ablations and discussion.
Training noise. We investigate the impact of three types of noise during the training phase. As shown in Tab. 2, training with multi-resolution noise significantly improves the depth prediction accuracy over using standard Gaussian noise. Furthermore, the gradual annealing of multi-resolution noise yields an additional improvement. We also noticed that training with multi-resolution noise leads to more consistent predictions given different initial noise at inference time and annealing further enhances this consistency.
Training data domain. To better understand the impact of the synthetic datasets used for our fine-tuning protocol, we ablate on a photorealistic street-scene Virtual KITTI , and a more diverse and higher-quality indoor dataset Hypersim . The results are shown in Tab. 3. When fine-tuned on a single synthetic dataset, the pretrained LDM can already be adapted for monocular depth estimation to a certain degree, while the more diverse and photorealistic data leads to better performance on both indoor and outdoor scenes. Interestingly, adding additional training data from a different domain not only improves the performance on the new domain but also brings improvements in the original domain.
Test-time ensembling. We test the effectiveness of the proposed test-time ensembling scheme by aggregating various numbers of predictions. As shown in Fig. 6, a single prediction of Marigold already yields reasonably good results. Ensembling 10 predictions reduces the absolute relative error on NYUv2 by and ensembling 20 predictions brings an improvement of . It has been observed as a systematic effect that the performance is constantly improved as the number of predictions increases, while the marginal improvement diminishes with more than 10 predictions.
Number of denoising steps. We evaluate the effect of the re-spaced inference denoising steps driven by the DDIM scheduler . The results are shown in Fig. 7. Although trained with 1000 DDPM steps, the choice of 50 steps is sufficient to produce accurate results during inference. As expected, we obtain better results when using more denoising steps. We observe that the elbow point of marginal returns given more denoising steps depends on the dataset but is always under 10 steps. This implies that one can further reduce the denoising steps to 10 or even less to gain efficiency while keeping comparable performance. Interestingly, this threshold is smaller than what is usually required for diffusion-based image generators , i.e., 50 steps.
Conclusion
We presented Marigold, a fine-tuning protocol for Stable Diffusion and a model for state-of-the-art affine-invariant depth estimation. Evaluation results of our method confirm the importance of rich scene understanding prior, which we harness from the pretrained text-to-image latent diffusion model. Future research directions could find ways to counter the limitations of the current approach, such as handling a higher depth range of the scenes, possibly including infinite depth prediction, as well as more efficient inference and automatic selection of denoising steps.
References
Appendix
In this supplementary material, we provide additional implementation details in Appendix A and present additional quantitative and qualitative results in Appendix B and Appendix C, respectively.
Appendix A Implementation Details
We train on two synthetic datasets, Hypersim and Virtual KITTI , whose images have different resolutions and aspect ratios. For each batch, we probabilistically choose the dataset and then draw samples from it. We ablate the Bernoulli parameter of dataset sampling in Sec. B.4.
A.2 Annealed Multi-Resolution Noise
In the standard multi-resolution noise, multiple Gaussian noise images are sampled to form a pyramid of resolutions and then subsequently combined by upsampling, weighted averaging, and renormalization. The weight for the -th pyramid level is computed as , where is a strength of influence of lower-resolution noise. To bring such noise closer to the Gaussian used in the original DDPM formulation, we propose to anneal the weight of levels based on the diffusion schedule. Specifically, we assign the -th level at timestep the weight , where is the total number of diffusion steps. Thus, a smaller weight is given to lower-resolution levels at timesteps closer to the noise-free end of the schedule. In addition to the ablation study in the main paper, we further demonstrate the effectiveness of annealing and other noise settings in Sec. B.3.
A.3 Alignment with Ground Truth Depth
Following the established evaluation protocol , we use least squares fitting over pixels with valid ground truth values to compute the scale and shift factors of the affine-invariant predictions. Note that, while some methods predict affine-invariant disparities , others (including ours) predict affine-invariant depth values . We apply least squares fitting accordingly, i.e. the disparities are aligned to the inverse ground truth depth.
A.4 Visualization in 3D
We compute the scale and shift scalars between the prediction and ground truth. Subsequently, we unproject pixels into the metric 3D space using the camera intrinsics. We manually estimate the scale, shift, and intrinsics of “in-the-wild” samples, where ground truth and camera intrinsics are unavailable. For some samples, camera intrinsics can also be extracted from the EXIF metadata. To visualize normals, we perform least squares plane fitting at each position, considering a neighborhood area of pixels around it.
Appendix B Experimental Results
To assess how well the pre-trained image variational autoencoder of Stable Diffusion works with depth maps, we tested it with 800 samples from the Hypersim training set. To this end, each sample is normalized to the operational range of VAE as explained in the main paper, and replicated three times to accommodate the RGB interface. Upon decoding the latents, the reconstructed depth map is derived by averaging the three RGB channels. Over the chosen set of depth maps, the Mean Absolute Error (MAE) of reconstructions is , which is safely below the current state-of-the-art depth estimation errors.
B.2 Consistency of Channels After VAE Decoder
To further understand the suitability of the Stable Diffusion latent space for depth representation, we evaluate the agreement of depth channels obtained from the VAE decoder during inference. We validate with the training split of NYUv2 and a subsampled Eigen training split of the KITTI dataset . As shown in Tab. S1, the channel-wise discrepancy resulting from decoding depth from the latent space is small relative to the value range of the decoder output, i.e., $$. This could be related to the ability of VAE to represent gray-scale RGB images.
B.3 Prediction Variance and Training Noise
Since Marigold is a generative model, the predictions vary depending on the initial noise starting the diffusion process. We evaluate the consistency of predictions of three models, trained differently, i.e., with Gaussian noise, multi-resolution noise, and annealed multi-resolution noise. We train with two synthetic datasets and validate with the training split of NYUv2 and a subsampled Eigen training split of the KITTI dataset . Specifically, we perform inference 10 times for each sample and compute pixel-wise statistics over the resulting depth predictions. Subsequently, we aggregate these statistics across entire datasets and report them in Tab. S2. As seen from the values, training with the multi-resolution noises increases the prediction consistency at inference, and the annealed version brings further improvement. Fig. S1 demonstrates predictions for a single sample with three models and varying starting noise.
B.4 Ratio of Mixed Training Datasets
To further investigate the impact of the synthetic datasets used in our fine-tuning protocol, we ablate the mixing ratio of the datasets, discussed in Sec. A.1. We train with two synthetic datasets, Hypersim and Virtual KITTI , and validate with the training split of NYUv2 and a subsampled Eigen training split of the KITTI dataset . As shown in Tab. S3, training with a mixture of these two synthetic datasets yields better results on both indoor and outdoor real datasets, than training with a single synthetic dataset. Interestingly, based on the higher-quality indoor dataset, Hypersim , adding a small portion (5%) of Virtual KITTI , a street-view dataset, can already increase the performance on the outdoor dataset. We find a sweet spot at around 10% where the performance is improved on both indoor and outdoor scenes. When the ratio of Virtual KITTI keeps increasing, the overall performance is impaired. This is likely caused by the varying scene diversity and rendering quality of these two datasets.
Appendix C Qualitative Comparisons
We present the gallery of “in-the-wild” images and corresponding predictions in Fig. S2. The input images are taken in daily life or downloaded from the internet. Our method, Marigold, predicts accurate depth maps, exhibiting better overall layout and fine details. We show the final predictions for each method, that is, depth for Marigold and LeReS, and disparity for MiDaS.
C.2 Test Datasets
We show additional qualitative comparisons with our competitors , on 5 test datasets . The depth maps are visualized in Fig. S3, and the normal maps can be found in Fig. S4. Marigold excels at capturing fine scene details and reflecting the global scene layout.