Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Kaixuan Wang, Hao Chen, Gang Yu, Chunhua Shen, Shaojie Shen

Details for Models

Details for ConvNet models. In our work, our encoder employs the ConvNext [liu2022convnet] networks, whose pretrained weight is from the official released ImageNet-22k pretraining. The decoder follows the adabins [bhat2021adabins]. We set the depth bins number to 256, and the depth range is [0.3m,150m][0.3m,150m]. We establish 4 flip connections from different levels of encoder blocks to the decoder to merge more low-level features. An hourglass subnetwork is attached to the head of the decoders to enhance background predictions.

Details for ViT models. We use dino-v2 transformers [oquab2023dinov2] with registers [darcet2023vision] as our encoders, which are pretrained on a curated dataset with 142M images. DPT [ranftl2021vision] is used as the decoders. For the ViT-S and ViT-L variants, the DPT decoders take only the last-layer normalized encoder features as the input for stabilized training. The giant ViT-g model instead takes varying-layer features, the same as the original DPT settings. Different from the convnets models above, we use depth bins ranging from [0.1m,200m][0.1m,200m] for ViT models.

Details for recurrent blocks. As illustrated in Fig 1. Each recurrent block updates hierarchical features maps {H1/14t,H1/7t,H1/4t}\{\mathbf{H}^{t}_{1/14},\mathbf{H}^{t}_{1/7},\mathbf{H}^{t}_{1/4}\} at {114,17,14}\{\frac{1}{14},\frac{1}{7},\frac{1}{4}\} scales and the intermediate predictions sD^ct\hat{\mathbf{D}}_{c}^{t}, N^t\hat{\mathbf{N}}^{t} at each iteration step tt. This block compromises three ConvGRU sub-blocks to refine feature maps at different scales, and two projection heads Gd\mathcal{G}_{d} and Gn\mathcal{G}_{n} to predict updates for depth and normal respectively. The feature maps are gradually refined from the coarsest (H1/14t\mathbf{H}^{t}_{1/14}) to the finest (H1/4t\mathbf{H}^{t}_{1/4}). For instance, the refined feature map at the 114\frac{1}{14} scale H1/14t+1\mathbf{H}^{t+1}_{1/14} is fed into the second ConvGRU sub-block to refine the 17\frac{1}{7}-scale feature map H1/7t\mathbf{H}^{t}_{1/7}. Finally, the projection heads Gd,Gn\mathcal{G}_{d},\mathcal{G}_{n} employs a concatenation of original predictions D^ct\hat{\mathbf{D}}_{c}^{t}, N^t\hat{\mathbf{N}}^{t} and the to finest feature map H1/4t\mathbf{H}^{t}_{1/4} to predict the update items ΔD^ct+1\Delta\hat{\mathbf{D}}_{c}^{t+1}, ΔN^t+1\Delta\hat{\mathbf{N}}^{t+1}. Both projection heads are composed of two linear layers with a sandwiched ReLU activation layer.

Resource comparison of different models. We compare the resource and performance among our model families in Tab. 1. All inference-time and GPU memory results are computed on an Nvida-A100 40G GPU with the original pytorch implemented models (No engineering optimization like TensorRT or ONNX). Generally, the enormous ViT-Large/giant-backbone models enjoy better performance, while the others are more deployment-friendly. In addition, our models built in classical en-decoder schemes run much faster than the recent diffusion counterpart [ke2023repurposing].

We collect over 1616M data from 18 public datasets for training. Datasets are listed in Tab. 2. When training the ConvNeXt-backbone models, we use a smaller collection containing the following 11 datasets with 88M images: DDAD [packnet], Lyft [lyftl5preception], DrivingStereo [yang2019drivingstereo], DIML [cho2021diml], Argoverse2 [Argoverse2], Cityscapes [Cordts2016Cityscapes], DSEC [Gehrig21ral], Maplillary PSD [MapillaryPSD], Pandaset [itsc21pandaset], UASOL[bauer2019uasol], and Taskonomy [zamir2018taskonomy]. In the autonomous driving datasets, including DDAD [packnet], Lyft [lyftl5preception], DrivingStereo [yang2019drivingstereo], Argoverse2 [Argoverse2], DSEC [Gehrig21ral], and Pandaset [itsc21pandaset], have provided LiDar and camera intrinsic and extrinsic parameters. We project the LiDar to image planes to obtain ground-truth depths. In contrast, Cityscapes [Cordts2016Cityscapes], DIML [cho2021diml], and UASOL [bauer2019uasol] only provide calibrated stereo images. We use raftstereo [lipson2021raft] to achieve pseudo ground-truth depths. Mapillary PSD [MapillaryPSD] dataset provides paired RGB-D, but the depth maps are achieved from a structure-from-motion method. The camera intrinsic parameters are estimated from the SfM. We believe that such achieved metric information is noisy. Thus we do not enforce learning-metric-depth loss on this data, i.e., LsilogL_{silog}, to reduce the effect of noises. For the Taskonomy [zamir2018taskonomy] dataset, we follow LeReS [yin2022towards] to obtain the instance planes, which are employed in the pair-wise normal regression loss. During training, we employ the training strategy from [yin2020diversedepth_old] to balance all datasets in each training batch.

The testing data is listed in Tab. 2. All of them are captured by high-quality sensors. In testing, we employ their provided camera intrinsic parameters to perform our proposed canonical space transformation.

2 Details for Some Experiments

where Dc\mathbf{D}_{c} and N\mathbf{N} are the predicted depth in the canonical space and surface normal, Dc∗\mathbf{D}^{\ast}_{c} and N∗\mathbf{N}^{\ast} are the groundtruth labels, LdL_{d}, LnL_{n}, and Ld−nL_{d-n} are the losses for depth, normal, and depth-normal consistency introduced in the main text.

Reconstruction of in-the-wild scenes. We collect several photos from Flickr. From their associated camera metadata, we can obtain the focal length f^\hat{f} and the pixel size δ\delta. According to \nicefracf^δ\nicefrac{{\hat{f}}}{{\delta}}, we can obtain the pixel-represented focal length for 3D reconstruction and achieve the metric information. We use meshlab software to measure some structures’ size on point clouds. More visual results are shown in Fig. 6.

Generalization of metric depth estimation. To evaluate our method’s robustness of metric recovery, we test on 7 zero-shot datasets, i.e. NYU, KITTI, DIODE (indoor and outdoor parts), ETH3D, iBims-1, and NuScenes. Details are reported in Tab. 2. We use the officially provided focal length to predict the metric depths. All benchmarks use the same depth model for evaluation. We don’t perform any scale alignment.

Evaluation on affine-invariant depth benchmarks. We follow existing affine-invariant depth estimation methods to evaluate 5 zero-shot datasets. Before evaluation, we employ the least square fitting to align the scale and shift with ground truth [leres]. Previous methods’ performance is cited from their papers.

Dense-SLAM Mapping. This experiment is conducted on the KITTI odometry benchmark. We use our model to predict metric depths, and then naively input them to the Droid-SLAM system as an initial depth. We do not perform any finetuning but directly run their released codes on KITTI. With Droid-SLAM predicted poses, we unproject depths to the 3D point clouds and fuse them together to achieve dense metric mapping. More qualitative results are shown in Fig. 5.

3 More Visual Results

Qualitative comparison of depth and normal estimation. In Figs 2, 3, we compare visualized depth and normal maps from the Vit-g CSTM_label model with ZoeDepth [bhat2023zoedepth], Bae etal [bae2021estimating], and Omnidata [eftekhar2021omnidata]. In Figs. 4, 8, 9, and 10, We show the qualitative comparison of our depth maps from the ConvNeXt-L CSTM_label model with Adabins [bhat2021adabins], NewCRFs [yuan2022new], and Omnidata [eftekhar2021omnidata]. Our results have much fine-grained details and less artifacts.

Reconstructing 360∘ NuScenes scenes. Current autonomous driving cars are equipped with several pin-hole cameras to capture 360∘ views. Capturing the surround-view depth is important for autonomous driving. We sampled some scenes from the testing data of NuScenes. With our depth model, we can obtain the metric depths for 6-ring cameras. With the provided camera intrinsic and extrinsic parameters, we unproject the depths to the 3D point cloud and merge all views together. See Fig. 7 for details. Note that 6-ring cameras have different camera intrinsic parameters. We can observe that all views’ point clouds can be fused together consistently.