Probabilistic and Geometric Depth: Detecting Objects in Perspective

Tai Wang, Xinge Zhu, Jiangmiao Pang, Dahua Lin

Introduction

3D object detection is an essential task for many robotic systems such as autonomous vehicles. Recent advanced methods in this field typically resort to various sensors, such as LiDAR , Radar , binocular vision , or their combinations for accurate depth information. Nevertheless, these perceptual systems are complicated, expensive, and difficult to maintain in complex environments. In contrast, monocular 3D detection, a setting that aims at perceiveing 3D objects from 2D monocular images, has drawn increasing attention due to its low costs. However, as the depth information is not directly manifest in the input, this task is inherently ill-posed, making the problem particularly challenging.

This paper starts from a systematic study about this problem on two authoritative benchmarks in a quantitative way. Although we already knew the depth information is critical to this task, the study surprisingly shows that inaccurate depth estimation blocks all the other localization predictions from improving the final results. As instance depth has shown to be the bottleneck, we can simplify monocular 3D detection as an instance depth estimation problem to tackle it essentially.

Previous methods first use an extra cumbersome depth estimation model to complement 2D detectors on depth information. The following methods simplify the frameworks by directly regarding depth as one dimension of the 3D localization task. However, they still use simple methods that estimate depth from isolated instances or pixels in a regression manner. We observe that aside from each object itself, other objects are co-existing in an image and the geometric relations across them can be valuable constraints to guarantee accurate estimation.

Motivated by these observations, we propose Probabilistic and Geometric Depth (PGD) that jointly leverages probabilistic depth uncertainty and geometric relationships across co-existed objects for accurate depth estimation. Specifically, as the preliminary depth estimation of each instance is usually inaccurate in this ill-posed setting, we incorporate a probabilistic representation to capture the uncertainty of the estimated depth. We first bucket the depth values into a set of intervals and calculate the depth by the expectation of the distribution (Fig. 1(a)). The average of top-k confidence scores from the distribution is taken as the uncertainty of the depth. To model the geometric relations, we further construct a depth propagation graph to enhance the estimations with their contextual relationship. The uncertainty of each instance depth provides useful guidance for the propagation therein. Benefiting from this overall scheme, we can easily identify the predictions with higher confidence, and more importantly, estimate their depths more accurately with the graph-based synergistic mechanism.

We implement the methods on a simple monocular 3D object detector FCOS3D . Despite the simplicity of the basic idea, our PGD results in significant improvements on KITTI and nuScenes with different benchmark settings and evaluation metrics. It achieves 1st place out of all monocular vision-only methods while still maintaining real-time efficiency. The simple yet effective method proves that with only designs tailored to depth, a 2D detector can be capable of detecting objects in perspective.

Related Work

2D Object Detection According to the base of initial guesses, modern 2D detection methods can be divided into two branches, anchor-based and anchor-free. Anchor-based methods benefit from the predefined anchors in terms of much easier regression, while anchor-free methods do not need complicated prior settings and thus have better universality. For simplicity, we take FCOS3D , the 3D adapted version of FCOS , as the baseline considering its capability of handling overlapped ground truths and scale variance problem.

Monocular Depth Estimation Monocular depth estimation is also a challenging ill-posed problem like monocular 3D detection. It aims at predicting dense and global depth field at pixel level given an RGB image. Early works predict depth from hand-crafted features with non-parametric optimization methods. With the rapid progress of CNNs, fully supervised methods , self-supervised methods based on stereo pairs and monocular videos gradually emerged. Although this problem has been explored for a long time, there are very few works studying it in a specific task, like detection, where the dense depth supervision is always not guaranteed and we only care about the accuracy of instance depth instead of the global depth field.

As for the reformulation of depth learning problems, there are a few attempts in this field. For example, DORN recasts the depth learning problem as ordinal regression and proposes a spacing-increasing discretization (SID) strategy to improve network training and reduce computations. It is similar to the underlying idea of our probabilistic representation for uncertainty modeling while different in terms of motivation and design details.

Monocular 3D Object Detection Monocular 3D detection is more complicated than the 2D case. The underlying problem is the inconsistency of input 2D data modal and the output 3D predictions.

Methods involving sub-networks Earlier work uses sub-networks to assist 3D detection. 3DOP and MLFusion use a depth estimation network while Deep3DBox uses a 2D object detector. They rely on the design and performance of these sub-networks, even external data and pre-trained models, which makes the training inconvenient and introduces additional system complexity.

Transform to 3D representation Another category is to convert the RGB input to 3D representations like OFTNet and Pseudo-Lidar . Although these methods have shown promising performance, they actually rely on dense depth labels and hence are not regarded as pure monocular approaches. There are also domain gaps between different depth sensors, making them hard to generalize smoothly to a new practical setting. Furthermore, the efficiency of processing a large number of point clouds is also a significant issue to deal with in practical applications.

End-to-end designs like 2D detection Recent work notices these drawbacks, and end-to-end frameworks are thus proposed. M3D-RPN implements a single-stage multi-class detector with an end-to-end region proposal network and depth-aware convolution. SS3D proposes to detect 2D key points and further predicts object characteristics with uncertainties. MonoDIS introduces a disentangling loss to reduce the instability of the training procedure. Some of them still have multiple training stages or post-optimization phases. In addition, they all follow anchor-based manners, and thus the consistency of 2D and 3D anchors is needed to be determined. In contrast, anchor-free methods do not need to make statistics on the given data and have better generalized ability to more various classes or different intrinsic settings, so we choose to follow this paradigm.

Nevertheless, all of these works rarely have customized designs for instance depth estimation in particular, and only take it as one common regression target for isolated points or instances. It actually hinders the breakthrough of this problem, which will be discussed in our quantitative study and specifically addressed in our approach.

Preliminary and Motivating Study

In this section, we aim at making an in-depth quantitative error analysis on top of a basic adapted monocular 3D detector to investigate the key challenge in the specific 3D detection setting.

Typically, conventional 2D detection expects the model to predict 2D bounding boxes and category labels for each object of interest, while a monocular 3D detector needs to predict 7-DoF 3D boxes given the same input. From this perspective of problem formulation, the main difference lies on the regression targets. An intuitive reason for the much worse performance of monocular 3D detection compared to 2D is that there exist much more difficult targets to regress in the localization. Hence, we choose a simple detector FCOS3D to study the specific problem, which keeps the well-developed designs for 2D feature extraction and is adapted for this 3D task with only basic designs for specific 3D detection targets. As shown in the left part of Fig. 3, there are overall two branches for classification and localization respectively. Formally, for the regression branch, the detector predicts 3D attributes, including offsets Δ\Deltax, Δ\Deltay to the projected 3D center, depths dd, 3D size w3Dw^{3D}, l3Dl^{3D}, h3Dh^{3D}, sin value of rotation θ\theta, direction class CθC_{\theta}, center-ness cc, and distances to four sides of 2D boxes ll, rr, tt, bb, for each location on the output dense map. We further equip it with a basic consistency loss between 3D and 2D localization, which will be detailed in the appendix.

On this basis, we apply this baseline on two representative benchmarks, KITTI and nuScenes, and replace the predictions with ground truths step by step to identify the performance bottleneck (Fig. 2). Unexpectedly, the inaccurate depth blocks all the other sub-task predictions from improving the overall detection performance, on both datasets under different metrics. Hence, current monocular 3D detection, especially 3D localization, can be reduced to the dominating instance depth estimation problem to a great extent, which will be the focus of our method to be presented next. See more details about the explanation of oracle analyses in the appendix.

Our Approach

Given images collected from similar cameras, previous work typically resorts to direct regression for instance depth estimation and expects the model to directly learn that objects with certain appearances and sizes always exist at locations with certain depths. Our baseline also follows this way. However, it is hard to learn due to the large variance and also obviously not enough for the accuracy needed in 3D detection. Given the inherent downside of hard regression for isolated points, in our approach, we aim at constructing an uncertainty-aware depth propagation graph to enhance the estimation from contextual connections among instances. Next, we will first elaborate on the adopted probabilistic representation and technical details of the constructed geometric graph, and finally present how we integrate these obtained depth estimations.

where DPD_{P} is the so-called probabilistic depth. It is equivalent to compute the expectation of the probabilistic distribution formed by softmax(DPM)softmax(D_{PM}). Apart from the DPD_{P}, we can further obtain the depth confidence score, denoted as sd∈SDs^{d}\in S_{D}, from the depth distribution of each instance. In practice, we take the average of top-2 confidence as the depth score for U=10mU=10m. It will be multiplied by the center-ness and classification score as the final ranking criterion for predictions during inference.

Subsequently, we fuse DRD_{R} and DPD_{P} with the sigmoid response of a data-agnostic single parameter λ\lambda:

Here DLD_{L} is regarded as a local depth estimation for each isolated instance, which together with the depth score derived from DPMD_{PM} serve as the foundation of constructing the depth propagation graph.

It is worth noting that our implementation is different from the typical way used in monocular depth estimation , which usually adopts a fine-grained quantization for the depth interval and further estimates the value with classification and residual regression. In comparison, our method is more memory-efficient, more straightforward for regressing continuous value, and provides a natural indicator for uncertainty estimation. Please refer to the appendix for empirical results about comparison with other depth interval division methods.

2 Depth Propagation from Perspective Geometry

With the depth prediction DLD_{L} for isolated instances and their depth confidence scores SDS_{D} for uncertainty estimation, we can further construct the propagation graph based on the contextual geometric relationship. Consider the typical driving scenarios: a general constraint can be leveraged, i.e., almost all the objects are on the ground. Early in , Hoiem et al. utilized the scene projection with this constraint to put objects in the context of the overall 3D scene by modeling the relationship of different elements. Here, targeting the depth estimation problem, we instead propose a geometric depth propagation mechanism with consideration of interdependence between instances. Next, we will first derive the perspective relationship between two instances, and then present the details of the graph-based depth propagation scheme with edge pruning and gating.

Perspective Relationship Consider in the general perspective projection, suppose the camera projection matrix PP is:

where ff is the focal length, cuc_{u} and cvc_{v} are the vertical and horizon position of camera in the image, bxb_{x}, byb_{y} and bzb_{z} denote the baseline with respect to the reference camera (non-zero in KITTI while zero in nuScenes). Note that we represent the focal length with a single ff considering most cameras share the same one for the uu and vv axis. Then a 3D point x3D=(x,y,z,1)T\mathbf{x^{3D}}=(x,y,z,1)^{T} in the camera coordinates can be projected to a point x2D=(u′,v′,1)T\mathbf{x^{2D}}=(u^{\prime},v^{\prime},1)^{T} in the image with:

To simplify the result, we replace v′v^{\prime} with v+cvv+c_{v}, then vv represents the distance to the horizon line (Down is the positive direction in Fig. 4). Then we get:

The relation for uu is similar. Considering the constraint that all the objects are on the ground, the bottom centers of objects always share the same yy (height in the camera coordinates), so we mainly consider this relation for vv next. Given two objects 1 and 2, the relationship between their center depths can be derived from Eqn. 5:

with which we can predict d2d_{2} given d1d_{1} precisely with the height difference between 3D centers. Besides, we can also leverage an approximation of this relationship given the assumption that objects share the same bottom height, then y2−y1y_{2}-y_{1} can be substituted by the difference of half heights of 3D boxes 12(h13D−h23D)\frac{1}{2}(h^{3D}_{1}-h^{3D}_{2}), defined as d1→2Pd_{1\to 2}^{P}.

In this relation, when h13D=h23Dh^{3D}_{1}=h^{3D}_{2}, v1d1=v2d2v_{1}d_{1}=v_{2}d_{2}, which is easy to understand, i.e., an object closer to vanishing line is farther away. It is a clear relationship connecting different instances but also yields errors. Suppose ∣(y2−y1)−12(h1−h2)∣=δ|(y_{2}-y_{1})-\frac{1}{2}(h_{1}-h_{2})|=\delta, the error of depth will be Δd=fv2δ\Delta d=\frac{f}{v_{2}}\delta. When δ=0.1m\delta=0.1m, v2=50v_{2}=50 (pixels), Δd\Delta d can be about 1.5m1.5m. Although it is acceptable for objects 30m30m away (corresponding with v2=50v_{2}=50), we also need a mechanism to avoid possible large errors. It consists of the edge pruning and gating scheme to be described next and the location-aware weight map to be mentioned in the Sec. 4.3.

Graph-Based Depth Propagation With the pairwise perspective relationship, we can estimate the depth of any object from the cues of other objects. Then we can construct a dense directed graph with two bidirectional edges between any two objects representing the depth propagation (Fig. 4). Formally, suppose we have NN predicted objects with indices from P={1,2,...,n}\mathcal{P}=\{1,2,...,n\}, we can estimate the depth of object ii given dj→iPd_{j\to i}^{P} for all the j∈Pj\in\mathcal{P}, defined as the geometric depth diG∈DGd_{i}^{G}\in D_{G}. Considering the computational efficiency and possible large errors mentioned previously, we propose an edge pruning and gating scheme to improve the propagation graph. From our observation, the same category of nearby objects can well satisfy the ”same ground” condition, so we select the following 3 most important factors to decide which edges are influential and reliable, including the depth confidence sjds^{d}_{j}, 2D distance score sij2Ds^{2D}_{ij}, and classification similarity sijclss^{cls}_{ij}. The latter two and the overall edge score sj→ies^{e}_{j\to i} are computed as follows:

where tij2Dt^{2D}_{ij} is the 2D distance between projected centers of object ii and jj, tmax2Dt^{2D}_{max} is set to the length of image diagonal, fi\bm{f_{i}} and fj\bm{f_{j}} are the output confidence vectors of two objects from classification branch and kk is the maximum number of edges to be kept after pruning (edges with top-k scores are kept). The edge score is then used for gating so that each node attends its edges with their importance:

Note that obtaining the geometric depth map DGD_{G} from this graph is free of learnable parameters. To avoid influencing the learning of other components, we cut off the gradients backpropagated from this computation and only focus on how to integrate DLD_{L} and DGD_{G}, which will be discussed next.

3 Probabilistic and Geometric Depth Estimation

The fused depth DD will replace the direct regressed DRD_{R} in the baseline and trained with the common smooth L1 loss in the same end-to-end way. Note that adding intermediate supervisions empirically makes the training more stable but does not bring any performance gains.

Experiments

In this section, we present our experimental setup and implementation details, and then make the quantitative analysis on the KITTI and nuScenes dataset with details of both performance and efficiency. Finally, detailed ablation studies are conducted to show the efficacy of each component in our method. Refer to the appendix for more qualitative analysis.

We evaluate our method on two datasets, KITTI and nuScenes . There are 7481/7518 samples for training/testing respectively on KITTI, and the training samples are generally divided into 3712/3769 samples as training/validation splits. We first validate our method on this popular benchmark. Nevertheless, the variety of scenes and categories is limited on KITTI, so we further test our approach on the large-scale nuScenes dataset. NuScenes consists of multi-modal data collected from 1000 scenes, including RGB images from 6 cameras, points from 5 Radars, and 1 LiDAR. It is split into 700/150/150 scenes for training/validation/testing. There are overall 1.4M annotated 3D bounding boxes from 10 categories. In addition, nuScenes uses different metrics, distance-based mAP and NDS, which can help evaluate our method from another perspective. See more explanations about metrics in the appendix.

2 Implementation Details

Network Architectures As shown in Fig. 3, our baseline framework basically follows the design of FCOS3D . Given the input image, we utilize ResNet101 as the feature extraction backbone followed by FPN for generating multi-level predictions. Detection heads are shared among multi-level feature maps except that three scale factors are used to differentiate some of their final regressed results, including offsets, depths, and sizes, respectively. For the hyperparameters in the depth estimation module, UU is set to 10m10m and kk is set to 5. The overall framework is built on top of MMDetection3D . Please refer to FCOS3D and appendix for the design of loss and other implementation details.

Training Parameters For all the experiments, we trained randomly initialized networks from scratch following end-to-end manners. Models are trained with SGD optimizer, in which gradient clip and warm-up policy are exploited with learning rate 0.001, number of warm-up iterations 500, warm-up ratio 0.33 and batch size 32/12 on 16/4 GTX 1080Ti GPUs for nuScenes/KITTI.

Data Augmentation We only implement image flip for augmentation, where offset and 2D targets are flipped for the 2D image while 3D boxes are transformed correspondingly in 3D space. No other augmentation (right image augmentation, cropping, resizing, etc.) methods are utilized.

3 Quantitative Analysis

We make quantitative analyses both on KITTI (Tab. 1 and 4, Fig. 1(c)) and much harder, less commonly validated nuScenes dataset (Tab. 2). It can be seen that our method achieves the state-of-the-art on both benchmarks with different settings and metrics while maintains outstanding speed.

We list part of early monocular methods with extra data or pre-trained models and recent image-only methods that have related results for comparison on the KITTI dataset. Only the results for car detection are compared here because the performance of small objects is always unstable due to their limited samples. Our framework based on the simple adapted FCOS3D achieves much better performance than others, especially considering M3D-RPN and RTM3D adopt stronger backbone and data augmentation. Furthermore, our method can run at the speed of 36Hz to achieve this, thanks to most of our modules not introducing extra computational costs to inference. It is an excellent trade-off between performance and efficiency.

Then for the nuScenes dataset, we also compare the results on the test set and validation set, respectively. On the test set, we first compared all the methods using RGB images as the input data. Our single model achieved the best performance among them with mAP 37.0% and NDS 43.2%, in which we particularly exceeded the previous best method more than 3% in terms of mAP. We also list benchmarks based on other data modality, including lightweight, real-time PointPillars with LiDAR, CenterFusion with RGB image and Radar, and CenterPoint ensemble results with all the sensors. It can be seen that although our method has a certain gap with the high-performance CenterPoint, it even surpasses PointPillars and CenterFusion on mAP, which shows that this ill-posed problem can be solved decently with enough data. At the same time, the methods using other modal data usually yield better NDS, mainly because the mAVE is smaller. The reason is that they can predict the speed of objects from continuous multi-frame point clouds or velocity measurement of Radar. In contrast, we only use the single-frame image in our experiments. So how to mine the speed information from consecutive frame images will be a direction worthy of exploring in the future. On the validation set, we compare our method with the best open-source center-based detector, CenterNet. Our method is not only much more efficient to train and inference (3 days to train the CenterNet vs. only one day to train our model with comparable performance), but also achieves better performance, especially in terms of the mAP and mAOE. On this basis, we finally achieved an improvement of about 9% on NDS. See more detailed results about depth estimation accuracy and per-class detection performance in the appendix.

4 Ablation Studies

Finally, we conduct ablation studies to validate the efficacy of our proposed key components on KITTI (Tab. 4) and nuScenes (Tab. 6). We can observe that local constraints can basically enhance the baseline, and our probabilistic and geometric depth further boost the performance significantly, especially in terms of mAP and translation error (mATE). Tab. 8 and Tab. 8 show more details of two core components for improving depth estimation with the metrics average precision under IOU≥\geq0.7. It can be seen that combining the probabilistic representation (prop. branch in Tab. 8) with direct regression (w/ direct) and leveraging the depth score in the inference (depth score) can finally make the most of this design. For geometric depth, the basic fusion with local estimation can not bring the desirable gain. Improving the propagation graph via edge pruning and gating (edge gating) and cutting off the unexpected gradients propagation (cut off grad.) can help remove possible noises and prompt the learning more focused on the final integration, thus making the overall scheme much more effective. As for alternative implementations, we compare feasible methods of computing the depth score from the probabilistic distribution (Tab. 6). Compared to other more complicated ways, normalized entropy and standard deviation, our exploited top-2 score can achieve decent results with better efficiency. See more results about different depth division methods and detailed analyses for these two datasets from other perspectives like the Precision-Recall curve in the appendix.

Conclusion

This paper targets the key challenge lying in monocular 3D object detection, instance depth estimation. Started from a basic adapted 3D detector, we firstly make in-depth oracle analyses. We surprisingly find that depth estimation is the dominating bottleneck for current 3D detection, especially in terms of localization. To tackle the discovered challenge, we propose a novel approach, Probabilistic and Geometric Depth (PGD), which leverages the geometric relationship in perspective to construct a graph connecting instance estimations with uncertainty and thus predicts depths more accurately. The efficacy of this solution is demonstrated on both KITTI and large-scale nuScenes datasets. In the future, we will further extend the geometric depth scheme to more general cases by relaxing the ”ground” assumption via 2D height regression or ground normal estimation, and validate this pipeline on other 2D detectors. How to better leverage temporal geometry information to address the difficulty of instance depth estimation is also a promising direction worthy of further exploration.

This work is supported in part by Centre for Perceptual and Interactive Intelligence Limited, in part by the GRF through the Research Grants Council of Hong Kong under Grants (Nos. 14208417, 14207319 and 14203518) and ITS/431/18FX, in part by CUHK Strategic Fund and CUHK Agreement TS1712093, in part by the Shanghai Committee of Science and Technology, China (Grant No. 20DZ1100800).

References

Implementation Details

This section first presents the adopted local geometric constraints between 2D and projected 3D bounding boxes in the enhanced baseline. Subsequently, we will elaborate on the details of training loss and inference procedure.

Our baseline FCOS3D only stiffly adjusts the output of networks to fit the requirements of 3D detection. There is no relationship or constraints between these predicted attributes, making this network hard to train, especially when the data is limited. Considering our detector can achieve 90% accuracy on 2D vehicle detection, we add 2D localization into our targets and use it to regularize 3D outputs. Actually, this closed-loop and self-supervised approach is also consistent with what humans do in the annotation procedure . In practice, as shown in Fig. 5, we add a consistency loss (GIoU loss) between our estimated 2D boxes and the exterior 2D boxes of 3D predictions to enhance our baseline, which is particularly important on the small KITTI dataset. Note that due to the difficulty of regressing accurate depth, we use the ground truth depth for deriving the 3D bounding boxes when computing the consistency loss.

Here we provide an example to show the intuition behind this design. Typically when the data is limited, it is hard for the network to direct regress different 3D targets (offset, depth, orientation, etc.) independently. For example, in Fig. 6, the orientation of nearby large objects predicted by our baseline can be very inaccurate (the top line in the figure) even though it can be easily rectified with simple verification. So we add the more reliable 2D localization into our targets to regularize our 3D predictions. It turns out that the simple local constraint could alleviate this problem in the learning procedure while does not introduce extra computational costs to inference. The improved results after adding this constraint can be seen in Fig. 6 (the bottom line).

2 Loss

Overall Loss Design We basically follow the loss design of FCOS3D except our proposed consistency loss and the adjustments for different datasets.

To have a brief review, firstly, we use the focal loss as the object classification loss:

where pp is the class probability of a predicted box, and we follow the common settings, α=0.25\alpha=0.25 and γ=2\gamma=2. For attribute classification on nuScenes, we use a simple softmax classification loss, denoted as LattrL_{attr}.

For regression branch, we use the smooth L1 loss for each regression target except centerness:

The weights of Δx,Δy,d,w,l,h,θ\Delta x,\Delta y,d,w,l,h,\theta error are 1 and the weights of vx,vyv_{x},v_{y} on nuScenes are 0.05. We use the softmax classification loss and binary cross entropy (BCE) loss for direction classification and centerness regression, denoted as LdirL_{dir} and LctL_{ct} respectively. For local geometric constraints, denote our predicted 2D boxes as B2D\bm{B_{2D}}, the minimum exterior 2D boxes of projected 3D boxes as Bproj\bm{B_{proj}}, then the consistency loss is:

NposN_{pos} is the number of positive predictions and βcls=βattr=βloc=βdir=βct=βgeo=1\beta_{cls}=\beta_{attr}=\beta_{loc}=\beta_{dir}=\beta_{ct}=\beta_{geo}=1. Note that the attribute loss LattrL_{attr} and velocity loss in the LlocL_{loc} are only required in the nuScenes experiments.

In addition, we use a much stronger uncertainty formulation for this multi-task learning problem as presented in . Specifically, referring to its formulation of maximum likelihood and homoscedastic uncertainty, we formulate the depth loss as:

Here D^\hat{D} and DD are the targets and predictions of depth, L1L_{1} represents the original smooth L1 loss with δ=3.0\delta=3.0 and σ\sigma is the variable for uncertainty. In practice, to make the learning easier, we train the network to predict the log variance s=logσ2s=log\sigma^{2} only for depth estimation, which is more numerically stable than directly predicting the variance. Correspondingly, exp(−s)exp(-s) serves as the weight of depth loss. In this way, the depth loss will be adaptively weighted relative to other regression losses. Additionally, the uncertainty exp(−s)exp(-s) can also be used as another confidence score to be multiplied when inference, such that predictions with more accurate depths will have particularly higher scores. Note that this strong uncertainty indicator can only bring a significant gain on KITTI experiments while seriously hurting the general performance as evaluated on the nuScenes dataset.

Alternative Depth Loss Designs Considering we have several intermediate depth predictions, such as DRD_{R}, DPD_{P} and DLD_{L} in Fig. 5, a natural idea is to add intermediate supervisions for these predictions to guarantee that each branch can learn meaningful information. So we further defined several depth L1 losses for these predictions and tried to replace the original depth loss in the LlocL_{loc} with their weighted summation. It turns out that although this approach can make the training procedure more stable, it does not bring any performance gain. We also find that the framework never overfits to only relying on one kind of estimation even with only supervision for the final prediction, as to be shown in Sec. 3.2. It indicates that these predictions and components indeed work together from complementary aspects.

3 Inference

The inference procedure is to forward the input image through the framework and obtain bounding boxes with their class scores, attribute scores (if necessary) and centerness predictions. We multiply the class score, the predicted centerness and the depth confidence score as the overall confidence for each prediction and conduct rotated Non-Maximum Suppression (NMS) in the bird view as most 3D detectors to get the final results.

Explanation of Oracle Analyses

In this section, we will explain more about our empirical analysis, from the specific settings to more details in the results.

First, we would like to emphasize one detail in our analysis, i.e., we replace the dense predictions from the direct output of detection head with oracles to purely observe the problems of our networks. In comparison, other alternatives exist, such as replacing the decoded dense output or predictions after post-processing, which can not reveal some entangling problem lying in the formulation. One example to show the difference between these two implementations is that we replace the offset with corresponding ground truth while the latter approach replaces the decoded X,YX,Y in the 3D space with targets.

2 Comparison of Different Metrics

Then we can see that it will also consider predictions with relatively inaccurate locations (like objects with the distance error larger than 2 meters but smaller than 4 meters). This difference is especially notable when discussing the improvements from depth score, which will be detailed in Sec. 3.2.

Finally we basically describe how the NuScenes Detection Score (NDS) is calculated. To begin with, we first define that predictions with center distance from the matching ground truth d2D≤2md_{2D}\leq 2m will be considered as true positives (TP) and thus introduce 5 True Positive metrics, Average Translation Error (ATE), Average Scale Error (ASE), Average Orientation Error (AOE), Average Velocity Error (AVE) and Average Attribute Error (AAE). Given these metrics, we compute the mean TP metric (mTP) over all categories:

Therefore, NDS is a combination of several decoupled metrics and could reflect the performance of 3D detectors from another perspective. See more details about the intermediate computation in its original paper .

3 Detailed Explanations and Conclusions

Due to the space limitation in the main paper, we do not discuss much about the results shown in Fig. 7. Next, we will analyze it in detail and summarize a series of important conclusions.

Basic Observations As shown in Fig. 7, we replace the predicted attributes with their ground truth values step by step and observe the performance improvements. We can see that:

1. With only one oracle (circle dots), only depth can bring a considerable improvement (green lines). It shows that with current depth estimation, other predicted attributes do not drag down the performance, while with other predictions, the current accuracy of depth estimation is far not enough.

2. With accurate depth, other oracles (triangle dots in the figures) could bring the expected performance gains. While with current depth estimation, even all the other predictions are accurate (green rhombus dots), the results are always disappointing, even almost like the baseline.

3. Although KITTI and nuScenes are different in terms of category variety and metrics, the trend of these curves is the same. The difference is reflected in the importance of localization and classification oracles. Localization is more important on the KITTI, which has less category variety and more strict metrics. Classification is another important factor apart from depth on nuScenes, e.g., our monocular predictions with location oracle is still not better than the best LiDAR-based methods. In contrast, with an accurate depth and classification map, the performance is almost ideal.

From these observations, we can conclude that the inaccurate depth blocks all the other sub-task predictions from improving the overall detection performance. Hence, as mentioned in the main paper, the current monocular 3D detection, especially 3D localization, can be actually reduced to the dominating instance depth estimation problem.

Comparison with Best LiDAR-Based Methods There is an interesting phenomenon not much related to depth estimation in the above analysis, i.e., the comparison with best LiDAR-based methods on nuScenes. We can see that classification is particularly important on nuScenes, and our monocular predictions with location oracles are still not better than the state-of-the-art LiDAR-based methods. This result is a little dataset-specific. We conjecture it is because the classification for ten categories on nuScenes is relatively hard, or the annotation is mainly conducted in the point clouds, leading to missing objects in the images.

Supplementary Experimental Results

In this section, we will show more experimental results to help further understand our approach. First, we will provide toy examples to explain and validate our derived pairwise perspective relationship in the depth propagation. Subsequently, we make more detailed analyses in quantitative and qualitative ways to reveal the working mechanism and effect of our method.

As shown in Fig. 8, we provide two samples with many objects in one image. We first have a brief review of the perspective relationship derived in the main paper. Given two objects 1 and 2, the relationship between their centers strictly satisfies:

Considering two objects share the same ground (bottom height), we can get the approximate relationship as follows:

where dd denotes the depth, vv denotes the distance between the projected 2D object center and the horizon line in the image, yy is the 3D height of object center and h3Dh^{3D} is the height of the 3D bounding box. Taking the left sample in Fig. 8 as an example, the depths of the 8 cars are {5.23,11.80,16.50,22.05,23.64,28.53,29.07,42.85}\{5.23,11.80,16.50,22.05,23.64,28.53,29.07,42.85\}. With our derived relationship, we can estimate them with only the first 2 accurate depths: {5.23,11.74,16.78,22.92,21.13,26.59,25.78,36.51}\{5.23,11.74,16.78,22.92,21.13,26.59,25.78,36.51\}. We can see that similar to the general case of depth estimation, our propagation mechanism also yields more notable errors for distant objects, which has been analyzed in the main paper (The effect of δ\delta over Δd\Delta d will be enlarged as the v2v_{2} decreases.)

Next, we can further observe the inconsistent bottoms problem shown in Fig. 8. We mark some representative instances in the figure. It can be seen that it is sometimes caused by the actual topography, like pedestrians and cars in the second sample. Nevertheless, the noise only exists between objects far away from each other most of the time. We conjecture this is related to the annotation pipeline, e.g., we tend to make use of nearby annotations when the information for labeling the current instance is inadequate. Alternatively, sometimes it is just because the LiDAR only sweeps the top part of the distant objects such that the annotator can not determine its bottom accurately.

In conclusion, although the ground constraint holds most of the time, it is still important to design mechanisms to avoid these possible noises and incorporate the geometric depth adaptively, such as the edge pruning/gating scheme and location-aware integration in the main paper.

2 Quantitative Analysis

Difference Between Datasets Here we mainly show the observation in the ablation study to explain the different effects from the same component on these two datasets. We take the depth score as an example. First, Tab. 7 in the main paper has shown the especially important role of depth score on the KITTI. However, it does not contribute much to the improvements on nuScenes. Specifically, it only brings about 0.3% increase on NDS by reducing the mATE instead of boosting the mAP. To figure out the reason, we take a closer look at the performance from the Precision-Recall (PR) curve. As shown in Fig. 9, we can see that the depth score (solid line) significantly improves the precision under low recall and strict matching thresholds (like 0.5 and 1.0 meters, blue and yellow lines) while influences the performance under high recall and less strict cases (like 2.0 and 4.0 meters, green and red lines). This problem is especially notable for large objects. It reveals the effect of depth score from another perspective, i.e., it can overly suppress those predictions with inaccurate depth, of which we should be tolerant under some circumstances, like distant and small objects. Therefore, designing a more suitable depth score with better interval division methods or other approaches can be a direction worthy of further exploration.

Mean AP for Multi-Class Detection on nuScenes To present the multi-class detection results more comprehensively, we provide the mean AP results (over all the matching thresholds) for each category on nuScenes in Tab. 9. We can see that our method shows the superiority especially on small (from pedestrian to barrier) and quite large objects (bus). Firstly, the better capability of handling objects with different scales should partly come from the leveraged well-developed backbone and FPN. Furthermore, our probabilistic and geometric depth also improves the accuracy of depth estimation, which is especially important for small objects.

Contributions of Each Depth Estimation To understand the role of each component for depth estimation more clearly, we make statistics about the fusion weights. Firstly, for local depth estimation, we find that the direct regression accounts for about 25.6% in the results, i.e., σ(λ)\sigma(\lambda) is about 0.256. It implies that the direct regression may be responsible for regressing the residual of the probabilistic estimation, which plays an auxiliary but important role according to the ablation study in the main paper (Tab. 7). On the other hand, for final integration, we make statistics for the location-aware weights σ(α)\sigma(\alpha) of predictions with matching ground truths before NMS on the validation set, and plot its distribution in Fig. 11 (higher value means more contribution from local estimation). We can see that although the preliminary local estimation plays a more important role in many cases, the propagated geometric depth does contribute a lot to the overall estimation. In addition, we also plot the scatter diagram of these weights with respect to the estimated depth and different categories (Fig. 11 and 12). We can see that the geometric depth contributes more to the estimation of very nearby (can be truncated in the image) and small objects like pedestrians, which is consistent with our common sense that these two cases are relatively hard such that we need to incorporate some contextual information in the reasoning procedure.

Ablation Studies for Alternative Depth Division Methods We also made ablation studies for alternative probabilistic depth settings, including the different settings for the depth unit UU and different division methods to bucket the depth value into intervals. First, Tab. 11 shows that more fine-grained division can not bring performance gains. As for the division methods, we test several alternatives as shown in Tab. 11, among which Log and Linear refer to the spacing-increasing discretization (SID) and linear-increasing discretization (LID) , respectively. We directly take their split points and compute the depth estimation with Eqn. 1 in the main paper. In contrast, Uniform Log means that we take the split points that are uniformly distributed in the log space as the base to compute the depth estimation in the log space with Eqn. 1, and then apply the exponential transformation to get the final result. We can see that although the simplicity, our adopted uniform division method achieves the best performance. Note that this ablation study is conducted with U=10mU=10m. There may be different conclusions if we exploit more fine-grained divisions or use classification and residual regression to implement the probabilistic depth estimation.

Ablation Studies for Geometric Depth Recall that we select three important factors for edge pruning and gating in the depth propagation graph. We also tried other alternatives for the distance score, including the height difference between 3D bottoms, the distance of 3D centers and our adopted 2D centers (Tab. 13). It can be observed that using the 2D centers yields the best performance. We conjecture that it is because the 3D criteria are based on the inaccurate depths such that they are less reliable than the disentangled 2D distance.

Depth Error Analysis We have validated the efficacy of our method in the main paper by comparing the detection performance of our method and the baseline, especially in terms of the improved mean average precision (mAP) and the mean translation error (mATE). Here we further prove its effectiveness with the depth error analysis. We make depth error statistics for the predictions (before NMS) which have corresponding ground truths on the KITTI validation set (Tab. 13). We can observe that our method significantly reduces the mean error of depth estimation, both on the absolute error and relative error ((Abs. and Rel. in Tab. 13).

3 Qualitative Analysis

Then we show some qualitative results on nuScenes in Fig. 13 by drawing the predicted 3D bounding boxes in the six-view images and the top-view point clouds. We compare the results predicted by our model and the baseline FCOS3D to demonstrate the improvements in terms of depth estimation intuitively. We can see that from the perspective of images, both detection results are appealing, especially for some small objects that are not labeled. For example, the barriers in the rear right camera are not labeled but detected by these two models. However, from the bird-eye-view, the depth accuracy of the two methods is notably different, especially for those objects marked with red circles: The accuracy is significantly improved by our proposed method. It is also in line with the quantitative results (the mATE is reduced remarkably) and further validates the efficacy of our method.