Delving into Localization Errors for Monocular 3D Object Detection

Xinzhu Ma, Yinmin Zhang, Dan Xu, Dongzhan Zhou, Shuai Yi, Haojie Li, Wanli Ouyang

Introduction

Remarkable progress has been achieved in 3D detection, especially for LiDAR/stereo-based approaches , along with the advances in deep neural networks. In contrast, the accuracy of 3D detection from only monocular images is obviously lower than that from LiDAR or stereo. In this work, we aim to quantitatively identify the problem and propose our solutions.

To investigate and quantify the underlying factors that restrict the performance of monocular 3D object detection, we conduct intensive diagnostic experiments for this task, inspired by the error identifying methods commonly used in the 2D detection scope. Specifically, we build our baseline model (see Section 3.2 for details) based on CenterNet and progressively replace predicted items with their ground-truth values. To better analyze the error patterns, we evaluate the results in a range-wise manner and show the summary of those experiments in Figure 1. Based on our investigation, we have the following three observations and corresponding designs.

Observation 1: The most striking feature in Figure 1 is the leap in performance when using ground-truth location, reaching a level similar to the state-of-the-art LiDAR-based methods, suggesting the localization error is the key factor in restricting monocular 3D detection. Furthermore, except for depth estimation, detecting the projected center of the 3D object also plays an important role in restoring the 3D position of the object. To this end, we revisit the misalignment between the center of the 2D bounding box and the projected center of the 3D object. Besides, we also confirm the necessity of keeping 2D detection related branches in monocular 3D detector. In this way, 2D detection is used as the correlated auxiliary task to help learning the features shared with 3D detection, which is different from the existing work in that discards 2D detection.

Observation 2: An apparent trend reflected in Figure 1 is that the detection accuracy significantly decreases with respect to the distance (the low performance of very close range objects will be discussed in supplementary materials). More importantly, all the models cannot output any true positive samples beyond a certain distance. We found that it is almost impossible to detect distant objects accurately with existing technologies due to the inevitable localization errors (see Section 4.4 for details). In this case, whether it is beneficial to add these samples into the training set becomes a question. In fact, there is a clear domain gap between ‘bad’ samples and ‘easy-to-detect’ samples and forcing the network to learn from those samples will reduce its representative ability for the others, which will thus impair the overall performance. Based on the observation above, we propose two schemes. The first scheme removes distant samples from the training set and the second scheme reduces the training loss weights of these samples.

Observation 3: We found that, except for localization error, there are also some other vital factors, such as dimension estimation, restricting monocular 3D detection (there is still 27.4% room for improvements even we use the ground-truth location). Existing methods in this scope tend to optimize each component of the 3D bounding box independently, and the studies in confirm the effectiveness of this strategy. However, the failure to consider the contribution of each loss item to the final metric (\ie3D IoU) may lead to sub-optimal optimization. To alleviate this problem, we propose an IoU oriented loss for 3D size estimation. The new IoU oriented loss dynamically adjust the loss weight for each side in sample level according its contribution rate to the 3D IoU.

In summary, the key contributions of this paper are as follows: First, we conduct intensive diagnostic experiments for monocular 3D detection. In addition to finding that the ‘localization error’ is the main problem restricting monocular 3D detection, we also quantify the overall impact of each sub-task. Second, we investigate the underlying reasons behind localization error, analyze the issues it might bring. Accordingly, we propose three novel strategies operating on annotations, training samples, and optimization losses to alleviate problems caused by localization error for boosting the detection.

Experimental results show the effectiveness of the proposed strategies. In particular, compared with existing best-performing monocular 3D object detection approaches, the proposed method achieves at least 1.6 points AP40{\rm AP}_{40} improvements on the bird’s view detection and 3D object detection in the KITTI dataset.

Related Work

Standard monocular 3D detection. Here we briefly review the ‘standard’ monocular 3D detection approaches only use the RGB images, annotations and camera calibrations provided by KITTI dataset. try to improve the representation ability of the models by introducing novel geometric constraints. OFTNet presents an orthographic feature transform to map image-based features into an orthographic 3D space. MonoDIS disentangles the loss for 2D/3D detection and jointly trains these two tasks in an end-to-end manner. M3D-RPN extends the region proposal network (RPN) with 3D box parameters. These works are orthogonal to our analysis to localization error and the proposed strategies for handling it.

Monocular 3D detection using additional data. To better estimate the 3D bounding boxes, many methods are proposed for effectively using additional data . Specifically, use the CAD models as shape templates to get better object geometry. Deep MANTA , which takes 3D detection as a key-points detection task, uses more detailed annotated locations of keypoints, \egwheels, as training labels. Besides, estimate the depth maps from off-the-shelf depth estimators trained from larger datasets, and use them to augment the input RGB images. In addition, propose to transform the estimated depth maps to pseudo-LiDAR representation, before applying existing LiDAR-based 3D detection designs, and achieve promising performance on KITTI benchmark. PatchNet analyzes the underlying mechanism behind pseudo-LiDAR representation and proposes its corresponding image representation based implementation. Recently, Kinematic3D propose to use 3D Kalman filter to capture the temporal cues from monocular videos. In contrast, our method does not use any extra data or annotation, and can still achieve better or competitive performance.

Misalignment between the definitions of object’s center. To recover the 3D object position, there are two groups of methods. The first group use 2D bounding box to obtain 3D position. In particular, CenterNet regards the center of the 2D bounding box as the projected 3D position in the image plane and back-project it to 3D space with the help of estimated depth and camera parameters. However, generally speaking, the center of the 2D box and the center of 3D box are not the same. regress an offset to compensate for the difference between them. As the second group, SMOKE removes the 2D detection and directly estimate 3D position using projected 3D center. This work considers the 2D related sub-tasks are redundant because 2D bounding boxes can be generated from 3D detection results. In this work, we revisit this problem and confirm that replacing the 2D center by the projected 3D center can improve the localization accuracy. Besides, we also find that 2D detection is necessary, because it helps to learn shared features for 3D detection.

Approach

Given are RGB images and the corresponding camera parameters, our goal is to classify and localize the objects of interest in 3D space. Each object is represented by its category, 2D bounding box B2D{\bf B_{2D}}, and 3D bounding box B3D{\bf B_{3D}}. Specifically, B2D{\bf B_{2D}} is represented by its center ci=[x′,y′]2D\mathbf{c^{i}}=[x^{\prime},y^{\prime}]_{2D} and size [h′,w′]2D[h^{\prime},w^{\prime}]_{2D} in the image plane, while B3D{\bf B_{3D}} is defined by its center [x,y,z]3D[x,y,z]_{3D}, size [h,w,l]3D[h,w,l]_{3D} and heading angle γ\gamma in the 3D world space.

2 Baseline Model

Architecture. We build our baseline model based on the anchor-free one-stage detector CenterNet . Specifically, we use standard DLA-34 as our backbone for a better speed-accuracy trade-off. On top of this, seven lightweight heads (implemented by one 3×33\times 3 conv layer and one 1×11\times 1 conv layer) are used for 2D detection and 3D detection. More design choices and implementation details can be found in the supplementary material.

2D detection. For 2D detection task, following , the proposed model outputs a heatmap to indicate the classification score and the coarse center c=(u,v){\bf c}=(u,v) of the object. In existing methods , c{\bf c} is supervised by the ground-truth 2D bounding box center. Another branch predict the offset oi=(Δui,Δvi){\bf o^{i}}=(\Delta u^{i},\Delta v^{i}) between the coarse center and the real center of 2D bounding box, and we can get the final 2D box center location ci=c+oi{\bf c^{i}}={\bf c}+{\bf o^{i}}. Finally, we use another branch to estimate the size [w′,h′]2D[w^{\prime},h^{\prime}]_{2D} of 2D bounding box.

where zz is the output of depth branch. Finally, the last two branches are used to predict the 3D size [h,w,l]3D[h,w,l]_{3D} and orientation γ\gamma, respectively.

Losses. There are seven loss terms in total, one for foreground/background sample classification, two (center and size) for 2D detection, and four (center, depth, size, and heading angle) for 3D detection. We adopt the modified Focal Loss used in for classification sub-task. We use L1 Loss without any anchor for center and size regression in 2D detection task. For the 3D detection task, uncertainty modeling (see Section B in the supplementary for details) is used for depth estimation; L1 loss is used for 3D center refinement; and multi-bin loss (we consider 12 non-overlap equal bins) is used for heading angle estimation. Lastly, for 3D size estimation, we use L1 loss in baseline (without anchor), and the proposed IoU loss in our model. The weights for all loss items are set to 1.

3 Error Analysis

In this section, we explore what restricts the performance of monocular 3D detection. Inspired by CenterNet and CornerNet in the 2D detection field, we conduct an error analysis for different prediction items on KITTI validation set via replacing each predictions with ground truth value and evaluating the performance. Specifically, we replace each output head with its ground truth according to the practice of . As shown in Table 1, if we replace projected 3D center cw{\bf c^{w}} predicted from baseline model with its ground-truth, the accuracy is improved from 11.12% to 18.97%. On the other hand, depth can improve the accuracy to 38.01%38.01\%. If we consider both depth and projected center, \iereplacing the predicted 3D locations [x,y,z]3D[x,y,z]_{3D} with ground-truth results, then the most obvious improvement is observed. Therefore, the low accuracy of monocular 3D detection is mainly caused by localization error. On the other hand, according to Equation 1, depth estimation and center localization jointly determine the position of the object in 3D world space. Compared with the ill-posed depth estimation from a monocular image, improving the accuracy of center detection is a more feasible way.

Table 2 shows localization errors introduced by inaccurate center detection. Furthermore, the mean shape of cars in KITTI dataset is [1.53m,1.63m,3.53m][1.53m,1.63m,3.53m] for [h,w,l]3D[h,w,l]_{3D}. Suppose that all other quantities are correct and the localization error is aligned with the length ll (resulting in the maximum tolerance), the IoU can be computed by:

where Δloc\Delta_{loc} represents the localization error. According to the official setting, the IoU threshold should be set to 0.7, thus the theoretically acceptable maximum error is 0.62m0.62m. However, an error of only 4-8 pixels in the image (1-2 pixel in 4×4\times down sampling feature map) will cause the object at 60 meters cannot be detected correctly. Coupled with the errors accumulated by other tasks such as depth estimation (Figure 3 shows the errors of depth estimation), it becomes an almost impossible task to accurately estimate the 3D bounding box of distant objects from a single monocular image, unless the depth estimation is accurate enough (not achieved to date).

To better show the importance of center localization, we show the localization error in 3D space caused by shifting the center in image plane in Table 2.

4 Revisiting Center Detection

Our design for center detection. For estimating the coarse center c{\bf c}, our design is simple. In particular, we 1) use the projected 3D center cw\mathbf{c^{w}} as the ground-truth for the branch estimating coarse center c{\bf c} and 2) force our model to learn features from 2D detection simultaneously. This simple design is from our analysis below.

Analysis 1. As shown in Figure 4, there is a misalignment between the 2D bounding box center ci\mathbf{c^{i}} and the projected center cw\mathbf{c^{w}} of the 3D bounding box. According to the formulation in Equation 1, the projected 3D center cw\mathbf{c^{w}} should be the key for recovering the 3D object center [x,y,z]3D[x,y,z]_{3D}. The key problem here is what should be the supervision for the coarse center c{\bf c}. Some works choose to use 2D box center ci{\bf c^{i}} as its label, which is not related to the 3D object center, making the estimation of the coarse center not aware of the 3D geometry of the object. Here we choose to adopt the projected 3D center cw\mathbf{c^{w}} as the ground-truth for the coarse center c{\bf c}. This helps the branch for estimating the coarse center aware of 3D geometry and more related to the task of estimating 3D object center, which is the key of localization problem (see Section E in supplementary materials for visualizations).

Analysis 2. Note that SMOKE also use the projected 3D center cw{\bf c^{w}} as the label of the coarse center c{\bf c}. However, they discard 2D detection related branches while we preserve them. In our design, the coarse center c{\bf c} supervised by the projected 3D center cw\mathbf{c^{w}} is also used for estimating the 2D bounding box center ci\mathbf{c^{i}}. With our design, we force a 2D detection branch to estimate an offset oi=ci−c{\bf o^{i}}={\bf c^{i}}-{\bf c} between the real 2D center and the coarse 2D center. This makes our model aware of the geometric information of the object. Besides, another branch is used to estimate the size of the 2D bounding box so that the shared features can learn some cues that benefit to depth estimation due to the perspective projection. In this way, the 2D detection serves as an auxiliary task that helps to learn better 3D aware features.

5 Training Samples

Different from which force network focus on the ‘hard’ samples, we argue that ignoring some extremely ‘hard’ cases can improve the overall performance for the monocular 3D detection task. Both the results shown in Figure 1 and the analysis conducted in Section 4.4 illustrate there is a strong relationship between the distance of the object and the difficulty of detecting it. According to this, two schemes are proposed on how to generate the object-level training weight wiw_{i} for sample ii.

Scheme 1, hard coding. This scheme discard all samples over a certain distance:

where did_{i} denotes the depth of sample ii, and ss is the threshold of depth which is set to 60 meters in our implementation. In this way, the samples with depth larger than ss will not be used in the training phase.

Scheme 2, soft coding. The other one is soft encoding, and we generate it using a reverse sigmoid-like function:

where cc and TT are the hyper-parameters to adjust the center of symmetry and bending degree, respectively. When c=sc=s and T→0T\to 0, it is equivalent to the hard encoding scheme. When T→∞T\to\infty, it is equivalent to using the same weight for all samples. By default, cc and TT are set to 60 and 1, and the empirical experiments in Section 4 find that scheme 1 and scheme 2 are both effective and have similar results.

6 IoU Oriented Optimization

Recently, some LiDAR based 3D detectors applied the IoU oriented optimization . However, determining the 3D center of object is an very challenging task for monocular 3D detection, and the localization error often reaches several meters (see Section 4.4). In this case, localization related sub-tasks (such as depth estimation) will overwhelm others (such as 3D size estimation), if we apply IoU based loss function directly. Moreover, depth estimation from monocular image itself an ill-posed problem, and this kind of contradiction will make the training process collapse. Disentangling each loss item and optimize them independently is a another choice , but this ignores the correlation of each component to the final result. To alleviate this problem, we propose a IoU oriented optimization for 3D size estimation. Specifically, suppose all prediction items except the 3D size s=[h,w,l]3D{\bf s}=[h,w,l]_{3D} are completely correct, then we can get (details for deriving can be found in supplementary materials):

Accordingly, we can adjust the weight of each side by its partial derivative \wrtIoU (in magnitude), and the loss function of the 3D size estimation can be modified to:

where ∣∣⋅∣∣1||\cdot||_{1} represent the L1L_{1} norm. Note that, compared with the standard 3D size loss Lsize′=∣∣s−s∗∣∣1\mathcal{L}_{size}^{\prime}=||{\bf s}-{\bf s^{*}}||_{1} used in the baseline model, our new loss’s magnitude is changed. To compensate it, we compute Lsize′\mathcal{L}_{size}^{\prime} once more, and dynamically generate the compensate weight ws=∣Lsize′/Lsize∣w_{s}=|\mathcal{L}_{size}^{\prime}/\mathcal{L}_{size}|, so that the mean value of the final loss function ws⋅Lsizew_{s}\cdot\mathcal{L}_{size} is equal to the standard one. By this way, the proposed loss can be regard as a re-distribution of the standard L1 loss.

7 Implementation

Training. We train our model on two GTX 1080Ti GPUs with a batch size of 16 in an end-to-end manner for 140 epochs. We use Adam optimizer with initial learning rate 1.25e−31.25e^{-3}, and decay it by ten times at 90 and 120 epochs. The weight decay is set to 1e−51e^{-5} and the warmup strategy is also used for the first 5 epochs. To avoid over-fitting, we adopt the random cropping/scaling (for 2D detection only) and random horizontal flipping. Under this setting, it takes around 9 hours for whole training process.

Inference. During the inference phase, we obtain the prediction results from the parallel decoders. To decoding the results, similar to , we conduct the efficient non-maxima suppression (NMS) on center detection results using a 3×33\times 3 max pooling kernel. Then, we recover 2D/3D bounding boxes according to encoding strategy introduced in Section 3.2 and use the score of center detection as the confidence of predicted results. Finally, we discard predictions with confidence less than 0.2.

Experimental Results

Dataset. We evaluate our method on the challenging KITTI dataset , which provides 7,481 images for training and 7,518 images for testing. Since the ground truth for the test set is not available and the access to the test server is limited, we follow the protocol of prior works to divide the training data into a training set (3,712 images) and a validation set (3,769 images). We conduct ablation studies based on this split and also report final results which trained on all 7,481 images and tested by KITTI official server.

Metrics. The KITTI dataset provides many widely used benchmarks for autonomous driving scenarios, including 3D detection, bird’s eye view (BEV) detection, and average orientation similarity (AOS). We report the Average Precision with 40 recall positions (AP40{\rm AP}_{40}) under three difficultly settings (easy, moderate, and hard) for those tasks. We mainly focus on the Car category, and also report the performances of the Pedestrian and Cyclist categories for reference. The default IoU threshold are 0.7, 0.5, 0.5 for these categories.

2 Main Results

Results on the KITTI test set. As shown in Table 3, we report our results of the Car category on KITTI test set. Overall, our method achieves superior results over previous methods across all settings under fair conditions. For instance, the proposed method obtains 2.47/2.27/1.64 improvements under easy/moderate/hard setting for 3D detection task. Besides, our method achieves 18.89/90.23 in BEV detection/AOS task under moderate setting, improving previous best results by 4.06/4.12 AP40{\rm AP}_{40}. Compared with the methods with extra data, the proposed method still get comparable performances, which further proves the effectiveness of our model.

Results on the KITTI validation set. We also present our model’s performance on the KITTI validation set in Table 4. Note that some methods directly use the pre-trained model provided by DORN as their depth estimator. However, the DORN’s training set overlaps with the validation set of KITTI 3D, so we are not comparing these methods here. We can find that the proposed model performs better than all previous methods in 3D detection task. For BEV detection task, our method outperforms all methods except for MonoPair. Compared with MonoPair, our method is better at detecting objects under strict conditions (0.7 IoU threshold), while MonoPair is slightly better at catching samples under loose conditions (0.5 IoU threshold). Also note that our method shows better performance consistency between the validation set and test set. This indicates that our method has better generalization ability, which is of great significance in autonomous/assisted driving.

Latency analysis. We test the proposed model on a single GTX 1080Ti GPU with a batchsize of 1 for runtime analysis. As shown in Table 3, the proposed method can run at 25 FPS, meeting the requirement of real-time detection. Specifically, our method runs 4×4\times faster than the two-stage detector M3D-RPN. Compared with MonoPair, which shares a similar framework as ours, our method can still save 16 ms for one image in the inference phase, mainly because: 1) we use standard DLA-34 as our backbone, instead of modified DL4-34 with DCN . 2) we apply fewer prediction heads in our model. 3) we don’t need any post-processing. SMOKE can run faster than our method. However, it only conducts 3D detection while the proposed method can perform 2D detection and 3D detection jointly.

Besides, although the detectors with pretrained depth estimator usually have promising performance, the additional depth estimator introduce lots of computational overheads (\egthe most commonly used DORN takes about 400 ms to process a standard KITTI image. See KITTI Depth Benchmark for more details).

3 Pedestrian/Cyclist Detection

Here we present the Pedestrian/Cyclist detection results on the KITTI test set in Table 5. Compared with cars, pedestrians/cyclists are more difficult to detect, and only provide the performances of those categories on KITTI test set. Specifically, the proposed method performs better than and gets comparable results with . But it is important to note that, since the number of training samples for those two categories is quite small, the performance may fluctuate to some extent.

4 Analysis

Accumulation of the proposed designs. Table 6 shows experimental results evaluating how the proposed designs contribute to the overall performance for this task. Our design in Section 3.4, which uses projected 3D center for supervising center detection and influencing 2D detection (‘+p.’ in Table 6), improves 3D detection accuracy by 1.5. The IoU loss design in Section 3.6 further improves the accuracy by 0.3. And the design for discarding distant samples in Section 3.5 leads to 0.7 improvement.

Supervision for coarse center detection and multi-task learning. We show the performance changes caused by center definition and multi-task learning in Table 7. Specifically, from setting a (used in ) and setting b (used in ) in the table, predicting an offset to compensate for the misalignment between 2D center and projected 3D center can improve the performance of 3D detection significantly. Then, using projected 3D center as the ground truth for coarse detection (setting d, our model) can further improve the performance. Besides, by comparing the setting c used in and the setting d in our design, we can find the performance of 3D detection benefits from multi-task learning (performing 2D detection and 3D detection jointly). Note that the accuracy of 2D detection under setting d is also better than that under setting c, which suggests generating 2D bounding boxes from 3D detection may reduce the quality of the 2D detection results. The above conclusions are also reflected in Table 3 and 4.

Training samples. From Table 8, we can find that both removing some samples from training set appropriately and reducing the training weights of them can improve overall performance. Note that those samples are only a small part of the whole training set and will not affect the representation learning of the network to the whole dataset. For example, in the 7,481 images in trainval set, only 1,301/767 samples beyond 60/65 meters, accounting for 4.5%/2.7% of the total 28,742 samples.

5 Qualitative Results

We visualize some representative outputs of the proposed method in Figure 5. To clearly show the object’s position in the 3D world space, we also visualize the LiDAR signals. We can observe that our model outputs remarkably accurate 3D bounding boxes for the cases at a reasonable distance. We also find that our model outputs some false positive samples, \egthe 3D box on the right in the sixth picture, and the foremost reason for that is the imprecise depth or center estimation. Note that the dimension and orientation estimation for those cases are still accurate.

Conclusion

In this paper, we systematically analyze the problems in monocular 3D detection and find the localization error is the bottleneck of this task. To alleviate this problem, we first revisit the misalignment between the center of the 2D bounding box and the projected center of 3D object. We argue that directly detecting projected 3D center can reduce the localization error and 2D detection is conducive to optimize 3D detection. Besides, we also find distant samples are almost impossible to detect accurately with the existing technologies, and discarding these samples from the training set will stop them from distracting the network. Finally, we also proposed an IoU oriented loss for 3D size estimation. Extensive experiments on the challenging KITTI dataset show the effectiveness of the proposed strategies.

Acknowledgement

This work was supported by SenseTime, the Australian Research Council Grant DP200103223, and Australian Medical Research Future Fund MRFAI000085.

References

Supplementary Material

A Overview

This document provides additional technical details, experimental results, theoretical analysis, and qualitative results to the main paper. Specifically, in Section B, we provide more details on the implementation of the depth estimation sub-task, and Section C shows the details and ablations about the proposed IoU oriented loss. Section D provides more discussion which is omitted in the main paper. Finally, Section E presents more visual results.

B Depth Estimation

Uncertainty modeling. Following , we model the heteroscedastic aleatoric uncertainty in the depth estimation sub-task. Specifically, we simultaneously predict the depth d\mathbf{d} and the standard deviation σ\mathbf{\sigma} (or variance σ2\mathbf{\sigma}^{2}):

where x\mathbf{x} is the input data and ff is a convolutional neural network parametrised by the parameters w\mathbf{w}. Then, we fix a Laplace likelihood to model the uncertainty, and the loss for the depth estimation sub-task can be formulated by:

where ∣∣⋅∣∣1||\cdot||_{1} denotes the L1 norm and d∗\mathbf{d}^{*} is the ground truth value for depth d\mathbf{d}. Similarly for the Gaussian likelihood:

where ∣∣⋅∣∣2||\cdot||_{2} denotes the L2 norm (please refer to for the derivation of Equation 8 and Equation 9). Note that the uncertainty modeling is not claimed as our contribution.

Experimental results. First, from Figure 6 and Table 9, we can find that uncertainty-based estimation improves the accuracy of depth map, thereby improving the overall performance of monocular 3D detection. Second, the experimental result also show that modeling uncertainty based on the Laplace distribution (all models in the main paper adopted this setting) is more suitable for our task than Gaussian distribution.

C IoU Oriented Loss

This section provides the proof of the following proposition, which is used in Equations 5 and 6 for IoU oriented optimization in Section 3.6.

Proposition. Suppose all predicted items except the 3D sizes (h,w,l)(h,w,l) are completely correct, the contribution ratio of each predicted side to the 3D IoU ∂IoU∂h:∂IoU∂w:∂IoU∂l\frac{\partial IoU}{\partial h}:\frac{\partial IoU}{\partial w}:\frac{\partial IoU}{\partial l} can be approximated to 1h:1w:1l\frac{1}{h}:\frac{1}{w}:\frac{1}{l}.

Proof. Given the above conditions, the 3D IoU metric can be formulated as:

where (h∗,w∗,l∗)(h^{*},w^{*},l^{*}) denotes the ground truth of 3D size (h,w,l)(h,w,l). With the different relationship between the prediction and the ground truth of the 3D size, we can obtain the following cases:

Case 1: If h≤h∗h\leq h^{*}, w≤w∗w\leq w^{*}, and l≤l∗l\leq l^{*}, the Equation 10 can be simplified as:

and we further compute the partial derivative of 3D IoU with respect to the variable hh as

where ∂IoU∂h\frac{\partial IoU}{\partial h} represents the partial derivative of 3D IoUIoU with respect to the variable hh, analogically for ∂IoU∂w\frac{\partial IoU}{\partial w} and ∂IoU∂l\frac{\partial IoU}{\partial l}. Then, combining the derivative of 3D IoU with respect to hh, ww, and ll, the contribution ratio of each predicted side can be given as:

Case 2: If h>h∗h>h^{*}, w>w∗w>w^{*}, and l>l∗l>l^{*}, the Equation 10 can be simplified as:

and similar to Equation 12 and 13, we can derive the same conclusion as Case 1.

Case 3: If h>h∗h>h^{*}, w≤w∗w\leq w^{*}, and l≤l∗l\leq l^{*}, then we represent the 3D IoU as:

By calculating the derivative of 3D IoU with respect to hh, ww, and ll respectively, we can get the contribution ratio of each predicted side:

Case 4: If h>h∗h>h^{*}, w>w∗w>w^{*}, and l≤l∗l\leq l^{*}, similarly, we can get the IoU formulation as:

Similar to previous steps, the formulation of each side’s contribution rate to the 3D IoU is given as:

The other cases are similar to Case 3 and Case 4. When h≈h∗h\approx h^{*}, w≈w∗w\approx w^{*}, and l≈l∗l\approx l^{*}, we can get the Equation 5 used in the main paper.

C.2 Experiments

We report the improvement introduced by the proposed loss function in the main paper. To further validate the effectiveness of it, we also implement the 3D GIoU loss for reference. Specifically, we add the 3D GIoU loss as a regularization item as in , investigating different weights considered in our baseline model, and the AP40{\rm AP}_{40} of cars on the moderate setting on KITTI validation set (Table 10) show that our IoU oriented optimization improves accuracy but 3D-GIoU with different weights does not.

D Performance for the Close Objects

The Figure 1 in the main paper provides lots of insights to us. Except for the observations analyzed in the main paper, we also found that the performance degrades for the very close object. Here we provides our analysis for this. In particular, there are three main reasons in total. a) The close-range objects tend to have larger center misalignment (see Figure 3 for the statistics). b) The objects at closer ranges are usually more truncated, \egthe red car (depth=3.7, truncation=0.88) and the black car (depth=6.2, truncation=0.34) in Figure 7. c) The training samples in the close range are fewer. For example, there are 5,979 cars in [5m,15m][5m,15m] and 6,707 cars in [10m,20m][10m,20m] on the KITTI trainval set, and the distribution for those samples are summarized in Table 11. Note that the KITTI annotate the difficulty of each samples according to its size of 2D bounding box, occlusion, and truncation. The instance with ‘unKnown’ tag usually means that it is extremely difficult to detect and is ignored in evaluation. With that in mind, the effective samples of those two ranges are 4,522 and 6,149. In summary, the low performance of the very close objects is caused by the limited training samples (c) and the large proportion of hard cases (a, b).

E More Visualizations

From Figure 4 in the main paper, we can see there is a misalignment between the center of the 2D bounding box and the projected center of the 3D object, especially for close objects (see Figure 3 and Figure 4). Accordingly, we propose our solution for this problem. Here we visualize the learned features of coarse center detection branch in Figure 7 to show the effectiveness of the proposed method. The qualitative results clearly show that using projected 3D center as ground truth can make the coarse center more accurate, thereby improving the localization accuracy.

E.2 Comparison of qualitative results

Visualizations in the image plane. We show more qualitative results of M3D-RPN (the best of all open-source standard monocular 3D detector) and the proposed method in Figure 8. We use red circle to highlight the main differences of each pair of images, and we can find that our method performs better than M3D-RPN for dense objects.

Visualizations in the 3D world space. We also visualize the 3D bounding boxes in the 3D world space for better presentation. As shown in Figure 9, the proposed model outputs better results than M3D-RPN, especially for the orientation estimation.

Representative failure case. We show a typical error pattern in monocular 3D object detection in Figure 10. We can observe that the projected 3D bounding boxes fit the object’s appearance tightly in the image plane. However, from the visualization results in the 3D world space, this is a clear false positive because the depth is inaccurate (the outline of the object can be perceived through the point clouds, best viewed with zooming in). Note that this problem is common in the monocular 3D detection task, which suggests that depth estimation is a key factor restricting this task.