Geometry-based Distance Decomposition for Monocular 3D Object Detection

Xuepeng Shi, Qi Ye, Xiaozhi Chen, Chuangrong Chen, Zhixiang Chen, Tae-Kyun Kim

Introduction

Object detection is a fundamental and challenging problem in computer vision. With the emergence of deep learning , 2D object detection has achieved great progress in the past years . However, it is yet insufficient for applications requiring 3D spatial information like autonomous driving. 3D object detection, which detects objects as 3D bounding boxes, has drawn much attention. Comparing with 3D object detection methods relying on expensive LiDAR sensors to provide depth information , monocular 3D object detection infers depth from monocular images with low computation and energy cost. The core challenge of monocular 3D object detection is inferring the distance of objects in the absence of explicit depth information. Given the visual appearance of an object, its spatial location can be inferred based on the imaging geometry as an inverse problem. Thus, the priors of object physical size, scene layout, and the imaging process of cameras are essential to exploit to recover the distance.

On one hand, such geometric priors have been exploited to predict the pose or distance of objects by their factors. In 6D object pose estimation, PVNet and SegDriven regress the 2D keypoints of objects. In category-level 6D object pose and size estimation, NOCS uses the normalized object coordinate space map. In stereo 3D object detection, Stereo R-CNN uses sparse 2D keypoints, yaw angle, object physical size, and a region-based photometric alignment using the left and right RoIs. These works recover the pose or distance by several factors, such as 2D keypoints, 2D bounding boxes, and object physical size, which achieves interpretable and robust pose or distance estimation.

On the other hand, most monocular 3D object detection methods deal with the challenging distance prediction by regressing it as a single variable. The learning-based methods directly learn a mapping from input images to the distance. The pseudo-LiDAR based methods first regress the depth map of an input image and then predict the distance of objects with the depth map. The 3D-anchor-based methods break the distance prediction into the region proposal and the offset regression. The only exceptions are , which recover the distance by minimizing the re-projection error between 3D bounding boxes and 2D bounding boxes or 2D keypoints. However, still lag behind those methods which regress the distance as a single variable.

Aiming to close this gap, we propose a novel geometry-based distance decomposition to recover the distance by its factors. Different from , we abstract objects as vertical lines at the center of 3D bounding boxes, and their visual projection as the projection of these vertical lines, then recover the distance by them based on the imaging geometry , as shown in Fig. 1. The decomposition is designed to be as simple as possible, yet effective and efficient to extract the most representative and stable factors of the distance of objects, i.e., the physical height and the projected visual height. The advantages of the decomposition are four-fold. 1) It makes the distance prediction interpretable. The physical height can be interpreted as an intrinsic attribute of objects, and the visual height can be interpreted as the extrinsic position in a scene. 2) The physical height and projected visual height are easy to estimate as revealed by our observations. 3) The decomposition maintains the self-consistency between the two heights, leading to robust distance prediction when both predicted heights are inaccurate. 4) The decomposition enables us to trace and interpret the causes of the distance uncertainty for different scenarios, by introducing an uncertainty-aware regression loss for the decomposed variables. In addition, our method can generalize to images with different camera intrinsics, because it reasons the distance only by the local information of objects and decouples the focal length from the distance prediction. This generalization ability is crucial to facilitate the deployment of the machine learning models of monocular 3D vision .

The contributions of our method are summarized below:

A novel geometry-based distance decomposition makes the distance prediction interpretable, accurate and robust.

Based on the decomposition, our method originally traces the causes of the distance uncertainty.

Our method directly predicts 3D bounding boxes from RGB images with a compact architecture, making the training and inference simple and efficient.

Our method achieves the state-of-the-art (SOTA) performance on the monocular 3D Object Detection and Bird’s Eye View tasks of the KITTI dataset , and can adapt to images with different camera intrinsics.

Related Work

2D object detection has achieved sustainable improvements in the past years. Notably, the two-stage frameworks, such as Faster R-CNN and Mask R-CNN , achieve dominated performance on several challenging datasets . Feature pyramid networks (FPN) has also been proposed to improve the 2D object detection performance. We adopt Faster R-CNN with FPN as our 2D object detection framework, because of its high accuracy and flexibility.

2 Monocular 3D Object Detection

Most monocular 3D object detection methods deal with the challenging distance prediction by regressing it as a single variable. The learning-based methods directly regress the distance of objects by adding distance branches to 2D object detectors, which are simple and efficient. The pseudo-LiDAR-based methods first predict the depth map of an input image using an external monocular depth estimator, then predict the distance of objects from the estimated depth map using a point-cloud-based 3D object detector. Though the explicit depth cues from the the estimated depth map can ease the distance prediction, the generalization of these methods is bounded by that of the monocular depth estimators . The 3D-anchor-based methods extend the 2D anchor boxes to the 3D anchor boxes by supplementing 3D bounding box templates, then predict the transformations from the 3D anchor boxes to the ground-truth 3D bounding boxes. The 3D anchor boxes can ease the distance learning. Unlike regressing the distance as a single variable in these methods, we propose a novel geometry-based distance decomposition to recover the distance by its factors.

A few works predict the distance of objects by its factors. Deep3Dbox abstracts objects as 3D bounding boxes and their visual projection as the four boundaries of projected 3D bounding boxes, then recovers the distance by minimizing the re-projection error between the four boundaries of projected 3D bounding boxes and 2D bounding boxes. The keypoint-based methods abstract objects as 3D bounding boxes and their visual projection as the eight projected corners of 3D bounding boxes, then recovers the distance by minimizing the re-projection error between the eight projected corners of 3D bounding boxes and the predicted eight projected corners. However, still lag behind those methods which regress the distance as a single variable. Similar to , our method also recovers the distance by its factors, but our decomposition is more simple and effective.

3 Geometry-based Object Pose Estimation

Geometric priors have been exploited to predict the pose or distance of objects by their factors. In 6D object pose estimation, PVNet and SegDriven regress 2D keypoints of objects, then optimize the estimation of the 6D pose by solving a Perspective-n-Point (PnP) problem. In category-level 6D object pose and size estimation, NOCS uses the normalized object coordinate space map together with the depth map in a pose fitting algorithm, to estimate the 6D pose and physical size of unseen objects. In stereo 3D object detection, stereo R-CNN first calculates the coarse distance from sparse 2D keypoints, yaw angle, and object physical size, then recovers the accurate distance by a region-based photometric alignment using the left and right RoIs. Inspired by these works, we propose a geometry-based distance decomposition for monocular 3D object detection, to recover the distance by its factors.

4 Uncertainty Estimation

There are two seminal works exploring uncertainties in deep learning for computer vision. The uncertainty-aware regression loss enables networks to re-balance samples and re-focus on more reasonable samples, which improves the overall accuracy. MonoLoco , MonoPair and UR3D regress the distance of objects with the uncertainty-aware regression loss to improve the distance regression accuracy. MonoDIS proposes a self-supervised confidence score to re-sort the predicted 3D bounding boxes. Kinematic3D proposes a self-balancing 3D confidence loss to both improve the 3D box regression accuracy and re-sort the predicted 3D bounding boxes. directly apply the uncertainty-aware losses to the distance. Instead, we apply the uncertainty-aware regression loss to the decomposed variables of the distance, which enables us to trace the causes of distance uncertainty for different scenarios.

Proposed MonoRCNN

We first present the basic framework, then two 3D-related detection heads, i.e., the 3D distance head and 3D attribute head. We detail the geometry-based distance decomposition and uncertainty-aware regression in the 3D distance head. We term our method as MonoRCNN, and the main architecture is illustrated in Fig. 2.

We address monocular 3D object detection, which predicts the 3D bounding boxes of objects from monocular RGB images. Two common assumptions are 1) only considering the yaw angle of 3D bounding boxes and setting the roll and pitch angle as zero, 2) per-image camera intrinsics are available both during training and inference. For a given RGB image, MonoRCNN reports all objects within concerned categories, and the output for each object is

class label cls\mathit{cls} and confidence score\mathit{score},

2D bounding box represented by the top-left and bottom-right corners, denoted as b=(x1,y1,x2,y2)\mathbf{b}=(x_{1},y_{1},x_{2},y_{2}),

the 2D projected center of the 3D bounding box, denoted as p=(p1,p2)\mathbf{p}=(p_{1},p_{2}),

the physical size of the 3D bounding box, denoted as m=(W,H,L)\mathbf{m}=(W,H,L), where W,H,LW,H,L are the physical width, height, and length, respectively,

the yaw angle of the 3D bounding box, denoted as a=(sin⁡(θ),cos⁡(θ))\mathbf{a}=(\sin(\theta),\cos(\theta)), where θ\theta is the allocentric pose of the 3D bounding box,

the distance of the center of the 3D bounding box, denoted as ZZ.

MonoRCNN predicts the 3D center (p1,p2,Z)(p_{1},p_{2},Z) in pixel coordinates, and convert it to camera coordinates using the projection matrix P\mathbf{P} during inference, formulated as

For the yaw angle prediction, MonoRCNN predicts sin⁡(θ)\sin(\theta) and cos⁡(θ)\cos(\theta), and convert them to θ\theta during inference.

MonoRCNN is built upon Faster R-CNN . We use a ResNet-50 with FPN as the backbone and RoIAlign to extract the crops of object features. For the training and inference of the 2D object detection network, we follow the pipelines in . To adapt to monocular 3D object detection, we add the 3D distance head and 3D attribute head.

2 3D Distance Head

The 3D distance head recovers the distance of objects and is based on our geometry-based distance decomposition. Specifically, we decompose the distance of an object ZZ, into the physical height HH, and the reciprocal of the projected visual height hrec=1hh_{rec}=\frac{1}{h}, which is formulated as

where ff denotes the focal length of the camera. We regress HH and hrech_{rec} separately and recover ZZ by them.

The decomposition makes the distance prediction interpretable. HH can be interpreted as an intrinsic attribute of objects, the estimation of which can be regarded as a fine-grained object classification problem. While given an object, hrech_{rec} can be interpreted as the extrinsic position in a scene, the estimation of which is a 2D regression problem in the image plane.

The physical height and visual height required by our decomposition are easy to estimate. Predicting the eight projected corners of 3D bounding boxes is challenging due to the occlusion, truncation, yaw angle variations, and extreme lighting conditions. As shown in Fig. 3, the visual height prediction is accurate in different challenging cases but the projected corner prediction fails. Moreover, the physical height is the simplest and the most stable variable among the physical size, according to the prediction error of the physical size of cars on the val subset of the KITTI validation split , shown in Tab. 1. The mean prediction error of the physical height is much smaller than that of the physical length. In addition, the prediction error of the physical length and width are influenced by the yaw angle due to single-view ambiguity, while the prediction error of the physical height is not. Our decomposition only uses the physical height, instead of the full physical size , to recover the distance, which improves the distance prediction accuracy.

The decomposition can also maintain the self-consistency during inference, leading to robust distance prediction when both predicted heights are inaccurate. With the objective of predicting the distance, the neural network can learn the correlation between HH and hrech_{rec} during training, then the learned correlation can serve as the self-consistency during inference. This correlation is detailed as below. In the context of monocular 3D object detection task, the physical height HH is a constant for a given object. The distance of the object to the camera ZZ can be modeled as a random variable since the object can appear in different locations in a scene. Similarly, the reciprocal of the length of the PCL of the object hrech_{rec} is also a random variable. While the variable ZZ is random for a specified object, we notice that this variable is expected to follow a same distribution for different objects, denoted as D\mathcal{D}. This is because the distribution of the location of an object is irrelevant to its fine-grained object type. For example, the spatial positions of cars on a street are not influenced by their car types. We formulate this as

By taking expectation on Eq. (3), we have

Eq. (4) shows that, for the training labels of different objects, the product between their HH and their expectation of hrech_{rec} is a constant. In other words, for different objects, the expectation of hrech_{rec} decreases with the increase of HH. Intuitively speaking, the larger the physical height of an object, the larger the average projected visual height of this object. This is the correlation between the training labels of HH and hrech_{rec}. The neural network can learn this correlation during training, as shown in Fig. 4. During inference, if the predicted HH is larger than the groundtruth, the learned correlation pushes the predicted hrech_{rec} to become smaller on average, and vice versa. Thus, our method can recover the accurate distance ZZ with the inaccurate HH and hrech_{rec}, i.e., maintain the self-consistency during inference.

2.2 Uncertainty-aware Regression

Based on the decomposition, we further trace and interpret the causes of the distance uncertainty for different scenarios. We modify the uncertainty-aware regression loss to regress HH and hrech_{rec}. The loss functions for HH and hrech_{rec} can be formulated as

where H^\hat{H} and h^rec\hat{h}_{rec} are the groundtruths, HH and hrech_{rec} are the predictions, λH\lambda_{H} and λhrec\lambda_{h_{rec}} are the positive parameters to balance the uncertainty terms, and σH\sigma_{H} and σhrec\sigma_{h_{rec}} are the learnable variables of uncertainties.

In Fig. 5, we show the uncertainties of the physical height and the projected visual height, i.e., σH\sigma_{H} and σhrec\sigma_{h_{rec}}, for objects at different distances on the val subset of the KITTI validation split . For both HH and hrech_{rec}, the uncertainties first decrease and then increase when objects go far away from the camera. At a close distance, the high uncertainties are mainly caused by the truncated views of objects. It is hard to make accurate prediction with partial observations. At a far distance, the high uncertainties are mainly caused by the coarse views of objects with fewer pixels representing them in images. The truncated and coarse views result in comparable increases for the uncertainties of HH. However, hrech_{rec} is with noticeable higher uncertainties for the coarse views than the truncated views. In other words, the accuracy of the distance estimation for faraway objects is heavily influenced by the accuracy of hrech_{rec}.

σhrec\sigma_{h_{rec}} is more discriminative for objects at different distance than σH\sigma_{H}, and fHσhrecfH\sigma_{h_{rec}} can represent the distance uncertainty. We use scorefHσhrec\frac{score}{fH\sigma_{h_{rec}}}, instead of scorescore, to sort the predicted boxes to improve the 3D object detection accuracy.

3 3D Attribute Head

3D attribute head predicts the physical size, yaw angle, and 2D keypoints, i.e., the projected center and corners of 3D bounding boxes. We use the L1L_{1} loss to directly regress the physical size and yaw angle, formulated as

where m^\hat{\mathbf{m}} and a^\hat{\mathbf{a}} are the groundtruths, and m\mathbf{m} and a\mathbf{a} are the predictions. For the keypoint regression, we normalize the keypoints by their proposal size. Let (x1,y1,x2,y2)(x_{1},y_{1},x_{2},y_{2}) denote the top-left and bottom-right corners of a proposal, and p^=(p^1,p^2)\hat{\mathbf{p}}=(\hat{p}_{1},\hat{p}_{2}) and p=(p1,p2)\mathbf{p}=(p_{1},p_{2}) denote the groundtruth keypoint and the predicted keypoint, respectively. Let t^\hat{\mathbf{t}} and t\mathbf{t} denote the normalized groundtruth keypoint and the normalized predicted keypoint, respectively, and t^\hat{\mathbf{t}} is defined as

The keypoint loss function can be formulated as

During inference, we transform the normalized predicted keypoint t\mathbf{t} to the predicted keypoint p\mathbf{p}. We only use the projected center of a 3D bounding box during inference, the losses of the eight projected corners are auxiliary losses during training.

4 Overall Loss

The overall training loss function for the detection heads is

where λcls\lambda_{cls} is 11, λbbox\lambda_{bbox} is 11, λsize\lambda_{size} is 33, λyaw\lambda_{yaw} is 55, λkpt\lambda_{kpt} is 55, λH\lambda_{H} is 0.250.25, and λhrec\lambda_{h_{rec}} is 11.

5 Implementation Details

The backbone of MonoRCNN is ResNet-50 with FPN and is pre-trained on ImageNet . We extract ROI features from P2, P3, P4 and P5 of the backbone, as defined in . We use five scale anchors of {32, 64, 128, 126, 512} with three ratios {0.5, 1, 2}, and tile the anchors on P4. Images are scaled to a fixed height of 512512 pixels for both training and inference. During training the batch size is 44, and the total iteration number is 1\times1051\text{\times}{10}^{5} and 2\times1052\text{\times}{10}^{5} on the training subset of the KITTI validation split and the KITTI official test split , respectively. We adopt the step strategy to adjust the learning rate. The initial learning rate is 0.010.01 and reduced by 1010 times after 60%60\%, 80%80\%, 90%90\% iterations. During training random mirroring is used as augmentation, and during inference no augmentation is used. We implement our method with PyTorch and Detectron2 . All the experiments run on a server with 2.22.2 GHz CPU and GTX Titan X.

Experiments

We first analyze the ablation studies and self-consistency on the KITTI validation split . Then we comprehensively benchmark MonoRCNN on the KITTI official test dataset . We also present the cross-dataset test results using a nuScenes cross-test set. Finally, we visualize qualitative examples on the KITTI dataset in Fig. 6, and the nuScenes cross-test set in Fig. 7 More qualitative examples can be found in the supplementary file..

The KITTI dataset provides multiple widely used benchmarks for computer vision problems in autonomous driving. The Bird’s Eye View (BEV) and 3D Object Detection tasks are used to evaluate the 3D localization performance. These two tasks are characterized by 74817481 training and 75187518 test images with 2D and 3D annotations for cars, pedestrians, cyclists, etc. Each object is assigned with a difficulty level, i.e., easy, moderate or hard, based on its visual size, occlusion level and truncation degree. We conduct experiments on two common data splits, the val split and the official test split . We only use the images from the left cameras for training. We report the AP∣R40\text{AP}|_{R_{40}} to compare the accuracy. We use the car class, the most representative class, and the official IoU criteria 0.70.7 for cars.

The nuScenes 3D object detection task requires detecting 1010 object classes in terms of full 3D bounding boxes, attributes and velocities. In this work, we focus on detecting the 3D bounding boxes of cars, to test the cross-dataset performance from KITTI to nuScenes . We use the script https://github.com/nutonomy/nuscenes-devkit/blob/master/python-sdk/nuscenes/scripts/export_kitti.py for converting nuScenes data to KITTI format to generate a cross-test set. The cross-test set consists of 60196019 front camera images from the official val subset.

2 Ablation Studies

We conduct ablation studies to examine how each proposed component affects the final performance. We evaluate the performance by first setting a baseline which utilizes our decomposition, then adding the uncertainty-aware loss , finally sorting the predicted boxes by scorefHσhrec\frac{score}{fH\sigma_{h_{rec}}}, as shown in Tab. 2. We also implement a model which directly regresses the distance to compare our method with the learning-based methods , and a model which recovers the distance by the eight projected corners and physical size to compare our method with the keypoint-based methods . From Tab. 2, we can see:

1) The geometry-based distance decomposition is effective. Comparing ‘D’ with ‘K’, we can see the geometry-based distance decomposition outperforms the keypoint-based model by a large margin, which supports its effectiveness.

2) The uncertainty-aware regression loss improves the accuracy. Comparing ‘D+U’ with ‘D’, we can see using the uncertainty-aware regression loss for HH and hrech_{rec} improves the accuracy.

3) Sorting by scorefHσhrec\frac{score}{fH\sigma_{h_{rec}}} is effective. Comparing ‘D+U+S’ with ‘D+U’, we can see sorting by scorefHσhrec\frac{score}{fH\sigma_{h_{rec}}} improves the accuracy, showing that scorefHσhrec\frac{score}{fH\sigma_{h_{rec}}} represents the 3D prediction quality better than scorescore.

3 Self-Consistency Comparisons

We analyze the self-consistency as shown in Tab. 4. From ‘Ours’, we can see that, the accuracy decreases if the predicted physical height for recovering the distance is replaced with the groundtruth physical height. This supports our method can maintain the self-consistency during inference. From ‘K’, we can see that, the accuracy increases if the predicted physical size for optimizing the distance is replaced with the groundtruth physical size. This shows there is no correlation, i.e., self-consistency, between the predicted projected corners and physical size.

4 Comparisons on the KITTI benchmark

We comprehensively benchmark MonoRCNN on the KITTI official test dataset in Tab. 3, and show the advantages of our method compared with existing methods. From Tab. 3, we can see:

1) MonoRCNN achieves the SOTA accuracy. Existing image-only methods are not able to emulate the methods using extra depth input, while our method pushes the forefront of image-only methods , by 3.17/1.753.17/1.75 in AP3D{}_{\text{3D}} and 2.72/1.082.72/1.08 in APBEV{}_{\text{BEV}} on the easy and moderate subset, and surpasses those depth-based methods by 1.59/0.931.59/0.93 in AP3D{}_{\text{3D}} and 0.45/0.790.45/0.79 in APBEV{}_{\text{BEV}} on the easy and moderate subset. With only single frame as input, our method is even comparable with a video-based method . Notice that MonoRCNN uses a ResNet-50 backbone , while uses a DenseNet-121 backbone and use DLA-34 backbones . Though DenseNet-121 and DLA-34 are more advanced than our ResNet-50 , MonoRCNN still outperforms . We also emphasize that MonoRCNN significantly outperforms RTM3D , a keypoint-based method, though uses additional images from right cameras for training. This supports that our distance decomposition is much better than the keypoint-based distance decomposition.

2) MonoRCNN is simple and efficient. MonoRCNN is an image-only method, thus is more simple and efficient than those depth-based and video-based methods. The monocular depth estimator in uses a heavy ResNet-101 backbone , and MonoRCNN runs 33 times faster than , and 55 times faster than .

5 Cross-Dataset Test

To evaluate the ability to generalize to images with different camera intrinsics, we conduct cross-dataset test by applying the model trained with the training subset of the KITTI val split to the nuScenes cross-test set. Since this paper is the first paper presenting cross-dataset test in monocular 3D object detection, we also provide the results of M3D-RPN using its official model https://github.com/garrickbrazil/M3D-RPN as a comparison. To focus on the accuracy of the distance prediction, we report the mean error of the distance prediction within different distance ranges for both methods in Tab. 5. The errors are calculated for recalled objects. The results show that our method achieves lower distance prediction errors. From errors in different distance intervals, we can see that the further the objects are, the better generalization of our method is. This is because our method reasons the distance by the local geometric variables of objects. We also visualize some qualitative examples in Fig. 7, and we can see our method achieves accurate distance prediction. The pseudo-LiDAR-based methods suffer from the difficulty of generalizing to images with different camera intrinsics, because it is difficult for monocular depth estimators to generalize to images with different camera intrinsics .

Conclusion

We have proposed a novel geometry-based distance decomposition which makes the distance prediction interpretable, accurate, and robust. Our method directly predicts 3D bounding boxes from RGB images with a compact architecture, thus is simple and efficient. The experimental results show that our method achieves the SOTA performance on the monocular 3D Object Detection and Bird’s Eye View tasks of the KITTI dataset, and can generalize to images with different camera intrinsics.

References