AutoShape: Real-Time Shape-Aware Monocular 3D Object Detection

Zongdai Liu, Dingfu Zhou, Feixiang Lu, Jin Fang, Liangjun Zhang

Introduction

Perceiving 3D shapes and poses of surrounding obstacles is an essential task in autonomous driving (AD) perception systems. The accuracy and speed performance of 3D objection detection is important for the following motion planning and control modules in AD. Many 3D object detectors have been proposed, mainly for depth sensors such as LiDAR or stereo cameras , which can provide the distance information of the environments directly. However, LiDAR sensors are expensive and stereo rigs suffer from on-line calibration issues. Therefore, monocular camera based 3D object detection becomes a promising direction.

The main challenge for monocular-based approaches is to obtain accurate depth information. In general, depth estimation from a single image without any prior information is a challenging problem and recent many deep learning-based approaches achieve good results . With the estimated depth map, pseudo LiDAR point cloud can be reconstructed via pre-calibrated intrinsic camera parameters and 3D detectors designed for LiDAR point cloud can be applied directly on pseudo LiDAR point cloud . Furthermore, integrates the depth estimation and 3D object detection network together following an end-to-end manner. However, heavy computation burden is one main bottleneck of such two-stage approaches.

To improve the efficiency, many direct regression-based approaches have been proposed (e.g., SMOKE , RTM3D ) and achieved promising results. By representing the object as one center point, the object detection task is formulated as keypoints detection and its corresponding attributes (e.g., size, offsets, orientation, depth, etc.) regression. With this compact representation, the computation speed of this kind of approach can reach 20∼\sim30 fps (frame per second). However, the drawback is also obvious. One center point representation ignores the detailed shape of the object and results in location ambiguity if its projected center point is on another object’s surface due to occlusion . To alleviate this ambiguity, other geometrical constraints have been used to improve the performance. RTM3D adds 8 more keypoints as additional constraints which are defined as the projected 2d location of the 3D bounding box’s corners. However, these keypoints have non-real context meanings and their 2D locations vary differently with the changing of the camera view-point, even the object’s orientation. As shown in the left of Fig. 1, some keypoints are on the ground and some are on the sky or trees. This makes the keypoints detection network extremely difficult to distinguish the keypoints or other image pixels.

In this paper, we propose a novel approach to learn the meaningful keypoints on the object surface and then use them as additional geometrical constraints for 3D object detection. Specifically, we design an automatic deformable model-fitting pipeline first to generate the 2D/3D correspondences for each object. Then, the center point plus several distinguished keypoints are learned from the deep neural network. Based on these keypoints and other regressed objects’ attributes (e.g., orientation angle, object dimension etc.), the object’s 3D bounding box can be solved with linear equations. The proposed framework can be trained in an end-to-end manner. Our contributions include:

We propose a shape-aware 3d object detection framework, which employs keypoints geometry constraints for 2D/3D regression to boost the detection performance.

We present a method for automatically fitting the 3D shape to the visual observations and then generating ground-truth annotations of 2D/3D keypoints pairs for the training network. Our source code and dataset will be made public for the community.

The effectiveness of our approach has been verified on the public KITTI dataset and achieved SOTA performance. More importantly, the proposed framework achieves real-time (2525 fps), which can be integrated into the AD perception module.

Related Work

Image-based 3D object detection becomes popular due to the cheap price of the camera sensors. Stereo-based approaches usually suffer from calibration issues between two camera rigs. Therefore, many 3D object detection approaches have proposed to use a single image frame. Generally, these approaches can be categorized into three types: depth-map-based, direct regression-based, and CAD model-based methods.

Depth-map-based methods usually need to estimate the depth map first. In and , the estimated depth map is transformed into point clouds, and then point-cloud-based 3D object detectors are employed for achieving the detection results. Rather than transforming the depth map into point clouds, many approaches propose using the depth estimation map directly in the framework to enhance the 3D object detection. In M3D-RPN and , the pre-estimated depth map has been used to guide the 2D convolution, which is called as “Depth-Aware Convolution”. Direct regression-based methods are proposed to estimate the objects’ 3D information via image domain directly, such as . Direct-based methods are much more efficient than depth-map-based methods because the depth-map computation procedure is not necessary.

In order to well benefit the prior knowledge, the shape information has been integrated into the CAD-based approaches. Deep MANTA and ApolloCar3D are two keypoints based methods, in which the 3D keypoints are pre-defined on the CAD model and their corresponding 2D points on the image plane are computed by the deep neural network. Then the 3D pose can be solved with a standard 2D/3D pose solver with these 2D/3D correspondences. Besides keypoints-based methods, dense-matching-based approaches are proposed in . In , Rendering-and-Compare loss is designed for optimizing the 3d pose estimation. While in and , the 3D pose estimation and reconstruction of each object are generated simultaneously with the deep neural network.

2 Data Labeling for 3D Object Detection

For easy representation, objects are usually described as 3D cuboids in deep learning frameworks while the shape information has been totally ignored. Manually label the object shape via only the image observation is extremely difficult and the annotation quality also can not be guaranteed. Many CAD model guided annotation approaches have been proposed to obtain the dense shape annotations. In , both the stereo image and the sparse LiDAR point cloud has been employed for generating the dense scene flow for both foreground and background pixels. For dynamic objects, 16 vehicle models are chosen as basic templates and then the dense annotation is achieved by finding an optimal 3D similarity transformation (e.g., the pose and scale of the 3D model) with three types of observations such as LiDAR points, dense disparity computed by SGM and labeled 2D/3D correspondences.

In , 66 keypoints are defined on the 3D CAD models and annotators label their corresponding 2D keypoints on the image. Based on the 2D/3D correspondences, the object poses can be obtained via a PnP solver. In , the authors apply a differentiable shape renderer to signed distance fields (SDF), leveraged together with normalized object coordinate spaces (NOCS) to automatically generate the dense 3D shape without the 3D bounding boxes annotation. Although the whole process is labor-free, the annotation quality is far-from the ground truth. Different from , we use the ground truth 3D bounding boxes as strong guidance for our 3D shape annotation generation process.

Problem Definition

Before the introduction of our proposed approach, a general description of image-based 3D object detection problem is introduced first.

Given an image, the task of pose estimation is to estimate the orientation and translation of objects in 3D. Specifically, 6D pose is represented by a rigid transformation (R,T)(\mathbf{R},\mathbf{T}) from the object coordinate system to the camera coordinate system, where R\mathbf{R} represents the 3D rotation and T\mathbf{T} represents the 3D translation.

Assuming a 3D object point Po\mathbf{P}_{o} (xo,yo,zo)(x_{o},y_{o},z_{o}) in the object coordinate system, transformed 3d point Pc\mathbf{P}_{c} (xc,yc,zc)(x_{c},y_{c},z_{c}) in camera coordinate can be obtained as

where R\mathbf{R} is rotation matrix, T\mathbf{T} is translation vector. Given the camera intrinsic matrix \mathbf{K}=\left[\begin{array}[]{ccc}f_{x}&0&c_{x}\\ 0&f_{y}&c_{y}\\ 0&0&1\\ \end{array}\right] , the projected image point p\mathbf{p} (u,v)(u,v) can be obtained as

Based on Eq. 1 and Eq. 2, the object pose R\mathbf{R} and T\mathbf{T} can be theoretically recovered with the geometric constraints between 3D points Po\mathbf{P}_{o} on the object and the projected 2D image points p\mathbf{p}.

2 Learning-based 3D Object Detection

In the era of deep learning, many approaches have been proposed to detect objects and directly regress their poses using neural networks, while geometric 2D/3D constraints have been ignored in the formulation. Image-based 3D object detection is a typical task, which aims at estimating the location, orientation of an object in the camera coordinate. Usually, an object is represented as a rotated 3D BBox as

in which r, t represent the object’s orientation, location in the camera coordinate and d is the dimension of the object. With the super expression ability of neural networks, all these parameters are regressed directly without imposing addition constraints. Indeed, both 3D object detection and pose estimation are essentially the same problem and (r, t) can be easily transformed from (R, T). Therefore, we explicitly employ geometric constraints in pose estimation formulation to improve the learning-based 3D object detection.

Proposed Method

In this section, we propose a general deep learning-based 3D object detection framework, which can employ the 2D/3D geometric constraints. To well explore the prior knowledge, CAD models are employed here. First, we pre-define several distinguished 3D keypoints on CAD models. Then, we propose to build the correlation between these 3D keypoints and their 2D projections on the image resorting to the deep learning network. Finally, the object pose can be easily solved with these geometrical constraints. More importantly, all the processes are implemented into the neural network, which can be trained in an end-to-end manner.

Assuming a 3D point Poi\mathbf{P}_{o}^{i} (xoi,yoi,zoi)(x_{o}^{i},y_{o}^{i},z_{o}^{i}) in object local coordinate, then its projection location (ui,vi)(u^{i},v^{i}) on the image plane can be obtained based on Eq. 1 and Eq. 2 as

In autonomous driving scenario, the road surface that the object lies on is almost flat locally, therefore the orientation parameters are reduced from three to one by keeping only the yaw angle ryr_{y} around the Y-axis. Therefore the rotation matrix R\mathbf{R} becomes as \left[\begin{array}[]{ccc}cos(r_{y})&0&sin(r_{y})\\ 0&1&0\\ -sin(r_{y})&0&cos(r_{y})\\ \end{array}\right] and Eq. 4 can be simplified as

However, in the real AD scenario, not all the keypoints can be seen from a certain camera viewpoint. For these keypoints which have been seriously occluded, the 2D/3D keypoints regression can not be well guaranteed. To well handle this kind of uncertainty, we propose to output an additional score to measure the confidence of each keypoint. And this score can be used as a weight during the pose calculation process. Specifically, for the 2n2n constrains in Eq. 6, an additional weights c={c1,c2,...,c2n}c=\{c_{1},c_{2},...,c_{2n}\} have been added to determine their importance during the pose calculation procedure. Therefore, Eq. 6 can be reformulated as

In this linear system, T\mathbf{T} represents the object location in the camera coordinate system, which can be solved by providing 2D/3D correspondences and the rotation angle ryr_{y}. Here, the 3D keypoints are defined in the local object’s coordinate varying in a relatively small range and the 2D keypoints are defined in the image domain. Both of them are easy for networks to learn. However, manual labeling of the ground truth for 2D and 3D keypoints is very costly and tedious. Therefore, we develop an auto-labeling pipeline by optimizing the 2D and 3D reprojection errors. The detailed annotation pipeline will be introduced in Section 5.

2 Network

An overview of our proposed framework is illustrated in Fig. 2. Here, we follow one-stage-based 3D object detection framework such as CenterNet for its inference efficiency. Our proposed framework is backbone independent and here we employ DLA-34 in our implementation. Given an image I\mathbf{I} with width WW and height HH, the output feature map will be 4 times smaller than I\mathbf{I} after passing through the backbone network. To well utilize the geometric constraints, the following information is required to be learned from the deep neural network.

Object Center: in anchor-free based object detection frameworks, the object center is essential information, which serves two functions: one is whether there is an object and the other is that if there exists an object, where is the center. Usually, these two functions are realized by a classification branch by distinguishing a pixel whether is an object center or not. The output of this branch will be W4×W4×C\frac{W}{4}\times\frac{W}{4}\times C, where CC is the number of classes.

Besides the classification, an additional “offset” regression branch is required to compensate for the quantization error during the down-sampling process. The output of this branch is W4×H4×2\frac{W}{4}\times\frac{H}{4}\times 2 to represent the offset in xx and yy direction respectively.

Object Dimension: a separate branch is used to regress the object dimension hh, ww, ll with the output size of W4×H4×3\frac{W}{4}\times\frac{H}{4}\times 3. Similar to other approaches, we don’t regress the absolute object’s size directly and regress a relative scale compared to the mean object size of each class. Details operation can be found in .

2D Keypoints: rather than directly detect these keypoints from the image, we regress nn ordered 2D offset coordinates for each object center. The benefit is that the number and order of keypoints for each object can be well guaranteed. In addition, the regression for the offset is easier by removing the object center. The output size of this branch is W4×H4×2n\frac{W}{4}\times\frac{H}{4}\times 2n.

3D Keypoints: similar to the 2D keypoints, we regress the 3d keypoints in the local object coordinate. In addition, all 3D keypoints values are normalized by object dimension (l,w,h)(l,w,h) in xx, yy, zz-direction respectively. By using this format, the 3D keypoints values are in a relatively small range, which will benefit the whole regression process. The output size of this branch is W4×H4×3n\frac{W}{4}\times\frac{H}{4}\times 3n.

Object Orientation: similarly, we regress local orientation angle with respect to the ray through the perspective point of 3D center following Multi-Bin based method . Here, 8 bins are used with the output size of W4×W4×8\frac{W}{4}\times\frac{W}{4}\times 8.

Keypoints Confidence Scores: for each keypoint, a couple of additional confidence scores have been regressed for measuring its contribution in the linear system for solving the object pose. For 2n2n constrains in Eq . 6, a feature map with size of W4×W4×2n\frac{W}{4}\times\frac{W}{4}\times 2n will be outputted.

3D IoU Confidence Score: rather using the classification score directly as the object detection confidence, we add one branch to regress the 3D IoU score in purpose. This score is supervised by the IoU between estimated Bbox and ground truth Bbox. Finally, the product of this score and the output classification score is assigned as the final 3D detection confidence score.

3 Loss Function

The overall loss contains the following items: a center point classification loss lml_{m} and center point offset regression loss loffl_{off}, a 2D keypoints regression loss l2Dl_{2D}, a 3D keypoints points regression loss l3Dl_{3D}, an orientation multi-bin loss lrl_{r}, a dimension regression loss lDl_{D}, a 3D IoU confidence loss lcl_{c} and a 3D bounding box IoU loss lIoUl_{IoU}. Specifically, the multi-task loss is defined as

where lml_{m} is the focal loss as used in , l2Dl_{2D} is a depth-guided l−1l-1 loss as used in , lDl_{D} and l3Dl_{3D} are L1 loss with respect to the ground truth. Orientation loss lrl_{r} is the Multi-Bin loss. 3D IoU confidence lcl_{c} is a binary cross-entropy loss supervised by the IoU between the predicted 3D BBox and ground truth. lIoUl_{IoU} is the IoU loss between the predicted 3D BBox and ground truth .

3D Shape Auto-Labeling

In this section, we will introduce how to automatically fit the 3D shape to the visual observations and then automatically generate ground-truth annotations of 2D keypoints and 3D locations in the local object coordinate for training the network. The main process is illustrated in 3. Different from existing methods which use only a few CAD models for 3D labeling (e.g., 11 in and 16 in ), we adopt a 3D deformable vehicle template that can represent arbitrary vehicle shape by adjusting the parameters. Therefore, the 3D shape labeling process can be formulated as an optimization problem that aims at computing the optimal parameter combination to fit the visual observations (i.e. 2D instance mask, 3D bounding box, and 3D LiDAR points).

In the real-world traffic scenarios, there are many different vehicle types (e.g., coupe, hatchback, notchback, SUV, MPV, etc.) and their geometric shapes vary significantly. To perform 3D shape fitting, a straight-forward solution is to build a 3D shape dataset and the fitting process can be regarded as model retrieval. However, dataset construction is labor-intensive, inefficient, and costly. Instead, we use a deformable 3D model for vehicle representation . Specifically, this template is composed of a set of PCA (Principal Components Analysis) basis rr. Any new 3D vehicle M(s)\mathcal{M}(s) can be represented as a mean shape model M0\mathcal{M}_{0} plus a linear combination of rr principal components with coefficient s=[s1,s2,...,sr]s=[s_{1},s_{2},...,s_{r}] as

where pkp_{k} and δk\delta_{k} are the principal component direction and corresponding standard deviation and sks_{k} is the coefficient of the kthk_{th} principal component. Based on the vehicle template, we can automatically fit an optimal 3D shape to the visual observations (details in Subsec. 5.2).

2 3D Shape Optimization

For each vehicle, our goal is to assign a proper 3D shape to fit the visual observations, including the 2D instance mask, the 3D bounding box, and the 3D LiDAR points. Specifically, the annotations of 2D instance mask IinsI_{ins} and the 3D bounding box BboxB_{box} are provided by the KINS dataset and the KITTI dataset , respectively. The annotation of 3D LiDAR points for each vehicle is much more complex. According to the labeled 3D bounding boxes, we first segment out the individual 3D points from the entire raw point cloud. Then we remove the ground points using the ground-plane estimation method (i.e. RANSAC-based plane fitting). Finally, we obtain the “clean” 3D points for each vehicle, which is represented as p={p0,...,pk}\mathbf{p}=\{p_{0},...,p_{k}\}.

The 3D shape annotation is to compute the best PCA coefficient s^\mathbf{\widehat{s}} and the object 6-DoF pose (R^,t^\widehat{\mathbf{R}},\widehat{\mathbf{t}}). Existing 3D object detection benchmarks (e.g., KITTI, Waymo) only label the yaw angle because they assume vehicles are on the road plane. However, we experimentally (an example has been given in Fig. 5) find that the other two angles (i.e. pitch and roll) can significantly improve the 3D shape annotation results. Therefore, the loss function is formulated as

which consists of 3D points loss L3D\mathbf{L}_{3D} and 2D instance loss L2D\mathbf{L}_{2D}. α\alpha, β\beta are two hype-parameters to balance these two constraints.

The function of Eq. 10 can be optimized by the gradient descent strategy. The vehicle’s center position and orientation are used for t\mathbf{t} and yaw angle initialization, while the pitch, roll angles, and PCA coefficients are initialized as zeros. Then we forward the pipeline and compute L2D\mathbf{L}_{2D} and L3D\mathbf{L}_{3D} and finally back-propagate gradients to update ss, R\mathbf{R}, tt. In Fig. 4, we depict the intermediate result during the optimization process.

Experimental Results

We implement the approach and evaluate it on the public KITTI 3D object detection benchmark.

Dataset: the KITTI dataset is collected from the real traffic environment from the Europe streets. The whole dataset has been divided into training and test two subsets, which consist of 7,4817,481 and 7,5187,518 frames, respectively. Since the ground truth for the test set is not available, we divide the training data into a train set and a val set as in , and obtain 3,7123,712 data samples for training and 3,7693,769 data samples for validation to refine our model. On the KITTI benchmark, the objects have been categorized into “Easy”, “Moderate”, and “Hard” based on their height in the image and occlusion ratio, etc.

Evaluation Metric: we focus on the evaluation on “Car” category because it has been considered most in the previous approaches. For evaluation, the average precision (AP) with Intersection over Union (IoU) is used as the metric for evaluation. Our AutoShape approach is compared with existing methods on the test set using APR40AP_{R_{40}} by training our model on the whole 7,4817,481 images. We evaluate on the val set for ablation by training our model on the train set using APR40AP_{R_{40}}.

Implementation Details: We implement our auto-labeling approach (Sec. 5) using differentiable renderer which is optimized by using the Adam optimizer with a learning rate of 0.002. To speed the optimization and save memory, we downsample the PCA model to 666666 vertices and 998998 faces. We set α\alpha and β\beta to 1.0 and 5.0, respectively. Our shape-aware 3D detection network uses DLA-34 as backbone. We pad the image size to 1280×3841280\times 384. 3D IoU confidence loss weight wcw_{c} and 3D IoU loss weight wiouw_{iou} are increased from 0 to 1 with exponential RAMP-UP strategy . We use Adam optimizer with a base learning rate of 0.0001 for 200 epochs and reduce by 10×10\times at 100 and 160 epochs. We project the ground truth to corresponding right image and use random scaling (between 0.6 to 1.4), random shifting in the image range, and color jittering for data augmentation. The network is trained on 2 NVIDIA Tesla V100 (16G) GPU cards and the batch size is set to 16. For the KITTI test set evaluation, we sample 16/48 keypoints from the 3D shape, 8 corner points, and 1 center to train network.

2 Data Auto-Labeling Evaluation

Our approach can automatically generate the 2D keypoints and their corresponding 3D locations in the local object coordinate which are employed as the supervision signal during the training process. To verify the quality of the labeling results, the 2d instance segmentation mean AP and 3D bounding box mean AP is used here for verification. Specifically, the 2D instance segmentation IoU is calculated using the projected mask by 3D models and the ground truth mask (from KINS ). We obtain the labeled 3D bounding box using the dimension of the 3D model with the optimized 6-DoF pose, which is compared to the ground-truth 3D bounding box provided by with the mean IoU score. Tab. 2 shows the detailed comparison results. The proposed method can achieve 0.86 for 2D mean AP and 0.76 for 3D mean AP, which justifies the effectiveness of our auto-labeling approach.

3 Evaluation for 3D Object Detection

The evaluation of the proposed approach with other SOTA methods for 3D detection detection on KITTI test set are given in Tab. 1. From the table, we can obviously find that the proposed method with 48 keypoints achieves 4 first places in 6 tasks with the AP∣R40AP|_{R_{40}} metric. We also report our method with 16 keypoints, which has faster inference time and keeps promising accuracy. In addition, most of the existing methods such as , need to estimate the depth map, resulting in a heavy computation burden in inference. In contrast, our method obtains the depth information by 3D shape-aware geometric constraints, which is more accurate with faster running speed. We achieve 25 FPS with an NVIDIA V100 GPU card with 16 keypoints configuration. Compared with baseline geometric constraint methods using 8 corners and 1 center point as keypoints for training, our method with 48 keypoints utilizes more shape-aware keypoints to construct stronger geometric constraints, getting +5.74%, +2.72%, 1.44%, +7.22%, +3.88%, +1.12% improvements for AP3DAP_{3D} and APBEVAP_{BEV} on “Easy”, “Moderate”, and “Hard” categories.

4 Qualitative Results

Qualitative results of 3D shape auto-labeling are shown in Fig 6. Each vehicle in the image is overlaid with a rendered 3D model optimized by our method. We can see the consistency of our labeled shape and the real object. We also visualize some representative results of our shape-aware model in Fig 7. Our model can predict object location accurately even for distant and truncated objects.

5 Ablation Studies

The Number of Keypoints: our shape-aware 3D detection network benefits from the geometric constraint of 2D-3D keypoints from the 3D shape. To better understand the effect of different numbers of the keypoints, we set it from 0 to 48 with an interval of 8. Note that the 8 corners and 1 center point are always maintained in this experiment and we vary the extra keypoints. As shown in Fig. 8, from 0 to 16, the network performance is significantly improved. From 16 to 48, however, we observe that the network performance is not sensitive to the number of the keypoints. The main reason is that more dense 2D shape points can be overlapped in the W4×H4\frac{W}{4}\times\frac{H}{4} heatmap during the regression process. Furthermore, with more keypoints, the network consumes more GPU memory for storage and computation, resulting in longer training and inference time. In practice, we set the number of extra keypoints to 16, which is a good compromise of accuracy and efficiency.

2D/3D Loss for Auto-Labeling: our auto-labeling approach (Sec. 5) can generate precise posed 3D shape for each 2D vehicle instance. The key technique is simultaneously optimizing the 2D/3D constraints (loss) for better matching. Here, we conduct an ablation study to justify the effectiveness of the 2D/3D loss. We first only use 3D point loss L3D\mathbf{L}_{3D} in the objective function. Then we only use the 2D mask loss L2D\mathbf{L}_{2D} for optimization. Finally, we take both 2D/3D loss into computation. Tab. 2 shows that using both 2D/3D loss get the best performance in the auto-labeling process. We further observe that the impact of 2D mask loss L2D\mathbf{L}_{2D} is more important than 3D point loss. By using both L2DL_{2D} and L3DL_{3D}, the labeling accuracy is improved to 86.35 and 76.92, resulting in better 3D detection performance. This correlation indicates that 3D detection performance can be significantly improved by using high-quality labeling data of 3D shapes.

Conclusion

In this paper, we present a framework for real-time monocular 3D object detection by explicitly employing shape-aware geometric constraints between 3D keypoints and their 2D projections on images. Both the 3D keypoints and 2D project points are learned from deep neural networks. We further design an automatic annotation pipeline for labeling object 3D shape, which can automatically generate the shape-aware 2D/3D keypoints correspondences for each object. Experimental results show our approach can achieve state-of-the-art detection accuracy with real-time performance. Our approach is general for other types of vehicles, and in the future, we are interested in validating the performances of our approach on other objects.

References

Supplemental Material

The 2D/3D keypoints regression is a critical component in the proposed framework, however, inaccurate regression of these keypoints is inevitable in the real AD scenario due to many reasons e.g., viewpoint change, occlusion, and labeling noise, etc. Especially, these prediction outliers will greatly affect the results of the linear system described in Eq. 6. In order to handle this problem, we propose to predict a confidence score for each keypoint and employ it as a weight for determining its contribution to the linear system. To verify the effectiveness of the prediction confidence, we set a series of ablations studies on the “Car” category.

We give the results in Tab. 3. From this table, we can see that the 3D object detection performance can be significantly improved by integrating the regressed key points confidences. More importantly, this improvement is independent of the number of keypoints. In addition, for further understanding the actual meaning of this predicted confidence, we have visualized them in Fig. 9. Interestingly, we find that these keypoints with high confidences usually come from the ground point (the intersection point between the tire and the ground) and these distinguished shape border points. These points will give more contribution to the object pose estimation.

2 Multi-classes Detection

Currently, the designed Autoshape model can’t generate the keypoints annotation for “Pedestrian” and “Cyclist” due to the lack of CAD models. Here, we simply transform the 3D keypoints from the mean “Car” template to the “Pedestrian” and “Cyclist” by normalize them first and re-scale them to the bounding box’s size of other categories. By generating these keypoints, then the object’s pose can easily solve as the “Car” category. We evaluate multi-class 3d detection on the KITTI test sever and the performances are shown in Tab. 4. From this table, we can find that the proposed framework performs relatively well even though the keypoints annotation is not very accurate of “Pedestrian” and “Cyclist”. Interestingly, we find that the cyclist gives much better results than the “Pedestrian” and this is because the “Cyclist” can be considered as a rigid object to some extent. On the contrary, the “Pedestrian” is a non-rigid object and the location of these keypoints varies a lot with different object pose.