SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation
Zechen Liu, Zizhang Wu, Roland Tóth
Introduction
Vision-based object detection is an essential ingredient of autonomous vehicle perception and infrastructure less robot navigation in general. This type of detection methods are used to perceive the surrounding environment by detecting and classifying object instances into categories and identifying their locations and orientations. Recent developments in 2D object detection have achieved promising performance on both detection accuracy and speed. In contrast, 3D object detection has proven to be a more challenging task as it aims to estimate pose and location for each object simultaneously.
Currently, the most successful 3D object detection methods heavily depend on LiDAR point cloud or LiDAR-Image fusion information (features learned from the point cloud are key components of the detection network). However, LiDAR sensors are extremely expensive, have a short service life time and too heavy for autonomous robots. Hence, LiDARs are currently not considered to be economical to support autonomous vehicle operations. Alternatively, cameras are cost-effective, easily mountable and light-weight solutions for 3D object detection with long expected service time. Unlike LiDAR sensors, a single camera in itself can not obtain sufficient spatial information for the whole environment as single RGB images can not supply object location information or dimensional contour in the real world. While binocular vision restores the missing spatial information, in many robotic applications, especially Unmanned Aerial Vehicles (UAVs), it is difficult to realize binocular vision. Hence, it is desirable to perform 3D detection on a monocular image even if it is a more difficult and challenging task.
Previous state-of-the-art monocular 3D object detection algorithms heavily depend on region-based convolutional neural networks (R-CNN) or region proposal network (RPN) structures . Based on the learned high number of 2D proposals, these approaches attach an additional network branch to either learn 3D information or to generate a pseudo point cloud and feed it into point-cloud-detection network. The resulting multi-stage complex process introduces persistent noise from 2D detection, which significantly increases the difficulty for the network to learn 3D geometry. To enhance performance, geometry reasoning , synthetic data and post 3D-2D processing have also been used to improve 3D object detection on single image. By the knowledge of the authors, no reliable monocular 3D detection method has been introduced so far to learn 3D information directly from the image plane avoiding the performance decrease that is inevitable with multi-stage methods.
In this paper, we propose an innovative single-stage 3D object detection method that pairs each object with a single keypoint. We argue and later show that a 2D detection , which introduces nonnegligible noise in 3D parameter estimation, is redundant to perform 3D object detection. Furthermore, 2D information can be naturally obtained if the 3D variables and camera intrinsic matrix are already known. Consequently, our designed network eliminates the 2D detection branch and estimates the projected 3D points on the image plane instead. A 3D parameter regression branch is added in parallel. This design results in a simple network structure with two estimation threads. Rather than regressing variables in a separate method by using multiple loss functions, we transform these variables together with projected keypoint to 8 corner representation of 3D boxes and regress them with a unified loss function. As in most single-stage 2D object detection algorithms, our 3D detection approach only contains one classification and regression branch. Benefiting from the simple structure, the network exhibits improved accuracy in learning 3D variables, has better convergence and less overall computational needs.
Second contribution of our work is a multi-step disentanglement approach for 3D bounding box regression. Since all the geometry information is grouped into one parameter, it is difficult for the network to learn each variable accurately in a unified way. Our proposed method isolates the contribution of each parameter in both the 3D bounding box encoding phase and the regression loss function, which significantly helps to train the whole network effectively.
Our contribution is summarized as follows:
We propose a one-stage monocular 3D object detection with a simple architecture that can precisely learn 3D geometry in an end-to-end fashion.
We provide a multi-step disentanglement approach to improve the convergence of 3D parameters and detection accuracy.
The resulting method outperforms all existing state-of-the-art monocular 3D object detection algorithms on the challenging KITTI dataset at the submission date November 12, 2019.
Related Work
In this section, we provide an in-depth overview of the state-of-the-art of 3D object detection based on the used sensor inputs. We first discuss LiDAR based and LiDAR-image fusion methods. After that, stereo image based methods are overviewed. Finally, we summarize approaches that only depend on single RGB images.
LiDAR-based 3D object detection methods achieve high detection precision by processing sparse point clouds into various representations. Some existing methods, e.g., , project point clouds into 2D Bird’s eye view and equip standard 2D detection networks to perform object classification and 3D box regression. Others methods, like , represent point clouds in voxel grid and then leverage 2D/3D CNNs to generate proposals. LiDAR-image fusion methods learn relevant features from both the point clouds and the images together. These features are then combined and fed into a joint network trained for detection and classification.
Stereo images based methods:
The early work 3DOP generates 3D proposals by exploring many handcrafted features such as stereo reconstruction, depth features, and object size priors. TLNet introduces a triangulation based learning network to pair detected regions of interests between left and right images. Stereo R-CNN creates 2D proposals simultaneously on stereo images. Then, the methods utilize keypoint prediction to generate a coarse 3D bounding box per region. A 3D box alignment w.r.t. stereo images is finally used on the object instance to improve the detection accuracy. Pseudo-LiDAR methods, e.g., , generate a “fake” point cloud and then feed these features into a point cloud based 3D detection network.
Monocular image based methods:
3D object detection based on a single perspective image has been extensively studied and it is considered to be a challenging task. A common approach is to apply an additional 3D network branch to regress orientation and translation of object instances, see . Mono3D generates 3D anchors by using massive amount of features via semantic segmentation, object contour, and location priors. These features are then evaluated via an energy function to accommodate learning of relative information. Deep3DBox introduces bins based discretization for the estimation of local orientation for each object and 2D-3D bounding box constrain relationships to obtain the full 3D pose. MonoGRNet subdivides the 3D object localization task into four tasks that estimate instance depth, 3D location of objects, and local corners respectively. These components are then stacked together to refine the 3D box in a global context. The network is trained in a stage-wise fashion and then trained end-to-end to obtain the final result. Some methods, like , rely on features detected in a 2D object box and leverage external data to pair information from 2D to 3D. DeepMANTA proposes a coarse-to-fine process to generate accurate 2D object proposals, which proposals are then used to match a 3D CAD model from an external annotated dataset. 3D-RCNN also uses 3D models to pair the outputs from a 2D detection network. They then recover the 3D instance shape and pose by deploying a render-and-compare loss. Other approaches, like , generate hand-crafted features by transforming region of interest on images to other representations. AM3D transforms 2D imagery to a 3D point cloud plane by combining it with a depth map. A PointNet is then used to estimate 3D dimensions, locations and orientations. The only one-stage method M3D-RPN proposes a standalone network to generate 2D and 3D object proposals simultaneously. They further leverage a depth-aware network and post 3D-2D optimization technique to improve precision. OFTNet maps the 2D feature map to bird-eye view by leveraging orthographic feature transform and regress each 3D variable independently. Consequently, none of the above methods can estimate 3D information accurately without generating 2D proposals.
Detection Problem
SMOKE Approach
In this section, we describe the SMOKE network that directly estimates 3D bounding boxes for detected object instances from monocular imagery. In contrast to previous techniques that leverage 2D proposals to predict a 3D bounding box, our method can detect 3D information with a simple single stage. The proposed method can be divided into three parts: (i) backbone, (ii) 3D detection, (iii) loss function. First, we briefly discuss the backbone for feature extraction, followed by the introduction of the 3D detection network consisting of two separated branches. Finally, we discuss the loss function design and the multi-step disentanglement to compute the regression loss. The overview of the network structure is depicted in Fig. 2.
We use a hierarchical layer fusion network DLA-34 as the backbone to extract features since it can aggregate information across different layers. Following the same structure as in , all the hierarchical aggregation connections are replaced by a Deformable Convolution Network (DCN) . The output feature map is downsampled 4 times with respect to the original image. Compared with the original implementation, we replace all BatchNorm (BN) operation with GroupNorm (GN) since it has been proven to be less sensitive to batch size and more robust to training noise. We also use this technique in the two prediction branches, which will be discussed in Sec. 4.2. This adjustment not only improves detection accuracy, but it also reduces considerably the training time. In Sec. 5.2, we provide performance comparison of BN and GN to demonstrate these properties.
2 3D Detection Network
Regression Branch:
This operation is the inverse of Eq. (1). In order to retrieve object dimensions , we use a pre-calculated category-wise average dimension computed over the whole dataset. Each object dimension can be recovered by using the residual dimension offset :
Inspired by , we choose to regress the observation angle instead of the yaw rotation for each object. We further change the observation angle with respect to the object head , instead of the commonly used observation angle value , by simply adding . The difference between these two angles is shown in Fig. 4. Moreover, each is encoded as the vector . The yaw angle can be obtained by utilizing and the object location:
Finally, we can construct the 8 corners of the 3D bounding box in the camera frame by using the yaw rotation matrix , object dimensions and location :
3 Loss Function
We employ the penalty-reduced focal loss in a point-wise manner on the downsampled heatmap. Let be the predicted score at the heatmap location and be the ground-truth value of each point assigned by Gaussian Kernel. Define and as:
For simplicity, we only consider a single object class here. Then, the classification loss function is constructed as
where and are tunable hyper-parameters and is the number of keypoints per image. The term corresponds to penalty reduction for points around the groundtruth location.
Regression Loss:
where represents the number of groups we define in the 3D regression branch. The multi-step disentangling transformation divides the contribution of each parameter group to the final loss. In Sec. 5.2, we show that this method significantly improves detection accuracy.
4 Implementation
In this section, we discuss the implementation of our proposed methodology in detail together with selection of the hyperparemeters.
We avoid applying any complicated preprocessing method on the dataset. Instead, we only eliminate objects whose 3D projected center point on the image plane is out of the image range. Note that the total number of projected center points outside the image boundary for the car instance is 1582. This accounts for only the 5.5% of the entire set of 28742 labeled cars
Data Augmentation:
Data augmentation techniques we used are random horizontal flip, random scale and shift. The scale ratio is set to 9 steps from 0.6 to 1.4, and the shift ratio is set to 5 steps from -0.2 to 0.2. Note that the scale and shift augmentation methods are only used for heatmap classification since the 3D information becomes inconsistent with data augmentation.
Hyperparameter Choice:
In the backbone, the group number for GroupNorm is set to 32. For channels less than 32, it is set to be 16. For Eq. (7), we set and in all experiments. Based on , the reference car size and depth statistics we use are and (measured in meters).
Training:
Our optimization schedule is easy and straightforward. We use the original image resolution and pad it to 1280 384. We train the network with a batch size of 32 on 4 Geforce TITAN X GPUs for 60 epochs. The learning rate is set at and drops at 25 and 40 epochs by a factor of 10. During testing, we use the top 100 detected 3D projected points and filter it with a threshold of 0.25. No data augmentation method and NMS are used in the test procedure. Our implementation platform is Pytorch 1.1, CUDA 10.0, and CUDNN 7.5.
Performance Evaluation
We evaluate the performance of our proposed framework on the challenging KITTI dataset. The KITTI dataset is a broadly used open-source dataset to evaluate visual algorithms on a driving scene considered representative for autonomous driving. It contains 7481 images for training and 7518 images for testing. The test metric is divided into easy, moderate and hard cases based on the height of the 2D bounding box of object instances, occlusion and truncation level. Frequently, the training set is split into 3712 training examples and 3769 validation examples as mentioned in . For the 3D detection task of our proposed method, the 3D Object Detection and Bird’s Eye View benchmarks are available for evaluation.
The 3D detection results of our proposed method on the split sets test and val are compared with the state-of-the-art single image-based methods in Tabs. 1 and 2. We principally focus on the car class since it has been at the focus of previous cooperative studies. For both tasks, the average precision (AP) with Intersection over Union (IoU) larger than 0.7 is used as the metric for evaluation. Note that as pointed out by , the official KITTI evaluation has been using 40 recall points instead of 11 recall points to measure the AP value since October 8, 2019. However, previous methods only report accuracy at 11 points on the val set. For fair comparison, we report the average precision on 40 points on the test set and on the val set.
Results on the test split, shown in Tab. 1, show that SMOKE outperforms all existing monocular methods on both 3D object detection and Bird’s eye view evaluation metrics. We achieve improvement in the moderate and hard sets and comparable results on the easy set in the 3D object detection task. For Bird’s eye view detection, we also achieve notable improvement on the moderate and hard sets. Compared with other methods that increase image size for better performance, our approach uses relatively low-resolution input and still achieves competitive results on the hard set in 3D detection. Next to these, SMOKE shows a significant improvement on detection speed. Without the time-consuming region proposal process and by the benefits of single-stage structure, our proposed method only needs 30ms to run on a TITAN XP. Note that we only compare our method with methods that directly learn features from images. Approaches based on hand-crafted features are not listed in the table. However, with respect to the val set of KITTI, the performance degrades as reported in Tab. 2. We argue that this is due to a lack of training objects. A similar problem has been reported in .
Estimation of object location in a monocular image is difficult since the incompleteness of spatial information. We evaluate the depth estimation of SMOKE using two different distance measures. In Fig. 5, the achieved depth error is displayed in intervals of 10 meters. The error is computed if the 2D bounding box of a detection with any of the ground truth objects has an IoU larger than 0.7. As shown in the figure, the depth estimation error increases as the distance grows. This phenomenon has been observed in many monocular image-based detection algorithms since small objects have large distance distribution. We compare our method with two other methods Mono3D and 3DOP on the same val set. The curve indicates that our proposed SMOKE method outperforms both methods largely on depth error. Especially at distances larger than 40m, our method achieves more robust and accurate depth estimation.
D Object Detection:
The 2D detection performance on the official KITTI test set is depicted in Tab. 3. Although the 2D bounding box is not directly regressed in the SMOKE network, we observe that our method achieves comparative results on the 2D object detection task. The 2D detection box is obtained as the smallest rectangle that encircles the projected 3D bounding box on the image plane. Unlike other approaches following a 2D3D structure, our proposed method reverse this process in a 3D2D fashion and outperforms many of the existing methods. This clearly shows that 3D object detection provides more abundant information than 2D detection, hence 2D proposals are redundant and not needed for 3D detection. Furthermore, our proposed method does not use extra data, complicated networks and high-resolution input compared to other methods.
2 Ablation Study
In this section, we show the results of experiments we conducted to compare different normalization choices, loss function, and rotation angle parameterizations. All experiments are performed on the train/val split on the KITTI dataset. Moreover, we use car class to evaluate our model.
We chose GN as the normalization strategy since it is less sensitive to batch size and cross-GPU training issues. We compare the performance difference in the 3D detection task of BN and GN used in the backbone network. As illustrated in Tab. 4, GN achieves significant improvement over BN on the val set. In addition, we notice that GN can save considerable time in training. For each epoch, GN consumes around 5 minutes while BN needs 8 minutes which takes 60% more time compared to GN.
Regression Loss:
Rotation Parametrization:
We compare the performance of SMOKE with respect to different representations of rotation. Following prior work , the orientation can be encoded as a 4D quaternion to formulate 3D bounding box. The result with this representation is illustrated in Tab. 6. We observe that our simple vectorial representation yields slightly better result than the quaternion representation on both 3D detection and Bird’s eye view evaluation.
3 Qualitative Results
Qualitative results on both the test and val sets are displayed in Fig. 6. For better visualization and comparison, we also plot the object localization in Bird’s eye view. The results clearly demonstrate that SMOKE can recover object distances accurately
Conclusion and Future Work
In this paper, we presented a novel single-stage monocular 3D object detection method based on projected 3D points on the image plane. Unlike previous methods, which depend on 2D proposals to estimate 3D information, our approach regresses 3D bounding boxes directly. This leads to a simple and efficient architecture. To further improve the convergence of regression loss, we proposed a multi-step disentanglement method to isolate the contribution of various parameter groups. In addition, our model does not need synthetic data, complicated pre/post-processing, and multi-stage training. In overall, we largely improve both the detection accuracy and speed on KITTI 3D object detection and Bird’s eye view tasks.
Our proposed SMOKE 3D detection framework achieves promising accuracy and efficiency, which can be further extended and used on autonomous vehicles and in robotic navigation. In the future, we aim at extending our method to stereo images and further improving the estimation of projected 3D keypoints and their depth.