Back-tracing Representative Points for Voting-based 3D Object Detection in Point Clouds
Bowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang, Dong Xu
Introduction
As one of the fundamental tasks that aims at understanding 3D visual world, 3D object detection would like to predict amodal 3D bounding boxes and associated semantic labels of objects in real 3D scenes. 3D object detection technologies would significantly benefit various downstream real world applications such as augmented reality, robotics and \etc. In this work, we focus on 3D object detection from point clouds. It is even more challenging because the irregular, sparse and orderless characteristics of this special 3D input make it a hard task to design reliable point-based 3D object detection systems by leveraging the recent progress in 2D object detection.
While earlier works resorted to reordering point clouds into regular forms , or applying predefined shape templates , VoteNet and its variants have shown a great success in designing end-to-end 3D object detection networks based on raw point clouds. VoteNet reformulates the traditional Hough voting process into a point-wise regression problem, and generates an object proposal by sampling a number of seed points from the input point cloud whose votes are within the same cluster. The aggregated feature in each vote cluster is then used to estimate the 3D bounding box (e.g. center, size and orientation) and the associated semantic label.
Therefore, the quality of the regressed votes principally determine the reliability of the generated proposals, and then the performance on the object detector. However, although the clustered vote centers are quite accurate, the votes are usually not as representative as our expectation. For example, as illustrated in Fig. 1, by retrieving the seed points of votes from the given vote clusters, these corresponding seed points either partially cover the underlying objects (Fig. 1(a)) or contain severe outliers from the cluttered background (Fig. 1(b)). Therefore as shown in Fig. 1(a), it is undoubted that we cannot accurately predict the bounding box of a long bookshelf if the votes only capture a small area surrounding the vote center. Likewise as shown in Fig. 1(b), the severe outliers make it impossible to accurately detect the chair based on the vote features. Moreover, these seed points are less informative due to the lack of knowledge from the votes, so that there will be less significant gains if we simply back-trace these seed features (as in conventional Hough voting ) to improve the voting-based 3D object detection methods.
However, in our point of view, back-tracing is still necessary and could partially address the aforementioned issues with a special design. To be specific, as shown in Fig. 2, we would like to backwardly generate (or trace) the virtual representative points from the center of each vote cluster, and use these virtual points to revisit their surrounding seed points. This generative back-tracing operation indicates possible object shape distributions around the vote center, while the revisited seed features provide complementary local structural clues that may not be fully discovered by the votes. This bottom-up and then top-down process can end up with a mutual interaction that associates the seed features and the vote features, which has the potential to enhance each other features and enable more robust object class prediction and more accurate bounding box regression.
To this end, we propose a new point cloud-based 3D object detection method, named as Back-tracing Representative Point Network (BRNet), by incorporating the end-to-end learnable back-tracing and revisiting operations into the voting-based framework. Specifically, we propose a representative points generation module that generatively samples uniformly distributed representative points within the 3D area of a candidate object, based on the features of a vote cluster center. The generated points can coarsely infer the object bounding boxes even though their sampling process is class-agnostic. The revisited seed points of each representative point are aggregated in a similar way as ROI grid pooling , but based on the spatial layout of the representative points. After fusing the aggregated features of the revisited seed points and the features of the vote cluster center, we obtain the refined proposals to eventually detect the objects. Note that the proposed bounding box regression scheme explicitly depends on the spatial distribution of the representative points, thus improves robustness with respect to shape variations within and across object categories.
The contributions of this work are three-fold: (1) the first 3D object detection network, named as BRNet, that successfully adapts the back-tracing step of Hough voting to 3D object detection. (2) an end-to-end learnable network that can generatively back-trace the representative points, reliably revisit the seed points, and then mutually refine the object proposals for more robust object classification and more accurate bounding box regression. (3) the state-of-the-art 3D object detection performance on two benchmark datasets, ScanNet V2 (50.9% in terms of mAP@0.50) and the SUN RGB-D (43.7% in terms of mAP@0.50).
Related Works
3D object detection on point clouds. Object detection from 3D point clouds is challenging due to the irregular, sparse and orderless characteristics of 3D points. Earlier attempts usually relied on projections onto regular grids such as multi-view images and voxel grids , or based on the candidates from RGB-driven 2D proposal generation or segmentation hypotheses , where the existing 2D object detection or segmentation methods based on regular image coordinates can be effortlessly adapted. Other approaches also studied how to exploit discriminative or generative shape templates , and high-order contextual potentials to regularize the proposal objectness , or used sliding shapes , or clouds of oriented gradients (COG) .
Thanks to PointNet , deep neural networks have become extensively employed onto raw point clouds. For instance, PointRCNN introduced a two-stage 3D object detector, which is analogous to the two-stage 2D object detection methods such as Faster RCNN . Inspired by the Hough voting strategy for 2D object detection and instance segmentation , VoteNet was built upon the backbone of PointNet++ and presented an end-to-end trainable 3D object detector. Later on, the extensions of VoteNet , such as MLCVNet , HGNet and 3DSSD , employed the contextual clues, the hierarchical graph neural networks and the feature-FPS sampling strategy to enable better generation of object proposals. However, these methods heavily depend on the unreliable vote clustering proposed in , which is inevitably affected by outliers and usually overlooks inlier seed points. H3DNet partially tackled this issue by introducing a hybrid set of overcomplete geometric primitives to refine the initial bounding boxes predicted by the clustered votes. But these primitives centers are learned with less accurate supervisions and also collected by a similar clustering strategy, thus may still fail to eliminate the outliers or capture sufficient geometric clues to infer the target objects. In this work, we show how to leverage the representative points back-traced from the vote centers to complementarily profile the target objects, which enables more discriminative categorization and more robust bounding box regression.
Anchor-free 2D object detection. The implementation of the back-tracing representative points in our BRNet adopts similar anchor-free localization strategies in 2D object detection. Unlike two-stage 2D object detectors such as Faster RCNN , SSD and YOLOv2 that generate proposals with the predefined anchors, the anchor-free detectors , especially the regression-based approaches , either directly regress borders , regress the object boundaries with an iterative dynamic sampling strategy , or regress 4D offsets as the surrogate of the localization results. Inspired by these methods, in our method, the back-tracing process relies on a class-agnostic offset regressor to retrieve the representative points that indicate the likely shape profile surrounding each vote center and thus provides more local structural clues for latter inference. Rather than localization constrained by predefined class-aware statistics, as in VoteNet and its successors, the proposed BRNet benefits more flexible regression without losing its discriminative power.
Back-tracing in voting-based object detection and instance segmentation. Leibe et al. applied the hough voting strategy for simultaneously 2D object detection and instance segmentation. The core part of this approach is a learned highly flexible representation for object shapes in a probabilistic extension of Generalized Hough Transform. Moreover, the work in combined the top-down clues available from object detection and the bottom-up power of Markov Random Fields (MRFs) when performing class-specific object detection and segmentation in 3D scenes. These methods rely on a top-down strategy such as back-tracing object hypotheses to enhance the bottom-up strategy such as Hough voting. Their mutual agreement enhances each other, and thus devotes to the success of more reliable object detection. The proposed BRNet also follows this idea with a new end-to-end trainable back-tracing process based on the representative points. Recently, as a 3D instance segmentation method, 3D-MPA applied a “direct” back-tracing strategy to cluster the surface points from the corresponding votes in one cluster. In contrast, our method alleviates the inherent partial coverage and outlier issues from the “generative” back-tracing strategy.
Methodology
In this section, we describe the technical details of our BRNet. Sec. 3.1 presents an overview of our method. In Sec. 3.2 to Sec. 3.5, we elaborate the network architecture and the learning objective of our BRNet.
BRNet consists of four main modules: (1) vote generation and clustering, (2) back-traced representative points generation, (3) seed point revisiting, and (4) proposal refinement and classification followed by standard 3D NMS. In the first module, we follow the same network and training strategy as in VoteNet to generate the seed points, the votes and the vote clusters. We will elaborate the other three modules in the following parts.
2 Generating Back-traced Representative Points
The conventional back-tracing step of Hough voting for identifying object boundaries is less reliable for amodal object detection from partial observations, as it just picks up seed points that contribute to the selected votes. For example, in VoteNet , these back-traced seed points can only capture local geometric area near the cluster center while containing the outliers from the cluttered background in the meantime. VoteNet circumvents this issue by removing the back-tracing step and using a PointNet-like set aggregation block just for votes, and then generates the object proposals and classifies them. However, the aforementioned incompleteness issue and the outliers within the votes (delivered from the seed points) are clearly harmful for the detection task. To this end, we argue that it is still beneficial to use back-tracing in point-based 3D object detection, but it requires a better tracing strategy to effectively find the representative seed points. In contrast to the conventional back-tracing strategy, we propose a representative point generation (RPG) module to backwardly regress the virtually generated representative points from the votes in a generative manner. The generated representative points are uniformly distributed within the potential 3D area of a candidate object, which can also indicate 3D object shapes when interacted with their actual surrounding seed points.
Network architecture and learning. The RPG module is implemented by using multi-layer perceptrons (MLP) with the ReLU activation function and batch normalization. It takes the feature from the vote center as the input, and its output is the set . We employ to map any real number to (0, ) on the output of . This module is supervised by the ground-truth (GT) offsets as the vote center can be assigned to a GT object, i.e.
where is used to balance the two terms.
3 Revisiting Seed Points
4 Proposal Refinement and Classification
5 The Learning Objective
In summary, the loss function of the entire framework of the newly proposed BRNet is defined as following:
Following the terms and label assignment strategy used in VoteNet , the loss terms , , indicate the per-point vote regression loss, the objectness loss and the semantic classification loss, respectively. is defined in Sec. 3.2. is used to supervise the residuals from the initial representative point sets to the final representative point sets:
Experiments
Datasets. We evaluate our method on two large-scale indoor scene datasets, i.e. SUN RGB-D and ScanNet V2 . SUN RGB-D consists of single-view indoor RGB-D images annotated with the oriented 3D bounding boxes and the semantic labels for categories. The point clouds are converted from the depth maps based on the provided camera parameters. The captured point clouds contain severe occlusions and holes, thus are challenging for 3D object detection. ScanNet V2 is a 3D mesh dataset about 3D reconstructed indoor scenes. It contains object categories with densely annotated axis-aligned bounding boxes. The scans in the ScanNet V2 dataset are more complete with more objects than those in the SUN RGB-D dataset. For both datasets, we use the same data preparation and training/validation split as in VoteNet .
Input and data augmentation. The input of our method is a point cloud randomly sub-sampled from the raw data of each dataset, i.e., points from a point cloud in the SUN RGB-D dataset, and points from a 3D mesh in the ScanNet V2 dataset. We also include the height feature to each point. To augment the training data, we add random flipping, rotating and scaling to the input point clouds, as the way employed by VoteNet .
Network training details. Our network is end-to-end optimized by using the Adam optimizer with the batch size as . The base learning rates are for the SUN RGB-D dataset and for the ScanNet V2 dataset. We train the network for epochs on both datasets. The cosine annealing learning rate strategy is adopted for learning rate decay. Based on PyTorch platform equipped with one NVIDIA GeForce RTX 2080 Ti GPU card, it takes around hours to train the model on the ScanNet V2 dataset, while it takes around 12 hours on the SUN RGB-D dataset.
Inference and evaluation. Our method takes the point clouds of the entire scenes as the inputs and outputs the object proposals. The proposals are post-processed by a 3D NMS module with an IoU threshold of . The evaluation follows the same protocol as in using mean average precision, especially mAP@ and mAP@.
2 Comparisons with the State-of-the-art Methods
We compare our method with a list of reference methods, for example the earlier attempts, such as COG , DSS and 3D-SIS , 2D-driven and F-PointNet , and GSPN , and the recent point cloud-based state-of-the-art methods such as VoteNet and its successors MLCVNet , HGNet and H3DNet .
Quantitative results. The comparison results are summarized in Table 1. Our method outperforms all baseline methods by remarkable performance gains, for example more than % and % improvement in terms of the mAP@ metric on the validation sets of ScanNet V2 and SUN RGB-D respectively. Note that mAP@ is a fairly challenging metric as it basically requires more than coverage in each dimension of a bounding box, which indicates that back-tracing representative points can significantly improve the localization accuracy. Notably, MLCVNet works well on the ScanNet dataset but achieves relatively poor performance on the SUN RGB-D dataset, while HGNet works well on the SUN RGB-D dataset but achieves poor result on the ScanNet dataset, especially in terms of the mAP@ metric. Our method works well on both datasets, which indicates its stronger generalization ability for different detection scenarios. ScanNet contains relative complete 3D reconstructed meshes, while SUN RGB-D consists of single-view RGB-D scans with severe occlusions and holes. Moreover, H3DNet ensembles PointNet++ backbones to achieve the reported result on the SUN RGB-D dataset, while our model only needs one backbone as the base feature extractor. It further validates it is effective to back-trace the representative points for reliably parsing the object proposals. As shown in Table 2, our method performs the best on classes among total classes from the ScanNet dataset in terms of mAP@. While our method only uses one PointNet++ backbone for point cloud feature extraction, it outperforms H3DNet with PointNet++ backbones. Moreover, it achieves better performance on the categories (e.g. “cabinet”, “chair”, “sofa”, “table”, “counter” and “desk”) with irregular sizes or shapes, as its back-tracing and revisiting process removes the outliers from the votes and enables better mutual agreement between the votes and the local object surfaces, whilst its class-agnostic regression strategy makes the estimation process robust to shape variations.
Qualitative results. In Fig. 4 and Fig. 6, we visualize the representative 3D object detection results, from our method and the baseline methods, such as VoteNet , MLCVNet and H3DNet . These results demonstrate that our method achieves more reliable detection results with more accurate bounding boxes and orientations. Our method also eliminates false positives and discovers more missing objects when compared with the baseline methodsMLCVNet does not provide a checkpoint for the SUN RGB-D dataset thus we cannot provide its visualization results on this dataset..
3 Ablation Study and Discussions
Class-agnostic bounding box regression. Our method regresses the representative points in a class-agnostic way, which are then converted to the proposal’s bounding boxes. Note VoteNet and its variants have to estimate the sizes of object proposals in a class-aware way. Thus these baseline methods usually output the object sizes that can only moderately vary around the class-aware templates, and tend to falsely detect the objects when their sizes are unusual. To validate this observation, we implement an alternative method that employs a similar regression strategy as in our method but shares the same network as VoteNet . We term this variant as “VoteNet+CA-Reg”. As shown in Table 3, this variant significantly outperforms VoteNet. As shown in Figure 5, we also observe that this alternative method works better for the categories with high intra-category variance in sizes, and the mAP@ gains of this alternative method over VoteNet on the SUN RGB-D dataset are positively related to size variances.
Back-tracing, revisiting and refinement. Back-tracing the representative points should also be combined with the subsequent revisiting and refinement modules. As shown in Table 3, we find this complete method has significant performance gains (% mAP improvement on ScanNet and % mAP improvement on SUN RGB-D in terms of mAP@) over the aforementioned baseline. The back-tracing operation gives rough estimation of the object extent, and the revisiting and refining operations further update the proposal features with the reliable seed features in the neighborhood, thus offering better chance to produce more accurate detection results. Moreover, as shown in Figure 7, the revisited seed points by our method compactly cover the object’s surface, while the corresponding seed points retrieved by the votes can only partially cover the surface, and also suffer from the outliers.
Moreover, to validate whether the seed points can help improve the object detection results, we consider another variant (termed as “VoteNet+Seed-Pts”) that VoteNet has its vote features fused with the corresponding seed points’ features. In comparison to VoteNet, this alternative method also achieves non-trivial gains on both datasets, especially on ScanNet V2 in terms of mAP@.
Sampling strategy of representative points. In Table 4, we compare different sampling strategies to generate our representative points. “Ray” means uniform sampling along directions between and the maximum offsets. “Grid” means uniform sampling within the 3D bounding box spanned based on the predicted offsets. “#Pts” is the number of sampled points. Our methods using different strategies are generally comparable.
Model size and speed. As listed in Table 5, our proposed method is efficient in comparison to VoteNet, and is faster than the current state-of-the-art H3DNet , when evaluated on both datasets. Its model size is marginally increased from that of VoteNet, and around smaller than that of H3DNet. Knowing that the proposed method has significant performance gains than these reference methods (as discussed in Sec. 4.2), its lightweight model validates that the proposed back-tracing strategy is significant for 3D object detection in point cloudsNote that MLCVNet does not provide a checkpoint for the SUN RGB-D dataset, we omit its comparison on this dataset..
Number of Backbones. Our BRNet can also be improved after using backbones, and it achieves the result of 51.8% in terms of mAP@ on ScanNet , which outperforms H3DNet ( backbones) with a remarkable margin (+3.7%).
Conclusion
In this work, we have introduced a new approach to improve the voting-based 3D object detection method by generatively and class-agnostically back-tracing the representative points. We revisit the seed points around the back-traced representative points and extract fine object surface features to generate the high-quality object proposals. Comprehensive ablation studies show the importance and effectiveness of the proposed back-tracing, revisiting and refinement operations. Qualitative and quantitative results further demonstrate that our method remarkably outperforms the existing methods while bringing negligible increases in model size and executive time compared with VoteNet .
Acknowledgements. This work was supported by Key Research and Development Program of Guangdong Province, China, under Grant No. 2019B010154003, and the National Natural Science Foundation of China under Grant No. 61906012. We thank Zizheng Que and Zinuo You for valuable discussions and feedback.
References
A Supplementary
This supplementary provides more quantitative results of our method (Sec. A.1), more qualitative results (Sec. A.1), and finally implementation details (Sec. A.3).
Finer performance evaluations. We try to evaluate our method using mean average precision with multiple IoU thresholds for finer performance evaluations in Table S1 and S2. We use mAP@0.25, mAP@0.50, mAP@0.75 to evaluate different methods, i.e. VoteNet , HGNet , MLCVNet , H3DNet and our BRNet .
Our method performs the best on the metrics mAP@0.50 and mAP@0.75. Notably, mAP@0.75 requires more than coverage in each dimension of a bounding box, which is very challenging for a detector. Our method gains %, %, % increase on mAP@0.25, mAP@0.50, mAP@0.75 compared with H3DNet using PointNet++ backbones and doubled input point clouds (i.e., points by our BRNet , and points by H3DNet) on the ScanNetV2 dataset. On more challenging evaluation metrics, our method has more gain, which shows the importance of our representative point generation, and its benefits for seed points revisiting and finer surface feature extraction to accurately detect objects with more reliable bounding boxes.
Per-category results. We show the per-category results on ScanNet V2 dataset with 3D IoU threshold 0.25 in Table S3, and the per-category results on SUN RGB-D with both 3D IoU thresholds 0.25 and 0.50 in Table S4 and S5. In terms of the accuracy about the object detection, our approach outperforms the baseline VoteNet and prior state-of-the-art method H3DNet significantly. For objects in the SUN RGB-D dataset, our approach can gain %, %, %, %, % increase on Bathtub, Bed, Dresser, Nightstand and Sofa compared with H3DNet . These improvements are achieved by using back-tracing and seed points revisiting to better capture object surface features.
A.2 More Qualitative Results
We provide more qualitative comparisons between our method and the top-performing reference methods, such as VoteNet , MLCVNet and H3DNet , on the ScanNet V2 and SUN RGB-D datasets, as shown in Fig. S2 and Fig. S3, respectively. Our method can generate high-quality and compact predicted bounding boxes compared with the other reference methods.
We also show two typical failure cases in Fig. S1. Our BRNet cannot avoid the existence of false positive predicted bounding boxes which appear on the hollow floor. Also, it is hard for our method to detect objects on the smooth wall, especially windows and pictures. We need to mention that these failure cases are also common, and hard for the reference methods. It is an interesting and significant future direction of our work to tackle these false positives when points are over sparse and increase the robustness when perceiving objects within the cluttered background.
A.3 Implementation Details
As mentioned in the main paper, the BRNet consists of four modules: (1) vote generation and clustering, (2) back-traced representative points generation, (3) seed points revisiting, and (4) proposal refinement and classification followed by 3D NMS. Here we elaborate the implementation details with respect to each module.
Vote generation and clustering. We follow the same network architecture and vote regression loss as in VoteNet .
Representative point generation. It has output sizes of , , for the three MLP layers, where is the number of heading bins for estimating the orientations, is the distance offsets from vote point to object surface (front/back/left/right/up/down) in the canonical coordinate centered at the vote point. Then we sample representative points on each skewed direction as the back-traced representative points, thus we have representative points per proposal.
Seed point revisiting. We use the set abstraction module (SA module) to aggregate seed points features within m radius surrounding a back-traced representative point. The SA module has the output size of , , for the MLP layers. After revisiting seed points, we get a dimensional feature vector for each representative point. We concatenate the representative point features in a predefined local-structure-aware order to a dimensional feature vector per proposal. The feature vector is then projected to -dimensional as the captured surface feature of the object proposal.
Proposal refinement and classification. The input is the dimensional fused feature vector which is the concatenation of -D vote cluster feature and -D revisited seed point feature. Then the fused feature is fed into a three-layer MLP, whose output sizes are , , . is the number of semantic classes, i.e., for the SUN RGB-D dataset and for ScanNet V2 dataset . In the first 9 channels, the first two are for objectness classification, the following one is for heading angle refinement and the last six are for distance offsets refinement.