RBGNet: Ray-based Grouping for 3D Object Detection
Haiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang, Qi Qian, Hongsheng Li, Bernt Schiele, Liwei Wang
Introduction
3D object detection is becoming an active research topic in computer vision, which aims to estimate oriented 3D bounding boxes and semantic labels of objects in 3D scenes. As a fundamental technique for 3D scene understanding, it plays a critical role in many applications, such as autonomous driving , augmented reality and domestic robots . Unlike the scenarios in the well-studied 2D image problems, 3D scenes are generally represented by point clouds, a set of unordered, sparse and irregular points captured by depth sensors (\eg, RGB-D cameras, LiDAR sensors), which makes it significantly different from traditional regular input data like images and videos.
Previous 3D detection approaches can be coarsely classified into two lines in terms of point representations, \ie, the grid-based methods and the point-based methods. The grid-based methods generally convert the irregular points to regular data structure such as 3D voxels or 2D bird’s eye view maps . Thanks to the great success of PointNet series , the point-based methods directly extract the point-wise features from the irregular and points. These point-wise features are generally enhanced by various feature grouping modules for predicting the 3D bounding boxes. However, these feature grouping strategies have not well explored the fine-grained surface geometry to help improve the performance of 3D box generation.
We argue that feature grouping module plays an important role in point-based 3D detectors, and how to better incorporate the foreground object geometry features to enhance the quality of point-wise features is the key to predict better 3D bounding boxes. As shown in Table 1, for the popular VoteNet point-wise 3D detector, by simply grouping the features of accurate object surface points to the features of their correct vote centers, the performance can be improved dramatically with a gain of 13.31 on mAP@0.25 for explicit usage of ground truth labels ( row of Table 1), and a gain of 8.79 for implicit usage of ground truth labels ( row of Table 1). Here the “explicit usage” indicates that the ground truth labels are not only utilized for grouping the object surface points but also for replacing the vote centers with ground truth centers, while the “implicate usage” means the ground truth labels are only used for grouping the object surface points. These facts inspire us to explore on designing a better feature representation for the surface geometry of foreground objects, to help the prediction of 3D bounding boxes.
Hence, we present a new 3D detection framework, RBGNet, which is a one-stage 3D detector for 3D object detection from raw point clouds. Our RBGNet is built on top of VoteNet , and we propose two novel strategies to boost the performance of 3D object detection by implicitly learning from foreground object features.
Firstly, we propose the ray-based feature grouping that could learn better feature representation of the surface geometry of foreground objects. The learned features are utilized to augment the cluster features for 3D boxes estimation. Specifically, we formulate a ray-based mechanism to capture the object surface points, where a number of rays are uniformly emitted from the cluster center with the determined angles (see Fig. 1). The far bounds of the rays are based on our predicted object scale of this cluster. Then a number of anchor points is densely sampled on each ray, where the aggregated local features of each anchor point are utilized to predict whether they are on the object surface to learn the geometry shape. Moreover, a coarse-to-fine strategy is proposed to generate different number of anchor points based on the sparsity of different regions. The learned features from all the anchor points will be finally aggregated to boost the features of cluster centers for predicting 3D bounding boxes. The experiments (Table 4) show that our ray-based feature grouping strategy can effectively encode the surface geometry of foreground objects and significantly improves 3D detection performance.
Secondly, we propose the foreground biased sampling strategy to allocate more foreground object points for predicting 3D boxes. We observe that the points on object surfaces are more useful than those on the background for 3D box estimation (similar observations are also mentioned by ), and row of Table 1 shows that by conducting farthest point sampling only on the ground truth foreground points, the performance of VoteNet could be boosted from 62.90 to 71.27 in terms of mAP@0.25. Therefore, we propose a simple but effective strategy to sample points biased towards object surface while still keeping the coverage rate of the whole scene. Specifically, we append a segmentation head to the point-wise features before each farthest point sampling, where the head will predict the confidence of each point being a foreground point. According to the ranking of their foreground scores, the input points are separated into foreground set and background set. And these two sets will apply farthest point sampling separately, where we sample most target points (\ie, 87.5% in our case) from the foreground set and a small number (\ie, 12.5%) from the background set to keep the coverage rate of the whole scene. Our foreground biased sampling can produce a more informative sampling of points over foreground objects surface for feature extraction, and the performance gains (Table 4) demonstrate its effectiveness.
In a nutshell, our contributions are three-fold: 1) We propose a novel ray-based feature grouping module to encode object surface points with determined rays, which can learn better surface geometry features of objects to boost the performance of point-based 3D object detectors. 2) We present foreground biased sampling module to focus feature learning of the network on foreground surface points while also keeping the coverage rate for the whole scene, which can incorporate more object points to benefit point-based 3D box generation. 3) Equipped with the above two modules, our proposed RBGNet framework outperforms state-of-the-art methods with remarkable margins both on ScanNetV2 and SUN RGB-D .
Related Work
3D Object Detection is challenging due to the irregular, sparse and orderless characteristics of 3D points. Most existing works could be classified into two categories in terms of point cloud representations, \ie, grid-based and point-based. Grid-based approaches transform point clouds to regular data, such as 2D grids or 3D voxels . 2D grid methods project point clouds to a bird view before proceeding to the rest of the pipeline. Voxel-based methods convert the point clouds into 3D voxels to be processed by 3D CNN or efficient 3D sparse convolution , which greatly facilitate 3D object detection. Popularized by PointNet and its variants , point-based methods have become extensively employed on estimating bounding box directly from raw points. Most of existing methods can be considered as a bottom-up manner, which requires point grouping step to obtain object features. Point R-CNN groups point features within the 3D proposals by point cloud region pooling. VoteNet applies Hough Voting to group the points that vote to the similar center region. Group-free implicitly groups point features by an attention module. Although these methods have explored various feature grouping strategies, they have not leveraged the surface geometry of foreground objects. We propose a novel ray-based feature grouping module to encode object shape distribution with determined rays, and the learned features are used to further boost 3D detection performance.
2D Shape Representation. Shape representation is of particular interest due to the ability to explicitly describe the 2D object shape with points. Polar Mask and ESE-SEG both use polar representation to model the object boundary and then regress the object locations as well as the length of rays emitting uniformly from the object centroids. However, these shape representations may fail to model 3D object surface, because of the limited expressive ability on concave shape, the difficulty of inner center definition. We design a ray-based 3D shape representation to effectively model object surface geometry.
Point Cloud Sampling. Sampling aims to represent the original point cloud in a sparse way, plays a key role in point cloud analysis. Farthest point sampling (FPS) has been widely used as a pooling operation , since it can uniformly sample distributed points. However, FPS is agnostic to downstream tasks by a predefined rule, foreground instances with few interior points may lose all points after sampling. 3DSSD applies a fused FPS based on feature and euclidean distance, but still does not focus on foreground points explicitly. To deal with the dilemma, we design a simple but effective strategy, foreground biased sampling, to sample more points on object surface while still keeping the coverage rate of the whole scene.
Methodology
This section describes the technical details of the proposed RBGNet detector. §3.1 briefly presents the overview of our approach. Next, §3.2 to §3.4 elaborate on the network design and the learning objective.
RBGNet is a one-stage 3D object detection framework aiming at more accurate bounding box estimation from irregular point clouds. As illustrated in Fig. 2, RBGNet consists of three major components: i) a backbone network with foreground biased sampling to extract feature representation from point clouds, ii) a ray-based feature grouping module to effectively capture the points on object surface and learn from the shape distribution to augment cluster feature and iii) a proposal and classification module followed by 3D non-maximum-suppression (NMS). Our paper mainly focuses on the sampling and grouping modules, so we follow the same proposal and classification strategy as in VoteNet to estimate final bounding boxes. We will describe the technical details in the following parts.
2 Ray-based Feature Grouping
VoteNet has shown tremendous success for 3D object detection. After getting the seed points from the backbone (PointNet++ ), it reformulates traditional Hough voting, and generates object candidates by grouping the seed points whose votes are within the same cluster. The aggregated feature is then used to estimate the 3D bounding boxes and associated semantic labels. However, the quality of the grouping principally determines the reliability of proposal features and detector performance. Some follow-up works are actually trying to solve this problem, but they have not well explored on the fine-grained surface geometry of foreground objects. To address this limitation, we propose the ray-based feature grouping module, which can effectively encode the shape distribution and learn better object features to enhance 3D detection performance.
We first illustrate the process of our proposed ray point representation, where two types of anchor points are generated on each ray to encode the object geometry around the cluster centers. These anchor points of all rays will be utilized for the final feature enhancement in §3.2.2.
The polar angle is split into bins, and each bin corresponds a round surface that is perpendicular to the -axis. The angle of bin is:
The number of rays (denoted as ) terminated on the round surface is calculated as follows:
where is a hyper-parameter to indicate the factor for the number of sampled rays in each round surface.
With and , a ray could be determined. Its azimuth angle and polar angle could be formulated as follows:
Our adopted strategy could generate more uniformly distributed rays to better cover the surrounding region of clusters. Given the polar-bin number , the number of rays is (\ie in our case). Note that more rays will be generated when the polar angle is closer to .
As for the far bounds of the rays of each cluster, all the rays are of the same length as the object scale , which is predicted based on the cluster features . Here, we explicitly supervise the object scale by regression loss
Firstly, in coarse stage, as for the ray, we sample a set of anchor points as
where is the number of anchor points sampled on each ray, and the anchor points are generated by stratified sampling to evenly partition the ray into bins.
To extract local feature of each anchor point, we apply set abstraction to aggregate the features of the seed points around each anchor point. The aggregated local features of anchor point is denoted as . Finally, we append a binary classification module for estimating the positive mask of point based on cluster feature and local feature as follows:
where the ground-truth of positive masks are calculated by applying ball query operation for each anchor point. We assign positive label to an anchor point if some surface points of its assigned GT object are within its ball query region, or the anchor point will be assigned with a negative label. Hence this point mask module could predict whether each anchor point belongs to its corresponding object or not.
Secondly, in fine stage, different from which computes sample probability from point density, our fine anchor points are biased towards the dense part of its corresponding object. To achieve this goal, we apply inverse transform sampling to uniformly generate some anchor points set on positive regions (predicted by the point mask module of coarse stage) of each ray. As adopted in the coarse branch, we also extract the local features and predict the positive mask for each fine anchor point.
Repeat this coarse-to-fine process on all rays, we obtain the coarse and fine local point feature set , , point mask set , and their corresponding positions of anchor points.
2.2 Feature Enhancement by Determined Rays
As discussed in §1, the fined-grained surface geometry of foreground objects plays a crucial role in generating accurate object proposals. The process of our proposed coarse-to-fine anchor point generation already encodes such a surface geometry features, since our predicted surface masks and learned local features could implicitly describe the object geometry. Here we propose to aggregate those informative features of the anchor points to enhance the quality of cluster features, where the order of rays plays an important roles in the feature aggregation.
To be specific, given the local features and , the point masks and of each anchor point, the local features of each anchor points will be masked by setting the features of negative anchor points to zeros. We denote the masked features as and .
To aggregate the learned features orderly based on the determined rays, we formulate a fusion stage to integrate point features in a predefined order of rays. The features of coarse and fine anchor points are aggregated with two separate branches. In coarse branch, the masked point features of ray, , are firstly fused into a single ray feature . It is implemented by concatenating the features of anchor points in order before being projected to a 32-dimensional features:
where means the concatenation operation. Then, in the same way, we concatenate all the ray features with a determined order and apply a two-layer MLP to generate a 128-dimensional coarse feature:
Note that the predefined order of both anchor points and rays are consistent for each proposal, but different ordering strategies do not affect the performance. The fine branch also adopts the same strategy as the coarse branch to generates a 128-dimensional feature .
Finally, the coarse and fine features are fused as:
The fused feature is finally combined with the cluster feature to improve the performance of 3D object detection.
In this way, our RBGNet models the surface geometry implicitly and roughly obtain the size and the position of a possible object in a class-agnostic way, which could greatly benefit the prediction of 3D bounding boxes.
3 Foreground Biased Sampling
The foreground points provide rich information on predicting their associated object locations and orientations, and force network to capture shape information for more accurate 3D box generation. However, the widely-adopted farthest point sampling algorithm in the backbone is agnostic to the downstream tasks and samples a lot of background points. It may bring negative effects for 3D detection. Therefore, we design a simple but effective strategy, Foreground Biased Sampling, to sample more points on foreground object surfaces while still keeping the coverage rate of the whole scene.
Given the point-wise features encoded by each set abstraction layer, we append a segmentation head for estimating the confidence of each points. The ground-truth segmentation mask is naturally provided by the 3D ground-truth boxes. To be specific, for example, after going through the first SA layer of standard PointNet++ , we obtain 2048 downsample point set with xyz and 128-dimensional feature . Then the segmentation head scores each point to be a foreground point or not as:
We sort the confidence scores, select top to form a foreground set and the rest are the background set . Due to the concentration of high score points, there is a trade-off between the recall of foreground points and the sampling coverage for the whole scene. Based on this observation, we apply farthest point sampling on foreground and background set separately, and combine them into the final sample set as follows:
where and , and are the sample number of foreground and background set. is the set of final sampled points, which contains more object points while still keeping the coverage rate of the whole scene. In our case, we sample most target points,(\ie, 87.5%) from the foreground set and a small number (\ie, 12.5%) from background set. For example, in downsample process of the SA layer (2048 1024), , and are 1024, 896 and 128, respectively.
We adopt the cross entropy loss for foreground segmentation. In inference, the confidence score is obtained by the margin between positive class and negative class.
4 Learning Objective
The loss function consists of foreground biased sampling , voting regression , ray-based feature grouping , objectness , bounding box estimation , and semantic classification losses.
Following the setting in VoteNet , we use the same label assignment and loss terms , , and . is a cross entropy loss used to supervise foreground sampling (see §3.3). is the sum loss of ray-based feature grouping module defined as follows:
Experiments
We evaluate our method on two large-scale indoor 3D scene datasets, \ie, ScanNet V2 and SUN RGB-D , and we follow the standard data splits for both of them.
SUN RGB-D is a single-view RGB-D dataset for 3D scene understanding, which consists of 5K RGB-D training images annotated with the oriented 3D bounding boxes and the semantic labels for 10 categories. Following the standard data processing in , we convert the depth images to point clouds using the provided camera parameters.
ScanNet V2 consists of richly-annotated 3D reconstructions of indoor scenes. It consists of 1513 training samples (reconstructed meshes converted to point clouds) with axis-aligned bounding box labels for 18 object categories. Compared to SUN RGB-D, its scenes are larger and more complete with more objects. We sample point clouds from the reconstructed meshes by following .
For both datasets, the evaluation follows the same protocol as in VoteNet using mean average precision(mAP) under different IoU thresholds, \ie, 0.25 and 0.5.
2 Implementation Details.
Network Architecture Details. For each 3D scene in the training set, we subsample 50000 points from the scene point cloud as the inputs. For the backbone and voting layers, we follow the same network structure of , but replace FPS with our proposed Foreground Biased Sampling (FBS) in SA layers. More network details about other parts are given in Appendix.
Training Scheme. Our network is end-to-end optimized by using the AdamW optimizer with the batch size 8 per-GPU and initial learning rate of 0.006 for ScanNet V2 and 0.004 for SUN RGB-D. We train the network for 360 epochs on both datasets, and the initial learning rate is decayed by 10x at the 240-th epoch and the 330-th epoch. The gradnorm clip is applied to stabilize the training dynamics.
3 Comparison with state-of-the-art methods.
For performance benchmarking, we compare with a wide range state-of-the-art methods on ScanNet V2 and SUN RGB-D. We follow the previous work and also report both best results and average results. ScanNet V2. The results are summarized in Table 2. With the same backbone network of a standard PointNet++, our approach achieves 70.2 mAP@0.25 and 54.2 mAP@0.5 using 66 rays and 256 object candidates, which is 2.5 and 3.3 better than previous best methods using the same backbones. With stronger backbones and more sampled object candidates just like , \ie, more channels and 512 candidates, our approach is also improved dramatically, achieving 70.6 mAP@0.25 and 55.2 mAP@0.5, which is still 1.5 and 2.4 better than . Notably, we also only use geometric input (point cloud) as previous works did. SUN RGB-D. We also evaluate our RBGNet against several competing approaches on SUN RGB-D dataset, which is also a standard benchmark for 3D object detection. The results are summarized in Table 3. Our base model with standard PointNet++ achieves 64.1 on mAP@0.25 and 47.2 on mAP@0.5, which outperforms all previous state-of-the-arts point-only methods.
4 Ablation Studies and Discussions
In this section, a set of ablative studies are conducted on ScanNet V2 dataset, to investigate effectiveness of essential components of our algorithm. We follows and report the average performance of 25 trials by default.
Effect of ray-based representation. We first ablate the effects of ray-based feature grouping in Table 4, 5 and 6. As evidenced in the first three rows in Table 4, with ray-based feature grouping, our model performs better, \ie, 66.2 69.0, 48.2 52.9 on mAP@0.25 and mAP@0.5. Note that our model is implemented based on a strong baseline. The first row in Table 4 is actually the VoteNet with corner loss regularization, vote sampling in vote aggregation layer and optimized hyper-parameters. Even on such a strong baseline (almost close to state-of-the-art already, 66.2 vs 66.6 on mAP@0.25), ray-based feature grouping module still boosts our model with a remarkable improvement.
Our ray-based feature grouping module also works well in a wide range of hyper-parameters, such as the number of rays. Table 5 shows its performance with different ray number. More rays can bring significant performance improvement, especially in the mAP@0.5. Compared with the setting without any rays, our ray-102 model performs much better on mAP@0.25 and mAP@0.5 by 2.8 and 4.9, respectively. For the recall of object points, the second column shows that more rays can capture surface points more completely. Considering the trade-off between memory usage and performance improvement, our model finally adopts the variant with 66 rays though more rays is better.
To further demonstrate the effectiveness of ray-based feature grouping module, we refer several grouping strategies in 3D object detection as baselines and compare with them. For a fair comparison, we only switch the feature aggregation mechanism while all other settings remain unchanged. Table 6 shows that our approach achieves more reliable detection results than others with a remarkable margin (1.5 on mAP@0.25 and 3.1 on mAP@0.5). Effect of foreground biased sampling. Table 4 also demonstrates the effectiveness of the foreground biased sampling strategy. We can observe that, it improves the performance in both settings with and without feature grouping module. This verifies the necessity of sampling more foreground points for 3D object detection tasks. To further ablate the effectiveness of FBS, we compare the foreground points recall of SA layers among different sub-sampling methods in Table 7. Our sampling strategy draws better performance with a large margin.
5 Inference Speed.
The realistic inference speed of our method is competitive with other state-of-the-art methods. For a fair comparison, all experiments are run on the same workstation (single NVIDIA Tesla V100 GPU, 256G RAM, and Xeon E5-2650 v3). The results are shown in Table. 8. Our method achieves better performance with a competitive speed.
Conclusion
In this paper, we have presented the RBGNet, a novel framework for 3D object detection from point clouds. We propose the ray-base feature grouping module, which can encode object surface geometry with determined rays and learn better geometric features to boost the performance of point-based 3D detectors. We also introduce the foreground biased sampling to sample more points on object surface while keeping the coverage rate for the whole scene. All of the above designs enable our model to achieve state-of-the-art performance on ScanNet V2 and SUN RGB-D benchmarks with remarkable performance gains.
Acknowledgments. Liwei Wang was supported by National Key RD Program of China (2018YFB1402600), BJNSF (L172037) and Alibaba Group through Alibaba Innovative Research Program. Project 2020BD006 supported by PKUBaidu Fund.
References
Appendix A Details on Fine Sampling
To be specific, we first normalize the coarse point masks as to produce a piecewise-constant probability density function (PDF). Then we translate it into the cumulative distribution function (CDF). Finally, sampling with the CDF at uniform steps concentrates samples around regions with positive coarse point masks.
To further illustrate it, we provide a demo case and visualize it in Fig.6. The number of coarse and fine anchor points on each ray is 8 and 10 respectively.
The predicted coarse point masks of ray are:
We compute the piece-wise PDF by normalizing :
Then we convert PDF to CDF,
Finally, sample 10 points based on the CDF at uniform steps and inverse them to original distribution. As demonstration in Fig.6, the relative distance of fine points from object center are as follows:
Appendix B Implementation Details.
As mentioned in the main paper, the RBGNet architecture consists of a backbone with foreground biased sampling, a voting layer, a ray-based feature grouping module and a proposal module.
The backbone network, based on the PointNet++ architecture, has four set abstraction layers and two feature up-sampling layers. We follow the same layer parameters (\egball-region radius, number of sample points and MLP channels) as VoteNet . To sample points biased towards object surface, we append a segmentation head for estimating the foreground confidence of each point. The detailed layer parameters are shown in Table 9. The voting module is the same as VoteNet. Note that, in training stage, we generate proposals from the votes by vote FPS (samples clusters based on votes’ XYZ), in test stage, we apply vote FPS on ScanNet V2 and seed FPS on SUN RGB-D (sample on seed XYZ and then find the votes corresponding to the sampled seeds).
The proposal module is a two-layer MLP. We follow on how to estimate the 3D bounding boxes, except for size prediction that we adopts class-agnostic head to regress bounding box size directly. The layer’s output has 5+2NH+3+NC where the first five channels are for objectness classification and center regression (relative to the vote cluster center), 2NH channels are for heading bins classification and offsets regression, 3 is the scale regression for height, width and length, NC is the number of semantic classes. In SUN RGB-D: NH = 12, NC = 10, and in ScanNet: NH = 1, NC = 18, due to the axis aligned bounding box.
B.2 RBGNet loss function details.
As mentioned in the main paper, our model is trained end-to-end with a multi-task loss including foreground biased sampling , voting regression , ray-based feature grouping , objectness , bounding box estimation , and semantic classification losses.
Following the setting in VoteNet , we use the same loss terms , , and , but is class-agnostic and contains an additional corner loss defined in for accurate bounding box estimation,
As discussed in §3.4, is a cross entropy loss used to supervise foreground sampling (see §3.2). is the sum loss of ray-based feature grouping module defined as follows:
B.3 Other Grouping Mechanisms
To further ablate the effectiveness of our ray-based feature grouping module, we refer several grouping strategies in 3D object detection as baselines and compare with them in §4.4. For a fair comparison, we only switch the feature aggregation mechanism while all other settings remain unchanged (\egbackbone with FBS, vote-FPS, proposal module). Here we give some detailed descriptions.
Voting. The voting mechanism is first introduced by VoteNet . In our implementation, it is actually the VoteNet equipped with FBS, corner loss, vote-FPS in test stage and optimized hyperparameters.
RoI-Pooling. For a given object proposal, the points within the predicted box are aggregated together. We adopt the similar implementation with , predict the bounding boxes based on the voting cluster features and aggregate the points within the corresponding boxes by max-pooling. Finally, the aggregated features and cluster features are concatenated for 3D object detection.
Back-Tracing. It is first formulated in . In our implementation, it is similar to RoI-Pooling described above, except that the prediction of bounding boxes is replaced by the maximum offsets of 6 directions. And then the points are aggregated by the balls uniformly sampled along the rays to enhance 3D bounding box estimation.
Group-free. We replace the ray-based feature grouping module with a transformer network adopted in Group-free . For a fair comparison, we use vote-FPS for initial object candidate sampling instead of KPS. Then, the same as group-free , we adopt the transformer as the decoder to leverage all the seed points to compute the object feature of each candidate.
Appendix C Ray-based Representation Discussion
There are many choices for anchor point generation, such as classification and regression .
Regression: Given the center and surface points of an object, they want to represent the shape by polar coordinates. The length of rays can be computed easily. Then, the model regresses the length of each ray and captures the points when the ray terminates somewhere.
Classification: Given an instance, model predicts far bounds of all rays and samples a fixed number of potential query points on each ray, and then extracts local features and classifies those points whether belong to corresponding object to generate reasonable anchor points. Our model adopts this way and we will discuss why do we choose it.
In 2D perception community, some methods also represent object shape by rays in regression way, which applies the angle and distance as the coordinate to locate points. However, due to the particular property of point clouds, this regression pipeline has many problems in 3D scenario, \ie, i) center is outside of the object, no intersection with the object surface at some angles, ii) limited expressive ability on concave shape, one ray may has multiple intersections. Compare to regression, classification pipeline is more reasonable to represent point clouds and doesn’t have the above problems, so we choose it to generate anchor points. In Table 15, we show the results of the two representations, classification approach performs better than regression.
Appendix D More Results
We evaluate per-category on ScanNet V2 and SUN RGB-D under different IoU thresholds. Table 10 and Table 11 report the results on 18 classes of ScanNetV2 with 0.25 and 0.5 box IoU thresholds respectively. Table 13 and Table 14 show the results on 10 classes of SUN RGB-D with 0.25 and 0.5 box IoU thresholds. Our approach outperforms the baseline VoteNet and prior state-of-the-art methods Group-free significantly in almost every category. These improvements are achieved by using ray-based feature grouping and foreground biased sampling to better encode object surface geometry.
D.2 Visualization of Positive Anchor Points
Fig.7 shows the scores of coarse anchor points predicted from our RBGNet in a typical SUN RGB-D scene. We clearly see that the high responses are almost on the object surface (bed, chair \etc) while low responses are on the empty space or background surface. This verifies that our method can really learn the shape distribution and boost point-based 3D detectors.
D.3 Quantitative Results
We provide more qualitative comparisons between our method and the top-performing reference methods, such as Group-free and VoteNet , on the ScanNet V2 and SUN RGB-D datasets. Please see Fig. 8 for more qualitative results.
Appendix E Limitations
Although our method achieves promising performance on multiple datasets, there are still some limitations. Compared with the previous approaches, the performance of RBGNet is significantly better in the case of a large number of rays. However, there is a trade-off between computational cost and performance improvement, as shown in main paper. In the future, we hope to discover approaches that can encode surface geometry more efficiently.