RangeDet:In Defense of Range View for LiDAR-based 3D Object Detection
Lue Fan, Xuan Xiong, Feng Wang, Naiyan Wang, Zhaoxiang Zhang
Introduction
LiDAR-based 3D object detection is an indispensable technology in the autonomous driving scenario. Though shared some similarities, object detection in the 3D sparse point cloud is fundamentally different from its 2D counterpart. The key is to efficiently represent the sparse and unordered point clouds for subsequent processing. Several popular representations include Bird’s Eye View (BEV) , Point View (PV) , Range View (RV) and fusion of them , which are shown in Fig.1. Among them, BEV is the most popular one. However, it introduces quantization error when dividing the space into the voxels or pillars, which is unfriendly for the distant objects that may only have few points. To overcome this drawback, the point view representation is usually incorporated. Point view operators can extract effective features from unordered point clouds, but they are difficult to scale up to large-scale point cloud data efficiently in autonomous driving scenes.
The range view is widely adopted in semantic segmentation task , but it is rarely used in object detection task individually. However, in this paper, we argue that the range view itself is the most compact and informative way for representing the LiDAR point clouds because it is generated from a single viewpoint. It essentially forms a 2.5D scene instead of a full 3D point cloud. Consequently, organizing the point cloud in range view misses no information. The compactness also enables fast neighborhood queries based on range image coordinates, while point view methods usually need a time-consuming ball query algorithm to get the neighbors. Moreover, the valid detection range of range-view-based detectors can be as far as the sensor’s availability, while we have to set a threshold for the detection range in BEV-based 3D detectors. Despite its advantages, an intriguing question raised, Why are the results of range-view-based LiDAR detection inferior to other representation forms?
Indeed some works have made attempts to make use of the range view representation from the pioneering work VeloFCN to LaserNet to the recently proposed RCD . However, there is still a huge gap between the pure range-view-based method and the BEV-based method. For example, on the validation split of Waymo Open Dataset (WOD) , they are still lower than state-of-the-art methods by a large margin (more than 20 points 3D AP in vehicle class).
To liberate the power of range view representation, we examine the designs of the current range-view-based detectors and found several overlooked facts. These points seem simple and obvious, but we find that the devils are in the details. Properly handling these challenges is the key to high-performance range-view-based detection.
First, the challenge of detecting objects with sparse points in BEV is converted to the challenge of scale variation in the range image. Though there have been many methods in 2D object detection tried to address this issue, this challenge is never seriously considered in the range-view-based 3D detector.
Second, unlike in 2D image, though the convolution on range image is conducted on 2D pixel coordinates, while the output is in the 3D space. This point suggests an inferior design in the current range-view-based detectors: both the kernel weight and aggregation strategy of standard convolution ignore this inconsistency, which leads to severe geometric information loss even from the very beginning of the network.
Third, the 2D range view is naturally more compact than 3D space, which makes feature extractions in range-image-based detectors more efficient. However, how to utilize such characteristics to improve the performance of detectors is ignored by current range-image-based designs.
In this paper, we propose a pure range-view-based framework – RangeDet, which is a single-stage anchor-free detector designated to address the aforementioned challenges. We analyze the defects of the existing range-view-based 3D detector and point out the aforementioned three key challenges that need to be addressed. For the first challenge, we propose a simple yet effective Range Conditioned Pyramid to mitigate it. For the second challenge, we propose Meta-Kernel to capture 3D geometric information from 2D range view representation. For the third one, we use weighted Non-Maximum Suppression to remedy the issue. In addition to these techniques, we also explore how to transfer common data augmentation techniques from 3D space to the range view. Combining all the techniques, our best model achieves comparable results with state-of-the-art works in multiple views. And we surpass previous pure range-view-based detectors by a margin of 20 3D AP in vehicle detection. Interestingly, in contrast to common belief, RangeDet is more advantageous for farther or small objects than BEV representation.
Related Work
BEV-based 3D detectors. Several approaches for LiDAR-based 3D detection discretize the whole 3D space. 3DFCN and PIXOR encode handcrafted features into voxels, while VoxelNet is the first to use end-to-end learned voxel features. SECOND accelerates VoxelNet by sparse convolution. PointPillars is aggressive in feature reduction that it applies PointNet to collapse the height dimension first and then treat it as a pseudo-image.
Point-view-based 3D detectors. F-PointNet first generates frustums corresponding to 2D Region of Interest (ROI), then use PointNet to segment foreground points and regress the 3D bounding boxes. PointRCNN generates 3D proposals directly from the whole point clouds instead of 2D images for 3D detection with point clouds by using PointNet++ both in proposal generation and refinement. IPOD and STD are both two-stage methods which use the foreground point cloud as a seed to generate proposals and refine them in the second stage.
Range-view-based 3D detectors. VeloFCN is a pioneering work in range image detection, which projects point cloud to 2D and applies 2D convolutions to predict 3D box for each foreground point densely. LaserNet uses a fully convolutional network to predict a multimodal distribution for each point to generate the final prediction. Recently, RCD addresses the challenges in range-view based detection by learning a dynamic dilation for scale variation and soft range gating for the “boundary blur” issue as pointed in Pseudo-LiDAR .
Multi-view-based 3D detectors. MV3D is the first work to fuse features in frontal view, BEV, and camera view for 3D object detection. PV-RCNN jointly encodes point and voxel information to generate high-quality 3D proposals. MVF endows a wealth of contextual information from different perspectives for each point to improve the detection of small objects.
2D detectors. Scale variation is a long-standing problem in 2D object detection. SNIP and SNIPER rescale proposals to a normalized size explicitly based on the idea of image pyramids. FPN and its variants build feature pyramids, which have become the indispensable component for modern detectors. TridentNet constructs weight-shared branches but using different dilation to build scale-aware feature maps.
Review of Range View Representation
In this section, we quickly review the range view representation of LiDAR data.
For a LiDAR with beams and times measurement in one scan cycle, the returned values from one scan form a matrix, called range image (Fig. 1). Each column of the range image shares an azimuth, and each row of the range image shares an inclination. They indicate the relative vertical and horizontal angle of a returned point w.r.t the LiDAR original point. The pixel value in the range image contains the range (depth) of the corresponding point, the magnitude of the returned laser pulse called intensity and other auxiliary information. One pixel in the range image contains at least three geometric values: range , azimuth , and inclination . These three values then define a spherical coordinate system. Fig. 2 illustrates the formation of the range image and these geometric values.
The commonly-used point cloud data with Cartesian coordinates is actually decoded from the spherical coordinate system:
where denote the Cartesian coordinates of points. Note that range view is only valid for the scan from one viewpoint. It is not available for general point cloud since they may overlap for one pixel in the range image.
Unlike other LiDAR datasets, WOD directly provides the native range image. Except for range and intensity values, WOD also provides another information called elongation . The elongation measures the extent to which the width of the laser pulse is elongated, which helps distinguish spurious objects.
Methodology
In this section, we first elaborate on three components of RangeDet. Then the full architecture is presented.
In 2D detection, feature-pyramid-based methods such as Feature Pyramid Network (FPN) are usually adopted to address the scale variation issue. We first construct the feature pyramids as in FPN which is illustrated in Fig. 4. Although the construction of the feature pyramid is similar to that of FPN in 2D object detection, the difference lies in how to assign each object to a different layer for training. In the original FPN, the ground-truth bounding box is assigned based on its area in the 2D image. Nevertheless, simply adopting this assignment method ignores the difference between the 2D range image and 3D Cartesian space. A nearby passenger car may have similar area with a far away truck but their scan patterns are largely different. Therefore, we designate the objects with a similar range to be processed by the same layer instead of purely using the area in FPN. Thus we name our structure as Range Conditioned Pyramid (RCP).
2 Meta-Kernel Convolution
Compared with the RGB image, the depth information endows range images with a Cartesian coordinate system, however standard convolution is designed for 2D images on regular pixel coordinates. For each pixel within the convolution kernel, the weights only depend on the relative pixel coordinates, which can not fully exploit the geometric information from the Cartesian coordinates. In this paper, we design a new operator which learns dynamic weights from relative Cartesian coordinates or more meta-data, making the convolution more suitable to the range image.
For better understanding, we first disassemble standard convolution into four components: sampling, weight acquisition, multiplication and aggregation.
1) Sampling. The sampling locations in standard convolution is a regular grid , which has relative pixel coordinates. For example, a common sampling grid with dilation 1 is:
For each location on the input feature map , we usually sample feature vectors of its neighbors , using im2col operation.
3) Multiplication. We decompose the matrix multiplication of the standard convolution into two steps. The first step is pixel-wise matrix multiplication. For each sampling point , its output is defined as
4) Aggregation. After multiplication, the second step is to sum over all the in , which is called channel-wise summation.
In summary, the standard convolution can be presented as:
In our range view convolution, we expect that the convolution operation is aware of the local 3D structure. Thus, we make the weight adaptive to the local 3D structure via a meta-learning approach.
For weight acquisition, we first collect the meta-information of each sampling location and denote this relationship vector as . usually contains relative Cartesian coordinates, range value, etc. Then we generate the convolution weight based on . Specifically, We apply a Multi-Layer Perceptron (MLP) with two fully-connected layers:
For multiplication, instead of matrix multiplication, we simply use element-wise product to obtain as follows:
We do not use matrix multiplication because our algorithm runs on large-scale point clouds, and it costs too much GPU memory to save a weight tensor with shape . Inspired by the depth-wise convolution, the element-wise product eliminates the dimension from the weight tensor, which is much less memory-consuming. However, there is no cross-channel fusion in the element-wise product. We leave it to the aggregation step.
For aggregation, instead of channel-wise summation, we concatenate all , and pass it to a fully-connected layer to aggregate the information from different channels and different sampling locations.
Summing it up, the Meta-Kernel can be formulated as:
where is the aggregation operation containing concatenation and a fully-connected layer. Fig. 3 provides a clear illustration of Meta-Kernel.
Comparison with point-based operators. Although shares some similarities with point-based convolution-like operators, Meta-Kernel has three significant differences from them. (1) Definition space. Meta-Kernel is defined in 2D range view, while others are defined in the 3D space. So Meta-Kernel has regular neighborhood, and point-based operators have an irregular neighborhood. (2) Aggregation. Points in 3D space are unordered, so the aggregation step in point-based operators is usually permutation-invariant. Max-pooling and summation are widely adopted. neighbors in the RV are permutation-variant, which is a natural advantage for Meta-Kernel to adopt concatenation and fully-connected layer as the aggregation step. (3) Efficiency. Point-based operators involve time-consuming key-point sampling and neighbor query. For example, downsampling 160K points to 16K with Farthest Point Sampling (FPS) takes 6.5 seconds in a single 2080Ti GPU, which is also analyzed in RandLA-Net . Some point-based operators, such as PointConv , KPConv and the native version of Continuous Conv , generate a weight matrix or feature matrix for each point, so they face severe memory issue processing large-scale point cloud. These disadvantages make it impossible to apply point-based operators to large-scale point clouds (more than points) in autonomous driving scenarios.
For a clear comparison, we summarize the differences between several closely related work and our Meta-Kernel convolution in Table 1.
3 Weighted Non-Maximum Suppression
As mentioned earlier, how to utilize the compactness of range view representation to improve the performance of range-image-based detectors is an important topic. In common object detectors, a proposal inevitably has a random deviation from the mean of the proposal distribution. The straightforward way to get a proposal with small deviation is to choose the one with the highest confidence. While a better and more robust way to eliminate the deviation is using the majority votes of all the available proposals. An off-the-shelf technique just fits our need – weighted NMS . Here comes an advantage of our method: the nature of compactness makes RangeDet feasible to generate proposals in the full-resolution feature map without huge computation cost, however it is infeasible for most BEV-based or point-view-based methods. With more proposals, the deviation will be better eliminated.
We first filter out the proposals whose scores are less than a predefined threshold 0.5, and then sort the proposals as in standard NMS by their predicted scores. For the current top-rank proposal , we find the proposals whose IoUs with are higher than 0.5. The output bounding box for is a weighted average of these proposals, which can be described as:
4 Data Augmentation in Range View
Data augmentation is a critical technique to improve the performance of LiDAR-based 3D object detectors. Random global rotation, Random global flip and Copy-Paste are three typical ones. Although they are straightforward in 3D space, it’s non-trivial to transfer them to RV and preserve the structure of RV.
Rotation of point clouds can be regarded as translation of range images along the azimuth direction. Flipping in 3D space corresponds to the flipping with respect to one or two vertical axes of range images (We provide a clear illustration in supplementary materials). From the leftmost column to the rightmost, the span of azimuth is . So, unlike the augmentation of 2D RGB-image, we calculate the new coordinate of each point to keep it consistent with its azimuth. For Copy-Paste , the objects are pasted on the new range image with their original vertical pixel coordinates. Because of the non-uniform vertical angular resolution, we can only keep the structure of RV and avoid objects largely deviating from the ground by this treatment. Besides, a car in the distance should not be pasted in the front of a nearby wall. So we carry out “range test” to avoid such a situation.
5 Architecture
Overall pipeline. The architecture of RangeDet is shown in Fig. 4. The eight input range image channels include range, intensity, elongation, x, y, z, azimuth, and inclination, as described in Sec. 3. Meta-Kernel is placed in the second BasicBlock. Feature maps are downsampled to stride 16, and upsampled to full resolution gradually. Next, we assign each ground-truth bounding box to the layers of stride in RCP according to the range of the box center. All the positions whose corresponding points are in ground-truth 3D bounding boxes are treated as positive samples, otherwise negative. At last, we adopt Weighted NMS to de-duplicate the proposals and generate high-quality results.
RCP and Meta-Kernel. In WOD, the range of a point varies from 0m to 80m. According to the distribution of points in the ground-truth bounding boxes, we divide $[0,15),[15,30),$. We a use two-layer MLP with 64 filters to generate weights from relative Cartesian coordinates. ReLU is adopted as activation.
IoU Prediction head. In the classification branch, we adopt a very recent work – varifocal loss to predict IoU between the predicted bounding box and the ground-truth bounding box. Our classification loss is defined as:
where is the number of valid points, and is the point index. is the varifocal loss of each point:
where is the predicted score, and is the IoU between the predicted bounding box and the ground-truth bounding box. and play a similar role as in focal loss .
Regression head. The regression branch also contains four Conv as in the classification branch. We first formulate the ground-truth bounding box containing pointHere, a point is actually a location in the feature map and corresponds to a Cartesian coordinate. For a better understanding, we still call it a point. as to denote the coordinates of the bounding box center, dimension and orientation, respectively. The Cartesian coordinate of point is . We define the offsets between the point and the center of bounding box containing point as . For point , we regard its azimuth direction as its local -axis which is the same as in LaserNet . And we formulate such transformation as follows (Fig. 5 provides a clear illustration):
where denotes the azimuth of point , and is the transformed coordinate offset to be regressed. Such a transformed target is appropriate for range-image-based detection since an object’s appearance in the range image doesn’t change with the azimuth in a fixed range. Thus, it’s reasonable to make regression targets azimuth-invariant. So for each point, we regard azimuth direction as local -axis.
We denote the point ’s ground-truth targets set as . So the regression loss is defined as
where is the predicted counterpart of . is the number of ground-truth bounding boxes, and is the number of points in the bounding box which contains the point . The total loss is the sum of and .
Experiments
We conduct our experiments on large-scale Waymo Open Dataset (WOD), which is the only dataset that provides native range images. We report LEVEL_1 average precision in all experiments for comparing with other methods. Please refer to supplemental material for the detailed results and configuration of the pipeline. Experiments in Table 2, Table 4 and Table 10 using the entire training dataset. And we uniformly sample 25% training data (k frames) for other experiments.
We conduct extensive experiments to ablate Meta-Kernel in this section. These experiments do not involve data augmentation. We build our baseline by replacing Meta-Kernel with a 2D convolution.
Different input features. Table 3 shows the results of different meta information as input. Not surprisingly, using relative pixel coordinates (E4) only brings marginal improvements compared with the baseline, demonstrating the necessity of Cartesian information (coordinates or range) in kernel weight.
Different locations to place Meta-Kernel. We place the Meta-Kernel at stages with different strides. The results are shown in Table 5, which demonstrates that Meta-Kernel is more prominent at a lower level. This result is reasonable since the low-level layers have a closer association with geometric structure, where the Meta-Kernel takes a vital role.
Performance on small objects. Boundary information is more crucial for small objects in range view, for example pedestrian, to avoid being diluted by background than large objects. Meta-Kernel enhances the boundary information by capturing local geometric features, so it is especially powerful in small objects detection. Table 6 shows the significant effectiveness.
Comparison with point-based operators. We discussed the main differences between Meta-Kernel and point-based operators in Sec. 4.2. For a fair comparison, we implement some typical point-based operators on the 2D range image with fixed neighborhood just like our Meta-Kernel. Please refer to supplementary materials for the implementation details. Some operators such as KPConv , PointConv are not implemented due to huge memory costs. These methods all obtain inferior results as Table 7 shows. We owe it to the strategies they used for aggregation in unordered point clouds, which will be elaborated next.
Different ways of aggregation. Instead of concatenation, we try max-pooling and summation in a channel-wise manner just like other point-based operators, and Table 8 shows the results. Performance significantly drops when using max-pooling or summation as they treat the features from different locations equally. These results demonstrate the importance of keeping and utilizing the relative orders in range view. Note that other views cannot adopt concatenation due to the disorder of point clouds.
2 Study of Range Conditioned Pyramid
Instead of conditioning on the range, we try three other strategies to assign bounding boxes: azimuth span, projected area and visible area. The azimuth span of a bounding box is proportional to its width in the range image. The projected area is the area of a box projected into the range image. The visible area is the area of visible object parts. Note that area is the standard assign criterion in 2D detection. For a fair comparison, we keep the number of ground-truth boxes in a certain stride consistent between these strategies. Results are shown in Table 9. We owe the inferior results to the pose change as well as occlusion, which makes the same object fall into different layers with different pose or occlusion conditions. Such a result demonstrates that it is not enough to only consider the scale variation in the range image, since some other physical features, such as intensity, density, change with the range.
3 Study of Weighted Non-Maximum Suppression
To support our claims in Sec. 8, we apply weighted NMS in two typical voxel-based methods – PointPillars and SECOND based on the strong baselines in MMDetection3Dhttps://github.com/open-mmlab/mmdetection3d. Compared with RangeDet, the others generate fewer proposals due to memory and computational limits, which degrades the performance of weighted NMS as Table 10 shows.
4 Ablation Experiments
We further conduct ablation experiments on the components we use. Table 2 summarizes the results. Meta-Kernel is effective and robust in different settings. Both RCP and Weighted NMS significantly improve the performance of our whole system. Although IoU prediction is a common practice of recent 3D detectors , it has a considerable effect on RangeDet, so we ablate it in Table 2.
5 Comparison with State-of-the-Art Methods
Table 4 shows that RangeDet outperforms other pure range-view-based methods, and is slightly behind the state-of-the-art BEV-based two-stage method. Among all the results, we observe an interesting phenomenon: In contrast to the stereotype that range view is inferior in long-range detection, RangeDet outperforms most other compared methods in the long-range metric (i.e. 50m - inf), especially in the pedestrian class. Unlike in the range view, the pedestrian is very tiny in BEV. This again verifies the superiority of the range view representation and the effectiveness of our remedies to the inconsistency between range view input and 3D Cartesian output space.
6 Runtime Evaluation
On Waymo Open Dataset, our model achieves 12 FPS evaluated on a single 2080Ti GPU without deliberate optimization. Note that our method’s runtime speed is not affected by the expansion of the valid detection distance, while the speed of BEV-based methods will quickly slow down as the maximum detection distance expands.
Conclusion
We present RangeDet, a range-view-based detection framework consisting of Meta-Kernel, Range Conditioned Pyramid, and weighted NMS. With our special designs, RangeDet utilizes the nature of range view to overcome a couple of challenges. RangeDet achieves comparable performance with state-of-the-art multi-view-based detectors.