M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D Object Detection with Transformers
Tianrui Guan, Jun Wang, Shiyi Lan, Rohan Chandra, Zuxuan Wu, Larry Davis, Dinesh Manocha
Introduction
3D object detection is a fundamental problem in computer vision and many applications, including autonomous driving , augmented reality and robotics . Moreover, different methods have been proposed for various sensors, including monocular cameras, depth cameras, LiDAR, and radars . 2D object detection deals with detecting objects from RGB images and videos, while 3D object detection utilizes point cloud-based representations obtained from LiDARs or other sensors. Moreover, it is known that point cloud data obtained from LiDAR sensors tends to be more accurate than RGB images and videos . Consequently, point clouds are being widely used for scene understanding in autonomous driving and AR.
The previous state-of-the-art methods for 3D object detection base on different networks . However, there are two key limitations:
Ineffective point cloud representations: The three major techniques used to process point clouds are based on voxels , raw point clouds , and bird’s-eye-view . Each representation has a unique advantage and it has been shown that combining these representations can result in terms of detection accuracy . However, fusing these representations is non-trivial. First, the architectures corresponding to the VoxelNets, the PointNets, and the 2D convolutional neural networks are different. Moreover, raw point clouds need to be converted to voxels and pixels before techniques based on VoxelNets and 2D convolutional neural networks can be applied. The differences between the inputs of these three neural models can result in semantic gaps. Previous works tend to use feature concatenation and attention modules to fuse multi-representation features. However, the correlation between features of different representations has not been addressed.
Insufficient modeling of multi-scale features: Fusing multi-scale feature maps is a well-known technique used for improving the detection performance of 2D object detection. In terms of 3D object detection, current approaches tend to use multi-scale feature pyramids . However, fusing these multiple feature pyramids is non-trivial because the higher resolution and the larger receptive fields are conflicting . Existing methods fuse the multi-scale features using bi-linear down-sampling/up-sampling and concatenation. Although these approaches can improve the accuracy by a large margin, there are many challenges with respect to the underlying fusion method in terms of the correlation between feature maps of different scales.
A key issue in terms of designing a good approach for object detection is exploiting the correlation between different representations and the large size of the receptive fields. Our approach is motivated by use of transformers , a form of neural network based on attention that has been used in natural language processing . Specifically, transformers use multi-head attention to narrow the semantic gap between different representations by adapting to informative features and eliminating noise.
Another key aspect of 3D object detection is to model the mutual relationships between different points in the point cloud data . Modeling mutual relationships can enhance the ability to recognize the fine-grained patterns and can generalize to complex scenes. Prior works in 3D object detection have modeled these relationships using multi-layer perceptrons , farthest point sampling layers , max pooling layers , and graph convolutional networks . However, a key challenge is to model these mutual relationships along with fusing different representations and multi-scale features.
We present M3DETR, a novel two-stage architecture for 3D object detection task. Given raw 3D point cloud data, our approach can localize static and dynamic obstacles with state-of-the-art accuracy. As illustrated in Figure 1, the key components of our approach are M3 transformers, which are used to combine different feature representations. Conceptually, each of the M3 transformers are used for aggregating point cloud representations, multi-scale representations and mutual relationships among a subset of points in the point cloud data.
M3DETR is the first unified architecture for 3D object detection with transformers that accounts for multi-representation, multi-scale, mutual-relation models of point clouds in an end-to-end manner.
M3DETR is robust and insensitive with respect to the hyper-parameters of transformer architectures. We test multiple variants with different transformer blocks designs. We demonstrate improved performance of M3DETR regardless of hyper-parameters.
Our unified architecture achieves state-of-the-art performance on KITTI 3D Object Detection Dataset and Waymo Open Dataset . We outperform the previous state-of-the-art approaches by 2.86% mAP for car class on the Waymo validation set and 1.48% mAP for all classes on the Waymo test set.
Related Work
Multi-representation modeling. Existing techniques for modeling 3D point cloud data include bird-eye-view (BEV), volumetric, and point-wise representations. Generally, BEV-based approaches first project 3D point clouds into 2D BEV space and then adopt the standard 2D object detectors to generate 3D object proposals from projected 2D feature maps. To deal with the irregular format of input point clouds, voxel-based architectures use equally spaced 3D voxels to encode the point clouds such that the volumetric representation can be consumed by the region proposal network (RPN) . Inspired by the PointNet/PointNet++ approach , which is invariant under transformation, extend this method to the task of 3D object detection and directly process the raw point clouds to infer 3D bounding boxes. However, these methods are typically limited due to either information loss or high computation cost. Recently, many approaches have combined the advantages of speed (of voxel-based representation) and efficiency (of point-based representation) by fusing point-voxel features for 3D object prediction.
Multi-scale modeling. Modeling multi-scale features is an important procedure in deep learning-based computer vision because it is able to enlarge the receptive field and increase resolution. In 3D representation, modeling multi-scale features is also popular and important. PointNet++ proposes the set abstraction module to model local features of a cluster of point clouds. To model multi-scale patterns of point clouds, they use 3 different sampling ranges and radii with 3 parallel PointNets and thus fuse the multi-scale. In 3D object detection, adopt different detection heads with multi-scale feature maps to handle both large and small object classes.
Mutual-relation modeling. 2D Convolutional Neural Networks are commonly used to process mutual relations in 2D images. As point clouds are scattered and lacking structure, passing information from one point to its neighbors is not trivial. PointNets proposes the set abstraction module to model the local context by using the subsampling and grouping layers. After this well-known work, many convolution-like operators on point clouds have been proposed to model the local context and the mutual relation between points. Recently, transformers have been introduced in PointNets to model mutual relation. However, those previous works mainly focus on the local and global contexts of point clouds by applying mutual-relation transformers on points. Instead, our approach not only models the mutual relation between points, but it also models multi-scale and multi-representation features of point clouds.
Transformers in computer vision. Inspired by their success in machine translation , transformer-based architectures have recently become effective and popular in a variety of computer vision tasks. Particularly, the design of self-attention and cross-attention mechanisms in transformers has been successful in modeling dependencies and learning richer information. leverages the direct application of transformers on the image recognition task without using convolution. apply transformers to eliminate the need for many handcrafted components in conventional object detection and achieve impressive detection results. explore transformers on image and video synthesis tasks. investigates the self-attention networks on 3D point cloud segmentation task.
Recently, there have been several joint representation fusion research attempts. PointPainting proposes a novel method that accepts both images and point clouds as inputs to do 3D object detection by appending 2D semantic segmentation labels to LiDAR points. In the visual question answering task, jointly fuses and reasons over three different modality representations. combines multimodal information to solve robust emotion recognition problems. Building on a multi-representation and multi-scale transformer, our proposed model addresses the voxel-wise, point-wise, and BEV-wise feature representation gap and enables effective cross-representation interactions with different levels of semantic features. Coupled with a point-wise mutual-relation transformer, our framework learns to capture deeper local-global structures and richer geometric relationships among point clouds.
Our Approach
M3DETR takes point cloud data as input and generates 3D boxes for different object categories with state-of-the-art accuracy as shown in Figure 2. Our goal is to perform multi-representation, multi-scale, and mutual-relation fusion with transformers over a joint embedding space. Our method consists of three main steps:
Generate feature embeddings for different point cloud representations using VoxelNet, PointNet, and 2D ConvNet.
Fuse these embeddings using M3 transformer that leverages multi-representation and multi-scale feature embedding and models mutual relationships between points.
Perform 3D detection using detection heads network, including RPN and R-CNN stages.
Our network processes the raw input point clouds and encodes them into three different embedding spaces, namely, voxel-, point-, and BEV-based feature representations. We discuss the embedding process for each representation in detail.
2 Multi-representation, Multi-scale, and Mutual-relation Transformer
Once the three feature embedding sequences, , and are generated, they are able to dynamically and intelligently attend to each other, generating final cross-representations, cross-scales, and cross-points descriptive feature representations as shown in Figure 2. In the remainder of this section, we will briefly review transformers basics followed by discussing the multi-representation, multi-scale, and mutual-relation transformer layers.
Now, we present the proposed transformers used to capture the inter- and intra- interactions among input features. Specifically, we propose two stacked transformer encoder layers named M3 Transformers, as shown in Figure 3: (1) the multi-representation and multi-scale transformer, and (2) the mutual-relation transformer.
The multi-representation and multi-scale transformer layer takes 6 different inputs: , , , , , . As different inputs may have different feature dimensions, we use the single-layer perceptron to apply feature reduction on the input features to align the feature dimensions of each feature embeddings. The outputs of the feature reduction layer are , , , , , , where the output dimension of each feature is equivalent to .
After the feature reduction layer, the multi-representation and multi-scale transformer layer takes as inputs and generates self-attention features , ,,, , , which corresponds to , ,,, , . We visualize the multi-representation and multi-scale transformer in Figure 3.
Mutual-relation transformer layer. Inspired by , inter-points feature fusion within a spatial neighboring space is leveraged in the mutual-relation layer of the transformer. Our goal is to attend to and aggregate the neighboring information in an attention manner for each point with an enriched feature.
The mutual-relation transformer takes point-wise features of keypoints as inputs and uses the multi-head self-attention head to model the mutual relationship between keypoints. The outputs of the mutual-relation transformer are .
Comparison with previous point-based transformers Prior work has explored the transformer application on the task of point cloud processing. First, leverage the inherent permutation invariance of transformers to capture local context within the point cloud on the shape classification and segmentation tasks, while our M3DETR mainly investigates the strong attention ability of transformers between input embeddings for the 3D object detection task. Pointformer is the approach most related to our method because we both address the 3D object detection task by capturing the dependencies among points’ features. However, Pointformer adopts a single PointNet branch to extract points feature, while M3DETR considers all three different representations and also applies the transformer to learn aggregated representation-based features.
3 Detection Heads Network
After we obtain the enriched embedding from the M3 transformer, the detection network is composed of two stages that predict 3D bounding box class, localization, and orientation in a coarse-to-fine manner, including RPN and R-CNN. Please refer to PV-RCNN for more details.
RPN: Region Proposal Networks take the deep semantic features produced by the 2D ConvNets as inputs and generate high-quality 3D object proposals , as shown in Figure 2. A 3D object box is parameterized as , where is the center of the box, is the dimension of the box, and is the orientation in bird’s-eye-view. Similar to the conventional RPN in 2D object detection each position on the deep feature map is placed by predefined bounding box anchors denoted as (, , , , , , ). Then initial proposals are generated by predicting the relative offset of an object’s 3D ground truth bounding box (, , , , , , ).
R-CNN: R-CNN serves as the second stage and it takes the initial region proposals from RPN as input to conduct further proposal refinement. For each input proposal box, the RoI-grid pooling module is adopted to extract the corresponding proposal-specific grid points’ features from the transformer-based embeddings . Compared with previous works, M3DETR leverages the richer embedding information from the learned transformers for the fine-grained proposal refinement. As the main component to extract refined features, the RoI-grid pooling module uniformly samples grid points per 3D proposal. For each grid point, the output feature is generated by applying a PointNet-block on a small number of surrounding keypoints, , within its spatial surrounding region with a radius of . Specifically, keypoints are the subset of input points that are sampled using the Furthest-point-sampling algorithm to cover the entire point set.
Finally, the refined representations of each proposal are first passed to two fully connected layers and then generate the box 3D Intersection-over-Union (IoU) guided confidence scoring of class prediction and location refinement of regression targets, . Compared with the traditional classification-guided box scoring, 3D IoU guided confidence scoring considers the IoU between the proposal box and its corresponding ground truth box. Empirically, show that it achieves better results compared with the traditional classification confidence based techniques.
4 Loss Functions
In this section, we define the loss function. The bounding box regression target for both RPN and R-CNN stages is calculated as the relative offsets between the anchors and the ground truth as: , , , , where = . Similar to , the focal loss is applied for the classification loss, . Smooth L1 loss is adopted for the box localization regression target’s losses, and . In addition, 3D IoU loss is used for .
Similar to PV-RCNN , we formally define a multi-task loss for both the RPN and R-CNN stages,
where is the model’s estimated class probability for an anchor box, and , and are chosen to balance the weights between classification loss, IoU loss and regression loss for RPN stage and R-CNN stage. We adopt the default , and from the parameters of focal loss .
Experiments
In this section, we evaluate M3DETR both qualitatively and quantitatively on the Waymo Open Dataset and the KITTI Dataset in the task of LiDAR-based 3D object detection. Our main results include achieving a state-of-the-art accuracy on these datasets and robustness to hyper-parameter tuning.
Waymo Open Dataset: The Waymo Open Dataset is a large-scale autonomous driving dataset containing scenes of s duration each, with scenes for training and scenes for validation. Each scene is sampled at a frequency of Hz. Overall, the dataset includes M labeled objects and thus we only use one fifth of the training scenes for the following experiment. We consider LiDAR data as the input to our approach. The evaluation protocol on the Waymo dataset consists of the mean average precision (mAP) and mean average precision weighted by heading (mAPH). For each object category, the detection outcomes are evaluated based on two difficulty levels: LEVEL_1 denotes the annotated bounding box with more than 5 points and LEVEL_2 represents the annotated bounding box with more than 1 point.
KITTI dataset The KITTI 3D object detection benchmark is another popular dataset for autonomous driving. It contains training and testing LiDAR scans. We follow the standard split on the training ( samples) and validation sets ( samples). For each object category, the detection outcomes are evaluated based on three difficulty levels based on the object size, occlusion state, and truncation level.
2 Implementation Details
Closely following the codebasehttps://github.com/open-mmlab/OpenPCDet., we use PyTorch to implement our M3 transformer modules and integrate them into the PV-RCNN network . More details about backbone and detection heads network will be introduced in the supplementary materials.
M3 Transformer We project the embeddings obtained from the backbone network with different scales and representations to 256 channels, as the input of the multi-representation and multi-scale transformer requires, and project the output features back to their original dimensions before passing into the mutual-relation transformer. Due to the GPU memory constraint, we experiment with two types of MHSA module designs: 2 encoder layers with 4 attention heads and 1 encoder layer with 8 attention heads.
Training Parameters Models are trained from scratch on NVIDIA P6000 GPUs. We use the Adam optimizer with a fixed weight decay of and use a one-cycle scheduler proposed in . For the Waymo Open Dataset, we train our models for epochs with a batch size of scenes and a learning rate , which takes around hours. For the KITTI dataset, we train our models for epochs with a batch size of scenes per and a learning rate , which takes around hours.
3 Results
Waymo Open Dataset: We first present our object detection results for the vehicle, pedestrian, and cyclist classes on the test set of Waymo Open Dataset in Table 1 compared with PV-RCNN . We evaluate our method at both LEVEL_1 and LEVEL_2 difficulty levels. Note that we reproduce the baseline PV-RCNN with a single frame input since they recently adopt two-frames input on test set. As we can see, PV-RCNN achieves 69.57% and 63.65% on average in LEVEL_1 mAP and LEVEL_2 mAP, respectively, while M3DETR improves them by 1.48% and 1.85%, respectively. Without bells and whistles, our approach works better than PV-RCNN . Furthermore, we compare our framework on the vehicle class for different distances with state-of-the-art methods, including StarNet , PointPillars , RCD , Det3D , RangeDet and PV-RCNN . In Table 2, M3DETR outperforms PV-RCNN significantly in both LEVEL_1 and LEVEL_2 difficulty levels across all distances, demonstrating the effectiveness of newly proposed framework. Moreover, we visualize the detection results of M3DETR in Figure 4.
Compared with the PV-RCNN shown in Figure 5, M3DETR successfully captures the inter- and intra- interactions among input features and effectively helps the model generate high-quality box proposals. To the best of our knowledge, M3DETR achieves the state-of-the-art in the Vehicle class in both LEVEL_1 and LEVEL_2 difficulty levels among all the published papers with a single frame LiDAR input.
We also evaluate the overall object detection performance with an IoU of 0.7 for Vehicle class on the full Waymo Open Dataset validation set as in Table 3, further proving that our architecture is more efficient for jointly modeling the input features.
KITTI dataset: We compare our approach with the state-of-the-art methods on the KITTI test set . We compute the mAP on three difficult types of both car and cyclist classes in 3D detection metric. Table 4 shows that M3DETR achieves state-of-the-art performance and outperforms the previous work by a large margin especially on the cyclist class. In particular, HotSpotNet achieves in the “easy” categories of 3D detection metric, while M3DETR improves these results by significant .
4 Ablation Studies
To demonstrate the individual benefits of the multi-representation, multi-scale, and mutual-relation layers of the M3 transformer, we perform ablation experiments and tabulate the results in Table 5. All experiments are conducted on the validation set of KITTI dataset.
With the single multi-representation and multi-scale transformer layer, we can achieve 4.07% and 1.71% on the moderate difficulty in car class with 11 and 40 recall positions, respectively compared with the PV-RCNN baseline. On the other side, with the single mutual-relation transformer layer, the performance gain are 4.47% and 1.99% compared with the PV-RCNN baseline. Without hyper-parameter tuning, M3DETR benefits from unifying multiple point cloud representations, feature scales, and model mutual-relations simultaneously which results in the best performance.
5 Robustness of M3DETR
To demonstrate the robustness of M3DETR to hyper-parameter tuning, we perform a series of tests by varying the sampling size, number of detection heads, and number of transformer encoder layers. We present the results of these tests in Figure 6, where we observe that M3DETR performs consistently well for the “car” category with IoU threshold of 0.7 for both 11 and 40 recall positions on the KITTI validation set.
Conclusions
In this paper, we present M3DETR, a novel transformer-based framework for object detection with LiDAR point clouds. M3DETR is designed to simultaneously model multi-representation, multi-scale, mutual-relation features through the proposed M3 Transformers. Overall, the first transformer integrates features with different scales and representations, and the second transformer aggregates information from all keypoints. Experimental results show that M3DETR outperforms previous work by a large margin on the Waymo Open Dataset and the KITTI dataset. Without bells and whistles, M3DETR is demonstrated to be invariant to the hyper-parameters of transformer.
Acknowledgement. This work was supported in part by ARO Grants W911NF1910069, W911NF2110026 and U.S. Army Grant No. W911NF2120076.
References
Appendix A More Our Approach Details
We further discuss our approach in the following.
For the voxel-wise feature extraction from raw point clouds input, there are two steps, voxelization using voxelization layer and feature extraction using 3D sparse convolutions. We denote the size of each discretized voxel as , where indicate the length, width, and height of the voxel grid and represents the channel of the voxel features. We adopt the average of the point-wise features from all the points to represent the whole non-empty voxel feature. After voxelization, the input feature is propagated through a series of sparse cubes, including four consecutive blocks of 3D sparse convolution with downsampled sizes of , , , , using convolution operations of stride 2. Specifically, each sparse convolutional block includes a 3D convolution layer followed by a LayerNorm layer and a ReLU layer.
A.2 Multi-head Self-attention Basics
Building on the attention mechanism, Multi-head Self-attention (MHSA) with heads and the input matrix is defined as follows:
Appendix B More Implementation Details
As mentioned in the Section 4.2, we give more details on our implementation for reproduction of our result. Our code will also be released later, including trained models that can match the performance that was included in the paper.
The 3D voxel CNN branch consists of 4 blocks of 3D sparse convolutions with output feature channel dimensions of 16, 32, 64, 64. Those 4 different voxel representations of different scales, as well as point features from point cloud input, are used to refine keypoint features by PointNet through set abstraction and voxel set abstraction . The number of sampled keypoints is 2,048 for both Waymo and KITTI. In order to sample the keypoints effectively and accurately, we uses FPS on points within the range of top 1000 or 1500 initial proposals with a radius of 2.4.
Waymo: The voxel size is [0.1, 0.1, 0.15], and we focus on the input LiDAR point cloud range with [-75.2, 75.2], [-75.2, 75.2], and meters in x, y, and z axis, respectively. The 2D ConvNets output size is .
KITTI: The voxel size is [0.05, 0.05, 0.1], and we focus on the input LiDAR point cloud range with [0, 70.4], , and meters in x, y, and z axis, respectively. The 2D ConvNets output size is .
B.2 Detection Heads
The RPN anchor size for each object category is set by computing the average of the corresponding objects from the annotated training set. The RoI-grid pooling module samples grid points within each initial 3D proposal to form a refined features. The number of surrounding points used to extract the grid point’s feature, M, is 16.
During the training phase, 512 proposals are generated from RPN to R-CNN, where non-maximum suppression (NMS) with a threshold of 0.8 is applied to remove the overlapping proposals. In the validation phase, 100 proposals are fed into R-CNN, where the NMS threshold is 0.7.