3D Object Detection with Pointformer

Xuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li, Gao Huang

Introduction

3D object detection in point clouds is essential for many real-world applications such as autonomous driving and augmented reality . Compared to images, 3D point clouds can provide detailed geometry and capture 3D structure of the scene. On the other hand, point clouds are irregular, which can not be processed by powerful deep learning models, such as convolutional neural networks directly. This poses a big challenge for effective feature learning.

The common feature processing methods in 3D detection can be roughly categorized into three types, based on the form of point cloud representations. Voxel-based approaches gridify the irregular point clouds into regular voxels and are followed by sparse 3D convolutions to learn high dimensional features. Though effective, voxel-based approaches face the dilemma between efficiency and accuracy. Specifically, using smaller voxels gains more precision, but suffers from higher computational cost. Conversely, using larger voxels misses potential local details in the crowded voxels.

Alternatively, point-based approaches , inspired by the success of PointNet and its variants, consume raw points directly to learn 3D representations, which mitigates the drawback of converting point clouds to some regular structures. Leveraging learning techniques for point sets, point-based approaches avoid voxelization-induced information loss and take advantage of the sparsity in point clouds by only computing on valid data points. Nevertheless, due to the irregularity of point cloud data, point-based learning operations have to be permutation-invariant and adaptive to the input size. To achieve this, it learns simple symmetric functions (e.g. using point-wise feedforward networks with pooling functions) which highly restricts its representation power.

Hybrid approaches attempt to combine both voxel-based and point-based representations. leverages PointNet features at the voxel level and a column of voxels (pillar) level respectively. deeply integrate voxel features and PointNet features at the scene level. However, the fundamental difference between the two representations could pose a limit on the effectiveness of these approaches for 3D point-cloud feature learning.

To address the above limitations, we resort to the Transformer models, which have achieved great success in the field of natural language processing. Transformer models are very effective at learning context-dependent representations and capturing long range dependencies in the input sequence. Transformer and the associate self-attention mechanism not only meet the demand of permutation invariance, but also are proved to be highly expressive. Specifically, proves that self-attention is at least as expressive as convolution. Currently, self-attention has been successfully applied to classification and 2D object detection in computer vision. However, the straightforward application of Transformer to 3D point clouds is prohibitively expensive because computation cost grows quadratically with the input size.

To this end, we propose Pointformer, a backbone for 3D point clouds to learn features more effectively by leveraging the superiority of the Transformer models on set-structured data. As shown in Figure 2, Pointformer is a U-Net structure with multi-scale Pointformer blocks. A Pointformer block consists of Transformer-based modules that are both expressive and friendly to the 3D object detection task. First, a Local Transformer (LT) module is employed to model interactions among points in the local region, which learns context-dependent region features at an object level. Second, a coordinate refinement module is proposed to adjust centroids sampled from Furthest Point Sampling (FPS) which improves the quality of generated object proposals. Third, we propose Local-Global Transformer (LGT) to integrate local features with global features from higher resolution. Finally, Global Transformer (GT) module is designed to learn context-aware representations at the scene level. As illustrated in Figure 1, Pointformer can capture both local and global dependencies, thus boosting the performance of feature learning for scenes with multiple cluttered objects.

Extensive experiments have been conducted on several detection benchmarks to verify the effectiveness of our approach. We use the proposed Pointformer as the backbone for three object detection models, CBGS , VoteNet , and PointRCNN , and conduct experiments on three indoor and outdoor datasets, SUN-RGBD , KITTI , and nuScenes respectively. We observe significant improvements over the original models on all experiment settings, which demonstrates the effectiveness of our method.

In summary, we make the following contributions:

We propose a pure transformer model, Pointformer, which serves as a highly effective feature learning backbone for 3D point clouds. Pointformer is permutation invariant, local and global context-aware.

We show that Pointformer can be easily applied as the drop-in replacement backbone for state-of-the-art 3D object detectors for the point cloud.

We perform extensive experiments using Pointformer as the backbone for three state-of-the-art 3D object detectors, and show significant performance gains on several benchmarks including both indoor and outdoor datasets. This demonstrates that the versatility of Pointformer as 3D object detectors are typically designed and optimized for either indoor or outdoor only.

Related Work

Feature learning for 3D point clouds. Prior work includes feature learning on voxelized grids, direct feature learning on point clouds and the hybrid of the two. 3D sparse convolution is very effective on voxel grids. For direct feature learning, PointNet and PointNet++ learn point-wise features and region features using feed-forward networks and simple symmetric functions (e.g. max) respectively. PCCN generalizes convolution to non-grid structured data by exploiting parameterized kernel functions that span the full continuous vector space. EdgeConv exchanges local neighborhood information and acts on graphs dynamically computed in each layer of the network. Hybrid methods combine both types of features at the local level or at the network level .

Transformers in computer vision. Image GPT is the first to adopt the Transformers in 2D image classification task for unsupervised pretraining. Further, ViT extends this scheme to large scale supervised learning on images. For high level vision tasks, DETR and Deformable DETR leverage the advantages of Transformers in 2D object detection. Set Transformer uses attention mechanisms to model interactions among elements in the input set. In the field of 3D vision, PAT designs novel group shuffle attentions to capture long range dependencies in point clouds. To the best of our knowledge, we are the first to propose a pure Transformer model for 3D points clouds feature learning with carefully designed Transformer blocks and a positional encoding module to capture geometric and rich context information.

3D object detection in point clouds. Detectors are designed either with point clouds as the only input or fusing multiple sensor modalities such as LiDAR and camera . Their backbones are designed with the aforementioned feature learning approaches. We focus on point cloud only object detection. In this category, VoxelNet divides the point cloud into voxels, followed by 3D convolutions to extract features. VoteNet devises a novel 3D proposal mechanism using deep Hough voting, before H3DNet makes further investigations on geometric primitives. In addition, MLCVNet focuses more on contextual information aggregation based on VoteNet, and PointGNN exploits graph learning methods in point cloud detection. We show that our novel Transformer based model, Pointformer, can be used as a drop-in replacement for voxel-based detector, CBGS and point-based detectors, VoteNet and PointRCNN .

Pointformer

Feature learning for 3D point clouds needs to confront its irregular and unordered nature as well as its varying size. Prior work utilizes simple symmetric functions, e.g., point-wise feedforward networks with pooling functions , or resorts to the techniques in graph neural networks by aggregating information from the local neighborhood . However, the former is not effective in incorporating local context-dependent features beyond the capability of the simple symmetric functions; the latter focuses on the message passing between the center point and its neighbors while neglecting the feature correlations among the neighbor points. Additionally, global representations are also informative but rarely used in 3D object detection tasks.

In this paper, we design Transformer-based modules for point set operations which not only increase the expressiveness of extracting local features, but incorporate global information into point representations as well. As shown in Figure 2, a Pointformer block mainly consists of three parts: Local Transformer (LT), Local-Global Transformer (LGT) and Global Transformer (GT). For each block, LT first receives the output from its previous block (high resolution) and extracts features for a new set with fewer elements (low resolution). Then, LGT uses the multi-scale cross-attention mechanism to integrate features from both resolutions. Lastly, GT is adopted to capture context-aware representations. As for the up-sampling block, we follow PointNet++ and adopt the feature propagation module for its simplicity.

We first revisit the general formulation of the Transformer model. Let F ⁣= ⁣{fi}F\!=\!\{f_{i}\} and X ⁣= ⁣{xi}X\!=\!\{x_{i}\} denote a set of input features and their positions, where fif_{i} and xix_{i} represent the feature and position of token ii, respectively. Then, a Transformer block comprises of a multi-head self-attention module and feedforward network:

where Wq,Wk,WvW_{q},W_{k},W_{v} are projections for query, key and value. mm is the index of MM attention heads and dd is the feature dimension. PE(⋅){\rm PE(\cdot)} is the positional encoding function for input positions, and FFN(⋅){\rm FFN(\cdot)} represents a position-wise feed-forward network. σ(⋅)\sigma(\cdot) is a normalization function and SoftMax is mostly adopted.

In the following sections, for simplicity, we use

to represent the basic Transformer block (Eq.(1)∼\sim Eq.(4)). Readers can refer to for further details.

2 Local Transformer

where F={fi∣i∈N(xct)}F=\{f_{i}|i\in\mathcal{N}(x_{c_{t}})\} and X={xi∣i∈N(xct)}X=\{x_{i}|i\in\mathcal{N}(x_{c_{t}})\} denote the set of features and coordinates in the local region with centroid xctx_{c_{t}}.

Compared to the existing local feature extraction modules in , the proposed Local Transformer has several advantages. First, the dense self-attention operation in the Transformer block greatly enhances its expressiveness. Several graph learning based approaches can be approximated as special cases of the LT module with learned parameter space carefully designed. For instance, a generalized graph feature learning function can be formulated as:

where most of the models utilize summation as the aggregation function A\mathcal{A} and the operation ⊕\oplus is chosen from {Concatenation, Plus, Inner-product}. Therefore, the edge function eije_{ij} is at most a quadratic function of {xi,xj,fi,fj}\{x_{i},x_{j},f_{i},f_{j}\}. For a one-layer Transformer block, the learning module can be formulated with the inner-product self-attention mechanism as follows:

where dd is the feature dimension of fif_{i} and fjf_{j}. We can observe that the edge function is also a quadratic function of {xi,xj,fi,fj}\{x_{i},x_{j},f_{i},f_{j}\}. With sufficient number of layers in FFN{\rm FFN}s, the graph-based feature learning module has the same expressive power as a one-layer Transformer encoder. When it comes to Pointformer, as we stack more Transformer layers in the block, the expressiveness of our module is further increased and can extract better representations.

Moreover, feature correlations among the neighbor points are also considered, which are commonly omitted in other models. Under some circumstances, neighbor points can be even more informative than the centroid point. Therefore, by leveraging message passing among all points, features in the local region are equally considered, which makes the local feature extraction module more effective.

3 Coordinate Refinement

Furthest point sampling (FPS) is widely used in many point cloud frameworks, as it can generate a relatively uniform sampled points while keeping the original shape, which ensures that a large fraction of the points can be covered with limited centroids. However, there are two main issues in FPS: (1) It is notoriously sensitive to the outlier points, leading to highly instability especially when dealing with real-world point clouds. (2) Sampled points from FPS are a subset of original point clouds, which makes it challenging to infer the original geometric information in the cases that objects are partially occluded or not enough points of an object are captured. Considering that points are mostly captured on the surface of objects, the second issue may become more critical as the proposals are generated from sampled points, resulting in a natural gap between the proposal and ground truth.

To overcome the aforementioned drawbacks, we propose a point coordinate refinement module with the help of the self-attention maps. As shown in Figure 3, we first take out the self-attention map of the last layer of the Transformer block for each attention head. Then, we compute the average of the attention maps and utilize the particular row for the centroid point as a weight vector:

where MM represents the number of attention heads and A(m)A^{(m)} is the attention map for the mthm_{\rm th} head. Lastly, the refined centroid coordinates are computed as weighted average of all points in the local region:

where wkw_{k} is the kthk_{\rm th} entry of WW. With the proposed coordinate refinement module, centroid points are adaptively moving closer to object centers. Moreover, by utilizing the self-attention map, our module introduces little computational cost and no additional learning parameters, making the refinement process more efficient.

4 Global Transformer

Global information representing scene contexts and feature correlations between different objects is also valuable in the detection tasks. Prior work using PointNet++ or sparse 3D convolution to extract high level features for 3D point clouds enlarges the receptive field as the depth of their networks increases. However, this has limitations on modeling long-range interactions.

As a remedy, we leverage the power of Transformer modules on modeling non-local relations and propose a Global Transformer to achieve message passing through the whole point cloud. Specifically, all points are gathered to a single group P\mathcal{P} and serves as input to a Transformer module. The formulation for GT is summarized as follows:

By leveraging the Transformer on the scene level, we can capture the context-aware representations and promote message passing among different objects. Moreover, global representations can be particularly helpful for detecting objects with very few points.

5 Local-Global Transformer

Local-Global Transformer is also a key module to combine the local and global features extracted by the LT and GT modules. As shown in Figure 2, the LGT adopts a multi-scale cross-attention module and generates relations between low resolution centroids and high resolution points. Formally, we apply cross attention similar to the encoder-decoder attention used in Transformer. The output of LT serves as query and the output of GT from the higher resolution is used as key and value. With the LL-layer Transformer block, the module is formulated as:

where Pl\mathcal{P}^{l} (keypoints, the output of LT in Figure 2) and Ph\mathcal{P}^{h} (the input of a Pointformer block in Figure 2) represent subsamples of point cloud P\mathcal{P} from low and high resolution respectively. Through the Local-Global Transformer module, we utilize whole centroid points to integrate global information via an attention mechanism, which makes the feature learning of both more effective.

6 Positional Encoding

Positional encoding is an integral part of Transformer models as it is the only mechanism that encodes position information for each token in the input sequence. When adapting Transformers for 3D point cloud data, positional encoding plays a more critical role as the coordinates of point clouds are valuable features indicating the local structures. Compared to the techniques used in natural language processing, we propose a simple and yet efficient approach. For all Transformer modules, coordinates of each input point are firstly mapped to the feature dimension. Then, we subtract the coordinates of the query and key points and use relative positions for encoding. The encoding function is formalized as:

7 Computational Cost Reduction

Since Pointformer is a pure attention model based on Transformer blocks, it suffers from extremely heavy computational overhead. Applying a conventional Transformer to a point cloud with nn points consumes O(n2)O(n^{2}) time and memory, leading to much more training cost.

Some recent advances in efficient Transformers have mitigated this issue , among which Linformer reduces the complexity to O(n)O(n) by low-rank factorization of the original attention. Under the hypothesis that the self attention mechanism is low rank, i.e. the rank of the n×nn\times{}n attention matrix

is much smaller than nn, Linformer projects the nn-dimension keys and values to the ones with lower dimension k≪nk\ll{}n, and kk is closer to the rank of AA. Therefore, the ii-th head in the projected multi-head self-attention is

Compared with the Taylor expansion approximation technique used in MLCVNet , Linformer is easier to implement in out method. We thus adopt it to replace the Transformer layers in the vanilla Pointformer. Practically, we map the number of points nn to k=nrk=\frac{n}{r}, where rr is a factor controlling the number of projected dimensions. We apply this mapping in Local Transformer, Global Transformer and Local-Global Transformer blocks. By setting an appropriate factor rr for each block, there would be a significant boost in both time and space consumption with little performance decay.

Experimental Results

In this section, we use Pointformer as the backbone for state-of-the-art object detection models and conduct experiments on several indoor and outdoor benchmarks. In Sec. 4.1, we introduce the implementation details of the experiments. In Sec. 4.2 and Sec. 4.3, we show the comparison results on indoor and outdoor datasets respectively. In Sec. 4.4, we conduct extensive ablation studies to analyze our proposed Pointformer model. Finally, we show qualitative results in Sec. 4.5. More analysis and visualizations are provided in the appendix.

Datasets. We adopt SUN RGB-D and ScanNet V2 for indoor 3D detection benchmark. SUN RGB-D has 5K training images annotated with oriented 3D bounding boxes for 37 object categories and ScanNet V2 has 1513 labeled scenes with 40 semantic classes. We follow the same setting in VoteNet and report performance on the 10 classes on SUN RGB-D and 18 classes on ScanNet V2. For outdoor datasets, we choose KITTI and nuScenes for evaluation. KITTI contains 7,481 training samples and 7,518 test samples for autonomous driving. NuScenes contains 1k different scenes with 40K key frames, which has 23 categories and 8 attributes. We follow the evaluation protocol proposed along with the datasets.

Experimental setups. We use the Pointformer as the backbone for three 3D detection models, including VoteNet , PointRCNN and CBGS . VoteNet is a point-based approach for indoor datasets, while PointRCNN and CBGS are adopted for outdoor datasets. PointRCNN is a classic approach for autonomous driving detection and CBGS is the champion of nuScenes 3D detection Challenge held in CVPR 2019. For a fair comparison, we adopt the same detection head, number of points for each resolution, hyperparameters and training configurations as baseline models.

2 Outdoor Datasets

KITTI. We first evaluate our method comparing with PointRCNN on KITTI’s 3D detection benchmark. PointRCNN uses PointNet++ as its backbone with four set abstraction layers. Similarly, we adopt the same architecture, while switching the set abstraction layer in PointNet++ with the proposed Transformer block. The comparison results on the KITTI test server are shown in Table 1.

For the car category, we also report the performance of 3D detection results on the val split as shown in Table 3. As we can observe, by adopting Pointformer, our model achieves consistent improvements comparing to the original PointRCNN. Especially in the hard difficulty, our method shows the most promising result with 1.5% AP improvement. We believe the better performance on hard objects is attributed to the higher expressiveness of local Transformer module. For hard objects which are often small or occluded, GT captures context-dependent region features, which contributes to the bounding box regression and classification.

Additionally, we evaluate the performance of proposal generation network by calculating the recall of 3D bounding box with various number of proposals and 3D IoU threshold. As shown in Table 4, our backbone module significantly enhances the performance of proposal generation network under almost all the settings. Analyzing the figures vertically, we observe that our backbone shows better performance when the number of RoIs are relatively small. As stated in Sec.3, the GT and LGT help to capture context-aware representations and models the relations among different objects (proposals). This provides additional references for locating and reasoning the bounding boxes. Therefore, despite the lack of RoIs, we can still improve the performance of the proposal generation module and achieve higher recall.

NuScenes. We also validate the effectiveness of Pointformer on the nuScenes dataset, which greatly extends KITTI in dataset size, number of object categories and number of annotated objects. Furthermore, nuScenes suffers from severe class imbalance issues, making the detection task more difficult and challenging. In this part, we adopt CBGS, the champion of nuScenes 3D detection Challenge held in CVPR 2019, as the baseline model and show the comparison results when replacing the backbone with Pointformer. We summarize the results in Table 2. As we can observe, by utilizing Pointformer as the backbone, our model achieves 0.8 higher mAP than baseline. For 8 of 10 classes, our model shows better performance, which demonstrates the effectiveness of Pointformer on larger and more challenging datasets.

3 Indoor Datasets

We evaluate our Pointformer accompanied by VoteNet on SUN RGB-D and ScanNet V2. We follow the same hyperparameters on the backbone structure as VoteNet. Followed by the Pointformer blocks, two feature propagation(FP) modules proposed in PointNet++ serve as upsamplers to increase the resolution for the subsequent detection heads.

SUN RGB-D. We report the average precision(AP) over 10 common classes in SUN RGB-D, as shown in Table 5. Compared with the PointNet++ in VoteNet , our Pointformer provides a significant boost with 2%\% mAP over the implementation in MMDetection3D . On some categories with large and complex objects like dresser or bathtub, Pointformer shows its splendid capability on extracting non-local information by a sharp increase over 5%\% AP, which we attribute to the GT module in Pointformer.

ScanNet V2. We report the average precision(AP) over 18 classes in ScanNet V2, as shown in Table 6. Compared with VoteNet, Pointformer outperforms its original version by 1.2%\% mAP with MMDetection3D.

4 Ablation Study

In this section, we conduct extensive ablation experiments to analyze the effectiveness of different components of Pointformer. All experiments are trained on the train split with PointRCNN detection head and evaluated on the val split with the car class.

Effects of each component. We validate the effectiveness of each Transformer component and the coordinate refinement module, and summarized the results in Table 7. The first row corresponds to the PointRCNN baseline and the last row is the full Pointformer model. By comparing the first row and second row, we can observe that easy objects benefit more from the local Transformer with 0.60.6 AP improvement. By comparing the second row and fourth row, we can see that global Transformer is more suitable for hard objects with 0.90.9 AP improvement. This observation is consistent with our analysis in Sec. 4.2. As for Local-Global Transformer and coordinate refinement, the improvement is similar under three difficulty settings.

Positional Encoding. Playing a critical role in Transformer, position encoding can have huge impact on the learned representation. As we have shown in Table 8, we compare the performance of Pointformer without positional encoding and with two approaches to position encoding (adding or concatenating positional encoding with the attention map). We can observe that Pointformer without positional encoding suffers from a huge performance drop, as the coordinates of points can capture the local geometric information.

5 Qualitative Results and Discussion

Qualitative results on SUN RGB-D. Figure 4 shows representative examples of detection results on SUN RGB-D with VoteNet + Pointformer. As we can observe, our model achieves robust results despite the challenges of clutter and scanning artifacts. Additionally, our model can even recognize the missing objects in the ground truth. For instance, the dresser in the left scene is only partially observed by the sensor. However, our model can still generate precise proposals for the object with proper bounding box sizes. Similar results are shown in the right scene, where the table in the front suffers from clutter because of the books on it.

Inspecting Pointformer with attention maps. To validate how modules in Pointformer affect learned point features, we visualize the attention maps from the GT module of the second last Pointformer block. We show the attention of the particular points in Figure 5. The second row shows the 50, 100, 200 points with highest attention values towards the points marked with star. We can observe that Pointformer first focuses on the local region of the same object, then spread the attention to other regions, and finally attends points from other objects globally. The overall attention map shows the average attention weights of all the points in the scene, indicating that our model mostly focuses on points on the objects. These visualization results show that Pointformer can capture local and global dependencies, and enhance message passing on both object and scene levels.

Conclusion

This paper introduces Pointformer, a highly effective feature learning backbone for 3D point clouds that is permutation invariant to points in the input and learns local and global context-aware representations. We apply Pointformer as the drop-in replacement backbone for state-of-the-art 3D object detectors and show significant performance improvements on several benchmarks including both indoor and outdoor datasets.

Comparing to classification and segmentation tasks including part-segmentation and semantic segmentation in prior work, 3D object detection typically involves more points (4×\times - 16×\times) in a scene, which makes it harder for Transformer-based models. For future work, we would like to explore extensions to these two tasks and other 3D tasks such as shape completion, normal estimation, etc.

Acknowledgments

This work is supported in part by the National Science and Technology Major Project of the Ministry of Science and Technology of China under Grants 2018AAA0100701, the National Natural Science Foundation of China under Grants 61906106 and 62022048, the Institute for Guo Qiang of Tsinghua University and Beijing Academy of Artificial Intelligence.

References

Appendix

Appendix A Architectures and Implementation Details

In this section, we discuss each component in our Pointformer models for indoor and outdoor settings in detail.

Indoor Datasets. First, the Local Transformer(LT) block is composed of a sequence of sampling and grouping operations, followed by a shared positional encoding layer and two self-attention transformer layers, with a linear shared Feed-Forward Network(FFN) in the end. As shown in Table 9, we use the same sampling and grouping parameters, and feature dimensions as those of PointNet++ , the backbone in VoteNet and H3DNet .

Second, the Local-Global Transformer(LGT) and the Global Transformer(LT) have fewer hyper-parameters than LT, where we adopt two self-attention layers in GT and one cross-attention layer in LGT for each Pointformer block. Since their massive attention computation may lead to overfitting, we apply dropout with the dropping probability 0.4 on SUN RGB-D and 0.2 on ScanNetV2 . As for the number of heads in multi-head attention, we set it to 8 on ScanNetV2 and 4 on SUN RGB-D. In our experiments, we found that noisy backgrounds in indoor datasets affect the LGT performance, by reducing 1∼\sim2%\%mAP. So we report the results on indoor datasets without the LGT module.

Finally, we implement our indoor models on the top of MMDetection3D, an open source toolbox 3D object detection. We follow the same hyper-parameters and data augmentation techniques as those of VoteNet. To train a Pointformer on SUN RGB-D, we use the AdamW optimizer with an initial learning rate of 3e-4 and weight decay factor of 0.05, and decay the learning rate by 0.3 at epoch 24 and 32 during the training of a total of 36 epochs. And for ScanNet, we use AdamW optimizer with 0.002 learning rate and set 0.1 weight decay. We decay the learning rate by 0.3 at epoch 32 and 40 during the training of 48 epochs.

Outdoor Datasets. We adopt the same structure of Transformer blocks as that for indoor datasets. Two self-attention layers with FFN are adopted in LT and GT, while only one cross-attention layer is utilized in LGT block. The number of heads are set to 8 for both KITTI and nuScenes .

We implement our outdoor models on top of OpenPCDet , an open source toolbox for LiDAR-based 3D object detection. We follow the same hyper-parameters as that of PointRCNN, including data augmentation, post-processing, etc. To train a Pointformer on KITTI, we use the Adam optimizer with an initial learning rate of 5e-3 and weight decay of 0.01.

Appendix B More Quantitative Results

In this section, we provide more results and analysis on SUN RGB-D and ScanNetV2 as shown in Table 11&12. With 0.5 IoU threshold, our proposed Pointformer achieves consistent improvements on both dataset. In Table 13, we use the one tower version H3DNet as baseline, showing our method can work well with the recent advanced model.

Appendix C More Ablation Studies

Parameter Efficiency. To further validate the effectiveness of Pointformer, we conduct experiments and compare the backbones with similar model parameters. We reduce the Transformer layers adopted in each block and refer the model as Pointformer(small). Similarly, we increase the FFN layers in PointNet++ and refer the model as PointNet++(large). As we have shown in Table 14, Pointformer achieves better results under both parameter budgets. Although our model suffers from a performance reduction when using fewer Transformer layers, we are still 0.5% to 1% AP higher for all difficulty levels. Additionally, PointNet++ shows little improvement with larger feature dimensions. By comparison, Pointformer can adapt to deeper models and use learning parameters more efficiently.

Computational Cost Reduction. As stated in Section 3.7, Transformer-based modules suffer from heavy computational cost and memory consumption. Therefore, we adopt the Linformer technique to improve model efficiency. The results are shown in Table 15 and we can observe that inference latency is decreased with little drop in performance.

Appendix D More Qualitative Results

We provide additional visualization results in this section. Figure 6 shows more visualized attention maps on SUN RGB-D dataset. Figure 7 and Figure 8 present qualitative results of detection models with Pointformer on ScanNetV2 and KITTI dataset, respectively.