LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling
Yaohua Zha, Naiqi Li, Yanzi Wang, Tao Dai, Hang Guo, Bin Chen, Zhi Wang, Zhihao Ouyang, Shu-Tao Xia
Introduction
3D point cloud perception, as a crucial application of deep learning, has achieved significant success across various areas such as autonomous driving, robotics, and virtual reality. Recently, point cloud self-supervised learning [50; 1; 53], capable of learning universal representations from extensive unlabeled point cloud data, has gained much attention. Among which, masked point modeling (MPM) [53; 33; 57; 7; 58], as an important self-supervised paradigm, has become mainstream in point cloud analysis and has gained immense success across diverse point cloud tasks.
The classical MPM [53; 33; 57], inspired by masked image modeling [2; 17; 51] (MIM), divides point clouds into patches and uses a standard Transformer backbone. It randomly masks some patches in the encoder input and combines the unmasked patch tokens with randomly initialized masked patch tokens in the decoder input. It predicts the geometric coordinates or semantic features of the masked patches from the decoder output tokens, enabling the model to learn universal 3D representations. Despite the significant success, two inherent issues of Transformers still limit their practical deployment.
The first issue is that the Transformer architecture leads to quadratic complexity and huge model sizes. As shown in Figure 1 (a) and (b), MPM methods like Point-MAE based on standard Transformer require 22.1M parameters and complexity exponentially grows with an increase in the length of input patches. However, in practical point cloud applications, models are often deployed on embedded devices such as robots or VR headsets, where strict constraints exist regarding the model’s size and complexity. In this context, lightweight networks such as PointNet++ are more popular in practical applications due to their lower parameter requirement (only 1.5M) even though they may have inferior performance.
Another issue is that when Transformers are used as decoders in Masked Point Modeling (MPM), their potential to reconstruct masked patches with lower information density is limited. In the decoder input of MPM, randomly initialized masked tokens with lower information density are typically concatenated with unmasked tokens with higher information density and fed into the Transformer-based decoder. The self-attention layers then learn to process these tokens of varying information density based on loss constraints. However, relying solely on the loss to learn this objective is challenging due to the lack of explicit importance guidance for different densities. Additionally, in Section 5.1, we further explain from an information theory perspective that the self-attention mechanism, as a higher-order processing function, can limit the model’s reconstruction potential.
To address the above issues, as shown in Figure 2, we first conducted a comprehensive analysis of the effects of different top-K attention on the performance of the Transformer model, emphasizing the idea that redundancy reduction is crucial for point cloud analysis. To this end, we propose a Locally constrained Compact point cloud Model (LCM), consisting of a locally constrained compact encoder and a locally constrained Mamba-based decoder, to replace the standard Transformer. Specifically, based on the idea of redundancy reduction, our compact encoder replaces self-attention with our local aggregation layers to achieve an elegant balance between performance and efficiency. The local aggregation layer leverages static local geometric constraints to aggregate the most relevant information for each patch token. Since static local geometric constraints only need to be computed once at the beginning and are shared across all layers, it avoids dynamic attention computations in each layer, significantly reducing complexity. Furthermore, it uses only two MLPs for information mapping, greatly reducing the network’s parameters.
In our decoder design, considering the varying information density between masked and unmasked patches in the inputs of MPM, our decoder introduces the State Space Model (SSM) from Mamba [9; 12; 26; 60; 14] to replace self-attention, ensuring linear complexity while maximizing the perception of point cloud geometry information from unmasked patches with higher information density. However, as discussed in section 7, the directly replaced SSM layer exhibits a strong dependence on the order of input patches. Inspired by our compact encoder, we migrate the idea of local constraints to the feedforward neural network of our Mamba-based decoder, proposing the Local Constraints Feedforward Network (LCFFN). This eliminates the need to explicitly consider the sequence order of input in SSM layers because the subsequent LCFFN can adaptively exchange information among geometrically adjacent patches based on their implicit geometric order.
Our LCM is a universal point cloud architecture designed based on the characteristics of the point cloud to replace the standard Transformer. It can be trained from scratch or integrated into any existing pretraining strategy to achieve an elegant balance between performance and efficiency. For example, the LCM model pre-trained based on the Point-MAE strategy requires only 2.7M parameters, which is about 10 efficient compared to the original Transformer with 22.1M. Furthermore, in terms of performance, compared to the Transformer, the LCM shows significant accuracy improvements of 2.24%, 0.87%, and 0.94% in the classification tasks of three variants of ScanObjectNN . Additionally, in the detection task of ScanNetV2 , there are also significant improvements of +5.2% on and +6.0% on .
We summarize the contributions of our paper as follows: 1) We propose a locally constrained compact encoder, which leverages static local geometric constraints to aggregate the most relevant information for each patch token, achieving an elegant balance between performance and efficiency. 2) We propose a locally constrained Mamba-based decoder for masked point modeling, which replaces the self-attention layer with Mamba’s SSM layer and introduces a locally constrained feedforward neural network to eliminate the explicit dependency of Mamba on the input sequence order. 3) Our locally constrained compact encoder and locally constrained Mamba-based decoder together constitute the efficient backbone LCM for masked point modeling. We combine LCM with various pretraining strategies to pre-train efficient models and validate our model’s superiority in efficiency and performance across various downstream tasks.
Related Work
Point Cloud Self-supervised Pre-training. Point cloud self-supervised pre-training has achieved remarkable improvement in many point cloud tasks. This approach first applies a pretext task to learn the latent 3D representation and then transfers it to various downstream tasks. PointContrast and CrossPoint initially explored utilizing contrastive learning [32; 40] for learning 3D representations, which achieved some success; however, there were still some shortcomings in capturing fine-grained semantic representations. Recently, masked point modeling methods [53; 33; 24; 5; 54] demonstrated significant improvements in learning fine-grained point cloud representations through masking and reconstruction. Many methods [7; 16; 4; 58; 37] have attempted to leverage multimodal knowledge to assist MPM in learning more generalized representations, yielding significant improvements. While the pre-trained models mentioned above have achieved tremendous success, they all rely on the Transformer architecture. In this paper, we focus on designing a more efficient architecture to replace the Transformer in these methods, significantly reducing computational and resource requirements.
Methodology
Standard Transformer architecture requires computing the correlation between each patch with all input patches, resulting in quadratic complexity. While this architecture performs well in language data, its effectiveness in point cloud data has been under-explored. Not all points are equally important. As illustrated in Figure 3, the key points for aircraft recognition are mainly distributed on the wings, while for vase recognition, they are primarily located on the bottom of the vase. Therefore, directly skipping the attention computation for less important points provides a straightforward solution.
We first replaced the computation of global attention for all patch tokens with calculations top-K attentions in both feature and geometric space. As shown in Figure 2, our empirical observations indicate that: 1) In self-attention, it is often more effective to use attention weights based on the top-K most important patch tokens rather than using all patch; 2) Compared to using top-K attention in a dynamic feature space, employing top-K attention in a static geometric space yields nearly identical representational capacity and offers the advantage of a smaller K value. Although this naive method of masking out unimportant attention still exhibits quadratic complexity, this redundancy reduction idea not only brings performance improvements but also provides a direction for further optimizing computational efficiency.
2 The Pipeline of Masking Point Modeling with LCM
The overall architecture of our Locally constrained Compact Model (LCM) is shown in Figure 4. The specific process is as follows.
Encoder. We employ our locally constrained compact encoder to extract features from the unmasked features . It consists of stacked encoder layers, each layer incorporating a local aggregation layer and a feedforward neural network, detailed in Figure 4. For the input feature of the -th layer, after adding its positional embedding , it feeds to the -th encoding layer to obtain the feature . Therefore, the forward process of each encoder layer is defined as:
Decoder. In the decoding phase, although various MPMs have different decoding strategies, they can generally be divided into feature-level or coordinate-level reconstruction, and their decoders mostly rely on the Transformer architecture. Here, we illustrate the decoding process of our locally constrained Mamba-based decoder using the coordinate-level reconstruction method Point-MAE as an example.
Reconstruction. We utilize the features decoded by the decoder to perform the 3D reconstruction. We employ multi-layer MLPs to construct coordinates reconstruction head and our reconstruction target is to recover the relative coordinates of the masked patches. We use the Chamfer Distance () as reconstruction loss. Therefore, our loss function is as follows
3 Locally Constrained Compact Encoder
The classical Transformer relies on the self-attention mechanism to perceive long-range correlations among all patches globally and has achieved great success in language and image domains. However, there remains uncertainty about whether directly transferring a Transformer-based encoder is suitable for point cloud data. Firstly, applications of point clouds are more inclined towards practical embedded devices such as robots or VR headsets. The hardware resources of these devices are limited, imposing higher limits on the model size and complexity, and the Transformer-based backbone demands significantly more resources than traditional networks, as illustrated in Table 1. Secondly, extensive research [36; 46; 30] and our empirical observation as illustrated in Figure 2 also indicate that the perception of local geometry in point cloud data far outweighs the need for global perception. Therefore, the computation of long-range correlations in self-attention leads to a considerable amount of redundant calculations. To address these practical issues, we propose a locally constrained compact encoder.
Our compact encoder consists of stacked compact encoder layers, each layer comprising a local aggregation layer (LAL) and a feed-forward network (FFN), as shown in Figure 5 (a). For the -th encoder layer, the output () of the preceding layer, added with the positional embedding and normalized by layer normal, is initially fed to the Local Aggregation Layer (LAL) for aggregating local geometric. Afterward, the result is added to the input residual, passed through layer normalization, and finally fed into a Feed-forward Network (FFN) to obtain the ultimate output feature (). This process can be formalized as follows,
where represents the LAM, and represents layer normalization, and represents the FFN.
4 Locally Constrained Mamba-based Decoder
The decoder for mask point modeling needs to recover information about masked patches based on the features extracted from unmasked patches by the encoder. A common approach is to concatenate the features of unmasked patches before randomly initialized features of masked patches as the input to the decoder, as shown in Figure 4. However, at this point, there is a significant difference in information density between features and . The Transformer architecture and our local aggregation layer both treat each token in the input as equally important initially, it works well when the information density of all tokens is similar. It does not adapt well to cases where there is a large difference in information density in the input.
To efficiently extract more geometric priors from unmasked features , we were inspired by the Mamba model in time sequence and proposed using a Mamba-based decoder. This decoder can extract more prior information from the preceding tokens in the sequence based on the input order to aid the learning of subsequent tokens. Initially, we simply replaced the self-attention layer in the original Transformer-based decoder with the state space model (SSM) layer from Mamba. We also sorted the input sequence based on the order of each patch’s center point coordinates, creating a naive Mamba-based decoder. Our experiments in section 8 revealed that although this naive decoder is efficient enough, the simple sorting method cannot effectively model the complex spatial geometry of point clouds and leads to a strong dependence on the order of input patches.
To ensure that the SSM fully perceives the spatial geometry of point clouds, we further introduced the concept of local constraints from the local aggregation layer into the feedforward neural network layer of our decoder, getting the Local Constraints Feedforward Network (LCFFN). By feeding the tokens outputted by the SSM layer into the LCFFN, the LCFFN can implicitly exchange information between geometrically adjacent patches based on their central coordinates. This eliminates the limitation in the SSM layer where explicit sequential input fails to perceive complex geometry fully. Finally, in Section 5.1, we also qualitatively explain from an information theory perspective that this Mamba-based architecture has greater reconstruction potential compared to the Transformer.
Our Mamba-based decoder consists of stacked decoder layers, each layer comprising a Mamba SSM layer and a local constraints feedforward network (LCFFN), as shown in Figure 5 (b). For the -th decoder layer, we first add the output () of the previous layer with the positional embeddings () and normalize it through layer normalization. Then, we use the Mamba SSM layer () to perceive geometry from unmasked features and predict masked features. Finally, in the LCFFN (), we further perceive shape priors based on the central coordinates of each token from its geometrically adjacent tokens. This process can be formalized as follows:
Experiments
We pre-training our LCM using five different pretraining strategies: Point-BERT , MaskPoint , Point-MAE , Point-M2AE , and ACT . For a fire comparison, we use ShapeNet as our pre-training dataset, encompassing over 50,000 distinct 3D models spanning 55 prevalent object categories. For the hyperparameter settings during the pretraining phase, we used the same settings as previous methods.
2 Fine-tuning on Downstream Tasks
We assess the performance of our LCM by fine-tuning our models on various downstream tasks, including object classification, scene-level detection, and part segmentation.
We initially assess the overall classification accuracy of our pre-trained models on both real-scanned (ScanObjectNN ) and synthetic (ModelNet40 ) datasets. ScanObjectNN is a prevalent dataset consisting of approximately 15,000 real-world scanned point cloud samples from 15 categories. These objects represent indoor scenes and are often characterized by cluttered backgrounds and occlusions caused by other objects. For the ScanObjectNN dataset, we sample 2048 points for each instance and report results without voting mechanisms. We applied simple scaling and rotation data augmentation of previous work [33; 7] in the downstream setting of ScanObjectNN. We reported the results of different models under our downstream setting, with marking the results. For the ModelNet40 dataset, due to space limitation, we will further analyze its results in section 5.4.
As presented in Table 1, our model has many exciting results. 1) Lighter, faster, and more powerful. When trained from scratch using supervised learning only, our LCM model demonstrates performance improvements of 1.55%, 0.51%, and 1.35% across three variant datasets compared to the Transformer architecture. Similarly, after pre-training (e.g., Point-MAE), our model outperformed the standard Transformer by 2.24%, 0.87%, and 0.94% across the three variants of the ScanObjectNN dataset. Notably, these improvements are achieved despite an 80% reduction in parameters and a 73% reduction in FLOPs. This improvement is exciting as it indicates that our architecture is better suited for point cloud data compared to the standard Transformer. Additionally, due to its extremely high efficiency, it provides strong support for the practical deployment of these pre-trained models. 2) Universal. We have replaced the original Transformer architecture with our LCM model in five different MPM-based pre-training methods. All experimental results are exciting as our model achieved universal performance improvements with fewer parameters and computations, highlighting the versatility of our model. In the future, we will further adapt to additional pre-training methods.
2.2 Object Detection
We further assess the object detection performance of our pre-trained model on the more challenging scene-level point cloud dataset, ScanNetV2 , to evaluate our model’s scene understanding capabilities. Following the previous pre-training work [24; 7], we use 3DETR as the baseline and only replace the Transformer-based encoder of 3DETR with our pre-trained compact encoder. Subsequently, the entire model is fine-tuned for object detection. In contrast to previous approaches [24; 7; 4], which necessitate pre-train on large-scale scene-level point clouds like ScanNet, our approach directly utilizes models pre-trained on ShapeNet. This further emphasizes the generalizability of our pre-trained models.
Table 4.2.2 showcases our experimental results, our compact model has shown significant improvements in scene-level point cloud data, such as Point-MAE achieving a 5.2% improvement in and a 6.0% improvement in compared to the Transformer. This improvement is remarkable, and we believe this is primarily due to the presence of a large number of background and noise points in the scene-level point cloud. Using a local constraint modeling approach effectively filters out unimportant background and noise, allowing the model to focus more on meaningful points.
2.3 Part Segmentation
We also assess the performance of LCM in part segmentation using the ShapeNetPart dataset , comprising 16,881 samples across 16 categories. We utilize the same segmentation setting after the pre-trained encoder as in previous works [33; 57] for fair comparison. As shown in Table 4.2.2, our LCM-based model also exhibits a clear boost compared to Transformer-based models. These results demonstrate that our model exhibits superior performance in tasks such as part segmentation, which demands a more fine-grained understanding of point clouds.
3 Ablation Study
Effects of Locally Constrained Compact Encoder. We explore the performance of our locally constrained compact encoder by comparing it with a Transformer-based encoder in classification, detection, and part segmentation. The results from Tables 1, Table 4.2.2, and Table 4.2.2, obtained solely through supervised learning from scratch, clearly demonstrate the advantages of our LCM encoder over the Transformer-based encoder in terms of performance and efficiency, particularly in detection tasks, with an improvement of up to 6.0% in the metric.
This substantial improvement is attributed to the compact encoder’s focused attention on the most crucial information for each point patch, such as local neighborhoods while disregarding unimportant details. This is similar to a redundancy-reducing compression concept, which is crucial for point cloud analysis, especially in large-scale scene-level point clouds where significant redundancy and noise points often exist. Our local constraint approach enables the model to focus on critical areas, leading to a combined improvement in efficiency and performance. Moreover, this redundancy-reducing concept helps our model avoid overfitting the training dataset. We provide detailed explanations of this phenomenon in the section 5.5.
Effects of the Network Structure of the Locally Constrained Encoder. As shown in Figure 5(a), each layer of our locally constrained compact encoder consists primarily of three parts: a locally constrained unit based on k-NN, MLPs mapping unit composed of Down MLP and Up MLP, and the final FFN layer. We explore the effects of each unit separately. Specifically, we train Encoders with different structures from scratch on the ScanObjectNN dataset and test their classification performance. As shown in Table 4.3, comparing A and B reveals that a simple two-layer MLP without local aggregation does not substantially improve the network’s performance. In contrast, the results of C and D compared to A and B demonstrate a significant performance improvement. This improvement is mainly attributed to the introduction of local geometric perception and aggregation. Comparing the results of C and D, the introduction of FFN brings a slight improvement. Therefore, FFN is not indispensable in our compact encoder, but we choose to incorporate FFN to further perform mapping. These experiments further indicate the necessity of local geometric perception and aggregation for point cloud feature extraction.
Effects of Locally Constrained Mamba-based Decoder. We further compared the impact of different decoder designs during the pre-training phase. Specifically, we compared the vanilla Transformer-based decoder, our LAL-based decoder, and the vanilla Mamba-based architecture, as well as their performance after incorporating LCFFN. As shown in Table 4.3, the results indicate that the Vanilla Transformer slightly outperforms the Vanilla Mamba in terms of performance, likely due to the limitation imposed by the simple geometric sequential input sequences on the Vanilla Mamba’s capabilities. After incorporating LCFFN, the Mamba decoder exhibits a significant improvement due to the introduction of implicit geometric order. In contrast, the Transformer’s improvement is slight because the geometric order is already implicitly captured by self-attention.
4 Limitation
The main limitation of our LCM is its lack of dynamic perception of importance for each point patch. Since our method primarily relies on local constraints in geometric space to identify important point patches for efficiency, this approach represents a static importance that may result in missing crucial parts in certain regions of point clouds, thereby limiting the model’s representational capacity in some aspects. In the future, we will further explore efficient strategies for dynamic importance.
5 Conclusion
In this paper, we propose a compact point cloud model, LCM, specifically designed for masked point modeling pre-training, aiming to achieve an elegant balance between performance and efficiency. Based on the idea of redundancy reduction, we propose focusing on the most relevant point patches ignoring unimportant parts in the encoder, and introducing a local aggregation layer to replace the vanilla self-attention. Considering the varying information density between masked and unmasked patches in the decoder inputs of MPM, we introduce a locally constrained Mamba-base decoder to ensure linear complexity while maximizing the perception of point cloud geometry information from unmasked patches. By conducting extensive experiments across various tasks such as classification and detection, we demonstrate that our LCM is a universal model with significant improvements in efficiency and performance compared to traditional Transformer models.
References
Appendix
Here, we provide an information-theoretic perspective for our decoder design, using mutual information to qualitatively demonstrate that the Mamba-based SSM can perceive more information from unmasked patches to predict masked patches compared to a Transformer-based self-attention. The mutual information between random variables and , , measures the amount of information that can be gained about a random variable from the knowledge about the other random variable . Therefore, based on the decoder input’s different information densities, we can simply divide the input into , representing unmasked patches with higher information density, and , representing randomly initialized masked patches with lower information density. As illustrated in Figure 6, after being processed by the decoder, and respectively yield outputs for unmasked patches and for masked patches. We reconstruct the masked points based on .
Ideally, needs to perceive sufficient geometric priors from both and to recover the masked points, more mutual information represents more recovery potential. Therefore, we would like to maximize the mutual information . In what follows, we demonstrate that the mutual information preserved by our proposed Mamba-based decoder is larger than that of the standard transformer decoder.
Let and denote the outputs of the Mamba-based and Transformer-based decoders respectively, denote the mutual information preserved by the Mamba-based decoder, and denote that of the Transformer-based decoder. We have .
The first step is to formalize the input-output relation of the two decoding structures. For the Mamba decoder, as defined in [12; 9], the output can be expressed as:
For the Transformer decoder, the attention mechanism can be expressed in the following matrix form:
Compared with the linear relation captured by , models higher-order interactions of the input variables. So for any given Mamba parameters and , there exists Transformer parameters and a function , such that .
As is a function of , forms a Markov chain. So and are independent when conditioned on , i.e., . According to the definition of conditional mutual information, this implies
On the other hand, by the chain rule of mutual information we have
Since we already show that , and mutual information is non-negative, we have
2 Additional Related Work
Deep Network Architecture for Point Cloud. Point clouds, as 3D data directly sampled from scanning devices, inherently exhibit irregularity and disorder. To employ deep neural networks for point cloud analysis, various structures [35; 36; 46; 22; 48; 38; 25; 55; 52; 21] have been developed. PointNet , a pioneer in point cloud analysis, introduced an MLP-based network to address the disorder of point clouds. Subsequently, PointNet++ further proposed adaptive aggregation of multiscale features on MLPs and incorporated local point sets for effective feature learning. DGCNN introduced the graph convolutional networks dynamically computing local graph neighboring nodes to extract geometric information. PointMLP suggested efficient point cloud representation solely relying on pure residual MLPs. Recently, many Transformer-based models [15; 30; 33; 54], benefiting from attention mechanisms, have achieved notable improvements in point cloud analysis. However, this led to a significant increase in model size, posing considerable challenges for practical applications. PointMamba first attempted to introduce the Mamba architecture based on the state space model to point clouds, but it still has high complexity and parameters. In this paper, we focus on designing more efficient point cloud architectures specific to pre-training models.
State Space Models. State Space Models [10; 11; 12; 13; 39] (SSMs) originate from classical control theory and have been introduced into deep learning as the backbone of state space transformations. They combine the parallel training capabilities of CNNs with the fast inference characteristics of RNNs, capturing long-range dependencies in sequences while maintaining linear complexity. The Structured State-Space Sequence model (S4) is a pioneer work for the deep state-space model in modeling the long-range dependency. S5 proposed based on S4 and introduces MIMO SSM and efficient parallel scan. GSS leverages the gating structure in the gated attention unit to reduce the dimension of the state space module. Recently, Mamba with efficient hardware design and selective state space, outperforms Transformers in terms of performance and efficiency. Subsequent works [60; 26; 45; 27; 20; 31] have attempted to introduce Mamba into the visual domain, achieving significant improvements. For example, Vision Mamba and VMamba directly apply Mamba to image processing and design corresponding scanning methods tailored for image data. As for point cloud, PointMamba is the first to introduce Mamba into point cloud analysis, traversing the input sequences from the x, y, and z geometric directions. In this paper, we introduce Mamba into the decoder for masked point modeling and discuss its advantages from an information-theoretic perspective. Additionally, we propose a locally constrained feedforward neural network for Mamba block to adaptively exchange information among geometrically adjacent patches based on their implicit geometry.
3 Implementation Details
In Figure 2, we replace the global attention computation of all patch tokens in Self-Attention with top-K attention computation in both feature space and geometric space to demonstrate the significant amount of redundant computation in the vanilla Transformer. Specifically, after computing all global attention, we further compute a mask matrix. We then add negative infinity to the attention values that need to be masked. After that, we calculate the softmax, where the attention values that were set to negative infinity will become 0, ensuring that the sum of the attention values of the unmasked top-K patches equals 1. We compute different top-K values in both feature space and geometric space, and pretrain the corresponding models. Subsequently, we fine-tune these pretrained models on the three variants of ScanObjectNN using the same top-K attention algorithm, evaluating their accuracy on classification tasks. To minimize error, we report the average accuracy over 10 repeated experiments.
Due to significant differences in the settings used for downstream fine-tuning tasks of point cloud classification on the ScanObjectNN dataset in previous self-supervised learning methods [53; 33; 7; 57; 24] , such as input point quantity, data augmentation, and the input of the classification task head, we conducted extensive experiments to obtain a performance-friendly downstream fine-tuning setting. Furthermore, we re-evaluated most of the previous methods under our setting, while also conducting a fair comparison between our LCM model and the previous Transformer model under our setting. We mark the results of our downstream fine-tuning setting with an " " in Table 1.
It can be observed that, compared to the results reported in the original paper, the fine-tuning results using our downstream settings have achieved significant performance improvements. For instance, Point-BERT has shown improvements of 5.34%, 3.62%, and 4.99% on the three variants of the ScanObjectNN dataset, respectively. This improvement is surprising, indicating that there is further potential to be explored in earlier self-supervised learning methods such as Point-BERT, Point-MAE, etc.
Due to the surprisingly lightweight and efficient of our LCM model, we were able to complete the pre-training tasks using just a single 24GB NVIDIA GeForce RTX 3090 GPU. For downstream classification and segmentation tasks, we used a single RTX 3090 GPU for each. For detection tasks, to accelerate training, we utilized four parallel RTX 3090 GPUs.
We pre-train and fine-tune Point-MAE for 3D object detection both on ScanNetV2 . In our detection experiments on ScanNetV2, we evaluate our model’s understanding of scene-level tasks. Specifically, in the downstream detection fine-tuning experiments, we use 3DETR as the baseline model and replace 3DETR’s pre-encoder and encoder with our embed layer and compact encoder, respectively, while keeping all other training settings identical to 3DETR. Unlike many previous methods [24; 7; 57] that require retraining models on ScanNet, we initialize the embed layer and compact encoder with models pre-trained directly on ShapeNet . While this may result in some loss of performance due to the gap between ShapeNet and ScanNet data, it demonstrates the universality of our pre-trained models.
4 Additional Experiments
ModelNet40 is a well-known synthetic point cloud dataset, comprising 12,311 meticulously crafted 3D CAD models distributed across 40 categories. Following previous work [53; 33; 57], for the ModelNet40 dataset, we sample 1024 points for each instance and report overall accuracy with voting mechanisms. In ModelNet40, we no longer differentiate between the results reported in the paper and our results, as we use the exact same downstream fine-tuning settings as previous methods [33; 24; 7; 57]. Table 6 presents our experimental results, and the overall conclusions are consistent with Section 4.2.1. Our LCM model outperforms the Transformer architecture in terms of both efficiency and performance, indicating the superiority of our model.
Effects of Locally Constrained K Value. We further explore the impact of using different numbers of neighbors K in local constraints on performance and efficiency. K=1 indicates no consideration of neighboring information. As K increases, the consideration of local geometry for each point patch also increases, but so does the computational complexity. We train object classification from scratch on the PB-RS-T50 variant of ScanObjectNN, and Figure 7 presents our ablation results. The area of the circle represents the computational floating-point operations (FLOPs). We found that a smaller K, such as 5, is sufficient to achieve satisfactory results in terms of performance and efficiency. Performance initially increases slowly, but when K exceeds a certain threshold, it tends to decline. This is mainly due to larger K values introducing excessive redundancy, thereby limiting the learning capacity.
Effects of K-NN Space. We further explored the impact of performing K-NN based on Euclidean distance in both the feature space and the geometric space of our compact encoder. Geometric K-NN in the geometric space imposes explicit geometric constraints, serving as a static importance measure that greatly benefits point cloud analysis. Searching for K-NN based on feature Euclidean distance in the feature space can be considered a simple form of dynamic importance. We analyzed the effect of this approach on point cloud classification from scratch on ScanObjectNN, evaluating geometric K-NN and feature K-NN at different K values.
As shown in Table 7, we found that feature K-NN performed consistently lower than geometric K-NN in almost all cases. This result suggests that the naive idea of assigning dynamic importance to point patches based on Euclidean distance in the feature space does not lead to substantial improvements. Efficient computation of dynamic importance for each point patch remains an area for further exploration.
Effects of Patch Order and LCFFN for Mamba-based Decoder. The ordering of input patches significantly impacts our Mamba-based Decoder. To more effectively illustrate this effect on Mamba’s SSM model, we analyze the issue from a different perspective. Specifically, we use our Mamba-based Decoder as an Encoder to directly extract features from the input point cloud and perform classification on ScanObjectNN. This substitution is straightforward, as our Mamba-based Decoder can also be viewed as an Encoder.
We trained our Mamba-based encoder from scratch for the classification task on the PB-RS-T50 variant of ScanObjectNN without using any data augmentation strategies, and we took the average of ten repeated experiments as the final result. We first experimented with a naive Mamba-based decoder using a traditional FFN to illustrate the impact of different sequence orders on the original Mamba. We selected four different patch ordering methods: sorting by the center point of the patch along the x-axis (X), y-axis (Y), and z-axis (Z), and Hilbert curve ordering (H), as shown by the orange curve in Figure 8 (a). Furthermore, we also conducted experiments with combinations sequences, combining these four orderings "H+X+Y+Z (HXYZ)", "X+Y+Z+H (XYZH)", "Y+Z+H+X (YZHX)", and "Z+H+X+Y (ZHXY)", as shown by the green curve in Figure 8 (a). Finally, based on the single-order sequence, we used our proposed LCFFN to demonstrate the performance of Mamba with added implicit geometric constraints, as shown by the yellow curve in Figure 8 (a). The experimental results, as illustrated in Figure 8, lead us to the following conclusions:
1) The performance of the Mamba model is greatly influenced by the different orders of input patches. The orange line represents the results for individual sequences, highlighting that different sequences have a significant impact on the final model performance. For example, the Y-order achieves the highest classification accuracy at 82.34%, while the Hilbert order performs the worst at 80.65%, resulting in a difference of 1.69%.
2) The more combinations of sequences, the better the representation of point cloud geometry, resulting in improved performance, but also increased computational complexity. The green line represents the combinations sequences. While different combinations sequences do affect the final model performance, the impact is relatively minor. This indicates that the Mamba model can compensate for information across different sequences, allowing it to capture nearly complete geometric information for each patch. Consequently, this significantly enhances the model’s performance. However, this approach leads to a significant increase in computational complexity due to the increase in the length of the input sequence, as shown in Figure 8 (b). The processing time for the sequences of the four orders is approximately longer than that of a single order.
3) Introducing LCFFN allows for better perception of point cloud geometry through implicit local geometric constraints, thereby mitigating the dependence on sequence order. The yellow line represents the experimental results of using LCFFN to replace FFN for single-order input. It can be observed that the overall classification accuracy is significantly improved, surpassing the combinations sequence in the y-order and showing only slight differences from the combinations sequence in other orders. Moreover, in terms of runtime efficiency, as shown in Figure 8 (b), our single-order + LCFFN method exhibits a considerable improvement compared to the combinations sequence, indicating the superiority of our design.
5 Additional Visualization
Effects of the Compact Encoder from the Perspective of Overfitting. While our compact encoder has fewer parameters compared to Transformer-based encoders, its performance surpasses that of Transformer-based encoders , as analyzed in subsection 4.2.1. One significant reason for this lies in the reduced risk of overfitting in downstream tasks due to the redundancy reduction. Given the challenging nature of acquiring point cloud data, existing point cloud datasets for downstream tasks are often small, such as ScanObjectNN and ModelNet40 , each comprising just over 10,000 point clouds, and ScanNetV2 with only 1,000 scenes. These dataset sizes are much smaller than those commonly found in image and language tasks. Therefore, fine-tuning in these size-limited datasets can be more prone to overfitting when considerable redundancy exists in the computation.
We visualize the training and testing curves for different encoders on the classification task in ScanObjectNN and the detection task in ScanNetV2 in Figures 9 and 10. Figure 9 illustrates the classification and detection curves for our compact encoder and a Transformer-based encoder after pretraining. It can be observed that during training, the classification accuracy and AP25 metric of the Transformer-based encoder are significantly higher than those of our compact encoder. However, during testing, our compact encoder exhibits superior performance compared to the Transformer encoder. This starkly indicates that the Transformer-based encoder tends to overfit the training set, demonstrating poorer generalization. Conversely, our compact encoder displays stronger generalization capabilities, indicating the superiority of the design of our compact encoder.
Figure 10 displays the classification and detection curves of our compact encoder and the Transformer-based encoder trained from scratch. In comparison to its counterpart in Figure 9, although it shows a slower convergence, the overfitting issue of the Transformer-based encoder still emerges in the late stages of training, reaffirming our conclusion. Meanwhile, the phenomenon of slow convergence in Figure 10 is reasonable as it is an encoder trained from scratch without a better initialization.
We further used t-SNE to visualize the feature distributions extracted by our LCM model and the Transformer. In Figure 11, we visualized the two-dimensional (2D) feature distributions of the two models, pretrained using Point-MAE, when directly transferred to the test set of the ModelNet40 dataset without downstream fine-tuning. In Figure 12, we visualized the 2D feature distributions of the two pre-trained models after fine-tuning on the most challenging variant of the ScanObjectNN dataset, PB-RS-T50, using its test set.
In the 2D t-SNE visualizations, instances from the same category tend to be distributed in relatively clear and tight clusters. The compactness of the feature distributions of different instances from the same category can be viewed as the model’s ability to represent features of the same category. A more compact distribution indicates a stronger modeling capability. As shown in Figure 11 and Figure 12, our LCM model achieves more compact feature distributions for instances of the same category compared to the Transformer model in most cases, indicating that our LCM model has a stronger ability to model the general representations of the same category.