Learning 3D Representations from 2D Pre-trained Models via Image-to-Point Masked Autoencoders

Renrui Zhang, Liuhui Wang, Yu Qiao, Peng Gao, Hongsheng Li

Introduction

Driven by huge volumes of image data , pre-training for better visual representations has gained much attention in computer vision, which benefits a variety of downstream tasks . Besides supervised pre-training with labels, many researches develop advanced self-supervised approaches to fully utilize raw image data via pre-text tasks, e.g., image-image contrast , language-image contrast , and masked image modeling . Given the popularity of 2D pre-trained models, it is still absent for large-scale 3D datasets in the community, attributed to the expensive data acquisition and labor-intensive annotation. The widely adopted ShapeNet only contains 50k point clouds of 55 object categories, far less than the 14 million ImageNet and 400 million image-text pairs in 2D vision. Though there have been attempts to extract self-supervisory signals for 3D pre-training , raw point clouds with sparse structural patterns cannot provide sufficient and diversified semantics compared to colorful images, which constrain the generalization capacity of pre-training. Considering the homology of images and point clouds, both of which depict certain visual characteristics of objects and are related by 2D-3D geometric mapping, we ask the question: can off-the-shelf 2D pre-trained models help 3D representation learning by transferring robust 2D knowledge into 3D domains?

To tackle this challenge, we propose I2P-MAE, a Masked Autoencoding framework that conducts Image-to-Point knowledge transfer for self-supervised 3D point cloud pre-training. As shown in Figure 1, aided by 2D semantics learned from abundant image data, our I2P-MAE produces high-quality 3D representations and exerts strong transferable capacity to downstream 3D tasks. Specifically, referring to 3D MAE models in Figure 2 (Left), we first adopt an asymmetric encoder-decoder transformer as our fundamental architecture for 3D pre-training, which takes as input a randomly masked point cloud and reconstructs the masked points from the visible ones. Then, to acquire 2D semantics for the 3D shape, we bridge the model gap by efficiently projecting the point cloud into multi-view depth maps. This requires no time-consuming offline rendering and largely preserves 3D geometries from different perspectives. On top of that, we utilize off-the-shelf 2D models to obtain the multi-view 2D features along with 2D saliency maps of the point cloud, and respectively guide the pre-training from two aspects, as shown in Figure 2 (Right).

Firstly, different from existing methods to randomly sample visible tokens, we introduce a 2D-guided masking strategy that reserves point tokens with more spatial semantics to be visible for the MAE encoder. In detail, we back-project the multi-view semantic saliency maps into 3D space as a spatial saliency cloud. Each element in the saliency cloud indicates the semantic significance of the corresponding point token. Guided by such saliency cloud, the 3D network can better focus on the visible critical structures to understand the global 3D shape, and also reconstruct the masked tokens from important spatial cues.

Secondly, in addition to the recovering of masked point tokens, we propose to concurrently reconstruct 2D semantics from the visible point tokens after the MAE decoder. For each visible token, we respectively fetch its projected 2D representations from different views, and integrate them as the 2D-semantic learning target. By simultaneously reconstructing the masked 3D coordinates and visible 2D concepts, I2P-MAE is able to learn both low-level spatial patterns and high-level semantics pre-trained in 2D domains, contributing to superior 3D representations.

With the aforementioned image-to-point guidance, our I2P-MAE significantly accelerates the convergence speed of pre-training and exhibits state-of-the-art performance on 3D downstream tasks, as shown in Figure 3. Learning from 2D ViT pre-trained by CLIP , I2P-MAE, without any fine-tuning, achieves 93.4% classification accuracy by linear SVM on ModelNet40 , which has surpassed the fully fine-tuned results of Point-BERT and Point-MAE . After fine-tuning, I2P-MAE further achieves 90.11% classification accuracy on the hardest split of ScanObjectNN , significantly exceeding the second-best Point-M2AE by +3.68%. The experiments fully demonstrate the effectiveness of learning from pre-trained 2D models for superior 3D representations.

Our contributions are summarized as follows:

We propose Image-to-Point Masked Autoencoders (I2P-MAE), a pre-training framework to leverage 2D pre-trained models for learning 3D representations.

We introduce two strategies, 2D-guided masking and 2D-semantic reconstruction, to effectively transfer the well learned 2D knowledge into 3D domains.

Extensive experiments have been conducted to indicate the significance of our image-to-point pre-training.

Related Work

Supervised learning for point clouds has attained remarkable progress by delicately designed architectures and local operators . However, such methods learned from closed-set datasets are limited to produce general 3D representations. Instead, self-supervised pre-training via unlabelled point clouds have revealed promising transferable ability, which provides a good network initialization for downstream fine-tuning. Mainstream 3D self-supervised approaches adopt encoder-decoder architectures to recover the input point clouds from the transformed representations, including point rearrangement , part occlusion , rotation , downsampling , and codeword encoding . Concurrent works also adopt contrastive pre-text tasks between 3D data pairs, such as local-global relation , temporal frames , and augmented viewpoints . More recent works adopt pre-trained CLIP for zero-shot 3D recognition , or introduce masked point modeling as strong 3D self-supervised learners. Therein, Point-BERT utilizes pre-trained tokenizers to indicate discrete point tokens, while Point-MAE and Point-M2AE apply Masked Autoencoders (MAE) to directly reconstruct the 3D coordinates of masked tokens. Our I2P-MAE also adopts MAE as the basic pre-training framework, but is guided by 2D pre-trained models via image-to-point learning schemes, which benefits 3D pre-training with diverse 2D semantics.

Masked Autoencoders.

To achieve more efficient masked image modeling , MAE is firstly proposed on 2D images with an asymmetric encoder-decoder transformer . The encoder takes as input a randomly masked image and is responsible for extracting its high-level latent representation. Then, the lightweight decoder explores informative cues from the encoded visible features, and reconstructs raw RGB pixels of the masked patches. Given its superior performance on downstream tasks, a series of follow-up works have been developed to improve MAE with customized designs: pyramid architectures with convolution stages , window attention by grouping visible tokens , high-level targets with semantic-aware sampling , and others . Following the spirit, Point-MAE and MAE3D extend MAE-style pre-training on 3D point clouds, which randomly sample visible point tokens for the encoder and reconstruct masked 3D coordinates via the decoder. Point-M2AE further modifies the transformer architecture to be hierarchical for multi-scale 3D learning. Our proposed I2P-MAE aims to endow masked autoencoding on point clouds with the guidance from 2D pre-trained knowledge. By introducing the 2D-guided masking and 2D-semantic reconstruction, I2P-MAE fully releases the potential of MAE paradigm for 3D representation learning.

D-to-3D Learning.

Except for jointly training 2D-3D networks , only a few existing researches focus on 2D-to-3D learning, and can be categorized into two groups. To eliminate the modal gap, the first group either upgrades 2D pre-trained networks into 3D variants for processing point clouds (convolutional kernels inflation , modality-agnostic transformer ), or projects 3D point clouds into 2D images with parameter-efficient tuning (multi-view adapter , point-to-pixel prompting ). Different from them, our I2P-MAE benefits from three properties. 1) Pre-training: We learn from 2D models during the pre-training stage, and then can be flexibly adapted to various 3D scenarios by fine-tuning. However, prior works are required to utilize 2D models on evert different downstream task. 2) Self-supervision: I2P-MAE is pre-trained by raw point clouds in a self-supervised manner and learns more general 3D representations. In contrast, prior works are supervised by labelled downstream 3D data, which might constrain the diverse 2D semantics into specific 3D domains. 3) Independence: Prior works directly inherit 2D network architectures for 3D training, which is memory-consuming and not 3D extensible, while we train an independent 3D network and regard 2D models as semantic teachers. Another group of 2D-to-3D methods utilize paired real-world 2D-3D data of indoor or outdoor scenes, and conduct contrastive learning for knowledge transfer. Our approach also differs from them in two ways. 4) Masked Autoencoding. I2P-MAE learns 2D semantics via the MAE-style pre-training without any contrastive loss between point-pixel pairs. 5) 3D Data Only. We require no real-world image dataset during the pre-training and efficiently projects the 3D shape into depth maps for 2D features extraction.

Method

The overall pipeline of I2P-MAE is shown in Figure 4. In Section 3.1, we first introduce I2P-MAE’s basic 3D architecture for point cloud masked autoencoding without 2D guidance. Then in Section 3.2, we show the details of utilizing 2D pre-trained models to obtain visual representations from 3D point clouds. Finally in Section 3.3, we present how to conduct image-to-point knowledge transfer for 3D representation learning.

As our basic framework for image-to-point learning, I2P-MAE refers to existing works to conduct 3D point cloud masked autoencoding, which consists of a token embedding module, an encoder-decoder transformer, and a head for reconstructing masked 3D coordinates.

Encoder-Decoder Transformer.

D-coordinate Reconstruction.

where H3D⁡(⋅)\operatorname{H_{3D}}(\cdot) denotes the head to reconstruct the masked 3D coordinates.

2 2D Pre-trained Representations

We can leverage 2D models of different architectures (ResNet , ViT ) and various pre-training approaches (supervised and self-supervised ones) to assist the 3D representation learning. To align the input modality for 2D models, we project the input point cloud onto multiple image planes to create depth maps, and then encode them into multi-view 2D representations.

To ensure the efficiency of pre-training, we simply project the input point cloud PP from three orthogonal views respectively along the x,y,zx,y,z axes. For every point, we directly omit each of its three coordinates and round down the other two to obtain the 2D location on the corresponding map. The projected pixel value is set as the omitted coordinate to reflect relative depth relations of points, which is then repeated by three times to imitate the three-channel RGB. The projection of I2P-MAE is highly time-efficient and involves no offline rendering , projective transformation , or learnble prompting . We denote the projected multi-view depth maps of PP as {Ii}i=13\{I_{i}\}_{i=1}^{3}.

D Visual Features.

D Saliency Maps.

3 Image-to-Point Learning Schemes

On top of the 2D pre-trained representations for the point cloud, I2P-MAE’s pre-training is guided by two image-to-point learning designs: 2D-guided masking before the encoder, and 2D-semantic reconstruction after the decoder.

where I2P⁡(⋅)\operatorname{I2P}(\cdot) denotes the 2D-to-3D back-projection operation in Figure 5. We apply a softmax function to normalize the MM points within S3DS^{\rm 3D}, and regard each element’s magnitude as the visible probability for the corresponding point token. With this 2D semantic prior, the random masking becomes a nonuniform sampling with different probabilities for different tokens, where the tokens covering more critical 3D structures are more likely to be preserved. This boosts the representation learning of the encoder by more focusing on significant 3D geometries, and provides the masked tokens with more informative cues at the decoder for better reconstruction.

D-semantic Reconstruction.

where H2D⁡(⋅)\operatorname{H_{2D}}(\cdot) denotes the head to reconstruct visible 2D semantics, parallel to H3D⁡(⋅)\operatorname{H_{3D}}(\cdot) for masked 3D coordinates. The final pre-training loss of our I2P-MAE is then formulated as LI2P=L3D+L2D\mathcal{L}_{\rm I2P}=\mathcal{L}_{\rm 3D}+\mathcal{L}_{\rm 2D}. With such image-to-point feature learning, I2P-MAE not only encodes low-level spatial variations from 3D coordinates, but also explores high-level semantic knowledge from 2D representations. As the 2D-guided masking has preserved visible tokens with more spatial significance, the 2D-semantic reconstruction upon them further benefits I2P-MAE to learn more discriminative 3D representations.

Experiments

In Section 4.1, we first introduce our pre-training settings and the linear SVM classification performance without fine-tuning. Then in Section 4.2, we present the results by fully fine-tuning I2P-MAE on various 3D downstream tasks. Finally in Section 4.3, we conduct ablation studies to investigate the characteristics of I2P-MAE.

We adopt the popular ShapeNet for self-supervised 3D pre-training, which contains 57,448 synthetic point clouds with 55 object categories. For fair comparison, we follow the same MAE transformer architecture as Point-M2AE : a 3-stage encoder with 5 blocks per stage, a 2-stage decoder with 1 block per stage, 2,048 input point number (NN), 512 downsampled number (MM), 16 nearest neighbors (kk), 384 feature channels (CC), and the mask ratio of 80%. For off-the-shelf 2D models, we utilize ViT-Base pre-trained by CLIP as default and keep its weights frozen during 3D pre-training. We project the point cloud into three 224×224224\times 224 depth maps, and obtain the 2D feature size H×WH\times W of 14×1414\times 14. I2P-MAE is pre-trained for 300 epochs with a batch size 64 and learning rate 10−310^{-3}. We adopt AdamW optimizer with a weight decay 5×10−25\times 10^{-2} and the cosine scheduler with 10-epoch warm-up.

Linear SVM.

To evaluate the transfer capacity, we directly utilize the features extracted by I2P-MAE’s encoder for linear SVM on the synthetic ModelNet40 and real-world ScanObjectNN without any fine-tuning or voting. As shown in Table 1, for 3D shape classification in both domains, I2P-MAE shows superior performance and exceeds the second-best respectively by +0.5% and +3.0% accuracy. Our SVM results (93.4%, 87.1%) can even surpass some existing methods after full downstream training in Table 2 and 3, e.g., PointCNN (92.2%, 86.1%), Transformer (91.4%, 79.86%), and Point-BERT (92.7%, 87.43%). In addition, guided by the 2D pre-trained models, I2P-MAE exhibits much faster pre-training convergence than Point-MAE and Point-M2AE in Figure 3. Therefore, the SVM performance of I2P-MAE demonstrates its learned high-quality 3D representations and the significance of our image-to-point learning schemes.

2 Downstream Tasks

After pre-training, I2P-MAE is fine-tuned for real-world and synthetic 3D classification, and part segmentation. Except ModelNet40 , we do not use the voting strategy for evaluation.

The challenging ScanObjectNN consists of 11,416 training and 2,882 test 3D shapes, which are scanned from the real-world scenes and thus include backgrounds with noises. As shown in Table 2, our I2P-MAE exerts great advantages over other self-supervised methods, surpassing the second-best by +2.93%, +2.76%, and +3.68% respectively for the three splits. This is also the first model reaching 90% accuracy on the hardest PB-T50-RS spilt. As the pre-training point clouds are synthetic 3D shapes with a large domain gap with the real-world ScanobjectNN, the results well indicate the universality of I2P-MAE inherited from the 2D pre-trained models.

Synthetic 3D Classification.

The widely adopted ModelNet40 contains 9,843 training and 2,468 test 3D point clouds, which are sampled from the synthetic CAD models of 40 categories. We report the classification accuracy of existing methods before and after the voting in Table 3. As shown, our I2P-MAE achieves leading performance for both settings with only 1k input point number. For linear SVM without any fine-tuning, I2P-MAE-svm can already attain 93.4% accuracy and exceed most previous works, indicating the powerful transfer ability. By fine-tuning the entire network, the accuracy can be further boosted by +.0.3% accuracy, and achieve 94.1% after the offline voting.

Part Segmentation.

The synthetic ShapeNetPart is selected from ShapeNet with 16 object categories and 50 part categories, which contains 14,007 and 2,874 samples for training and validation. We utilize the same segmentation head after the pre-trained encoder as previous works for fair comparison. The head only conducts simple upsampling for point tokens at different stages and concatenates them alone the feature dimension as the output. Two types of mean IoU scores, mIoUC and mIoUI are reported in Table 4. For such highly saturated benchmark, I2P-MAE can still exert leading performance guided by the well learned 2D knowledge, e.g., +0.29% and +0.25% higher than Point-M2AE concerning the two metrics. This demonstrates that the 2D guidance also benefits the understanding for fine-grained point-wise 3D patterns.

3 Ablation Study

In this section, we explore the effectiveness of different components in I2P-MAE. We utilize our final solution as the baseline for ablation and report the linear SVM classification accuracy (%) for comparison by default.

In Table 5, we experiment different masking strategies for the masked autoencoding of I2P-MAE. The first row represents our I2P-MAE with 2D-guided masking, which preserves more semantically important tokens to be visible for the encoder. Compared to the second row with random masking, the guidance of 2D saliency maps contributes to +0.4% and +0.9% classification accuracy respectively on the two downstream datasets. Then, we reverse the token scores in the spatial semantic cloud, and instead mask the most important tokens. As shown in the third row, the SVM results are largely harmed by -0.9% and -3.3%, demonstrating the significance of encoding critical 3D structures in the encoder. Finally, we modify the masking ratio by ±\pm 0.1, which controls the proportion between visible and masked tokens. The performance decay indicates that, the 2D-semantic and 3D-coordinate reconstruction are required to be well balanced for a properly challenging pre-text task.

D-semantic Reconstruction.

In Table 6, we investigate which groups of tokens are the best for learning 2D-semantic targets. The comparison of the first two rows reveals the effectiveness of reconstructing 2D semantics from visible tokens for point cloud pre-training, i.e., +0.5% and +2.4% classification accuracy. By only using 2D targets for either the visible or masked tokens (the 3rd{}^{\text{rd}} and 4th{}^{\text{th}} rows), we verify that the 3D-coordinate reconstruction still plays an important role in I2P-MAE, which learns low-level geometric 3D patterns and provides complementary knowledge to the high-level 2D semantics. However, if the 3D and 2D targets are both reconstructed from the masked tokens (the last row), the network is restricted to learn 2D knowledge from the nonsignificant masked 3D geometries, other than the more discriminative parts. Also, assigning two targets on the same tokens might cause 2D-3D semantic conflicts. Therefore, the best-performing configuration is to reconstruct 2D and 3D targets separately from the visible and masked point tokens (the first row).

Pre-training with Limited 3D Data.

In Table 7 and Figure 7, we randomly sample the pre-training dataset, ShapeNet , by different ratios, and evaluate the performance of I2P-MAE when 3D data is further deficient. Aided by 2D pre-trained models, I2P-MAE still achieves competitive downstream accuracy in low-data regimes, especially for 20% and 60%, which outperforms Point-M2AE by +1.3% and +1.0, respectively. Importantly, with only 60% of the pre-training, I2P-MAE (93.1%) outperforms Point-MAE (91.0%) and Point-M2AE (92.9%) with full training data. This indicates that our image-to-point learning scheme can effectively alleviate the need for large-scale 3D training datasets.

Effectiveness of Pre-training.

In Table 8, we compare the performance on different downstream tasks between training from scratch and fine-tuning after pre-training. For ScanObjectNN , the pure 3D pre-training without image-to-point learning (‘w/o 2D Guidance’) can improve the classification accuracy by +1.09%, and our proposed 2D-to-3D knowledge transfer further boosts the performance by +2.59%. Similar improvement can be observed on other downstream datasets, which demonstrates the significance of the pre-training of I2P-MAE.

Visualization

To ease the understanding of our approach, we visualize the input point cloud, random masking, spatial saliency cloud, 2D-guided masking, and the reconstructed 3D coordinates in Figure 6. Guided by the semantic scores from 2D pre-trained models (darker points indicate higher scores), the masked point cloud largely preserves the significant parts of the 3D shape, e.g., the main body of an airplane, the grip of a guitar, the frame of a chair and headphone. In this way, the 3D network can learn more discriminative features by reconstructing these visible 2D semantics.

Conclusion

In this paper, we propose I2P-MAE, a masked point modeling framework with effective image-to-point learning schemes. We introduce two approaches to transfer the well learned 2D knowledge into 3D domains: 2D-guided masking and 2D-semantic reconstruction. Aided by the 2D guidance, I2P-MAE learns superior 3D representations and achieves state-of-the-art performance on 3D downstream tasks, which alleviates the demand for large-scale 3D datasets. For future work, not limited to masking and reconstruction, we will explore more sufficient image-to-point learning for 3D masked autoencoders, e.g., point token sampling and 2D-3D class-token contrast. Also, we expect our pre-trained models to benefit wider ranges of 3D tasks, e.g., 3D object detection and visual grounding.

Appendix

In this section, we present the detailed model configuration and training settings for fine-tuning I2P-MAE on downstream tasks. All experiments are conducted on a single RTX 3090 GPU.

For both ModelNet40 and ScanObjectNN , we fine-tune I2P-MAE for 300 epochs with a batch size 32. We adopt AdamW optimizer with a learning rate 0.0005 and weight decay 0.05, and utilize cosine scheduler with a 10-epoch warm-up. We append a 3-layer MLP after I2P-MAE’s encoder as the classification head. For ScanObjectNN, we adopt max and average pooling to respectively summarize the point tokens from the encoder, and concatenate the two global features along the feature dimension for the head. I2P-MAE follows existing methods to take 2,048 points as input, and adopts random scaling with rotation as data augmentation. For ModelNet40, we element-wisely add the two global features for the classification head. I2P-MAE takes 1,024 points as input, and adopts random scaling with translation as data augmentation.

Part Segmentation.

On ShapeNetPart , we fine-tune I2P-MAE for 300 epochs with a batch size 16. We also adopt AdamW optimizer with a learning rate 0.0002 and weight decay 0.00005, and utilize cosine scheduler with a 10-epoch warm-up. For fair comparison, we experiment with the same segmentation head and training settings as Point-M2AE .

2 Additional Ablation Study

In Figure 8 and 9, we show the comparison of training I2P-MAE from scratch and fine-tuning after pre-training on two shape classification datasets. Our image-to-point pre-training can largely accelerate the convergence speed during fine-tuning and the final classification accuracy, indicating the effectiveness of the 2D-to-3D knowledge transfer.

Projected View Number.

In Table 10 (1st and 2nd rows), we show how the number of projected views affect the performance of I2P-MAE. As default, we project the point cloud into 3 views along the x,y,zx,y,z axes. For the view number 1 and 2, we enumerate all possible projected views along x,y,zx,y,z axes, and report the highest results in the table. As shown, using less views would harm the pre-training performance, which constrains 2D pre-trained models from ‘seeing’ complete 3D shapes due to occlusion. Instead, the 3D network can learn more comprehensive high-level semantics from the 2D representations of all three views.

D Sailency Cloud and 2D-semantic Target.

We first investigate how to aggregate multi-view 2D sailency maps as the sailency cloud for 2D-guided masking in Table 10 (3rd and 4th rows). Compared to assigning the maximum or minimum score to a certain point, averaging the 2D sailency scores from different views achieves the best performance. Then, we explore how to generate the 2D-semantic targets from multi-view 2D features in Table 10 (5th row). The results indicate that, concatenating 2D features between different views performs better than averaging them, which preserves more diverse 2D semantics for reconstruction.

Fine-tuning Settings.

In Table 11, we experiment different fine-tuning settings for downstream shape classification on the two datasets. For the point tokens from the encoder, ‘Max Only’ and ‘Ave Only’ denote applying either max or average pooling to summarize global features for the classification head. ‘Add’ or ‘Concat’ denotes to add or concatenate the two global features after max and average pooling. We observe that, ‘Add’ and ‘Concat’ perform the best for ModelNet40 and ScanObjectNN , respectively.

3 Few-shot Classification

We fine-tune I2P-MAE for few-shot classification on ModelNet40 in Table 9. Following previous work , we adopt the same training settings and few-shot dataset splits, i.e., 5-way 10-shot, 5-way 20-shot, 10-way 10-shot, and 10-way 20-shot. With limited downstream fine-tuning data, our I2P-MAE exhibits competitive performance among existing methods, e.g., +0.5% classification accuracy to Point-M2AE on the 10-way 20-shot split.

4 Additional Visualization

In Figure 10, we additionally visualize the input point cloud, random masking, spatial saliency cloud, 2D-guided masking, and the reconstructed 3D coordinates, respectively. As shown, the 2D-guided masking can preserve the semantically important 3D geometries guided by the spatial sailency cloud (darker points indicate higher scores). In this way, I2P-MAE can inherit more significant 2D knowledge through the 2D-semantic reconstruction of the unmasked visible parts.

References