PointCLIP: Point Cloud Understanding by CLIP
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, Hongsheng Li
Introduction
Deep learning has dominated computer vision tasks of both 2D and 3D domains in recent years, such as image classification , object detection , semantic segmentation , point cloud recognition and part segmentation . With 3D sensing technology developing rapidly, the growing demand for processing 3D point cloud data has boosted many advanced deep models with better local feature aggregator , geometry modeling and projection-based processing . Different from grid-based 2D image data, 3D point clouds suffer from space sparsity and irregular distribution, which hinder direct methods transfer from 2D domain. Additionally, large-scale newly captured point cloud data contain a large number of objects of “unseen” categories to the trained classifier. In this scenario, even the best-performing models might fail to recognize them and it is unaffordable to re-train every time when “unseen” objects arise.
Similar issues have been dramatically mitigated in 2D vision by Contrastive Vision-Language Pre-training (CLIP) , which proposed to learn transferable visual features with natural language supervisions. For zero-shot classification of “unseen” categories, CLIP utilizes the pre-trained correlation between vision and language to conduct open-vocabulary recognition and achieves promising performance. To further enhance the accuracy in few-shot settings, CoOp adopted learnable tokens to encode the text prompts, so that the classifier weights can be adaptively formed. From another perspective, CLIP-Adapter appends a lightweight residual-style adapter with two linear layers for better adapting image features. Tip-Adapter further boosts its performance while greatly reduces the training time. Both methods achieve significant improvements over zero-shot CLIP. Consequently, the problem of recognizing new unlabeled objects has been explored by CLIP in 2D. However, a question is naturally arised: Could such CLIP-based models be transferred to 3D domain and realize zero-shot classification for “unseen” 3D objects?
To address this issue, we propose PointCLIP, which transfers CLIP’s 2D pre-trained knowledge to 3D point cloud understanding. The first concern is to bridge the modal gap between unordered point clouds and the grid-based images that CLIP could process. Considering the need for real-time prediction in various scenarios, such as autonomous driving and indoor navigation , we propose to adopt online perspective projection without any post rendering , i.e., simply projecting each point onto a series of pre-defined image planes to generate scatter depth maps. The cost of this projection process is marginal in both time and computation, but reserves the original property of the point cloud from multiple views. On top of that, we apply CLIP to encode multi-view features of point cloud by the CLIP pre-trained visual encoder and obtain each view’s text-matched prediction independently via zero-shot classifier. Following CLIP, we place 3D category names into a hand-crafted template as prompts and generate the zero-shot classifier by CLIP’s textual encoder. As different views contribute differently to the recognition of entire scene, we obtain the final prediction for point cloud by weighted aggregation between views.
Although PointCLIP achieves cross-modality zero-shot classification without any 3D training, its performance still falls behind classical point cloud networks well-trained on full datasets. To eliminate this gap, we introduce a learnable inter-view adapter with bottleneck linear layers to better extract features from multiple views in few-shot settings. Specifically, we concatenate all views’ features and extract the compact global feature of the point cloud via interacting and summarizing cross-view information. Based on the global representation, adapted feature of each view is generated and added to their original CLIP-encoded feature via a residual connection. In this way, each view is equipped with the fused global feature and also combines newly adapted feature from the 3D few-shot dataset with 2D pre-trained CLIP’s encoding. During training, we only fine-tune this lightweight adapter and freeze CLIP’s both visual and textual encoders to avoid over-fitting, since only a few samples per class are given. Surprisingly, PointCLIP with an inter-view adapter with few-shot fine-tuning achieves comparable performance with some previous models well-trained with full datasets, which is a good trade-off between performance and cost.
Additionally, we observe that CLIP’s 2D knowledge, supervised by contrastive loss, is complementary to the close-set 3D supervisions. The PointCLIP with an inter-view adapter can be fine-tuned under few-shot settings to improve the performance of classical fully-trained 3D networks. Taking PointCLIP in 16-shot ModelNet40 and fully-trained PointNet++ as an example, we directly ensemble their predicted logits for testing. Surprisingly, the performance of PointNet++’s 89.71, is enhanced to 92.03 by PointCLIP with an accuracy of 87.20. Furthermore, we select CurveNet , the state-of-the-art 3D recognition model, as the ensembling baseline, and achieve performance boost from 93.84 to 94.08. In contrast, simply ensembling two models fully trained on ModelNet40 without PointCLIP only leads to performance loss. Therefore, PointCLIP could be regraded as a multi-knowledge ensembling module, which promotes 3D networks via 2D contrastive knowledge with limited additional training.
The contributions of our paper are as follows:
We propose PointCLIP to extend CLIP for handling 3D point cloud data, which achieves cross-modality zero-shot recognition by transferring 2D pre-trained knowledge into 3D.
An inter-view adapter is introduced upon PointCLIP via feature interaction among multiple views and improves the performance of few-shot fine-tuning.
PointCLIP can be utilized as a multi-knowledge ensembling module for enhancing performance of existing fully-trained 3D networks, which surpasses state-of-the-art performances.
Comprehensive experiments are conducted on widely adapted ModelNet10, ModelNet40 and the challenging ScanObjectNN, which indicate PointCLIP’s potential for 3D understanding.
Related Work
The objective of zero-shot learning is to enable recognition of “unseen” objects which are not adopted during training. Although zero-shot learning has drown much attention on 2D classification , only a few works explore how to conduct it in 3D domain. As the first attempt on point cloud, divides the 3D dataset into two parts: “seen” and “unseen” samples, and trains PointNet on the former but tests on the latter by measuring cosine similarities with category semantics. Based on this prior work, further mitigates the hubness problem resulted from low-quality extracted 3D features and introduces a triplet loss for better performance in transductive settings, which allows to utilize unlabeled “unseen” data at training time. Different from all above settings, which train the network on part of the 3D samples and predict on the others, PointCLIP achieves direct zero-shot recognition without any 3D training and conducts prediction on the whole point cloud datasets. Thus, our setting is more challenging for the domain gap between 2D pre-training and 3D application, but more urgent for practical problems.
Transfer Learning.
Transfer learning aims to utilize the knowledge from data-abundant domains to help the learning on data-scarce domains. For general vision, ImageNet pre-training can greatly assist downstream tasks, such as object detection and semantic segmentation . Also in natural language processing, representations pre-trained on web-crawled corpus via Mask Language Model achieves leading performance on machine translation and natural language inference . Without any fine-tuning, the recently introduced CLIP shows superior image understanding ability for “unseen” datasets. CLIP-Adapter , Tip-Adapter , ActionCLIP and WiSE-FT further indicate that the performance of CLIP can be largely improved by infusing domain-specific supervisions. Although the successes stories are encouraging, most of the existing methods conduct knowledge transfer within the same modality, namely, image to image , video to video or language to language . Different from them, our PointCLIP is able to efficiently transfer representations learned from 2D images to the disparate 3D point clouds, which motivates future researches on transfer learning across different modalities.
Deep Neural Networks for Point Cloud.
Existing deep neural networks for point cloud can be divided into point-based and projection-based methods. Point-based models process on raw points without any pre-transformation. PointNet and PointNet++ firstly encode each point with a Multi-layer Perceptron (MLP) and utilize max pooling operation to realize permutation invariance. Recent point-based methods propose more advanced local aggregators and architecture designs . Other than raw points, projection-based methods understand point cloud by transferring it to volumetric or multi-view data forms. Therein, multi-view methods project point cloud into images of multiple views and process them with 2D Convolution Neural Networks (CNN) pre-trained on ImageNet , such as MVCNN and others . Normally, such view-projected methods operate on offline-generated images which are projected from point-converted 3D meshes or required post-rendering for shades and textures, so they are costly and impractical to be adopted for real-time applications. On the contrary, we follow SimpleView , to naively project raw points onto image planes and set their pixel values according to the vertical distance. Such depth-map generation results in marginal time and computation costs, which meets the demand for efficient end-to-end zero-shot recognition.
Method
In Section 3.1, we first revisit Contrastive Vision-Language Pre-training (CLIP) for 2D zero-shot classification. Then in Section 3.2, we introduce our PointCLIP, which transfers 2D pre-trained knowledge into 3D. In Section 3.3, we provide PointCLIP with inter-view adapter for better performance under few-shot settings. In Section 3.4, we propose to ensemble PointCLIP with fully-trained classic 3D networks for multi-knowledge ensembling, which can achieve state-of-the-art performance.
CLIP is trained to match images with their corresponding natural language descriptions. There are two independent encoders in CLIP, respectively for visual and textual features encoding. During training, given a batch of images and texts, CLIP extracts their features and learns to align them in the embedding space with a contrastive loss. To ensure comprehensive learning, 400 million training image-text pairs are collected from the internet, which enables CLIP to align images with any semantic concepts in an open vocabulary for zero-shot classification.
2 Point Cloud Understanding by CLIP
A variety of large-scale datesets in 2D provide abundant samples to pre-train models for high-quality and robust 2D features extraction. In contrast, the widely-adopted 3D datasets are comparatively much smaller and have limited categories, e.g. ModelNet40 with 9,843 samples and 40 classes vs. ImageNet with 1 million samples and 1,000 classes. Thus, it is very difficult to obtain good pre-trained 3D networks for transfer learning. To alleviate this problem and explore the cross-modality power of CLIP, we propose PointCLIP to conduct zero-shot learning on point clouds based on the pre-trained CLIP.
Point cloud is a set of unordered points scattering in the 3D space, its sparsity and distribution greatly differ from grid-based 2D images. To convert point clouds into CLIP-accessible representations, we generate point-projected images from multiple views to eliminate the modal gap between 3D and 2D. In detail, if the coordinate of a point is denoted as in the 3D space, taking the bottom projection view as an example, its location on the image plane is following . In this way, the projected point cloud is a foreshortened figure, namely, small in the distance but big on the contrary, which is more similar to that in real photos. Other than applying convolution layers to pre-processing the one-channel depth map into three, we do not adopt any pre-convolution and directly set the pixel value equaling to in all three channels. Also, different from other off-line projection methods, whose projected images are generated from meshes or CAD models , our projected depth maps are from raw points and contain no color information but scattered depth values, which leads to marginal time and computation cost. With this lightweight cross-modality cohesion, CLIP’s pre-trained knowledge can be then utilized for point cloud understanding.
Zero-shot Classification.
where is a hyper-parameter weighing the importance of view . Each view encodes a different perspective of the point cloud feature, which is capable for independent zero-shot classification. Their summation further complements the information of different perspectives to obtain an overall understanding. The whole process of PointCLIP is non-parametric for the “unseen” 3D dataset, which pairs each point cloud with its category via CLIP’s pre-trained 2D knowledge and without any 3D training.
3 Inter-view Adapter for PointCLIP
Although PointCLIP achieves efficient zero-shot classification on point clouds, its performance is still incomparable to fully-trained 3D neural networks . We then consider a more common scenario where a few objects of each “unseen” category are contained in the newly collected data, and networks are required to recognize them under such few-shot settings. It is impractical to fine-tune the whole model, since the enormous parameters and insufficient samples would easily result in over-fitting. Therefore, referring to in Natural Language Processing (NLP) and CLIP-Adapter for fine-tuning pre-trained models on downstream tasks, we append a three-layer Multi-layer Perceptron (MLP) on top of PointCLIP, named inter-view adapter, to further enhance its performance under few-shot settings. For training, we freeze CLIP’s both visual and textual encoders and fine-tune the learnable adapter via a cross-entropy loss.
After the inter-view adapter, each view conducts classification with the adapted feature and the textual classifier . Same as zero-shot classification, all logits from all views are summarized to construct the final prediction, and the view weights can be learnable parameters here for more adaptive aggregation. Surprisingly, just fine-tuning this lightweight adapter with few-shot samples contributes to significant performance improvement, e.g. from 20.18 to 87.20 on ModelNet40 with 16 samples per category, less than 1/10 of the full data. This inspirational boost demonstrates the effectiveness and importance of feature adaption on 3D few-shot data, which greatly facilitates knowledge transfer from 2D to 3D. Consequently, PointCLIP with inter-view adapter provides a promising alternative solution for point cloud understanding. In some applications, there is no condition to train the entire model with large-scale fully annotated data, and fine-tuning only the three-layer adapter with few-shot data can achieve comparable performance.
4 Multi-knowledge Ensembling
Classical point cloud networks, such as the early PointNet and the recent CurveNet , are trained from scratch on 3D datasets by close-set supervision. In contrast, PointCLIP mostly inherits pre-trained priors from 2D vision-language learning, containing different aspects of knowledge. We then investigate if the two forms of knowledge can be ensembled together for joint inference. In practice, we first obtain the classical model, e.g. PointNet++ pre-trained from , and PointCLIP of either zero-shot or the adapter version. We conduct inferences of the two models and ensemble their predicted logits by simple addition as the final output. Beyond our expectation, aided by 16-shot fine-tuned PointCLIP of 87.20, PointNet++ of 89.71 is enhanced to 92.03 with a significant improvement of . In other words, ensembling of two low-score models can produce a much stronger one, which fully demonstrates the complimentary interaction of knowledge from the two models. Also, even with the zero-shot PointCLIP of 20.18, PointNet++ can still be improved to 92.10. In contrast, ensembling a pair of classical full-trained models would not enhance the performance, which indicates the importance of complimentary knowledge. We also implement this ensembling with other advanced networks and observe similar performance boosts, some of which achieve state-of-the-art performances. Therefore, PointCLIP can be utilized as a plug-and-play enhancement module to achieve robust point cloud understanding.
Experiments
We evaluate the zero-shot classification performance of PointCLIP on three well-known datasets: ModelNet10 , ModelNet40 and ScanObjectNN . For each dataset, we require no training data and adopt the full test set for evaluation. For the pre-trained CLIP model, we adopt ResNet-50 as the visual encoder and transformer as the textual encoder by default. We then project the point cloud from 6 orthogonal views: front, right, back, left, top and bottom, and each view has a relative weight value ranged from 1 to 10, shown in the fourth column of Table 1. As the point coordinates are normalized from -1 to 1, we set the 6 image planes at a fixed distance away from the coordinate center . This distance is shown as the first value of Proj.Settings in Table 1, and the larger distance leads to the denser points distributions on the image. The side length of projected square depth maps varies to different datasets, which is presented as the second value in Proj.Settings, and larger side length results in smaller projected object size. We then upsample all images to for alignment with CLIP’s settings. Also, we set the textual template as “point cloud depth map of a [CLASS].” to cater to the visual features of point clouds.
Performance.
In Table 1, we present performances of zero-shot PointCLIP for three datasets with their best-performing settings. Without any 3D training, PointCLIP is able to achieve a promising 30.23 on ModelNet10, which demonstrates the effective knowledge transfer from 2D to 3D. For ModelNet40 with 4 times the number of categories and ScanObjectNN of noisy real-world scenes, PointCLIP achieves slightly worse performances, 20.18 and 15.38, respectively, due to the lack of 3D downstream adaptions. As for the projection distances and image resolutions of Proj.Settings, their variances accord with the properties of different datasets. Compared to indoor ModelNet10, PointCLIP on ModelNet40 requires more details to recognize complex outdoor objects, such as airplanes and plants, and thus performs better with more scattered points and larger object size, namely, larger perspective projection distance and resolutions. In contrast, for ScanObjectNN, denser points and larger resolutions are required for filtering out the noise and reserving complex real-scene information. With respect to view weights, ModelNet10 and ModelNet40 of synthetic objects require all 6 views’ contributions to the final classification with different importance, but for ScanObjectNN which contains noisy points of floors and ceilings, the top and bottom views could hardly provide any information.
Ablations.
In Table 2, We conduct ablation studies of zero-shot PointCLIP concerning projection view numbers and the importance of each view on ModelNet40. For the number of projected views, we try 1, 4, 6, 8, 10 and 12The settings of views are in the Appendix. views, for increasingly capturing the multi-view information of point clouds, but more than 6 views would bring redundancy and lead to performance decay. To explore how different views impact the performance, we unify all relative weights to 3 and respectively increase each view’s weight to 9. As is shown in the table, projection from the right achieves the highest performance, which indicates its leading role, and the top and down views contribute relatively less to the zero-shot classification. In Table 4, we implement different visual backbones from ResNet to vision transformer , and RN5016 achieves the best performance of 23.78, which has 16 times more computations than ResNet-50. However, upgrading ResNet-50 to ResNet-101 with more parameters and deeper layers would not provide higher classification accuracy.
Prompt Design.
We present five prompt designs for zero-shot PointCLIP in Table 3. We observe that the naive “a photo of a [CLASS].” achieves 17.02 on ModelNet40, but simply inserting the word “point cloud” into it would hurt the performance. We then remove “a photo” and directly utilize “point cloud” as the subject, which benefits the accuracy by +1.66. Also, as the projected point cloud normally covers most of the image area, appending an adjective “big” could bring further performance improvement. Furthermore, we add the “depth map” to describe the projected images more relevantly, which contributes to the best-performing 20.18, demonstrating the importance of prompt choices.
2 Few-shot Classification
We experiment PointCLIP with the inter-view adapter under 1, 2, 4, 8, 16 shots also in the three datasets: ModelNet10 , ModelNet40 and ScanObjectNN . For -shot settings, we randomly sample point clouds from each category of the training set. We inherit the best projection settings from zero-shot experiments in Section 4.1. In contrast, considering both efficiency and performance, we adopt ResNet-101 as CLIP’s pre-trained visual encoder for stronger feature extraction, and increase the projected view numbers to 10, adding the views of upper/bottom-front/back-left corners, since the left view is proven to be the most informative for few-shot recognition in Table 2. In addition, we modify the prompt to “point cloud of a big [CLASS].”, which performs better in the few-shot experiments. For the inter-view adapter, we construct a residual-style Multi-layer Perceptron (MLP) consisting of three linear layers, as described in Section 3.3.
Performance.
In Figure 5, we present the few-shot performances of PointCLIP and compare it with 4 representative 3D networks: PointNet , PointNet++ , SimpleView and the state-of-the-art CurveNet . As we can see, PointCLIP with inter-view adapter surpasses all other methods for the few-shot classification. When there are only a small number of samples per category, PointCLIP has distinct advantages, exceeding PointNet by 25.49 and CurveNet by 12.29 on ModelNet40 with 1 shot. When given more training samples, PointCLIP still leads the performance, but the gap becomes smaller due to the limited fitting capacity of the lightweight three-layer adapter. For the detailed training settings, please refer to the Appendix.
Ablations.
In Table 2, we show the 16-shot PointCLIP under different projection views and explore how each view contributes on ModelNet40. Differing from the zero-shot version, 10 views of 16-shot PointCLIP performs better than 6 views, probably because the newly-added adapter is able to better utilize the information from more views and adaptively aggregate them. For the importance of views, we follow the configurations of our zero-shot version and observe the reversed conclusion that, the left view is the most informative here. Surprisingly, for different visual encoders in Table 4, ResNet-101 achieves the highest accuracy with less parameters than vision transformer or ResNet-5016. Table 3 lists the performance influence caused by prompt designs, and the “point cloud of a big [CLASS].” performs the best, which is slightly different from the analysis in Paragraph 4.1.
3 Multi-knowledge Ensembling
To verify the complementarity of blending pre-trained 2D priors with 3D knowledge, we aggregate the fine-tuned 16-shot PointCLIP of 87.20 on ModelNet40, respectively with fully-trained PointNet , PointNet++ , DGCNN , SimpleView and CurveNet , whose trained models are obtained from without any voting. We manually modulate the fusion ratio between PointCLIP and each model, and report the performance with the best Ratio in Table 5, which represents PointCLIP’s relative weight to the whole.
Performance.
As shown in Table 5, ensembling with PointCLIP improves the performances of all classical fully-trained 3D networks. The results fully demonstrate the complementarity of PointCLIP to existing fully-trained 3D models, and the performance gain is not simply achieved by ensembling models. These are surprising results to us, because the accuracy of 16-shot PointCLIP is lower than all other models trained with full datasets, but could still benefit their already high performances to be higher. Therein, the largest accuracy improvement is on PointNet++ from 89.71 to 92.10, and combining PointCLIP with the state-of-the-art CurveNet further achieves 94.08. Also, we observe that, for models with low baseline performances, PointCLIP’s logits need to account for a large proportion, but for the well-performing ones, such as CurveNet, their knowledge is supposed to play a dominant role in the ensembling.
Ablations.
We conduct ablation studies of ensembling two models fully trained on ModelNet40 without PointCLIP, and fuse their logits with the same ratio for simplicity. As is presented in Table 6, ensembling PointNet++ lowers the performance of RSCNN and CurveNet, and aggregating the highest two models, SimpleView and CurveNet, could not achieve better performance. Also, a pair of PointCLIP would hurt the performance. Hence, simply ensembling two models with the same training scheme normally leads to performance degradation, which demonstrates the significance of multi-knowledge interaction. In Table 7, we fuse zero-shot PointCLIP and the model fine-tuned by 8, 16, 32, 64 and 128 shots, respectively with CurveNet to explore their ensembling performances. As reported, zero-shot PointCLIP with only 20.18 could enhance CurveNet by +0.04. However, too much training on 3D dataset would adversely influence the ensembling accuracy. This is possibly caused by the too high similarity between two models, which cannot provide complementary knowledge as expected.
Conclusion and Limitation
We propose PointCLIP, which conducts cross-modality zero-shot recognition on point cloud without any 3D training. Via multi-view projection, PointCLIP efficiently transfers CLIP’s pre-trained 2D knowledge into the 3D domain. Under few-shot settings, we design a lightweight inter-view adapter to aggregate multi-view representations and generate adapted features. By fine-tuning such adapter and freezing all other modules, the performance of PointCLIP is largely improved. In addition, PointCLIP could serve as a plug-and-play module to provide complimentary information for the classical 3D networks, which surpasses state-of-the-art performance. Although PointCLIP realizes the transfer learning from 2D to 3D, how to utilize CLIP’s knowledge for other 3D tasks is still under explored. Our future work will focus on generalizing CLIP for wider 3D applications.
References
Appendix
Appendix A Datasets
We evaluate our PointCLIP on three well-known datasets: ModelNet10 , ModelNet40 and ScanObjectNN . Therein, ModelNet10 consists of 4,899 synthetic meshed CAD models with 10 indoor categories, 3,991 for training and 908 for testing. ModelNet40 is larger and contains 12,311 samples of 40 common categories, 9,843 for training and 2,468 for testing. In both datasets, we uniformly sample 1,024 points from each object as the network input. ScanObjectNN contains 2,321 training and 581 testing point clouds of 15 categories collected directly from real-world scans. Different from synthetic data with complete profiles, objects in ScanObjectNN are occluded at different levels and disturbed with background noise, so it is more challenging for accurate recognition.
Appendix B Implementation Details
For ablation studies of projected view numbers, we adopt different settings for zero-shot and few-shot PointCLIP. As the right view is the most important for zero-shot PointCLIP, we set the 12 views to: front, right, back, left, top, bottom, upper/lower right diagonal front/back (4 views) and upper left diagonal front/back (2 views). In contrast, few-shot PointCLIP achieves higher performance with left views, so we replace all the “left” settings above into “right”. For both versions, the view number of represents picking the first views for experiments.
For PointCLIP with inter-view adapter, we fine-tune it under 1, 2, 4, 8 and 16 shots with batch size 32 and learning rate 0.01 for 250 epochs. Stochastic Gradient Decent (SGD) with momentum 0.9 is adopted as the optimizer. We utilize a cosine scheduler for learning rate decay and Smooth Loss following . In ModelNet10 and ModelNet40, We apply random scaling and translation for training augmentation, but in the challenging ScanObjectNN, we append jitter and random rotation following . During training, we freeze CLIP’s both visual and textual encoders, and only fine-tune the inter-view adapter. For other compared models, we unfreeze all the parameters, and adopt the same data augmentation and loss functions reported in the papers.
Appendix C Supplementary Ablations
We adopt the inter-view adapter with three linear layers: one for global extraction and two for view-wise adapted features generation. Here, we explore other architectures of the adapter on 16-shot PointCLIP for ModelNet40 in Table 8. Specifically, w/o global denotes the adapter processing each view separately without interaction, and the w/o view-wise version repeats the global feature as each view’s adapted feature. The 2-layer adapter removes the linear layer after the global representation and the pre-layer version moves it before the global extraction. The results show that dropping or changing the original modules in the adapter would all hurt the performance, especially the inter-view extraction of global feature.
Adapted Features Fusion.
The view-wise adapted feature is generated by the adapter and then added to the original CLIP-encoded feature via a residual connection. On ModelNet40, we evaluate the performance of 16-shot PointCLIP with different fusion ratios , which denotes the proportion of adapted features. To show the effect of , we set all view weights the same. From the results in Table 9, different ratios actually lead to little performance variance and the of 0.6 perfoms better than others. Thus, we adopt 0.6 as the fusion ratio by default, which indicates the comparable contributions between 2D pre-trained knowledge and 3D learned knowledge.
Full Training Set.
We also fine-tune PointCLIP on full training set of ModelNet40 and present the results in Table 10. Likewise, we freeze both pre-trained visual and textual encoders in CLIP and only train the inter-view adapter. As expected, visual encoders with more parameters lead to higher accuracy, and only fine-tuning the appended lightweight adapter could achieve the performance of 92.01.
Fine-tuning Settings.
Under full training set of ModelNet40 , we further fine-tune different modules of PointCLIP in Table 11. Therein, we adopt ResNet-101 as the visual encoder, and fine-tuning without the adapter represent unfreezing the visual or textual encoder upon the zero-shot PointCLIP. As presented, unfreezing just the textual encoder normally hurts the performance, but training both encoders and all modules of PointCLIP achieves better performance of 91.40 and 91.89, respectively.
Appendix D Visualizations
We visualize some cases of ensembling PointCLIP with PointNet++ to reveal the effectiveness of enhancement. As shown in Figure 6, two models both predict correctly for the four samples, and the ensembled model preserves the prediction. As for samples in the second and the third rows, PointCLIP and PointNet++ show the complementary properties that the ensembled model would rectify one of their wrong predictions, which demonstrates the importance of knowledge interaction.