Image2Point: 3D Point-Cloud Understanding with 2D Image Pretrained Models
Chenfeng Xu, Shijia Yang, Tomer Galanti, Bichen Wu, Xiangyu Yue, Bohan Zhai, Wei Zhan, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka
Introduction
Point-cloud is an important visual representation for 3D computer vision. It is widely used in a variety of applications, including autonomous driving behley2019iccv ; caesar2020nuscenes ; yue2018lidar , robotics armeni2017joint ; pomerleau2015review ; xu2021you , augmented and virtual reality 3dwarehouse ; Wu_2015_CVPR ; 7273863 , etc. However, a point-cloud represents visual information in a significantly different way from a 2D image. Specifically, a point-cloud consists of a set of unordered points lying on the object’s surface, with each point encoding its spatial coordinates and potentially other features such as intensity. In contrast, a 2D image organizes visual features as a dense 2D RGB pixel array. Due to the representation differences, 2D image and 3D point-cloud understanding are treated as two separate problems. 2D image models and point-cloud models are designed to have different architectures and are trained on different types of data. No efforts have tried to directly transfer models from images to point-clouds.
Intuitively, both 3D point-clouds and 2D images are visual representations of the physical world. Their low-level representations are drastically different, but they can represent the same underlying visual concept. Furthermore, human vision has no problem understanding both representations. To connect images and point-clouds, previous works attempted to generate pseudo point-clouds by estimating the depth of mono/stereo images wang2019pseudo ; gur2019single ; yin2021virtual . However, depth estimation from a single image is a challenging problem in computer vision, which requires large-scale dense depth labels Ranftl2020 . Estimating depth from stereo images is easier but requires strict calibrated and synchronized stereo cameras, which limits the data scale. Therefore, it is interesting to ask whether we could use large-scale image models that were pretrained using supervised classification datasets (e.g., ImageNet1K/ImageNet21K classification) for point-cloud understanding.
Remarkably, the answer to the question above is positive. As we show in this work, 2D image models trained on image datasets can be transferred to understand 3D point-clouds with minimal effort. As illustrated in Figure 1, given the commonly-used image-pretrained models, such as 2D ConvNets he2016deep and vision transformers dosovitskiy2020image , we can easily convert them into various kinds of point-cloud models. In particular, a pretrained 2D ConvNet and vision transformer can be easily extended into projection-based, voxel-based, and transformer-based point-cloud models via copying weights or inflating weights carreira2017quo .
In this paper, we primarily focus on 3D ConvNets inflated from 2D pre-trained models. With the transformed point-cloud model (e.g., inflated 3D ConvNets), we add linear input and output layers to the network; and on a target point-cloud dataset, we only finetune the input and output layers, and batch normalization layers, while keeping the pretrained model weights untouched. We call such partially-finetuned-image-pretrained models as FIP-IO+BN (finetuning input, output, and BN layers). As we show, FIP-IO+BN can achieve competitive performance up to 90.8% top-1 accuracy on the ModelNet 3D Warehouse dataset, on top of ResNet50, outperforming previous point-cloud models that adopt task-specific model architectures and tricks.
Even though incorporating pretrained models is useful for tackling downstream tasks, point-cloud models are typically trained from scratch. Based on our discovery, we further investigate fully-finetuned-image-pretrained models (termed as FIP-ALL). We observe that FIP-ALL brings significant improvement on top of different kinds of point-cloud models transformed from image-pretrained models. Besides, we also find that it generalizes to PointNet++ qi2017pointnetplusplus which is pre-trained on images by ourselves. Specifically, FIP-ALL outperforms the training-from-scratch by a large margin on top of PointNet++, SimpleView, ViT-B-16, and ViT-L-16, respectively. In addition to the performance gain, FIP-ALL exhibits superior data efficiency with up to accuracy improvement in few-shot classification on the ModelNet 3D Warehouse dataset. Compared with training-from-scratch, FIP-ALL also dramatically speeds up the training by using 11.1 times fewer epochs to reach the same validation accuracy (e.g., 90% accuracy).
Finally, we theoretically explore the relationship between transferring knowledge between tasks of different modalities and neural collapse to shed light on why the transfer works. The analysis is based on extending the framework proposed in galanti2022on and is provided in Appendix C.
Related Work
In this section, we list the most prominent approaches for processing point-clouds.
The 3D convolution-based method is one of the mainstream point-cloud processing approaches which efficiently processes point-clouds based on voxelization. In this approach, voxelization is used to rasterize point-clouds into regular grids (called voxels). Then, we can apply 3D convolutions to the processed point-cloud. However, enamors empty voxels make lots of unnecessary computations. Sparse convolution is proposed to apply on the non-empty voxels liu2015sparse ; choy20194d ; tang2020searching ; zhou2020cylinder3d ; yan2018second ; feng2021simple , largely improving the efficiency of 3D convolutions.
The projection-based method attempts to project a 3D point-cloud to a 2D plane and uses 2D convolution to extract features wang2018fusing ; wu2017squeezeseg ; wu2018squeezesegv2 ; xu2020squeezesegv3 ; su2015multi ; lawin2017deep ; boulch2017unstructured . Specifically, bird-eye-view projection yang2018pixor ; lang2019pointpillars and spherical projection wu2017squeezeseg ; wu2018squeezesegv2 ; xu2020squeezesegv3 ; milioto2019rangenet++ have made great progress in outdoor point-cloud tasks.
Another approach is the point-based method, which directly processes the point-cloud data. The most classic methods, PointNet qi2016pointnet and PointNet++ qi2017pointnetplusplus , consume points by sampling the center points, group the nearby points, and aggregate the local features. Many works further develop advanced local-feature aggregation operators that mimic the 3D convolution operation to structured data xu2021you ; li2018pointcnn ; hua2018pointwise ; liu2019densepoint ; liu2020closer ; wang2017cnn ; li2018so ; komarichev2019cnn .
2 Pretraining in 2D and 3D Computer Vision
Pretraining in 2D computer vision is an effective approach using supervised dosovitskiy2020image ; girshick2014rich , self-supervised jing2020self ; goyal2021self , and contrastive learning he2020momentum ; bachman2019learning ; chen2020simple ; caron2020unsupervised ; chen2020improved ; hjelm2018learning . After pretraining on a large amount of data, a 2D model requires less computational and data resources for finetuning in order to obtain competitive performance on downstream tasks kataoka2020pre ; caron2019unsupervised ; chen2020big ; henaff2020data .
Pretraining in 3D computer vision has been studied similarly as pretraining in 2D vision: both self-supervised and contrastive pretraining xie2020pointcontrast ; hou2021exploring ; wang2020unsupervised show promising results. 3D point-clouds are difficult to annotate, and there is no large-scale annotated dataset available. To address this, previous works have tried to use model pretraining to improve data efficiency xu2020weakly . Recent works hou2020exploring ; zhang2021self explored using contrastive learning on point-clouds. Our work does not rely on long-time pretraining. Instead, we can directly take large amounts of open-sourced image-pretrained models for a variety of point-cloud tasks.
3 Cross-Modal Transfer Learning
Cross-modal transfer learning takes advantage of data from various modalities dai20183dmv ; liu20213d . For example, liu2021learning proposed pixel-to-point knowledge transfer (PPKT) from 2D to 3D which uses aligned RGB and RGB-D images during pretraining. Our work does not rely on joint image-point-cloud pretraining. Instead, we directly transfer an image-pretrained model to a point-cloud model with the simplest pretraining-finetuning scheme.
Some of the previous works for video and medical images carreira2017quo ; shan20183 have adopted the method of simply extending a pretrained 2D convolutional filter along time or depth direction for transferring to 3D models. However, the domain gaps between point-clouds and images are much more than that of videos/medical images and images. Between language and image modalities, transfer learning with minimal finetuning also shows a competitive performance lu2021pretrained ; radford2021learning .
4 Neural Collapse
Neural collapse (NC) Papyan24652 ; han2021neural is a recently discovered phenomenon in deep learning. It has been observed that during the training of deep overparameterized neural networks for standard classification tasks, the penultimate layer’s features associated with training samples belonging to the same class concentrate around their class means. Essentially, Papyan24652 observed that the ratio of the within-class variances and the distances between the class means converge to zero. In addition to that, it has also been observed that asymptotically the class means (centered at their global mean) are not only linearly separable, but are also maximally distant and located on a sphere centered at the origin up to scaling, and furthermore, that the behavior of the last-layer classifier (operating on the features) converges to that of the nearest-class-mean decision rule.
Recently, galanti2022on studied the relationship between neural collapse and transfer learning. They studied a transfer learning setting, where we intend to solve a target (classification) task, where only a limited amount of samples is available, so a model is pretrained and transferred from a source (classification) task. They showed that neural collapse extends beyond training and generalizes also to unseen test samples and new classes. In addition, it was shown that in the presence of neural collapse in the new classes, training a linear classifier on top of the learned penultimate layer requires only a few samples to generalize well. However, their empirical and theoretical analysis assumes that the source and target classes are i.i.d. samples (e.g., a random split of the classes in ImageNet). This implies that the two tasks share the same modality. Therefore, we suggest training an adaptor (e.g., a linear layer) along with retraining the normalization parameters as part of the transfer process. Intuitively, the adaptor takes samples of the second modality and translates them to representations that are interpretable by the pretrained model, such that it produces feature embeddings that are clustered into classes. In Appendix C, we extend the framework in galanti2022on to the case where the source and target tasks are of different modalities and theoretically analyze it.
Converting a 2D Image Model to a 3D Point-Cloud Model
In this paper, we primarily focus on the 3D sparse-convolution-based method to process point-clouds, since it can be extended to a wide range of point-cloud tasks. The other point-cloud models we use in this paper are byproducts of copying the weights of 2D image models, for example, 2D ConvNets he2016deep or vision transformers dosovitskiy2020image . In this section, we provide an in-depth introduction to how we transform the 2D ConvNets into 3D sparse ConvNets by inflation carreira2017quo .
As discussed in Section 2.1, we consider a set of points, where each point is represented by its 3D coordinates and additional features such as its intensity and RGB. We then voxelize/quantize these points into voxels according to their 3D space coordinates, following choy20194d . A voxel’s feature is inherited from the point that lies within the voxel. If there are multiple points associated with the same voxel, we average all points’ features and assign the mean to the voxel. If there is no point in the voxel, then we simply set the voxel’s feature to 0. With sparse convolution, the computation on empty voxels can be skipped.
Given a pretrained 2D ConvNet, we convert it to a 3D ConvNet that takes 3D voxels as input. The key element of this procedure is to convert 2D convolution filters to 3D, i.e., constructing 3D filters with the weights directly inherited from 2D filters. A 2D convolutional filter can be represented with a 4D tensor of shape , representing output dimension, input dimension, and two spatial kernel sizes, respectively. A 3D convolutional filter has an extra dimension, and its shape is . To better illustrate, we ignore the output and input dimensions and only consider a spatial slice of the 2D filter with shape . The simplest way to convert this 2D filter to 3D is to repeat the 2D filter times along a third dimension. This operation is the same as the inflation technique used by carreira2017quo to initialize a video model with a pretrained 2D ConvNet.
Besides convolution, other operations such as downsampling, BN, and nonlinear activation can be easily migrated to 3D. Our 3D model inherits the architecture of the original 2D ConvNet, but we also add a linear layer as the input layer and an output layer depending on the target task. For classification, we use a global average pooling layer followed by one fully connected layer to get the final prediction. For semantic segmentation, the output layer is a U-Net style decoder ronneberger2015u . The architecture of the input/output layers is described in more detail in Appendix B.6.
A note on image-to-video transfer.
It is noteworthy to mention that although inflation is commonly used in video domains, image-to-point-cloud transfer is fundamentally different from image-to-video transfer. Even though videos and point-clouds are both 3D data, they are represented with completely different visual modalities with different distributions. Intrinsically, 3D point-clouds are represented as a sparse set of points lying on object surfaces and parameterized by -coordinates, while videos are dense RGB arrays, where the two spatial arrays represent RGB images and the temporal array reflects how images evolve through time. Point-clouds are translation and rotation invariant or equi-variant, while for videos, the spatial and temporal dimensions are not interchangeable. In this paper, we surprisingly find that with simple operations such as inflation, the image-pretrained models can be directly used for point-cloud understanding under the situation that image and point-cloud are drastically different. The detailed experiments showing the feasibility and utility, and the discussion of why it works from the aspect of neural collapse are illustrated in Section 4 and Section 5, respectively.
Empirical Evaluation
To explore the image to point-cloud transfer, we study three settings: , (1) finetuning input, output, and batch normalization layers (FIP-IO+BN), (2) finetuning the whole pretrained network (FIP-ALL), and optionally (3) partially-finetuned-image-pretrained model, only finetuning input and output layers (FIP-IO). Under the three settings, we extensively explore the feasibility of transferring the image-pretrained model for point-cloud understanding and its benefits. The entire empirical evaluation is organized as four questions: (1) Can we transfer pretrained-image models to recognize point-clouds? (Section 4.1) (2) Can image-pretraining benefit the performance of point-cloud recognition? (Section 4.2) (3) Can image-pretrained models improve the data efficiency on point-cloud recognition? (Section 4.3) (4) Can image-pretrained models accelerate training point-cloud models? (Section 4.4)
We evaluate the transferred models on ModelNet 3D Warehouse classification Wu_2015_CVPR , S3DIS indoor segmentation armeni2017joint , and SemanticKITTI outdoor segmentation behley2019iccv tasks. ModelNet 3D Warehouse is a CAD model classification dataset that consists of point-clouds with 40 categories. CAD models in this benchmark come from 3D Warehouse 3dwarehouse . In this benchmark, we only utilize coordinates as features. S3DIS is a dataset collected from real-world indoor scenes and includes 3D scans of Matterport Scanners from 6 areas. It provides point-wise annotations for indoor objects like chair, table, and bookshelf, etc. SemanticKITTI dataset from KITTI Vision Odometry geiger2012cvpr is a driving scene dataset. It provides dense point-wise annotations for the complete 360 degrees field-of-view of the deployed automotive lidar, which is currently one of the most challenging datasets.
ResNet he2016deep series is used mostly throughout our experiments. Depending on the experiments, ResNets are pretrained on Tiny-ImageNet, ImageNet-1K, ImageNet-21K deng2009imagenet , and Fractal database (FractalDB) kataoka2020pre . Our pretrained models are directly downloaded from various sources, with detailed links provided in the Appendix A. To study the benefits of using pretrained image models, we also utilize PointNet++ qi2017pointnetplusplus , ViT dosovitskiy2020image , and SimpleView goyal2021revisiting as our baselines.
1 Can we transfer pretrained-image models to recognize point-clouds?
To evaluate the feasibility of transferring pretrained 2D image models to 3D point-cloud tasks, we conduct experiments on top of the ResNet series since there are abundant open-source pretrained ResNet available. In particular, we convert 2D ConvNets into 3D ConvNets using the procedure described in Section 3. We hypothesize that, if a pretrained 2D image model is capable of understanding point-clouds directly, we can see a non-trivial performance by only finetuning input and output layers of the transferred model. Further, as we gradually relax the frozen parameters, finetuning BN parameters as well, the transferred model can achieve better performance, even surpassing training-from-scratch.
We conduct two groups of experiments with FIP-IO and FIP-IO+BN, with the results shown in Figure 2. The first is to evaluate the performance as the trainable parameters gradually increase. As shown in Figure 2 (a), training no more than 0.3 % (345.5x fewer) of the whole parameters, the image pretraining even beats the training-from-scratch (100 % trainable parameters). Specifically, ResNet152 FIP-IO+BN with ImageNet1K pretraining improves training-from-scratch by 0.16 points, and ResNet50 FIP-IO+BN with ImageNet21K pretraining improves 0.48 points. Meanwhile, FIP-IO reaches a non-trivial performance. ResNet50 FIP-IO pretrained on ImageNet1K achieves 81.20 % top-1 accuracy, only 9.12 points worse than training-from-scratch with approximately 0.1 % trainable parameters.
Furthermore, to investigate the effect of different datasets, as shown in the right figure of Figure 2, we inflate ResNet50 pretrained from different image datasets, including Tiny-ImageNet, ImageNet1K, ImageNet21K, FractalDB1K, and FractalDB10K, then evaluate on the ModelNet 3D Warehouse.
We discover that, even if we only finetune the input and output layers while keeping the image-pretrained weights frozen, the FIP-IO pretrained from ImageNet1K, FractalDB1K, and FractalDB10K achieves competitive performance. Specifically, ResNet50 FIP-IO with ImageNet1K pretraining outperforms 3D ShapeNet Wu_2015_CVPR and DeepPano 7273863 , which were the state-of-the-arts in 2015, by 4.2 and 3.6 points respectively in top-1 accuracy on ModelNet 3D Warehouse. More importantly, with ImageNet21K pretrained model, ResNet50 FIP-IO+BN surpasses training-from-scratch by 0.48 points, even beating a variety of well-known methods including PointNet qi2016pointnet , MVCNN su2015multi , DGCNN wang2019dynamic , etc.
Notably, we find out the answer to "Can we transfer pretrained-image models to recognize point-clouds?": Yes. The pretrained 2D image models can be directly used for recognizing point-clouds. Surprisingly, the pretraining dataset is not restricted to natural but also synthetic images like those in FractalDB1K/10K.
2 Can image-pretraining benefit point-cloud recognition?
From the previous subsection, we find unexpectedly that the image-pretrained model can be directly used for point-cloud understanding. In this subsection, we investigate whether the image-pretrained model is helpful to improve the performance of point-cloud tasks. We use different baselines, including voxelization-based method (simply ResNet), point-based method (PointNet++ qi2017pointnetplusplus ), projection-based method (SimpleView goyal2021revisiting ), and current popular transformer-based method (ViT-B-16 and ViT-L-16 dosovitskiy2020image ), and fully finetune them on three point-cloud datasets: classification on ModelNet 3D Warehouse, indoor scene segmentation on S3DIS, and outdoor scene segmentation on SemanticKITTI, as shown in Table 1 and Table 2.
For PointNet++, we use ImageNet1K to pretrain: we break each image into pixels and regard it as a point-cloud. For ViT, we directly use the open-source pretrained model and finetune it on ModelNet 3D Warehouse. All the implementation details are illustrated in Appendix A.
Table 1 presents performance on ModelNet 3D Warehouse dataset. We observe that FIP-ALL improves all baselines steadily and significantly. Besides, pretraining brings more improvements to deeper models. For example, ResNet18 can only be improved by 0.13% top-1 accuracy, but pretraining on ImageNet1K leads to 0.81 points top-1 accuracy improvement on top of ResNet152. Moreover, larger pretrained datasets also lead to better performance. Specifically, ResNet50 FIP-ALL from ImageNet21K can reach 91.05% top-1 acc, with 0.73 points improvement over training-from-scratch. Such FIP-ALL significantly outperforms a series of well-known methods such as qi2016pointnet ; qi2017pointnetplusplus ; klokov2017escape ; wang2019dynamic ; su2015multi ; li2018so .
We also explore FIP-ALL on different architectures, as shown in the second group of Table 1. In particular, FIP-ALL on top of PointNet++, ViT-B-16, ViT-L-16, and SimpleView with image dataset pretraining improve the training-from-scratch by 0.88, 3.50, 4.18, 0.50 points, respectively. Especially for the current superior baseline in image recognition, ViT-B-16 and ViT-L-16, the improved performance is quite significant, revealing the huge potential of using image-pretrained models for point cloud recognition.
For the challenging indoor and outdoor scene segmentation, using ImageNet1K pretrained models (FIP-ALL on ImageNet1K) also improve the training-from-scratch consistently, as shown in Table 2. PointNet++ (resp. ResNet18) pretrained on ImageNet1K outperforms the training-from-scratch by 2.56 points (resp. 1.53 points) mIoU on S3DIS dataset. For SemanticKITTI, we utilize the commonly used projection-based method with 2D ConvNet HRNet. With ImageNet1K pretraining, we observe 3.41 points mIoU improvement, a large margin in such a challenging task. Since HRNetV2-W48 has rich pretrained models, we finetune Cityscapes pretrained HRNetV2-W48 and observe this enhances more (5.25% mIoU improvement over training from scratch). Even for the ResNet18 with a high from-scratch performance of 64.75% mIoU, the ImageNet1K pretraining can also bring 0.82 points mIoU improvement.
Finally, we compare the performance gain with the well-known point-cloud self-supervised method PointContrast xie2020pointcontrast , as presented in Table 3. We use the same model architecture and finetuning recipe, and the only difference is the pretraining weights. Note that the model architecture used in PointContrast does not have corresponding open-sourced image-pretrained weights, so we pretrain it by ourselves on ImageNet1K, with the standard ImageNet training recipe provided by Pytorch. We can observe that image-pretraining on ImageNet1K significantly boosts the training-from-scratch by 0.93 points, surpassing the PointContrast by at least 0.64 points.
Therefore, the answer to "Can image-pretraining benefit point-cloud recognition" is: Yes. Image-pretraining can indeed improve point-cloud recognition, which can generalize to a wide range of backbones and benefit a variety of challenging tasks.
3 Can image-pretrained models improve the data efficiency on point-cloud recognition?
Data efficiency is extremely important in point-cloud understanding due to the huge labor of collecting and annotating point-cloud data. In this subsection, we investigate whether the image-pretrained model can help to improve the data efficiency by conducting few-shot setting experiments, including 1-shot, 5-shot, and 10-shot.
In detail, for each class (ModelNet 3D Warehouse involves 40 classes), we randomly choose a few point-clouds as training data and still evaluate on the whole test set. We compare the results between training-from-scratch and FIP-ALL pretrained on the ImageNet1K dataset. The experimental results are shown in Table 4. We observe that FIP-ALL dramatically surpasses training-from-scratch on the low data regime (1-shot): pretraining on ImageNet1K brings 10.0, 6.0, and 9.9 points top-1 accuracy improvement for ResNet18, ResNet50, and ResNet152, respectively. For 5-shot and 10-shot settings, using ImageNet1K pretraining can still consistently improve the performance.
Furthermore, inspired by previous work chen2020big which proposed big self-supervised models are strong semi-supervised learners in 2D image recognition, we borrow the idea and propose an image-pretrained model is also a strong semi-supervised learner in point-cloud recognition. We also compared the image-pretrained model with the self-supervised pretrained model in this experiment. Specifically, we first take pretrained models from the previous self-supervised pretraining method PointContrast xie2020pointcontrast . PointContrast provides two ScanNet dai2017scannet pretrained models of architecture ResNet34 trained with hardest-contrastive loss and PointInfoNCE loss. Then, we finetune PointContrast on 1/5/10 shot of the labeled ModelNet 3D Warehouse dataset and regard it as a teacher model. Finally, we distill the teacher model to a randomly initialized student model. In detail, we pass in the rest of unlabeled ModelNet 3D Warehouse dataset and 1/5/10 shot of the labeled dataset into the teacher model to generate pseudo labels. We use softmax MSE loss as consistency loss between student model outputs and pseudo labels. When the data instance is labeled, we add an additional cross entropy loss as a class criterion between student output and the label.
To show the effectiveness of the image-pretrained model, we repeat the above experiment, only replacing self-supervised pretrained models with ResNet34 ImageNet1K pretrained models. Results are reported in Table 5. We observe that image-pretrained ResNet34 consistently outperforms PointContrast, and improves the baseline by a large margin with 11.9, 4.1, and 2.7 points on 1-shot, 5-shot, and 10-shot, respectively. The results in Table 5 show that an image-pretrained model is indeed a strong semi-supervised learner in point-cloud recognition.
However, in both Table 4 and Table 5, we observe that as the amount of training data increases, the performance gain becomes saturated. Therefore, our answer to "Can image-pretrained models improve the data efficiency on point-cloud recognition?" is: Yes. Image-pretrained models can improve the data efficiency on point-cloud recognition, especially on low data regime. When the training data increases, performance still improves, but the gain becomes marginal.
4 Can image-pretrained models accelerate point-cloud training?
We also investigate whether the image-pretrained model can accelerate training on the point-cloud domains. The results are shown in Fig. 3.
We discover that, after training only one epoch on ModelNet 3D Warehouse dataset, FIP-ALL pretrained on ImageNet1K achieves very impressive performance, yet the performance of training-from-scratch is still very low. For instance, after the first epoch, ResNet50 (resp. ResNet152) with training from scratch achieves 28.48% (resp. 13.94%) top-1 accuracy while ResNet50 (resp. ResNet152) with ImageNet1K pretraining reaches 80.11% (resp. 79.34%) top-1 accuracy. Moreover, to reach 90% top-1 accuracy, a non-trivial performance, FIP-ALL significantly accelerates the training by 2.14x (28 vs. 60 epoch), 11.1x (11 vs. 122 epoch), 2.95x (19 vs. 56 epoch) over training-from-scratch, on top of ResNet18, ResNet50, and ResNet152, respectively.
Therefore, our answer to “Can image-pretrained models accelerate point-cloud training?” is still positive. The image-pretrained models can significantly accelerate the training speed of point-cloud tasks.
Neural Collapse in Cross-Modal Transfer
In this section, we provide an explanation of why the image to point-cloud transfer works based on the recently observed phenomenon called neural collapse han2021neural ; Papyan24652 . galanti2022on in depth studied the relationship between neural collapse and transfer learning between two classification tasks of the same modality (image domain). Similar to this work, we focus on transferring pretrained models between domains of different modalities, i.e., from images to point-clouds.
As illustrated in Section 4, we can transfer models that were pretrained on images to the point-cloud domain. This motivates us to question whether the phenomenon of neural collapse generalization galanti2022on (see Section 2.4) is also evident in our case. Following galanti2022on , we explore the relationships between neural collapse and image-to-point transfer by calculating the class-distance normalized variance (CDNV). Informally, the CDNV measures the ratio between the within-class variances of the embeddings and the squared distance of their means (see Appendix B.5 for details). We measure the CDNV of the fine-tuned model on both train and test data of the point-cloud domain. Since neural collapse is essentially a clustering property of features learned by neural networks, we further examine the neural collapse using tSNE visualizations. The results are summarized in Fig. 4.
We observe that with finetuning much fewer (345.5x fewer) parameters in ResNet50 pretrained on ImageNet1K, both class-distance-normalized-variance and the clustering of tSNE are worse than training-from scratch, but still show relatively obvious clustering phenomenon. However, when we use the ResNet50 pretrained on ImageNet21K, the top-1 accuracy, and CDNV are significantly improved. More importantly, CDNV of ImageNet1K pretrained ResNet50 and ImageNet21K pretrained ResNet50 is lower than 1. This observation indicates although the image domain and point-cloud domain are quite different, the phenomenon of neural collapse generalization galanti2022on still exists in their transfer. More results and analysis are illustrated in Appendix B.5.
Moreover, the interesting discovery pushes us to think about the reason of cross-modal transfer having neural collapse. Inspired by galanti2022on , we briefly explain below. More detailed theoretical proof is presented in Appendix C.
In this work we focused on the problem of transferring knowledge between two tasks (source and target) consisting of two different modalities with different classes. Therefore, in the theoretical analysis, we have to deal with two separate modes of generalization: between classes and between modalities. In order to model this problem, we assume that the target and source tasks are decomposed of i.i.d. classes that are samples of two different distributions and (each stands for a different domain/modality). Each class is defined by a distribution over samples (e.g., samples of dog images). Given a target task (consisting of a set of randomly selected classes ), the pretrained model is evaluated after training an adaptor and a linear classifier on top of it. Its overall performance is measured in expectation over the selection of target tasks.
Conclusions
In this work, we use finetuned-image-pretrained models (FIP) to explore the feasibility of transferring image-pretrained models for point-cloud understanding and the benefits of using image-pretrained models on point-cloud tasks. We surprisingly discover that, with simply transforming a 2D pretrained ConvNet and minimal finetuning — input, output, and batch normalization layer (FIP-IO or FIP-IO+BN), FIP can achieve very competitive performance on 3D point-cloud classification, beating a wide range of point-cloud models that adopt a variety of tricks. Moreover, we find that when finetuning all the parameters of the pretrained models (FIP-ALL), the performance can be significantly improved on point-cloud classification, indoor and outdoor scene segmentation. Fully finetuned models generalize to most of the popular point-cloud methods. We also find that FIP-ALL can improve the data efficiency on few-shot learning and accelerate the training speed by a large margin. Additionally, we explore the relationships between neural collapse and cross modal transferring for our case, and shed light on why it works based on neural collapse. Compared with previous works that seek improvements from designing architectures and pretraining only on point-cloud modality, our work is not limited by the architecture design and the small-scale point-cloud dataset. We believe that image pretraining is one of the solutions to the bottleneck of point-cloud understanding and hope this direction can inspire the research community.
Acknowledgements
Co-authors from UC Berkeley were sponsored by Berkeley Deep Drive (BDD). Tomer Galanti’s contribution was supported by the Center for Minds, Brains and Machines (CBMM), funded by NSF STC award CCF-1231216.
References
Appendix A Implementation Details
Our experiments are conducted on ModelNet 3D Warehouse, S3DIS, and SemanticKITTI datasets. For the ModelNet 3D Warehouse dataset, we train all models on the train set and evaluate on the validation set. For the S3DIS, we train all models on area 1, 2, 3, 4, 6 and evaluate on area 5. For the SemanticKITTI dataset, we train all models on splits 00-10 except 08 which is used for evaluation. For each of the datasets, all ResNet series models use the same training scheme, and all experiments are implemented with PyTorch.
Training on ModelNet 3D Warehouse dataset.
In this case, coordinates of point-clouds are randomly scaled, translated, and jittered. We employ the SGD optimizer with momentum 0.9, weight-decay , and initial learning rate 0.1 with cosine learning rate scheduler. Each mini batch is set to 32, and models are trained for 300 epochs. For both training and inference phase, we only utilize coordinates without other features and set the voxel size to be 0.05. The experiments for ModelNet 3D Warehouse are all conducted on a Titan RTX GPU.
Training on the S3DIS dataset.
In this case, we concatenate all subparts of an indoor scene to train and validate on. Along directions, scenes are applied horizontal flip randomly. RGB features are randomly jittered, translated, and auto contrasted. Finally, we normalize and clip point-clouds. We set voxel size to 0.05, use SGD optimizer with momentum 0.9, weight-decay , and initialize learning rate to 0.1 with polynomial learning rate scheduler. Each mini batch is set to 3, and models are trained for 400 epochs on 2 Titan RTX GPUs.
Training on the SemanticKITTI dataset.
In this case, coordinates of each point-cloud are randomly scaled and rotated. We use SGD optimizer with momentum 0.9, weight-decay , and initial learning rate 0.24 with cosine warmup learning rate scheduler. Each mini batch is set to 2, and models are trained for 15 epochs on 4 Titan RTX GPUs. For both training and inference phases, we utilize coordinates as well as intensity feature and set voxel size to 0.05.
Most of our pretrained models were taken from open-sources https://pytorch.org/vision/stable/models.htmlhttps://github.com/Alibaba-MIIL/ImageNet21Khttps://github.com/hirokatsukataoka16/FractalDB-Pretrained-ResNet-PyTorchhttps://github.com/HRNet/HRNet-Semantic-Segmentation/tree/pytorch-v1.1https://github.com/rgeirhos/Stylized-ImageNethttps://github.com/wielandbrendel/bag-of-local-features-models, so we do not need to take time and computational resources for pretraining. We use torchsparsehttps://github.com/mit-han-lab/torchsparse to produce sparse 3D convolutions.
Details on Section 4.1
In this section, we take the ResNet architecture, inflate the pretrained models of different image datasets, and add linear input and output layers as shown in Section B.6. The ResNet50 was pretrained on ImageNet1K and is taken from the original PyTorch example. We use the same training recipe provided by PyTorch to train the ResNet50 on Tiny-ImageNet. The pretrained ResNet50 on ImageNet21K was taken from ridnik2021imagenet21k .
Details on Sections 4.2, 4.3, and 4.4.
In these sections, the pretrained ResNet models are taken from the same sources as those in Section 4.1.
For pretraining PointNet++ on ImageNet1K, we utilize the PointNet++ SSG version qi2017pointnetplusplus . We break the image into pixels and regard the group of pixels as a point-cloud with coordinates of positions in the original image and appending to all pixels. Then, we set center sampling number to 1024 and 256 for first and second stage, and the radius is set into 8 and 64, respectively. For each center point, we query 64 neighboring points. The training recipe is also provided by PyTorch.
For ViT models, we directly take the pretrained weights from dosovitskiy2020image . To apply it on ModelNet 3D Warehouse, we sample 256 centers and group 64 nearby points, regarding these as “point-cloud patches”. Then, we use a linear embedding to project the point-cloud patches into a sequence, and ViT processes them the same as image patches. Except for the linear embedding and the final output classifier, all the models are kept the same as the original version. For the experiments on S3DIS and SemanticKITTI, the architectural detail of ResNet18 is shown in A.4 listing 2.
For SimpleView model, all the experiment settings are the same as goyal2021revisiting . The only difference is whether to use the pretrained ResNet18. For HRNetV2-W48, we directly use the ImageNet1K and Cityscape pretrained models from SunXLW19 .
We conduct three trials on the few-shot experiments. For each trial, we change the random seed but keep all the other settings the same. To plot the training speed curve, we directly use the training log without any other changes, such as smoothing.
Appendix B Additional Experiments
For the first group of experiments, ResNet50 FIP either has IO or IO+BN finetuned. In addition to these two experimental settings, we also investigate finetuning input, output layers, and mean, variance of normalization layers, while fixing the convolution layer weights, normalization layer weights, and bias. The full experiment results with this extra setting are reported in Table 6 and 7. We can observe that compared with only finetuning input and output layers, updating mean and variance can also largely improve the performance of point-cloud recognition. As suggested in Section 2.4, we train batch normalizations to enhance the adaption between modalities.
B.2 Ablation study of inflating towards different directions.
We conduct experiments of inflating filters along different directions with the illustration figure shown in Figure 5 and the results shown in Table 8. We find that the performance is different when using different inflation methods. In particular, with ResNet50 pretrained on ImageNet1K, inflating along the x axis and the y axis leads to better performance compared with inflating along axis for both FIP-IO and FIP-IO+BN. More importantly, the minimally finetuned FIP-IO+BN with inflating along the and axis even surpasses the training-from-scratch.
B.3 Ablation study of loading different stages of the image-pretrained model.
We investigate the effect of loading different subsets of stages. The results are shown in Table 9. In detail, we load the pretrained weights partially while keeping the other weights randomly initialized. We observe that excluding the weights of the first stage achieves the best performance, bringing 0.77 points improvement.
B.4 Stability analysis of the semi-supervised experiment.
For the semi-supervised experiment, we change the random seed and calculated the mean and standard deviation of three trials for each setting as shown in Table 10.
B.5 Neural Collapse in the Embedding Layer.
Papyan24652 characterized neural collapse as training dynamics of overparameterized neural networks in which the feature embeddings of samples from the same class tend to concentrate around their class means. In this section, we briefly define neural collapse and evaluate it in our current setting. We refer the reader for additional details in Papyan24652 ; galanti2022on .
Essentially, this quantity measures to what extent the deviations of the embeddings of samples coming from and are smaller than the distance between their means. Intuitively, if the deviations are very small in comparison with the distances, then we expect the embeddings to be clustered with respect to their class labels. Note, that this quantity is also scale-invariant, i.e., if we multiply by , then, the CDNV would not change for any pair .
According to the definition in galanti2022on , neural collapse is defined in the following manner
where is the embedding function after epochs of training . Intuitively, during train time, the feature embeddings of samples of the same class tend to concentrate around their class-means in comparison with their distance from the other classes.
As can be seen in Table 11, across all of the experiments, the values of the CDNV are lower than , meaning that the standard deviations of the embeddings per class are smaller in comparison with the distances between class means. Therefore, we encounter a scenario where the embeddings of samples are fairly separated into classes. In addition, we observe that the degree of collapse generalizes well to new samples, as the CDNV on the train and validation data are relatively similar.
B.6 Details of used architectures.
Appendix C Theoretical Analysis
To motivate our approach, in this section we provide some theoretical support for the transfer of image domain to point-cloud domain described in the main text. Instead of specifying image and point-cloud in the analysis, we begin by introducing a formal framework for analyzing transfer learning between different modalities and classes. Note that the modalities should have grounded “relationships”, or as illustrated below, have meaningful task similarity. For example, image and point-cloud are both visual representations of the real world that share mutual characteristics, such as content and shape that could be captured using neural networks encoders.
We begin by introducing a formal framework for analyzing transfer learning between different modalities and classes. Then, we analyze a simple toy example, in which it is possible to perfectly translate one modality to the other using a linear mapping. Finally, we consider a more realistic case, in which we assume that the two domains share a ‘mutual semantic space’ that encodes the content within samples from the two domains.
We extend the transfer learning setting in galanti2022on . We consider the problem of training a generic feature representation on a source classification task and transferring it to a target task. The two classification problems correspond to different modalities (i.e., two different kinds of data representations) and consist of different sets of classes.
The source task.
Tasks similarity.
In general, transferring between the source and target tasks is meaningless if the two tasks are extremely unrelated to each other. For instance, we should not expect to have any guarantee to transfer knowledge from very different tasks, such as voice separation and image segmentation.
Evaluation process.
As a next step, we would like to evaluate the performance of the pretrained feature map on the target task. To do so, we evaluate its expected performance over the distribution of binary classification target tasks
where and are the outputs of a learning algorithm that trains and to fit to the dataset while freezing . For simplicity, in this work we focus on and denote , even though the analysis could be readily extended to . Note that several implementations of mappings are possible. For simplicity, in this work, we choose to be the ‘nearest empirical mean classifier’ and to be an empirical risk minimizer. Formally, for a given embedding function , we consider the linear function . In this case, forms a nearest empirical mean classification rule that classifies a vector as if it is closer to the empirical embeddings mean than to . In addition, , where and is the corresponding nearest empirical mean classifier. Notice that while the feature map is evaluated on the distribution of target tasks determined by classes taken from , the training of , as described above, is fully agnostic of this target.
Notation.
Appendix D Theoretical Results
where are i.i.d. uniformly distributed over . The empirical Rademacher complexity can lead to tighter bounds than those based on other measures of complexity such as the VC-dimension koltchinskii2004rademacher . It also has the added advantage that it is data-dependent and can be measured from finite samples.
where if and are spherically symmetric and otherwise.
where consists of the samples in excluding their labels and . By union bound over , with probability at least over the selection of , the following inequality holds uniformly for all ,
Hence, with probability at least over the selection of ,
In particular, since the loss function is bounded in $1-\frac{1}{2n}S$, we obtain the following inequality
Finally, we can take expectation over the selection of on both sides of the inequality to obtain that
Finally, since for any given distribution , we have for , by Proposition 5 in galanti2022on , we obtain that
D.2 Case 2
In the previous section, we assumed existence of a mapping , such that for . However, this assumption is typically violated in practice pmlr-v48-reed16 ; pmlr-v37-xuc15 ; DBLP:journals/corr/abs-1907-01341 . Therefore, following the Unsupervised Domain Adaptation literature bendavid ; DBLP:conf/alt/Mansour09 , we use a relaxed assumption that there is a ‘shared representation space’ for both domains. Informally, the two domains can be mapped to a shared space, in which classification into classes is possible. Variations of this assumption are algorithmically and theoretically used in multiple areas of computer vision huang2018munit ; Benaim2019DomainIntersectionDifference ; CycleGAN2017 ; radford2021learning ; Liu2021ContrastiveMF ; ALBEF .
where if and are spherically symmetric and otherwise.
The proof of this theorem is based on the analysis of galanti2022on , the theory of Unsupervised Domain Adaptation bendavid ; DBLP:conf/alt/Mansour09 ; DBLP:conf/colt/MansourMR09 and Rademacher complexities Mohri:2012:FML:2371238 .
By union bound over , with probability at least over the selection of , for any pair , we have
Hence, with probability at least over the selection of ,
Since the loss function is bounded in $1-\frac{1}{2m}S$, we have the following
Next, we can take expectation over the selection of on both sides of the inequality to obtain that
Finally, by Proposition 5 in galanti2022on we obtain that