Self-Supervised Pretraining of 3D Features on any Point-Cloud
Zaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan Misra
Introduction
Pretraining visual features on large labeled datasets is a pre-requisite to achieve good performance when access to annotations is limited . More recently, self-supervised pretraining has become a popular alternative to supervised pretraining especially for tasks where annotations are time-consuming, such as detection and segmentation in images or tracking in videos . In 3D computer vision, single-view depth scans are easy to acquire while reconstructed 3D scenes and annotations are difficult to obtain. Reconstructing a 3D scene requires registering and aligning multiple depth maps captured in a static environment, which may fail easily if there exists fast camera motion or odometry drift . The resulting 3D scene is a point cloud composed of thousands of 3D points that need to be annotated individually in the case of segmentation or by groups for detection, which is time-consuming and laborious. For example, it takes around 22 minutes to annotate a single scene in ScanNet . Thus, 3D data acquisition and annotation both require significantly more effort than images or videos, and there is no existing weak supervision, like image tags, to shortcut the process . This cumbersome annotation process results in a lack of large annotated 3D datasets. However, consumer-grade depth sensors have become more easily accessible and simpler to use, \eg, in phones , leading to large quantities of single-view raw depth maps. These depth maps can be leveraged to pretrain 3D features, but, there is surprisingly little work that can be applied.
Pretraining on depth maps faces several challenges: first, the absence of any supervision at scale requires the use of a self-supervised signal. Second, the absence of alignments between depth maps excludes methods using multi-view constraints, such as finding correspondences between views . Finally, the representation of a 3D scene varies depending on the application—segmentation typically uses a voxel representation , whereas detection uses point-clouds —and therefore the ideal pretraining method should be easily applicable to any 3D representation.
In this paper, we introduce a simple contrastive framework, DepthContrast, to simultaneously learn different representations of any 3D data. Our approach is based on the Instance Discrimination method by Wu et al. , applying to depth maps. We side-step the need of registered point clouds or correspondences, by considering each depth map as an instance and discriminating between them, even if they come from the same scene. For 3D representations with different architectures, we jointly learn features by considering the different representations as data augmentations that are processed with their associated networks .
Our contributions can be summarized as follows:
We show that single view 3D depth scans can be used to learn powerful feature representations using self-supervised learning.
We show that joint training of different input representations like points and voxels is important for learning good representations, and a naive application of contrastive learning may not yield good results.
Our method is applicable across different model architectures, indoor/outdoor 3D data, single/multi-view 3D data. We also show that it can be used to pretrain high capacity 3D architectures which otherwise overfit on tasks like detection and segmentation.
We show performance improvements over nine downstream tasks, and set a new state-of-the-art for two object detection tasks (ScanNet and SUNRGBD). Our models are efficient few-shot learners.
Related Work
Our method builds on the work from the self-supervised learning literature, with 3D data as an application. In this section, we give an overview of the recent advances in both self-supervision and 3D representations.
Self-supervised learning for images. Self-supervised learning is a well studied problem in machine learning and computer vision . There are many classes of methods for learning representations - clustering , GANs , pretext tasks \etc. Recent advances have shown that self-supervised pretraining is a viable alternative to supervised pretraining for 2D recognition tasks. Our work builds upon contrastive learning where models are trained to discriminate between each instance with no explicit classifier . These instance discrimination methods can be extended to multiple modalities . Our method extends the work of Wu et al. to multiple 3D input formats following Tian et al. using a momentum encoder instead of a memory bank.
Self-supervised learning for 3D data. Most methods on self-supervised learning focus on single 3D object representation with different applications to reconstruction, classification or part segmentation . Recently, Xie et al. proposed a self-supervised method to build representations of scene level point clouds. Their method relies on the complete 3D reconstruction of a scene with point-wise correspondences between the different views of a point cloud. These point-wise correspondences requires post-processing the data by registering the different depth maps into a single 3D scene. Their method can only be applied to static scenes that have been registered, which greatly limits the applications of their work. We show a simple self-supervised method that learns state-of-the-art representations from single-view 3D data, and can also be applied to multi-view data.
Representations of 3D scenes. There are multiple ways to represent 3D information in different vectorized forms such as point-clouds, voxels or meshes. Point-cloud based models are widely used in classification and segmentation tasks , 3D reconstruction and 3D object detection . Since many 3D sensors acquire data in terms of 3D points, point clouds are a convenient input for deep networks. However, since using convolution operations on point-clouds directly is difficult , voxelized data is another popular input representation. 3D convolutional models are widely used in 3D scene understanding . There are also efforts to combine different 3D input representations . In this work, we propose to jointly pretrain two architectures for points and voxels, that are PointNet++ for points and Sparse Convolution based U-Net for voxels.
3D transfer tasks and datasets. We use shape classification, scene segmentation, and object detection as the recognition tasks for transfer learning. Shape classification techniques are widely evaluated on the ModelNet dataset, which we use. It contains synthetic 3D data and each sample contains exactly one object. We also evaluate on complete 3D scenes using the more general 3D scene understanding task. Scene-centric datasets can be broadly divided into indoor scens , and outdoor (self-driving focussed) scenes . We use these datasets and evaluate the performance of our methods on the indoor detection , scene segmentation , and outdoor detection tasks .
Approach
We develop a scalable pretraining method, DepthContrast, for 3D representations that uses unprocessed single-view or multi-view depth maps without human annotations. Our method, illustrated in fig. 2, is based on the instance discrimination framework of Wu et al. with a momentum encoder . DepthContrast learns 3D representations across multiple 3D input formats like points and voxels, and across different 3D architectures by using an extension of contrastive learning to multiple data formats .
Given a dataset containing samples , we wish to learn a function that produces useful representations of the input sample. Our method uses 3D data where can be represented by point coordinates or voxelsPoints in a depth map are a set, but for simplicity we denote them as a matrix. Our method does not rely on any specific ordering of the points.. We apply a data augmentation sampled randomly from a large set of augmentations , to obtain an augmented sample . The augmented sample is input to a deep network that extracts unit-norm global features by pooling over the 3D spatial coordinates. We setup an instance discrimination problem where the features and obtained from two data augmented versions of sample must be similar to each other, and different from features obtained using other (negative) samples in the dataset. We use a contrastive loss to achieve this goal:
where is the temperature that controls the smoothness of the softmax distribution. This loss encourages features from different augmentations of the same scene to be similar, while being dissimilar to features of other scenes. Thus, it learns features that focus on discriminative regions of a scene that make it different from other scenes in the dataset.
Minimal assumptions on input data. Our method, by design, makes minimal assumptions about the input , \ie, it is an unprocessed single-view depth map. It does not require careful sampling of overlapping multi-view 3D inputs or object centric depth maps . These minimal assumptions enable us to learn from large scale single-view 3D depth maps in § 4, indoor and outdoor 3D depth maps obtained from a variety of sensors without relying on 3D calibration in § 5.3.
Momentum encoder. As using a large number of negatives is important for contrastive learning , we use the method of He et al. where the features of the other augmentation and negative samples in eq. 1 are obtained using a momentum encoder and a queue respectively. This allows us to use a large number of negative samples without increasing the training batch size.
2 Extension to Multiple 3D Input Formats
Multiple input formats are commonly used to represent 3D data - point clouds, voxels, meshes \etcand have their specific deep learning architectures and applications. Our self-supervised method can be naturally extended to accommodate these input formats and architectures. For each input format , we denote the corresponding input sample as , the format-specific encoder network as , and the extracted feature as . Extending eq. 1, we can minimize a single objective that performs instance discrimination within and across input formats :
When the input formats are identical, this objective reduces to the within format loss of eq. 1, and when this objective aligns the feature representations obtained across formats using different network architectures . As illustrated in fig. 2, we use two popular input formats - point clouds and voxels, and train these format-specific models with a single joint loss function
Similar techniques have been explored in the context of different modalities of data, \eg, color and grayscale images , audio and video \etc. While these methods use different modalities, our extension uses the same 3D data and only changes the input format.
3 Model Architecture
Point input. We use PointNet++ as the backbone network which takes as input the XYZ coordinates of the 3D data. PointNet++ employs a U-Net structure which has four layers of feature extraction and down-sampling, and two layers of feature aggregation and up-sampling. Our network takes as input K points from the scene represented as a K matrix of XYZ coordinates. The network’s final layer produces dimensional per-point features for points after aggregation. We obtain the scene level dimensional feature in eq. 2 by global max pooling to these last layer features, followed by a two layer MLP as in and L2 normalization. We increase the size of the output feature in the last two upsampling (fp) layers from in to . This improves the performance of the baselines, \eg, training from scratch, as well as our method and we use this improved architecture throughout the paper.
Voxel input. We use a sparse convolution U-Net model as the backbone for the voxel 3D input. The network takes a 3D occupancy grid as the input representation of the 3D data. Our U-Net consists of four layers of feature extraction and pooling and four layers of feature aggregation and up-sampling. We use a voxel size of to voxelize the input data and following , use the voxel occupancy grid and the RGB values as input to the model. To obtain the scene level dimensional feature ( eq. 2), similar to the point input, we apply global max-pooling to the last layer feature, followed by a two layer MLP and L2 normalization.
Using global feature representations. Our instance discrimination formulation learns global feature representations from the input 3D data , \ie, the feature representation is obtained by pooling over the 3D coordinates of the input. On the other hand many downstream 3D tasks require local per-point predictions like segmentation. We show in § 4 that despite using global features, our pretraining benefits such prediction tasks.
4 Data Augmentation for 3D
Data augmentation is as an essential component of our framework. We first adopt standard pointcloud data argumentation methods proposed in , which are random point up/down sampling, random flip in xy axis, and random rotation. However, after adding these methods, it is still easy for the network to distinguish different training instances. Thus, we add two new data augmentation methods: random cuboid and random drop patches. Inspired by the random crop in 2D images , we define a random cuboid augmentation that extracts random cuboids from the input point cloud. Cuboids are sampled using a random scale of the original scene, and a random aspect ratio . We also drop (erase) cuboids to force the network learn local geometric features. The dropped cuboid is randomly cropped with of the scene scale. The performance boost from each augmentation is analyzed in § 5. For voxelized inputs, in addition to all the point augmentations, we use the augmentations from .
5 Implementation Details
We use K negatives for contrastive learning in eq. 3 and a momentum of for the momentum encoder following . As noted in § 3.3, we follow Chen et al. and use an additional non-linear projection (two layer MLP) and L2 normalization after the network to obtain the features . The features are dimensional and we use a temperature value of while computing the non-parametric softmax in eq. 1. We use a standard SGD optimizer with momentum , cosine learning rate scheduler starting from to and train the model for epochs with a batch size of .
Experiments
We evaluate DepthContrast pretraining by transfer learning, \ie, fine-tuning on downstream tasks and datasets. As table 1 shows, we use a diverse set of 3D understanding tasks like object classification, semantic segmentation, and object detection. We first study a single input 3D format and a single network architecture in § 4.2 showing that DepthContrast’s performance improves with large data and higher capacity models, and benefits finetuning with limited labels. Finally, in § 4.3, we show the benefits of our pre-training across different 3D input formats (points and voxels).
Pretraining Details. We use single-view depth map videos from the popular ScanNet dataset and term it as ScanNet-vid. ScanNet-vid contains about 2.5 million RGB-D scans for more than 1500 indoor scenes. Following the train/val split from , we extract around 190K RGB-D scans (one frame every 15 frames) from about 1200 video sequences in the train set. We do not use camera calibration or 3D registration methods and operate directly on single-view depth maps. We use our data augmentation described in § 3.4 and use the training objectives from § 3. Additional details are provided in § 3.5 and the supplemental material.
We evaluate our pretrained model by transfer learning and finetune it on different downstream datasets and tasks summarized in table 1. We use diverse downstream datasets - full scenes/object centric; using different 3D sensors; single/multi-view; real/synthetic; indoor/outdoor. On these diverse datasets, we use three major tasks, which are classification, semantic segmentation, and object detection. These tasks test different aspects of the pretrained model - while object detection and semantic segmentation use local features, classification is performed on global features.
2 Pretraining with Point Input Format
We pretrain a PointNet++ model using the instance discrimination objective in eq. 1 on the single-view depth maps from ScanNet-vid. We study the transfer performance of the pretrained model on object detection using the VoteNet framework that uses a PointNet++ backbone network. In table 2 we report the detection results by finetuning the VoteNet model with different backbone initializations. We use the implementation of for finetuning and report the detection performance using the mean Average Precision at IoU=0.25 (AP25) metric. Training from scratch or random initialization is standard practice in VoteNet and serves as a baseline for comparing other pretraining methods. As a supervised pretraining baseline, we use the VoteNet model trained on ScanNet detection. Since the supervised baseline is pretrained specifically on object detection, it serves as a strong baseline.
Scratch training provides competitive results on the larger detection datasets like ScanNet and SUNRGBD , however, its performance on the smaller S3DIS dataset is low. In comparison, supervised pretraining provides large gains in the detection performance across all datasets. DepthContrast outperforms training from scratch on all the four datasets, and improves performance by 12.1% mAP on the small S3DIS dataset that has only 200 labeled training samples. We further analyze label efficiency of our model in § 4.2.4. Interestingly, despite using no labels during pretraining, DepthContrast is better than the detection-specific supervised pretraining for two datasets (SUNRGBD and Matterport3D). Our method also outperforms recent work by significant margins.
We use DepthContrast for training higher capacity models. We follow the practice in 2D self-supervised learning and increase the capacity of the PointNet++ model by multiplying the channel width of all the layers by . We pretrain all models on the ScanNet-vid dataset and measure their transfer performance in fig. 3. Training large models from scratch provides some benefit, but quickly leads to reduced or plateauing performance. We observe overfitting on the small datasets like S3DIS where increasing the model capacity does not improve performance. On the other hand, our self-supervised pretraining on ScanNet-vid reduces this overfitting and performance improves or stays the same for larger models. This suggests that pretraining is crucial for training large 3D detection models.
2.2 Using More Pretraining Data
We increase the pretraining data by using readily available single-view 3D data from the Redwood-vid dataset . Redwood-vid contains over 23 million depth scans from RGB-D videos taken in both indoor and outdoor settings. As this dataset is extremely large, we use a subset of 2500 video sequences consisting of 10 categories and extract 370K RGB-D scans. More importantly, the Redwood-vid dataset does not contain camera extrinsic parameters and thus cannot be registered to get a multi-view dataset which is a necessity for prior self-supervised methods .
Combining the Redwood-vid and ScanNet-vid datasets allows us to triple our pretraining data. We pretrain all models on this combined dataset and report their performance (AP25) in fig. 3. DepthContrast’s performance improves with both model capacity and number of pretraining samples across all four detection datasets. The higher capacity models show a larger improvement in performance particularly on the smaller S3DIS dataset. This suggests that DepthContrast can leverage large amounts of single-view 3D data to obtain better and higher capacity 3D models.
2.3 State-of-the-art Detection Frameworks
We use two state-of-the-art detection frameworks - H3DNet and VoteNet and study the benefit of using our pretrained model. We use our PointNet++ model pretrained on the combined Redwood-vid and ScanNet-vid dataset and transfer it using these detection frameworks. The detection results in table 3 show that our pretrained model achieves state-of-the-art performance on SUNRGBD and ScanNet. In particular, as the gains are larger on stricter mAP at IoU=0.5, our pretrained models result in detection models that are better at localization.
2.4 Label Efficiency of Pretrained Models
Pretraining allows models to be finetuned with small amount of labeled data. In table 2, we observe that small labeled datasets benefit more from pretraining. We study the label efficiency of DepthContrast pretrained models by varying the amount of labeled data used for finetuning. While varying the data, we draw 3 independent samples and report average results. We use the PointNet++ models pretrained on ScanNet-vid ( § 4.2) and report the detection performance in fig. 1. DepthContrast pretraining provides large gains in performance at every setting. On both the ScanNet and SUNRGBD datasets, our model with just samples gets the same performance as training from scratch with the full dataset. When using samples for finetuning, our pretrained models provide a gain of over 10% mAP. This shows that our pretraining is label efficient and can improve performance especially on tasks with limited supervision.
Does pretraining benefit tail classes? 3D detection datasets like SUNRGBD and ScanNet exhibit a long tailed distribution where many ‘tail’ classes have few training instances. In the SUNRGBD dataset, the ‘tail’ classes like bathtub, toilet, dresser have less than 200 training instances, while classes like chair have over 9000 instances. fig. 4 shows the gain of our pretrained model over the scratch model across object classes on the SUNRGBD dataset. Our pretraining improves the performance of classes with fewer instances, \ie, the tail classes, by AP. This suggests that DepthContrast pretraining can partially address the long tailed label distributions of current 3D scene understanding benchmarks.
3 Pretraining with Multiple Input Formats
We pretrain DepthContrast using both the point and voxel input formats and use two format-specific encoders - PointNet++ for points and UNet for voxels. As explained in eq. 3, when using multiple 3D input formats, we can define two loss terms - a within format loss and an across format loss. To analyze which of these loss terms matter for pretraining, we consider three variants - (1) Within format which independently trains format-specific models for each input format and is a straightforward application of instance discrimination to 3D; (2) Across format which trains the format-specific models jointly using the second term of eq. 3; (3) Ours which trains the format-specific models jointly using our combined loss function. We evaluate the pretrained models by transfer learning. As in § 4.2, we finetune the pretrained point input format PointNet++ models on SUNRGBD and ScanNet detection using the VoteNet framework. We finetune the voxel UNet models on segmentation using the framework from Spatio-Temporal Segmentation which uses a UNet backbone network. The results are summarized in table 4.
Compared to training from scratch, the within format pretraining only provides a benefit for the point input format PointNet++ models. For the voxel models, this pretraining does not improve consistently over training from scratch, which is in line with observations from recent work . This shows that a naive application of instance discrimination to 3D representation learning may not yield good pretrained models. The across format loss improves performance for both the point and voxel models, suggesting the benefit of using multiple input formats. Our proposed loss function that combines both the within and across format losses provides the best transfer performance. The gains are particularly significant on the voxel format model which improves by over the within format loss. In the supplemental material, we show that this benefit of joint training over the within format loss also holds across different pretraining data and architectures. Although our method only uses single-view unprocessed depth scans, our results on the voxel transfer tasks are comparable to the recent PointContrast method that uses multi-view point clouds and pointwise correspondences. We note that our UNet architecture is different from since their architecture underfit on our self-supervised pretraining task. We provide results with their architecture in the supplement.
Analysis
In this section we present a series of experiments designed to understand DepthContrast better. We first pretrain point format (PointNet++) models on the ScanNet-vid dataset following the settings from § 4.2. We use two transfer tasks for evaluation - (1) object detection on SUNRGBD using VoteNet where we finetune the full model and test the quality of the pretraining; (2) object classification on ModelNet where we keep the model fixed and only train linear classifiers on fixed features, thus testing the quality of the learned representations . Finally, we also evaluate DepthContrast’s generalizability to outdoor 3D data.
Data augmentations play an important role for self-supervised representation learning and have been studied extensively in the case of 2D images . However, the impact of data augmentation for 3D representation learning is less well understood. Thus, we analyze the effect of our proposed augmentations from § 3.4 on transfer performance. We train different DepthContrast point models with the same training setup and only vary the data augmentation used. Our results are summarized in table 5.
The widely used VoteNet augmentations from supervised learning perform worse than our proposed augmentations. Our augmentations lead to both a better feature representation: a gain of accuracy on ModelNet classification, and a better pretrained model: mAP on SUNRGBD detection. In our experiments, we consistently observe gains from our improved data augmentation on all the downstream tasks from § 4 which underscores the importance of designing good data augmentation.
2 Impact of Single-view or Multi-view 3D Data
Our self-supervised method does not make assumptions on the input data and can use single-view depth maps without 3D preprocessing as input. We study whether pretraining on reconstructed multi-view 3D scenes impacts the downstream performance. We use the ScanNet dataset which contains multi-view 3D data obtained by 3D registration of the ScanNet-vid depth maps. As another single-view dataset, we pretrain on the Redwood-vid dataset from § 4.2.2. We pretrain DepthContrast point models on these datasets and compare their performance by transfer learning in table 6.
The transfer performance is similar when models are pretrained on ScanNet-vid or ScanNet. Since ScanNet-vid and ScanNet only differ in the 3D preprocessing involved, the result suggests that DepthContrast is not sensitive to single-view or multi-view input data. This is not surprising given that our objective does not rely on multi-view information. Pretraining on the single-view Redwood-vid dataset also gives similar performance suggesting that DepthContrast is robust to different data distributions during pretraining. All the DepthContrast models outperform the scratch model.
3 Generalization to Outdoor LiDAR data
We test DepthContrast’s generalization to outdoor LiDAR data by pretraining on the Waymo Open Dataset where we extract 79K single-view scans from the videos. We use the same data augmentation parameters from § 3.4 and only modify random cuboid to work on the full scale of the (depth) dimension of the scene. We use the standard LiDAR-specific model architectures as our format-specific encoders - PointnetMSG for point clouds and Spconv-UNet for voxels. Similar to § 3.5, we obtain features from these models after global max pooling and a two layer MLP. The models are optimized jointly with eq. 3 using both within and across-format losses. For transfer learning, we use the standard KITTI object detection benchmark, and PointRCNN and Part- for down-stream models. We report results on the cyclist class since it has fewer examples in the training set compared to the other classes. We provide results for other classes and finetuning details in the supplemental material. Similar to § 4.2.4, while varying the fraction of pretraining data, we report average performance across 3 independent samplings of the data. fig. 5 shows that our pretrained models outperform training from scratch especially when finetuning on fewer training samples. For Spconv-Unet, we achieve a 20% gain with 5% of labeled data. This suggests that DepthContrast pretraining generalizes across multiple input formats, and our proposed data augmentation generalizes to different depth sensors and scene types.
Conclusion
We propose DepthContrast- an easy to implement self-supervised method that works across model architectures, input data formats, indoor/outdoor 3D, single/multi-view 3D data. DepthContrast pretrains high capacity models for 3D recognition tasks, and leverages large scale 3D data that may not have multi-view information. We show state-of-the-art performance on detection and segmentation benchmarks, outperforming all prior work on detection. We provide crucial insights that make our simple implementation work well - training jointly with multiple input data formats, and designing a generalizable data augmentation scheme. We hope DepthContrast helps future work in 3D self-supervised learning.
References
Supplemental Material
We provide the model architecture details in appendix A. Hyper-parameters used in training and fine-tuning for PointNet++ and UNet and additional results are shown in appendix B. We also show hyper-parameters used in training and fine-tuning for PointnetMSG and Spconv-UNet and results for other categories of detection in KITTI in appendix C.
Appendix A Architecture Details
PointNet++ model used in §§ 4, 5.1 and 5.2. As shown in Table 7, PointNet++ contains four set abstraction layers and two feature up-sampling layers, designed in . Each SA layer is specified by , where represents number of output points, represents the ball-region radius of the reception field, represents the feature channel size of the i-th layer in the MLP. Each feature up-sampling (FP) layer upsamples the point features by interpolating the features on input points to output points, as designed in . Each FP layer is specified by where is the output of the i-th layer in the MLP. In § 4.2.1 of the main paper, we create higher capacity versions of the PointNet++ model by increasing the channel width. For PointNet++ , and , we multiply the feature size by of each layer in the MLP of the four set abstraction layers. When applying PointNet++ in VoteNet and H3DNet, we adjust the rest of the model accordingly based on different point feature size.
PointnetMSG model on LiDAR data used in § 5.3. Table 8 shows the architecture details for PointnetMSG, which processes lidar point cloud for PointRCNN detection model. PointnetMSG contains four multi-scale set abstraction layers and four feature up-sampling layers. Multi-scale set abstraction samples points in different scales and process them with different MLPs. In here, we adopt the architecture design in , in which each set abstraction layer contains two point features produced with two ball-region radius and two MLPs. As shown in table 8, each SA layer is specified by , where indicates the ball-region radius for each scale and indicates the feature channel size of the i-th layer in the MLP of each scale. We directly apply the learnt PointnetMSG in PointRCNN for detection evaluation in KITTI.
UNet model used in § 4.3. fig. 6 shows the network architecture for UNet. It mainly contains four encoding resblock and four decoding resblock. We use sparse convolution and sparse resblock designed in . For each sparse resblock, we first apply sparse convolution or deconvolution, depending on encoding or decoding, with kernel size 2 and stride 2. Then, we apply N number of sparse convolution layers with kernel size 3 and stride 1. D represents the output feature dimension. Since sparse convolution takes variable sized input and output, we do not specify the number of voxels in each layer here. We directly apply the learnt UNet backbone for different scene segmentation tasks.
Spconv-UNet model on LiDAR data used in § 5.3. fig. 7 shows the network architecture for Spconv-UNet used for Part detection model. It mainly contains four sparse blocks for encoding and four sparse upblocks for decoding. We use sparse convolution and sparse resblock designed in . For each sparse block/upblock, we show the number of convolution layers N and output feature dimension D. We directly adopted the architecture design in Part . For KITTI detection evaluation, we directly load the learnt backbone for fine-tuning.
Appendix B Training Details
As mentioned in the main paper, we use a standard SGD optimizer with momentum 0.9, cosine learning rate scheduler starting from 0.12 to 0.00012 and train the model for 1000 epochs with a batch size of 1024. We observed that pretraining for 400 epochs already gives good results, and 1000 epochs for pretraining only slightly improve the model.
B.2 Experimental Details for PointNet++
For all the VoteNet fine-tuning evaluations, we use the original configurations , where we apply Adam optimizer and use a base learning rate 0.001 with a 10 weight decrease at 80, 120 and 160 epochs. The model is trained for 180 epochs in total. Since the training set of S3DIS only contains 200 training instances, we train for 360 epochs. We use a batch size of 8 for ScanNet, Matterport3D, and S3DIS, and a batch size of 16 for SUNRGBD. We use the same configuration for training from scratch and fine-tuning, and we only load the pretrained PointNet++ backbone during fine-tuning.
For the H3DNet fine-tuning, we only use one backbone network instead of the original four . For initial learning rate and decays, we use the original configurations. We found that with PointNet++ backbone, we are able to re-produce the previous results reported in the paper with one backbone network. We use the same configuration for training from scratch and fine-tuning, and we only load the pretrained PointNet++ backbone during fine-tuning.
In § 5 of the main paper, we used the ModelNet dataset for transfer learning. We trained linear classifiers on fixed features for this task and measured the classification accuracy.
Full finetuning. We now show results on the same task but using finetuning, \ie, all the parameters of the backbone model are updated. We use the PointNet++ backbone to extract per-point feature and apply max-pooling to get the final global feature vector. We then apply one linear layer to get the final class labels. We use SGD+momentum optimizer with an initial learning rate 0.01. We use multistep LR scheduler with 10 weight decrease at 8, 16 and 24 epochs. We train for 28 epochs. We apply the data augmentation used in during training. We use the same configuration for training from scratch and fine-tuning, and we only load the pretrained PointNet++ backbone during fine-tuning. For linear probing, we fix the pretrained weight and only fine-tune the last linear layer. In table 9, we compare the fine-tuning and train from scratch results for both linear probing and full fine-tuning. The pretrained PointNet++ model provides consistent improvements.
B.2.2 Label efficiency of PointNet++
We follow the same settings as § 4.2.4 of the main paper and evaluate the label efficiency of the PointNet++ pretrained using DepthContrast. In Figure 1 (main paper) we showed that DepthContrast pretraining was label efficient on the ScanNet and SUNRGBD datasets. In fig. 8, we show the label efficiency plots for Matterport3D and S3DIS downstream detection tasks. Our results are consistently better than training from scratch across all the detection benchmarks used in the paper.
B.3 Experimental Details for UNet
For all the UNet fine-tuning evaluation, we use SGD+momentum optimizer with an initial learning rate 0.1. We use Polynomial LR scheduler with a power factor of 0.9. Weight decay is 0.0001. For voxel size, we use 0.05(5cm) for S3DIS and Synthia and 0.04(4cm) for ScanNet. We use the original data augmentation techniques in . We use batch size 48 for S3DIS and 56 for ScanNet and Synthia. We train the model with 8 V100 GPUs with data parallelism for 20000 iterations for all three tasks. We use the same configuration for training from scratch and fine-tuning, and we only load the pretrained UNet backbone during fine-tuning.
For ScanNet scene segmentation, due to memory issue, we increased the voxel size from the default 0.02(2cm) to 0.04(4cm), which leads to different scratch training results compared to . Although we are pretraining and fine-tuning on the same dataset, our approach still provides improvements as shown in table 10.
B.3.2 Label efficiency of UNet
We pretrain a UNet model using DepthContrast on the ScanNet-vid dataset. We finetune this model for scene segmentation task. In fig. 9, we show the data efficiency plot for S3DIS scene segmentation dataset. Our approach provides consistent performance boost for different percentages of training data used.
Appendix C Experimental Details for Lidar Data
We now present the experimental details when using DepthContrast on LiDAR 3D data (Section 5.3 of the main paper).
We use the original configuration from PointRCNN and Part to get the scratch training results for different splits of training dataset. For PointRCNN, we use AdamW optimizer with an initial learning rate 0.01, weight decay 0.01, and momentum 0.9. For Part , we also use AdamW optimizer with an initial learning rate 0.003, weight decay 0.01, and momentum 0.9. We use batch size 24 for PointRCNN and batch size 16 for Part . We train both models for 80 epochs, and the learning rate will drop by at 35 and 45 epochs. For both methods, we apply the same data augmentation and processing pipelines in .
We use the same configuration for training from scratch and fine-tuning for 5%, 10%, 20% and 50% of the labeled training data. For 100% training data, we observed over-fitting issue for classes with fewer training instances, such as cyclist and pedestrian. Thus, we increased the initial learning rate by . For PointRCNN, we also decreased the number of training epochs to 60 and set the first weight drop at 30 epochs. After modifying those parameters, we observed that the performance of scratch training didn’t change.
C.2 Results on KITTI
In the main paper, due to space constraints, we only showed the results on the cyclist class in the KITTI dataset. We present results on the remainder classes.
We show the label efficiency evaluation results for car detection in KITTI in fig. 10. For simplicity, we only show the results at moderate difficulty level. The results at other difficulty levels maintain similar pattern. DepthContrast provides a performance boost with fewer training instances, especially with 5% of labeled data.
In fig. 11, we show the label efficiency evaluation results for pedestrian detection in KITTI. Our pretraining provides consistent gain over scratch for PointRCNN model. For Part , our pretraining also provides significant performance boost with fewer training instances.
Advantage of joint training over within-format loss. Similar to Section 4.3 of the main paper, we analyze if training jointly with the within and across-format losses provides better transfer performance. Specifically, we compare (Equation 3 of the main paper) our full joint loss with the within-format only loss. We pretrain separate models with these losses and evaluate them on the KITTI dataset.
In fig. 12, we show the label efficiency evaluation results for pedestrian detection. We can see that with joint loss training, the model provides more performance gain with fewer training instances, which proves that the benefit of joint training holds across different pretraining data and architectures.