PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding

Saining Xie, Jiatao Gu, Demi Guo, Charles R. Qi, Leonidas J. Guibas, Or Litany

Introduction

Representation learning is one of the main driving forces of deep learning research. In 2D vision, the finding that pre-training a network on a rich source set (e.g. ImageNet classification) can help boost performance once fine-tuned on the usually much smaller target set, has been key to the success of many applications. A particularly important setting is when the pre-training stage is unsupervised, as this opens up the possibility to utilize a practically infinite train set size. Unsupervised pre-training has been remarkably successful in natural language processing , and has recently attracted increasing attention in 2D vision .

In the past few years, the field of 3D deep learning has witnessed much progress with an ever-increasing number of 3D representation learning schemes . However, it still falls behind compared to its 2D counterpart as evidently, in all 3D scene understanding tasks, ad-hoc training from scratch on the target data is still the dominant approach. Notably, all existing representation learning schemes are tested either on single objects or low-level tasks (e.g. registration). This status quo can be attributed to multiple reasons: 1) Lack of large-scale and high-quality data: compared to 2D images, 3D data is harder to collect, more expensive to label, and the variety of sensing devices may introduce drastic domain gaps; 2) Lack of unified backbone architectures: in contrast to 2D vision where architectures such as ResNets have proven successful as backbone networks for pre-training and fine-tuning, point cloud network architecture designs are still evolving; 3) Lack of a comprehensive set of datasets and high-level tasks for evaluation.

The purpose of this work is to move the needle by initiating research on unsupervised pre-training with supervised fine-tuning in deep learning for 3D scene understanding. To do so, we cover four important ingredients: 1) Selecting a large dataset to be used at pre-training; 2) identifying a backbone architecture that can be shared across many different tasks; 3) evaluating two unsupervised objectives for pre-training the backbone network, and 4) defining an evaluation protocol on a set of diverse downstream datasets and tasks.

Specifically, we choose ScanNet as our source set on which the pre-training takes place, and utilize a sparse residual U-Net as the backbone architecture in all our experiments and focus on the point cloud representation of 3D data. For the pre-training objective, we evaluate two different contrastive losses: Hardest-contrastive loss , and PointInfoNCE – an extension of InfoNCE loss used for pre-training in 2D vision. Next, we choose a broad set of target datasets and downstream tasks that includes: semantic segmentation on S3DIS , ScanNetV2 , ShapeNetPart and Synthia 4D ; and object detection on SUN RGB-D and ScanNetV2. Remarkably, our results indicate improved performance across all datasets and tasks (See Table 1 for a summary of the results). In addition, we found a relatively small advantage to pre-training with supervision. This implies that future efforts in collecting data for pre-training should favor scale over precise annotations.

Our contributions can be summarized as follows:

We evaluate, for the first time, the transferability of learned representation in 3D point clouds to high-level scene understanding.

Our results indicate that unsupervised pre-training improves performance across downstream tasks and datasets, while using a single unified architecture, source set and objective function.

Powered by unsupervised pre-training, we achieve a new state-of-the-art performance on 6 different benchmarks.

We believe these findings would encourage a change of paradigm on how we tackle 3D recognition and drive more research on 3D representation learning.

Related work

Deep neural networks are notoriously data hungry. This renders the ability to transfer learned representations between datasets and tasks extremely powerful. In 2D vision it has led to a surge of interest in finding optimal pretext unsupervised tasks . We note that while many of these tasks are low-level (e.g. pixel or patch level reconstruction), they are evaluated based on their transferability to high-level tasks such as object detection. Being much harder to annotate, 3D tasks are potentially the biggest beneficiaries of unsupervised- and transfer-learning. This was shown in several works on single object tasks like reconstruction, classification and part segmentation . Yet, generally much less attention has been devoted to representation learning in 3D that extends beyond the single-object level. Further, in the few cases that did study it, the focus was on low-level tasks like registration . In contrast, here we wish to push forward research in 3D representation learning by focusing on transferability to more high-level tasks on more complex scenes.

In this work, we focus on learning useful representation for point cloud data. Inspired by the success in 2D domain, we conjecture that an important ingredient in enabling such progress is the evident standardization of neural architectures. Canonical examples include VGGNet and ResNet/ResNeXt . In contrast, point cloud neural network design is much less mature, as is apparent by the abundance of new architectures that have been recently proposed. This has multiple reasons. First, is the challenge of processing unordered sets . Second, is the choice of neighborhood aggregation mechanism which could either be hierarchical , spatial CNN-like , spectral or graph-based . Finally, since the points are discrete samples of an underlying surface, continuous convolutions have also been considered . Recently Choy et al. proposed the Minkowski Engine , an extension of submanifold sparse convolutional networks to higher dimensions. In particular, sparse convolutional networks facilitate the adoption of common deep architectures from 2D vision, which in turn can help standardize deep learning for point cloud. In this work, we use a unified U-Net architecture built with Minkowski Engine as the backbone network in all experiments and show it can gracefully transfer between tasks and datasets.

PointContrast Pre-training

In this section, we introduce our unsupervised pre-training pipeline. First, to motivate the necessity of a new pre-training scheme, we conduct a pilot study to understand the limitations of existing practice (pre-training on ShapeNet) in 3D deep learning (Section 3.1). After briefly reviewing an inspirational local feature learning work Fully Convolutional Geometric Features (FCGF) (Section 3.2), we introduce our unsupervised pre-training solution, PointContrast, in terms of pretext task (Section 3.3), loss function (Section 3.4), network architecture (Section 3.5) and pre-training dataset (Section 3.6).

Previous works on unsupervised 3D representation learning mainly focused on ShapeNet , a dataset of single-object CAD models. One underlying assumption is that by adopting ShapeNet as the ImageNet counterpart in 3D, features learned on synthetic single objects could transfer to other real-world applications. Here we take a step back and reassess this assumption by studying a straightforward supervised pre-training setup: we simply pre-train an encoder network on ShapeNet with full supervision, and fine-tune it with a U-Net on a downstream task (S3DIS semantic segmentation). Following the practice in 2D representation learning, we use full supervision here as an upper bound to what could be gained from pre-training. We train a sparse ResNet-34 model (details to follow in Section 3.5) for 200 epochs. The model achieves a high validation accuracy of 85.4% on ShapeNet classification task. In Figure 1, we show the downstream task training curves for (a) training from scratch and (b) fine-tuning with ShapeNet pre-trained weights. Critically, one can observe that ShapeNet pre-training, even in the supervised fashion, hampers downstream task learning. Among many potential explanations, we highlight two major concerns:

Domain gap between source and target data: Objects in ShapeNet are synthetic, normalized in scale, aligned in pose, and lack scene context. This makes pre-training and fine-tuning data distributions drastically different.

Point-level representation matters: In 3D deep learning, the local geometric features, e.g. those encoded by a point and its neighbors, have proven to be discriminative and critical for 3D tasks . Directly training on object instances to obtain a global representation might be insufficient.

This led us to rethink the problem: if the goal of pre-training is to boost performance across many real-world tasks, exploring pre-training strategies on single objects might offer limited potential. (1) To address the domain gap concern, it might be beneficial to directly pre-train the network on complex scenes with multiple objects, to better match the target distributions; (2) to capture point-level information, we need to design a pretext task and corresponding network architecture that is not only based on instance-level/global representations, but instead can capture dense/local features at the point level.

2 Revisiting Fully Convolutional Geometric Features (FCGF)

Here we revisit a previous approach FCGF designed to learn geometric features for low-level tasks (e.g. registration) as our work is mainly inspired by FCGF. FCGF is a deep learning based algorithm that learns local feature descriptors on correspondence datasets via metric learning. FCGF has two major ingredients that help it stand out and achieve impressive results in registration recall: (1) a fully-convolutional design and (2) point-level metric learning. With a fully-convolutional network (FCN) design, FCGF operates on the entire input point cloud (e.g. full indoor or outdoor scenes) without having to crop the scene into patches as done in previous works; this way the local descriptors can aggregate information from a large number of neighboring points (up to the extent of receptive field size). As a result, point-level metric learning becomes natural. FCGF uses a U-Net architecture that has a full-resolution output (i.e. for NN points, the network outputs NN associated feature vectors), and positive/negative pairs for metric learning are defined at the point level.

Despite having a fundamentally different goal in mind, FCGF offers inspirations that might address the pretext task design challenges: A fully-convolutional design will allow us to pre-train on the target data distributions that involve complex scenes with a large number of points, and we could define the pretext task directly on points. Under this perspective, we pose the question: Can we repurpose FCGF as the pretext task for high-level 3D understanding?

3 PointContrast as a Pretext Task

FCGF focuses on local descriptor learning for low-level tasks only. In contrast, a good pretext task for pre-training aims to learn network weights that are universally applicable and useful to many high-level 3D understanding tasks. To take the inspiration of FCGF and create such pretext tasks, several design choices need to be revisited. In terms of architecture, since inference speed is a major concern in registration tasks, the network used in FCGF is very light-weight; Contrarily, the success of pre-training relies on over-parameterized networks, as clearly evidenced in other domains . In terms of dataset, FCGF uses domain-specific registration datasets such as 3DMatch and KITTI odometry , which lack both scale and generality. Finally, in terms of loss design, contrastive losses explored in FCGF are tailored for registration and it is interesting to explore other alternatives.

In Algorithm 1, we summarize the overall pretext task framework explored in this work. We name the framework PointContrast, since the high-level strategy of this pretext task is, contrasting—at the point level—between two transformed point clouds. Conceptually, given a point cloud x sampled from a certain distribution, we first generate two views x1\textbf{x}^{1} and x2\textbf{x}^{2} that are aligned in the same world coordinates. We then compute the correspondence mapping MM between these two views. If (i,j)∈M(i,j)\in M then point xi1\textbf{x}_{i}^{1} and point xj2\textbf{x}_{j}^{2} are a pair of matched points across two views. We then sample two random geometric transformations T1T_{1} and T2T_{2} to further transform the point clouds into two views. The transformation is what could make the pretext task challenging as the network needs to learn certain equivariance to the geometric transformation imposed. In this work, we mainly consider rigid transformation including rotation, translation and scaling. Further details are provided in Appendix. Finally, a contrastive loss is defined over points in two views: we minimize the distance for matched points and maximize the distance of unmatched points. This framework, though coming from a very different motivation (metric learning for geometric local descriptors), shares a strikingly similar pipeline with recent contrastive-based methods for 2D unsupervised visual representation learning . The key difference is that most work for 2D focuses on contrasting instances/images, while in our work the contrastive learning is done densely at the point level.

4 Contrastive Learning Loss Design

The first loss function, hardest-contrastive loss we try, is borrowed from the best-performing loss design proposed in FCGF , which adopts a hard negative mining scheme in traditional margin-based contrastive learning formulation,

Here P\mathcal{P} is a set of matched (positive) pairs of points xi1\textbf{x}_{i}^{1} and xj2\textbf{x}_{j}^{2} from two views x1\textbf{x}^{1} and x2\textbf{x}^{2}, and fi1\textbf{f}_{i}^{1} and fj2\textbf{f}_{j}^{2} are associated point features for the matched pair. N\mathcal{N} is a randomly sampled set of non-matched (negative) points which is used for the hardest negative mining, where the hardest sample is defined as the closest point in the L2\mathcal{L}_{2} normalized feature space to a positive pair. [x]+[x]_{+} denotes function max⁡(0,x)\max(0,x). mp=0.1m_{p}=0.1 and mn=1.4m_{n}=1.4 are margins for positive and negative pairs.

4.2 PointInfoNCE Loss

Here we propose an alternative loss design for PointContrast. InfoNCE proposed in is widely used in recent unsupervised representation learning approaches for 2D visual understanding. By modeling the contrastive learning framework as a dictionary look-up process , InfoNCE poses contrastive learning as a classification problem and is implemented with a Softmax loss. Specifically, the loss encourages a query qq to be similar to its positive key k+k^{+} and dissimilar to, typically many, negative keys k−k^{-}. One challenge in 2D is to scale the number of negative keys .

However, in the domain of 3D, we have a different problem: usually the real-world 3D datasets are much smaller in terms of instance count, but the number of points for each instance (e.g. a indoor or outdoor scene) can be huge, i.e. 100100K+ points even from one RGB-D frame. This unique property of 3D data property, together with the original motivation to modelling point level information, inspire us to propose the following PointInfoNCE loss:

Here P\mathcal{P} is the set of all the positive matches from two views. In this formulation, we only consider points that have at least one match and do not use additional non-matched points as negatives. For a matched pair (i,j)∈P(i,j)\in\mathcal{P}, point feature fi1\textbf{f}^{1}_{i} will serve as the query and fj2\textbf{f}^{2}_{j} will serve as the positive key k+k^{+}. We use point feature fk2\textbf{f}^{2}_{k} where ∃(⋅,k)∈P\exists(\cdot,k)\in\mathcal{P} and k≠jk\neq j as the set of negative keys. In practice, we sample a subset of 4096 matched pairs from P\mathcal{P} for faster training.

Compared to hardest-contrastive loss, the PointInfoNCE loss has a simpler formulation with fewer hyperparameters. Perhaps more importantly, due to a large number of negative distractors, it is more robust against mode collapsing (features collapsed to a single vector) than the hardest-contrastive loss. In our experiments, we find that hard-contrastive loss is unstable and hard to train: the representation often collapses with extended training epochs (which is also observed in FCGF ).

5 A Sparse Residual U-Net as Shared Backbone

We use a Sparse Residual U-Net (SR-UNet) architecture in this work. It is a 34-layer U-Net architecture that has an encoder network of 21 convolution layers and a decoder network of 13 convolution/deconvolution layers. It follows the 2D ResNet basic block design and each conv/deconv layer in the network is followed by Batch Normalization (BN) and ReLU activation. The overall U-Net architecture has 37.8537.85M parameters. We provide more information and a visualization of the network in Appendix. The SR-UNet architecture was originally designed in that achieved significant improvement over prior methods on the challenging ScanNet semantic segmentation benchmark. In this work, we explore if we can use this architecture as a unified design for both the pre-training task and a diverse set of fine-tuning tasks.

6 Dataset for Pre-training

For local geometric feature learning approaches, including FCGF , training and evaluation are typically conducted on domain and task-specific datasets such as KITTI odometry or 3DMatch . Common registration datasets are typically constrained in either scale (training samples collected from just dozens of scenes), or generality (focusing on one specific application scenario, e.g. indoor scenes or LiDAR scans for self-driving cars), or both. To facilitate future research on 3D unsupervised representation learning, in our work we utilize the ScanNet dataset for pre-training, aiming to address the scale issue. ScanNet is a collection of ∼\sim1500 indoor scenes. Created with a light-weight RGB-D scanning procedure, ScanNet is currently the largest of its kind.Admittedly, ScanNet is still much smaller in scale compared to 2D datasets.

Here we create a point cloud pair dataset on top of ScanNet for the pre-training framework shown in Figure 2. Given a scene x, we extract pairs of partial scans x1\textbf{x}^{1} and x2\textbf{x}^{2} from different views. More precisely, for each scene, we first sub-sample RGB-D scans from the raw ScanNet videos every 25 frames, and align the 3D point clouds in the same world coordinates (by utilizing estimated camera poses for each frame). Then we collect point cloud pairs from the sampled frames and require that two point clouds in a pair have at least a 30%30\% overlap. We sample a total number of 870K point cloud pairs. Since the partial views are aligned in ScanNet scenes, it is straightforward to compute the correspondence mapping MM between two views with nearest neighbor search.

Although ScanNet only captures indoor data distributions, as we will see in Section 4.4, surprisingly it can generalize to other target distributions. We provide additional visualizations for the pre-training dataset in Appendix.

Fine-tuning on Downstream Tasks

The most important motivation for representation learning is to learn features that can transfer well to different downstream tasks. There could be different evaluation protocols to measure the usefulness of the learned representation. For example, probing with a linear classifier , or evaluating in a semi-supervised setup . The supervised fine-tuning strategy, where the pre-trained weights are used as the initialization and are further refined on the target downstream task, is arguably the most practically meaningful way of evaluating feature transferability. with this setup, good features could directly lead to performance gains in downstream tasks.

Under this perspective, in this section we perform extensive evaluations of the effectiveness of PointContrast framework by fine-tuning the pre-trained weights on multiple downstream tasks and datasets. We aim to cover a diverse suite of high-level 3D understanding tasks of different natures such as semantic segmentation, object detection and classification. In all experiment, we use the same backbone network, pre-trained on the proposed ScanNet pair dataset (Section 3.6) using both PointInfoNCE and Hardest-Constrastive objectives.

In Section 3.1 we have observed that weights learned on supervised ShapeNet classification are not able to transfer well to scene-level tasks. Here we explore the opposite direction: Are PointContrast features learned on ScanNet useful for tasks on ShapeNet? To recap, ShapeNet is a dataset of synthetic 3D objects of 55 common categories. It was curated by collecting CAD models from online open-sourced 3D repositories. In , part annotations were added to a subset of ShapeNet models segmenting them into 2-5 parts. In order to provide a comparison with existing approaches, here we utilize the ShapeNetCore dataset (SHREC 15 split) for classification, and the ShapeNet part dataset for part segmentation, respectively. We uniformly sample point clouds of 1024 points from each model for classification and 2048 points for part segmentation. Albeit containing overlapping indoor object categories with ScanNet, this dataset is substantially different as it is synthetic and contains only single objects. We also follow recent works on 3D unsupervised representation learning to explore a more challenging setup: using a very small percentage (e.g. 1%-10%) of training data to fine-tune the pre-trained model.

As shown in Table 2 and Table 3, for both datasets, the effectiveness of pre-training are correlated with the availability of training data. In the ShapeNet classification task (Table 2), pre-training helps most where less training data is available, achieving a 4.0%4.0\% improvement over the training-from-scratch baseline with the hardest-negative objective. We also note that ShapeNet is a class-imbalanced dataset and the minority (tail) classes are very infrequent. When using 100% of the training data, pre-training provides a class-balancing effect, as it boosts performance more on underrepresented (tail) classes. Table 3 shows a similar effects of pre-training on part segmentation performance. Notably, using SR-UNet backbone architecture already boosts performance; yet, pre-training is able to provide further gains, especially when training data is scarce.

2 S3DIS Segmentation

Stanford Large-Scale 3D Indoor Spaces (S3DIS) dataset comprises 3D scans of 6 large-scale indoor areas collected from 3 office buildings. The scans are represented as point clouds and annotated with semantic labels of 13 object categories. Among the datasets used here for evaluation S3DIS is probably the most similar to ScanNet. Transferring features to S3DIS represents a typical scenario for fine-tuning: the downstream task dataset is similar yet much smaller than the pre-training dataset. For the commonly used benchmark split (“Area 5 test”), there are only about 240 samples in the training set. We follow for pre-processing, and use standard data augmentations. See Appendix for details.

Results are summarized in Table 4. Again, merely switching the SR-UNet architecture, training from scratch already improves upon prior art. Yet, fine-tuning the features learned by PointContrast achieves markedly better segmentation results in mIoU and mAcc. Notably, the effect persists across both loss types, achieving a 2.7% mIoU gain using Hardest-Contrastive loss and an on-par improvement of 2.1% mIoU for the PointInfoNCE variant.

3 SUN RGB-D Detection

We now attend to a different high-level 3D understanding task: object detection. Compared to segmentation tasks that estimate point labels, 3D object detection predicts 3D bounding boxes (localization) and their corresponding object labels (recognition). This calls for an architectural modification as the SR-UNet architecture does not directly output bounding box coordinates. Among many different choices , we identify the recently proposed VoteNet as a good candidate for three main reasons. First, VoteNet is designed to work directly on point clouds with no additional input (e.g. images). Second, VoteNet originally uses PointNet++ as the backbone architecture for feature extraction. Replacing this with a SR-UNet requires a minimal modification, keeping the proposal pipeline intact. In particular, we reuse the same hyperparameters. Third, VoteNet is the current state-of-the-art method that uses geometric features only, rendering an improvement markedly useful. We evaluate the detection performance on the SUN RGB-D dataset , a collection of single view RGB-D images. The train set contains 5K images annotated with amodal, 3D oriented bounding boxes for objects from 37 categories.

We summarize the results in Table 5. We find that by simply switching in the backbone network, our baseline result (training from scratch) with the SR-UNet architecture achieves worse results (-1.4% mAP@0.25). This may be attributed to the fact that VoteNet design and hyperparameter settings were tailored to its PointNet++ backbone. However, PointContrast gracefully closes the gap by showing a +3.1% gain on mAP@0.5, which also sets a new state-of-the-art in this metric. The performance gain with a harder evaluation metric (mAP@0.5) suggests that PointContrast pre-training can greatly help localization.

4 Synthia4D Segmentation

Synthia4D is a large synthetic dataset designed to facilitate the training of deep neural networks for visual inference in driving scenarios. Photo-realistic renderings are generated from a virtual city, allowing dense and precise annotations of 13 semantic classes, together with pixel-accurate depth. We follow the train/val/test split as prescribed by in the clean setting. In the context of this work, Synthia4D is especially interesting since it is probably the most distant from our pre-training set (outdoor v.s. indoor, synthetic v.s. real). We test the segmentation performance using 3D SR-UNet on a per-frame basis.

PointContrast pre-training brings substantial improvement over the baseline model trained from scratch (+2.3% mIoU) as seen in Table 6. PointInfoNCE performs noticeably better than the hardest-contrastive loss. With unsupervised pre-training, the overall results are much better than the previous state-of-the-art reported in . Note that in it has been shown that adding the temporal learning (i.e. using a 4D network instead of a 3D one) brings additional benefit. To use 3D pre-trained weights for a 4D network with an additional temporal dimension, we can simply inflate the convolutional kernels, following the standard practice in 2D video recognition . We leave it as future work.

5 ScanNet: Segmentation and Detection

Although typically the source dataset for pre-training and the target dataset for fine-tuning are different, because of the specific multi-view contrastive learning pipeline for pre-training, PointContrast can likely learn different representations (e.g. invariance/equivariance to rigid transformations or robustness to noise) compared to directly training with supervision. Thus it is interesting to see whether the pre-trained weights can further improve the results on ScanNet itself. We use ScanNet semantic segmentation and object detection tasks to test our hypothesis. For the segmentation experiment, we use the SR-UNet architecture to directly predict point labels. For the detection experiment, we again follow VoteNet and simply switch the original backbone network with the SR-UNet without other modifications to the detection head (See Appendix for details).

Results are summarized in Table 7 and Table 8. Remarkably, on both detection and segmentation benchmark, models pre-trained with PointContrast outperform those trained from scratch. Notably, PointInfoNCE objective performs better than the Hardest-contrastive one, achieving a relative improvement of +1.9% in terms of segmentation mIoU and +2.6% in terms of detection mAP@0.5. Similar to SUN RGB-D detection, here we also observe that PointContrast features help most for localization as indicated by the larger margin of improvement for mAP@0.5 than mAP@0.25.

6 Analysis Experiments and Discussions

In this section, we show additional experiments to provide more insights on our pre-training framework. We use S3DIS segmentation for the experiments below.

While the focus of this work is unsupervised pre-training, a natural baseline is to compare against supervised pre-training. To this end, we use the training-from-scratch baseline for the segmentation task on ScanNetV2 and fine-tune the network on S3DIS. This yields an mIoU of 71.2%, which is only 0.3% better than PointContrast unsupervised pre-training. We deem this a very encouraging signal that suggests that the gap between supervised and unsupervised representation learning in 3D has been mostly closed (cf. years of effort in 2D). One might argue that this is due to the limited quality and scale of ScanNet, but even at this scale the amount of labor involved in annotating thousands of rooms is large. The outcome of this complements the conclusion we had so far: not only should we put resources into creating large-scale 3D datasets for pre-training; but if facing a trade-off between scaling the data size and annotating it, we should favor the former.

A recent study in 2D vision suggests that simply by training from scratch for more epochs might close the gap from ImageNet pre-training. We conduct additional experiments to train the network from scratch with 2×2\times and 3×3\times schedules on S3DIS, relative to the 1×1\times schedule of our default setup (10K iterations with batch size 48). We found that validation mIoU does not improve with longer training. In fact, the model exhibits overfitting due to the small dataset size, achieving 66.7%66.7\% and 66.1%66.1\% mIoU at 20K and 30K iteration, respectively. This suggests that potentially many of the 3D datasets could fall into the “breakdown regime” where network pre-training is essential for good performance.

To show that the multi-view design in PointContrast is important, we try a different variant where instead of having partial views x1\mathbf{x}^{1} and x2\mathbf{x}^{2}, we directly use the reconstructed point cloud x\mathbf{x} (a full scene in ScanNet) PointContrast. We still apply independent transformations T1T_{1} and T2T_{2} to the same x\mathbf{x}. We tried different variants and augmentations such as random cropping, point jittering, and dropout. We also tried different transformations for T1T_{1} and T2T_{2} of different degrees of freedom. However, with the best configuration we can get a validation mIoU on S3DIS of 68.3568.35, which is just slightly better than the training from scratch baseline of 68.1768.17. This suggests that the multi-view setup in PointContrast is critical. Potential reasons include: much more abundant and diverse training samples; natural noise due to camera instability as good regularization, as also observed in .

Conclusions

We have demonstrated an extensive evaluation of the transferability of learned representations in 3D point clouds to high-level 3D understanding tasks. With the help of our unsupervised pre-training framework PointContrast, we achieve state-of-the-art results across 6 different benchmarks and demonstrate that the learned representation can generalize across domains. We hope these findings will encourage more research on 3D representation learning.

The authors would like to thank Chris Choy for his help in setting up the segmentation experiments with MinkowskiEngine. O.L. and L.G. were supported in part by NSF grant IIS-1763268, a Vannevar Bush Faculty Fellowship, and a grant from the SAIL-Toyota Center for AI Research.

References

Appendix 0.A Visualization of the SR-UNet Architecture

Here we show the SR-UNet architecture that is used as a shared backbone in our paper for both the pre-training and the fine-tuning phases. This U-Net architecture was originally proposed in for ScanNet semantic segmentation.

Appendix 0.B Visualization of the ScanNet Point Cloud Pair Dataset

Appendix 0.C ShapeNet Supervised Training Details

We use a sparse ResNet network that has an identical structure to the encoder part of the SR-UNet in Appendix 0.A. We use Adam optimizer, and add standard data augmentations including rotation, scaling and translation, following . We perform a grid search over the learning rate, weight decay, voxel size (for sparse convolution), and the number of input points. The best performing model configuration is learning rate 0.004, voxel size 0.01, weight decay 1e-5, batch size 512 and 2048 input points. The 85.4% accuracy is to our knowledge the best results that have been reported on this SHREC benchmark split. We use 8 Titan-V100 GPU with data parallelism to train the model. We train the model for 200 epochs and the training takes around 8 hours.

Appendix 0.D Details on PointContrast Pre-training

The transformations T1\mathbf{T}_{1} and T2\mathbf{T}_{2} applied to two views x1\mathbf{x}^{1} and x2\mathbf{x}^{2} in our experiments involves a random rotation ( to 360∘360^{\circ}) along an arbitrary axis (applied independently to both views). We apply scale augmentation to both views (0.8×\times to 1.2×\times of the input scale). We have experimented with other augmentations such as translation, point coordinate jittering, and point dropout and did not find noticeable difference in fine-tuning performances.

D.2 Details on Loss Functions

For the hardest-contrastive loss, the positive sample size is 1024 and the hardest negative sample size is 256. More details can be found in . For the PointInfoNCE loss, we provide a detailed PyTorch-like pseudo-code (and explanatory comments) in Algorithm 2.

Appendix 0.E S3DIS Segmentation Experimental Details

Here we provide training details for S3DIS semantic segmentation task. We use the widely adopted Area 5 Test (Fold 1) split for training and testing. For all the PointContrast variants (Training from scratch, Hardest-contrastive Pretrained, and PointInfoNCE Pretrained) we use the same hyperparameter settings. Specifically we train the model with 8 V100 GPUs with data parallelism for 10,000 iterations. Batch size is 48. Batch normalization is applied independently on each GPU. We use SGD+momentum optimizer with an initial learning rate 0.8. We use Polynomial LR scheduler with a power factor of 0.9. Weight decay is 0.0001 and voxel size is 0.05 (5cm). We use the same data augmentation techniques in such as color hue/saturation augmentation and jittering, as well as scale augmentations (0.9×\times to 1.1×\times). In Table 9 we show detailed per-category performance breakdown for our models and previous approaches.

Appendix 0.F Synthia4D Segmentation Experimental Details

Here we provide training details for Synthia4D semantic segmentation task. As mentioned in the main paper, we only use 3D sparse convnet without any temporal aggregation mechanisms such as 4D kernels and temporal CRF. For all the PointContrast variants (Training from scratch, Hardest-contrastive Pretrained, and PointInfoNCE Pretrained) we use the same hyperparameter settings, and those are mostly identical the S3DIS experiments. Specifically we train the model with 8 V100 GPUs with data parallelism for 15,000 iterations. Batch size is 72. Batch normalization is applied independently on each GPU. We use SGD+momentum optimizer with an initial learning rate 0.8. We use Polynomial LR scheduler with a power factor of 0.9. Weight decay is 0.0001 and voxel size is 0.05 (5cm). We also use the same data augmentation techniques in in color space and point coordinate space. In Table 10 we show detailed per-category performance breakdown for our models and results reported in .

Appendix 0.G ScanNet Segmentation Experimental Details

For ScanNet segmentation task, we train the model with 8 V100 GPUs with data parallelism for 15,000 iterations. Batch size is 48. We use SGD+momentum optimizer with an initial learning rate 0.8. We use Polynomial LR scheduler with a power factor of 0.9. Weight decay is 0.0001 and voxel size is 0.025 (2.5cm). We also use the same data augmentation techniques in in color space and point coordinate space. In Table 10 we show detailed per-category performance breakdown for our models and results reported in .

Appendix 0.H ScanNet and SUN RGB-D Detection Details

For the 3D object detection experiments, we mostly follow the configurations in VoteNet framework after switching in the SR-UNet backbone architecture. We train the model on 1 GPU, with batch size 64 for SUN RGB-D and 32 for ScanNet. Learning rate is 0.001 and we use Adam optimizer. The input points are subsampled before voxelization, we use 20000 points for SUN RGB-D and 40000 points for ScanNet. The voxel size is 2.5cm for ScanNet and 5cm for SUN RGB-D. In Table 15 we show more results reported by previous methods. In Table 12 and Table 13, we show per-category AP performance for PointContrast models agains training from scratch results, under AP@0.5 metric.

Appendix 0.I PointContrast vs FCGF for low- and high-level tasks

We take the best performing FCGF model released in that achieves a high registration feature matching recall (FMR) of: 0.958. However, this model does not perform well for S3DIS segmentation. On the other hand, the PointContrast model that performs best for segmentation achieves a lower FMR when applied to the registration task. We conclude that low-level tasks and high-level tasks in 3D might require different design choices.