Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic Segmentation

Li Jiang, Shaoshuai Shi, Zhuotao Tian, Xin Lai, Shu Liu, Chi-Wing Fu, Jiaya Jia

Introduction

3D point cloud semantic segmentation is a fundamental and essential perception task for many downstream applications . Existing deep-learning-based methods for the task heavily rely on the availability and quantity of labeled point cloud data for the model training. However, 3D point-level labeling is time-consuming and labor-intensive. Compared with point cloud labeling, point cloud collection requires much less effort, mainly by means of 3D scanning followed by some data post-processing. Hence, we are motivated to explore semi-supervised learning (SSL) for improving the data efficiency and performance of deep segmentation models with unlabeled point clouds.

While SSL has been widely explored for tasks on 2D images , it is rather underexplored for 3D point clouds. To achieve SSL, a common strategy is consistency regularization , which aligns features of the same image/pixel under different perturbations for maintaining the prediction consistency when exploiting unlabeled data. Our method shares this common ground in SSL by encouraging similar and robust features for matched 3D point pairs with different transformations. Yet, inspired by the contrastive loss applied in self-supervised learning , we further enhance the feature representation by proposing the guided point contrastive loss to additionally enlarge the distance between inter-category features by using the semantic predictions as guidance in the semi-supervised setting.

Contrastive learning starts with works on 2D images, and is recently extended by PointContrast to 3D point clouds as a pre-training pretext task in a self-supervised setting. The point contrastive loss encourages the matched positive point pairs to be similar in the embedding space while pushing away the negative pairs. Yet, without any label, negative pairs in the same category may also be sampled, especially for large objects (e.g., sofa) and redundant background classes (e.g., floor and wall); these negative pairs actually weaken the features’ discriminative ability. Unlike PointContrast, we leverage a few labeled point clouds to optimize the network model for producing point-level semantic predictions, and meanwhile, utilize the predicted semantic scores and labels for the unlabeled data to guide the contrastive loss computation. Our pseudo-label guidance helps alleviate the side effect of intra-class negative pairs in feature learning, while our confidence guidance utilizes the semantic scores to reduce the chance of feature worsening. Also, we propose a category-balanced sampling strategy to exploit pseudo labels to mitigate the class imbalance issue in point sampling, helping to preserve point samples from rare categories and to improve the feature diversity in contrastive learning. As revealed in the t-SNE visualizations in Fig. 1, the model equipped with our pseudo guidance learns more discriminative point-wise features.

We follow the conventional practice in SSL to conduct experiments with a small portion of labeled data and a larger portion of unlabeled data and then evaluate how effective an SSL method improves the performance with the unlabeled data. Excellent performance for both indoor (ScanNet V2 and S3DIS ) and outdoor (SemanticKITTI ) scenes are obtained, showing the effectiveness of our semi-supervised method, which surpasses the supervised-only models with 5%, 10%, 20%, 30%, and 40% labeled data by a large margin consistently on all three datasets. Also, we experiment with 100% labeled data, in which the labeled set is also fed into the unsupervised branch with our guided point contrastive loss as an auxiliary feature learning loss. In this case, the accuracy of our method still exceeds the baseline with only supervised cross entropy loss, showing that without extra unlabeled data, our guided point contrastive loss also helps to refine the feature representation and model’s discriminative ability.

We adopt semi-supervised learning to 3D scene semantic segmentation, demonstrating that unlabeled point clouds can help to enhance the feature learning in both indoor and outdoor scenes.

We extend contrastive learning to 3D point cloud semi-supervised semantic segmentation with pseudo-label guidance and confidence guidance.

We propose a category-balanced sampling strategy to alleviate the point class imbalance issue and to increase the embedding diversity.

Related Works

Point cloud segmentation. Various approaches have been explored for 3D semantic segmentation. Voxel-based approaches utilize 3D convolutional neural networks by transforming irregular point clouds to regular 3D grids. Other approaches explore the sparsity of voxels for high-resolution 3D representations with OctNet or sparse convolution . Pioneered by PointNet , point-based approaches directly learn point features from raw point clouds with assorted hierarchical local feature aggregation strategies . KPConv defines a kernel function on points for conducting convolutions on local points. There are also works, e.g., , that incorporate graph convolutions for point feature learning.

To train the network model, these fully-supervised approaches require data with point-wise labels, which are time-consuming and tedious to prepare as well as error-prone. Hence, in this work, we incorporate unlabeled point clouds in network training for improving the data efficiency in 3D point cloud semantic segmentation.

Semi-supervised learning (SSL) aims to improve a model by learning from unlabeled data, in addition to labeled data. Existing works on SSL mainly focus on image classification and image semantic segmentation . Consistency regularization is a common strategy for SSL, emphasizing that the model predictions should be consistent for different perturbations applied to the same input. Π\Pi-model , a simplified version of Γ\Gamma-model , encourages consistent model outputs for different dropouts and augmentations on the same input, while temporal ensembling and Mean Teacher adopt the exponential moving average strategy to stabilize the predictions for consistency regularization.

Recent SSL methods for image segmentation show that pixel-wise consistency could be achieved by perturbing the input images or the intermediate features , or by feeding the same image to different models . Pseudo-label-based self-training is another approach for SSL, in which we first train a model with labeled data then refine it by generating pseudo labels on the unlabeled data for further training. Some other works also adopt a generative adversarial network for SSL image segmentation to incorporate unlabeled images for learning.

Though many SSL works have been proposed for images, SSL for point cloud scenes is rather underexplored. Currently, there are two works for 3D detection that leverage unlabeled scenes by Mean-Teacher framework or by quality-aware pseudo labeling . Compared with 3D box annotations, point-wise dense annotations for 3D point cloud segmentation are more resource-intensive. Hence, we propose a novel SSL framework for the task, demonstrating the feasibility of incorporating unlabeled point clouds to improve the performance of segmenting 3D points.

Contrastive Learning is a widely-used approach for unsupervised learning . Its core idea is the contrastive loss that encourages the features of the query samples to be similar to those of the positive key samples, while being dissimilar with those of the negative key samples. A common choice of contrastive loss is InfoNCE , which measures the similarity by a dot product. PointContrast proposes the PointInfoNCE loss for point-level unsupervised representation learning, and their follow-up work proposes a ShapeContext-like spatial partition for location-aware contrastive learning. Supervised contrastive learning is also proposed recently to better align the intra-class features with labeled data. In this work, we extend contrastive learning to support semi-supervised point cloud segmentation and propose to incorporate point-wise pseudo labels to the contrastive loss for better distinguishing the positive and negative samples and collectively using both the labeled and unlabeled point clouds for learning a more effective representation.

Our Approach

The key in SSL is to learn feature representation from unlabeled data, which is a common goal shared by both unsupervised learning and SSL . When it comes to 3D semantic segmentation, feature learning in point level is of critical importance. Hence, we begin this section by revisiting and analyzing label-free point-level feature learning in an unsupervised setting and then in SSL.

Point-level contrast in self-supervised learning. PointContrast firstly proposes the point-level self-supervised strategy for pre-training with unlabeled point clouds. It extends the InfoNCE loss to points as PointInfoNCE loss for contrastive learning on 3D scenes:

where Mp\mathbf{M}_{p} is the index set of randomly-sampled positive pairs (one-to-one matched points) across two point clouds perturbed from the same input; Eu1\mathbf{E}^{u1} and Eu2\mathbf{E}^{u2} are feature embeddings of the two point clouds; and τ\tau is a temperature hyperparameter. For point ii in the first point cloud u1u1, (i,j)∈Mp(i,j)\in\mathbf{M}_{p} is a positive pair, whose feature embeddings (Eiu1,Eju2)(\mathbf{E}^{u1}_{i},\mathbf{E}^{u2}_{j}) are encouraged to be similar, while \{(i,k)|(\cdot,k)\in\mathbf{M}_{p},k<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo mathvariant="normal">≠</mo></mrow><annotation encoding="application/x-tex">\neq</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">=</span></span></span></span></span></span>j\} are negative point pairs. Point ii is called the anchor point; its feature embedding is enforced to be dissimilar with the feature embeddings of all its negative points. PointContrast serves as a pretext task for pre-training and validates the effectiveness of point-level contrastive loss in point cloud self-supervised learning.

Point-level consistency in SSL. Consistency regularization is a widely-used strategy to exploit unlabeled data for enhancing feature robustness. Hence, we define a simple baseline with consistency regularization. For point-level consistency, inspired by 2D SSL , one may enforce a corresponding point pair with different augmentations to have similar feature representation by minimizing the Mean-Squared Error (MSE) between the feature embeddings of the points. Formally, the loss in the unsupervised branch in SSL with MSE can be expressed as

where M\mathbf{M} is the index set of all matched point pairs across point clouds u1u1 and u2u2 perturbed from the same input. In SSL, we combine LuL_{u} with the following supervised cross entropy loss LlL_{l} on labeled data for model training:

where NlN^{l} is the number of points in the given labeled point cloud; Sl\mathbf{S}^{l} is the predicted semantic scores; and Yl\mathbf{Y}^{l} represents the ground-truth labels.

Discussion. Though both PointInfoNCE and our SSL baseline could learn from unlabeled point clouds and benefit 3D semantic segmentation (see Table 6), they have several drawbacks: (i) Negative point pairs of same category could worsen the feature learning: In the unsupervised setting of PointInfoNCE, a negative point pair (i,k)(i,k) may come from the same semantic category, so pushing away their embeddings (Eiu1,Eku2)(\mathbf{E}^{u1}_{i},\mathbf{E}^{u2}_{k}) could degrade the feature learning. (ii) Points from the same category are likely to be sampled, especially for large objects or for common categories such as road: Random sampling could easily produce unfavorable negative point pairs that actually come from the same category. (iii) Feature distance for both intra- and inter-class should be considered: In our SSL baseline, only paired intra-class features are constrained to be similar. However, the inter-class feature distance should also be enlarged to better improve the semantic segmentation.

To mitigate the above problems, we focus on exploring and leveraging the information from labeled point clouds to better guide the feature learning from unlabeled point clouds for improving the 3D scene semantic segmentation.

2 Pseudo Guidance on Contrastive Learning

Now, we focus on the setting of semi-supervised learning (SSL) for 3D point cloud semantic segmentation, in which we could leverage some labeled data to train the model to produce semantic predictions for unlabeled scenes. We accordingly propose the Guided Point Contrastive Learning framework for SSL-based point cloud segmentation and leverage the semantic predictions as pseudo guidance for improving the contrastive learning on unlabeled point clouds. Fig. 2 shows the overall architecture of our framework, which consists of a supervised branch and an unsupervised branch. In this section, we focus on our guided contrastive loss in the unsupervised branch.

Formally, for a pair of perturbed point clouds (Pu1,Pu2)(\mathbf{P}^{u1},\mathbf{P}^{u2}) from the same unlabeled data, we can generate their pseudo labels (Y^u1,Y^u2)(\hat{\mathbf{Y}}^{u1},\hat{\mathbf{Y}}^{u2}) and label confidences (Cu1,Cu2)(\mathbf{C}^{u1},\mathbf{C}^{u2}) from the predicted semantic scores (Su1,Su2)(\mathbf{S}^{u1},\mathbf{S}^{u2}) as follows:

where ∗* is u1u1 or u2u2 and σ\sigma denotes the softmax function.

We then denote Mp\mathbf{M}_{p} as the set of matched positive point pairs across point clouds u1u1 and u2u2 perturbed from the same input. For the negative point sets, instead of using points in Mp\mathbf{M}_{p} as in , we separately sample negative points to ensure negative samples can also come from un-matched regions. We denote the negative point sets sampled from point clouds u1u1 and u2u2 as Mnu1⊆{1,2,...,Nu1}\mathbf{M}_{n}^{u1}\subseteq\{1,2,...,N^{u1}\} and Mnu2⊆{1,2,...,Nu2}\mathbf{M}_{n}^{u2}\subseteq\{1,2,...,N^{u2}\}, respectively, where Nu1N^{u1} and Nu2N^{u2} denote the number of points in the associated point clouds.

Guided contrastive loss. Given the positive point pair set Mp\mathbf{M}_{p} and negative point sets Mnu1\mathbf{M}_{n}^{u1} and Mnu2\mathbf{M}_{n}^{u2}, our guided contrastive loss Lu(i,j)L_{u}^{(i,j)} for the positive point pair (i,j)∈Mp(i,j)\in\mathbf{M}_{p} could be represented as

where γ\gamma is a confidence threshold. Note that the loss is computed on Pu1\mathbf{P}^{u1} and Pu2\mathbf{P}^{u2} separately. For each side, the features from the other side is detached to stop gradients and is thus treated as constant references for better optimizing the features on the current side.

By leveraging the semantic predictions in Eq. (4), we propose two pseudo guidances in Eq. (5) to guide the feature learning from unlabeled point clouds, which are illustrated in Fig. 3 and discussed below:

Pseudo-label guidance: GG is the pseudo-label guidance for filtering negative point pairs with the same pseudo labels, which is defined as

As shown in Fig. 3, for an anchor point on ‘sofa’, many negative samples are also on ‘sofa’. Pushing away such an intra-category negative point pair could adversely affect the feature learning (Fig. 3 left). By incorporating our proposed pseudo-label guidance on contrastive learning, only negative feature pairs with different semantic predictions are forced to be dissimilar (Fig. 3 right).

The overall guided contrastive loss is the average of the losses for the positive pairs in Mp\mathbf{M}_{p}:

3 Category-balanced Sampling

The computation cost of the guided contrastive loss is highly correlated with the positive / negative point numbers. As there is a large number of points in each point cloud (e.g., around 100k - 1000k for an indoor scene and around 100k for an outdoor LiDAR frame), usually we could not take all the points as positive or negative samples. Hence, for data with imbalanced category distribution, some categories with a very small number of points may have little chance to be sampled with random sampling, while some large-portion categories are often sampled redundantly, thus affecting the feature diversity in contrastive learning. To mitigate this problem, we propose a simple but effective sampling strategy—Category-Balanced Sampling (CBS).

Category-balanced sampling for positive pairs. We define the category set as C\mathcal{C}, so the number of categories is ∣C∣|\mathcal{C}|. For a pair of perturbed point clouds, we denote the set of all matched point pairs as M\mathbf{M} and reorganize M\mathbf{M} by the predicted categories of the 1st1^{st} points in the pairs. The number of matched point pairs in each category c∈Cc\in\mathcal{C} is denoted as NMcN_{M}^{c}. To sample a total number of KpK_{p} positive point pairs from M\mathbf{M} to form Mp\mathbf{M}_{p} for our guided contrastive loss, our CBS strategy evenly samples positive pairs from each category. Specifically, for category cc, the number of selected positive pairs NMpcN_{M_{p}}^{c} is calculated as

Then, additional Kp−∑cNMpcK_{p}-\sum_{c}N_{M_{p}}^{c} pairs are sampled from all categories to ensure the total number of positive pairs is KpK_{p}.

Category-balanced sampling for negative point set. To enhance the sample diversity of negative points, we conduct CBS by collecting negative samples from scenes in the entire training set instead of just from the current scene, since some categories may even be absent in a specific scene or batch. Precisely, we maintain a category-aware negative embedding memory bank of size ∣C∣×B×CE|\mathcal{C}|\times B\times C_{E}, where BB is the bank length for each category and CEC_{E} is the channel number of feature embeddings. Each category in the memory bank is updated with a “First-In, First-Out” strategy to ensure the bank contains the latest feature embeddings. The number of updated embeddings for each category at each iteration is set to BuB_{u}. Then, to collect negative points, at each iteration, we evenly sample ⌊Kn∣C∣⌋\left\lfloor\frac{K_{n}}{|\mathcal{C}|}\right\rfloor points from each category’s memory bank to form a total number of KnK_{n} negative feature embeddings for contrastive learning.

Our proposed CBS strategy generates both category-balanced positive pairs and negative points. It enables a more effective contrastive feature learning from the unlabeled point clouds, as shown later in Table 5.

4 Overall Architecture

Overall objective. The overall objective of our semi-supervised framework is a combination of losses in the supervised and unsupervised branches:

where LuL_{u} is our guided contrastive loss to enhance the feature learning with the unlabeled point clouds; LlL_{l} is a common cross-entropy loss for semantic segmentation; and λ\lambda is a hyperparameter to adjust the loss ratio.

Experiments

We present evaluations on our guided contrastive learning framework with both indoor and outdoor scenes. We use the mean Intersection-over-Union (mIoU) and mean accuracy (mAcc) as the evaluation metrics in the experiments.

We use both indoor (ScanNet V2 and S3DIS ) and outdoor (SemanticKITTI ) datasets in our evaluations:

ScanNet V2 is a popular indoor 3D point cloud dataset that contains 1,613 3D scans with point-wise semantic labels. The whole data is split into a training set (1201 scans), a validation set (312 scans), and a testing set (100 scans). There are totally 20 categories for semantic segmentation.

S3DIS is another commonly-used indoor 3D point cloud dataset for semantic segmentation. It has 271 point cloud scenes across six areas, and there are in total 13 categories in the point-wise annotations. We follow the common split in previous works to utilize Area 5 as the validation set and adopt the other five areas as the training set.

SemanticKITTI is a large-scale outdoor point cloud dataset for 3D semantic segmentation in an autonomous driving scenario, where each scene is captured by the Velodyne-HDLE64 LiDAR sensor. The dataset contains 22 sequences that are divided into a training set (10 sequences with ∼\sim19k frames), a validation set (1 sequence with ∼\sim4k frames), and a testing set (11 sequences with ∼\sim20k frames). There are 19 categories for semantic segmentation.

SSL training set partition. Following the conventional practice in SSL, we employ existing datasets in our evaluations and split the training set into labeled and unlabeled sets with five different ratios of labeled data, i.e., {5%, 10%, 20%, 30%, 40%}. For SemanticKITTI, considering that adjacent frames could have very similar contents, when we split the dataset, we try our best to ensure that labeled and unlabeled data do not come from the same sequence. However, to achieve a specific labeled ratio, we may need to cut at most one sequence into two parts, the front part for labeled set and the latter part for unlabeled set.

1.2 Augmentations for Semi-Supervised Learning

We adopt random crop as one of our augmentation operations. Since indoor and outdoor scenes have very different point distributions due to the use of different capturing devices, we perform different crop operations on them.

Augmentations for indoor scenes. For indoor scenes, the crop augmentation is implemented by randomly cropping a square region of size 3.53.5m ×\times 3.53.5m in the top-down view. For each unlabeled scene, we crop it twice and guarantee an overlap between the two cropped point clouds to build a point-to-point correspondence in the overlapping region. Besides random crop, we adopt random rotation ( - 2π2\pi) around the z-axis (vertical axis) and random flip. Following the released code of , we also adopt the elastic operation.

Augmentations for outdoor scenes. For outdoor scenes, we propose a sector-range crop centered at the origin that follows the beam pattern in LiDAR point clouds. Specifically, we randomize a heading angle in range [0,2π][0,2\pi] as the center direction of the sector and further randomize a field-of-view angle in range [23π,2π][\frac{2}{3}\pi,2\pi] to form the cropping sector. For each unlabeled scene, two sectors are cropped with a guaranteed overlap for setting up a point-to-point correspondence. Besides the sector-based crop, we adopt the commonly-used random flip, random rotation (−π4-\frac{\pi}{4} - π4\frac{\pi}{4}), and random scale (0.95 - 1.05) augmentations.

1.3 Implementation Details

Network details. For both indoor and outdoor scenes, we utilize the sparse-convolution-based U-Net as the backbone network for 3D semantic segmentation. The encoder applies sparse convolution layers with a stride of 2 to downsample the input volume six times, while the decoder gradually upsamples the volume back to the original size with six deconvolutions. Submanifold sparse convolutions with a stride of 1 are used in the U-Net to encode the features. The projector is a multi-layer perception that maps the features to an embedding space. For voxelizing the input point clouds, the voxel size is set to 22cm for the indoor scenes and 1010cm for the outdoor scenes.

Training details. For ScanNet V2, we train our SSL framework from scratch using an SGD optimizer. The learning rate is initialized as 0.2 and decayed with the poly policy with a power of 0.9. The batch size is 16, i.e., 16 labeled scenes and 16 unlabeled scenes. For S3DIS, we apply the Adam optimizer with an initial learning rate of 0.02. We keep the same number of training iterations for different settings to train for 3636k iterations on ScanNet V2 and 88k iterations on S3DIS using eight GPUs. For a more stable semi-supervised training, we train the model with only supervised loss at the beginning 200 iterations. For outdoor scenes, the segmentation network is first pretrained on the labeled set by an Adam optimizer with a batch size of 48 and a learning rate of 0.02 for 1616k iterations on eight GPUs. Then, we train the network with our SSL framework on the labeled and unlabeled sets for another 1818k iterations by an Adam optimizer with a batch size of 24 and a learning rate of 0.002. The cosine annealing strategy is utilized to decay the learning rate. The loss ratio λ\lambda for the guided contrastive loss is set to 0.1, while the temperature τ\tau in the loss is set to 0.1. The confidence threshold γ\gamma is 0.75.

2 Main Results

To demonstrate the effectiveness of our method on exploiting unlabeled data, we take the strong sparse-convolution-based U-Net as our backbone and follow the conventional practice in SSL to compare our semi-supervised models with models that are fully trained with only labeled point clouds, separately using {5%, 10%, 20%, 30%, 40%} of the training set as the labeled data.

Table 1 summarizes the quantitative results on ScanNet V2, S3DIS, and SemanticKITTI in terms of mIoU and mAcc. For all three datasets, both indoor and outdoor, our semi-supervised models consistently outperform the supervised-only ones for all ratios, showing that our SSL model is able to effectively leverage the unlabeled data to improve the embedding features and so the segmentation performance. On all the datasets, the performance gap increases with the relative amount (ratio) of unlabeled data. Given 5% labeled data and 95% unlabeled data, our semi-supervised method improves the mIoU relatively by 13.9%, 17.8%, and 20.1% in ScanNet V2, S3DIS, and SemanticKITTI, respectively. Further, we present some qualitative results in Fig. 4, which shows that our model helps improve the segmentation quality with the unlabeled data.

Additionally, we conduct experiments on the 100% ratio, in which the whole training set is taken as the labeled set and simultaneously fed also into the unsupervised branch as the unlabeled set. Our guided point contrastive loss serves as an auxiliary constraint for feature learning in the 100% setting. As shown in the last column of Table 1, without extra unlabeled data, our method can still boost the network performance by enhancing the feature representation and model discriminative power via the guided contrastive learning in the unsupervised branch. We also compare our 100% results with recent state-of-the-art methods in Table 2. Our supervised-only baseline model is already competitive among these methods, while our approach can further improve the prediction quality, which achieves excellent performance on all three datasets.

Transductive learning on 100% ratio. Unlike inductive learning that aims at generalizing the model to unseen testing set, in transductive learning, the testing set is given and also observed in training. We extend the experiments on 100% ratio to a transductive form by incorporating the testing data as part of the unlabeled data. We observe that the performance gets higher in the transductive form, as shown in Table 3. The transductive model has comparable performance (70.2%) as the SOTA’s in the SemanticKITTI CodaLab benchmark (Single Scan).

3 Ablation Studies

Pseudo guidance & CBS. We conduct ablation studies on our pseudo guidance and CBS design using the 20% labeled data on ScanNet V2. Table 4 shows the contribution of each component in our method. Without the pseudo guidance, the effect of point contrastive loss in exploiting unlabeled data is limited. By avoiding potential intra-category negative pairs, our pseudo-label guidance provides the most significant performance gain (1.5%) over the vanilla point contrastive loss. Our confidence guidance, which promotes the feature learning quality, further improves the mIoU from 65.9% to 66.4%. Our full model with CBS to enhance feature diversity leads to the highest performance 66.7%.

CBS on different datasets. For further analyzing the effect of our CBS, we compare CBS with random sampling on ScanNet V2 (indoor) and SemanticKITTI (outdoor) with 20% labeled data. Table 5 reports the results. Compared with indoor scenes, outdoor data suffers more from the category imbalance problem. We count the number of points in each category of the training set for both datasets. In SemanticKITTI, 6 out of the 19 categories, i.e., ‘bicycle’, ‘motorcycle’, ‘person’, ‘bicyclist’, ‘motorcyclist’, and ‘traffic-sign’, have fewer than 1‰ points in the training set. The most rare ‘motorcyclist’ category has only 0.04‰. However, in ScanNet V2 training set, the point category distribution is more balanced; the most rare category ‘sink’ still has 2.75‰ points. Hence, our CBS contributes a larger increase for SemanticKITTI. For a category with very sparse points, it is hard to be sampled with random sampling. CBS increases the probability of selecting samples from these categories and thus improves the feature diversity.

Partial views vs. random crop. PointContrast suggests that the multi-view design is critical in improving the quality of the pretrained model. For a scene in ScanNet V2, the multi-view design samples two partial views of the scene instead of cropping the reconstructed point cloud. We also try this multi-view strategy in our unsupervised branch. However, the experimental results on ScanNet V2 (with 20% labeled data) show that applying partial views instead of random crop even lowers the performance from 66.7% to 63.2%. The possible reason for the performance drop is that the multi-view design could widen the discrepancy between the labeled and unlabeled sets, introducing imprecise semantic predictions in the unsupervised branch.

The projector. The projector is essential for contrastive learning, as discussed in . If we remove the projector, the performance (mIoU) drops from 66.7% to 65.0% on ScanNet V2 dataset with 20% labeled data.

4 Analysis on various Semi-supervised Strategies

Apart from our guided contrastive learning, we also experiment with some other semi-supervised strategies, and the performance comparison is shown in Table 6.

Consistency regularization. With consistency regularization as the semi-supervised strategy, we apply the MSE loss as in Eq. (2) or the cosine similarity loss to align the matched features under different perturbations. The results surpass the supervised-only model by 1.0% and 0.9%, respectively. Our method achieves better performance with an improvement of 2.7% with stronger constraints on the features by not only keeping the consistency between matched points but also pushing away the points with different semantic predictions in the embedding space.

PointInfoNCE loss. Directly applying the PointInfoNCE loss in PointContrast as the unsupervised loss does not perform well in the semi-supervised setting. The high probability of sampling negative pairs within the same category adversely affects the feature learning.

Self-training. Pseudo-label-based self-training is another alternative strategy for SSL. We have also tried it by first training a model with the labeled set, loading it to produce pseudo labels for the unlabeled set, and then constraining the semantic predictions in the unsupervised branch with the pseudo labels by a cross-entropy loss. Self-training also performs well in exploiting unlabeled data (65.5%), but our method can still attain a higher performance (66.7%).

Conclusion

In this work, we present a semi-supervised framework to take advantage of unlabeled point clouds to perform 3D semantic segmentation in a data-efficient manner. With our guided point contrastive loss, the network can learn more discriminative features by leveraging our pseudo-label and confidence guidance. Also, we propose the category-balanced sampling to benefit contrastive learning with more diverse feature embeddings. Experimental results show the effectiveness of our approach to exploit unlabeled 3D data and to improve the model generalization ability.

Acknowledgments The project is supported in part by the Research Grants Council of the Hong Kong Special Administrative Region (Project no. CUHK 14206320).

References