A new weakly supervised approach for ALS point cloud semantic segmentation
Puzuo Wang, Wei Yao
Introduction
As an important data source of active remote sensing, airborne laser scanning (ALS) data depicts a precise three-dimensional representation of large scale out-door scenes. While 3D coordinates and associated attributes (e.g. laser reflectance and return count information) are usually contained in ALS point clouds, to fully interpret a complex geographical scene, the key step is to acquire semantic information as a valuable cue utilized in a variety of remote sensing applications, such as land cover survey (Yan et al., 2015), forest monitor (Yao et al., 2012), change detection (Okyay et al., 2019) and 3D mapping (Zhang et al., 2018). Semantic segmentation, or classification, usually assigning a label to each point, is an indispensable solution for point cloud parsing .
Over past decades, point cloud semantic segmentation is always a research hot-spot in scene understanding. Initially, former studies focused on developing rule-based methods to distinguish different categories of land covers (Antonarakis et al., 2008; Zhou, 2013). Hand-crafted features were extracted based on characteristics of specific point clouds, and the hard threshold was used to classify different objects. While these methods were unsupervised, the performance relied heavily on effective feature extraction and suitable threshold settings, which makes them hard to generalize to new areas. The application of machine learning methods to the classification of point clouds improved the accuracy of results. In classical machine learning methods (Guo et al., 2011; Weinmann et al., 2015), well-designed hand-crafted features were still required, but they were fed into supervised classifiers to optimize model parameters automatically. The development of convolutional neural networks (CNNs) has apparently driven the progress of point cloud semantic segmentation tasks, and cutting-edge methods based on deep learning have been newly developed, achieving state-of-the-art results (Charles et al., 2017; Qi et al., 2017; Wang et al., 2019; Thomas et al., 2019). Most of them have focused on designing new network structures or convolution kernels based on the characteristics of point cloud data, without taking into account the high costs paid to secure the availability of labels. Actually, the supervised methods rely on a large amount of precise data annotations, which triggers the issue of data hungry (Gao et al., 2021). Data hungry refers to the demand on large number of labeled data for supervised learning to achieve leading results. To collect precise annotations is usually associated with heavy workloads, even requiring extremely meticulous efforts for an expert operator to complete the task. Moreover, the labeling of ALS point clouds is particularly difficult, usually demanding the operator to confirm the category of a point from multiple perspectives. The occlusions caused by the scan pattern of ALS systems lead to data voids, which makes it even harder to determine the exact label of points in occluded areas. In addition, the discrete data structure of point clouds in 3D space increases the difficulty of visual interpretation.
Although data labeling is a difficult and time-consuming job, the collection cost of massive unlabeled point cloud data is greatly reduced, thanks to the advances in LiDAR technology and diversified data acquisition platforms. So far, multiple ALS point clouds benchmark datasets have been released for the task of semantic segmentation. For instance, a new benchmark, Hessigheim 3D Benchmark (H3D) (Kölle et al., 2021), contains tens of millions of high-resolution 3D point clouds, the density of which is about 800 pts/m². Confronted with massive point clouds subject to annotation, we naturally ask a question, whether promising results can be achieved without the necessity to label the entire scene, and how is the performance of methods affected under such condition? An illustrated comparison of fully labeled and sparsely labeled data is shown in Fig. 1. If competitive classification results could be achieved by only using incomplete labels, the workload of data annotation will be considerably reduced, contributing significantly to the efficiency in real-life applications.
Semi- and weakly supervised learning are commonly used methods addressing situations, in which the scarcity of labels prevails. As two concepts are often mixed up in studies, we use weakly supervised learning or weak supervision in this study to represent the situation in which label information is incomplete and deficient. Recently, comprehensive reviews of weakly supervised learning methods are performed (Zhou, 2017; Van Engelen and Hoos, 2020). Under weak supervision, it can be categorized into two types, transductive and inductive mode. For transductive one, unlabeled data are exactly test one and also used as co-training set in parallel to parameterize the model. In contrast, inductive learning follows the normal process of a full supervision scheme, training a model separately with labeled samples and generalizing to unseen areas. The mode of inductive learning is considered for investigation in our study due to high applicability. By applying a weakly supervised method, it is meaningful to choose a suitable weak label annotation strategy for the purpose of reducing labeling workload. In image processing, weak labels are represented as a few labeled images (Dong and Xing, 2018), a few labeled pixels (Bearman et al., 2016), labels of image patches (Yao et al., 2016), or a few bounding boxes or categories that appear in images (Kolesnikov and Lampert, 2016). For conventional computer vision tasks, image-level labels are the most accessible annotation, whereby adding a portion of labeled pixels could further contribute to an improved classification result. However, it is different for ALS point cloud processing, which usually covers a large-area earth surface. The data has to be subdivided into small blocks to enable training. Thus, compared with annotating scene-level labels for large number of extracted subsets, directly labeling few points across the whole data seems more practical. Given a fixed budget number of labeled points, how to choose an effective labeling strategy is another issue. Xu and Lee (2020) presented theoretical and experimental analysis to demonstrate the superiority of discrete labels against conventional ones. Thus, the strategy to initialize spatially discontinuous weak-labels is adopted in this study. As for weakly supervised learning methods, while image processing witness a large number of relevant works, there are quite few studies related to point cloud processing. Guinard and Landrieu (2017); Yao et al. (2020) proposed a transductive learning framework to classify point clouds with sparse weak labels, which was deemed less practicable. A weakly supervised strategy for point cloud semantic segmentation were developed in (Xu and Lee, 2020). However, at the least 10% of labels in a scene were needed to achieve a satisfactory classification result in experiments, so the considerable workload of labeling work was still required, especially for datasets with high density. Our previous study (Wang and Yao, 2021) proposed a pseudo-label-assisted approach to point cloud semantic segmentation using limited annotations, but the training process was inefficient. Additionally, the classification result lacks of robustness when only adopting pseudo-labels.
In this study, we propose a plug-and-play weakly supervised framework for ALS point cloud semantic segmentation, which is integrated with diverse deep network architectures to leverage information in unlabeled data. Firstly, considering that there is no ground-truth information for unlabeled points, we take advantages of the entropy of predictions during training, and entropy regularization is adopted to reduce the class overlap in prediction probability and improve prediction confidence. Then, we develop an ensemble prediction constraint to acquire more robust results by making a contrastive pair, the prediction at current training step with its ensemble value during the whole training process. An online soft pseudo-labeling strategy is proposed to create extra supervisory sources and further improve the accuracy of classification result. Pseudo-labels are generated from the ensemble prediction, and their weights contributing to the loss value is based on the reliability calculated from the predicted probability. We use KPConv (Thomas et al., 2019), a point convolution network, as backbone network. Experiments in three ALS datasets indicate that competitive results are achieved compared to those under the full supervision scheme with only 1‰ of labels. Our main contributions are summed up as follows:
We propose a plug-and-play weakly supervised framework for ALS point cloud semantic segmentation, which can be flexibly integrated with mainstream backbone networks by reducing the reliance on label abundance to achieve a competitive result;
The entropy regularization is introduced to penalize class overlap in predictive probability caused by scant annotations and improve the confidence of predictions;
A consistency constraint is designed by minimizing the contrast between current prediction and the ensemble one to increase the prediction robustness;
An online soft pseudo-labeling strategy is proposed to supply additional supervisory sources, which enables the training process to be completed in an efficient and parameter-free mode.
The remainder of this paper is organized as follows: In Section 2, we systematically review deep learning-based methods for ALS point cloud semantic segmentation and semi- and weakly supervised learning for image and point cloud semantic segmentation. Our methodology is described in detail in Section 3. Section 4 presents the datasets and weak-label settings. In Section 5, we conduct extensive experimental analysis to compare and analyze the effectiveness of proposed method. Concluding remarks are provided for future work in Section 6.
Related work
No matter of full or weak supervision, it is essential to develop a powerful semantic segmentation network structure to extract representative features. Owing to the irregular data distribution in point clouds, the network structure is not as uniform as standard 2D convolutional neural network (CNN), leading to different types of networks. We summarize three main types, including projection-based, voxel-based, and point-based ones.
Due to the structural irregularity of point clouds, 2D CNNs cannot be directly utilized for point cloud processing. To achieve it, early studies projected point clouds onto images and applied mature 2D CNNs to the classification task. Classical point cloud features were usually considered, such as color, intensity and height. Hu and Yuan (2016) proposed a ALS point cloud filtering algorithm by transforming each point into a image. The color value of each pixel was generated based on three types of height differences in a neighborhood of the center point. In Yang et al. (2017), hand-crafted features of point clouds were extracted to generate images, which were fed into a 2D CNN to achieve point cloud classification. Yang et al. (2018) achieved a better classification result by extension of developing a multi-scale CNN. Meanwhile, a similar study was proposed in Zhao et al. (2018). One problem of above studies is that each point needs to be converted into one image, associated with low computational efficiency. Boulch et al. (2018) solved the problem by generating multi-view images of point clouds and fed them into a 2D semantic segmentation network. Then, the classification of each point was acquired with back-projection. Rizaldy et al. (2018) utilized the image generation strategy in Hu and Yuan (2016) and implemented a fully convolutional network to perform ALS ground filtering, which significantly reduced the computational cost.
1.2 Voxel-based methods
Such methods voxelize point clouds and use 3D CNNs to process voxel data. Compared with the projection-based method, the voxel-based method can inherently retain the 3D structural information of the point cloud. VoxNet (Maturana and Scherer, 2015) transformed points to 3D voxels and achieved 3D object recognition by implementing a 3D CNN. VoxelNet (Zhou and Tuzel, 2018) proposed a 3D detection network and presented an efficient strategy to process sparse point structure. To reduce memory footprint and computation consumption during voxelization process, several point cloud organization structures were integrated in 3D CNNs, such as KD-trees (Klokov and Lempitsky, 2017) and octree (Wang et al., 2017). In addition, voxels and projections were combined in some studies. Qi et al. (2016) analyzed 3D volumetric CNNs versus multi-view CNNs and proposed two new architectures of volumetric CNNs for 3D object classification. Qin et al. (2019) combined voxel and pixel representation-based networks to classify ALS data. The disadvantage of projection-/voxel-based methods is that point clouds need to be converted into regularized data formats, which inevitably destroyed original geometric structure.
1.3 Point-based methods
Recently, point-based networks have established as mainstream method for point cloud semantic segmentaiton. PointNet and PointNet++ (Charles et al., 2017; Qi et al., 2017) were pioneers in the development of shared multilayer perceptrons (MLPs) to directly analyze point clouds. Randla-net (Hu et al., 2020) analyzed different point cloud downsampling methods and proposed an computational efficient network. For ALS, different studies had made improvements based on exploiting the characteristics of the data. Yousefhussien et al. (2018) introduced Pointnet into ALS data classification and implemented a multi-scale fully convolutional network. Li et al. (2020) proposed a dense connected network for ALS data classification. Hand-crafted features were integrated into the network, and an elevation-attention module was designed to further enhance the representation of semantic features. In Huang et al. (2020), a multi-scale network was developed, combined with a manifold-based feature embedding module and a graph-structured optimization method. GraNet (Huang et al., 2021) proposed a local convolution module and a global attention module to mine local and global dependencies in point clouds. Graph convolution network is another branch of point-based methods, which constructs a graph through relative spatial positions between points for feature extraction and fusion. In Wang et al. (2019), a dynamic graph was constructed and an edge convolution was proposed for local feature extraction. Landrieu and Simonovsky (2018) segmented point clouds into clusters and proposed a graph-based network. Inspired from convolution kernels in 2D CNNs, Thomas et al. (2019) proposed a point convolution network, in which point kernels were designed to learn local geometric information.
2 Weakly supervised methods
We firstly review weakly supervised methods in image processing where a number of pioneering works were developed. Entropy regularization (Grandvalet and Bengio, 2004), or entropy minimization has proved useful to semi-supervised learning. Pseudo-label, assigning annotations to unlabeled data based on the predictions of current model, is a simple and efficient method to improve the performance of the classification model under weak supervision. An early pseudo-label study was presented by Lee (2013). Iscen et al. (2019) developed a soft pseudo-label method, and the weight was calculated from the entropy of predicted probabilities. He et al. (2021) reduced the bias of pseudo-labels caused by the longtailed class distribution on real-world semantic segmentation datasets. Meanwhile, some methods considered the consistency constraint by creating a contrastive sample. A simple strategy was to conduct data augmentation and add a loss function to constrain the feature similarity. Laine and Aila (2017) utilized exponential moving average value of prediction during training for comparison, while ensemble model parameters were directly considered in Tarvainen and Valpola (2017). Miyato et al. (2018) proposed an adversarial training to generate more targeted comparison samples. Several proven semi-supervision strategies were combined to further improve the accuracy (Berthelot et al., 2019; Sohn et al., 2020).
2.2 Point cloud processing
Until now, there have been few works that used weakly supervised methods to classify point cloud data. Wei et al. (2020) applied a point class activation map to classify point clouds using only scene-level labels. However, we argue that it is not practical for point cloud data, particularly in outdoor scenes that cover a wide region, because the complete data must be divided into numerous small blocks and categories contained in each block must be specified. In comparison, assigning labels to a few number of points within the entire scene is a more desirable approach. Polewski et al. (2016) used an active learning method to detect standing dead trees from ALS data combined with infrared images. Lin et al. (2020) proposed an active and incremental learning strategy for ALS data semantic segmentation, and manual annotation was iteratively added for training. Nonetheless, the setting of weak labels is used to annotate all points falling into tiles, and manual intervention is required during training. A weakly supervised point cloud semantic segmentation framework was recently proposed by Xu and Lee (2020), and an approximate result of fully supervised learning was obtained using 10 of labels. However, the used weak labels were a spatial aggregation of downsampled full scene labels, signifying still a high workload of labeling. Guinard and Landrieu (2017) utilized the point cloud segmentation method to improve the classification accuracy with very few labels, but the result largely depended on the segmentation accuracy and the classification of the point cloud was limited to the same area where weak labels were initialized, which was in nature of transductive learning. Yao et al. (2020) introduced a pseudo-labeling method into point cloud semantic segmentation. However, similar to Guinard and Landrieu (2017), the framework was also a transductive learning scheme, and the performance of the model has not been verified on unseen data. Our previous study (Wang and Yao, 2021) introduced pseudo-label method into a inductive learning framework and designed a adaptive threshold to generate pseudo-labels. In Hu et al. (2021), a semantic query network was proposed to share sparse weak-label information in spatial domain by interpolating features from neighboring points.
Methodology
KPConv(Thomas et al., 2019) is used in this study as backbone network because of its state-of-the-art results achieved on several open datasets. KPConv resorts to the idea of convolution kernels from image processing by extending deformable kernel points to adapt the local features of point clouds. We choose the rigid point convolution kernel and use the same network architecture as KPConv for our semantic segmentation task. The encoder network comprises five convolutional layers, embedding the ResNet-like structure. Skip links are used in the decoder network, and features are passed by the nearest sampling.
2 Entropy regularization
In order to present the idea of entropy regularization, we first introduce the incomplete training on labeled data. The training process of incomplete supervision is similar to that of full supervision, and the only difference lies in the design of the loss function. As only a small number of points are given label information, we calculate the loss of these points and perform backpropagation. The softmax cross-entropy is a commonly-used loss function in supervised semantic segmentation of point clouds, denoted as:
where and are the prediction and label of point , respectively.
Notwithstanding lack of annotation information for loss calculation, class-wise posterior probability can be predicted for unlabeled points by feeding them into the network. Entropy regularization (Grandvalet and Bengio, 2004) was proposed based on the conclusion that the information contained in unlabeled data decreases as classes overlap. The entropy is a measure of class overlap and invariant to the parameterization of the model, which is related to the usefulness of unlabeled data where prediction is ambiguous. Hence, this measure can be utilized to predict well separated classes of unlabeled data. By reducing the class overlap, entropy regularization decreases the uncertainty and supply predictions with a high confidence. Given a target point cloud, the Shannon entropy of each point is used and denoted as:
The entropy regularization is proposed to minimize entropy of posterior probability, and the loss is calculated as an averaged value:
Supervised classification loss and entropy regularization loss are respectively calculated on labeled and unlabeled points, delivered to simultaneous optimization during training. A comparison of normalized entropy map of classification result on training set is shown in Fig. 3, where a number of points is shown to have a relatively high entropy when initializing the model training with limited weak labels, implying a underfitted model. By contrast, entropy regularization can reduce the entropy for most of points, providing a map similar to that under full supervision.
3 Ensemble prediction constraint
where is the coefficient to balance the weighting between ensemble and new values, set to 0.9 in this study. Mean square error (MSE) is utilized to describe the consistency cost, denoted as:
4 Online soft pseudo-labeling
As a simple yet efficient semi-supervised method, pseudo-label can alleviate the problem of limited annotations. The pseudo-label method is a learning form where a classifier is trained to produce predictions, and then retrained by taking inferred classes for unlabeled data as true labels. By increasing the number of pseudo-labels, we intend to reproduce the class distributions in the feature space at scene level. Vanilla pseudo-labeling methods face two problems. Firstly, it appears to inevitably have incorrect predictions in generating and updating pseudo-labels, arising from the underfitted model trained by weakly labeled data. Therefore, a criteria has to be designed to identify reliable predictions and overcome the inaccuracy, which is usually an empirical study considering properties of different networks and data sets. In addition, it needs to iteratively update the pseudo-labels during the entire training process for optimized performance, thus leading to a less efficient training process. In this study, we propose an online soft pseudo-labeling method attempting to solve these two problems, which enables all unlabeled points to be involved in pseudo-label training and maintain quasi the same training speed as the baseline.
Generally, a prediction with high posterior probability is more likely to be correct. Thus, a commonly used method is to select labels with predicted probabilities exceeding a fixed threshold. However, it is difficult to choose a generic threshold applicable to different datasets. In order to better reveal the class distribution and association across whole scene, we derive and soften pseudo-labels on all unlabeled data by associating them with different weights based on the class distribution uncertainty. In this study, the Shannon entropy of predicted probability is utilized to measure the uncertainty, and larger value it represents higher uncertainty. We follow the strategy in (Iscen et al., 2019), and the weight for point is defined as:
where is defined in Equation 2, and is used to normalize to because it is the maximum value of according to the principle of maximum entropy.
4.2 Online training
We argue that true weak labels from ground truth are essential to model training. Thus, to retain their influence, the loss functions of ground truth and pseudo-labels are formulated separately. The loss of pseudo-labels is calculated by weighted cross entropy using instant predictions, denoted as:
All three constraints are based on prediction probability during training. Both ER and OSPL can generate high confident predictions, but in a different way. The function of ER is to reduce class overlap in probability distribution, so it particularly benefits unlabeled points with high uncertainty . OSPL further extends by generating one-hot labels to reduce the uncertainty of unlabeled data. As soft pseudo-label assigns higher weights to more reliable predictions, OSPL takes advantage of unlabeled points with high confidence and low entropy . Hence, ER reduces the number of and produces more , and OSPL mainly utilizes to generate reliable pseudo-labels for further improving the trained model. As for EPC, it is utilized to ensure the robustness of predictions by taking ensemble values as reference, forming a smoothness term to constrain ER and OSPL. It is due to the reason that ER and OSPL may also strengthen the impact of incorrect predictions in unlabeled data, and the consistency constraint could offset such confirmation bias.
5 Training Process
All losses proposed in previous sections participate in backpropagation simultaneously with different weights. The combined optimization problem is presented as:
where represents model parameters, and , , are weighting factors. A ramp-up weight for unsupervised loss components is used in this study. In detail, and are set equal in our experiments, defined with a Gaussian curve following (Laine and Aila, 2017), where T increases linearly from 0 to 1 during ramp-up period, while is set zero during ramp-up period. So, the entire training process consisted of two stages, and each of stage lasts for 100 epochs in our experiments. Stage 1 is the ramp-up period, and , , are all set 1 during stage 2. The training process is detailed in Algorithm 1.
Experiment
Three ALS point cloud datasets are chosen in this study for evaluation and analysis, including the ISPRS Vaihingen 3D Semantic Labeling benchmark (ISPRS) (Cramer, 2010; Rottensteiner et al., 2012), the Large-scale ALS data for Semantic labeling in Dense Urban areas (LASDU) (Ye et al., 2020; Li et al., 2013), and the Hessigheim 3D Benchmark (H3D) (Kölle et al., 2021).
The dataset contains ALS point clouds obtained from the Leica ALS50 system and the co-aligned infrared camera for extracting auxiliary color information. ALS data and aerial images are obtained from Stuttgart region of Germany. The point density of the data is between 4 and 7 points/m². The image data covers the entire area, with a ground sampling distance of 8 cm. There are 9 categories in the dataset, including powerline, low vegetation, impervious surfaces, car, fence, roof, facade, shrub, and tree. The dataset is divided into two parts for training and testing, and the number of points for the training and testing sets is 753,859 and 411,721, respectively. The number of points varies considerably for different classes and are mainly concentrated in the following four categories: low vegetation, impervious surfaces, roof, and tree.
The dataset contains multiple scan data, and the point spacing in the overlap area is very small. To remove redundant points in overlapping ALS strips and maintain an even point density, we set the subsampling grid size to d = 0.4 m and assign labels to deleted points according to the nearest neighbor point in training and testing. The format of utilized features is {X, Y, Z, Intensity, IR, R, G}.
The dataset is obtained from Leica ALS70 system onboard an aircraft, and the study area is located in the valley along the Heihe River in the northwest of China, which covers an urban area of around 1 km². The average point density is approximately 3–4 points per/m². The whole area is divided into four connected sections, two of which are used as training set, and the remaining two as test set. The number of points is approximately 3.12 million, whereby the training and test parts contain around 0.59 million, 1.13 million, 0.77 million, and 0.62 million, respectively. Five categories are predefined in the dataset, including ground, artifacts, low vegetation, trees and buildings. Fig. 7 shows the training and testing sets of LASDU dataset.
Considering the even distribution and relatively low density of point clouds of the LASDU dataset, raw data is directly utilized for training and test. The format of utilized features is {X, Y, Z, Intensity}.
The dataset comprises high density LiDAR point cloud of around 800 points/m² enriched by RGB image with a GSD of 2-3 cm, acquired from a Riegl VUX-1LR Scanner and two oblique-looking Sony Alpha 6000 cameras mounted on a RIEGL Ricopter platform. The area of interest is a village of Hessigheim, Germany. The entire area is divided into three connected sections for training, validation and test, respectively. The training and validation sets are used in this study, for which the number of points are approximately 59.4 million and 14.5 million, respectively. Eleven semantic categories are predefined, including low vegetation, impervious surface, vehicle, urban furniture, roof, facade, shrub, tree, soil/gravel, vertical surface, and chimney. Fig. 7 shows the color map of training and validation sets of H3D dataset.
Considering the high density of raw data, we set the subsampling grid size to 0.1 m for training and test to improve computational efficiency while preserving point cloud structural details. In inference process, predictions of raw test data are from nearest neighbor interpolation. The format of utilized features is {X, Y, Z, R, G, B}.
2 Configuration for weak labels
To evaluate the performance of our weakly supervised method, weak labels are selected under conditions of different ratios. We randomly initialize point labels for each category with (quasi) the same number, which alleviates the issue of strong class imbalance. In addition, as the exact number of each category remains unknown before data annotation, the proposed weak-label configuration is more applicable to real-world tasks, and no prior knowledge of the probability of class occurrence is contained in weak labels. It should be noted that the number of weak labels for each category is not exactly the same due to the presence of extreme class imbalance in ALS data. Even when the total number of weak labels is quite small, a large proportion of points may be already labeled for certain classes, which is hard to be considered weakly supervised task. We follow three weak-label selection criteria in our experiments:
Weak labels are randomly selected from the training data;
The number of weak labels in each category can not exceed of the number of that category;
Less weak-label cases are contained as subset of the more weak-label cases. For instance, the selected labels in 1‰ weak-label setting are fully included in the 2‰ weak-label setting.
Different weak-label ratios are considered in our experiments, presented in Table 1. We first set a maximum number of weak labels for each of single categories. Then, weak labels are initialized by designed criteria. To approach the exact settings as far as possible, the maximum number can be adjusted. Weak-label initialization for ISPRS and H3D datasets is conducted in subsampled data.
3 Implementation details
We follow the strategy in KPConv to generate mini-batch of point clouds for training. An illustration of mini-batch generation is presented in Fig. 8. One batch of training sample is defined as a subset of point clouds contained in a circular area, whose radius is related to the point cloud density. As the number of points in each batch is different, the batch size is not fixed, and a upper bound of the sum number of points at each training step is set according to the memory limitation of the graphics card. In our experiments, the radius and number limit for ISPRS and LASDU dataset are the same, 30 m and 120,000, respectively, while those for H3D dataset are 5 m and 90,000. At each epoch, we train 80 steps for ISPRS dataset, and the number for LASDU and H3D dataset are 200 and 400, respectively. During training, different from random selection or uniform block, the location of circle at each batch is determined by a statistical function. Before the training process started, an initial random number is given to each point in training sets, referred to as potential value. Then, the point with minimum potential value is selected as the central position of next batch, and potential value of points in the batch is added by a number between zero to one according to the distance to the center. In this way, the number of times that each point is fed into the model is controlled in a close range, while it increases variety of training samples. An example of mini-batch generation during training is shown in Fig. 8(a). We further augment the batch data by randomly rotating points around the z-axis, scaling and adding noise offset in coordinates. During test process, though we have found that the batch generation strategy of training process can benefit the accuracy by increasing the number of mini-batches, to improve the efficiency and maintain a fair comparison with other methods, we create uniform blocks to generate subsets of test data. In detail, circle is still utilized with the radius as settings in training, impelling 50% overlap between adjacent batches, as shown in Fig. 8(b). Then, each point is tested for approximately three times. We use the default parameters of the KPConv segmentation network, adopting Stochastic gradient Descent optimizer, with a momentum of 0.98 and an initial learning rate of . All models are implemented in the framework of PyTorch and trained on a single GeForce GTX 1080Ti 11 GB GPU.
4 Evaluation metrics
We use the overall accuracy (OA) and F1 score to evaluate the performance of our method. OA is the percentage of predictions correctly classified where the F1 score is the harmonic mean of the precision and recall, presented as:
where , , and are true positives, false positives, and false negatives, respectively.
Results and discussion
We first compare the performance of our method under weak-label settings. It is worth noting that here the baseline method refers to KPConv which only imposes on weak labels. A comparison of experimental results on ISPRS dataset is illustrated in Fig. 9. From the figure, we can see a rise in OA and average F1 score for no matter baseline or our method when gradually increasing the number of weak labels, which is in line with the general perception of relation between accuracy and number of annotations. Compared to baseline, two evaluation indices have significantly improved using our method. Both OA and average F1 score witness rapid growth when the proportion of weak labels goes from 5‱ to 1‰, and the numbers are 83.0% and 70.0%, marginally below those under full supervision. In the 5‰ of weak-label setting, our method achieves an even better result than corresponding full supervision scheme. Considering the trade-off between annotation workload and model performance, we believe that it is a very promising result using 1‰ of labels, and the classification result and error map is presented in Fig.10. From the figure, it shows that the majority of points are classified correctly, and it maintains a relatively smooth boundary between different objects. Misclassified points show the characteristics of concentrated distribution, and mainly belong to vegetation class including low vegetation, shrub and tree, due to the similar geometric and color information. To provide a detailed information that demonstrate the effectiveness of our method, we visually compare classification results in several local regions in Fig. 11. Due to the limited number of initial weak labels, there are some points that should be easily recognized but are misclassified. Some tree points are wrongly classified to facade by baseline from the first row, and they are corrected by our method. The rest of two rows show our method can rectify most of errors at roof category.
We compare our method with both published results under full supervision and open-sourced weakly supervised methods. It should be mentioned that all these methods are based on deep learning networks. We first introduce deep networks under full supervision. The NANJ2 (Zhao et al., 2018) method generated images from hand-crafted point cloud features and proposed a multi-sacle CNN based classification method. The WhuY4 (Yang et al., 2018) method also transformed points into images and performed the classification. The RandLA-Net (Hu et al., 2020) method proposed a fast point cloud semantic segmentation network by randomly sampling points at each network layer, and a local feature aggregation module was designed to extract semantic features. The KPConv method is the backbone network used for baseline in this study, and we also assess its performance under full supervision. The GANet (Li et al., 2020) method proposed a dense connected network, where a geometry-aware convolution and a elevation-attention module were developed to generate the discriminative deep features. The GraNet (Huang et al., 2021) method proposed a local encoding convolution and a global attention module to combine local and long-range information. Due to currently limited works related to point cloud weakly supervised learning, we choose two open sourced methods for implementation on utilized datasets. The MT (Tarvainen and Valpola, 2017) method introduced a consistency constraint between predictions produced by current model parameters and their EMA values. MT is originally proposed for image processing tasks, and thus integrated with our baseline, KPConv, to conduct experiments. The Xu & Lee (Xu and Lee, 2020) method proposed several modules under weak supervision, including a siamese branch, an inexact supervision branch and a smooth branch.
Quantitative comparison results are listed in Table 2. Using 1‰ of labels, compared with baseline, OA and F1 score of every category have considerably improved using our method. By contrast, there is only a slightly rise in OA and average F1 score for MT. This may be related to the characteristics of ISPRS dataset. As the number of points is fairly small, each point is fed into the network at a high frequency. The model converges quickly using limited weak labels, hence there is hardly difference between model parameters and their EMA value, which limits the efficacy of the consistency constraint. Xu & Lee achieves a unsatisfactory classification performance. While one reason is the gap between different backbone networks, we argue that the inexact supervision branch cannot cope with fairly sparse weak labels, which is discussed in Section 5.4. As for methods under full supervision, we can first see that there is only a small gap for the F1 score of each category between our method and KPConv, which means our weakly supervised method can strengthen the model to extract deep features similar to those from full supervision. NANJ2 and WhuY4 achieve a higher OA, but avrage F1 score is still lower than our method. Two novel methods, GANet and GraNet, achieved more accurate classification result. Compared to them, the gap are mainly attributed to classes of fence and shrub.
1.2 LASDU dataset
Following the same comparison strategy, we first present the increment by our method, presented in Fig. 12. Compared to baseline, OA and average F1 score are considerately improved by our method. However, the growth trends of the two indicators are different. While average F1 score has improved by about 4% under different weak label ratios, the OA increases up to 87.3% in 5‱ of weak labels, but showing only a marginal rise when increasing the number of labels. Additionally, both OA and average F1 score using 5% of labels are still lower than that of full supervision. As there are only 5 main classes in the dataset, some of which contain indeed subcategories, the intraclass feature dissimilarity is comparatively large. Then, it is less feasible for weak labels to capture the overall information of the category under the same ratio. We still show a classification map using 1‰ of labels, presented in Fig.13. It can be seen that four main categories are distinguished well, whereas more errors are shown in artifacts. Despite this, our method performs much better than baseline, and Fig. 14 illustrates a detailed results in several regions. Our method rectified most of classification noises from baseline, so that the overall results exhibit high smoothness.
Several other methods are compared here, and quantitative results are listed in Table 3. Similarly, we analyze weakly supervised learning methods first. Using 1‰ of labels, our method achieves a considerable rise in OA and F1 score of every category compared with baseline. MT acquires a much better result than baseline on LASDU dataset, but worse than our method, and the gap mainly exists in the classes of tree and ground. By contrast, the performance of Xu & Lee is still poor. Comparing with results under full supervision, we can see that our method outperforms GrabNet, slightly below RandLA-Net. Our backbone netowrk, KPConv, acquires the highest result. The classification error of artifacts from our method is the main issue leading to the performance gap.
1.3 H3D dataset
The performance improvement incurred by our method is presented in Fig. 15, which shows that there is a rise in OA and average F1 score for most of weak-label settings. Our method attains the same OA as that of full supervision, and marginally lower average F1 score using 1‰ of weak labels. However, slightly poorer results can be obtained using 2‰ of labels. The reason may be that 1‰ of weak labels is adequate for saturating the performance limit of the backbone network, leading results to have a small range of deviation. Another noteworthy finding is that the average F1 score of baseline at 2‰ of weak labels almost reaches the level of full supervision. Owing to fact that the density of H3D dataset is greatly higher than other two datasets, a large number of weak labels exist. Thus, we believe that for H3D dataset a large number of redundant annotations exist as well. In Fig. 16, we provide the classification result using 1‰ of weak labels. Due to the high point density and the distinct geometric structure of objects, the classification result also maintain good boundary information. We further make a detailed comparison between baseline and our method in Fig. 17, where the first row shows a few trees. Due to the similar feature information under weak labels, some points are misclassified to shrub, which are corrected by our method. Meanwhile, classification errors on roof are also rectified as shown in the other two rows.
Quantitative comparison results are listed in Table 4, where we analyze the results using 1‰ and 1‱ of labels. Under two weak-label settings, our method achieves a considerable rise in OA and F1 score for most of categories compared with baseline. The performance gap between our method and MT reduces on this dataset. Since MT updates model parameters per step, it may help maintain high robustness in case of large number of training samples, thus achieving better result. Xu & Lee performs poorly at some classes, such as chimney and shrub. Using 1‰ of labels, our method obtains the comparable result with RandlA-Net, while the result of shrub and vertical surface poses main problem leading to lower average F1 score compared to KPConv.
2 Ablation study
The effectiveness of each module of proposed method is discussed in this section, and a comparison of the experimental results is presented in Table 5. The experiments are conducted using 1‰ of labels for ISPRS and LASDU dataset, and 1‱ for H3D dataset. We first analyze the improvement by each of individual components. From the table, it shows considerable increases for all three modules on all datasets. On ISPRS dataset, both ER and OSPL increase OA by 5%+ and average F1 score by 10%+ compared with the baseline. By contrast, the growth rate in EPC is smaller, and the reason is similar to MT. On LASDU dataset, the increment for OA and average F1 score in every module equals to about 4% and 5%. On H3D dataset, we can see the OA of EPC and OSPL have obviously improved compared to that of ER. Comparing the performance of each individual modules using different datasets, it indicates that as the number of weak labels in the dataset increases, the effect of ER becomes gradually insignificant, whereas that of EPC becomes increasingly obvious and that of OSPL remains stable. It could be caused by different characteristics of each module. The increased number of weak labels will alleivate the underfitted training model. Then, usable information contained in unlabeled samples decreases, which leads to the degradation effect of ER. By contrast, though OSPL also directly utilizes unlabeled samples to generate pseudo-labels, the accuracy of pseudo-labels grows as the number of weak labels is increased, thus allowing OSPL to maintain a steady rise in evaluation indices. As for EPC, large amount of points gives rise to great diversity in mini-batches, which contributes more robust ensemble value to consistency constraint. Despite the effect of each individual module, it still shows a further rise in evaluation metrics when combining these modules. Due to the similarity, we combine ER and OSPL to compare with our method. From the table, it shows that the combined version outperforms individual components on all three datasets, and our method achieves the best result in all settings, which demonstrates the effectiveness of our method.
3 Complexity and runtime analysis
We consider baseline, MT and our method, and the comparison is presented in Table 6. All three methods use KPConv as the backbone network. Our method makes use of output predictions during training, incurring no increase in model size. As for MT, though it needs to create shadow variables to save EMA value of model parameters, those variables do not participate in loss calculation and backpropagation. Thus, both our method and MT have the same model complexity compared to baseline, and the comparison is not listed. The running time refers to the whole training process, including a period of validation after each epoch. Under the same implementation condition, our method takes up 21% and 20% more running time on ISPRS and LASDU dataset compared to baseline, respectively, and for the case of H3D dataset it is increased by up to 30%. The reason is that EMA value of predictions at each step needs to be calculated and saved. By contrast, MT is obviously more demanding in terms of running time because not only model parameters need to be updated at each training step, but also two forward propagation are required to generate contrastive prediction pairs. Thus, our method shows better running efficiency compared to MT. It should be noticed that only training time is discussed here. Owing to fact that there is no modification in network structure design, the testing process of all three methods remains the same.
4 Comparison to exploiting contextual information
Contextual information describes overall class distribution in a scene. In this study, it refers specifically to existed scene-level labels in one training sample. Liu et al. (2020) proposed a contextual loss for point cloud scene understanding. In Xu and Lee (2020), it was utilized as inexact supervision using sparse weak labels. The rational behind that branch was that an object category should be contained even if only one weak label of the category is present in the sample. While experimental results showed the effectiveness using 10% of total labels in that work, we argue that it could produce incorrect information using fairly limited weak labels, which undermines the model training. When the number of weak labels continues to decrease, one situation occurs that no annotation for some existent categories in the sample can be acquired. The incorrect contextual information will lead to degraded effect because the loss function rewards the probability of existed classes while penalize that of nonexistent ones. A further experiment is presented in Table 7, where the effect of contextual loss is compared using datasets in this study. From the table, we can see a huge drop in accuracy on ISPRS dataset and a decline for other two. The characteristics of the dataset is the main reason, since low density, small number of points and relatively large number of categories of ISPRS dataset makes it difficult for weak labels to reflect complete scene-level label information.
5 Limitation
Our method focuses on utilizing potential information in unlabeled data and proposes three prediction-constraint strategies to improve classification accuracy under weak supervision. Though it is a lightweight framework, which can be easily integrated to diverse networks, we believe that a pre-designed network architecture targeted for weak supervision could further promote the development of this field. An experiment on ISPRS dataset is illustrated in Table 8, in which two proven network, KPConv and RandLA-Net, are considered for comparing their performance under 1‰ of labels. While there exists only a small performance gap between two fully supervised models, a large difference is observed under weak supervision. Thus, we argue that different networks are more sensitive to respond to weakly supervised point cloud classification tasks, especially for those fairly sparse labels. Though the attention mechanism, which successfully applied in many studies, is considered in our experiments, no direct solution to changing the network architecture is presented. We integrate DualAttention (Fu et al., 2019) module into KPConv with an intention to enhance the ability of feature representation. Considering the limitation of computational resource, it is only added to the last layer of the encoder network, while experimental results do not show a substantial improvement from Table 9. Thus, it seems that it does not bring much benefit when merely adding a general attention module, and it becomes our future work to propose a network closely compatible to weak supervision.
Conclusion
In this study, we investigate semantic segmentation of ALS point clouds using sparse annotations and propose an efficient weakly supervised framework, which is compatible with current point cloud classification networks. An entropy regularization module is introduced to reduce class overlap in predictions and improve the classification confidence. An ensemble prediction constraint, where a consistency loss is calculated between prediction confidence at current training step and its ensemble value, is proposed to strength the robustness of predictions. Additionally, an online soft pseudo-labeling module is developed to take advantage of output predictions of training samples and supply extra supervisory sources. We perform comprehensive experiments to evaluate our method using three benchmark ALS datasets. Our method significantly improves OA and average F1 score compared with baseline and achieves comparable results against the full supervision competitors using only 1‰ of labels. Experimental results demonstrate that our method can help to largely reduce the workload of data annotation as only sparse labels randomly generated across the scene are needed to attain promising results.
In the future, we would like to solve the limitation of our method as discussed before. A well-designed network architecture closely compatible to weak supervision could better exploit contextual relations between sparse weak labels. In addition, random selection may not the best way to initialize weak labels. Unsupervised pre-training could assist with prior knowledge and enable the model to achieve better performance using the same number of annotations.
Acknowledgements
The ISPRS dataset was provided by the German Society for Photogrammetry, Remote Sensing, and Geoinformation (DGPF). The LASDU dataset was provided by College of Surveying and Geo-informatics, Tongji University and Photogrammetry and Remote Sensing, Technical University of Munich. The H3D dataset was provided by Institute for Photogrammetry, University of Stuttgart.