Unsupervised Multi-Task Feature Learning on Point Clouds
Kaveh Hassani, Mike Haley
Introduction
Point clouds are sparse order-invariant sets of interacting points defined in a coordinate space and sampled from surface of objects to capture their spatial-semantic information. They are the output of 3D sensors such as LiDAR scanners and RGB-D cameras, and are used in applications such as human-computer interactions , self-driving cars , and robotics . Their sparse nature makes them computationally efficient and less sensitive to noise compared to volumetric and multi-view representations.
Classic methods craft salient geometric features on point clouds to capture their local or global statistical properties. Intrinsic features such as wave kernel signature (WKS) , heat kernel signature (HKS) , multi-scale Gaussian curvature , and global point signature ; and extrinsic features such as persistent point feature histograms and fast point feature histograms are examples of such features. These features cannot address semantic tasks required by modern applications and hence are replaced by the unparalleled representation capacity of deep models.
Feeding point clouds to deep models, however, is not trivial. Standard deep models operate on regular-structured inputs such as grids (images and volumetric data) and sequences (speech and text) whereas point clouds are permutation-invariant and irregular in nature. One can rasterize the point clouds into voxels but it demands excessive time and memory, and suffers from information loss and quantization artifacts .
Some recent deep models can directly consume point clouds and learn to perform various tasks such as classification , semantic segmentation , part segmentation , image-point cloud translation , object detection and region proposal , consolidation and surface reconstruction , registration , generation , and up-sampling . These models achieve promising results thanks to their feature learning capabilities. However, to successfully learn such features, they require large amounts of labeled data.
A few works explore unsupervised feature learning on point sets using autoencoders and generative models, e.g., generative adversarial networks (GAN) , variational autoencoders (VAE) , and Gaussian mixture models (GMM) . Despite their good feature learning capabilities, they suffer from not having access to supervisory signals and targeting a single task. These shortcomings can be addressed by self-supervised learning and multi-task learning, respectively. Self-supervised learning defines a pretext task using only the information present in the data to provide a surrogate supervisory signal whereas multi-task learning uses the commonalities across tasks by jointly learning them .
We introduce a multi-task model that exploits three regimes of unsupervised learning including self-supervision, autoencoding, and clustering as its target tasks to jointly learn point and shape features. Inspired by , we show that leveraging joint clustering and self-supervised classification along with enforcing reconstruction achieves promising results while avoiding trivial solutions. The key contributions of our work are as follows:
We introduce a multi-scale graph-based encoder for point clouds and train it within an unsupervised multi-task learning setting.
We exhaustively evaluate our model under various learning settings on ModelNet40 shape classification and ShapeNetPart segmentation tasks.
We show that our model achieves state-of-the-art results w.r.t prior unsupervised models and narrows the gap between unsupervised and supervised models.
Related Work
PointNet is an MLP that learns point features independently and aggregates them into a shape feature. PointNet++ defines multi-scale regions and uses PointNet to learn their features and then hierarchically aggregates them. Models based on KD-trees spatially partition the points using kd-trees and then recursively aggregate them. RNNs are applied to point clouds by the assumption that “order matters” and achieve promising results on semantic segmentation tasks but the quality of the learned features is not clear.
CNN models introduce non-Euclidean convolutions to operate on point sets. A few models such as RGCNN , SyncSpecCNN and Local Spectral GCNN operate on spectral domain. These models tend to be computationally expensive. Spatial CNNs learn point features by aggregating the contributions of neighbor points. Pointwise convolution , Edge convolution , Spider convolution , sparse convolution , Monte Carlo convolution , parametric continuous convolution , feature-steered graph convolution , point-set convolution , -convolution , and spherical convolution are examples of these models. Spatial models provide strong localized filters but struggle to learn global structures .
A few works train generative models on point sets. Multiresolution VAE introduces a VAE with multiresolution convolution and deconvolution layers. PointGrow is an auto-regressive model that can generate point clouds from scratch or conditioned on given semantic contexts. It is shown that GMMs trained on PointNet features achieve better performance compared to GANs .
A few recent works explore representation learning using autoencoders. A simple autoencoder based on PointNet is shown to achieve good results on various tasks . FoldingNet uses an encoder with graph pooling and MLP layers and introduces a decoder of folding operations that deform a 2D grid onto the underlying object surface. PPF-FoldNet projects the points into point pair feature (PPF) space and then applies a PointNet encoder and a FoldingNet decoder to reconstruct that space. AtlasNet extends the FoldingNet to multiple grid patches whereas SO-Net aggregates the point features into SOM node features to encode the spatial distributions. PointCapsNet introduces an autoencoder based on dynamic routing to extract latent capsules and a few MLPs that generate multiple point patches from the latent capsules with distinct grids.
2 Self-Supervised Learning
Self-supervised learning defines a proxy task on unlabeled data and uses the pseudo-labels of that task to provide the model with supervisory signals. It is used in machine vision with proxy tasks such as predicting arrow of time , missing pixels , position of patches , image rotations , synthetic artifacts , image clusters , camera transformation in consecutive frames , rearranging shuffled patches , video colourization , and tracking of image patches and has demonstrated promising results in learning and transferring visual features.
The main challenge in self-supervised learning is to define tasks that relate most to the down-stream tasks that use the learned features . Unsupervised learning, e.g., density estimation and clustering, on the other hand, is not domain specific . Deep clustering models are recently proposed to learn cluster-friendly features by jointly optimizing a clustering loss with a network-specific loss. A few recent works combine these two approaches and define deep clustering as a surrogate task for self-supervised learning. It is shown that alternating between clustering the latent representation and predicting the cluster assignments achieves state-of-the-art results in visual feature learning.
3 Multi-Task Learning
Multi-task learning leverages the commonalities across relevant tasks to enhance the performance over those tasks . It learns a shared feature with adequate expressive power to capture the useful information across the tasks. Multi-task learning has been successfully used in machine vision applications such as image classification , image segmentation , video captioning , and activity recognition . A few works explore self-supervised multi-task learning to learn high level visual features . Our approach is relevant to these models except we use self-supervised tasks in addition to other unsupervised tasks such as clustering and autoencoding.
Methodology
Clustering function maps the latent variable into categories such that and . This function encourages the encoder to generate features that are clustering-friendly by pushing similar samples in the feature space closer and pushing dissimilar ones away. It also provides the model with pseudo-labels for self-supervised learning through its hard cluster assignments.
Classifier function predicts the cluster assignments of the latent variable such that the predictions correspond to the hard clusters assignments of . In other words, maps the latent variable into predicted categories such that . This function uses the pseudo-labels generated by the clustering function, i.e., cluster assignments, as its proxy train data. The difference between the cluster assignments and the predicted cluster assignments provides the supervisory signals.
where and . The centroid matrix is initialized randomly. It is noteworthy that: (i) when assigning cluster labels, the centroid matrix is fixed, and (ii) the centroid matrix is updated epoch-wise and not batch-wise to prevent the learning process from diverging.
For the classification function, we minimize the cross-entropy loss between the cluster assignments and the predicted cluster assignments as follows.
where and are the cluster assignments and the predicted cluster assignments, respectively.
We use Chamfer distance to measure the difference between the original point cloud and its reconstruction. Chamfer distance is differentiable with respect to points and is computationally efficient. It is computed by finding the nearest neighbor of each point of the original space in the reconstructed space and vice versa, and summing up their Euclidean distances. Hence, we optimize the decoding loss as follows.
where and, and are the original and reconstructed point sets, respectively. and denote the number of point sets in the train set and the number of points in each point set, respectively.
Let’s denote the clustering, classification, and decoding objectives by , , and , respectively. we define the multi-task objective as a linear combination of these objectives: and train the model based on that. The training process is shown in Algorithm 1.
We first randomly initialize the model parameters and assume an arbitrary upper bound for the number of clusters. We show through experiments that the model converges to a fixed number of clusters by emptying some of the clusters. This is especially favorable when the true number of categories is unknown. We then randomly select point sets from the training data and feed them to the randomly initialized encoder and set the extracted features as the initial centroids. Afterwards we optimize the model parameters w.r.t the multi-task objective using mini-batch stochastic gradient descent. Updating the centroids with the same frequency as the network parameters can destabilize the training. Therefore, we aggregate the learned features and the cluster assignments within each epoch and update the centroids after an epoch is completed.
2 Architecture
Inspired by Inception and Dynamic Graph CNN (DGCNN) architectures, we introduce a graph-based architecture shown in Figure 1 which consists of an encoder and three task-specific decoders. The encoder uses a series of graph convolution, convolution, and pooling layers in a multi-scale fashion to learn point and shape features from an input point cloud jittered by Gaussian noise. For each point, it extracts three intermediate features by applying graph convolutions on three neighborhood radii and concatenates them with the input point feature and its convolved feature. The first three features encode the interactions between each point and its neighbors where as the last two features encode the information about each point. The concatenation of the intermediate features is then passed through a few convolution and pooling layers to learn another level of intermediate features. These point-wise features are then pooled and fed to an MLP to learn the final shape feature. They are also concatenated with the shape feature to represent the final point features. Similar to , we define the graph convolution as follows:
where is the learned feature for point based on its neighbor contributions, are the nearest points to the in Euclidean space, is a nonlinear function parameterized by and is the concatenation operator. We use a shared MLP for . The reason to use both and is to encode both global information () and local interactions () of each point.
To perform the target tasks, i.e., clustering, classification, and autoencoding, we use the following. For clustering, we use a standard implementation of K-means to cluster the shape features. For self-supervised classification, we feed the shape features to an MLP to predict the category of the shape (i.e., cluster assignment by the clustering module). And for the autoencoding task, we use an MLP to reconstruct the original point cloud from the shape feature. This MLP is denoising and reconstructs the original point cloud before the addition of the Gaussian noise. All these models along with the encoder are trained jointly and end-to-end. Note that all these tasks are defined on the shape features. Because a shape feature is an aggregation of its corresponding point features, learning a good shape feature pushes the model to learn good point features too.
Experiments
We optimize the network using Adam with an initial learning rate of 0.003 and batch size of 40. The learning rate is scheduled to decrease by 0.8 every 50 epochs. We apply batch-normalization and ReLU activation to each layer and use dropout with . To normalize the task weights to the same scale, we set the weights of clustering ( ), classification (), and reconstruction() to 0.005, 1.0, 500, respectively. For graph convolutions, we use neighborhood radii of 15, 20, and 25 (as suggested in ) and for normal convolutions we use 11 kernels. We set the upper bound number of clusters () to 500. We also set the size of the MLPs in prediction and reconstruction tasks to and , respectively. Note that the size of the last layers correspond to the upper bound number of clusters (500) and the reconstruction size (6144: 20483). Following we set the shape and point feature sizes to 512 and 1024, respectively.
For preprocessing and augmentation we follow and uniformly sample 2048 points and normalize them to a unit sphere. We also apply point-wise Gaussian noise of and shape-wise random rotations between degrees along -axis and random rotations between [-20, +20] degrees along and axes.
The model is implemented with Tensorflow on a Nvidia DGX-1 server with 8 Volta V100 GPUs. We used synchronous parallel training by distributing the training mini-batches over all GPUs and averaging the gradients to update the model parameters. With this setting, our model takes 830s on average to train one epoch on the ShapeNet (i.e, 55k samples of size 20483). We train the model for 500 epochs. At test time, it takes 8ms on an input point cloud with size 20483.
2 Pre-training for Transfer Learning
Following the experimental protocol introduced in , we pre-train the model across all categories of the ShapeNet dataset (i.e., 57,000 models across 55 categories) , and then transfer the trained model to two down-stream tasks including shape classification and part segmentation. After pre-training the model, we freeze its weights and do not fine-tune it for the down-stream tasks.
Following , we use Normalized Mutual Information (NMI) to measure the correlation between cluster assignments and the categories without leaking the category information to the model. This measure gives insight on the capability of the model in predicting category level information without observing the ground-truth labels. The model reaches an NMI of 0.68 and 0.62 on the train and validation sets, respectively which suggests that the learned features are progressively encoding category-wise information.
We also observe that the model converges to 88 clusters (from the initial 500 clusters) which is 33 more clusters compared to the number of ShapeNet categories. This is consistent with the observation that “some amount of over-segmentation is beneficial” . The model empties more than 80% of the clusters but does not converge to the trivial solution of one cluster. We also trained our model on the 10 largest ShapeNet categories to investigate the clustering behavior where the model converged to 17 clusters. This confirms that model converges to a fixed number of clusters which is less than the initial upper bound assumption and is more than the actual number of categories in the data.
To investigate the dynamics of the learned features, we selected the 10 largest ShapeNet categories and randomly sampled 200 shapes from each category. The evolution of the features of the sampled shapes visualized using t-SNE (Figure 2) suggests that the learned features progressively demonstrate clustering-friendly behavior along the training epochs.
3 Shape Classification
To evaluate the performance of the model on shape feature learning, we follow the experimental protocol in and report the classification accuracy on transfer learning from the ShapeNet dataset to the ModelNet40 dataset (i.e., 13,834 models across 40 categories divided to 9,843 and 3,991 train and test samples, respectively). Similar to , we extract the shape features of the ModelNet40 samples from the pre-trained model without any fine-tuning, train a linear SVM on them, and report the classification accuracy. This approach is a common practice in evaluating unsupervised visual feature learning and provides insight about the effectiveness of the learned features in classification tasks.
Results shown in Table 1 suggest that our model achieves state-of-the-art accuracy on the ModelNet40 shape classification task compared to other unsupervised feature learning models. It is noteworthy that the reported result is without any hyper-parameter tuning. With random hyper-parameter search, we observed an 0.4 absolute increase in the accuracy (i.e., 89.5%). The results also suggest that the unsupervised model is competitive with the supervised models. Error analysis reveals that the misclassifications occur between geometrically similar shapes. For example, the three most frequent misclassifications are between (table, desk), (nightstand, dresser), and (flowerpot, plant) categories. A similar observation is reported in and it is suggested that stronger supervision signals may be required to learn subtle details that discriminate these categories.
To further investigate the quality of the learned shape features, we evaluated them in a zero-shot setting. For this purpose, we cluster the learned features using agglomerative hierarchical clustering (AHC) and then align the assigned cluster labels with the ground truth labels (ModelNet40 categories) based on majority voting within each cluster. The results suggest that the model achieves 68.88% accuracy on the shape classification task with zero supervision. This result is consistent with the observed NMI between cluster assignments and ground truth labels in the ShapeNet dataset.
4 Part Segmentation
Part segmentation is a fine-grained point-wise classification task where the goal is to predict the part category label of each point in a given shape. We evaluate the learned point features on the ShapeNetPart dataset , which contains 16,881 objects from 16 categories (12149 train, 2874 test, and 1858 validation). Each object consists of 2 to 6 parts with total of 50 distinct parts among all categories. Following , we use mean Intersection-over-Union (mIoU) as the evaluation metric computed by averaging the IoUs of different parts occurring in a shape. We also report part classification accuracy.
Following , we randomly sample 1% and 5% of the ShapeNetPart train set to evaluate the point features in a semi-supervised setting. We use the same pre-trained model to extract the point features of the sampled training data, along with validation and test samples without any fine-tuning. We then train a 4-layer MLP on the sampled training sets and evaluate it on all test data. Results shown in Table 2 suggest that our model achieves state-of-the-art accuracy and mIoU on ShapeNetPart segmentation task compared to other unsupervised feature learning models. Also comparisons between our model (trained on 5% of the training data) and the fully supervised models are shown in Table 3. The results suggest that our model achieves an mIoU which is only 8% less than the best supervised model and hence narrows the gap with supervised models.
We also performed intrinsic evaluations to investigate the consistency of the learned point features within each category. We sampled a few shapes from each category, stacked their point features, and reduced the feature dimension from 1024 to 512 using PCA. We then co-clustered the features using the AHC method. The result of co-clustering on the airplane category is shown in Figure 3. We observed a similar consistent behavior over all categories. We also used AHC and hierarchical density-based spatial clustering (HDBSCAN) methods to cluster the point features of each shape. We aligned the assigned cluster labels with the ground truth labels based on majority voting within each cluster. A few sample shapes along with their ground truth part labels, predicted part labels by the trained MLP, AHC, and HDBSCAN clustering are illustrated in Figure 4. As shown, HDBSCAN clustering results in a decent segmentation of the learned features in a fully unsupervised setting.
5 Ablation Study
We first investigate the effectiveness of the graph-based encoder on the shape classification task. In the first experiment, we replace the encoder with a PointNet encoder and keep the multi-task decoders. We train and test the network with the same transfer learning protocol which results in a classification accuracy of 86.2%. Compared to the graph-based encoder with accuracy of 89.1%, this suggests that our encoder learns better features and hence contributes to the state-of-the-art results that we achieve. To investigate the effectiveness of the multi-task learning, we compare our result against the results reported on a PointNet autoencoder (i.e., single reconstruction decoder) which achieves classification accuracy of 85.7%. This suggests that using multi-task learning improves the quality of the learned features. The summary of the results is shown in Table 4.
We also investigate the effect of different tasks on the quality of the learned features by masking the task losses and training and testing the model on each configuration. The results shown in Table 5 suggest that the reconstruction task has the highest impact on the performance. This is because contrary to , we are not applying any heuristics to avoid trivial solutions and hence when the reconstruction task is masked both clustering and classification tasks tend to collapse the features to one cluster which results in degraded feature learning.
Moreover, the results suggest that masking the cross-entropy loss degrades the accuracy to 87.6% (absolute decrease of 1.5%) whereas masking the k-means loss has a less adverse effect (degraded loss of 88.3%, i.e., absolute decrease of 0.8%). This implies that the cross-entropy loss (classifier) plays a more important role than the clustering loss. Furthermore, the results indicate that having both K-means and cross-entropy losses along with the reconstruction task yields the best result (i.e., accuracy of 89.1%). This may seems counter-intuitive as one may assume that using the clustering pseudo-labels to learn a classification function would push the classifier to replicate the K-means behavior and hence the k-means loss will be redundant. However, we think this is not the case because the classifier introduces non-linearity to the feature space by learning non-linear boundaries to approximate the predictions of the linear K-means model which in turn affects the clustering outcomes in the following epoch. K-means loss on the other hand, pushes the features in the same cluster to a closer space while pushing the features of other clusters away.
Finally, we report some of our failed experiments:
We tried K-Means++ to warm-start the cluster centroids. We did not observe any significant improvement over the randomly selected centroids.
We tried soft parameter sharing between the decoder and classifier models. We observed that this destabilizes the model and hence we isolated them.
Similar to , we tried stacking more graph convolution layers and recomputing the input adjacency to each layer based on the feature space of its predecessor layer. We observed that this has an adverse effect on both classification and segmentation tasks.
Conclusion
We proposed an unsupervised multi-task learning approach to learn point and shape features on point clouds which uses three unsupervised tasks including clustering, autoencoding, and self-supervised classification to train a multi-scale graph-based encoder. We exhaustively evaluated our model on point cloud classification and segmentation benchmarks. The results suggest that the learned features outperform prior state-of-the-art models in unsupervised representation learning. For example, in ModelNet40 shape classification tasks, our model achieved the state-of-the-art (among unsupervised models) accuracy of 89.1% which is also competitive with supervised models. In the ShapeNetPart segmentation task, it achieved mIoU of 77.7 which is only 8% less than the state-of-the-art supervised model. For future directions, we are planning to: (i) introduce more powerful decoders to enhance the quality of the learned features, (ii) investigate the effect of other features such as normals and geodesics, and (iii) adapt the model to perform semantic segmentation tasks too.