Rotation Invariant Convolutions for 3D Point Clouds Deep Learning

Zhiyuan Zhang, Binh-Son Hua, David W. Rosen, Sai-Kit Yeung

Introduction

Recent 3D deep learning has led to great progress in solving scene understanding problems like object classification, semantic and instance segmentation with high accuracies by training a neural network with 3D data. Researches in this area have been continuing to grow and diverse as 3D data becomes more widely and easily available from consumer devices.

Among various data presentations, 3D point cloud is a strong candidate for scene understanding tasks thanks to its availability, compactness, and robustness compared to volumetric or image representations. Point clouds can be acquired by various methods and hardware including multiple view geometry in dual-lens cameras and structured light or time-of-flight sensing in depth and LiDAR cameras. However, learning with point clouds is deemed challenging because a point cloud does not contain a regular structure such as that in an image or a volume. Performing convolution on point cloud therefore requires some special designs in the convolution operator that takes care of this irregularity. A wide body of works have recently been proposed to solve this problem, demonstrating state-of-the-art performance in scene understanding with point cloud data.

Nevertheless, there remains a fundamental problem with existing convolution operator with point clouds: most previous works do not allow the input point cloud to be rotation invariant. During training, data is simply augmented with some rotations which can cause the network not able to generalize well to unseen rotations. A few convolution operators that allows rotation invariance exist but consistent predictions with arbitrarily rotated data are still not achieved.

In this work, we propose a novel convolution operator for point clouds that can achieve high accuracies in scene understanding tasks while still preserving the rotation invariance property. Particularly, our convolution is based on low-level geometric features that are translation and rotation invariant. Such features are used in tandem with a binning approach that addresses point ordering issue in point cloud convolution, resulting in a single convolution that is robust to both issues. In summary, our contributions are:

A robust feature extraction scheme suitable for convolution that supports both rotation and translation invariant features based on low-level geometric cues;

A novel convolution operator that is agnostic to both point cloud rotations and point orders. To address the point ordering issue, we devise a simple binning approach that can be seamlessly combined with the feature extraction step;

A compact convolutional neural network based on the proposed convolution for object classification and object part segmentation. We demonstrate highly consistent and accurate performance under different rotations.

Related Works

The availability of 3D object and scene datasets has made scene understanding in 3D feasible. Common tasks such as object classification, semantic segmentation, and retrieval can now achieve highly accurate results. We briefly summarize the development of 3D deep learning below.

3D deep learning is more diverse compared to image-based deep learning because there are various representations for learning with 3D data. In early stage of 3D deep learning, volume representation , or multiple view images are often adopted for neural networks since they are straightforward extensions from learning with images. However, such representations do not scale well due to large memory requirements and limited resolution in representing 3D geometry.

Recently, PointNet sparked the research interest in deep learning with 3D point clouds by showing that it is possible to learn features of a point set with a special network that is robust to input point orders. This opens the capability for object classification and semantic segmentation with point clouds. Several subsequent works are built along this line of research. Alternatives to make convolution operator compatible to point cloud is to summarize point features into a regular grid and apply a traditional convolution , performing convolution on a local space such as tangent planes , learning to transform point clouds into a canonical latent space . Such techniques perform competitively to PointNet while being able to exploit features from a local region on the point cloud.

The trend of deep learning with point cloud data has been continuing to grow diversely. Recent methods explores convolution kernels that exploit geometric features , add edges on top of points , parameterize convolution using polynomials , and leverage shape context . Some methods are specially design to be lightweight for real-time applications , or to combine with recurrent neural network and sequence model . Some methods exploit hierarchical structures and clustering for scalability , mapping point cloud to two dimensional space , applying spectral analysis , or addressing non-uniform point distribution .

Our method is a part of this trend. We explore how to perform convolution on local point features and at the same time achieve rotation invariance. Compared to deep learning with images, rotation invariance is an important property and a more critical issue for robustness because in 3D, there is no convention about how to align 3D shapes. In geometric deep learning , one can achieve rotation invariance with geodesic convolution on Riemannian manifolds with angular maxpooling . Such convolution, however, needs shape surfaces to operate. By contrast, our convolution is for point sets, and defined directly in the Euclidean space.

The most relevant work to ours is the concurrent work by Rao et al. . They showed that point clouds can be mapped to an icosahedral lattice on which a rotation invariance convolution can be implemented. The key difference here is that we do not need a spherical domain for rotation invariance. Instead, we define convolution with rotation invariant features, which is much simpler and intuitive. In addition, there are a few previous works about learning local descriptors from point clouds for feature matching , some of which can be rotation invariant. These works are however orthogonal to ours mainly because they are targeted for point cloud registration.

Rotation Invariant Convolution

In this section, we detail the RIConv operator construction procedure. Our goal is to seek a simple but efficient way to perform traditional convolution on features extracted from an input point cloud. We design a feature extraction scheme such that the local features are invariant to both translation, rotation, and point orders. Different from previous works that rely on a spherical convolution for rotation invariance , we show that it is possible to achieve rotation invariance directly in the Euclidean space by utilizing low-level geometric cues.

Our feature extraction can be explained as in Figure 1. Given a reference point pp (red), KK nearest neighbors are determined to construct a local point set. The centroid of the point set is denoted as mm (blue). We use vector \vvpm\vv{pm} as a reference to extract translation and rotation invariant features for all points in the local point set. Particularly, for a point xx in this set, its features are defined as

Here, d0d_{0} and d1d_{1} represent the distances from xx to pp and to mm, respectively. α0\alpha_{0} and α1\alpha_{1} represent the angles from xx towards pp and mm, as shown in Figure 1. Since such low-level geometric features are invariant under rigid transformations, they are very well suited for our need to make a translation invariant convolution with rotation invariance property. Note that the reference vector \vvpm\vv{pm} can also serve as a local orientation indicator and we will use it to build a local coordinate system for convolution, in the subsequent step.

A caveat from the feature extraction scheme is that the reference vector \vvpm\vv{pm} can degenerate when pp and mm become a single point. Such cases occur when the neighbors are distributed evenly around the reference point. In such case, we select the farthest point to pp as mm to avoid singularity. In fact, within such a smooth distribution, points that are equidistant to the reference point are expected to have similar features, and thus the degeneration does not negatively affect the features.

2 Convolution Operator

After obtaining rotation invariant features, we are now ready to detail the main idea of our convolution in Figure 2. A key issue here is how to perform convolution that is agnostic to input point orders. PointNet extracts a global feature vector from the entire input point cloud by maxpooling the features from a shared MLP. Here, we build our convolution on local features and use a binning approach with shared MLP to solve this issue. This idea is relevant to shell based convolution in that both apply binning to resolve the ordering issue of point sets and output fixed size features.

Particularly, we start by sampling a set of representative points through farthest point sampling strategy which is able to generate uniformly distributed points. From each of which we perform a set of K-nearest neighbors to obtain local point sets. For each point, the rotation invariant features are extracted as described in the previous section. The features are lifted to a high-dimensional space by a shared multi-layer perceptron (MLP).

To proceed with convolution, we have to define an order so that kernel weights in the convolution can be applied to the corresponding points. Here we devise a simple binning approach and turn the convolution into 1D. Such process has been shown to be highly efficient for local feature learning . In this work, the steps are as follows. We use the reference vector \vvpm\vv{pm} and split the point distribution into N cells along this vector. The feature of each cell is maxpooled from all points participating in the cell. As the cells are ordered, convolution thus becomes possible. We apply a 1D convolution on the fixed-size feature vector from the cells to obtain the output features of our operator. All steps are summarized in Algorithm 1 (see Appendix).

In addition, traditional convolutional neural networks often allows downsampling and upsampling to manipulate the spatial resolution of the input. We build this strategy into our convolution by simply treating the reference point set as the downsampling/upsampling points.

Neural Networks

We use our convolution operator as the core to build neural networks for two common scene understanding tasks: object classification and object part segmentation. These two tasks are commonly used to benchmark the performance of deep learning with point cloud data . Our network is shown in Figure 3.

The object classification network consists of three rotation invariant convolution operators followed by a classifier to output labels for the input point cloud. As our convolution operator is already designed to handle arbitrary rotation and point orders, we can simply place each convolution one after another. By default, each convolution is followed by a batch normalization and an ReLU activation.

The object part segmentation network follows an encoder-decoder architecture with skip connections similar to U-net . We assume a general condition that the object category is unknown when part segmentation is performed. The classification network acts as the encoder, yielding the features in the latent space that can be subsequently decoded into part labels.

In the decoding stage, after each feature is concatenated by skip connections, we apply a MLP before passing the features for deconvolution. Our deconvolution is basically similar to convolution except that it gradually outputs denser points with less feature channels until the output reaches the original number of points.

Unless otherwise mentioned, we use 1024 points for classification, and 2048 points for part segmentation, respectively. In the encoding stage, the point cloud is downsampled to 256256, 128128, and 6464, respectively for classification task, and 512512, 128128, and 3232, respectively for segmentation task. The nearest neighbor size is set to 6464, 3232, and 1616 respectively for the three layers of convolutions. We empirically set the number of bins for handling point orders in each convolution as 44, 22, 11, respectively, which strikes a good balance between accuracy and speed. This setting ensures that each bin contains 1616 points approximately. In general, the neighborhood has to be large enough for capturing the point distribution and features robustly but not too large that causes too much overhead.

Experimental Results

We report our evaluation results in this section. We implemented our network in Tensorflow . We use a batch size of 3232 for classification training and 1616 for segmentation training. The optimization is done with an Adam optimizer. The initial learning rate is set to 0.001. Our training is executed on a computer with an Intel(R) Core(TM) i7-6900K CPU equipped with a NVIDIA GTX 1080 GPU.

We evaluate the proposed convolution and neural network with two tasks: object classification and object part segmentation. The point cloud size is 1024 for classification and 2048 for segmentation. It takes about 3 hours for the training to converge for classification, and about 18 hours for part segmentation. Unless otherwise stated, for object classification, we train for 250250 epochs. The network usually converges within 150150 epochs. For object part segmentation, we train for 300300 epochs, and the network usually converges within 200200 epochs.

Following Esteves et al. , we perform experiments in three cases: (1) training and testing with data augmented with rotation about gravity axis (z/z), (2) training and testing with data augmented with arbitrary SO3 rotations (SO3/SO3), and (3) training with data by z-rotations and testing with data by SO3 rotations (z/SO3). The first case is commonly adopted by previous methods in handling rotated point clouds, and the last two cases are for evaluating rotation invariance. In general, it is expected that a convolution with rotation invariance should generalize well in case (3) even though the network is not trained with data augmented with SO3 rotations.

In general, our result demonstrates the effectiveness of the rotation invariant convolution we proposed. Our networks yield very consistent results despite that our networks are trained with a limited set of rotated point clouds and tested with arbitrary rotations. To the best of our knowledge, there is no previous work for point cloud learning that is able to achieve the same level of consistency despite that some methods demonstrated good performance when trained with a particular set of rotations. We detail our evaluations below.

The classification task is trained on the ModelNet40 variant of the ModelNet dataset . ModelNet40 contains CAD models from 40 categories such as airplane, car, bottle, dresser, etc. By following Qi et al. , we use the preprocessed 9,8439,843 models for training and 2,4682,468 models for testing. The input point cloud size is 1024, with each point represented by (x,y,z)(x,y,z) coordinates in the Euclidean space.

We followed Li et al. and use multiple feature vectors to train the classifier. Particularly, our network outputs 6464 feature vectors of length 512512 to the classifier. Each of these vectors is passed through an mlpmlp implemented by fully connected layers, resulting in 64×4064\times 40 category predictions. During training, we apply cross entropy loss to all such predictions. During testing, we take the mean of such predictions to obtain the final category prediction. In Section 5.3, we further evaluate this strategy and show that it leads to better performance than networks with a single feature vector.

The evaluation results are shown in Table 1. Following the work of , we perform experiments in three cases: training and testing with data rotated about the gravity axis (z/z), training and testing with arbitrary SO3 rotations (SO3/SO3), and training with z-rotations and testing with SO3 rotations (z/SO3). The first case is commonly adopted by previous methods in handling rotated point clouds, and the last two cases are for evaluating rotation invariance.

We use two criteria for evaluation: accuracy and accuracy standard deviation. Accuracy is a common metric to measure the performance of the classification task. In addition, accuracy deviation measures the consistency of the accuracy scores in three tested cases. In general, it is expected that methods that are rotation invariant should be insusceptible to the rotation used in the training and testing data and therefore has a low deviation in accuracy.

As can be seen, our method performs favorably to the state-of-the-art techniques. On one hand, our method achieves very good accuracy in all cases despite that there are no clear winner for all cases in our experiment. On the other hand, and more importantly, our method has the lowest accuracy deviation. Previous methods exhibit large accuracy deviations especially in the extreme z/SO3 case. This case is exceptionally hard for methods that rely on data augmentation to handle rotations . In our observation, such techniques are only able to generalize within the type of rotation they are trained with, and generally fail in the z/SO3 test. This applies to both voxel-based and point-based learning techniques. By contrast, our method has almost no performance difference in three test cases, which confirms the robustness of the rotation invariant geometric cues in our convolution. We also evaluate the accuracy of the classification task per object category. Please see the full results in the supplemental document.

The capability to handle rotation invariance also has a great effect on the number of network parameters. For networks that rely on data augmentation to handle rotations, it requires more parameters to ‘memorize’ the rotations. Networks that are designed to be rotation invariant, such as spherical CNN and ours, have very compact representations. In terms of number of trainable parameters, our network has 0.70 millions (0.70M) of trainable parameters, which is the most compact network in our evaluations. Among the tested methods, only spherical CNN (0.5M) and PointCNN (0.6M) have similar compactness. Our network has 5×5\times less parameters than PointNet (3.5M), about 2×2\times less than PointNet++ (1.4M). The well balance between trainable parameters, accuracy and accuracy deviations makes our method more robust for practical use.

2 Object Part Segmentation

We also evaluated our method with the object part segmentation task that aims to predict the part label for each input point. In this task, we train and test with the ShapeNet dataset that contains 16,88016,880 CAD models in 1616 categories. Each model is annotated with 22 to 66 parts, resulting in a total of 5050 object parts. We follow the standard train/test split with 14,00614,006 models for training and 2,8742,874 models for testing, respectively.

The evaluation results are shown in Table 2. As can be seen, our method outperforms previous methods significantly in z/SO3 test case and achieves similar performance in SO3/SO3 case. This result aligns well with the performance reported in the object classification task. Our method also has consistent performance for both rotation cases, which empirically confirms the rotation invariance in our convolution. Visualization of our prediction and the ground truth object parts are shown in Figure 4. It is easy to observe that our predictions are the closest to the ground truth. Table 5 and Table 6 further report per-class accuracies for both SO3/SO3 and z/SO3 case. Our method performs best in 3 out of 16 categories in SO3/SO3 case, and 15 out of 16 categories in z/SO3 case.

3 Evaluations of Network Designs

In this section, we perform experiments on object classification to analyze the performance and justify the design of the proposed convolution operator and network architecture. Inspired by the fact that there are negligible improvement after 150150 epochs of training (Section 5.1), we only train the networks with 160160 epochs in this ablation study.

We first experiment by turning on/off different components in our network. The result of this experiment is shown in Table 3. In this table, the Base column indicates a simple network similar to that in Figure 3 but only contains RIConv operators to extract local features for classification. The MLP indicates the use of an MLP layer to lift rotation invariant features to a high-dimensional feature space. The next two columns indicate the geometric attributes used in RIConv. The last row shows that when all components are used, we achieve the best accuracy of 86.5%86.5\%. Without high-dimensional feature learning by MLP (first row), the performance drops by almost 3%3\%. If we either use angle or distance features (second and third row), the accuracy also drops about 1%1\%. This confirms that our network architecture is plausible and yield good performance.

Number of Layers.

We vary the number of convolution layers as follows. Let us denote the convolution layers in our network in Figure 3 with L0L_{0}, L1L_{1}, L2L_{2} from left to right. Here we compare our current architecture with those that have L2L_{2} or L1L_{1} and L2L_{2} removed, or have an additional convolution L−1L_{-1} added before L0L_{0}. Note that we skip point sampling in L−1L_{-1} to keep the same number of input points. The results in Table 4 (first section) shows the accuracy when the number of layers vary from 1 to 4. We can see that with only 1 layer, the accuracy drops dramatically to 46.8%, which means a single convolution cannot extract effective features. With more convolutions, the accuracy is improved but this comes with the cost of longer training time. Thus, in this work, we choose the architecture of 3 layers for best speed and accuracy balance.

Number of Input Points.

We evaluated our network with point clouds of input sizes from 128128 to 10241024 points. Particularly, we retrained and tested the network with point clouds of corresponding number of points. The results are shown in Table 4 (middle section). It shows that our network generalizes well to different input size.

Number of Features for Classifiers.

For object classification, our network outputs 6464 vectors of length 512512 to the classifier. We compared this strategy with the one in PointNet which only outputs a single vector of 512 by maxpooling all features of all points. The results in Table 4 (last section) shows that more output feature vectors yield slightly higher accuracy. Such boost is due to the fact that multiple vectors can convey richer features from different latent spaces that facilitate feature clustering in the classifier.

4 Limitations

Our method is not without limitations. First, the geometric features we used is by no means complete. It is possible to use other more sophisticated low-level geometry features such as curvature to design the convolution. Second, while our convolution is robust and consistent to arbitrary rotations, when there is no rotation or simple rotations as in the z/z case in the classification task, our method is less accurate compared to state-of-the-art classification. This is because the original point coordinates are not retained in low-level geometric feature extraction, trading some discriminative features for rotation invariance.

We perform an additional experiment in which we remove the proposed geometric features, and replace them with the original 3D coordinates of the input point cloud. This makes our convolution no longer robust to SO3 rotations but in return, the convolution features are more discriminative. This allows us to achieve state-of-the-art accuracy (91.8% overall accuracy) in the classification task. Fusing original coordinates and geometric features into the same feature space would be therefore a very interesting extension to this work.

Conclusion

We presented a novel convolution operator for point cloud feature learning that can handle point clouds with arbitrary rotations. Given a point set as input, we determine a reference orientation based on a reference point and the centroid, from which rotation invariant features built upon geometric cues such as distances and angles can be constructed for each point. Combined with a binning strategy, our method handles both rotation invariance and point order issue in a single convolution. We then built a simple yet effective end-to-end convolutional neural network for point cloud classification and segmentation. Experiments demonstrate that our method achieves good performance on both classification and segmentation tasks with the best consistency with arbitrary rotation test cases. This is in contrast to existing methods that often perform quite inconsistently for different types of rotations.

Our method leads to several potential future researches. First, the low-level rotation invariance features for convolution are hand-crafted, which we aim to generalize by applying unsupervised learning to learn such features. Second, our convolution could be beneficial to more scene understanding applications such as object detection and retrieval. It would be also of great interest to extend our method to achieve invariance to rigid and non-rigid transformations.

Acknowledgement. The authors acknowledge support from the SUTD Digital Manufacturing and Design Centre (DManD) funded by the Singapore National Research Foundation. This project is also partially supported by Singapore MOE Academic Research Fund MOE2016-T2-2-154 and Singapore NRF under its Virtual Singapore Award No. NRF2015VSGAA3DCM001-014.

References