PointCNN: Convolution On $\mathcal{X}$-Transformed Points
Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, Baoquan Chen
Introduction
Spatially-local correlation is a ubiquitous property of various types of data that is independent of the data representation. For data that is represented in regular domains, such as images, the convolution operator has been shown to be effective in exploiting that correlation as the key contributor to the success of CNNs on a variety of tasks lecun2015deep . However, for data represented in point cloud form, which is irregular and unordered, the convoralution operator is ill-suited for leveraging spatially-local correlations in the data.
In this paper, we propose to learn a -transformation for the coordinates of input points , with a multilayer perceptron Rumelhart_1986 , i.e., . Our aim is to use it to simultaneously weight and permute the input features, and subsequently apply a typical convolution on the transformed features. We refer to this process as -Conv, and it is the basic building block for our PointCNN. The -Conv for , , and in Figure 1 can be formulated as in Eq. 1b, where the s are matrices, as in this figure. Note that since and are learned from points of different shapes, they can differ so as to weight the input features accordingly, and achieve . For and , if they are learned to satisfy , where is the permutation matrix for permuting into , then can be achieved.
From the analysis of the example in Figure 1, it is clear that, with ideal -transformations, -Conv is capable of taking the point shapes into consideration, while being invariant to ordering. In practice, we find that the learned -transformations are far from ideal, especially in terms of the permutation equivalence aspect. Nevertheless, PointCNN built with -Conv is still significantly better than a direct application of typical convolutions on point clouds, and on par or better than state-of-the-art neural networks designed for point cloud input data, such as PointNet++ Qi_NIPS17 .
Section 3 contains the details of -Conv, as well as PointCNN architectures. We show our results on multiple challenging benchmark datasets and tasks in Section 4, together with ablation experiments and visualizations for a better understanding of PointCNN.
Related Work
CNNs have been very successful for leveraging spatially-local correlation in images — pixels in 2D regular grids lecun1998gradient . There has been work in extending CNNs to higher dimensional regular domains, such as 3D voxels Wu_CVPR15 . However, as both the input and convolution kernels are of higher dimensions, the amount of both computation and memory inflates dramatically. Octree Riegler_CVPR17 ; Wang_SIGGRAPH17 , Kd-Tree Klokov_ICCV17 and Hash Shao_arXiv18 based approaches have been proposed to save computation by skipping convolution in empty space. The activations are kept sparse in SubmanifoldSparseConvNet to retain sparsity in convolved layers. Hua_CVPR18 and BenShabat_arXiv18 partition point cloud into grids and represent each grid with grid mean points and Fisher vectors respectively for convolving with 3D kernels. In these approaches, the kernels themselves are still dense and of high dimension. Sparse kernels are proposed in Li_NIPS16 , but this approach cannot be applied recursively for learning hierarchical features. Compared with these methods, PointCNN is sparse in both input representation and convolution kernels.
Stimulated by the rapid advances and demands in 3D sensing, there has been quite a few recent developments in feature learning from 3D point clouds. PointNet Qi_CVPR17 and Deep Sets Zaheer_NIPS17 proposed to achieve input order invariance by the use of a symmetric function over inputs. PointNet++ Qi_NIPS17 and SO-Net Li_CVPR18 apply PointNet hierarchically for better capturing of local structures. Kernel correlation and graph pooling are proposed for improving PointNet-like methods in Shen_CVPR18 . RNN is used in Huang_CVPR18 for processing features aggregated by pooling from ordered point cloud slices. Wang_arXiv18_mit proposed to leverage neighborhood structures in both point and feature spaces. While these symmetric pooling based approaches, as well as those in Dieleman_ICML16 ; Zaheer_NIPS17 ; ravanbakhsh2016deep , have guarantee in achieving order invariance, they come with a price of throwing away information.
Su_CVPR18 ; Atzmon_SIGGRAPH18 ; Tatarchenko_CVPR18 propose to first “interpolate” or “project” features into predefined regular domains, where typical CNNs can be applied. In contrast, the regular domain is latent in our method. CNN kernels are represented as parametric functions of neighborhood point positions to generalize CNNs for point clouds in Wang_CVPR18_Deep ; Groh_arXiv18 ; Xu_arXiv18 . The kernels associated with each point are parametrized individually in these methods, while the -transformations in our method are learned from each neighborhood, thus could potentially by more adaptive to local structures.
Besides as point clouds, sparse data in irregular domains can be represented as graphs, or meshes, and a few works have been proposed for feature learning from such representations Monti_CVPR17 ; Yi_CVPR17 ; Maron_SIGGRAPH17 . We refer the interested reader to Bronstein_17 for a comprehensive survey of work along these directions. Spectral graph convolution on a local graph is used for processing point clouds in Wang_arXiv18_mg .
A line of pioneering work aiming at achieving equivariance has been proposed to address the information loss problem of pooling in achieving invariance hinton2011transforming ; sabour2017dynamic . The -transformations in our formulation, ideally, are capable of realizing equivariance, and are demonstrated to be effective in practice. We also found similarity between PointCNN and Spatial Transformer Networks Jaderberg_NIPS15 , in the sense that both of them provided a mechanism to “transform” input into latent canonical forms for being further processed, with no explicit loss or constraint in enforcing the canonicalization. In practice, it turns out that the networks find their ways to leverage the mechanism for learning better. In PointCNN, the -transformation is supposed to serve for both weighting and permutation, thus is modelled as a general matrix. This is different than that in Cruz_CVPR17 , where a permutation matrix is the desired output, and is approximated by a doubly stochastic matrix.
PointCNN
The hierarchical application of convolutions is essential for learning hierarchical representations via CNNs. PointCNN shares the same design and generalizes it to point clouds. First, we introduce hierarchical convolutions in PointCNN, in analogy to that of image CNNs, then, we explain the core -Conv operator in detail, and finally, present PointCNN architectures geared toward various tasks.
Before we introduce the hierarchical convolution in PointCNN, we briefly go through its well known version for regular grids, as illustrated in Figure 2 upper. The input to grid-based CNNs is a feature map of shape , where is the spatial resolution, and is the feature channel depth. The convolution of kernels of shape against local patches of shape from , yields another feature map of shape . Note that in Figure 2 upper, , , and . Compared with , is often of lower resolution () and of deeper channels (), and encodes higher level information. This process is recursively applied, producing feature maps with decreasing spatial resolution ( in Figure 2 upper), but deeper channels (visualized by increasingly thicker dots in Figure 2 upper).
The representative points should be the points that are beneficial for the information “projection” or “aggregation”. In our implementation, they are generated by random down-sampling of in classification tasks, and farthest point sampling in segmentation tasks, since segmentation tasks are more demanding on a uniform point distribution. We suspect some more advanced point selections which have shown promising performance in geometry processing, such as Deep Points Wu_SIGGRAPH15 , could fit in here as well. We leave the exploration of better representative point generation methods for future work.
2 𝒳𝒳\mathcal{X}-Conv Operator
where is a multilayer perceptron applied individually on each point, as in PointNet Qi_CVPR17 . Note that all the operations involved in building -Conv, i.e., Conv, , matrix multiplication , and , are differentiable. Accordingly. -Conv is differentiable, and can be plugged into a neural network for training by back propagation.
Lines 4-6 in Algorithm 1 are the core -transformation as described in Eq. 1b in Section 1. Here, we explain the rationale behind lines 1-3 of Algorithm 1 in detail. -Conv is designed to work on local point regions, and the output should not be dependent on the absolute position of and its neighboring points, but on their relative positions. To that end, we position local coordinate systems at the representative points (line 1 of Algorithm 1, Figure 3b). It is the local coordinates of neighboring points, together with their associated features, that define the output features. However, the local coordinates are of a different dimensionality and representation than the associated features. To address this issue, we first lift the coordinates into a higher dimensional and more abstract representation (line 2 of Algorithm 1), and then combine it with the associated features (line 3 of Algorithm 1) for further processing (Figure 3c).
Lifting coordinates into features is done through a point-wise , as in PointNet-based methods. Differently, however,the lifted features are not processed by a symmetric function. Instead, along with the associated features, they are weighted and permuted by the -transformation that is jointly learned across all neighborhoods. The resulting is dependent on the order of the points, and this is desired, as is supposed to permute according to the input points, and therefore has to be aware of the specific input order. For an input point cloud without any additional features, i.e., is empty, the first -Conv layer uses only . PointCNN can thus handle point clouds with or without additional features in a robust uniform fashion.
For more details about the -Conv operator, including the actual definition of , and Conv, please refer to Supplementary Material Section 1.
3 PointCNN Architectures
From Figure 2, we can see that the Conv layers in grid-based CNNs and -Conv layers in PointCNN only differ in two aspects: the way the local regions are extracted ( patches vs. neighboring points around representative points) and the way the information from local regions is learned (Conv vs. -Conv). Otherwise, the process of assembling a deep network with -Conv layers highly resembles that of grid-based CNNs.
Figure 4a depicts a simple PointCNN with two -Conv layers that gradually transform the input points (with or without features) into fewer representation points, but each with richer features. After the second -Conv layer, there is only one representative point left, and it aggregates information from all the points from the previous layer. In PointCNN, we can roughly define the receptive field of each representative point as the ratio , where is the neighboring point number, and is the point number in the previous layer. With this definition, the final point “sees” all the points from the previous layer, thus has a receptive field of — it has a global view of the entire shape, and its features are informative for semantic understanding of the shape. We can add fully connected layers on top of the last -Conv layer output, followed by a loss, for training the network.
Note that the number of training samples for the top -Conv layers drops rapidly (Figure 4a), making it inefficient to train them thoroughly. To address this problem, we propose PointCNN with denser connections (Figure 4b), where more representative points are kept in the -Conv layers. However, we aim to maintain the depth of the network, while keeping the receptive field growth rate, such that the deeper representative points “see” increasingly larger portions of the entire shape. We achieve this goal by employing the dilated convolution idea from grid-based CNNs in PointCNN. Instead of always taking the neighboring points as input, we uniformly sample input points from neighboring points, where is the dilation rate. In this case, the receptive field increases from to , without increasing actual neighboring point count or kernel size.
In the second -Conv layer of PointCNN in Figure 4b, dilation rate is used, thus all the four remaining representative points “see” the entire shape, and all of them are suitable for making predictions. Note that, in this way, we can train the top -Conv layers more thoroughly, as much more connections are involved in the network, compared to PointCNN in Figure 4a. In test time, the output from the multiple representative points is averaged right before the to stabilize the prediction. This design is similar to that of Network in Network Lin_ICLR14 . The denser version of PointCNN (Figure 4b) is the one we used for classification tasks.
For segmentation tasks, high resolution point-wise output is required, and this can be realized by building PointCNN following Conv-DeConv Noh_ICCV15 architecture, where the DeConv part is responsible for propagating global information into high resolution predictions (see Figure 4c). Note that both the “Conv” and “DecConv” in the PointCNN segmentation network are the same -Conv operator. The only differences between the “Conv” and “DeConv” layers is that the latter has more points but less feature channels in its output vs. its input, and its higher resolution points are forwarded from earlier “Conv” layers, following the design of U-Net Ronneberger_MICCAI15 .
To train the parameters in -Conv, it is evidently not beneficial to keep using the same set of neighboring points, in the same order, for a specific representative point. To improve generalization, we propose to randomly sample and shuffle the input points, such that both the neighboring point sets and order may differ from batch to batch. To train a model that takes points as input, points are used for training, where denotes a Gaussian distribution. We found that this strategy is crucial for successful training of PointCNN.
Experiments
We conducted an extensive evaluation of PointCNN for shape classification on six datasets (ModelNet40 Wu_CVPR15 , ScanNet dai2017scannet , TU-Berlin Eitz_SIGGRAPH12 , Quick Draw ha2017neural , MNIST, CIFAR10), and segmentation task on three datasets (ShapeNet Parts Yi_SIGGRAPHAsia16 , S3DIS armeni20163d , and ScanNet dai2017scannet ). The details of the datasets and how we convert and feed data into PointCNN, are described in Supp. Material Section 2, and the PointCNN architectures for the tasks on these datasets can be found in Supp. Material Section 3.
We summarize our 3D point cloud classification results on ModelNet40 and ScanNet in Table 1, and compare to several neural network methods designed for point clouds. Note that a large portion of the 3D models from ModelNet40 are pre-aligned to the common up direction and horizontal facing direction. If a random horizontal rotation is not applied on either the training or testing sets, then the relatively consistent horizontal facing direction is leveraged, and the metrics based on this setting is not directly comparable to those with the random horizontal rotation. For this reason, we ran PointCNN and reported its performance in both settings. Note that PointCNN achieved top performance on both ModelNet40 and ScanNet.
We evaluate PointCNN on the segmentation of ShapeNet Parts, S3DIS, and ScanNet datasets, and summarize the results in Table 2. More detailed segmentation result comparisons can be found in Supplementary Material Section 4. We note that PointCNN outperforms all the compared methods, including SSCN graham20173d , SPGraph Landrieu_arXiv17 and SGPN Wang_CVPR18 , which are specialized segmentation networks with state-of-the-art performance. Note that the part averaged IoU metric for ShapeNet Parts is the one used in ShapeNet17 . Compared with mean IoU, the part averaged IoU puts more emphasis on the correct prediction of small parts.
Sketches are 1D curves in 2D space, thus can be more effectively represented with point clouds, rather than with 2D images. We evaluate PointCNN on TU-Berlin and Quick Draw sketches, and present results in Table 4, where we compare its performance with the competitive PointNet++, as well as image CNN based methods. PointCNN outperforms PointNet++ on both datasets, with a more prominent advantage on Quick Draw (25M data samples), which is significantly larger than TU-Berlin (0.02M data samples). On the TU-Berlin dataset, while the performance of PointCNN is slightly better than the generic image CNN AlexNet krizhevsky2012imagenet , there is still a gap with the specialized Sketch-a-Net Yu_IJCV17 . It is interesting to study whether architectural elements from Sketch-a-Net can be adopted and integrated into PointCNN to improve its performance on the sketch datasets.
Since -Conv is a generalization of Conv, ideally, PointCNN should perform on par with CNNs, if the underlying data is the same, but only represented differently. To verify this, we evaluate PointCNN on the point cloud representation of MNIST and CIFAR10, and show results in Table 4. For MNIST data, PointCNN achieved comparable performance with other methods, indicating its effective learning of the digits’ shape information. For CIFAR10 data, where there is mostly no “shape” information, PointCNN has to learn mostly from the spatially-local correlation in the RGB features, and it performed reasonably well on this task, though there is a large gap between PointCNN and the mainstream image CNNs. From this experiment, we can conclude that CNNs are still the better choice for general images.
2 Ablation Experiments and Visualizations
To verify the effectiveness of the -transformation, we propose PointCNN without it as a baseline, where lines 4-6 of Algorithm 1 are replaced by Conv. Compared with PointCNN, the baseline has less trainable parameters, and is more “shallow” due to the removal of in line 4 of Algorithm 1. For a fair comparison, we further propose PointCNN w/o -W/D, which is wider/deeper, and has approximately the same amount of parameters as PointCNN. The model depth of PointCNN w/o (deeper) also compensates for the decrease in depth caused by the removal of from PointCNN. The comparison results are summarized in Table 5. Clearly, PointCNN outperforms the proposed variants by a significant margin, and the gap between PointCNN and PointCNN w/o is not due to model parameter number, or model depth. With these comparisons, we conclude that -Conv is the key to the performance of PointCNN.
To verify this, we show T-SNE visualization of , and of randomly picked representative points from the ModelNet40 dataset in Figure 5, each with one color, and consistent in the sub-figures. Note that is quite “blended”, which indicates that the features from different representative points are not discriminative against each other (Figure 5a). While is better than , it is still “fuzzy” (Figure 5b). In Figure 5c, are “concentrated” by , and the features of each representative point become highly discriminative. To give an quantitative reference of the “concentration” effect, we firstly compute the feature centers of different representative points, then classify all the feature points to the representative points they belong to, based on nearest search to the centers. The classification accuracies are 76.83%, 89.29% and 94.72% for , and , respectively. With the qualitative visualization and quantitative investigation, we conclude that though the “concentration” is far from reaching a point, the improvement is significant, and it explains the performance of PointCNN in feature learning.
We implemented PointCNN in tensorflow tensorflow , and use ADAM optimizer Kingma_ICLR14 with an initial learning rate for the training of our models. As shown in Table 6, we summarize our running statistics based with the model for classification with batch size 16, 1024 input points on nVidia Tesla P100 GPU, in comparison with several other methods. PointCNN achieves 0.031/0.012 second per batch for training/inference on this setting. In addition, the model for segmentation with input points has M parameters runs on nVidia Tesla P100 with batch size at 0.61/0.25 second per batch for training/inference.
Conclusion
We proposed PointCNN, which is a generalization of CNN into leveraging spatially-local correlation from data represented in point cloud. The core of PointCNN is the -Conv operator that weights and permutes input points and features before they are processed by a typical convolution. While -Conv is empirically demonstrated to be effective in practice, a rigorous understanding of it, especially when being composited into a deep neural network, is still an open problem for future work. It is also interesting to study how to combine PointCNN and image CNNs to jointly process paired point clouds and images, probably at the early stages. We open source our code at https://github.com/yangyanli/PointCNN to encourage further development.
Yangyan would like to thank Leonidas Guibas from Stanford University and Mike Haley from Autodesk Research for insightful discussions, and Noa Fish from Tel Aviv University and Thomas Schattschneider from Technical University of Hamburg for proof reading. The work is supported in part by National Key Research and Development Program of China grant No. 2017YFB1002603, the National Basic Research grant (973) No. 2015CB352501, National Science Foundation of China General Program grant No. 61772317, and “Qilu” Young Talent Program of Shandong University.
References
𝒳𝒳\mathcal{X}-Conv Details
We implement in Line 2 of Algorithm 1 with two fully connected (FC) layers, each followed by ELU Clevert_ICLR16 activation function and batch normalization (BN) Ioffe_ICML15 , i.e., . We set to .
Conv in Line 6 of Algorithm 1, if implemented with typical convolution, has trainable parameters. We implemented it with separable convolution chollet2016xception , which has trainable parameters, where is the depth multiplier, and we use in our implementation. Separable convolution reduces both parameter number and computation compared with that of a typical convolution.
is used to harvest the global position information of the representative points in the last -Conv layer. It is implemented similar to , i.e., . We set to . The dimensional output of is concatenated with the dimensional output of the last -Conv layer for further processing.
In our implementation, a nearest neighbor search is applied for extracting the neighboring points. This assumes a more or less uniform distribution of input points. For point clouds with non-uniform distribution, a radius search can be applied first, and then points can be randomly picked out of the radius search results.
In theory, the -transformation can be applied on either the features or the kernels. We opt to apply it on the features, in which way the follow up operation is a standard convolution operation that is highly optimized by popular deep learning frameworks.
Dataset Details
We conducted extensive evaluation of PointCNN on datasets of various types and scales. Here we introduce the details of the datasets, as well as how we pre-process and feed them into PointCNN:
Object datasets: ModelNet40 Wu_CVPR15 and ShapeNet Parts Yi_SIGGRAPHAsia16 .
ModelNet40 is composed of 3D mesh models from categories, with a training/testing split. Both the gravity and “facing” directions of the models are mostly aligned in the dataset. In the “Pre-aligned” setting, the models are used for training and testing, without random horizontal rotations. In which way, the relative consistent “facing” direction is leveraged by the network. In the “Unaligned” setting, random horizontal rotations are explicitly applied on either the training or the testing models, not as a data augmentation, but to “forget” the relative consistent “facing” directions thus better approximate the scenarios in real world applications, where the “facing” direction of the objects are often unknown. We use the point cloud conversion of ModelNet40 provided by Qi_CVPR17 as our input, where points are sampled from each mesh, and we further sample points to train a model for testing with points on the classification task.
ShapeNet Parts contains models (/ training/testing split) from shape categories, each annotated with 2 to 6 parts and there are 50 different parts in total. Each point sampled from the models is associated with a part label. The task is to predict the part label for each point, thus a segmentation task, and can be treated as a dense point-wise classification problem. The category label for each model is given, and can be used for trimming irrelevant predictions, same as that in graham20173d . points are sampled from each point cloud to train a model for testing with input points on the segmentation task. Each testing point cloud is sampled multiple times to make sure all the points are evaluated at least ( in our experiments) times at testing time.
Indoor scene datasets: S3DIS armeni20163d and ScanNet dai2017scannet . While ModelNet40 and ShapeNet models are mostly made by 3D modeling tools, S3DIS and ScanNet are from real scans of indoor environments.
S3DIS contains 3D scans from Matterport scanners in areas including rooms. Each point with RGB features in the scan is annotated with one of the semantic labels from 13 categories. The task is segmentation. The data is firstly split by room, and then the rooms are sliced into m by m blocks, with m padding on each side. The points in the padding areas serve as context of the internal points, and themselves are not linked to loss in the training phase, nor used for prediction in the testing phase. Each block is moved to a local coordinate system defined by its center. Random horizontal rotations are applied on the sliced blocks for data augmentation. The rotated blocks are handled in the same way as the object point clouds in ShapeNet Parts.
ScanNet contains scanned and reconstructed indoor scenes, with scenes for training/testing in semantic voxel labeling of categories. We firstly prepare data in the same way as that of S3DIS to train a segmentation model, and the segmentation results on testing data are then converted into semantic voxel labeling, as that in Qi_NIPS17 , for a fair comparison with previous methods. The training/testing object instances from the categories in ScanNet are also used for evaluating classification task. Note that ScanNet comes with RGB information for each point. However, they are not used in previous methods. To make fair comparisons, we do not use them either.
2D sketch datasets: TU-Berlin Eitz_SIGGRAPH12 and Quick Draw ha2017neural . Similar to surfaces in 3D space, line sketches in 2D are inherently of less dimension than the ambient space, and can be represented as point cloud, thus we consider 2D sketches good arena for evaluating neural networks that are designed to consume point cloud data. TU-Berlin has sketches from categories, with sketches from each category, where are used for training and the rest for testing. Quick Draw is the largest available sketch dataset, with sketches from categories, each with training/testing samples. We sample points from the sketch stokes to train a model for testing with points on sketch classification task.
Image datasets: MNIST and CIFAR10. MNIST and CIFAR10 are widely used for sanity check of image CNNs. Since PointCNN is a generalization of CNNs, we would like to evaluate PointCNN on the point cloud representation of MNIST and CIFAR10. For MNIST, we randomly sample foreground pixels and convert them into point cloud representation, with the gray-scale pixel value as the input feature. For CIFAR10, we randomly sampled pixels out of the pixels for converting into point cloud with RGB features. Note that there is “shape” information in the MNIST point cloud, sine the point cloud follow the digits’ structure, but this is not the case for the CIFAR10 point cloud, where the points are mostly the same blob for all the data samples.
PointCNN Model Zoo
In Figure 1, we list the PointCNNs used for classification and segmentation tasks on multiple benchmark datasets. PointCNNs are easy to implement, setup, and tune. Larger are used for layers with more abstract/semantic information, such as the top layers in classification networks, and middle layers in “Conv-DeConv” segmentation networks. To relax the memory demand, smaller s are used at layers with large number of representative points, such as bottom layers of classification networks, and top and bottom layers of segmentation networks. Deeper PointCNN with larger receptive field in the last -Conv layer are used for larger or harder datasets. The skip-links, together with the dilation parameter , make it easy to fuse information from different scales (receptive fields), as illustrated in (d) and (e), which is essential for segmentation tasks.
Detailed Segmentation Results
We show detailed segmentation result comparisons on ShapeNet Parts in Table 1, we can see our approach achieves the best overall performance and are best on 7 of the 16 categories.
We show detailed segmentation result comparisons on S3DIS in Table 2, we can see our approach achieves the best overall performance and are best on 6 of the 13 categories. The detailed segmentation result comparisons on S3DIS Area 5 are summarized in Table 3, as some of the literatures only report the performance on this area.