Relation-Shape Convolutional Neural Network for Point Cloud Analysis

Yongcheng Liu, Bin Fan, Shiming Xiang, Chunhong Pan

Introduction

Recently, the analysis of 3D point cloud has drawn a lot of attention, as it has many applications such as autonomous driving and robot manipulation. However, this task is very challenging, since it is difficult to infer the underlying shape formed by these irregular points (see Fig. 1 for detail).

For this issue, much effort is focused on replicating the remarkable success of convolutional neural network (CNN) on regular grid data (e.g., image) analysis alexnet; VGG, to irregular point cloud processing c2_pointnet2; c23; c21; c5_generalcnn; c6_synccnn; c10_splatnet; c14_scn. Some works transform point cloud to regular voxels modelnet40; vox2; c45 or multi-view images multiview1; c37; multiview2 for easy application of classic grid CNN. These transformations, however, usually lead to much loss of inherent geometric information in 3D point cloud, as well as high complexity.

To directly process point cloud, PointNet c1_pointnet independently learns on each point and gathers the final features for a global representation. Though impressive, this design ignores local structures that have been proven to be important for abstracting high-level visual concepts in image CNN visualize. To solve this problem, some works partition point cloud into several subsets by sampling c2_pointnet2 or superpoint c8_superpoint. Then a hierarchy is built to learn contextual representation from local to global. Nevertheless, this extremely relies on effective inductive learning of local subsets, which is quite intractable to achieve.

To this end, we propose a relation-shape convolutional neural network (aliased as RS-CNN). The key to RS-CNN is learning from relation, i.e., the geometric topology constraint among points, which in our view can encode meaningful shape information in 3D point cloud.

Specifically, each local convolutional neighborhood is constructed by taking a sampled point xx as the centroid and the surrounding points as its neighbors N(x)\mathcal{N}(x). Then, the convolutional weight is forced to learn a high-level relation expression from predefined geometric priors, i.e., intuitive low-level relation between xx and N(x)\mathcal{N}(x). By convoluting in this way, an inductive representation with explicit reasoning about the spatial layout of points can be obtained. It discriminatively reflects the underlying shape that irregular points form thus is shape-aware. Furthermore, it can benefit from geometric priors, including the invariance to points permutation and the robustness to rigid transformation (e.g., translation and rotation). With this convolution as a basic operator, a hierarchical CNN-like architecture, i.e., RS-CNN, can be developed to achieve contextual shape-aware learning for point cloud analysis.

The key contributions are highlighted as follows:

A novel learn-from-relation convolution operator called relation-shape convolution is proposed. It can explicitly encode geometric relation of points, thus resulting in much shape awareness and robustness;

A deep hierarchy equipped with the relation-shape convolution, i.e., RS-CNN, is proposed. It can extend regular grid CNN to irregular configuration for achieving contextual shape-aware learning of point cloud;

Extensive experiments on challenging benchmarks across three tasks, as well as thorough empirical and theoretical analysis, demonstrate RS-CNN achieves the state of the arts.

Related Work

View-based and volumetric methods. View-based methods represent a 3D shape as a group of 2D views from different angles. Recently, many works multiview1; c37; multiview2; multiview3; c49; c51 have been proposed to recognize these view images with deep neural networks. They often finetune a pre-trained image-based architecture for accurate recognition. However, 2D projections could cause loss of shape information due to self-occlusions, and it often demands a huge number of views for decent performance.

Volumetric methods convert the input 3D shape into a regular 3D grid, over which classic CNN can be employed modelnet40; vox2; c45. The main limitation is the quantization loss of the shape due to the low resolution enforced by 3D grid. Recent space partition methods like K-d trees c26 or octrees c29; vox3; c53 rescue some resolution issues but still rely on the subdivision of a bounding volume rather than a local geometric shape. In contrast to these methods, our work aims to process 3D point cloud directly.

Deep learning on point cloud. PointNet c1_pointnet pioneers this route by independently learning on each point and gathering the final features with max pooling. Yet this design neglects local structures, which have been proven important for the success of CNN. To remedy this, PointNet++ c2_pointnet2 suggests a hierarchical application of PointNet to multiple subsets of point cloud. Local structure exploitation with PointNet is also investigated in PCPNet; c9_kcnet. In addition, Superpoint c8_superpoint is proposed to partition point cloud into geometric elements. Graph convolution network is applied on a local graph created by neighboring points c14_scn; c19; c30. However, these methods do not explicitly model the local spatial layout of points, thus acquiring less shape awareness. By contrast, our work captures the spatial layout of points by learning a high-level relation expression among points.

Some works map point cloud to a high-dimensional space to facilitate the application of classic CNN. SPLATNet c10_splatnet maps the input points onto a sparse lattice, then processing with bilateral convolution bcl. PCNN c16_eocnn extends the function over point cloud to a continuous volumetric function over ambient space. These methods could cause loss of geometric information, while our method directly operates on point cloud without introducing such loss.

Another key issue is the irregularity of points. Some works focus on analyzing symmetric functions that are equivariant to point sets learning c1_pointnet; c6_synccnn; c24; c20. Some other works c1_pointnet; c27 develop alignment network for the robustness to rigid transformation in 3D space. However, the alignment learning is a suboptimal solution for this issue. Some traditional descriptors like Fast Point Feature Histograms can be invariant to translation and rotation, yet they are often less effective for high-level shape understanding. Our method that learns on geometric relation among points is naturally robust to rigid transformation, whilst being highly effective due to the powerfulness of deep nets.

Relation learning. To learn a data-dependent weight from relation has been explored in the field of image and video analysis. Spatial transformer stn learns a transition matrix to align 2D images. Non-local network non-local learns long-term relation across video frames. Relation networks relation_detection learn position relation across objects. DFN dfn_net proposes general dynamic filters that inspire many subsequent works.

There are also some works focusing on the relation learning in 3D point cloud. DGCNN c22 captures similar local shapes by learning point relation in a high-dimensional feature space, yet this relation could be unreliable in some cases. Wang et al. Param_conv propose a parametric continuous convolution that is based on computable relation among points, but they do not explicitly learn from local to global like classic CNN. By contrast, our method learns a high-level relation expression from geometric priors in 3D space, and performs contextual local-to-global shape learning.

Shape-Aware Representation Learning

The core of point cloud analysis is to discriminatively represent the underlying shape with robustness. Here we learn contextual shape-aware representation for this goal, by extending regular grid CNN to irregular configuration with a novel relation-shape convolution (RS-Conv).

Local-to-global learning, which has gained remarkable success in image CNN alexnet; VGG, is a promising solution for contextual shape representation. However, it extremely relies on shape-aware inductive learning from irregular point subsets, which remains a quite intractable problem.

2 Properties

RS-Conv in Eq. (3) can maintain four decent properties:

Points interaction. Points are not isolated and nearby points form a meaningful shape in geometric space. Thus their inherent interaction is critical for discriminative shape awareness. Our solution of relation learning explicitly encode the geometric relation among points, naturally capturing the interaction of points.

3 Revisiting 2D Grid Convolution

The proposed RS-Conv is a generic formulation of 2D grid convolution for relation reasoning. We clarify this with a neighborhood (convolution kernel) of 3×33\times 3 on a 2D-grid feature map, as illustrated in Fig. 3. Specifically, the summation function ∑\sum is a specific instance of the aggregation function A\mathcal{A}. Moreover, note that wjw_{j} always implies a fixed positional relation between xix_{i} and its neighbor xjx_{j} in the regular grid. For example, w1w_{1} always implies the top-left relation with xix_{i}, and w2w_{2} implies the right-above relation with xix_{i}. In other words, wjw_{j} is actually constrained to encode one kind of regular grid relation in the learning process. Therefore, our RS-Conv with relation learning is more general and can be applied to model 2D grid spatial relationship.

4 RS-CNN for Point Cloud Analysis

Using RS-Conv (Fig. 2) as a basic operator and adopting a uniform sampling strategy, a hierarchical shape-aware learning architecture like classic CNN, namely, RS-CNN, can be developed for point cloud analysis as

Our RS-CNN applied in the classification and segmentation of point cloud is illustrated in Fig. 4. In both tasks, RS-CNN is used for learning a group of hierarchical shape-aware representation. The final global representation followed by three fully connected (FC) layers is configured for classification. For segmentation, the learned multi-level representation is successively upsampled by feature propagation c2_pointnet2 to generate per-point predictions. Both of them can be trained in an end-to-end manner.

5 Implementation Details

Our RS-CNN is implemented using Pytorch https://github.com/Yochengliu/Relation-Shape-CNN. The Adam optimization algorithm is employed for training, with a mini-batch size of 3232. The momentum for BN starts with 0.90.9 and decays with a rate of 0.50.5 every 2020 epochs. The learning rate begins with 0.0010.001 and decays with a rate of 0.70.7 every 2020 epochs. The weight of RS-CNN is initialized using the techniques introduced by He et al. conf_iccv_HeZRS15.

Experiment

In this section, we arrange comprehensive experiments to validate the proposed RS-CNN. First, we evaluate RS-CNN for point cloud analysis on three tasks (Sec 4.1). We then provide detailed experiments to carefully study RS-CNN (Sec 4.2). Finally, we visualize the shape features that RS-CNN captures and analyze the complexity (Sec 4.3).

Shape classification. We evaluate RS-CNN on ModelNet40 classification benchmark modelnet40. It is composed of 9843 train models and 2468 test models in 40 classes. The point cloud data is sampled from these models by c1_pointnet. We uniformly sample 1024 points and normalize them to a unit sphere. During training, we augment the input data with random anisotropic scaling in the range [-0.66, 1.5] and translation in the range [-0.2, 0.2], as in c26. Meanwhile, dropout technique dropout with 50% ratio is applied in FC layers. During testing, similar to c1_pointnet; c2_pointnet2, we perform ten voting tests with random scaling and average the predictions.

We test the robustness of RS-CNN on sampling density, by using sparser points of number 1024, 512, 256, 128 and 64 as the input to a model trained with 1024 points. As in c2_pointnet2, random input dropout technique is applied for a fair comparison. Fig. 5 shows the test results, where the compared methods are PointNet c1_pointnet, PointNet++ c2_pointnet2, PCNN c16_eocnn and DGCNN c22. As can be seen, it is more difficult for shape recognition when points get sparser. Even so, RS-CNN is still considerably robust. It achieves nearly consistent robustness as PointNet++, whilst showing superior performance on each density.

Shape part segmentation. Part segmentation is a challenging task for fine-grained shape analysis. We evaluate RS-CNN for this task on ShapeNet part benchmark c54 and follow the data split in c1_pointnet. This dataset contains 16881 shapes with 16 categories, and is labeled in 50 parts in total. As in c1_pointnet, we randomly pick 2048 points as the input and concatenate the one-hot encoding of the object label to the last feature layer. During testing, we also apply ten voting tests using random scaling. Except for standard IoU (Inter-over-Union) on each category, we also report two types of mean IoU (mIoU) that are averaged across all classes and all instances, respectively.

Normal estimation. Normal estimation in point cloud is a crucial step for numerous applications, such as surface reconstruction and rendering. This task is very challenging since it requires a higher level of reasoning, which goes beyond the underlying shape recognition. We take normal estimation as a supervised regression task, and achieve it using the segmentation network. The cosine-loss between the normalized output and ground truth normal is applied for regression training. ModelNet40 dataset is used for evaluation, with uniformly sampled 1024 points as the input.

The quantitative results are summarized in Table 3. RS-CNN outperforms other advanced methods on this task with a lower error of 0.15. This significantly reduces the error of PointNet++ (0.29) by 48.3%. Fig. 7 shows some normal estimation examples, where our RS-CNN with geometric relation learning can obtain more decent predictions. However, RS-CNN could also be less effective for some intractable shapes, such as spiral stairs and intricate plants.

2 RS-CNN Design Analysis

Ablation study. The results are summarized in Table 4. The baseline (model A) is set to learn without geometric relation encoding, but with a shared three-layer MLP as feature transformation function T\mathcal{T} in Eq. (1).

To investigate the impact of the number of input points on RS-CNN, we also train the network with 2048 points but find no improvement (model H). In addition, to compare with the baseline (model A) more fairly, we set a new baseline (model I) that works with all the techniques but relation learning. It gets an accuracy of 90.1%, which RS-CNN can also surpass by 3.5%. We speculate that RS-CNN with geometric relation reasoning can acquire more discriminative shape awareness, and this awareness can be greatly enhanced by multi-scale relation learning.

Aggregation function A\mathcal{A}. Three symmetric functions: max pooling (max), average pooling (avg.) and summation (sum), are employed to study the effect of A\mathcal{A} on RS-CNN. Table 5 summarizes the results. As can be seen, with M\mathcal{M} using three layers, max pooling achieves the best performance while average pooling and summation get the same accuracy. The reason may be that max pooling can select the biggest feature response, thus keeping the most expressive representation and removing redundant information.

Mapping function M\mathcal{M}. The results of M\mathcal{M} deployed with different layers are summarized in the first three rows of Table 5. One can see that the best accuracy of 93.6% is obtained by a shared three-layer MLP, and it decreases by 0.9% when increasing the number of layers. The reason might be that M\mathcal{M} with four layers brings some difficulty for network training. Noticeably, RS-CNN can also get a decent accuracy of 92.4% with M\mathcal{M} using only two layers. This verifies the powerfulness of relation learning for underlying shape capturing from point cloud.

As can be seen, all the methods are invariant to permutation. However, PointNet is vulnerable to both translation and rotation while PointNet++ is sensitive to rotation. By contrast, our RS-CNN with geometric relation learning is invariant to these perturbations, making it powerful for robust shape recognition.

3 Visualization and Complexity Analysis

Visualization. Fig. 8 visualizes the shape features learned by the first two layers of RS-CNN on ModelNet40 dataset. As it shows, the features learned by the first layer mostly respond to edges, corners and arcs, while the ones in the second layer capture more semantical shape parts like airfoils and heads. This verifies RS-CNN can learn progressive shape-aware representation for point cloud analysis.

Complexity Analysis. Table 8 summarizes the space (number of params) and the time (floating point operations/sample) complexity of RS-CNN in classification with 1024 points as the input. Compared with PointNet c1_pointnet, RS-CNN reduces the params by 59.7% and the FLOPs by 32.9%, which shows its great potential for real-time applications, e.g., scene parsing in autonomous driving.

Conclusion

In this work, RS-CNN, namely, Relation-Shape Convolutional Neural Network, which extends regular grid CNN to irregular configuration for point cloud analysis, has been proposed. The core to RS-CNN is a novel convolution operator, which learns from relation, i.e., the geometric topology constraint among points. In this way, explicit reasoning about the spatial layout of points can be made to obtain discriminative shape awareness. Moreover, the decent properties of geometric relation can also be acquired, such as robustness to rigid transformation. As a consequence, RS-CNN equipped with this operator can achieve contextual shape-aware learning, making it highly effective. Extensive experiments on challenging benchmarks across three tasks, as well as thorough empirical and theoretical analysis, have demonstrated RS-CNN achieves the state of the arts.

References

A Outline

This supplementary material provides further investigations for the proposed RS-CNN. Specifically, three issues on the construction of local neighborhood are discussed in Sec B. More details of the relation learning on 2D views of 3D point cloud are presented in Sec C. All the experiments are conducted on ModelNet40 dataset.

B Construction of Local Neighborhood

In the above process, there are mainly three issues worth further investigation: (1) How should N(xi)\mathcal{N}(x_{i}) be selected? (2) Is it suitable to simply aggregate all the relation between xix_{i} and N(xi)\mathcal{N}(x_{i})? (3) Is it reasonable to select the sampled point xix_{i} as the centroid? They are explored as follows.

(1) Selection of the neighbors N(xi)\mathcal{N}(x_{i}). Two strategies, k-nearest neighbor (k-NN) and random picking in the ball (Random-PIB), are investigated for this issue. Table 9 summarizes the results. Note that the number of neighbors is set to be equal for a fair comparison. As it shows, Random-PIB obtains better classification accuracy. The reason may be k-NN would suffer selection inhomogeneity in some cases, which is adverse to shape-aware learning (the aggregated relation may only focus on dense points and ignore sparse points that are essential for the underlying shape). By contrast, Random-PIB can have a better coverage of points even in the case of inhomogeneous distribution.

(3) Selection of the centroid. Three types of the centroid: the sampled point xix_{i}, the average of N(xi)\mathcal{N}(x_{i}) and random picking in N(xi)\mathcal{N}(x_{i}), are studied for this issue. Besides, a strategy that fuses all of them is also studied. The results are summarized in Table 11, where the first two strategies obtain the same decent accuracy while random picking performs less well. The reason may be that random picking requires RS-CNN to reason the spatial layout of points from various topological connections, which is quite difficult.

More details of the relation learning on 2D views of point cloud (the fourth part in Sec 4.2) are provided in this section. As illustrated in Fig. 9 in this material, the relation among points in the 2D view can also reflect the underlying shape. Therefore, we are interested in how powerfully the proposed RS-CNN to acquire shape awareness from only 2D-view relation of points.