3D-MiniNet: Learning a 2D Representation from Point Clouds for Fast and Efficient 3D LIDAR Semantic Segmentation

Iñigo Alonso, Luis Riazuelo, Luis Montesano, Ana C. Murillo

I INTRODUCTION

Autonomous robotic systems use sensors to perceive the world around them. RGB cameras and LIDAR are very common due to the essential data they provide. One of the key building blocks of autonomous robots is semantic segmentation. Semantic segmentation assigns a class label to each LIDAR point or camera pixel. This detailed semantic information is essential for decision making in real-world dynamic scenarios. LIDAR semantic segmentation provides very useful information to autonomous robots when performing tasks such as Simultaneous Localization And Mapping (SLAM) , autonomous driving or inventory tasks , especially for identifying dynamic objects. In these scenarios, it is critical to have models that provide accurate semantic information in a fast and efficient manner, which is particularly challenging working with 3D LIDAR data. On one hand, the commonly called point-based approaches tackle this problem directly executing 3D point-based operations, which is computationally expensive to operate at high frame rates. On the other hand, approaches that project the 3D information into a 2D image (projection-based approaches) are more efficient but do not exploit the raw 3D information. Recent results on fast and parameter-efficient semantic segmentation models are facilitating the adoption of semantic segmentation in real-world robotic applications .

This work presents a novel fast and parameter-efficient approach for 3D LIDAR semantic segmentation that consists of three modules (as detailed in Sec. III). The main contribution relies on our 3D-MiniNet module. 3D-MiniNet runs the following two steps: (1) It learns a 2D representation from the 3D point cloud (following previous works on 3D object detection [15, 16, zhou2020end]); (2) It computes the segmentation through a fast 2D fully convolutional neural network.

Our best configuration achieves state-of-the-art results in well known public benchmarks (SemanticKITTI and KITTI dataset ) while being faster and more parameter efficient that prior work. Figure 1 shows how 3D-MiniNet achieves better precision-speed trade-off than previous methods. The main novelties with respect to existing approaches, that facilitate these improvements, are:

An extension of MiniNet-v2 for 3D LIDAR semantic segmentation: 3D-MiniNet.

A validation of 3D-MiniNet on the SemanticKITTI benchmark and KITTI dataset .

The proposed projection module learns a rich 2D representation through different operations. It consists of four submodules: a context feature extractor, a local feature extractor, a spatial feature extractor and the feature fusion. We provide a detailed ablation study on this module showing how each proposed components contributes to improve the final performance of 3D-MiniNet. Besides, we implemented a fast version of the point neighbor search based on a sliding-window on the spherical projection in order to compute it at an acceptable frame-rate. All the code and trained models are available online https://sites.google.com/a/unizar.es/semanticseg/.

II RELATED WORK

Current 2D semantic segmentation state-of-the-art methods are deep learning solutions . Semantic segmentation architectures are evolved from convolutional neural networks (CNNs) architectures for classification tasks, adding a decoder on top of the CNN. Fully Convolutional Neural Networks for Semantic Segmentation (FCNN) carved the path for modern semantic segmentation architectures. The authors of this work propose to upsample the learned features of classification CNNs using bilinear interpolation up to the input resolution and compute the cross-entropy loss per pixel. Another of the early approaches, SegNet , proposes a symmetric encoder-decoder structure using the unpooling operation as upsampling layer. More recent works improve these earlier segmentation architectures by adding novel operations or modules proposed initially within CNNs architectures for classification tasks. FC-DenseNet follows DenseNet work using dense modules. PSPNet uses ResNet as its encoder and introduces the Pyramid Pooling Module incorporated at the end of the CNN allowing to learn effective global contextual priors. Deeplab-v3+ is one of the top-performing architectures for segmentation. Its encoder is based on Xception , which makes use of depthwise separable convolutions and atrous (dilated) convolutions .

With respect to efficiency, ENet set up certain basis which following works, such as ERFNet , ICNet , have built upon. The main idea is to work at low resolutions, i.e., quick downsampling, and to focus the computation on the encoder having a very light decoder. MiniNetV2 uses a multi-dilation depthwise separable convolution, which efficiently learns both local and global spatial relationships. In this work, we take MiniNetV2 as our backbone and adapt it to capture information from raw LIDAR points.

II-B 3D Semantic Segmentation

There are three main groups of strategies to approach this problem: point-based methods, 3D representations and projection-based methods.

Point-based methods work directly on raw point clouds. The order-less structure of the point clouds prevents standard CNNs to work on this data. The pioneer approach and base of the following point-based works is PointNet . PointNet proposes to learn per-point features through shared MLP (multi-layer perceptron) followed by symmetrical pooling functions to be able to work on unordered data. Lots of works have been later proposed based on PointNet. Following with the point-wise MLP idea, PoinNet++ groups points in an hierarchical manner and learns from larger local regions. The authors also propose a multi-scale grouping for coping with the non-uniformity nature of the data. In contrast, other approaches propose different types of operations following the convolution idea. Hua et al. propose to bin neighboring points into kernel cells for being able to perform point-wise convolutions. Other works resort to graph networks to capture the underlying geometric structure of the point cloud. Loic et al. use a directed graph to capture the structure and context information. For this, the authors represent the point cloud as a set of interconnected superpoints.

II-B2 3D representations

There are different kinds of representations of the raw point cloud data which have been used for 3D semantic segmentation. SegCloud makes use of a volumetric or voxel representation, which is a very common way for encoding and discretizing the 3D space. This approach feeds the 3D voxels into a 3D-FCNN . Then, the authors introduce a deterministic trilinear interpolation to map the coarse voxel predictions back to the original point cloud and apply a CRF as a final step. The main drawback of this voxel representation is that 3D-FCNN has very slow execution times for real-time applications. Su et al. proposed SPLATNet, making use of another type of representation: Permutohedral Lattice representation. This approach interpolates the 3D point cloud to a permutohedral sparse lattice and then bilateral convolutional layers are applied to convolve on occupied parts of the representation. LatticeNet was later proposed improving SPLATNet proposing its DeformSlice module for re-projecting the lattice feature back to the point cloud.

II-B3 Projection-based Methods

This type of approaches rely on projections of the 3D data into a 2D space. For example, TangentConv proposes to project the neighboring points into a common tangent plane where they perform convolutions. Another type of projection-based method is the spherical representation. This strategy consists of projecting the 3D points into a spherical projection and has been widely used for LIDAR semantic segmentation. This representation is a 2D projection that allows the application of 2D images operations, which are very fast and work very well on recognition tasks. SqueezeSeg and its posterior improvement SqueezeSegV2 , based on SqueezeNet architecture , show that very efficient semantic segmentation can be done through this projection. The more recent work from Milioto et al. combines the DarkNet architecture with a GPU based post-processing method for real-time semantic segmentation.

Projection-based approaches tend to be faster than other representations, but they lose the potential of learning 3D features. LuNet is a recent work which proposes to learn local features using point-based operations before projecting into the 2D space. Our novel projection module tackles with this issue by including a context feature extractor based on point-based operations. Besides, we build a faster and more parameter-efficient architecture and a faster implementation of LuNet’s neighbor search method.

III 3d-mininet: lidar point cloud segmentation

Our novel approach for LIDAR semantic segmentation is summarized in Fig. 2. It consists of three modules: (A) fast 3D point neighbor search, (B) 3D-MiniNet, which takes PP groups of NN points and outputs the segmented point cloud and, (C) the KNN-based post-processing which refines the final segmentation.

There are two main issues that typically prevent point-based models to run at an acceptable frame-rate compared to projection-based methods: 3D point neighbor search is a required, but slow, operation and performing 3D operations is slower than using 2D convolutions. In order to alleviate these two issues, our approach includes a fast point neighbor search proxy (subsection III-A), and a module to minimize expensive point-based operations, which takes raw 3D points and outputs a 2D representation to be processed with a 2D CNN (subsection III-B1).

We perform the point neighbor search in the spherical projection space using a sliding-window approach. Similarly to a convolutional layer, we get groups of pixels, i.e., projected points, by sliding a k×kk\times k window across the image. The generated groups of points have no intersection, i.e., each point belongs only to one group. This step generates PP point groups of NN points each (N=k2N=k^{2}), where all points from the spherical projection are used (P×N=W×HP\times N=W\times H).

III-B 3D-MiniNet

3D-MiniNet consists of two modules, as represented in Fig. 3: the proposed projection module, which takes the raw point cloud and computes a 2D representation, and our efficient backbone network based on MiniNetV2 to compute the semantic segmentation.

The goal of this module is to transform raw 3D points to a 2D representation that can be used for efficient segmentation. The input of this module if the output of the point neighbor search described in the previous subsection. It is a set of PP groups, where each group contains NN points with C2C_{2} features each, gathered through the sliding-window search on the spherical projection as explained in the previous subsection.

The following three kinds of features are extracted from the input data (see left part of Fig. 3 for a visual description of this proposed module) and fused in the final module step:

The first feature is a PointNet-like local feature extraction (see projection learning module (a) of Fig. 3). It runs four linear layers shared across the groups followed by a BatchNorm and LeakyRelu . We follow PointPillars implementation of these shared linear layers using 1×11\times 1 convolutions across the tensor resulting in very efficient computation when handling lots of point groups.

The second feature extraction (projection learning module (b) of Fig. 3) learns context information from the points.This is a very important module because although context information can be learned through the posterior CNN, point-based operations learn different features than convolutions. Therefore, this module helps learning a richer representation with information than might not be learned through the CNN.

The input of this context feature extractor is the output of the second linear layer of the local feature extractor (giving the last linear layer as input would drop significantly the frame-rate due to the high number of features). This tensor is maxpooled (in order to complete the PointNet-like operation which work on unordered points) and then, our fast neighbor search is run to get point groups. In this case, three different groupings (using our point neighbor search) are performed with a 3×33\times 3 sliding window with different dilation rates of 1, 2, 3 respectively. Dilation rates, as in convolutional kernels , keep the number of grouped points low while increasing the receptive field allowing a faster context learning. We use zero-padding and a stride of 11 for keeping the same size. After every grouping we perform a linear, BatchNorm and LeakyRelu. The outputs of these two feature extractor modules are concatenated and applied a maxpool operation over the NN dimension. This maxpool operation keeps the feature with higher response along the neighbor dimension, being order-invariant with respect to the neighbor dimension. The maxpool operation also makes the learning robust to pixels with no point information (spherical projection coordinates with no point projected).

The last feature extraction operation is a convolutional layer of kernel 1×N1\times N (projection learning module (c) of Fig. 3). Convolutions can extract features of each point with respect to the neighbors when there is an underlying spatial structure which is the case, as the point groups are extracted from a 2D spherical projection. In the experiment section, we take this feature extractor as our baseline without the two others which is equivalent of performing only standard convolutions on the spherical projection.

All implementation details, such as the number of features of each layer, are specified in Sect. IV. The experiments in Sect. V show how each part of this learning module contributes to improve 3D-MiniNet’s performance.

III-B2 2D Segmentation Module (MiniNet Backbone)

Similarly to MiniNetV2, we also include a second convolutional branch to extract fine-grained information, i.e., high-resolution low-level features. The input of this second branch is the spherical projection. The number of layers and features at each layer is specified in Sect. IV-B.

III-C Post-Processing

In order to cope with the miss-predictions of non-projected 3D points, we follow Milioto et al. post-processing method. All 3D points get a new semantic label based on K Nearest Neighbors (KNN). The criteria for selecting the K nearest points is not based on the relative euclidean distances but on relative depth values. Besides, the search is narrowed down based on 2D spherical coordinate distances. Milioto et al. implementation is GPU-based and is able to run in 7ms keeping the frame-rate high.

IV Experimental setup

This section details the setup used in our experimental evaluation.

The SemanticKITTI dataset is a recent large-scale dataset that provides dense point-wise annotations for the entire KITTI Odometry Benchmark . The dataset consists of over 43000 scans from which over 21000 are available for training (sequences 00 to 10) and the rest (sequences 11 to 21) are used as test set. The dataset distinguishes 22 different semantic classes from which 19 classes are evaluated on the test set via the official online platform of the benchmark. As this is the current most relevant and largest dataset of single-scan 3D LIDAR semantic segmentation, we perform our ablation study and our more thorough evaluation on this dataset.

SqueezeSeg work provided semantic segmentation labels exported from the 3D object detection challenge of the KITTI dataset . It is a medium-size dataset split into 8057 training scans and 2791 validation scans.

IV-B Settings

We set the resolution of the spherical projection to 2048×642048\times 64 for the SemanticKITTI dataset and 512×64512\times 64 for the KITTI (same resolution than previous works to be able to make fair comparisons). We set a 4×44\times 4 window size with a stride of 44 and no zero-padding for our fast point neighbor search leading to 8192 groups of 3D points for the SemanticKITTI data and 2048 groups for the KITTI data. Our projection module is fed with these groups and generates a learned representation of resolution 512×16512\times 16 for the SemanticKITTI configuration and 128×16128\times 16 for the KITTI.

For the K Nearest Neigbors post-process method , we set as 7×77\times 7 the windows size of the neighbor search on the 2D segmentation and we set KK to 7.

We train the different 3D-MiniNet configurations for 500 epochs with batch size of 3, 6 and 8 for 3D-MiniNet, 3D-MiniNet-small, and 3D-MiniNet-tiny respectively (different due to memory constraints). We use Stochastic Gradient Descent (SGD) optimizer with an initial learning rate of 4⋅10−34\cdot 10^{-3} and a decay of 0.99 every epoch. For the optimization, we use the cross-entropy loss function, see eq. 2.

where MM is the number of labeled points and CC is the number of classes. Yc,m{Y}_{c,m} is a binary indicator (0 or 1) of point mm belonging to a certain class cc and y^c,m\hat{y}_{c,m} is the CNN predicted probability of point mm belonging to a certain class cc. This probability is calculated by applying the soft-max function to the networks’ output. To account for class imbalance, we use the median frequency class balancing, as applied in SegNet . To smooth the resulting class weights, we propose to apply a power operation, wc=(ftfc)iw_{c}=(\frac{f_{t}}{f_{c}})^{i}, with fcf_{c} being the frequency of class cc and ftf_{t} the median of all frequencies. We set ii to 0.25.

During the training, we randomly rotate and shift the whole 3D point cloud. We randomly invert the sign for X and Z values for all the point cloud. We also drop some points. The rotation angle is a Gaussian distribution with mean 0 and standard deviation (std) of 40º. The shifts we perform are Gaussian distributions with mean 0 and std of 0.35, 0.35 and 0.01 (meters) for the X, Y, Z axis (being Z the height). The percentage of dropped points is a uniform distribution between 0 and 10.

V Results

The projection module is the main novelty from our approach. This subsection shows how each part helps to improve the learned representation. For this experiment, we use 3D-MiniNet-small configuration.

Table I shows the ablation study of our proposed module, measuring the mIoU, speed and learning parameters needed with each configuration. The first row and baseline is working on the spherical projection using a convolution as the projection method, i.e., just a downsampling in that case.

As the projection used is neither rotation nor shift invariant, performing this data augmentation helps to our network generalization as first row shows. Second row shows the performance using only 1×N1\times N convolutions in the learning layers with the 5-channel input (C1C_{1}) used in RangeNet which we establish as our baseline, i.e, our spatial feature extractor. The third row shows the performance if we replace the 1×N1\times N convolution for point-based operations, i.e, our local feature extractor. These results point that MLP operations work better for 3D points but take more execution time. The fourth row combines both the convolution and local MLP operation. Combining convolutions and MLP operations increases performance due to the different type of features learned by each type of operation as explained in Sect. III-B1.

The attention module also increases the performance with almost no extra computational effort. It reduces the feature space into a specified number of features, learning which features are more important. The sixth row shows the results adding our context feature extractor. Context is also learned later through the FCNN via convolutions but here, the context feature extractor learns different context through with MLP operations. Context information is often very useful in semantic tasks, e.g., for distinguishing between a bicyclist, a cyclist and a motorcyclist. This context information gives a boost higher than the other feature extractors showing its relevance. Finally, increasing the number of features of each point with features relative to the point group (C2C_{2}) also leads to better performance without decreasing the frame-rate and without adding any learning parameter.

V-B Benchmarks results

This subsection presents quantitative and qualitative results of 3D-MiniNet and comparisons with other relevant works.

Table II compares our method with several point-based approaches (rows 1-4), 3D representation methods (row 5) and projection-based approaches (rows 6-11) measuring the mIoU, the processing speed (FPS) and the number of parameters required by each method. As we can see, point-based methods for semantic segmentation of LIDAR scans tend to be slower than projection ones without providing better performance. As LIDAR sensors such as Velodyne usually work at 5-20 FPS, only RandLA-Net and projection-based approaches are currently able to process in real time the full amount of data made available by the sensor.

Looking at the different configurations of 3D-MiniNet, it gets state-of-the-art using fewer parameters and being faster (3D-MiniNet-small-KNN) beating both RandLANet (point-based method), SPLATNet (3D representation) and RangeNet53-KNN (projection-based). Besides, 3D-MiniNet-KNN configuration is able to get even better performance although it needs more parameters than RandLANet. If efficiency can be traded off for performance, smaller versions of Mininet also obtain better performance metrics at higher frame-rates. 3D-MiniNet-tiny is able to run at 98 fps and, with only a 9%9\% drop in mIoU (46.9%46.9\% compared to the 29%29\% of SqueezeSeg version that runs at 90 fps).

The post-processing method applied shows its effectiveness improving the results the same way it improved RangeNet. This step is crucial to correctly process points that were not included in the spherical projection, as discussed in more detail in Sect. III.

The scans of the KITTI dataset have a lower resolution (64x512) as we can see in the evaluation reported in Table III. 3D-MiniNet also gets state-of-the-art performance on LIDAR semantic segmentation on this dataset. Our approach gets considerably better performance than SqueezeSeg versions (+10-20 mIoU). 3D-MiniNet also gets better performance than LuNet and DBLiDARNet which were the previous best methods on this dataset.

Note that in this case, we did not evaluate the KNN post-processing since this dataset only provides 2D labels.

The experiments show that projection-based methods are more suitable for the LIDAR semantic segmentation with a good speed-performance trade-off. Besides, better results are obtained when including point-based operations to extract both context and local information from the 3D raw points into the 2D projection.

Fig. 4 shows a few examples of 3D-MiniNet inference on test data. The supplementary video includes inference results on a full sequencehttps://www.youtube.com/watch?v=5ozNkgFQmSM. As test ground-truth is not provided for the test set (evaluation is performed externally on the online platform), we can only show visual results with no label comparison.

Note the high quality results on our method in relevant classes such as cars, as well as in challenging classes such as traffic signs. In the supplementary video we can also appreciate some of the 3D-MiniNet failure cases. As it could be expected, the biggest difficulties happen distinguishing between classes with similar geometric shapes and structures like building and fences.

VI CONCLUSIONS

In this work, we propose 3D-MiniNet, a fast and efficient approach for 3D LIDAR semantic segmentation. 3D-MiniNet projects the 3D point cloud into a 2-Dimensional space and then learns the semantic segmentation using a fully convolutional neural network. Differently from common projection-based approaches that perform a predefined projection, 3D-MiniNet learns this projection from the raw 3D points, learning both local and context information from point-based operations, showing very promising and effective results. Our ablation study shows how each part of the proposed approach contributes to the learning of the representation. We validate our approach on the SemanticKITTI and KITTI public benchmarks. 3D-MiniNet gets state-of-the-art results while being faster and more efficient than previous methods.

References