Lite-HDSeg: LiDAR Semantic Segmentation Using Lite Harmonic Dense Convolutions

Ryan Razani, Ran Cheng, Ehsan Taghavi, Liu Bingbing

I Introduction

Perception is one of the primary tasks for autonomous driving. In perception for autonomous vehicles, one main challenge can be generalized as object recognition on road. Examples of such objects include but not limited to cars, cyclists, pedestrians, traffic signs, derivable area and sidewalks etc. Due to lack of bounding box abstraction for many of these classes and the need to recognize all these objects in one unified end-to-end approach, a semantic segmentation approach applies, which attempts to predict the class label or tag, for each pixel of an image or point of a point cloud.

To achieve this task, many sensors can be leveraged in robotic systems, such as cameras (mono, stereo), Light Detection And Ranging (LiDAR), ultrasonic, and Radio Detection And Ranging (RADAR) sensors. Among these sensors, LiDAR becomes the sole choice of sensor for perception tasks such as 33D semantic segmentation because of its active sensing nature with high resolution of 33D sensor readings.

During the past few years, and with availability of recent new datasets such as SemanticKITTI , many new methods have been proposed to address the problem of LiDAR semantic segmentation. Although different in design and solution proposal, many of the methods can be categorized into these four groups: point-based methods such as , image-based methods such as , voxel-based methods such as and graph-based methods such as . Section II is dedicated to covering different methods in literature on the subject of LiDAR semantic segmentation.

Overall, point-based methods can achieve better accuracy at the cost of high computational complexity and memory consumption. A more reasonable and realistic choice is image-based or voxel-based methods where raw and unordered point cloud is projected onto a structured plane. Lite-HDSeg is an end-to-end LiDAR semantic segmentation method that is solely based on range-image (a transformation of 33D LiDAR point clouds into an image using spherical projection). Thereby, it benefits from minimal prepossessing and efficient implementation using standard CNNs. The contributions of this paper are summarized as follows:

A novel residual Inception-like Context Module (ICM) to capture global and local context information in full 360360 degrees LiDAR scan. Incorporating ICM at the beginning of the neural network improves learning local and global features for the task of semantic segmentation as it is shown in Section IV-C;

A new and improved encoder-decoder CNN model on a new encoder named Lite-HDSeg, a residual network decoder, a multi-class SPN and a modified boundary loss which is trained end-to-end using the spherical data representation (range-image). Lite-HDSeg surpasses state-of-the-art methods in accuracy vs. runtime trade-off on a large-scale public benchmark, SemanticKITTI (shown in Fig. 1);

A thorough analysis on the LiDAR semantic segmentation performance based on different ablation studies, and qualitative and quantitative results.

The rest of the paper is organized as follows. In Section II, a brief history of recent and related works is given. Section III presents the proposed Lite-HDSeg network with detailed description of its components and novel techniques used to design the model. Section IV demonstrates a comprehensive experimental results and thorough qualitative and quantitative evaluation on the public benchmark, SemanticKITTI . Finally, some conclusions are drawn in Section V.

II Related Work

Point-based methods. A natural approach to extract features from point cloud data is to treat it in its original form, i.e., unordered point cloud, directly process it using various approaches such as a neural network. In this way, all data points will be processed and a label is predicted for each input data point. The pioneering methods of this group are PointNet and PointNet++ which use shared Multi-Layer Perceptrons (Shared MLPs) to extract features for different tasks such as semantic segmentation.

Much research focuses to efficiently implement point-based methods without sacrificing performance . However, point-based approaches remain successful in tasks with only small input point clouds . KPConv develops a flexible way to use any number of kernel points and thus can be extended to deformable convolutions that can use these flexible kernel points to learn local geometry. More recently, KPRNet has combined ResNext and to achieve better results, however, the computational complexity of the model makes it less attractive for autonomous driving/robotic applications.

Image-based methods. To benefit from 22D convolutional neural networks such as , one can project 33D LiDAR data onto a 22D surface using spherical projection . Using spherical projection turns the point cloud segmentation task into image segmentation problem, in which a prediction can be made for each of the projected data points.

Image-based methods, in specific those which use spherical projection, perform better in terms of segmentation accuracy vs computational complexity trade-off. This is due to leveraging 22D convolutional neural networks (CNNs), such as residual blocks , depth-separable convolutions , and dilated convolutions . Therefore, spherical projection methods have gained attention in recent years and been used in many methods such as SqueezeSeg series , RangeNet++ , DeepTemporalSeg , SalsaNet , SalsaNext , etc. SalsaNext inherits the encoder-decoder architecture from SalsaNet , which contains a series of residual blocks in the encoder and upsamples features in the decoder part. An extra layer, added before the encoder, brings in more global context information, together with other techniques such as average pooling and dropouts. A new loss term, Lovász-Softmax loss , is also used in SalsaNext which can optimize the mean Intersection over Union (IoU) metric directly. With all the above-mentioned changes, SalsaNext achieves a higher score on SemanticKITTI leaderboard than other publicly available methods in this category.

Voxel-based methods. Voxel-based methods use 33D data points and project them into cubical voxels of known size. Each voxel, which may contain more than one data point, carries different information, such as maximum height, mean, variance, and other information related to the data points associated with each voxel. Voxel-based methods usually take voxels as input and predict each voxel with one semantic label. SEGCloud is one of the early data driven methods to use voxel representation for point cloud semantic segmentation. ScanComplete takes in a 33D scan of a scene and outputs a complete 33D model with a semantic label for every voxel with a specially designed 33D CNN model, whose filter kernels are invariant to the overall scene size. More recently, the Minkowski CNN has been proposed which can efficiently handle 33D sparse convolutions and other standard neural network layers and thus reduces the computational costs of 33D CNNs significantly. A neural architecture search (NAS) model based on also introduced recently, i.e. SPVNAS , that can achieve state-of-the-art accuracy at the cost of high computational cost. Although efficient 33D CNNs are available, voxel-based methods suffer from loosing granular data and are prone to detecting smaller objects on the scene, specially with larger voxels.

Graph-based methods. A graph representation can be reconstructed from point cloud data, where a vertex represents a point or a group of points, and edges represent adjacency relationship between vertices. Using graph representation, authors in introduced a neural network model that benefits from geometrically homogeneous elements from the scanned scene using a structure called superpoint graph (SPG). More recently, GACNet proposed a graph convolution with new kernels that can be shaped dynamically to adapt to the structure of an object. In , authors designed a neural network by proposing a technique to extract features, benefiting from a hierarchical graph framework to extract point/edge level features. Although theoretically sound, graph-based methods are not suitable for real-time robotic applications due to their complex nature, and more research needs to be done to make them a practical choice.

III Proposed Method

where WW and HH are the width and height of the range-image. The vertical field of view of the sensor is f=fup+fdownf=f_{up}+f_{down}. Symbols uu and vv are the horizontal and vertical coordinates of the corresponding pixel in the bitmap image, respectively.

Using (1), one can generate an input of arbitrary size WW and HH with channels (x,y,z,rem,r)(x,y,z,rem,r), where (x,y,z)(x,y,z) are the Cartesian coordinates of a data point, remrem is the remission or intensity reading from the sensor and rr is the range, respectively. In the following subsections, a novel neural networks architecture and its design is explained to address this problem along with experimental results and ablation studies.

The proposed network structure is based on a) multi-scale convolutional learning module, Inception-like Context Module (ICM), to extract multi-scale contextual features from range-image to be processed by the encoder-decoder network, b) lite version of HarDNet as the encoder with less dense connections to achieve better performance (see Fig. 4(b)), c) CAM module in the encoder to collect nearby context from neighborhood feature maps, d) Multi-class Convolutional Spatial Propagation Network (MCSPN) before the last layer of convolutions in the encoder-decoder CNN to refine the masks of each class for the segmentation task and e) a dedicated boundary loss in addition to cross entropy and Lovász loss loss for the segmentation task of LiDAR point cloud data. The combination of all above-mentioned contributions demonstrate that segmentation results can substantially be improved over the range map with a large margin.

The general block diagram of the proposed model is shown in Fig. 2 and the details of the network architecture are described in subsections below.

We present a multi-scale convolutional learning module, Inception-like Context Module (ICM), to extract multi-scale contextual features from range-image to be processed by the encoder-decoder network targeted at semantic segmentation. The proposed context feature extractor consists of concatenation of several multi-scale convolution layers with residual connections and larger receptive fields to extract rich global information which is essential in learning complex correlations between classes with different size. This yields to gather the spatial fine-grained information along with global context information. A detailed visual description with numerical values for channels, kernels and dilation ratios can be found in Fig. 3(a).

III-B Lite-HDSeg Encoder

Using shortcuts or skip connections proved to be an effective design choice in tasks such as semantic segmentation. Shortcuts enable implicit supervision to make networks deepen without degradation. DenseNets proposed a shortcut scheme where all preceding layers are concatenated together. Authors in showed that this novel shortcuts scheme improves the deep supervision and achieves good results in image segmentation task. log-DeseNets and SparseNet both attempt to sparsify the DenseNets by sparsely connecting layers together. Although smaller set of shortcuts can achieve fast convergence speed, they both need to increase the growth rate (output channel width) to recover from connection pruning. One of the most successful design in this domain is Harmonic Dense Net (HDNet) . HarDNet proposed shortcuts following power-of-two-th harmonic waves achieving a significant reduction in cost, while keeping an acceptable performance of the model. As shown in Fig. 4(a), layer kk connects to layer k−2nk-2^{n} if 2n2^{n} divides kk, where nn is a non-negative integer, k−2n⩾0k-2^{n}\geqslant 0 and the multiplier mm serves as a low-dimensional compression factor. Further details can be found in §3 of . Overall, compared to the dense blocks, which connects all the layers iteratively in a block, HD block (Harmonic Dense block) has much deeper structure yet, uses less skip connections (shortcuts) according to the harmonic-pattern to propagate the gradients. HD block was inspired from the log-DenseNet which cut the network complexity yet, kept the effective skip connectivity that ensured the propagation of gradient.

Our proposed Lite-HDSeg uses a series of simplified HD blocks for the encoder by removing some of the skip connections, while keeping the critical connections (Enc11–Enc55). This is due to the fact that the dense skip connections introduce many redundant and repeated channels. This systematic pruning in the skip connections reduces the computational complexity and improves memory efficiency, which are the cornerstone of real-time applications. In the proposed Light HD Block (LHD), each convolution layer fif_{i} takes a direct input from at most ⌊log(i)5⌋\lfloor\frac{log(i)}{5}\rfloor number of previous layers, and these input layers are apart from layers with base 55, i.e.:

where ⌊⋅⌋\lfloor\cdot\rfloor is the floor function, ∣| is or operation, and ⌊⋅⌉\lfloor\cdot\rceil is the nearest integer function. For example, the feature maps of layer ii is concatenated from layers

Finally, the overall complexity of a Lite HDBlock is

which is pruned significantly compared to the DenseNet and HarDNet . Our experimental results show that pruning redundant connections in HD block not only increases the model runtime efficiency, but also improves the model performance substantially.

III-C Context Aggregation Module

Context Aggregation Module (CAM) is adopted from the SqueezeSegV2 and is used after each LHD block as an attention map to emphasize the feature maps locally. As stated in , the noise that is in the nature of LiDAR point cloud (missing points in the range-image, sensor jitters and mirror reflections) can cause a CNN model to perform poorly. Since the CAM module collects the nearby context from neighborhood feature maps, these characteristics of the sensor can be compensated at the time of training, generating features that are robust against LiDAR specific noise.

III-D Decoder

The proposed decoder is based on residual blocks (Res-block). The proposed residual block is a combination of conv-blocks in a specific order and increased kernel size as shown in Fig. 3(b) with a skip connection from the encoder with the same spatial size. The Resblock in Fig. 3(b) along with an up-sampling layer is used four times (Dec4, Dec3, Dec2 and Dec1) to generate an output with the spatial size of the input range-image. The final layer in the decoder is a convolutional layer to generate output with the depth size C13C_{13}.

III-E Mutli-class Spatial Propagation Network

To improve the performance of segmentation result, we adapt convolutional spatial propagation network and extend this framework to support multiple classes. In convolutional propagation network, a convolutional neural network is trained to learn explicit affinity between each pixel and its neighbors (in 22D image). Then, the segmentation mask is refined according to the learned affinity. Unlike dense CRF in which each pixel is connected with others, the proposed method only contains three way connection, as summarized in the following equation,

where hh is the propagated hidden feature map, kk is current pixel propagation position, N\mathit{N} is the neighborhood pixel for one direction, c∈Cc\in\mathcal{C} is the semantic class, and tt is the propagation time step. Parameter xk,t,cx_{k,t,c} is the original output feature map from the decoder, and pk,t,cp_{k,t,c} is the propagated affinity map of one direction (for example, left to right). The final output feature map is then fused from linearly propagated feature map from last time step through four directions. This propagation refines the semantic of small objects and boosts the overall performance. Finally, the output of the network is passed through a final convolutional layer followed by a softmaxsoftmax function, where the final predictions are being made for each pixel and loss is calculated.

III-F Loss Function

To address the specific data type used in this work, the labeled data, and its lack of a balanced ground truth labels for all classes, four different loss terms are used, i.e., weighted cross entropy, Lovász loss, boundary loss, and weight decay regularization, which are parameterized as Lwce{L}_{wce}, LlsL_{ls}, LbdL_{bd} and LregL_{reg}, respectively. The weighted cross entropy (WCE) loss is suitable where a neural network model deals with multiclass classification problem much like semantic segmentation. The second loss term that is used in this work is Lovász loss . As shown in , the Lovász loss is an effective additional loss term that can be used for different machine learning tasks such as object detection and segmentation.

As stated in , relying on pixel-level losses alone such as cross-entropy for predicting curvilinear structures is insufficient. In specific to semantic segmentation, there exist complex boundaries between different classes of various shapes which makes cross-entropy and other advanced losses such as Lovász loss sub-optimal for the task. To solve this problem, proposed to add an extra term to the conventional loss terms, manually extracting boundaries and forcing the network to focus on those regions more than other parts. Therefore, we use boundary loss to account for the error in the boundary of different classes. The loss term Lbd{L}_{bd}, used for the task of LiDAR semantic segmentation, can be calculated as,

where PcP^{c} and RcR^{c} are the precision and recall of the predicted boundary map ypredby^{b}_{pred} with respect to the ground truth boundary map ygtby^{b}_{gt} for class cc, respectively. yy represents the ground truth and y^\hat{y} is the model output for each class. The boundary map is thus calculated as,

where pool(⋅,⋅)pool(\cdot,\cdot) is a pixel-wise max-pooling operation for the binary map within a sliding window, i.e. θ\theta, to extract the vicious boundary.

Finally, the total loss is computed by combining the previous loss terms with a weight decay regularizer and is defined as,

where α\alpha, β\beta, γ\gamma, and λ\lambda are the weights associated to weighted cross entropy, Lovász-Softmax, boundary loss and regularization term, respectively.

III-G Post-processing

To project the prediction labels from range-images back to 33D representation, a KNN post-processing step is employed. Every point is given a new label based on its KK closest points. However, instead of finding the nearest neighbours in the complete unordered point cloud, a sliding window approach is used to sub-sample the point cloud .

IV Experimental Results

To evaluate the performance of the proposed method and to provide experimental results, including ablation studies, a large-scale SemanticKITTI dataset is used. SemanticKITTI is the largest available dataset with fully labeled LiDAR point clouds of various scenes. Due to this, many of the recent works such as evaluated their methods based on SemanticKITTI, making it a standard dataset for evaluating LiDAR semantic segmentation methods. Statistical analysis and more information about SemanticKITTI can be found in supplementary documents.

Mean Intersection over Union (mIoU\mathbf{mIoU}) is used as our evaluation metric to evaluate and compare our methods with others. It is the most popular metric for evaluating semantic point cloud segmentation and can be formalized as

where TPcTP_{c} is the number of true positive points for class cc, FPcFP_{c} is the number of false positives, and FNcFN_{c} is the number of false negatives. As the name suggests, the IoUs for each class are calculated and then the mean is taken.

We trained our network for 100100 epochs using stochastic gradient descent with the initial learning rate of 0.010.01 which was decayed by 0.010.01 after each epoch. The batch size was set to 66 and the dropout probability to 0.20.2. The number of channels [C0→C13C_{0}\rightarrow C_{13}] used in Lite-HDSeg architecture were assigned as 55, 2424, 3232, 4848, 9696, 192192, 320320, 720720, 12801280, 17441744, 842842, 448448, 192192, 6464, respectively. The height and width of the projected image were set to H=64H=64, W=2048W=2048. The values of 1.01.0, 1.51.5, 1.01.0, 1.01.0 were used for the weights of the total loss, α\alpha, β\beta, γ\gamma, λ\lambda, respectively. In the post-processing stage, the window size of the neighbor search was set to 55, with k=5k=5, σ=1\sigma=1, and a cutoff of 1m1m. Moreover, we applied augmentation, such as random rotation, transformation, flipping around the yy-axis with probability of 0.5, to the data to avoid overfitting.

IV-B Results

The numerical comparison of the described techniques is illustrated in Table I for recent available and published methods of projection-based and point-wise categories. For each class, the best achieved accuracy in terms of IoU is highlighted among different methods. As shown in Table I, the proposed method achieves the state-of-the-art performance on SemanticKITTI test set. Lite-HDSeg outperforms existing methods by 4.34.3% (mIoU) and most of the classes in terms of IoU within its category, while being real-time. It is worth noting that Lite-HDSeg not only outperforms all projection-based methods, but also outperforms real-time implementations of SPVNAS with 63.763.7 and 60.360.3 mIoU accuracy (see Fig. 1). To better visualize the improvements made by our proposed model, a sample of qualitative result is shown in Fig. 5. We compare SalsaNext vs. Lite-HDSeg in terms of the generated error map in the same data frame in Fig. 5, with Lite-HDSeg showing a noticeable improvement against SalsaNext.

IV-C Ablation Study

In order to understand the effectiveness of different proposed techniques, a thorough ablation study has been done and the results are shown in Table II for Lite-HDSeg. The mIoU\mathbf{mIoU} scores of all ablated networks are based on our training and the reported numbers are the mIoU\mathbf{mIoU} for sequence 88 on SemanticKITTI for all the experiments.

The ablation study starts with the baseline model, Salsanext , which is the closest method proposed in terms of performance to Lite-HDSeg. Next, we exchange the backbone with the proposed encoder-decoder in Section III. Afterwards, we add the ICM, CAM, boundary loss and MCSPN one by one, to show the effectiveness of each block. As it is shown in Table II, all the contributions stated in the table effectively improve our proposed encoder-decoder with the biggest jump of 2.3%2.3\% when adding MCSPN. It is worth noting that, although extra modules are proposed in the design of Lite-HDSeg, because of the novel reduction scheme in the complexity of the harmonic dense block (lite HD block), the overall computational complexity of Lite-HDSeg remains at a competitive range.

V Conclusions

In this paper, we presented Lite-HDSeg, a novel real-time CNN model for semantic segmentation of a full 33D LiDAR point cloud. Lite-HDSeg has a new encoder-decoder architecture based on light-weight harmonic dense convolutions and residual blocks with a novel contextual module called ICM and Multi-class SPN. We trained our network with boundary loss to emphasize the semantic boundaries. To show the performance of the proposed method, Lite-HDSeg was evaluated on the public benchmark, SemanticKITTI. Experiments show that the proposed method outperforms all real-time state-of-the-art semantic segmentation approaches in terms of accuracy (mIoU\mathbf{mIoU}), while being less complex. More specifically, Lite-HDSeg introduces a novel design with accuracy improvements from each of the blocks, i.e., the new context module (ICM), the new encoder-decoder, MCSPN and boundary loss, to benefit not only semantic segmentation task, but also instance segmentation and object detection.

References