TORNADO-Net: mulTiview tOtal vaRiatioN semAntic segmentation with Diamond inceptiOn module

Martin Gerdzhev, Ryan Razani, Ehsan Taghavi, Bingbing Liu

Introduction

Semantic segmentation of point clouds, specially on data collected from LiDARs is becoming a crucial task in many applications such as robotics, autonomous driving, etc. Semantic segmentation is a key component of a larger scope of tasks called scene understanding. The role of scene understanding is even more pronounced for autonomous systems such as self-driving cars, where safety is paramount and errors in the perception stack can lead to bad planning and accidents.

While semantic segmentation has been used on images in other domains, such as medical imaging and surveillance systems, it has seen limited applications in self-driving due to the difficulty of data annotation of the different sensors. An autonomous vehicle is equipped with many primary sensors, including cameras, LiDARs, RADARs and possibly other secondary sensors such as sonars, which help the task of perception for an autonomous system. The earliest works on data labeling in robotics and autonomous driving have been done on cameras and images due to their wide availability and somewhat easy labeling procedure.

To enable 33D data collected from LiDARs to be fully utilized in perception modules, one needs a labeled point cloud at the point level (semantic) to be available. To this date, SemanticKITTI proved to be one of the best LiDAR datasets with point wise labels available to researchers and industries. Due to this reason, the focus of this work is mainly on literature available on LiDAR semantic segmentation using SemanticKITTI as a benchmark dataset. It is worth noting that, due to wider field of view, precise distance measurements and light-invariance, most autonomous car platforms use LiDAR sensors for scene understanding on the road.

LiDAR point clouds present a number of challenges as compared to images - they are unstructured, sparse and their density varies with distance. Some methods try to tackle the problem by operating on the point clouds directly , while others try to utilize the approaches from the image domain, by projecting the point clouds onto images (Bird Eye View, or Frontal projection) or 3D voxels, and applying convolutions on the structured representation .

This work builds upon several approaches like multiple views projections similar to MVF and encoder-decoder networks like SalsaNext to propose an end-to-end model that achieves state of the art results on SemanticKITTI . With TORNADO-Net we introduce the following contributions:

A pillar based learning module that learns and extracts features on BEV data representation of LiDAR point clouds;

A novel global context module, named Diamond feature extractor, that processes the spherical range-image LiDAR data and extracts rich features suitable for semantic segmentation;

A novel loss function based on total variation denoising techniques that improves the overall accuracy of point cloud semantic segmentation models;

Circular padding to account for the LiDAR data with 360°360\degree horizontal field of view;

An analysis on the semantic segmentation performance using different architectures through an extensive ablation study;

The rest of the paper is organized as follows. In Section 2, a brief review of recent and related works is given. Section 3 describes the proposed neural network model and the new loss functions in detail. Training details and experimental results including qualitative, quantitative, and ablation studies can be found in Section 4, along with the benchmarks from all available and published methods in the literature for reference. Finally, conclusions and future work are presented in Section 5.

Related Work

Until recently, there have been few approaches that focus on LiDAR point cloud based semantic segmentation due to the lack of large-scale labeled datasets. Some of these early methods include PointNet , SqueezeSeg and DeepTemporalSeg . The introduction of the SemanticKITTI dataset has spurred the development of novel LiDAR-based segmentation methods .

Based on how the point cloud is represented, methods can be grouped into methods that operate on points directly (point-wise), or methods that project the point cloud into a different, easier to work with structure (projection-based). Point-wise methods like PointNet , KPConv and RandLA-Net don’t require any preprocessing or transformation. Although capable of processing smaller point clouds, they are less useful for large point clouds (specially those which collect 360°360\degree data) due to large memory requirements and slow inference speeds. To address some of these problems, some methods like SPG utilize a superpoint graph, which is formed by geometrically consistent elements. Some of the more prevailing methods, however, aim to project the pointcloud into either a 3D voxel grid, or a 2D image (Bird-Eye-View (BEV) or spherical Range View (RV)), and utilize convolutional operators that work on structured data . The projection-based methods usually have achieved higher accuracy, while maintaining a much faster inference time. Therefore, TORNADONet aims at solving the problem of LiDAR semantic segmentation using projection-based techniques that combines multiple projections - first to BEV, and then to RV similar to to extract complementary features and achieve state-of-the-art results.

In addition to the neural network model that is designed to address a specific problem, the choice of a loss function can play a crucial role to the accuracy of the model. For semantic segmentation tasks, cross entropy (CE) loss is one of the most widely used loss functions. However, since there can be a big class imbalance with many classes being over-represented in the data, the loss functions that take the class frequency into account, such as weighted cross entropy (WCE) or Focal loss , can improve the performance. While these losses work well in a variety of tasks, they do not optimize the same criterion that is used to evaluate the performance of semantic segmentation, i.e., intersection-over-union or Jaccard Index (IoU).

Since IoU is a discrete function and cannot be optimized for directly, surrogate functions that are differentiable have been proposed. Lovász-Softmax loss and Dice loss , optimize the Jaccard and Dice coefficients, respectively in order to maximize the mean IoU(mIoU).

These losses however do not take into account the local neighbourhoods of the pixels/points and can lead to noisy predictions. To address this problem, regularization functions, such as, the Total Variation regularizer, have been used to regularize the overall loss by the smoothness of the prediction over neighboring pixels or data points. More recently, the TV regularizer was used to address the denoising problem in image applications . A major target in image denoising is to preserve important image features, such as edges, while removing noise. In our work, we use this concept and propose a novel TV loss function that takes into account neighbouring information.

To achieve better results, it is also common to use a weighted combination of multiple loss functions. For example, SalsaNext uses a combination of WCE loss and Lovász-Softmax loss. In this paper, a combination of WCE loss, Lovász-Softmax loss and the novel TV loss is proposed to help TORNADO-Net achieve state-of-the-art accuracy in LiDAR semantic segmentation.

TORNADONet

In this paper, a novel and intuitive NN architecture is introduced to solve the problem of LiDAR semantic segmentation benefiting from information extraction in different views. Although the encoder-decoder model processes range image similar to , the pillar-projection-learning module (PPL) learns and extracts information in the Bird’s Eye View (BEV). This series of different projections as shown in Fig. 1, help the model pick up features that are otherwise difficult to extract. Results and ablation studies on the SemanticKITTI dataset benchmark show the effectiveness of the proposed CNN model. The details of the new architecture are explained in the subsections below.

The proposed model starts off with the PPL module (see Figure 2) in which the raw LiDAR point cloud PP is processed by mapping the unordered points to BEV. We follow the approach similar to where points are grouped into pillars and their normalized x,y,zx,y,z pillar coordinates are appended to the point cloud features. The point cloud goes through a FC layer to extract better features and is then mapped to the BEV pillars. All points within a pillar are processed through another FC layer and then after a pooling operation each pillar is represented by a single feature vector. The projected pillars then undergo a series of strided convolutions followed by upsampling layers in order to capture neighbouring information. All points are then augmented with the features from the pillar that they belong to.

The PPL thus generates rich point features that can be used for different applications. For the purpose of LiDAR semantic segmentation, range images have shown better and more accurate results with reasonable computational complexity. Because of this reason, in the proposed model and after applying PPL, the features are projected onto the range-image for further processing.

where (u,v)(u,v) are the coordinates of a given point cloud (x,y,z)(x,y,z) in spherically transformed data, hereafter range-image, WW is the desired horizontal resolution and HH is the desired vertical resolution. Moreover, the vertical field of view of the sensor can be described as f=fup+fdownf=f_{up}+f_{down}. In its simplest form, the channels of the new range-image can be filled with (x,y,z,rem,r)(x,y,z,rem,r), where remrem is remission reading of points and rr is the range. If desired, other channels can be added to the range-image to address the much needed features for a specific task. As the point cloud is already processed using a PPL block, in this NN model, we use extracted features CDC_{D} and project them onto the range-image. Empty pixels in the transformed data can be masked.

1.2 Diamond contextual block

Here we propose a new global context module to process the output of PPL, named diamond context block (DCB). DCB uses regular 2DD convolutions and provides context at different scales with efficient computation that can be used in various CNN models without loss. A generic block diagram of DCB is shown in Figure 3. DCB consists of three diamond shape convolution blocks. Each block is a combination of 3×33\times 3, 5×55\times 5 and 7×77\times 7 22D convolutions. Moreover, to carry the local features, a skip connection with 1×11\times 1 22D convolution connects the input to the output of the second diamond block. Finally, the third diamond block is concatenated with the skip connection of its input to generate the final feature tensor for further processing. Our ablation studies show that DCB enhances semantic segmentation on SemanticKITTI.

1.3 Encoder-Decoder

After the raw point cloud is processed by PPL and DCB, the feature tensor is fed to an encoder-decoder CNN similar to what is proposed in . The architecture of the proposed encoder-decoder is illustrated in Figure 1. The input to the network is the spherical projection of the extracted features from PPL+DCB in section 3.1.2. The encoder-decoder part of TORNADONet is built upon the base SalsaNet model which follows the standard encoder-decoder architecture with a bottleneck compression rate of 1616. As opposed to the original implementation of SalsaNet with series of ResNet blocks , we use the blocks introduced in with dilations in the convolutions both on the decoder and encoder parts of the network.

2 Loss function

One of the major contributions of this paper comes in the design of a new loss function using ideas introduced in , and . In this section, we first briefly review the common loss functions used in semantic segmentation tasks. Then the proposed loss function is introduced.

The weighted cross entropy loss can be written as,

where νi\nu_{i} is the frequency of each class, and P(y^i)P(\hat{y}_{i}) and P(yi)P(y_{i}) are the corresponding predicted and ground truth probability. Note that P(yi)P(y_{i}) acts as indicator for the correct label class.

This loss is suitable where a NN model deals with multiclass classification problem much like semantic segmentation. Moreover, because it minimizes the distance between the two probability distributions, namely, predicted and actual, WCE makes sure that the difference between P(y^i)P(\hat{y}_{i}) and P(yi)P(y_{i}) is being minimized.

The Lovász-Softmax loss can be expressed as:

where JJ is the Lovász extension of IoU, e(c)\boldsymbol{e}(c) is the vector of errors for class cc, e(c)∈p\boldsymbol{e}(c)\in^{p}, and pp is the number of pixels considered. JJ denotes a piece-wise linear function, with a global minimum. It has been shown in various works such , that Lovász loss is an effective additional loss term that can be used for different machine learning tasks such as object detection and segmentation. Hence, in the process of training the proposed model, Lovász loss will be combined with other losses to achieve better overall accuracy of the trained model.

Total variation (TV) regularization has been used in the literature to address the denoising problem in images . This technique forms a significant preliminary step in many computer vision tasks, such as object detection and object recognition (i.e. object localization and classification). A major concern in designing image denoising models is to preserve important image features, such as edges, while removing noise from a digital image.

The same problem can be recognized in machine learning when one tries to identify pixel or point level classification for images or point cloud data points. Although the task is not directly denoising, classifying a pixel in an image or a data point in a point cloud relies heavily on the information provided by the neighboring pixels or data points. Hence, designing a loss function based on TV regularization seems viable. In , a total variation regularizer was introduced, which regularizes the overall loss by the smoothness of the prediction over neighboring pixels or data points, resulting in a better optimization for any machine learning task such as image or point cloud semantic segmentation.

Although the technique introduced in and is mathematically sound and effective, it is only a regularization term. In this implementation, the TV regularizer is modified to a new loss function. Moreover, in combination with a standard CNN for semantic segmentation its effectiveness is shown for the task of LiDAR semantic segmentation.

Given an image-like prediction and label for the ground truth, one can write the TV loss as follows:

where i,ji,j are indexes for the pixel location, Δi\Delta i and Δj\Delta j are the step sizes in row and column directions, respectively. The ground truth label is denoted by yy and y^\hat{y} is the prediction of the network. In this expression, the quantities Y(Δi),(j)Y_{(\Delta i),(j)} and Y(i),(Δj)Y_{(i),(\Delta j)} represent the difference in pixel values of the current pixel with its vertical and horizontal neighbours, respectively. They are computed as,

For the task of LiDAR semantic segmentation on range-image, equation 4 can be simplified to the following for L1,1L_{1,1} norm and one sided neighbours, i.e., Δi,Δj=1,p,q=1\Delta i,\Delta j=1,p,q=1:

The total loss that is used to train the proposed model is a combination of Lovász, WCE, and TV losses introduced above. In order to balance the effect of each loss in the training, different weights are introduced for each loss term that can be accomodated as hyperparameters. The final loss can be formulated as

where βls\beta_{ls}, βwce\beta_{wce} and βtv\beta_{tv} are weights for Lovász loss, WCE loss, and TV loss, respectively.

3 Training Details

Before feeding the raw pointclouds PP into the model, they are truncated to the range , , in the xx, yy, and zz directions respectively. This is followed by a series of augmentations. We adopt an augmentation scheme similar to , where we do random point dropping, global pointcloud rotation and translation, and flipping along the xx-axis. We drop up to 20% of the points. For the rotations, we randomly sample angles in the range ,,, and $degreesfordegrees forx,,yandandzaxesrespectively.Similarly,forthetranslations,wesampleintherange<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><moseparator="true">,</mo></mrow><annotationencoding="application/x−tex">,</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.3em;vertical−align:−0.1944em;"></span><spanclass="mpunct">,</span></span></span></span></span>,axes respectively. Similarly, for the translations, we sample in the range <span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo separator="true">,</mo></mrow><annotation encoding="application/x-tex">,</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.3em;vertical-align:-0.1944em;"></span><span class="mpunct">,</span></span></span></span></span>,forforx,,y,and, andz$ directions. Each augmentation is independently applied with a probability of 0.5.

The model was trained using the Adam optimizer with a one-cycle learning rate scheduler for 50 epochs. The maximum learning rate was set to 0.004, division factor of 10 and cosine annealing phase split of 0.3. A weight decay of 10−410^{-4} was also used. The models were trained on 4 Tesla V100 with per GPU batch sizes of 4 for the low-res and 3 for the high-res models respectively.

The voxel size of the PPL block was set to [0.3125,0.3125,10][0.3125,0.3125,10] leading to a voxel grid of [512×512][512\times 512]. We used C=64C=64, CP=7C_{P}=7 , CD=192C_{D}=192 for filter sizes, and the height and width of the projected image were set to H=64H=64 , W=2048W=2048, except for the high-res model where H=128H=128.

The weights for the losses were set to βls=1.5\beta_{ls}=1.5, βwce=1.0\beta_{wce}=1.0, and βtv=7.5\beta_{tv}=7.5.

In the post-processing stage, the KNN used a kernel size of 5 for the low-res model and kernel size of 11 for the hi-res model, K=5K=5, σ=1\sigma=1, and a cutoff of 1m for all models.

Experimental Results

We use the SemanticKITTI dataset to evaluate the performance of our model and compare it with the state-of-the-art methods. It provides over 43K43K frames with dense point-wise annotations for the entire KITTI Odometry Benchmark across 2222 sequences. The dataset consists of 22 distinct semantic classes and is divided into two sets. The first set, referred to as the training and validation set, includes the sequences of (0000 - 1010), while the second set referred to as the test set, includes the sequences of (1111 - 2121). The validation set is sequence 0808 upon which the ablation studies are based upon. The point-wise labels for the first set are publicly available, however, the labels for the second set are not provided and are kept hidden for competition purposes.

In order to evaluate the results of the trained methods, the mIoU\mathbf{mIoU} metric is used. mIoU\mathbf{mIoU} is the most popular metric for evaluating semantic point cloud segmentation. It can be formalized as,

where TPcTP_{c} is the number of true positive points for class cc, FPcFP_{c} is the number of false positives, and FNcFN_{c} is the number of false negatives.

We compare the numerical experiments of our work with the existing methods in Table 1. It demonstrates the class wise IoU, Frames Per Second (FPS), and mean IoU for different approaches. We categorized the methods into two classes of point-wise and projection-based methods. In each category, the best IoU per class is selected. As shown, the proposed method achieves the state-of-the-art result, outperforming all the previous methods in mIoU and almost all the classes in its category. While the proposed method achieves high accuracy it is also faster than most point-based methods, making it applicable to the real-time systems.

2 Qualitative Analysis

In order to better understand the results produced by the proposed method, TORNADO-Net, in this section, four different samples from the SemanticKITTI validation set are provided in Figure 4. The results are presented along with ground truth labels for each case presenting the quality of the results in different situations where different objects/scene are available. Case 11 shows how well TORNADO-Net can perform where there exist many dynamic objects such as cars (parked or otherwise). Cases 22 and 33 depicts the successful prediction for road, vegetation, pole and other structural objects. As it is shown, TORNADO-Net can handle most of the cases well and in many cases, distinguishing between ground truth and prediction is hard.

However, there are situations for which TORNADO-Net fails to predict the correct class in a scene. Case 44 presents one of the more common failures seen in our analysis. In this case, the truck on the middle-right part of the scene is predicted partially as a car and partially as a truck. This is mostly due to limited number of training samples and a fair amount of shared features between different types of vehicles such as overall structure, intensity reflection, etc. Nevertheless, the authors think these issues should be addressed even with limited labeled data and is part of the future research on this topic. For more thorough analysis of the results a video of the predictions on sequence 88 will be available publicly as supplementary material.

3 Ablations

As shown in Table 2, each of our contributions amounts to an improvement in the overall mIoU. We receive the biggest increases with the introduction of the TV Loss. Circular padding mostly had a positive effect (about 0.5%). The PPL block is also a key component introducing a 1.5% jump at the cost of extra parameters and reduction in speed. Using a hi-res version of the same model leads to a further improvement in the mIoU at the cost of speed. KNN post-processing is also a key component leading to a 2-3% jump. The high-res model benefits from a higher KNN search value due to the sparser projection of points onto the image.

Conclusion

Semantic segmentation has been a subject of interest in many fields, such as autonomous driving. While other deep learning techniques are promising on the LiDAR semantic segmentation task, they are complex systems to implement. However, other methods that are real-time systems lack accurate performance. The main objective of this work was to propose a novel deep neural network, TORNADO-Net, for 3D LiDAR point cloud semantic segmentation. We leveraged a multi-view (bird-eye and range) projection feature extraction with an encoder-decoder ResNet architecture with a novel diamond context block. Moreover, TV loss was introduced along with Lovász-Softmax, and WCE loss to efficiently train the network. We evaluated the proposed method on the SemanticKITTI benchmark and were able to achieve state-of-the-art results on its published leaderboard, outperforming all the previous methods.

References