S3CNet: A Sparse Semantic Scene Completion Network for LiDAR Point Clouds

Ran Cheng, Christopher Agia, Yuan Ren, Xinhai Li, Liu Bingbing

Introduction

Scene understanding is a challenging component of the autonomous driving problem, and is considered by many as a foundational building block of a complete self-driving system. The construction of maps, the process of locating static and dynamic objects, and the response to the environment during self-driving are inseparable from the agent’s understanding of the 3D scene. In practice, scene understanding is typically approached by semantic segmentation. When working with sparse LiDAR scans, scene understanding is not only reflected in the semantic segmentation of the 3D point cloud but also includes the prediction and completion of certain regions, that is, semantic scene completion. Scene completion is the premise of 3D semantic map construction and is a hot topic in current research. However, as semantic segmentation based on 2D image content has reached very mature levels, methods that infer complete structure and semantics from 3D point cloud scenes, which is of significant import to robust perception and autonomous driving, are only at preliminary development and exploration stages. Limited by the sparsity of point cloud data and the lack of features, it becomes very difficult to extract useful semantic information in 3D scenes. Therefore, the understanding of large scale scenes based on 3D point clouds has become an industry endeavor.

Semantic segmentation and scene completion of 3D point clouds are usually studied separately , but with the emergence of large-scale datasets such as ScanNet and SemanticKITTI , researchers have discovered a deep intertwining of an object’s semantics with its underlying geometry, and since, have begun exploiting this with the joint learning of semantic segmentation and scene completion to boost model performance . For instance, speculating that an object occluded by vehicles and surrounded by leaves is a trunk simplifies the task of inferring it’s shape. Conversely, inferring the shape of a pole-like object forms a prior on it’s semantic class being a trunk rather than a wall. While previous semantic scene completion methods built on dense 2D or 3D convolutional layers have done well in small-scale indoor environments, they have struggled to maintain their accuracy and efficiency in outdoor environments for several reasons. For one, dense 2D convolutional methods that thrived in the feature rich 2D image space are no longer sufficient when tackling large and sparse LiDAR scans that contain far fewer geometric and semantic descriptors. Furthermore, the dense 3D convolution becomes extremely wasteful in terms of computation and memory since the majority of the 3D volume of interest is in fact empty. Thereby, our main contributions are listed as the following: (a) a sparse tensor based neural network architecture that efficiently learns features from sparse 3D point cloud data and jointly solves the coupled scene completion and semantic segmentation problem; (b) a novel geometric-aware 3D tensor segmentation loss; (c) a multi-view fusion and semantic post-processing strategy addressing the challenges of distant or occluded regions and small-sized objects. Given a single sparse point cloud frame, our model predicts a dense 3D occupancy cuboid with semantic labels assigned to each voxel cell (as shown in Fig. 1), generating rich information of the 3D environment that is not contained in the original input such as gaps between LiDAR scans, occluded regions and future scenes.

In order to effectively complete occluded voxel regions from LiDAR scans, we focus on exploiting the geometrical relationship of the 3D points both locally and globally. In this work, we utilize point-wise normal vectors as a geometrical feature encoding to guide our model in filling the gaps according to the object’s local surface convexity. We also leverage a LiDAR-based flipped Truncated Signed Distance Function (fTSDF ) computed from a spherical range image as a spatial encoding to differentiate free, occupied and occluded space of a scene. As for future scenes, because these regions are far from the vehicle and are primarily road or other forms of terrain, we propose a 2D variant of the sparse semantic scene completion network to support the construction of the 3D scene via multi-view fusion with Bird’s Eye View (BEV) semantic map predictions. To tackle sparsity, we leveraged the Minkowski Engine , an auto-differentiation library for sparse tensors to build our 2D and 3D semantic scene completion network. We have also adopted a combined geometric inspired semantic segmentation loss to improve the accuracy of semantic label predictions. Since our network is trained in a complex real-world autonomous driving dataset with 20 classes of dynamic and static objects, and the input data is simply a voxelized LiDAR point cloud appended with geometrical and spatial feature encodings, our model can be deployed on-the-go with various LiDAR sensors. We demonstrate this by applying our method to unseen real-world voxel data, which yields reasonable qualitative results. Our experiments show that our model outperforms all baseline methods by a large margin, with exceptional performance in the prediction of small, under-represented class categories such as bicycles, pedestrians, traffic signs and more.

Related Works

We review the related works across four major areas: volume reconstruction, point cloud segmentation, semantic scene completion, and multi-view fusion.

Volume Reconstruction. There are several approaches to inferring complete volumetric occupancy of shapes and scenes from partial or sparse geometric data. Efficient methods based on object symmetry and plane fitting apply for small non-complex completion tasks. In larger scenes with irregular objects, the former are supplanted by methods that fit 3D mesh models to object instances based on the local scene geometry . Yet, a lack of diversity within the 3D model library often leads to incomplete reconstruction, and expanding the library slows down retrieval. Attempts to simplify the process to 3D bounding box fitting neglects the local geometry of objects . Other studies process grid-octree data with CNNs to predict high resolution outputs , but are tailored to reconstructing individual objects rather than entire scenes.

Point Cloud Segmentation. A variety of methods segment range images constructed by spherical projection of a point cloud with deep networks, and back-project the predicted semantic classes onto the corresponding points in 3D space . The small sized range image tensor that these methods operate on leads to unparalleled speed, but ∼25%\sim 25\% of the original point cloud is unrecoverable after the initial spherical projection. Alternative approaches project point clouds onto a bird’s eye view perspective, computing pillar features (voxels) from 3D points to construct a BEV map and process it with deep CNNs . Another emergent stream aims to segment point clouds by hierarchically extracting features from the 3D points directly to capture local and global context of the scan . While these methods have had reasonable success, they are typically an order of magnitude slower than their non-direct counterparts.

Semantic Scene Completion. Most mainstream methods are built on small indoor scenes with dense depth maps or dense point clouds, which effectively reduces the impact of sparseness on scene completion. The representative methods are SSCNet and ScanNet . Their approach transforms dense depth maps into a volumetric TSDF signal and passes it through a dense 3D network. Extensions are made in TS3D by incorporating semantic segmentation from RGB images into the construction of the input TSDF volume. In outdoor scenes, these methods are limited by the vast increase in sparsity and their network architectures that predict at reduced voxel resolutions. Improvements are made in SATNet , which maintains high output resolution with a series of dilated convolutional modules (TNet) to capture global context, and propose configurations for early and late fusion of semantic information. Behley et al. 2019 produced a state-of-the-art system by integrating semantic priors from color images and LiDAR segmentation with the TSDF volume, and adopting a TNet network backbone. However, their memory intensive design requires users to infer on 6 equal parts of the input before fusing the semantic completed partial scenes, deterring potential real-time usage. Furthermore, the dependence on color images introduces instability in the overall system under low-light and poor weather conditions. Therefore, the design of a semantic scene completion method based solely on 3D point cloud data is motivated by the need for faster inference speeds, a smaller memory footprint, and robustness to extreme conditions.

Multi-view Fusion. Fusing semantic and geometric features across various view-points and dimensions has been explored for a variety of tasks. Several of the aforementioned semantic scene completion methods demonstrate the lifting of image-space semantic labels into 3D space . Similarly, Hane et al. 2013 proposed a joint optimization strategy for semantic image segmentation and scene reconstruction, using the completed 3D scene as a geometric prior on the corresponding image pixels to enforce spatially consistent segmentations. Extensions were made to support city-scale reconstruction at reasonable memory costs with an adaptive multi-resolution model . Dai and Nießner 2018 designed an 2D-3D network that fuses semantic features from RGB-D images with a differential backprojection layer, achieving multiple view-point reconstruction in indoor settings. Meanwhile, SLAM-based approaches aim to produce temporally consistent reconstructions by tracking motion states over sequential frames and mapping predicted semantics into 3D space.

Method

We describe our methods for LiDAR-based semantic scene completion in large outdoor driving scenes. After a brief system overview, we present our unique procedure for computing key spatial features from an input LiDAR scan, the detailed design of our networks, fusion module and refinement module, and a novel loss function incorporating a geometric-aware 3D segmentation loss.

The entire system pipeline is shown in Fig. 2. From a single LiDAR scan, we construct two sparse tensors that encapsulate the scene into memory efficient 2D and 3D representations. Each tensor is passed through their corresponding semantic scene completion network, 2D S3CNet or 3D S3CNet, to semantically complete the scene in the respective dimension. We propose a dynamic voxel fusion method to further densify the reconstructed scene with the predicted 2D semantic BEV map (detailed discussion in Section 3.3). This offsets the significant memory demands on the 3D network - exponential sparsity growth in 3D space makes it difficult to complete classes at range. Using a sparse tensor spatial propagation network , we refine the semantic labels in noisy regions of the fused 2D-3D predictions.

2 Spatial Feature Engineering

Sparse 2D Feature. We define the sparse 2D tensor as the set of non-empty pillars approximating the point distribution along the xx-yy plane (BEV). Each pillar, pi,jp_{i,j}, is a 7-dimensional vector encoding the mean, minimum, and maximum heights and intensities of points within the voxel, and the point density; all values are normalized.

Sparse 3D Feature. The sparsity of the raw point cloud makes it difficult to extract spatial features, and thus, we transform it into a range image via spherical projection. As the quality of extracted spatial features are sensitive to noise, we retrieve a smooth range image by performing dynamic depth completion using the dilation method of Ku et al. 2018. This enables robust extraction of a 3-dimensional normal surface feature for each pixel that is reversely assigned to the points in 3D space. Further, we maintain a memory efficient sparse 3D tensor by modifying the sign-flipped TSDF approach of Song et al. 2017, and compute TSDF values from the smooth range image, storing only the coordinates within the truncation range of existing LiDAR points. All other features for TSDF generated coordinates are zero-padded.

3 Network Architecture

We adopt the Minkowski Engine as our sparse tensor auto-differentiation framework to build the entire system. A sparse tensor can be defined as a hash-table of coordinates and their corresponding features: x=[Cn×d,Fn×m]\mathbf{x}=[\mathbf{C}_{n\times d},\mathbf{F}_{n\times m}]. Here, nn are the number of non-empty voxels, dd and mm are the dimension of coordinates and features, respectively. The sparse convolutional layer is thus:

Where KK is the kernel size and ND(u,K,Cin)\mathcal{N}^{D}(\mathbf{u},K,\mathcal{C}^{in}) are the set of offsets that are at most ⌈12(K−1)⌉\lceil\frac{1}{2}(K-1)\rceil away from u\mathbf{u}, the current coordinate. Unlike the conventional convolution, this generalized convolution (introduced by Choy et al. 2019 et al) suits generic input and output coordinates, and arbitrary kernel shapes. It allows extending a sparse tensor network to extremely high-dimensional spaces and dynamically generating coordinates for generative tasks. The arbitrary input shape and multi-dimensional support enables us deploy the same network layout in different dimensional spaces. Hence, our 2D and 3D networks share the same set of network components (built off sparse convolution, transposed convolution and pooling layers) with differing coordinate dimensions.

After extracting the features as described in Section 3.2, we create the sparse 3D tensor by voxelizing the point cloud and extracting the coordinates and features of non-empty or TSDF generated points. If multiple points occupy the same voxel, their features are averaged. The voxel resolution is 0.2m in each dimension; this applies to the sparse 2D tensor as well. The spatial extent of our predictions in the xx-yy-zz directions are [0m, 51.2m], [-25.6m, 25.6m], [-2m, 4.4m], respectively. Discretizing the volume yields a sized tensor, upon which our sparse 3D network will predict the set of occupied coordinates and a probability distribution over the 20 possible semantic categories.

As illustrated in Fig. 4, our 3D network is assembled from four primary building blocks: encode, decode, dilation, and spatial propagation. We implement a squeeze re-weight (SR) layer to model inter-channel dependencies and improve generalization. Each encoder block contains a context aggregation module (CAM) which captures context in a large receptive field, improving robustness to dropout noise. We adapt these modules for sparse tensor support, and observe improvements to both the segmentation and completion tasks. The decode blocks utilize sparse transposed convolutions capable of generating new coordinates with the outer product of the weight kernel and the input coordinates. As new coordinates in 3D are generated by the cube of the kernel size, we preserve memory with pruning modules that remove redundant coordinates throughout scene completion - training supervision is provided by ground truth filter masks (Fig. 4, target_\_s2-16). The dilation block features a sparse atrous spatial pyramid pooling (ASPP) module to trade-off accurate localization (small field-of-view) with context assimilation (large field-of-view), our dilation rates are . The spatial propagation block contains two parts, 3D spatial propagation module and guidance convolution network. The guidance network produces affinity matrices used to guide the spatial propagation network to deform the sparse tensor into a desired 3D shape.

Loss function. When training our 2D network, we accomplish healthy BEV completion and balanced learning of under-represented class categories by combining pixel-wise focal loss, weighted cross entropy loss. Per-voxel binary cross entropy loss is used to train the pruning modules.

Here, pp is the predicted label and yy is the ground truth label. The weighting factors α\alpha, β\beta, and ω\omega are empirically set to 0.5, 0.5, and 1, respectively. We define a novel loss function to train our 3D network. It consists of a completion term (voxelized binary cross entropy) and a geometric-aware 3D tensor segmentation loss. The two terms are balanced by an λ\lambda constant empirically set to 0.35.

The completion loss is expressed below, where LBCE\mathcal{L}_{\mathbf{BCE}} is the binary cross entropy loss over volumetric occupancy. Note that this is applied to train the pruning modules at various scale spaces.

Below is the geometric-aware 3D tensor segmentation loss computed over the final output tensor.

The terms MLGAM_{LGA}, ξ\xi and η\eta describe signals computed from the same local cubic neighborhood of a given coordinate (As shown in Fig. 5). The Local Geometric Anisotropy defined by Li et al. 2019, MLGA=∑i=1K(cp⊕cqi)M_{LGA}=\sum_{i=1}^{K}{(c_{p}\oplus c_{qi})}, is a discrete smoothness promoting signal that penalizes the prediction of many classes within a local neighborhood. Although it ensures locally consistent predictions in homogeneous regions (i.e. road center, middle of a wall), it may incorrectly divert the model from inferring boundaries between separate objects. We thus introduce η\eta, which accounts for the local arrangement of classes based on the volumetric gradient, and downscales MLGAM_{LGA} when the neighborhood contains structured predictions of different classes. To smoothen-out the loss manifold, we include a continuous entropy signal ξ=−∑c∈C′P(c)log(P(c))\xi=-\sum_{c\in C^{\prime}}{P(c)log(P(c))}, where P(c)P(c) is the distribution of class cc amongst all classes C′C^{\prime} in the local neighborhood.

Eq. (5) thus decomposes into two multiplicative factors with the cross-entropy loss which explicitly models the relationship between a predicted voxel and it’s local neighborhood. Intuitively, the smooth local entropy term, ξ\xi, down-scales the loss in easily classified homogeneous regions, enabling the network to attend to non-uniform regions (e.g. class boundaries) as learning progresses. However, a measure of non-uniformity alone is insufficient in that non-uniform regions should be more heavily penalized if the predicted neighborhood lacks structure. This motivates the inclusion of ηMLGA\eta M_{LGA} which also considers spacial arrangement of classes and down-scales the loss in structured cubic regions. In combination, we acquire a smooth loss manifold when the local prediction is close to the ground truth as well as uncluttered, with sharp increases when the local cubic neighbourhood is noisy and far off from the ground truth. This speeds up convergence while reducing the chance of stabilizing in local optimums.

Multi-view Fusion and Spatial Propagation.

The task of 2D semantic scene completion is far less complex than it’s 3D counterpart because it does not account for the local structure of objects along the zz dimension, relegating the task to semantic filling in the bird’s eye view plane. This allows our sparse 2D network to more easily understand the global distribution pattern of semantic classes even at far distances, yielding accurate predictions of roads, terrain, and side walks. When confronted with heavy noise or occlusions, such as Fig. 6, our 3D model may lack the confidence to complete certain regions. To remedy this, we express the predicted 3D tensor as a stack of BEV images, and attempt to fill in each empty voxel with the corresponding prediction in the 2D network. The lifting algorithm operates as follows: (1) split the 3D volume into layers along the z-dimension (32 layers in this application); (2) locate the slice with the maximum occurrence of the desired class; (3) fill each voxel with the associated 2D pixel prediction if there exists such a class in the voxel’s n×nn\times n (n=3n=3) neighborhood, otherwise attempt to fill in the above voxel until the highest layer is exceeded; (4) repeat for all classes. Hence, by lifting the 2D voxel-wise semantic completion into 3D space, we improve both completion and segmentation metrics of the 3D task at minimal cost. To mitigate any additional post-fusion noise, we apply a spatial propagation network that refines the segmentation results in 3D space.

Experiments

We evaluate our system on the SemanticKITTI dataset , which contains 8,728 voxelized LiDAR scans and the corresponding ground truth labels. We use the public train/val/test split defined in the SemanticKITTI API, and provide qualitative visualizations of our predictions. For evaluation metrics, we follow using mean IoU over positive classes and binary per-voxel completion IoU. Experiments are deployed on Nvidia GP100 GPUs (16GB Graphic RAM); we use two GPUs per model and train for 50 epochs each. The average inference time for 3D S3CNet is 0.5s (frame with 40,000±\pm500 points), and training requires around 120 hours. For 2D S3CNet, the average inference time is 0.05s resulting in 50 training hours. The training scheme for the best performing 2D and 3D models are: (3D) Adam optimizer, 0.0025 learning rate, 0.0005 weight decay, and (2D) SGD optimizer, 0.001 learning rate, and 0.0005 weight decay. Both experiments also incorporate an exponential learning rate scheduler with a decay rate of 0.9 every 10 epochs.

Qualitative Results. Fig. 9 depicts the inferences of our system on three SemanticKITTI test set samples and the corresponding RGB images. We notice that our method correctly captures most classes with exceptional detail and consistency, including small object categories like people and traffic signs in challenging scenes. The middle column highlights an open intersection scene with heavy front-facing occlusion, complicating the task of inferring distant mid-intersection vehicles. For 2D S3CNet, we demonstrate the prediction and ground truth BEV map in static scenes and dynamic scenes with moving vehicles (as shown in Fig. 8).

Quantitative results. Our primary results are shown in Table. 1, where we compare our overall system to several state-of-the-art methods in the SemanticKITTI test set benchmark. Note the approaches without citations are non-published works. At the time of writing our proposed S3CNet considerably outperforms all others, achieving a mean IoU score of 29.5%\% (a +23.9%\% improvement on the previous leading method) and a +66.6%\% over the baseline method . For 2D S3CNet, we conduct extra experiments on the nuScenes dataset . Ground truth 2D semantic scenes are created from object bounding boxes and cropped HD map data, and aligned with LiDAR frames at an identical spatial extent and voxel resolution to SemanticKITTI. We produced 1,166,187 frames of valid LiDAR scan and semantic scene label pairs and we split the train, validation, test into a 14:3:3 ratio. The total training time of our 2D S3CNet for 50 epochs on the nuScene dataset is ∼\sim350 hours. Table. 2 shows that our 2D S3CNet outmatches several well-known LiDAR segmentation baselines (adapted for BEV predictions) on the segmentation component of the SSC task, with comparable results on the completion component for both the SemanticKITTI and nuScenes datasets.

Ablation study. We investigate the individual contribution of all components in our system. As shown in Table. 3, we conduct various experiments on the 2D and 3D SSC task by modifying or removing core components of the system and track the resulting effect on mean IoU and completion IoU scores on the SemanticKITTI validation set. For both the 2D and 3D networks, we observe a mean IoU drop after removing the CAM and SR modules. Since the decode blocks are mainly responsible for completion, removing SR modules results in a more severe drop in completion IoU compared to CAM which are only present in the encoder. Eliminating spatial features decreases overall performance by a large amount; the system losses geometric priors that guide completion and spatial priors that distinguish free from occluded space. Adopting Lovasz-softmax achieves the highest mean IoU increase, since it directly optimizes for the mean IoU metric (Jaccard index). In our experiments, focal loss combined with binary cross entropy loss provides no performative advantage over the baseline weighted cross entropy loss for both the 2D and 3D networks. Post-processing modules like Multi-view Fusion and Spatial Propagation Network demonstrate very high contribution to the final results - without MVF the system performance degrades to that of the focal loss baseline. A key distinction between MVF and SPN is the comparative impact of MVF on completion IoU to the SPN on segmentation mean IoU, respectively.

Data augmentation. To increase model robustness, we integrate a series of data augmentation techniques into the training of our 2D and 3D networks. For both 2D and 3D tasks we apply uniform random cropping and dropout to the LiDAR point cloud, as well as uniform random translation of ±\pm0.1m in all three dimensions. On the 2D datasets, random rotations of ±45∘\pm 45^{\circ} are applied only on the yaw angle, however, we apply random rotation of ±10∘\pm 10^{\circ} on any two of the three Euler angles (roll, pitch, yaw) at a time when training the 3D model.

Conclusion

In this paper, we presented a Sparse Semantic Scene Completion Network, S3CNet, capable of efficiently reconstructing large outdoor scenes and predicting semantic voxel-wise labels from a single LiDAR scan. To complement S3CNet, we designed a novel geometric-aware sparse tensor segmentation loss that promotes class-consistent predictions in homogeneous regions while encouraging locally structured multi-class segmentations along object boundaries. In combination with Multi-View Fusion and Spatial Propagation post-processing modules, our method achieves state-of-the-art results on the SemanticKITTI test set benchmark by a large margin. We also adapt several leading LiDAR segmentation networks as baselines for 2D semantic scene completion, and demonstrate that the 2D variant of S3CNet outperforms these baselines on two large-scale datasets. Furthermore, we include a detailed ablation study highlighting the contribution of each individual component to the overall system. Future work holds the extension of our sparse tensor method to real-time speeds with improved performance across the board, additional investigation on useful spatial feature encodings, and a learning-based multi-view fusion technique to enable end-to-end learning.

We would like to thank Professor John K. Tsotsos for reviewing this work, as well as the anonymous reviewers for their valuable suggestions.

References

Appendix A: Extra Qualitative Results

We present additional qualitative results on the SemanticKITTI training set. As we can see in Fig. 9, our model well captures both completion and semantic segmentation characteristics of different scenes. Because a single ground truth label was constructed from LiDAR scans across several time steps (in order to densify the scene), dynamic objects were filtered out to avoid labelling noise. For instance, in the bottom-right most sample, while the bus was filtered out of the ground-truth scene, our well engineered features enabled our model to detect the object with the correct class label. Another example of this is visible in the top-right most sample, where our model detects a moving bicyclist in front of a vehicle. These object labels are learned from static scenes, but are successfully inferred in dynamic scenes which reflects positively on our model’s generalization ability.

NuScenes 2D qualitative results: we project the 2D semantic BEV map back to 3D lidar points according the respective x and y coordinates, and overlay the semantic point cloud data on the rgb images to show the model’s qualitative results. As we can see from Fig. 10, our predictions cover most of the road surface and precisely detects vehicles, pole-like objects and pedestrians.

Appendix B: Experiment Configurations

We provide the training scheme for all 2D experiments in Table. 4, which includes the adapted LiDAR segmentation baselines that have undergone multiple optimization iterations to achieve the stated results on the SemanticKITTI and NuScenes datasets. Note that all experiments use identical input data and augmentation configurations; changes only occur to the model, the loss function, optimizer, scheduler, and supporting hyperparameters - much of which are based on the training scheme proposed by the original works.

We also list our experiment configurations including the loss function, hyperparameters, scheduler, and provide the best validation mean IoU for all 3D model experiments (see Table. 5). We implemented a list of state-of-the-art methods (not currently on the SemanticKITTI benchmark) and trained them on the SemanticKITTI dataset for 50 epochs. Amongst the competitors, our model achieved the best mean IoU on the validation split (sequence 08).