PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
Gaku Narita, Takashi Seno, Tomoya Ishikawa, Yohsuke Kaji
I INTRODUCTION
Geometric and semantic scene understanding in 3D environments has an important role in autonomous robotics and context-aware augmented reality (AR) applications. Geometric scene understanding such as visual simultaneous localization and mapping (SLAM) and 3D reconstruction has been widely discussed since the early days of both the robotics and computer vision communities. In recent years, semantic mapping, which not only reconstructs the 3D structure of a scene but also recognizes what exists in the environment, has attracted much attention because of the great progress of deep neural networks.
Semantic mapping systems could take a variety of approaches in terms of geometry and semantics. When we think about robotic and AR applications that deeply interact with the real world, what kind of properties are required for the ideal semantic mapping system? In terms of geometry, it needs to be able to reconstruct a large-scale scene, not sparsely but densely. Additionally, the 3D reconstruction desirably needs to be represented as a volumetric map, not just point clouds or surfels, because it is difficult to directly utilize point clouds and surfels for robot–object collision detection or robot navigation. In terms of semantics, which we mainly focus on in this paper, we believe that it is important for the mapping system to have a holistic scene understanding capability, that is to say, dense semantic labeling as well as individual object discrimination. This is because densely labeled semantics is a crucial cue for intelligent robot navigation, and also, discriminating individual objects is essential for robot–object interaction.
Turning our eyes to the field of 2D image recognition, an image understanding task called panoptic segmentation has been proposed recently . In the panoptic segmentation task, semantic classes are defined as a set of stuff classes (amorphous regions, such as floors, walls, the sky and roads) and thing classes (countable objects, such as chairs, tables, people and vehicles) and one needs to predict class labels on stuff regions and both class labels and instance IDs on thing regions, where the predictions should be performed for each pixel. Extending this point of view to 3D mapping, in this paper we propose the PanopticFusion system. To the best of our knowledge, it is the first semantic mapping system that realizes scene understanding at the level of stuff and things. Our system incrementally performs large-scale 3D surface reconstruction online, as well as dense class label prediction on the background region and segmentation and recognition of individual foreground objects, as shown in Fig. 1.
Our approach first passes the incoming RGB frame to 2D semantic and instance segmentation networks and obtains a panoptic label image in which class labels are assigned to stuff pixels and instance IDs to thing pixels. The predicted panoptic labels and depth measurements are integrated into the volumetric map. Before integration, we keep the consistency of instance IDs, which possibly change from frame to frame, by referring to the volumetric map at that moment. In addition, we regularize the map using a fully connected CRF model with respect to panoptic labels. For CRF inference, we propose a unary potential approximation using limited information stored in the map. We also present a map division strategy that achieves a significant reduction in computational time without a drop in accuracy.
We evaluated the performance of our system on the ScanNet v2 dataset , a richly annotated large-scale dataset for indoor scene understanding. The results revealed that PanopticFusion is superior or comparable to the state-of-the-art offline 3D DNN methods in the both 3D semantic and instance segmentation tasks. Note that our system is not limited to indoor scenes. Finally, we demonstrated a promising AR application using the 3D panoptic map generated by our system.
The main contributions of this paper are the following:
The first reported semantic mapping system that realizes scene understanding at the level of stuff and things.
Large-scale 3D reconstruction and labeled mesh extraction thanks to the use of a spatially hashed volumetric map representation.
Map regularization using a fully connected CRF with a novel unary potential approximation and map division strategy.
Superior or comparable results in both 3D semantic and instance segmentation tasks, in comparison with the state-of-the-art offline 3D DNN methods.
II RELATED WORK
Previously proposed representative semantic mapping systems related to our PanopticFusion system are shown in Table I. These systems can be divided into two categories from the perspective of semantics: the dense labeling approach and the object-oriented approach.
The dense labeling approach builds a single 3D map and assigns a class label or a probability distribution of class labels to each surfel or voxel to realize a dense 3D semantic segmentation. Hermans et al. utilize random decision forests for 2D semantic segmentation and transfer the inferred probability distributions to point clouds with a Bayesian update scheme. Extending the approach of Hermans et al. , SemanticFusion improves the recognition performance by using CNNs for 2D semantic segmentation and makes use of ElasticFusion for a SLAM system to generate a globally consistent map. Xiang et al. presented KinectFusion-based volumetric mapping with novel data associated RNNs for improving the segmentation accuracy. While these methods realize dense scene understanding, they suffer from the drawback that they are not able to distinguish individual objects in the scene.
Methods adopted in the early days of the object-oriented approach leverage 3D model databases. SLAM++ performs point pair feature-based object detection and feeds the detected objects into a pose graph. Tateno et al. proposed a 3D object detection and pose estimation system that combines unsupervised geometric segmentation and global 3D descriptor matching. These methods, however, require the shapes of objects in the scene to be exactly the same as the 3D models in the database. Recently, several studies on the object-oriented approach using a CNN-based 2D object detector have been reported. Sünderhauf et al. and Nakajima et al. combine a 2D object detector and unsupervised geometric segmentation in order to detect objects in point clouds or a surfel map. MaskFusion , Fusion++ and MID-Fusion introduced an object-oriented map representation that individually builds 3D maps for each object based on 2D object detection. The object-oriented map representation enables tracking of individual objects and an object-level pose graph optimization . However, the quantitative recognition performance of these methods is not clear because they mainly evaluate the camera trajectory accuracy. Furthermore, they focus on foreground objects, resulting in a lack of semantics and/or geometry of background regions.
In contrast to these related studies, PanopticFusion realizes holistic scene reconstruction and dense semantic labeling with the ability to discriminate individual objects. Our system builds a single volumetric map, similar to dense labeling approaches, yet each voxel stores neither class labels nor class probability distributions but DNN-predicted panoptic labels in order to seamlessly manage both stuff and things semantics. The class labels of foreground objects can be restored by a probability integration process. In addition, our 3D reconstruction leverages the truncated signed distance field (TSDF) volumetric map with the voxel hashing data structure , which allows us to reconstruct a large-scale scene as well as extract labeled meshes by using marching cubes , in contrast to the 3D maps of previous methods, which are based on point clouds , surfels and a fixed-sized voxel grid . It should be noted that, with 3D DNN methods that directly apply deep networks to 3D data such as point clouds or voxel grids, high recognition performance has been reported . Nevertheless, with those methods, it is basically necessary to reconstruct the whole scene in advance, requiring offline processing, which could limit their application to robotics and AR. On the contrary, PanopticFusion is an online and incremental framework.
III METHOD
Fig. 2 shows the system overview of PanopticFusion. Our system first feeds an incoming RGB frame into 2D semantic and instance segmentation networks and obtains pixel-wise panoptic labels by fusing the two outputs (Section III-C). The panoptic labels are carefully tracked by referring to the volumetric map at that moment (Section III-D) and are integrated into the map with depth measurements (Section III-E). Probability distributions of class labels for foreground objects are also incrementally integrated (Section III-F). In addition, online map regularization with a fully-connected CRF model is performed for a further improvement of the recognition accuracy. Note that camera poses with respect to the volumetric map are given by an external vSLAM, and labeled meshes are extracted by using marching cubes .
III-B Volumetric Map
We use the TSDF-based volumetric map representation with a voxel hashing approach , which manages spatially hashed small regular voxel grids called voxel blocks. This approach is memory efficient compared with a single voxel grid approach like the original KinectFusion and enables us to reconstruct large-scale scenes. Our implementation is based on voxblox , which is a CPU-based TSDF mapping system, but we extend it to integrate the semantics.
III-C 2D Panoptic Label Prediction
III-D Panoptic Label Tracking
III-E Volumetric Integration
Here, denotes the distance between the voxel and the surface boundary, and a quadric weight that takes the reliability of depth measurements into account. Similar updating is applied to the voxel color .
In contrast, if those panoptic labels do not coincide, we decrement the weight:
III-F Thing Label Probability Integration
The thing label predicted by Mask R-CNN is frequently uncertain even while the segmentation mask is accurate, especially in the case where a small part of the object is visible. Hence we probabilistically integrate thing labels instead of assigning a single label to each foreground object:
Weighting the probability distributions with the detection confidence allows the final distribution to preferentially reflect reliable detections.
III-G Online Map Regularization
While the integration scheme described above yields a reliable 3D panoptic map, it is possible to further improve the recognition accuracy by using a map regularization with a fully connected CRF model. A fully connected CRF with Gaussian edge potentials has been widely used in 2D image segmentation since an efficient inference method was proposed . Recently, several studies that apply it to a 3D map, such as surfels or occupancy grids, have been reported . In those approaches, CRF models are constructed with respect to class labels whose number is fixed, whereas we consider the CRF with respect to panoptic labels whose number depends on the scene and is theoretically not limited. Here we are faced with two problems: how to properly compute unary potentials for panoptic labels, and how to infer a CRF whose number of labels is potentially large within a practical time.
While it is non-trivial which unary potentials should be used for a panoptic label CRF, we use a negative logarithm of a probability distribution following a standard class label CRF:
We utilize a linear combination of Gaussian kernels for pairwise potentials because the efficient inference method can be applied:
Here, is a simple Potts model. As in , we chose the following two kernels which regularize the map with respect to voxel colors and locations, respectively:
III-G2 Unary Potential Approximation
Previous approaches assigned a probability distribution to each surfel or voxel, which can be used directly to compute unary potentials; in contrast, from the viewpoint of memory efficiency, we store only a single label in each voxel. Therefore, we approximate the unary potentials using only a single label, and weights stored in a voxel, based on a certain assumption described as follows.
In addition, from the TSDF update scheme in Eq. (7) we have,
Consequently, the probability that the current panoptic label in the voxel is actually correct can be calculated as,
It is unfortunately not possible to calculate the exact probability that the voxel takes a label other than the current label because the map does not record all the information about previously integrated labels. Therefore, we approximate the probability as follows, where denotes the number of panoptic labels in the map:
Finally, we obtain the unary potential from Eq. (13), (19) and (20). In spite of the approximated approach, it realizes quantitative and qualitative improvements in recognition accuracy, as shown in Section IV-C.
III-G3 Map Division for Online Inference
The computational complexity of the inference algorithm proposed by Krähenbühl et al. is , where and are the numbers of voxels and panoptic labels, respectively. In our problem setting, however, is theoretically limitless and could in practice be large, e.g. several hundreds, which would make online inference impracticable. To solve this problem, we present a map division strategy. When we divide the volumetric map into spatially contiguous submaps, the number of panoptic labels in each submap can be expected to be . Hence, the total computational complexity could be reduced to . The map is divided by the block-wise region growing approach based on the predefined maximum number of voxel blocks. The division process has little effect on computational time.
IV EVALUATION
We employed ResNet-50 for the backbone of PSPNet. The network was initialized with the ADE20K pre-trained weights, and was then fine-tuned using a SGD optimizer for 30 epochs with a learning rate of 0.01 and a batch size of 2. We leveraged ResNet-101-FPN for the Mask R-CNN’s backbone. After initialization with MS COCO pre-trained weights, the network was fine-tuned by 4-step alternating learning using an ADAM optimizer for 25 epochs with a learning rate of 0.001 and a batch size of 1We used a publicly available implementation of and for PSPNet and Mask R-CNN, respectively..
IV-B Quantitative and Qualitative Results
Fig. 3 shows examples of 3D panoptic maps generated by our system. Unfortunately, there are no semantic mapping systems or 3D DNNs that can recognize a 3D scene at the level of stuff and things. Therefore, we evaluated the performance on two sub-tasks, 3D semantic segmentation and instance segmentation, for a quantitative comparison. In this evaluation, we used the hidden test set of ScanNet v2. We show the results in Tables II and III. In the tables, the state-of-the-art methods that apply 3D DNNs to points or volumetric grids are listed. Note that the methods of leverage RGB images with associated camera poses as well. Our system that uses only 2D-based recognition modules surprisingly achieves comparable or superior performance compared with those methods, thanks to the careful integration of multi-view predictions. In terms of the class-wise accuracy, the results revealed that our system has advantages especially in the case of small objects such as sinks and pictures, and objects that are confusing to recognize only by their geometry, such as beds, bookshelves, and curtains. In Table II, several semantic segmentation methods outperform our system because of their large receptive fields in 3D space. However, these methods basically need to reconstruct the entire scene in advance, assuming offline process, while our system is an online and incremental framework. How to apply 3D DNNs to partial observations and how to integrate them into an online mapping system are left for future work.
Additionally, we evaluated 3D panoptic quality on the open test set of ScanNet v2, although there are no quantitatively comparable methods. We employed the evaluation criteria originally proposed in . Note that the quality was evaluated with respect to each vertex instead of each pixel, and, as with the ScanNet 3D semantic instance benchmark, we ignored the predicted things with less than 100 vertices. We show the panoptic quality (PQ) as well as the segmentation quality (SQ) and recognition quality (RQ) in Table IV. We hope these results will invigorate research in this field.
IV-C Evaluation of Map Regularization
In this section, we evaluate the map regularization proposed in Section III-G. First, we evaluated the effects of the map division on the recognition accuracy and computational time. We used the open test set for the recognition accuracy and typical scenes in ScanNet v2 for the computational time. The result is shown in Fig. 4. Note that, in this experiment, we applied regularization to the pre-generated map as a post process to evaluate solely the effects of CRF.
As can been seen, the recognition performance was improved by the map regularization with the proposed unary potential approximation regardless of whether or not map division was used. The results also show that the map division strategy drastically reduced the computational time without a decrease in recognition performance, compared with the case of building a CRF model for a whole map.
Based on the above results, our online system employed map regularization with the map division strategy. We chose a maximum number of voxel blocks of 25 because of the better recognition accuracy and acceptable computational time. Table IV shows the difference in recognition performance due to whether or not map regularization was used in online processing. This result shows that the map regularization improved the recognition performance even when the system ran online. Note that the scores of almost all the classes were boosted by the proposed regularization. See Fig. 5 for qualitative effects of the map regularization.
IV-D Run-time Analysis
Table V shows computational times for each component of our system, which are measured on scene0645_01, a typical large-scale scene in ScanNet v2 (shown in Fig. 1). PSPNet and Mask R-CNN each run on GPUs, and the other components are processed on a CPU. All components are basically processed in parallel. The throughput of our system is around 4.3 Hz, which is determined by Mask R-CNN, the bottleneck process of our system. Although our current implementation is not highly optimized, our system is able to run at a rate allowing interaction. Note that the computational time except for the map regularization does not depend on the scale of scenes nor the number of things because we utilize the raycasting approach for the integrations. The processing time of the map regularization increases to about 10 seconds at the end of the sequence, but it could be reduced by processing only the voxel blocks near the camera frustum.
V APPLICATIONS
In this section, we demonstrate a promising AR application utilizing a 3D panoptic map generated online by the proposed system. A 3D panoptic map reconstructed as 3D meshes allows us to realize the following visualizations according to the context of the scene:
Path planning on stuff regions such as floors and walls.
Interaction with individual objects, or the thing regions.
Interaction appropriate for the semantics of each region.
Natural occlusion and collision visualization.
We show an example of an AR application utilizing the above visualizations in Fig. 6. Humanoids and insect-type robots are able to locomote on the floor and wall meshes, respectively, according to the automatic path planning. Additionally, the semantics of each object realizes context-aware interactions such that humanoids sit and lie on chairs and sofas, respectively, and CG objects appear on tables. Moreover, we can naturally visualize the occlusion effects, which are important for AR, because the 3D meshes of the scene are extracted. Note that, taking advantage of the accurately recognized 3D panoptic map, we can easily estimate the poses of seats of chairs and sofas, and top panels of tables by using simple normal- and curvature-based segmentation and plane detection.
We believe that our system is useful not only for AR scenarios but also for autonomous robots that explore scenes and manipulate objects.
VI CONCLUSIONS
In this paper, we have introduced a novel online volumetric semantic mapping system at the level of stuff and things. It performs dense semantic labeling while discriminating individual objects, as well as large-scale 3D reconstruction and labeled mesh extraction thanks to the use of a spatially hashed volumetric map representation. This was realized by pixel-wise panoptic label prediction and its volumetric integration with careful label tracking. In addition, we constructed a fully connected CRF model with respect to panoptic labels and inferred it online with a novel unary potential approximation and a map division strategy, which further improved the recognition performance. The experimental results showed that our system outperformed or compared well with state-of-the-art offline 3D DNN methods in terms of both 3D semantic and instance segmentation. In future work, we plan to extend our system to ensure global consistency against long-term pose drift, to perform high-throughput mapping by network reduction, and to support dynamic environments.
We believe that the stuff and things-level semantic mapping will open the way to new applications of intelligent autonomous robotics and context-aware augmented reality that deeply interact with the real world.