Multi-view PointNet for 3D Scene Understanding

Maximilian Jaritz, Jiayuan Gu, Hao Su

Introduction

The field of 3D perception is evolving at a fast pace, with recent major improvements on tasks such as semantic segmentation and object detection. This is crucial to applications in robotics and AR/VR, where 3D data are typically captured as depth maps or point clouds, along with 2D images from RGB cameras. A central problem of those applications is how we can efficiently fuse data from the 2D and 3D domains. This is quite challenging, because there usually is no one-to-one mapping between 2D and 3D data, and also the neighborhood definitions in 2D and 3D are different for convolution. More critically, while neighboring pixels are defined by the discrete grid, 3D points are defined at non-uniform continuous locations. Additionally, 3D sensors mostly deliver a much lower resolution than 2D cameras. For example, when the point cloud from a Velodyne HDL-64 Lidar is projected into the camera image, it covers only 5.9% of the pixels .

Point cloud based neural networks have been shown to generate powerful geometry cues for 3D scene understanding. However, not all objects can be distinguished by their shape, especially when they have flat surfaces such as doors, refrigerators and curtains. Therefore, additional color information should be leveraged, but recent results have shown that naively feeding colored point cloud (XYZRGB) to point cloud based networks does only marginally improve the performance over simple point cloud input (XYZ).

We argue that, because RGB cameras have much higher spatial resolution than 3D sensors in most realistic settings, it is better to compute image features in 2D first before lifting the 2D information to 3D. Like so, it is possible to gather additional information from higher resolution images and it is also natural from a sensor fusion perspective, to push modality centric features from different sources to 3D for their combination.

As different representations in 3D exist (voxel, point-cloud, multi-view, etc.), their respective scene understanding methods evolve in parallel. For voxel-based methods, there have been works on how to fuse geometry and image data coherently. However, for the point cloud domain, the common practice is to sparsely copy RGB information to points and there lacks a systematic exploration of how to conduct the fusion more effectively. In order to address this significant drawback of point cloud based methods, we propose MVPNet (Multi-View PointNet), where we first compute 2D image features on multiple, heuristically selected frames, then lift those features to 3D and adaptively aggregate them into the original point cloud (XYZ). Finally, the multi-view augmented point cloud is fed into PointNet++ for semantic segmentation.

There are several advantages: First, the lifted 2D features contain contextual information thanks to the receptive field of the 2D network. Second, the complementary RGB and geometry features are jointly processed in canonical 3D space. And third, our flexible approach can be added to any 3D network.

In this paper, we focus on exploring the 2D-3D fusion problem, a key component for 3D scene parsing. While the key message of our exploration can be concisely summarized as doing early feature fusion is better, the significant performance improvement from baselines is in fact obtained through extensive trials. In the experiment section, we made a rich set of ablation studies so as to compare design choices and inform our discoveries to the vision community.

In summary, the key contributions are as follows:

We propose a simple and fast framework that takes 3D point cloud and 2D RGB-D frames as input and fuses complementary features in the canonical point cloud space for the task of 3D semantic segmentation.

Our method outperforms previously published point cloud based networks by using additional dense image information while handling occlusions.

We provide insights to the design choices in dense-2D/sparse-3D point cloud fusion based on extensive experiments, and showcase its excellent robustness to very sparse point clouds.

Related Work

Several works have shown that lifting 2D features to 3D leads to better performance than just lifting RGB values. In , multiple 2D image feature maps are unprojected to 3D, voxel-volumes are created, combined by max-pooling and then fed into a 3D CNN. In , 2D image features are gathered at nearest neighbor locations defined by a lidar point cloud to build a dense bird view map. These approaches use pixel-level 2D-3D correspondences to lift low-level features as opposed to where only high-level 2D object proposals are lifted to 3D frustums. In this work, we establish pixel-to-point correspondences to lift 2D features to the canonical 3D point cloud space instead of voxel or birdview . The advantage is that once all modalities are represented in a 3D point cloud, correspondence between two data points is precisely defined by distance in the continuous domain without discretization errors.

D Networks

CNNs are the state-of-the-art on 2D RGB images, but competing network families exist for 3D data: 3D CNNs make use of the voxel representation where the raw point cloud data is transformed into a discrete grid of cells and in practice most of the cells are empty and only voxels that lie on the object surface are occupied. On the other hand, point cloud based networks can directly take point clouds as input. In our work, we use point cloud based networks, because of their inherent sparsity as compared to voxel-based methods.

D Semantic Segmentation

The aim of 3D semantic segmentation is to predict a label for every point in a 3D point cloud. PointNet leverages shared Multi Layer Perceptrons (MLPs) to compute point-wise features and uses max-pooling to obtain features for the global point cloud. This works very well for single objects in the ShapeNet dataset for the task of part segmentation. For whole scene analysis, PointNet++ is more suited, because it has set abstraction layers to create a hierarchical network structure akin to CNNs which scales much better to larger point sets. Voxel-based methods include SegCloud , 3DMV and Submanifold Sparse Convolution . The latter defines a very efficient way to deal with sparsely populated voxels by restricting computations to active voxels. Different with 3DMV , we exploit the fusion of multi-view and geometry information in point cloud space and achieve much better performance. In addition, we report the mIoU for all the ablation studies instead of the segmentation accuracy. SPLATNet takes point clouds and images as input and projects them on a permutohedral lattice for convolution and 2D-3D fusion. In our approach, we focus on fusing multi-view features with an aggregation module directly in the canonical point cloud space and achieve higher mIoU (64.1) than SPLATNet (39.3) on the ScanNet benchmark.

D Instance Segmentation

The task of 3D instance segmentation is more precise than 3D object detection: Instead of regressing boxes, point masks which describe the exact shape of each object are predicted. Proposal based approaches like Mask R-CNN are the state-of-the-art in 2D and have been extended to 3D by leveraging voxels and 3D box-proposals . Alternatively to proposing boxes, point clouds can be generated as proposals . Another strategy is clustering based on predicted semantic labels or a similarity matrix which can be learned . We extend MVPNet to instance segmentation using R-PointNet .

MVPNet

Our MVPNet is designed to effectively fuse complementary information from multiple RGB-D frames and 3D point cloud in order to achieve better 3D scene understanding on real-world data, like ScanNetV2 . The primary task is 3D semantic segmentation, where the goal is to predict a semantic label for each point in the input point cloud. Our pipeline is illustrated in Fig. 2. We also showcase an extension to 3D instance segmentation in Sec. 4.5.

The data of each scene consists of a sequence of RGB-D frames and a point cloud. The input point cloud, denoted as Ssparse\mathcal{S}_{\text{sparse}}, is sparse compared with the resolution of images. This can be seen in Fig. 3 by comparing the density of the sparse point cloud with the unprojected views. Following PointNet++ , we divide the whole scene into chunks (around 90 chunks for an average scene). For each chunk, the most MM informative views (RGB-D frames) are selected to maximize the coverage of the input point cloud (Sec. 3.2). Those views (RGB) are then fed into a 2D encoder-decoder network in order to compute MM feature maps (Sec. 3.3). To augment the sparse input point cloud Ssparse\mathcal{S}_{\text{sparse}}, pixels with valid depth in each 2D feature map are first lifted to a 3D point cloud and then a dense point cloud Sdense\mathcal{S}_{\text{dense}} is obtained by concatenating all the MM unprojected point clouds. Given the image features associated with Sdense\mathcal{S}_{\text{dense}}, our feature aggregation module samples the kk nearest neighboring points in Sdense\mathcal{S}_{\text{dense}} and adaptively combines them to form the new feature for the point in Ssparse\mathcal{S}_{\text{sparse}} (Sec. 3.4). Finally, we leverage PointNet++ to process the multi-view feature augmented point cloud from a 3D geometric perspective.

2 View Selection

In ScanNetV2 , the RGB-D frames come as video stream with strong overlap between consecutive frames. It would be redundant and computationally expensive to process them all. Therefore, we make a selection of 1 to 5 views, which maximize contained information, to fuse with the point cloud of the scene.

In the preprocessing step, the overlaps between the scene point cloud and all the unprojected RGB-D frames of the video stream are computed. To reduce computation, we downsample the point cloud (red points in Fig. 3). During training we use the overlap information to select the RGB-D frames on-the-fly with a greedy algorithm. The image which overlaps with the most yet uncovered points is selected. We found that this straightforward but efficient method can achieve very high coverage even with few frames, leading to better results with same computation.

3 2D Encoder-Decoder Network

We feed the selected RGB images into a 2D encoder-decoder network based on U-Net to compute image feature maps. In our implementation, the size of the input image is equal to that of the output feature map, and fixed to 160×120160\times 120. With the relatively low resolution, we found UNet to be better suited in terms of memory, speed, and performance than other 2D semantic segmentation architectures such as DeepLabv3 , PSPNet , optimized for a much higher resolution. We pretrain the 2D encoder-decoder network on the task of 2D segmentation on ScanNetV2 in order to bootstrap the training of the whole pipeline. More details can be found in Sec. 4.2.

4 2D-3D Feature Lifting Module

In order to obtain the 3D coordinates for the feature maps that have been computed with the RGB images and the 2D encoder-decoder network, we unproject the corresponding depth maps using the camera instrinsics and poses. Consider MM 2D feature maps of size H×W×CfeatH\times W\times C_{\text{feat}}, then each one is lifted to a point cloud of size NRGB×CfeatN_{\text{RGB}}\times C_{\text{feat}}, where NRGB<HWN_{\text{RGB}}<HW is a hyperparameter that corresponds to the number of unprojected pixels in each RGB image. By concatenating all the MM unprojected points together, we yield a dense point cloud Sdense\mathcal{S}_{\text{dense}} of size MNRGB×CfeatMN_{\text{RGB}}\times C_{\text{feat}}.

For semantic segmentation, the labels have to be predicted for the input point cloud Ssparse\mathcal{S}_{\text{sparse}}. Thus, we have to transfer the features from the unprojected point cloud Sdense\mathcal{S}_{\text{dense}} to Ssparse\mathcal{S}_{\text{sparse}}. Therefore, we use our feature aggregation module which includes a shared MLP inspired by in order to distill a new feature for each point in Ssparse\mathcal{S}_{\text{sparse}} from its kk nearest neighbors in Sdense\mathcal{S}_{\text{dense}}

where hi\textbf{h}_{i} is the distilled feature at point xi\textbf{x}_{i} in Ssparse\mathcal{S}_{\text{sparse}}, fj\textbf{f}_{j} the semantic feature at one of the kk nearest neighbors points xj\textbf{x}_{j} in Sdense\mathcal{S}_{\text{dense}}, and fdist(xi,xj)\textbf{f}_{\text{dist}}(\textbf{x}_{i},\textbf{x}_{j}) the distance feature between the two points which we define as

We define multi-view feature augmented point cloud as the resulting features associated with 3D coordinates. Note that the whole 2D-3D feature lifting module is differentiable, which enables end-to-end training of our MVPNet.

5 3D Fusion Network

To fuse multi-view image features and geometry information, we employ PointNet++ as backbone. The original PointNet++ consumes both the coordinates and its corresponding features, such as normal or color. For 3D semantic segmentation, it encodes the input point cloud with set abstraction layers hierarchically, and decodes the output semantic prediction through feature propagation layers. The 3D coordinates of input points are concatenated to the output features of each set abstraction layer.

We adopt early fusion, where the image features are concatenated to the geometry (XYZ) and then given as input to PointNet++. Thus, the network is able to fully exploit the image features from a geometric perspective. We also investigated intermediate fusion and late fusion. In late fusion, the image features are concatenated after the final feature propagation layer in PointNet++, right before the segmentation head. In intermediate fusion, we introduce separate encoder branches for geometry and image features whose outputs are then concatenated and fed into the decoder. Additionally, the decoder leverages the intermediate outputs of two encoder branches through skip connections. The different fusion strategies are illustrated in Fig. 4.

Experiments

In this section we cover experiments on the ScanNetV2 dataset, but additional results on S3DIS can be found in the supplementary where we improve over previous methods by 4.16 mIoU.

The ScanNetV2 dataset features indoor scenes like offices and living rooms for which a total of 2.5M frames were captured with the internal camera of an IPad and an additionally mounted depth camera. The data for each scan consists of an RGB-D sequence with associated poses, a whole scene mesh, as well as semantic and instance labels. There are 1201 training and 312 validation scans that were taken in 706 different scenes, thus each scene was captured about 1 to 3 times. The test set contains 100 scans with hidden ground truth, used for the benchmark.

2 Implementation Details

For the task of 3D semantic segmentation, we follow the same chunk-wise pipeline as PointNet++ . During training, one chunk (1.5m×1.5m1.5m\times 1.5m in xy-plane, parallel to ground surface) is randomly selected from the whole scene if it contains more than 30% annotated points. Random rotation along the up-axis is applied for data augmentation. During testing, the network predicts all the chunks with a stride of 0.5m0.5m in a sliding-window fashion through the xy-plane. A majority vote is conducted for the points that have predictions from multiple chunks.

We downsample the images and depth maps to a resolution of 160×120160\times 120. Random horizontal flip is applied to augment images during training. We fix the number of unprojected points per RGB-D frame to 8192 of a total 19200 pixels for a resolution of 160×120160\times 120. Note that even though many pixels are not lifted to 3D, they are still essential for the 2D feature computation as they lie in the receptive field of unprojected pixels.

The backbone of the 2D Encoder network is an ImageNet-pretrained VGG16 with batch normalization and dropout. For ablation studies and submissions, we also experiment with VGG19 and ResNet34 . The custom 2D Decoder network is a lightweight variant of U-Net . Batch normalization and ReLU are added after each convolution layer in the decoder.

For each chunk, 8192 points are sampled from the input point cloud and augmented by the views selected by the method described in Sec. 3.2. For the feature aggregation module, we use a two-layer MLP with 128 and 64 channels. To predict the semantic labels for the multi-view feature augmented point cloud, we use PointNet++ with single-scale grouping (SSG) as our 3D backbone. However, note that our MVPNet can adapt to any 3D network.

Each epoch consists of 20000 randomly sampled chunks, and the batch size of chunks is 6. The network is trained with the SGD optimizer for 100 epochs. We use a weight decay of 0.0001 and a momentum of 0.9. The learning rate is 0.01 for the first 60 epochs, and then divided by 10 every 20 epochs. MVPNet is trained on a single GTX 1080Ti.

3 Results for 3D Semantic Segmentation

We evaluate our MVPNet on the ScanNetV2 3D semantic label benchmark. The evaluation metric is the average IoU (mIoU) over 20 classes. For submission, we ensemble 4 models of MVPNet, which consumes 5 views and uses ResNet34 as 2D backbone.

Tab. 1 shows our performance compared to the published state-of-the-art point cloud based methods on the test set. Our MVPNet outperforms all the published point cloud based methods, like PointConv and PointCNN , by a large margin. This confirms the effectiveness of our approach of elevating 2D image features to 3D for geometric fusion, especially for classes with flat shapes, i.e. refrigerator, picture, curtain and the like, which lack discriminative geometric cues for point cloud based networks. Qualitative results are shown in Fig. 6 and a failure case in Fig. 7.

Tab. 2 shows our performance compared to the published state-of-the-art voxel based methods on the test set. 3DMV is a joint 2D-3D network similar to ours, but in the voxel-based domain, and does not match our performance and inference time (500 s/scene vs. 3.35 s/scene).

Although there exists a gap between our result and SCN , our MVPNet is much more robust to low resolution point clouds as detailed in Sec. 4.4. This is relevant to robotic applications where, due to sensor limitations, point clouds are always much sparser than images. Furthermore, we compare with SCN in terms training/inference time and number of parameters in Tab. 3. It shows that MVPNet is comparable with highly engineered SCN. However, MVPNet is able to converge in 20 hours on a GTX 1080Ti (incl. pretrain of 2D encoder-decoder), while it takes 12 days for heavyweight SCN for the same GPU model, or 4 days even with a V100.

4 Robustness to Varying Point Cloud Density

Real-world 3D sensors such as lidars and depth cameras have much lower resolution than 2D RGB cameras and the point cloud density also varies with view point angle, lighting conditions, object-to-sensor distance and object surface reflectivity. Thus, it is important for algorithms to be robust to varying point cloud density at test time, because training data can hardly cover all cases. To examine robustness to sparsity, we uniformly subsample the whole scene point cloud and feed it to the networks that were trained on full resolution. We report the results in Fig. 5. While our MVPNet is hardly affected at lower resolutions, the performance of SparseConvNet (SCN) suffers severely.

We attribute the performance difference mainly to two factors. First, the quality of our image features is not deteriorated when the point cloud is downsampled, because they are computed in the original dense 2D image. During 2D-3D lifting, even if few unprojected image pixels are finally used at coarse point cloud resolution, the image features can maintain their quality thanks to the receptive field of the 2D encoder-decoder. This is not the case for SCN where only the sparse RGB information at each 3D point is used. The second factor is related to the different neighborhood definition in voxel grids and point clouds. Voxel-based methods such as SCN have fixed neighbors defined by the discrete grid. Point-based methods on the other hand, use continuous locations and in each network layer neighbors are sampled, e.g. with ball query, that adapt naturally to the local point distribution. This enables our PointNet++ based approach to cope well with varying point cloud density.

5 Extension to 3D Instance Segmentation

In order to assess the ability of our MVPNet to be applied to a different task, we extend it to 3D Instance Segmentation on ScanNetV2. We use trained model of MVPNet for semantic segmentation to predict all scenes of the train/val/test set and save the features of the last layer before the segmentation head to disk. We modified R-PointNet in order to take the semantic features from MVPNet as input and yield a significant improvement from 38.8 to 47.1 mAP on the validation set which demonstrates the versatility of MVPNet.

Ablation Studies

To analyze our design choices and provide more insights, we conduct ablation studies on the validation set of ScanNetV2 and report the average IoU (mIoU) over 20 classes. The 2D encoder-decoder network is frozen in order to accelerate training, since we observe no significant improvement with end-to-end training. Unless stated otherwise, our 2D backbone is VGG16 and our 3D backbone contains 4 set abstraction and 4 feature propagation layers. The numbers of centroids are 1024, 256, 64, 16 respectively.

We define coverage as the ratio of points in the input point cloud that have at least one unprojected neighbor point with image features at a distance less then 0.1m. Tab. 4 shows how the number of views affects coverage and mIoU. We removed the feature aggregation module for this experiment. With 1 view the coverage reaches 68.1%, and with 3 frames already exceeds 90%. More views lead to higher mIoU, but introduce more computation. We choose 3 frames as default in this trade-off.

2 Feature Aggregation Module

In the following we study the parameters of our Feature Aggregation Module, defined in eq. 1, which distills features from the unprojected point cloud Sdense\mathcal{S}_{\text{dense}}, obtained from multiple views, into the input point cloud Ssparse\mathcal{S}_{\text{sparse}}. We report our results in Tab. 5 for 1 and 3 views, 1 or 3 nearest neighbors feature sampling, with or without MLP, and we also try maximum instead of sum as feature pooling function.

For the 1-view case, using 3 nearest neighbors instead of only 1 increases the performance by at least 0.8 mIoU. Due to the limited coverage of a single view, far-away image features are sampled for uncovered points and multiple neighbors might alleviate this problem by analyzing feature consistency between them.

For the 3-view case, we find that the number of nearest neighbors does not affect performance. This might be because coverage is already very high (92.9%) which means that, as opposed to the 1-view case, features can always be sampled from close-by. We also try summation instead of maximum to pool the features which does not significantly change the results. Using an MLP can slightly improve, maybe because it can transform 2D image features to an embedding space more consistent with the 3D representation. Our final choice for all other experiments is 3 nearest neighbors with MLP and sum aggregation.

3 Fusion

In this section we want to answer the question how to best fuse geometry and image features with point cloud based networks and give an insight about the strength of each modality. In Tab. 6 we report our quantitative results on the validation set. Our PointNet++ baseline yields 54.5 mIoU with XYZ only and 57.8 mIoU with additional color information.

In order to assess the strength of multi-view vs. geometry features, we conduct an experiment with 3 views where we unproject the output semantic labels of the pretrained 2D encoder-decoder to 3D and attribute the nearest neighbor 2D label to each 3D point in the input point cloud. This multi-view 2D CNN approach can already achieve similar performance as PointNet++ on colored point clouds, which confirms the benefit of features computed on dense 2D images before 2D-3D lifting.

Next, we study three fusion strategies introduced in Sec. 3.5. We could yield slightly better performance with the late fusion approach (+1.2 mIoU) than with the 2D CNN baseline. Intermediate fusion leads to much better results (+7.6 mIoU) than late fusion. Early fusion can reach the best score (+7.8 mIoU), and uses less parameters and computation compared to intermediate fusion. The observation is different from the voxel-based method 3DMV, where geometric features and image features are concatenated late, at roughly 2/3 in the network.

Moreover, we investigate whether it is necessary to add geometric features (XYZ coordinates) in the early fusion approach or if the image features are sufficient. In fact, PointNet++ already induces a geometric hierarchy and the dimension of the XYZ coordinates is much smaller compared to that of the image features (3 vs. 64). Nonetheless, the obtained result (-2.2 mIoU without XYZ) proves the contrary and indicates that MVPNet actually benefits from geometric features as complementary information to images.

4 Stronger backbone

To investigate the effect of stronger 2D backbones, we replace VGG16 with VGG19 and ResNet34. Tab. 7 shows the results of 2D CNN baselines and our MVPNet with different backbones on the validation set. Intuitively, stronger backbones lead to higher mIoU, and ResNet34 performs best. Due to runtime performance we choose VGG16 as backbone for ablation experiments and ResNet34 for best performance.

As to the stronger 3D backbone, we double the numbers of sampled centroids to (2048, 512, 128, 32), which increases the mIoU by 1.4. We also try to replace single-scale with multi-scale grouping (MSG) in PointNet++, but observe no improvement. As the image features already contain contextual information, it might not be necessary to process multiple scales explicitly with the MSG version.

Conclusion

While we can outperform state-of-the-art point based approaches by a significant margin using sliding window processing, methods that take the whole scene as input, e.g. the very well implemented SCN , have a clear advantage.

We have proposed a framework to fuse 2D multi-view images and 3D point clouds in an effective way by computing image features in 2D first, lifting them to 3D, and then fuse complementary geometry and image information in canonical 3D space. Comprehensive experiments are conducted on the ScanNetV2 Semantic Segmentation benchmark, which prove the advantage of calculating image features from multi-view images, and verify the superior robustness of our approach against voxel-based methods.

Acknowledgements

We thank Li Yi and Wang Zhao to have provided us with the code of R-PointNet. We thank Valeo to have supported Maximilian Jaritz for his visit at UC San Diego.

References

Appendix A 2D Encoder Decoder Architecture

Fig. 8 illustrates the architecture of the 2D encoder-decoder network inspired by U-Net . We use VGG16 as encoder and initialize with ImageNet pre-trained weights. In the decoder, convolution is used to fuse concatenated features from skip connections, and transposed convolution for upsampling.

Appendix B Comparison with SparseConvNet (SCN)

For lightweight SCN (small U-Net with 5cm-cubed voxels), we refer readers to the released codesgithub.com/facebookresearch/SparseConvNet.

Appendix C Experiments on S3DIS

We evaluate on Stanford Indoor 3D (S3DIS) using images and xyx-maps from 2D-3D Semantics (2D-3D-S) . Table 8 shows that MVPNet improves over previous methods by 4.16 mIoU.

Appendix D More Ablation Studies

For the ScanNetV2 3D semantic label benchmark, we employ MVPNet with 5 views and use ResNet34 as the 2D backbone. The numbers of centroids are 2048, 512, 128, 64 respectively.

Tab. 9 shows the comparison among several variants of the submission version. The stronger 2D backbone (ResNet34) improves the mIoU by 0.7 against the weaker 2D backbone (VGG19). Moreover, we also experiment with training MVPNet with class weights, which boosts the mIoU (+0.7) as the evaluation metric favors more balanced predictions. To achieve the best performance (68.3), we ensemble 4 models of MVPNet with ResNet34.