Deep Parametric Continuous Convolutional Neural Networks
Shenlong Wang, Simon Suo, Wei-Chiu Ma, Andrei Pokrovsky, Raquel Urtasun
Introduction
Discrete convolutions are the most fundamental building block of modern deep learning architectures. Its efficiency and effectiveness relies on the fact that the data appears naturally in a dense grid structure (e.g., 2D grid for images, 3D grid for videos). However, many real world applications such as visual perception from 3D point clouds, mesh registration and non-rigid shape correspondences rely on making statistical predictions from non-grid structured data. Unfortunately, standard convolutional operators cannot be directly applied in these cases.
Multiple approaches have been proposed to handle non-grid structured data. The simplest approach is to voxelize the space to form a grid where standard discrete convolutions can be performed . However, most of the volume is typically empty, and thus this results in both memory inefficiency and wasted computation. Geometric deep learning and graph neural network approaches exploit the graph structure of the data and model the relationship between nodes. Information is then propagated through the graph edges. However, they either have difficulties generalizing well or require strong feature representations as input to perform competitively. End-to-end learning is typically performed via back-propagation through time, but it is difficult to learn very deep networks due to the memory limitations of modern GPUs.
In contrast to the aforementioned approaches, in this paper we propose a new learnable operator, which we call parametric continuous convolution. The key idea is a parameterized kernel function that spans the full continuous vector space. In this way, it can handle arbitrary data structures as long as its support relationship is computable. This is a natural extension since objects in the real-world such as point clouds captured from 3D sensors are distributed unevenly in continuous domain. Based upon this we build a new family of deep neural networks that can be applied on generic non-grid structured data. The proposed networks are both expressive and memory efficient.
We demonstrate the effectiveness of our approach in both semantic labeling and motion estimation of point clouds. Most importantly, we show that very deep networks can be learned over raw point clouds in an end-to-end manner. Our experiments show that the proposed approach outperforms the state-of-the-art by a large margin in both outdoor and indoor 3D point cloud segmentation tasks, as well as lidar motion estimation in driving scenes. Importantly, our outdoor semantic labeling and lidar flow experiments are conducted on a very large scale dataset, containing 223 billion points captured by a 3D sensor mounted on the roof of a self-driving car. To our knowledge, this is 2 orders of magnitude larger than any existing benchmark.
Related Work
Deep learning approaches that exploit 3D geometric data have recently become populer in the computer vision community. Early approaches convert the 3D data into a two-dimensional RGB + depth image and exploit conventional convolutional neural networks (CNNs). Unfortunately, this representation does not capture the true geometric relationships between 3D points (i.e. neighboring pixels could be potentially far away geometrically). Another popular approach is to conduct 3D convolutions over volumetric representations . Voxelization is employed to convert point clouds into a 3D grid that encodes the geometric information. These approaches have been popular in medical imaging and indoor scene understanding, where the volume is relatively small. However, typical voxelization approaches sacrifice precision and the 3D volumetric representation is not memory efficient. Sparse convolutions and advanced data structures such as oct-trees have been used to overcome these difficulties. Learning directly over point clouds has only been studied very recently. The pioneer work of PointNet , learns an MLP over individual points and aggregates global information using pooling. PointNet++ , the follow-up, improves the ability to capture local structures through a multi-scale grouping strategy.
Graph Neural Networks:
Graph neural networks (GNNs) are generalizations of neural networks to graph structured data. Early approaches apply neural networks either over the hidden representation of each node or the messages passed between adjacent nodes in the graph, and use back-propagation through time to conduct learning. Gated graph neural networks (GGNNs) exploit gated recurrent units along with modern optimization techniques, resulting in improved performance. In , GGNNs are applied to point cloud segmentation, achieving significant improvements over the state-of-the-art. One of the major difficulties of graph neural networks is that propagation is conducted in a synchronous manner and thus it is hard to scale up to graphs with millions of nodes. Inference in graphical models as well as recurrent neural networks can be seen as special cases of graph neural networks.
Graph Convolution Networks:
An alternative formulation is to learn convolution operations over graphs. These methods can be categorized into spectral and spatial approaches depending on which domain the convolutions are applied to. For spectral methods, convolutions are converted to multiplication by computing the graph Laplacian in Fourier domain . Parameterized spectral filters can be incorporated to reduce overfitting . These methods are not feasible for large scale data due to the expensive computation, since there is no FFT-like trick over generic graph. Spatial approaches directly propagate information along the node neighborhoods in the graph. This can be implemented either through low-order approximation of spectral filtering, or diffusion in a support domain . Our approach generalizes spatial approaches in two ways: first, we use more expressive convolutional kernel functions; second, the output of the convolution could be any point in the whole continuous domain.
Other Approaches:
Edge-conditioned filter networks use a weighting network to communicate between adjacent nodes on the graph conditioned on edge labels, which is primarily formulated as relative point locations. In contrast, our approach is not constrained to a fixed graph structure, and has the flexibility to output features at arbitrary points over the continuous domain. In a concurrent work, uses similar parametric function form to aggregate information between points. However, they only use shallow isotropic gaussian kernels to represent the weights, while we use expressive deep networks to parameterize the continuous filters.
Deep Parametric Continuous CNNs
Standard CNNs use discrete convolutions (i.e., convolutions defined over discrete domain) as basic operations.
In contrast, continuous convolutions can be defined as
Continuous convolutions require the integration in Eq. (1) to be analytically tractable. Unfortunately, this is not possible for real-world applications, where the input features are complicated and non-parametric, and the observations are sparse points sampled over the continuous domain.
Motivated by monte-carlo integration we derive our continuous convolution operator. In particular, given continuous functions and with a finite number of input points sampled from the domain, the convolution at an arbitrary point can be approximated as:
2 From Convolutions to Deep Networks
In this section, we first design a new convolution layer based on the parametric continuous convolutions derived in the previous subsection. We then propose a deep learning architecture using this new convolution layer.
Let be the number of input points, be the number of output points, and the dimensionality of the support domain. Let and be predefined input and output feature dimensions respectively. Note that these are hyperparameters of the continuous convolution layer analogous to input and output feature dimensions in standard grid convolution layers. Fig. 1 depicts our parametric continuous convolutions in comparison with conventional grid convolution. Two major differences are highlighted: 1) the kernel function is continuous given the relative location in support domain; 2) the input/ouput points could be any points in the continuous domain as well and can be different.
Deep Parametric Continuous CNNs:
Using the parametric continuous convolution layers as building blocks, we can construct a new family of deep networks which operates on unstructured data defined in a topological group under addition. In the following discussions, we will focus on multi-diumensional euclidean space, and note that this is a special case. The network takes the input features and their associated positions in the support domain as input. Then the hidden representations are generated from successive parametric continuous convolution layers. Following standard CNN architectures, we can add batch normalization, non-linearities and residual connections between layers. Pooling can also be employed over the support domain to aggregate information. In practice, we find adding residual connection between parametric continuous convolution layers is critical to help convergence. Please refer to Fig. 2 for an example of the computation graph of a single layer, and to Fig. 3 for an example of the network architecture employed for our indoor semantic segmentation task.
Learning:
All of our building blocks are differentiable, thus our networks can be learned through back-prop:
3 Discussions
Standard grid convolution are computed over a limited kernel size to keep locality. Similarly, locality can be enforced in our parametric continuous convolutions by constraining the influence of the function to points close to , i.e.,
Efficient Continuous Convolution:
Special Cases:
Many previous convolutional layers are special cases of our approach. For instance, if the points are sampled over the finite 2D grid we recover conventional 2D convolutions. If the support domain is defined as concatenation of the spatial vector and feature vector with a gaussian kernel , we recover the bilateral filter. If the support domain is defined as the neighboring vertices of a node we recover the first-order spatial graph convolution .
Experimental Evaluation
We demonstrate the effectiveness of our approach in the tasks of semantic labeling and motion estimation of 3D point clouds, and show state-of-the-art performance. We conduct point-wise semantic labeling experiments over two datasets: a very large-scale outdoor lidar semantic segmentation dataset that we collected and labeled in house and a large indoor semantic labeling dataset. To our knowledge, these are the largest real-world outdoor and indoor datasets that are available for this task. The datasets are fully labeled and contain 137 billion and 629 million points respectively. The lidar flow experiment is also conducted on this dataset with ground-truth 3D motion label for each point.
We use the Stanford large-scale 3D indoor scene dataset and follow the training and testing procedure used in . We report the same metrics, i.e., mean-IOU, mean class accuracy (TP / (TP + FN)) and class-wise IOU. The input is six dimensional and is composed of the xyz coordinates and RGB color intensity. Each point is labeled with one of 13 classes shown in Tab. 1.
Competing Algorithms:
We compare our approach to PointNet and SegCloud . We evaluate the proposed end-to-end continuous convnet with eight continuous convolution layers (Ours PCCN). The kernels are defined over the continuous support domain of 3D Euclidean space. Each intermediate layer except the last has 32 dimensional hidden features followed by batchnorm and ReLU nonlinearity. The dimension of the last layer is 128. We observe that the distribution of semantic labels within a room is highly correlated with the room type (e.g. office, hallway, conference room, etc.). Motivated by this, we apply max pooling over all the points in the last layer to obtain a global feature, which is then concatenated to the output feature of each points in the last layer, resulting in a 256 dimensional feature. A fully connected layer with softmax activation is used to produce the final logits. Our network is trained end-to-end with cross entropy loss, using Adam optimizer.
Results:
As shown in Tab. 1 our approach outperforms the state-of-the-art by 9.3% mIOU and 9.6% mACC. Fig. 4 shows qualitative results. Despite the diversity of geometric structures, our approach works very well. Confusion mainly occurs between columns vs walls and window vs bookcase. It is also worth noting that our approach captures visual information encoded in RGB channels. The last row shows two failure cases. In the first one, the door in the washroom is labeled as clutter whearas our algorithm thinks is door. In the second one, the board on the right has a window-like texture, which makes the algorithm predict the wrong label.
2 Semantic Segmentation of Driving Scenes
We first conduct experiments on the task of point cloud segmentation in the context of autonomous driving. Each point cloud is produced by a full sweep of a roof-mounted Velodyne-64 lidar sensor driving in several cities in North America. The dataset is composed of snippets each having 300 consecutive frames. The training and validation set contains 11,337 snippets in total while the test set contains 1,644 snippets. We report metrics on a subset of the test set which is generated by sampling 10 frames from each snippet to avoid bias brought due to scenes where the ego-car is static (e.g., when waiting at a traffic light). Each point is labeled with one of seven classes defined in Tab. 2. We adopt mean intersection-over-union (meanIOU) and point-wise accuracy (pointAcc) as our evaluation metrics.
Baselines:
We compare our approach to the point cloud segmentation network (PointNet) and a 3D fully convolutional network (3D-FCN) conducted over a 3D occupancy grid. We use a resolution of 0.2m for each voxel over a 160mx80mx6.4m range. This results in an occupancy grid encoded as a tensor of size 800x400x32. We define a voxel to be occupied if it contains at least one point. We use ResNet-50 as the backbone and replace the last average pooling and fully connected layer with two fully convolutional layers and a trilinear upsampling layer to obtain dense voxel predictions. The model is trained from scratch with the Adam optimizer to minimize the class-reweighted cross-entropy loss. Finally, the voxel-wise predictions are mapped back to the original points and metrics are computed over points. We adapted the open-sourced PointNet model onto our dataset and trained from scratch. The architecture and loss function remain the same with the original paper, except that we removed the point rotation layer since it negatively impacts validation performance on this dataset.
Our Approaches:
Results:
As shown in Tab. 2, by exploiting sophisticated feature via 3D convolutions, 3D-FCN+PCCN results in the best performance. Fig. 5 shows qualitative comparison between models. As shown in the figure, all models produce good results. Performance differences often result from ambiguous regions. In particular, we can see that the 3D-FCN model oversegements the scene: it mislabels a background pole as vehicle (red above egocar), nearby spurirous points as bicyclist (green above egocar), and a wall as pedestrian (purple near left edge). This is reflected in the confidence map (as bright regions). We observe a significant improvement in our 3D-CNN + PCCN model, with all of the above corrected with high confidence. For more results and videos please refer to the supplementary material.
Model Sizes:
We also compare the model sizes of the competing algorithms in Tab. 2. In comparison to the 3D-FCN approach, the end-to-end continuous convolution network’s model size is eight times smaller , while achieving comparable results. And the 3D-FCN+PCCN is just 0.01MB larger than 3D-FCN, but the performance is improved by a large margin in terms of mean IOU.
Complexity and Runtime
We benchmark the proposed model’s runtime over a GTX 1080 Ti GPU and Xeon E5-2687W CPU with 32 GB Memory. The forward pass of a 8-layer PCCN model (32 feature dim in each layer with 50 neighbours) takes 33ms. The KD-Tree neighbour search takes 28 ms. The end-to-end computation takes 61ms. The number of operations of each layer is 1.32GFLOPs.
Generalization:
To demonstrate the generalization ability of our approach, we evaluate our model, trained with only North American scenes, on the KITTI dataset , which was captured in Europe. As shown in Fig. 6, the model achieves good results, with well segmented dynamic objects, such as vehicles and pedestrians.
3 Lidar Flow
We also validate our proposed method over the task of lidar based motion estimation, refered to as lidar flow. In this task, the input is two consecutive frames of lidar sweep. The goal is to estimation the 3D motion field for each point in the first frame, to undo both ego-motion and the motion of dynamic objects. The ground-truth ego-motion is computed through a comprehensive filters that take GPS, IMU as well as ICP based lidar alignment against pre-scaned 3D geometry of the scene as input. And the ground-truth 6DOF dynamics object motion is estimated from the temporal coherent 3D object tracklet, labeled by in-house annotators. Combining both we are able to get the ground-truth motion field. Fig. 7 shows the colormapped flow field and the overlay between two frames after undoing per-point motion. This task is crucial for many applications, such as multi-rigid transform alignment, object tracking, global pose estimation, etc. The training and validation set contains 11,337 snippets while the test set contains 1,644 snippets. We use 110k frame pairs for training and validation, and 16440 frame pairs for testing. End-point error, and outlier percentage at 10 cm and 20 cm are used as metric.
Competing Algorithms:
We compare against the 3D-FCN baseline using the same architecture and volumetric representation as used in Sec. 4.2. We also adopt a similar 3D-FCN + PCCN architecture with 7 residual continuous convolution layers added as a polishing network. In this task, we remove the ReLU nonlinearity and supervise the PCCN layers with MSE loss at every layer. The training objective function is mean square error loss between the ground-truth flow vector and the prediction.
Results:
Tab. 3 reports the quantitative results. As shown in the table, our 3D-FCN+PCCN model outperforms the 3D-FCN by 0.351cm in end-point error and our method reduces approximately of the outliers. Fig. 18 shows sample flow predictions compared with ground truth labels. As shown in the figure, our algorithm is able to capture both global motion of the ego-car including self rotation, and the motion of each dynamic objects in the scene. For more results please refer to our supplementary material.
Conclusions
We have presented a new learnable convolution layer that operates over non-grid structured data. Our convolution kernel function is parameterized by multi-layer perceptrons and spans the full continuous domain. This allows us to design a new deep learning architecture that can be applied to arbitrary structured data, as long as the support relationships between elements are computable. We validate the performance on point cloud segmentation and motion estimation tasks, over very large-scale datasets with up to 200 bilion points. The proposed network achieves state-of-the-art performance on all the tasks and datasets.
References
Appendix A Generalization
We show the generalization ability of our proposed model by training over one dataset and test it over another in our supplementary video. To be specific, we have used the following configurations:
Train our proposed semantic labeling network on the driving scene data (several north America cities), and test it on KITTI (Europe).
Train our proposed semantic labeling network on the driving scene data (non-highway road), and test it on a lidar sequence mounted on top of a high-way driving truck (highway road).
Train our proposed lidar flow network on the driving scene data (several north America cities), and test it on KITTI (Europe).
Under all the settings, our algorithm is able to generalize well. Fig. 9 shows our lidar flow model’s performance on KITTI. As shown in Fig. 9, our Lidar flow model generalizes well to the unseen KITTI dataset. From left to right, the figures show the most common scenarios with moving, turning, and stationary ego-car. The model produces covincing flow predictions in all three cases. Fig. 10 includes some additional segmentation results on KITTI.
For more results over several sequences, please refer to our supplementary video.
Appendix B Lidar Flow Data and Analysis
Please refer to Fig. 11 for an illustration.
Ground-truth Motion Analysis
We also conduct an analysis over the ground-truth motion distribution. In Fig. 12 we show the 2D histogram of the GT 3D translation component along and axis respectively. We also show the motion distribution across different object types, e.g. static background, vehicle and pedestrian. As we can see, different semantic types have different motion patterns. And the heaviest density of distribution is on the y-axis, which suggests the forward motion is the major motion pattern of our ego-car.
Ground-truth Validation and Visualization
Appendix C More Results
In this section, we show additional qualitative results of the proposed algorithm over all the tasks.
Fig. 15 and Fig. 16 show more qualitatitive results over the stanford dataset. As the figure shown, in most cases our model is able to predict the semantic labels correctly.
C.2 Semantic Segmentation for Driving Scenes
Fig. 17 shows additional results for semantic labeling in driving scenes. As shown, the results capture very small dynamics, e.g. pedestrians and bicyclists. This suggests our model’s potential in object detection and tracking. The model is also able to distinguish between road and non-road through lidar intensity and subtle geometry structure such as road curbs. This validates our model’s potential in map automation.
More specifically, we see that most error occur on road boundaries (bright curves in error map).
C.3 Lidar Flow
We show additional results on Lidar flow estimation in Fig. 18. Unlike the visualization in the main submission, we visualize the colored vector in order to better depicts the magnitudes of the motion vector. As shown in the figure, our model is able to capture majority flow field. The majority of the error happens at the object boundary. This suggests that a better support domain that includes both space and intensity features could be potentially used to boost performance.
Appendix D Activations
We visualize the activation maps of the trained PCCN network over a single lidar frame from the driving scene dataset for segmentation. Fig. 19 depicts the activation map at layer 1 of PCCN. As we can see, at the early conv layer the method mainly captures low-level geometry details, e.g. the z coordinate, the intensity peak, etc. Fig. 20 shows the activation map at layer 8 of PCCN. The conv layers begin to capture information with more semantic meaning, e.g. the road curb and the dynamic objects.
Appendix E Point Cloud Classification
To verify the applicability of the proposed parameteric continuous convolution over global prediction task, we conduct a simple point cloud classification task on the ModelNet40 benchmark. This dataset contains CAD models from 40 categories. The state-of-the-art and most representative algorithms conducted on ModelNet40 are compared . We randomly sampled 2048 points for each training and testing sample over the 3D meshes and feed the point cloud into our neural network. The architecture contains 6 continuous convolution layers with 32-dimensional hidden features, followed by two layers with 128-dimensions and 512 dimensions respectively. The output of the last continuous convolution layer is fed into a max pooling layer to generate the global 512-dimensional feature, followed by two fc layers to output the final logits. Tab. 4 reports the classification performance. As we can see in the table, the performance is comparable with PointNet and slightly below PointNet++. Here we use a naive global max pooling to aggregate global information for our method. We expect to achieve better results with more comprehensive and hierachical pooling strategies.