Distinctive 3D local deep descriptors
Fabio Poiesi, Davide Boscaini
I Introduction
Encoding local 3D geometric information (e.g. coordinates, normals) into compact descriptors is key for shape retrieval , face recognition , object recognition and rigid (six degrees-of-freedom) registration . Learning such encoding from examples using deep neural networks has outperformed hand-crafted methods . These approaches have been designed to encode geometric information either from meshes or from point clouds . Our method belongs to the latter category and can be used to rigidly register point clouds without requiring an initial alignment (Fig. 1).
Existing solutions to compute compact 3D descriptors can be categorised into one-stage and two-stage methods. Although both categories share the objective of making descriptors invariant to point-cloud rigid transformations, one-stage methods encode the local geometric information of a patch (collection of locally sampled points) using the points of the patch directly. Differently, two-stage methods firstly estimate a local reference frame (LRF) from the points within the patch to rigidly transform the patch to a canonical frame, then they encode the information of the canonicalised points into a compact descriptor. Given two corresponding patches of two non-aligned point clouds, if we canonicalise them through their respective LRFs we should obtain two identical overlapping patches. Hence, the encoding method should be simpler to design than that of one-stage methods. However, noise, occlusions and point clouds reconstructed with different sensors make the LRF estimation challenging . Modelling descriptors to be robust to different sensors (e.g. RGB-D, laser scanner) and to different environments (e.g. indoor, outdoor) is also a challenge . One-stage learning-based methods can achieve rotation invariance by encoding the local geometric information with a set of geometric relationships, such as points, normals and point pair features , and then by learning descriptors via a PointNet-based deep network in order to achieve permutation invariance with respect to the set of the input points . Alternatively, 3D convolutional neural networks (ConvNets) can be used locally, to process patches around interest points , or globally, to process whole point clouds . In two-step methods, LRFs can be computed with hand-crafted or learning-based methods. After LRF canonicalisation, points can be transformed into a voxel grid, where each voxel encodes the density of the points within . Then, descriptors can be encoded from this voxel representations using a 3D ConvNet learnt with a Siamese approach .
In this paper we present a novel two-stage method where compact descriptors are learnt end-to-end from canonicalised patches. To mitigate the problem of incorrectly estimated LRFs, we learn an affine transformation that refines the canonicalisation operation by minimising the Euclidean distance between points through the Chamfer loss . Similarly to , we learn descriptors with a PointNet-based deep neural network through a Siamese approach, but differently from (i) we use LRFs to canonicalise patches, (ii) our descriptors encode local information only, thus promoting robustness to clutter, occlusions, and missing regions, and (iii) we use a hardest contrastive loss to mine for quadruplets , thus improving metric learning. Differently from and , points are consumed directly by our network without adding augmented hand-crafted features or performing prior voxelisations. We train our network using the 3DMatch dataset that consists of indoor scenes reconstructed with RGB-D sensors . We achieve state-of-the-art results on the 3DMatch test set and on its augmented version, namely 3DMatchRotated , employed to assess descriptor rotation invariance. We significantly outperform existing approaches in terms of generalisation ability to different sensor modalities (RGB-D laser scanner) and to different environments (indoor bedrooms outdoor forest) using the ETH dataset . Moreover, we validate DIP generalisation ability to another sensor modality (RGB-D smartphone) by capturing three overlapping indoor point clouds with the Visual-SLAM system of an ARCore-based App we have developed to reconstruct the environment. Notably, DIPs can successfully and robustly be used also to align these point clouds. The source code and the reconstruction App are publicly available.
II Our approach
Fig. 2 shows our PointNet-based architecture , where the three main modifications that allow us to produce DIPs are in the Transformation Network, the Bottleneck and the Local Response Normalisation layer.
Next, let us take two corresponding patches and extracted from two overlapping point clouds and , and then compute , and , respectively. () will be the same regardless of the permutations of the points in (). We observed that the corresponding values of and can be interpreted as the correspondences between points in and . Accordingly, their corresponding max values and quantify how reliable these correspondences are. Then, we found that the norm of can be effectively used to quantify the reliability of , e.g. to lower the importance of, or discard, patches extracted from flat surfaces. It turns out that good descriptors can be selected imposing the condition
Fig. 3 shows an example of global signatures computed from two pairs of corresponding patches (green) that are extracted from two overlapping point clouds (blue and grey) from the 3DMatch dataset .
The first case shows two patches extracted from a flat surface (wall), whereas the patches in the second case are extracted from more structured surfaces (bed). From each patch we randomly sample 256 points and pass them through the network to obtain their respective global signatures (Eq. 1). This figure shows the correspondences between the points of the corresponding patches such that , where . There are a few things we can observe in this example. First, values of are on average higher when the patches are extracted on structured surfaces. Second, the percentage of correspondences above the threshold is larger when the patches are extracted on structured surfaces ( vs. ). Lastly, we can see that the patches extracted on flat surfaces have lower . Fig. 4 shows the distribution of the values for 20K patches randomly sampled from three point clouds. We can see that low values of are distributed on flat surfaces (poor information) and along borders (incomplete information). Differently, has higher value near corners and on objects.
i.e. the area under the probability density function to the left of is .
Metric layers and Local Response Normalisation After max pooling, is processed by a series of MLP layers acting as metric layers to learn distinctive embeddings for our descriptors. We use a Local Response Normalisation (LRN) layer to produce unitary-length descriptors as we found it works well in practice . LRN consists of a L2 normalisation of the last MLP layer’s -dimensional output.
II-B Loss functions
The objective of our training is to produce descriptors whose reciprocal distance in the embedding space is minimised for corresponding patches of different point clouds. To this end, we train our network following a Siamese approach that processes pairs of corresponding descriptors using two branches with shared weights . Each branch independently calculates a descriptor for a given patch. We learn the parameters of the network by minimising the linear combination of two losses, aiming at two different goals. The first goal is to geometrically align two patches under the learnt affine transformation. The second goal is to produce compact and distinctive descriptors via metric learning.
Chamfer loss Given two patches , we want to minimise the distance between each point and its nearest neighbour . Therefore we use the Chamfer loss on the output of TNet as
Hardest-contrastive loss Our metric learning is performed through negative mining using the hardest-contrastive loss . Given a pair of anchors , we mine the hardest-negatives , and define the loss as
where is the set of the anchor pairs and is the set of descriptors (opportunely sampled) used for the hardest-negative mining extracted from a minibatch. and are the margins for positive and negative pairs, respectively. takes the positive part of its argument.
III Experimental validation
We evaluate the distinctiveness of DIPs using the indoor 3DMatch dataset , and assess DIP generalisation on the outdoor ETH dataset and on a new indoor dataset we collected with a smartphone. We explain how patches are extracted and given as input to our deep network. The training pipeline is shown in Fig. 5. Our method is developed in Pytorch 1.3.1 . We compare our method with 14 state-of-the-art methods and carry out a thorough ablation study.
DIPs are learnt from patches that are extracted from point cloud pairs whose overlap region is greater than a threshold . Let and be the overlap regions. During training we know the ground-truth transformation that register to . Point correspondences between and can be determined either by using the 3DMatch toolbox , or by using a nearest neighbourhood search after applying to . We use the latter approach by seeking nearest points from to within a radius of . Corresponding points in and are the candidate anchors used by the hardest-contrastive loss (Eq. 6).
Farthest Point Sampling Anchor sampling is key to allow for an effective minimisation of Eq. 6, and typically this is carried out with random sampling . Such random sampling may lead to cases where anchors and negatives are sampled spatially close to each other. A solution can be disregarding negatives within a certain radius from an anchor by computing the Euclidean distances amongst all the anchors within a minibatch in order to determine whether to penalise for the distance between descriptors in the embedding space (Eq. 5 in ). Including points within the radius would force the network to learn distinctive descriptors of region with similar geometric structures, thus making training unstable. Therefore to avoid computing the Euclidean distances amongst all the anchors within the minibatch , we efficiently sample anchors having the largest distance amongst themselves using Farthest Point Sampling (FPS) . Specifically, we sample points within using FPS and then search for the nearest neighbour counterparts in . These points are the anchors that construct the minibatch on which the hardest-negative mining is performed. In our experiments we use .
where the operation defines the application of to each element of such that . Analogously, the same operations are performed for .
III-B Datasets, training and testing setup
We use Stochastic Gradient Descent with an initial learning rate of that decreases by a factor every 15 epochs. The eight test scenes consists of 1117 point cloud pairs. As in , testing is performed by randomly sampling 5K points from each point cloud. To evaluate DIP’s rotation invariance ability, we follow the evaluation of and create an augmented version of 3DMatch, namely 3DMatchRotated: each point cloud is rotated by an angle sampled uniformly between around all the three axes independently. Unless otherwise stated we use .
ETH dataset We use the ETH dataset to assess the ability of DIPs to generalise across sensor modalities (RGB-D laser scanner) and on different scenes (indoor outdoor) . To this end we use the same model trained on the 3DMatch dataset (no fine tuning). The ETH dataset consists of four outdoor scenes, namely Gazebo-Summer, Gazebo-Winter, Wood-Summer and Wood-Autumn, containing partially overlapping, sparse and dense vegetation point clouds. Differently from the 3DMatch dataset we subsample point clouds using a voxel size of m. We set the patch kernel size m. For a fair comparison, the evaluation procedure follows verbatim , i.e. random sampling of 5K points.
VigoHome dataset To evaluate DIPs on another sensor modality (RGB-D smartphone RGB), we have created a new dataset, namely VigoHome, by reconstructing the inside of a house using a Visual-SLAM smartphone App we developed with ARCore (Android) . We captured three zones, namely livingroom-downstairs (94K points), bedroom-upstairs (43K points), and bathroom-upstairs (83K points). The stairs between the three zones is the overlapping region of the point clouds. We calculated their transformations to a common reference frame and determined the point correspondences to evaluate the registration: two points of a point cloud pair are corresponding if they are nearest neighbours within a 0.1m-radius. We subsample point clouds using voxels of 0.01m and set m.
III-C Comparison and ablation study setup
We compare DIPs against 14 alternative descriptors: Spin , SHOT , FPFH , USC , CGF , 3DMatch , Folding , PPFNet , PPF-FoldNet , DirectReg , CapsuleNet , PerfectMatch , FCGF , and D3Feat . In our ablation study we train the model for five epochs on a subset of 3DMatch’s scenes, i.e. Chess and Fire, and test on Home2 and Hotel3. We chose these test scenes because we found them to be sufficiently challenging.
III-D Evaluation metrics
Feature-matching recall We use the feature-matching recall (FMR) to quantify the descriptor quality . FMR does not require RANSAC as it directly averages the number of correctly matched point clouds across datasets. Only recall is measured, as the precision can be improved by pruning correspondences . FMR is defined as
where is the number of matching point cloud pairs having (overlap between each other). is a pair of corresponding points found in the descriptor (embedding) space via a mutual nearest-neighbour search . is the set that contains all the found pairs in the overlap regions and , respectively. is the ground-truth transformation alignment between and . is the indicator function. cm and are set based on the theoretical analysis that RANSAC will find at least three corresponding points that can provide the correct with probability 99.9% using no more than iterations . In addition to , we also report mean () and standard deviation () of before applying .
Registration recall We measure the registration recall for the transformation estimated with RANSAC . The registration recall quantifies the miss-rate by measuring the distance between corresponding points for each point cloud pair using the estimated transformation based on ground-truth point correspondence information. The registration recall is defined as
III-E Quantitative analysis and comparison
Tab. I reports DIP results in comparison with alternative descriptors on 3DMatch and 3DMatchRotated datasets. Results show that DIPs achieve state-of-the art results and that are rotation invariant as FMR is almost the same for both datasets. When FMR is measured at different values of , DIPs are more distinctive than the alternatives (Fig. 6). Interestingly, DIPs largely outperform PPFNet descriptors that are also computed with a PointNet-based backbone. We believe that this occurs on the one hand as a result of DIP’s LRF canonicalisation, in fact PPFNet’s FMR drops in 3DMatchRotated as no canonicalisation is performed, and on the other hand because DIPs encode only the local geometric information, which makes them more generic and distinctive across different scenes as opposed to PPFNet descriptors that instead encode contextual information too. Amongst the eight tested scenes, the worst performing one is Lab that will analyse in detail later. Our Python implementation takes ms to process a DIP using an i7-8700 CPU at 3.20GHz with a NVIDIA GTX 1070 Ti GPU and 16GB RAM. 93% of the execution time is for the LRF estimation. Once a patch is canonicalised the deep network processes the descriptor in ms. Note that the LRF estimation can be parallelised and implemented in C++ to reduce the execution time.
Tab. II reports the registration recall results, where the descriptor distinctiveness we have observed in Fig. 6 is reflected on the estimated transformations. On average, DIPs outperform all the other descriptors. We can see that the Lab scene mentioned before is the worst performing one. This occurs because Lab contains several point clouds of partially reconstructed objects and flat surfaces. Two examples are shown in Fig. 7. The first one is a failed registration due to the lack of informative geometries in the scene. The second one is a successful registration, where the kitchen appliances produced more distinctive descriptors than the first case.
We further evaluate the registration recall and assess DIP robustness following the comparative ablation study proposed in , where the registration recall is measured as a function of a decreasing number of sampled points used by RANSAC to estimate the transformation. Tab. III shows that DIPs on average have a superior robustness than the alternatives.
III-F Ablation study
Tab. IV reports the results of our ablation study on the implementation choices. Here we can see how the three modules, i.e. TNet, LRF and LRN, affect FMR. TNet learns to compensate for incorrectly estimated LRFs. But we can see that without LRF, TNet cannot learn the complete transformation to canonicalise the patches (FMR drops on 3DMatchRotated). LRF is key to make DIPs rotation invariant. Following and , we can notice how learning unitary-length descriptors improve FMR. Lastly, as expected, the more the capacity to encode the information in descriptors of larger dimension, the better the performance. However, we used 32-dimensional descriptors throughout all the experiments in order to compare results with existing descriptors.
III-G Generalisation ability: comparison and analysis
Tab. V reports the results obtained on the ETH dataset . We can see that DIPs on average largely outperform the alternative descriptors. Second to DIPs are PerfectMatch’s descriptors that, as DIPs, use LRF canonicalisation. However, differently from DIPs, PerfectMatch’s descriptors are learnt from hand-crafted representations, namely voxelised smoothed density value . We argue that letting the network learn the encoding from the points directly (end-to-end), leads to a much greater robustness and generalisation ability. We can also observe that the application of improves the performance. Fig. 8 shows an example of result from Gazebo-Summer. Although the sensor modality and the structure of the environment is very different from that of 3DMatch, DIPs maintain their distinctiveness and can be successfully used to register two point clouds reconstructed with a laser scanner.
Fig. 9 shows results on our dataset, i.e. VigoHome. In each point cloud we included the corresponding reference frame, which is where each mapping session started. The result of a successful registration estimated using DIPs is shown in the bottom-right corner. We can notice that the structure of the environment largely differs from that of 3DMatch and ETH datasets, and that the distribution of the points on the surfaces is much noisier that that in the 3DMatch dataset. To quantify the registration results, we used a similar evaluation of that used in Sec. III-F. For each number of sampled points we run RANSAC 100 times and compute the registration recall. Tab. VI shows that with only 5K points sampled from each point cloud, 85% of the times the three point clouds are correctly registered. A correct registration takes about to be processed. We deem this a great result because it is achieved with DIPs learnt on the 3DMatch dataset. As additional comparison, we have also quantified the registration recall using FPFH descriptors . However, we found that the registration fails regardless the parameters used. So we have intentionally not included the results obtained with FPFH in the table.
IV Conclusions
We presented a novel approach to learn local, compact and rotation invariant descriptors end-to-end through a PointNet-based deep neural network using canonicalised patches. The affine transformation embedded in our network is learnt with the specific goal of improving patch canonicalisation. We showed the importance of this step through our ablation study. Results showed that DIPs achieve comparable performance to the state-of-the-art on the 3DMatch dataset, but that outperform the state-of-the-art by a large margin in terms of generalisation to different sensors and scenes. We further confirmed this by capturing a new indoor dataset using the Visual-SLAM system of ARCore (Android) running on an off-the-shelf smartphone. We observed that we can achieve good generalisation because DIPs are learnt end-to-end from the points without any hand-crafted preprocessing after canonicalisation. Our future research direction is to improve the canonicalisation operation .
Acknowledgment
This research has received funding from the Fondazione CARITRO - Ricerca e Sviluppo programme 2018-2020.