The Perfect Match: 3D Point Cloud Matching with Smoothed Densities

Zan Gojcic, Caifa Zhou, Jan D. Wegner, Andreas Wieser

Introduction

3D point cloud matching is necessary to combine multiple overlapping scans of a scene (e.g., acquired using an RGB-D sensor or a laser scanner) into a single representation for further processing like 3D reconstruction or semantic segmentation. Individual parts of the scene are usually captured from different viewpoints with a relatively low overlap. A prerequisite for further processing is thus aligning these individual point cloud fragments in a common coordinate system, to obtain one large point cloud of the complete scene.

Although some works aim to register 3D point clouds based on geometric constraints (e.g., ), most approaches match corresponding 3D feature descriptors that are custom-tailored for 3D point clouds and usually describe point neighborhoods with histograms of point distributions or local surface normals (e.g., ).

Since the comeback of deep learning, research on 3D local descriptors has followed the general trend in the vision community and shifted towards learning-based approaches and more specifically deep neural networks . Although the field has seen significant progress in the last three years, most learned 3D feature descriptors are either not rotation invariant , need very high output dimensions to be successful or can hardly generalize to new domains . In this paper, we propose 3DSmoothNet, a deep learning approach for 3D point cloud matching, which has low output dimension (16 or 32) for very fast correspondence search, high descriptiveness (outperforms all state-of-the-art approaches by more than 2020 percent points), is rotation invariant, and does generalize across sensor modalities and from indoor scenes of buildings to natural outdoor scenes.

We propose a new compact learned local feature descriptors for 3D point cloud matching that is efficient to compute and outperforms all existing methods significantly. A major technical novelty of our paper is the smoothed density value (SDV) voxelization as a new input data representation that is amenable to fully convolutional layers of standard deep learning libraries. The gain of SDV is twofold. On the one hand, it reduces the sparsity of the input voxel grid, which enables better gradient flow during backpropagation, while reducing the boundary effects, as well as smoothing out small miss-alignments due to errors in the estimation of the local reference frame (LRF). On the other hand, we assume that it explicitly models the smoothing that deep networks typically learn in the first layers, thus saving network capacity for learning highly descriptive features. Second, we present a Siamese network architecture with fully convolutional layers that learns a very compact, rotation invariant 3D local feature descriptor. This approach generates low-dimensional, highly descriptive features that generalize across different sensor modalities and from indoor to outdoor scenes. Moreover, we demonstrate that our low-dimensional feature descriptor (only 16 or 32 output dimensions) greatly speeds up correspondence search, which allows real-time applications.

Related Work

This section reviews advances in 3D local feature descriptors, starting from the early hand-crafted feature descriptors and progressing to the more recent approaches that apply deep learning.

Pioneer works on hand-crafted 3D local feature descriptors were usually inspired by their 2D counterparts. Two basic strategies exist depending on how rotation invariance is established. Many approaches including SHOT , RoPS , USC and TOLDI try to first estimate a unique local reference frame (LRF), which is typically based on the eigenvalue decomposition of the sample covariance matrix of the points in neighborhood of the interest point. This LRF is then used to transform the local neighborhood of the interest point to its canonical representation in which the geometric peculiarities, e.g. orientation of the normal vectors or local point density are analyzed. On the other hand, several approaches resort to a LRF-free representation based on intrinsically invariant features (e.g., point pair features). Despite significant progress, hand-crafted 3D local descriptors never reached the performance of hand-crafted 2D descriptors. In fact, they still fail to handle point cloud resolution changes, noisy data, occlusions and clutter .

Learned 3D Local Descriptors

The success of deep-learning methods in image processing also inspired various approaches for learning geometric representations of 3D data. Due to the sparse and unstructured nature of raw point clouds, several parallel tracks regarding the representation of the input data have emerged.

One idea is projecting 3D point clouds to images and then inferring local feature descriptors by drawing from the rich library of well-established 2D CNNs developed for image interpretation. For example, project 3D point clouds to depth maps and extract features using an auto-encoder. use a 2D CNN to combine the rendered views of feature points at multiple scales into a single local feature descriptor. Another possibility are dense 3D voxel grids either in the form of binary occupancy grids or an alternative encoding . For example, 3DMatch , one of the pioneer works in learning 3D local descriptors, uses a volumetric grid of truncated distance functions to represent the raw point clouds in a structured manner. Another option is estimating the LRF (or a local reference axis) for extracting canonical, high-dimensional, but hand-crafted features and using a neural network solely for a dimensionality reduction. Even-though these methods manage to learn a non-linear embedding that outperforms the initial representations, their performance is still bounded by the descriptiveness of the initial hand-crafted features.

PointNet and PointNet++ are seminal works that introduced a new paradigm by directly working on raw unstructured point-clouds. They have shown that a permutation invariance of the network, which is important for learning on unordered sets, can be accomplished by using symmetric functions. Albeit, successful in segmentation and classification tasks, they do not manage to encapsulate the local geometric information in a satisfactory manner, largely because they are unable to use convolutional layers in their network design . Nevertheless, PointNet offers a base for PPFNet , which augments raw point coordinates with point-pair features and incorporates global context during learning to improve the feature representation. However, PPFNet is not fully rotation invariant. PPF-FoldNet addresses the rotation invariance problem of PPFNet by solely using point-pair features as input. It is based on the architectures of the PointNet and FoldingNet and is trained in a self-supervised manner. The recent work of is based on PointNet, too, but deviates from the common approach of learning only the feature descriptor. It follows the idea of in trying to fuse the learning of the keypoint detector and descriptor in a single network in a weakly-supervised way using GPS/INS tagged 3D point clouds. does not achieve rotation invariance of the descriptor and is limited to smaller point cloud sizes due to using PointNet as a backbone.

Arguably, training a network directly from raw point clouds fulfills the end-to-end learning paradigm. On the downside, it does significantly hamper the use of convolutional layers, which are crucial to fully capture local geometry. We thus resort to a hybrid strategy that, first, transforms point neighborhoods into LRFs, second, encodes unstructured 3D point clouds as SDV grids amenable to convolutional layers and third, learns descriptive features with a siamese CNN. This strategy does not only establish rotation invariance but also allows good performance with low output dimensions, which speeds up correspondence search.

Method

In a nutshell, our workflow is as follows (Fig. 2 & 3): (i) given two raw point clouds, (ii) compute the LRF of the spherical neighborhood around the randomly selected interest points, (iii) transform the neighborhoods to their canonical representations, (iv) voxelize them with the help of Gaussian smoothing, (v) infer the per point local feature descriptors using 3DSmoothNet and, for example, use them as input to a RANSAC-based robust point cloud registration pipeline.

A core requirement for a generally applicable local feature descriptor is its invariance under isometry of the Euclidian space. Since achieving the rotation invariance in practice is non-trivial, several recent works choose to ignore it and thus do not generalize to rigidly transformed scenes . One strategy to make a feature descriptor rotation invariant is regressing the canonical orientation of a local 3D patch around a point as an integral part of a deep neural network inspired by recent work in 2D image processing . However, find that this strategy often fails for 3D point clouds. We therefore choose a different approach and explicitly estimate LRFs by adapting the method of . An overview of our method is shown in Fig. 2 and is described in the following.

The x-axis x^p\mathbf{\hat{x}}_{\mathbf{p}} is computed as the weighted vector sum

Intuitively, the weight αi\alpha_{i} favors points lying close to the interest point thus making the estimation of x^p\mathbf{\hat{x}}_{\mathbf{p}} more robust against clutter and occlusions. βi\beta_{i} gives more weight to points with a large scalar projection, which are likely to contribute significant evidence particularly in planar areas . Finally, the y-axis y^p\mathbf{\hat{y}}_{\mathbf{p}} completes the left-handed LRF and is computed as y^p=x^p×z^p\mathbf{\hat{y}}_{\mathbf{p}}=\mathbf{\hat{x}}_{\mathbf{p}}\times\mathbf{\hat{z}}_{\mathbf{p}}.

Smoothed density value (SDV) voxelization

Network architecture

Our network architecture (Fig. 3) is loosely inspired by L2Net , a state-of-the-art learned local image descriptor. 3DSmoothNet consists of stacked convolutional layers that applies strides of 22 (instead of max-pooling) in some convolutional layers to down-sample the input . All convolutional layers, except the final one, are followed by batch normalization and use the ReLU activation function . In our implementation, we follow and fix the affine parameters of the batch normalization layer to 1 and 0 and we do not train them during the training of the network. To avoid over-fitting the network, we add dropout regularization with a 0.30.3 dropout rate before the last convolutional layer. The output of the last convolutional layer is fed to a batch normalization layer followed by an l2l2 normalization to produce unit length local feature descriptors.

Training

We train 3DSmoothNet (Fig. 3) on point cloud fragments from the 3DMatch data set . This is an RGB-D data set consisting of 62 real-world indoor scenes, ranging from offices and hotel rooms to tabletops and restrooms. Point clouds obtained from a pool of data sets are split into 54 scenes for training and 8 scenes for testing. Each scene is split into several partially overlapping fragments with their ground truth transformation parameters TT.

Consider two fragments Fi\mathcal{F}_{i} and Fj\mathcal{F}_{j}, which have more than 30%30\% overlap. To generate training examples, we start by randomly sampling 300300 anchor points pa\mathbf{p}^{\text{a}} from the overlapping region of fragment Fi\mathcal{F}_{i}. After applying the ground truth transformation parameters Tj()T_{j}() the positive sample pp\mathbf{p}^{\text{p}} is then represented as the nearest-neighbor pp=:nn(pa)∈Tj(Fj)\mathbf{p}^{\text{p}}=:\text{nn}(\mathbf{p}^{\text{a}})\in T_{j}(\mathcal{F}_{j}), where nn()\text{nn}() denotes the nearest neighbor search in the Euclidean space based on the l2l2 distance. We refrain from pre-sampling the negative examples and instead use the hardest-in-batch method for sampling negative samples on the fly. During training we aim to minimize the soft margin Batch Hard (BH) loss function

The BH loss is defined for a mini-batch X\mathcal{X}, where Xia\mathbf{X}_{\text{i}}^{a} and Xip\mathbf{X}_{\text{i}}^{p} represent the SDV voxel grids of the anchor and positive input samples, respectively. The negative samples are retrieved as the hardest non-corresponding positive samples in the mini-batch (c.f. Eq. 8). Hardest-in-batch sampling ensures that negative samples are neither too easy (i.e, non-informative) nor exceptionally hard, thus preventing the model to learn normal data associations .

Results

Our 3DSmoothNet approach is implemented in C++ (input parametrization) using the PCL and in Python (CNN part) using Tensorflow . During training we extract SDV voxel grids of size W=H=D=0.3 mW=H=D=0.3~{}m (corresponding to ), centered at each interest point and aligned with the LRF. We use rLRF=3Wr_{\text{LRF}}=\sqrt{3}W to extract the spherical support S\mathcal{S} and estimate the LRF. We obtain the circumscribed sphere of our voxel grid and use the points transformed to the canonical frame to extract the SDV voxel grid. We split each SDV voxel grid into 16316^{3} voxels with an edge w=W16w=\frac{W}{16} and use a Gaussian smoothing kernel with an empirically determined optimal width h=1.75w2h=\frac{1.75w}{2}. All the parameters wew slected on the validation data set. We train the network with mini-batches of size 256256 and optimize the parameters with the ADAM optimizer , using an initial learning rate of 0.0010.001 that is exponentially decayed every 50005000 iterations. Weights are initialized orthogonally with 0.60.6 gain, and biases are set to 0.010.01. We train the network for 2020 epochs.

We evaluate the performance of 3DSmoothNet for correspondence search on the 3DMatch data set and compare against the state-of-the-art. In addition, we evaluate its generalization capability to a different sensor modality (laser scans) and different scenes (e.g., forests) on the Challenging data sets for point cloud registration algorithms data set denoted as ETH data set.

Comparison to state-of-the-art

We adopt the commonly used hand-crafted 3D local feature descriptors FPFH (33 dimensions) and SHOT (352 dimensions) as baselines and run implementations provided in PCL for both approaches. We compare against the current state-of-the-art in learned 3D feature descriptors: 3DMatch (512 dimensions), CGF (32 dimensions), PPFNet (64 dimensions), and PPF-FoldNet (512 dimensions). In case of 3DMatch and CGF we use the implementations provided by the authors in combination with the given pre-trained weights. Because source-code of PPFNet and PPF-FoldNet is not publicly available, we report the results presented in the original papers. For all descriptors based on the normal vectors, we ensure a consistent orientation of the normal vectors across the fragments. To allow for a fair evaluation, we use exactly the same interest points (provided by the authors of the data set) for all descriptors. In case of descriptors that are based on spherical neighborhoods, we use a radius that yields a sphere with the same volume as our voxel. All exact parameter settings, further implementation details etc. used for these experiments are available in supplementary material.

1 Evaluation on the 3DMatch data set

The test part of the 3DMatch data set consists of 88 indoor scenes split into several partially overlapping fragments. For each fragment, the authors provide indices of 50005000 randomly sampled feature points. We use these feature points for all descriptors. The results of PPFNet and PPF-FoldNet are based on a spherical neighborhood with a diameter of 0.6m0.6\text{m}. Furthermore, due to its memory bottleneck, PPFNet is limited to 20482048 interest points per fragment. We adopt the evaluation metric of (see supplementary material). It is based on the theoretical analysis of the number of iterations needed by a robust registration pipeline, e.g. RANSAC, to find the correct set of transformation parameters between two fragments. As done in , we set the threshold τ1=0.1m\tau_{1}=0.1m on the l2l2 distance between corresponding points in the Euclidean space and τ2=0.05\tau_{2}=0.05 to threshold the inlier ratio of the correspondences at 5%5\%.

Output dimensionality of 3DSmoothNet

A general goal is achieving the highest matching performance with the lowest output dimensionality (i.e., filter number in the last convolutional layer of 3DSmoothNet) to decrease run-time and to save memory. Thus, we first run trials to find a good compromise between matching performance and efficiency for the 3DSmoothNet descriptorsRecall that for correspondence search, the brute-force implementation of nearest-neighbor search scales with O(DN2)\mathcal{O}(DN^{2}), where DD denotes the dimension and NN the number of data points. The time complexity can be reduced to O(DNlog⁡N)\mathcal{O}(DN\log{N}) using tree-based methods, but still becomes inefficient if DD grows large (”curse of dimensionality”).. We find that the performance of 3DSmoothNet quickly starts to saturate with increasing output dimensions (Fig. 6). There is only marginal improvement (if any) when using more than 6464 dimensions. We thus decide to process all further experiments only for 1616 and 3232 output dimensions of 3DSmoothNet.

Comparison to state-of-the-art

Results of experimental evaluation on the 3DMatch data set are summarized in Tab. 1 (left) and two hard cases are shown in Fig. 4. Ours (16) and Ours (32) achieve an average recall of 92.8%92.8\% and 94.7%94.7\%, respectively, which is close to solving the 3DMatch data set. 3DSmoothNet outperforms all state-of-the-art 3D local feature descriptors with a significant margin on all scenes. Remarkably, Ours (16) improves average recall over all scenes by almost 2020 percent points with only 1616 output dimensions compared to 512512 dimensions of PPF-FoldNet and 352352 of SHOT. Furthermore, Ours (16) and Ours (32) show a much smaller recall standard deviation (STD), which indicates robustness of 3DSmoothNet to scene changes and hints at good generalization ability. The inlier ratio threshold τ2=0.05\tau_{2}=0.05 as chosen by results in ≈55k\approx 55\text{k} iterations to find at least 3 correspondences (with 99.9%99.9\% probability) with the common RANSAC approach. Increasing the inlier ratio to τ2=0.2\tau_{2}=0.2 would decrease RANSAC iterations significantly to ≈850\approx 850, which would speed up processing massively. We thus evaluate how gradually increasing the inlier ratio changes performance of 3DSmoothNet in comparison to all other tested approaches (Fig. 7). While the average recall of all other methods drops below 30%30\% for τ2=0.2\tau_{2}=0.2, recall of Ours (16) (blue) and Ours (32) (orange) remains high at 62%62\% and 72%72\%, respectively. This indicates that any descriptor-based point cloud registration pipeline can be made more efficient by just replacing the existing descriptor with our 3DSmoothNet.

Rotation invariance

We take a similar approach as to validate rotation invariance of 3DSmoothNet by rotating all fragments of 3DMatch data set (we name it 3DRotatedMatch) around all three axis and evaluating the performance of the selected descriptors on these rotated versions. Individual rotation angles are sampled arbitrarily between [0,2π][0,2\pi] and the same indices of points for evaluation are used as in the previous section. Results of Ours (16) and Ours (32) remain basically unchanged (Tab 1 (right)) compared to the non-rotated variant (Tab 1 (left)), which confirms rotation invariance of 3DSmoothNet (due to estimating LRF). Because performance of all other rotation invariant descriptors remains mainly identical, too, 3DSmoothNet again outperforms all state-of-the-art methods by more than 2020 percent points.

Ablation study

To get a better understanding of the reasons for the very good performance of 3DSmoothNet, we analyze the contribution of individual modules with an ablation study on 3DMatch and 3DRotatedMatch data sets. Along with the original 3DSmoothNet, we consider versions without SDV (we use a simple binary occupancy grid), without LRF and finally without both, LRF and SDV. All networks are trained using the same parameters and for the same number of epochs. Results of this ablation study are summarized in Tab. 2. It turns out that the version without LRF performs best on 3DMatch because most fragments are already oriented in the same way and the original data set version is tailored for descriptors that are not rotation invariant. Inferior performance of the full pipeline on this data set is most likely due to a few wrongly estimated LRF, which reduces performance on already oriented data sets (but allows generalizing to the more realistic, rotated cases). Unsurprisingly, 3DSmoothNet without LRF fails on 3DRotatedMatch because the network cannot learn rotation invariance from the data. A significant performance gain of up to more than 99 percent points can be attributed to using an a SDV voxel grid instead of the traditional binary occupancy grid.

2 Generalizability across modalities and scenes

We evaluate how 3DSmoothNet generalizes to outdoor scenes obtained using a laser scanner (Fig. 1). To this end, we use models Ours (16) and Ours (32) trained on 3DMatch (RGB-D images of indoor scenes) and test on four outdoor laser scan data sets Gazebo-Summer, Gazebo-Winter, Wood-Autumn and Wood-Summer that are part of the ETH data set . All acquisitions contain several partially overlapping scans of sparse and dense vegetation (e.g., trees and bushes). Accurate ground-truth transformation matrices are available through extrinsic measurements of the scanner position with a total-station. We start our evaluation by down-sampling the laser scans using a voxel grid filter of size 0.02m0.02\text{m}. We randomly sample 50005000 points in each point cloud and follow the same evaluation procedure as in Sec 4.1, again considering only point clouds with more than 30%30\% overlap. More details on sampling of the feature points and computation of the point cloud overlaps are available in the supplementary material. Due to the lower resolution of the point clouds, we now use a larger value of W=1 mW=1~{}\text{m} for the SDV voxel grid (consequently the radius for the descriptors based on the spherical neighborhood is also increased). A voxel grid with an edge equal to 1.5 m1.5~{}\text{m} is used for 3DMatch because of memory restrictions. Results on the ETH data set are reported in Tab 3. 3DSmoothNet achieves best performance on average (right column), Ours (32) with 79.0%79.0\% average recall clearly outperforming Ours (16) with 48.2%48.2\% due to its larger output dimension. Ours (32) beats runner-up (unsupervised) SHOT by more than 1515 percent points whereas all state-of-the-art methods stay significantly below 30%30\%. In fact, Ours (32) applied to outdoor laser scans still outperforms all competitors that are trained and tested on the 3DMatch data set (cf. Tab. 3 with Tab. 1).

3 Computation time

We compare average run-time of our approach per interest point on 3DMatch test fragments to in Tab. 4 (ran on the same PC with Intel Xeon E5-1650, 32 GB of ram and NVIDIA GeForce GTX1080). Note that input preparation (Input prep.) and inference of are processed on the GPU, while our approach does input preparation on CPU in its current state. For both methods, we run nearest neighbor correspondence search on the CPU. Naturally, input preparation of 3DSmoothNet on the CPU takes considerably longer (4.2 ms versus 0.5 ms), but still the overall computation time is slightly shorter (4.6 ms versus 5.0 ms). Main drivers for performance are inference (0.3 ms versus 3.7 ms) and nearest neighbor correspondence search (0.1 ms versus 0.8 ms). This indicates that it is worth investing computational resources into custom-tailored data preparation because it significantly speeds up all later tasks. The bigger gap between Ours(16 dim) and Ours(32 dim), is a result of the lower capacity and hence lower descriptiveness of the 16-dimensional descriptor, which becomes more apparent on the harder ETH data set, but can also be seen in additional experiments in supplementary material. Supplementary material also contains additional experiments, which show the invariance of the proposed descriptor to changes in point cloud density.

Conclusions

We have presented 3DSmoothNet, a deep learning approach with fully convolutional layers for 3D point cloud matching that outperforms all state-of-the-art by more than 2020 percent points. It allows very efficient correspondence search due to low output dimensions (16 or 32), and a model trained on indoor RGB-D scenes generalizes well to terrestrial laser scans of outdoor vegetation. Our method is rotation invariant and achieves 94.9%94.9\% average recall on the 3DMatch benchmark data set, which is close to solving it. To the best of our knowledge, this is the first learned, universal point cloud matching method that allows transferring trained models between modalities. It takes our field one step closer to the utopian vision of a single trained model that can be used for matching any kind of point cloud regardless of scene content or sensor.

References

Supplementary Material

In this supplementary material we provide additional information about the evaluation experiments (Sec. 6.1, 6.2 and 6.3) along with the detailed per-scene results (Sec. 6.4) and some further visualizations (Fig. 8 and 9). The source code and all the data needed for comparison are publicly available at https://github.com/zgojcic/3DSmoothNet.

This section provides a detailed explanation of the evaluation metric adopted from and used for all evaluation experiments throughout the paper.

Consider two point cloud fragments P\mathcal{P} and Q\mathcal{Q}, which have more than 30%30\% overlap under ground-truth alignment. Furthermore, let all such pairs form a set of fragment pairs F={(P,Q)}\mathcal{F}=\{(\mathcal{P},\mathcal{Q})\}. For each fragment pair the set of correspondences obtained in the feature space is then defined as

where f(p)f(\mathbf{p}) denotes a non-linear function that maps the feature point p\mathbf{p} to its local feature descriptor and nn() denotes the nearest neighbor search based on the l2l2 distance. Finally, the quality of the correspondences in terms of average recall RR per scene is computed as

where TfT_{f} denotes the ground-truth transformation alignment of the fragment pair f∈Ff\in\mathcal{F}. τ1\tau_{1} is the threshold on the Euclidean distance between the correspondence pair (i,j)(i,j) found in the feature space and τ2\tau_{2} is a threshold on the inlier ratio of the correspondences . Following we set τ1=0.1m\tau_{1}=0.1\text{m} and τ2=0.05\tau_{2}=0.05 for both, the 3DMatch as well as the ETH data set. The evaluation metric is based on the theoretical analysis of the number of iterations kk needed by RANSAC to find at least n=3n=3 corresponding points with the probability of success p=99.9%p=99.9\%. Considering, τ2=0.05\tau_{2}=0.05 and the relation

the number of iterations equals k≈55000k\approx 55000 and can be greatly reduced if the number of inliers τ2\tau_{2} can be increased (e.g. k=860k=860 if τ2=0.2\tau_{2}=0.2).

2 Baseline Parameters

In order to perform the comparison with the state-of-the-art methods, several parameters have to be set. To ensure a fair comparison we set all the parameters relative to our voxel grid width WW which we set as W3DMatch=0.3mW_{\textit{3DMatch}}=0.3\text{m} and WETH=1mW_{\textit{ETH}}=1\text{m} for 3DMatch and ETH data sets respectively. More specific, for the descriptors based on the spherical support we use a feature radius rf=34πWr_{f}=\sqrt{\frac{3}{4\pi}}W that yields a sphere with the same volume as our voxel grid and for all voxel-based descriptors we use the same voxel grid width WW. For descriptors that require, along with the coordinates also the normal vectors, we use the point cloud library (PCL) built-in function for normal vector computation, using all the points in the spherical support with the radius rn=rf2r_{n}=\frac{r_{f}}{2}. Tab. 5 provides all the parameters that were used for the evaluation. If some parameters are not listed in Tab 5 we use the original values set by the authors. For the handcrafted descriptors, FPFH and SHOT we use the implementation provided by the original authors as a part of the PCLhttps://github.com/PointCloudLibrary/pcl. We use the PCL version 1.8.11.8.1 x6464 on Windows 1010 and use the parallel programming implementations (omp) of both descriptors. For 3DMatch we use the implementation provided by the authorshttps://github.com/andyzeng/3dmatch-toolbox on Ubuntu 16.04 in combination with the CUDA 8.08.0 and cuDNN 5.15.1. Finally, for CGF we use the implementation provided by the authorshttps://github.com/marckhoury/CGF on a PC running Windows 1010. Note that we report the results of PPFNet and PPF-FoldNet as reported by the authors in the original papers, because the source code is not publicly available. Nevertheless, for the sake of completeness we report the feature radius rfr_{f} and the k-nearest neighbors knk_{n} used for the normal vector computation, which were used by the authors in the original works. For the 3DRotatedMatch and 3DSparseMatch data sets we use the same parameters as for the3DMatch data set.

The authors of the 3DMatch descriptor provide along with the source code and the trained model also the precomputed truncated distance function (TDF) representation and inferred descriptors for the 3DMatch data set. We use this descriptors directly for all evaluations on the original 3DMatch data set. For the evaluations on the 3DRotatedMatch, 3DSparseMatch and ETH data sets we use their source code in combination with the pretrained model to infer the descriptors. When analyzing the 3DSparseMatch data set results, we noticed a discrepancy. The descriptors inferred by us achieve better performance than the provided ones. We analyzed this further and determined that the TDF representation (i.e. the input to the CNN) is identical and the difference stems from the inference using their provided weights. In the paper this is marked by a footnote in the results section. For the sake of consistency, we report in this Supplementary material all results for 3DMatch data set using the precoumpted descriptors and the results on all other data set using the descriptors inferred by us.

3 Preprocessing of the benchmark data sets

The authors of 3DMatch data set provide along with the point cloud fragments and the ground-truth transformation parameters also the indices of the interest points and the ground-truth overlap for all fragments. To make the results comparable to previous works, we use these indices and overlap information for all descriptors and perform no preprocessing of the data.

DSparseMatch data set

In order to test the robustness of our approach to variations in point density we create a new data set, denoted as 3DSparseMatch, using the point cloud fragments from the 3DMatch data set. Specifically, we first extract the indices of the interest points provided by the authors of the 3DMatch data set and then randomly downsample the remaining points, keeping 50%50\%, 25%25\% and 12.5%12.5\% of the points. We consider two scenarios in the evaluation. In the first scenario we use one of the fragments to be registered with the full and the other one with the reduced point cloud density (Mixed), while in the second scenario we evaluate the descriptors on the fragments with the same level of sparsity (Both).

ETH data set

For the ETH data set we use the point clouds and the ground-truth transformation parameters provided by the authors of the data set. We start by downsampling the point clouds using a voxel grid filter with the voxel size equal to 0.02m0.02\text{m}. The authors of the data set also provide the ground-truth overlap information, but due to the downsampling step we opt to compute the overlap on our own as follows. Let pi∈P\mathbf{p}_{i}\in\mathcal{P} and qi∈Q\mathbf{q}_{i}\in\mathcal{Q} denote points in the point clouds P\mathcal{P} and Q\mathcal{Q}, which are part of the same scene of ETH data set, respectively. Given the ground-truth transformation TPQT_{\mathcal{P}}^{\mathcal{Q}} that aligns the point cloud Q\mathcal{Q} with the point cloud P\mathcal{P}, we compute the overlap ψP,Q\psi_{\mathcal{P},\mathcal{Q}} relative to point cloud P\mathcal{P} as

where nn denotes the nearest neighbor search based on the l2l2 distance in the Euclidean space and τψ\tau_{\psi} thresholds the distance between the nearest neighbors. In our evaluation experiments, we select τψ=0.06m\tau_{\psi}=0.06\text{m}, which equals three times the resolution of the point clouds after the voxel grid downsampling, and consider only the point cloud pairs for which both ψP,Q\psi_{\mathcal{P},\mathcal{Q}} and ψQ,P\psi_{\mathcal{Q},\mathcal{P}} are bigger than 0.30.3. Because no indices of the interest points are provided we randomly sample 50005000 interest points that have more than 1010 neighbor points in a sphere with a radius r=0.5mr=0.5\text{m} in every point cloud. The condition of minimum ten neighbors close to the interest point is enforced in order to avoid the problems with the normal vector computation.

4 Detailed results

Detailed per scene results on the 3DMatch data set are reported in Tab. 6. Ours (32 dim) consistently outperforms all state-of-the-art by a significant margin and achieves a recall higher than 89%89\% on all of the scenes. However, the difference between the performance of individual descriptors is somewhat masked by the selected low value of τ2\tau_{2}, e.g. same average recall on Hotel 3 scene achieved by Ours (16 dim) and Ours (32 dim). Therefore, we additionally perform a more direct evaluation of the quality of found correspondences, by computing the average number of correct correspondences established by each individual descriptor (Tab 7). Where the term correct correspondences, denotes the correspondences for which the distance between the points in the coordinate space after the ground-truth alignment is smaller than 0.1m0.1\text{m}. Results in Tab. 7 again show the dominant performance of the 3DSmoothNet compared to the other state-of-the-art but also highlight the difference between Ours (32 dim) and Ours (16 dim). Remarkably, Ours (32 dim) can establish almost two times more correspondences than the closest competitor.

DRotatedMatch data set

We additionally report the detailed results on the 3DRotatedMatch data set in Tab 8. Again, 3DSmoothNet outperforms all other descriptor on all the scenes and maintains a similar performance as on the 3DMatch data set. As expected the performance of the rotational invariant descriptors is not affected by the rotations of the fragments, whereas the performance of the descriptors, which are not rotational invariant drops to almost zero. This greatly reduces the applicability of such descriptors for general use, where one considers the point cloud, which are not represented in their canonical representation.

DSparseMatch data set

Tab 9 shows the results on the three different density levels (50%50\%, 25%25\% and 12,5%12,5\%) of the 3DSparseMatch data set. Generally, all descriptors perform better when the point density of only one fragments is reduced, compared to when both fragments are downsampled. In both scenarios, the recall of our approach drops marginally by max 11 percent point and remains more than 20 percent points above any other competing method. Therefore, 3DSmoothNet can be labeled as invariant to point density changes.