UnsupervisedR&R: Unsupervised Point Cloud Registration via Differentiable Rendering

Mohamed El Banani, Luya Gao, Justin Johnson

Introduction

Consider the two scenes depicted in Fig 1. How are they related? What is the layout of the room they depict? Aligning partial views of a scene into a single whole is essential to understanding one’s environment and is a key component of numerous robotics tasks such as SLAM and SfM. Recent approaches have leveraged supervised learning to develop end-to-end systems that outperform traditional methods in both accuracy and speed . However, with the rising prevalence of cameras with depth sensors, we can expect a new stream of raw RGB-D data without the annotations needed for supervision. How can we leverage this data for unsupervised learning of point cloud registration?

The common approach to point cloud registration relies on correspondence extraction and geometric model fitting. Traditional approaches relied on hand-crafted features and robust estimators such as RANSAC . While those approaches work well, their performance is limited by their inability to flexibly adapt to different data distributions. Recent work leverages supervised learning to address those limitations by learning to extract feature descriptors , finding better correspondences , and training more efficient robust estimators . However, accurate pose annotation can be challenging to attain automatically, due to sensor error or reliance on traditional SfM pipelines with no convergence guarantees .

Meanwhile, self-supervised visual learning has made remarkable progress in learning semantic and 3D features. The key idea is to use natural transformations in the data as indirect supervision. RGB-D video provides us with this supervision since successive frames capture different views of the same scene. In this case, aligning two point clouds from nearby frames is not only about achieving good geometric consistency, but also showing good photometric consistency between the two views. By achieving both photometric and geometric consistency, we can train a system using RGB-D image pairs without relying on additional supervision.

We propose using view synthesis between RGB-D images as a task for learning point cloud registration. Given two RGB-D video frames, we extract features from each frame to generate a feature point cloud where each point is represented by both a 3D coordinate and a feature vector. The extracted features serve as descriptors for correspondence estimation. The model is trained end-to-end using photometric and geometric consistency losses between the input and rendered frames. Through using differentiable components, we back-propagate the losses to the feature encoder to learn features that allow us to estimate unique correspondences and accurately register the two views.

We evaluate our model on ScanNet ; a large indoor scene dataset. We find that our model outperforms the traditional registration pipeline with visual or geometric descriptors (§ 4.1). Furthermore, it performs on-par with supervised geometric registration approaches despite being unsupervised; supporting our claim that RGB-D self-supervision can alleviate the need for pose annotation. Finally, we analyze our model through several ablations (§ 4.2).

In summary, our contributions are as follows:

We propose an unsupervised approach to point cloud registration via differentiable alignment and rendering;

We show how a differentiable variant of Lowe’s ratio test is sufficient for correspondence matching;

We empirically demonstrate our approach’s efficacy against traditional & supervised registration approaches;

We validate our design choices by evaluating our model with several ablations.

Related Work

Feature Descriptors. Early work on feature point extraction can be traced back to using corners for stereo matching . This work culminated in patch-based feature 2D descriptors and geometric features based on histograms of local 3D relationships . Those descriptors have been very popular due to being efficient to compute, relatively robust, and data-agnostic. More recently, there has been an interest in leveraging convolutional neural networks to extract good visual descriptors and geometric descriptors . Relevant to our work are approaches that use geometric transformations to learn visual features. This has been commonly done by using known pose or correspondences between large collections of images or point clouds . We extend this work by using existing transformation in RGB-D video data and relying on consistency losses instead of pose supervision.

Correspondence Estimation and Fitting. Early work on image and point cloud registration assume perfect correspondences . ICP relaxes this assumption for closely aligned points by introducing the simple heuristic of assuming the closest point is the correspondence . However, extending to real-world settings requires the ability to determine such correspondences from the raw input or extracted features. Early work uses feature similarity and heuristic approaches to determine correspondence and robust estimators such as RANSAC to handle noise and outliers in the correspondences . For a review, see . More recent approaches advance this idea by learning differentiable functions for weighting the correspondences . Finally, there have been recent self-supervised approaches for registering object point clouds . Those approaches operate on dense point clouds that are either augmented and sampled for partial views with known pose and correspondences. Hence, while the setup might be self-supervised, the methods still require ground-truth annotation. We are inspired by this line of work, but differ from it in two key ways: (1) we take RGB-D images as input, not keypoints and descriptors or 3D scenes; (2) our approach is unsupervised, while those approaches require pose or correspondence supervision.

Differentiable SfM. There has been a large number of recent approaches that replace the traditional SfM pipeline with end-to-end learning approaches . Related to our work are approaches that propose unsupervised learning of depth and camera motion. This is typically done through learning two CNNs: a pose network and a depth network, that are trained to minimize a consistency loss between video frames. While CNN pose estimators have shown a lot of success on outdoor scenes, they have been challenged by cases with larger and more erratic camera motions (\egvideo from a hand-held device) . Similar to those approaches, we train an end-to-end system using photometric and geometric consistency losses. Unlike that work, we are interested in pointcloud registration with larger camera motions and learn features for correspondence alignment of RGB-D scans.

View Synthesis. View synthesis is the task of generating views of the scene from image inputs. One line of work focuses on synthesizing views with small camera motions . NeRF and its variants learn a rendering function for a specific scene from a large collection of multi-view images. While the goal of that work is highly photo-realistic renderings, we are primarily interested in utilizing view synthesis as a training task to enforce photometric consistency. Similar to our goals are approaches that synthesize views for unsupervised 3D learning of object shape and depth . Closest to our work is Wiles et al. who train a model to for depth estimation and view synthesis with the goal of generating highly photo-realistic views of the scene. Our work complements this earlier work since we learn pose while they learn depth.

Method

The goal of this work is to build a system that can learn point cloud registration from RGB-D video without any supervision. Our approach, shown in Fig. 2, is based on the traditional registration pipeline as it similarly extracts feature descriptors, finds correspondences, and finds the best alignment. We adapt this pipeline by operating directly on the images and learning our own features, as well as using photometric and geometric consistency losses to learn those features. We first present a high-level sketch of our approach before explaining each stage in more detail. Architectural details are presented in the appendix and our code is available at https://github.com/mbanani/unsupervisedRR.

Given two RGB-D images of the scene and the camera’s intrinsic matrix, we first extract 2D features for each image and project them into two feature point clouds. We extract correspondences between the two point clouds and rank the correspondences based on their uniqueness. We then use a differentiable optimizer to align the top kk correspondences and estimate the 6-DOF transformation between them. Finally, we render the point cloud from the two estimated viewpoints to generate an RGB image for each view. We use photometric and geometric consistency losses between the RGB-D inputs and outputs and back-propagate through our entire pipeline.

1 Point Cloud Generation

2 Correspondence Estimation

where D(p,q)D(p,q) is a distance-metric defined on the feature space. In our experiments, we use cosine distance to determine the closest features.

We extract such correspondences for all the points in both P\mathcal{P} and Q\mathcal{Q} since correspondence is not guaranteed to be bijective. As a result, we have two sets of correspondences, CP→Q\mathcal{C}_{\mathcal{P}\to\mathcal{Q}} and CQ→P\mathcal{C}_{\mathcal{Q}\to\mathcal{P}}, where each set consists of NN pairs.

Determining the quality of each correspondence is a challenge faced by any correspondence-based geometric fitting approach. Extracting correspondences based on only the nearest neighbor will result in many false positives due to falsely matching repetitive pairs or non-mutually visible portions of the images.

The standard approach is to estimate a weight for each correspondence that captures the quality of this correspondence. Recent approaches estimate a correspondence weight for each match using self-attention graph networks , PointNets , and CNNs . In our experiments, we found that a much simpler approach based on Lowe’s ratio test works well without requiring any additional parameters in the network. The basic intuition behind the ratio test is that unique correspondences are more likely to be true matches. As a result, the quality of correspondence (p,qp)(p,q_{p}) is not simply determined by D(p,qp)D(p,q_{p}), but rather between the ratio rr which is defined as

where qp,iq_{p,i} is the ii-th nearest neighbor to point pp in Q\mathcal{Q}. Since 0≤rp≤10\leq r_{p}\leq 1 and a lower ratio indicates a better match, we weigh each correspondence by w=1−rw=1-r.

In the traditional formulation, one would define a distance ratio threshold for inlier vs outliers. Instead, we rank the correspondences using their ratio weight and pick the top kk correspondences. We pick an equal number of correspondences from CP→Q\mathcal{C}_{\mathcal{P}\to\mathcal{Q}} and CQ→P\mathcal{C}_{\mathcal{Q}\to\mathcal{P}}. Additionally, we keep the weights for each correspondence to use in the geometric fitting step. Hence, we end up with a correspondence set M={(p,q,w)i:0≤i<k}\mathcal{M}=\{(p,q,w)_{i}:0\leq i<k\} where k=400k{=}400.

3 Geometric Fitting

Given a set of correspondences M\mathcal{M}, we would like to find the transformation, T∗∈SE(3)\mathcal{T^{*}}\in\text{SE(3)} that would minimize the error between the correspondences

where the error E(M,T)E(\mathcal{M},\mathcal{T}) is defined as:

This can be framed as a weighted Procrustes problem and solved using a weighted variant of Kabsch’s algorithm .

While the original Procrustes problem minimizes the distance between a set of unweighted correspondences , Choy et al. have shown that one can integrate weights into this optimization. This is done by calculating the covariance matrix between the centered and weighted point clouds, followed by calculating the SVD on the covariance matrix. For more details, see .

Integrating weights into the optimization is important for two reasons. First, it allows us to build robust estimators that can weigh correspondences based on our confidence in their uniqueness. More importantly, it makes the optimization differentiable with respect to the weights, allowing us to backpropagate the losses back to the encoder for feature learning.

While this approach is capable of integrating the weights into the optimization, it can still be sensitive to outliers with non-zero weights. We take inspiration from RANSAC and use random sampling to mitigate the problem of outliers. More specifically, we sample tt subsets of M\mathcal{M}, and use Equation 3 to find tt candidate transformations. We then choose the candidate that minimizes the weighted error on the full correspondence set. Since the tt optimizations on the correspondence subsets are all independent, we are able to run them in parallel to make the optimization more efficient. We deviate from classic RANSAC pipelines in that we choose the transformation that minimizes a weighted error, instead of maximizing inlier count, to avoid having to define an arbitrary inlier threshold.

It is worth noting that the model can be trained and tested with a different number of random subsets. In our experiments, we train the model with 10 randomly sampled subsets of 80 correspondences each. At test time, we use 100 subsets with 20 correspondences each. We evaluate the impact of those choices on performance and run time in § 4.2.

4 Point Cloud Rendering

The final step of our approach is to render the RGB-D images from the aligned point clouds. This provides us with our primary learning signals: photometric and depth consistency. The core idea is that if the camera locations are estimated correctly, the point cloud renders will be consistent with the input images. We use differentiable rendering to project the colored point clouds onto an image using the estimated camera pose and known intrinsics. Our pipeline is very similar to Wiles et al. .

A naive approach of simply rendering both point clouds suffers from a degenerate solution: the rendering will be accurate even if the alignment is incorrect. An extreme case of this would be to always estimate cameras looking in opposite directions. In that case, each image is projected in a different location of space and the output will be consistent without alignment. We address this issue by forcing the network to render each view using only the other image’s point cloud, as shown in Fig. 4. This forces the network to learn consistent alignment as a correct reconstruction requires the mutually visible parts of the scene to be correctly aligned. This introduces another challenge: how to handle the non-mutually visible surfaces of the scene?

While view synthesis approaches hallucinate the missing regions to output photo-realistic imagery , earlier work in differentiable SfM observed that the gradients coming from the hallucinated region negatively impact the learning . Our solution to this problem is to only evaluate the loss for valid pixels. Valid pixels, as shown in Fig 4, are ones for which rendering was possible; \ie, there were points along the viewing ray for those pixels. This is important in this work since invalid pixels can occur due to two reasons: non-mutually visible surfaces and pixels with missing depth. While the first reason is due to our approach, the second reason for invalid pixels is governed by current depth sensors which do not produce a depth value for each pixel.

In our experiments, we found that pose networks are very susceptible to the issues above; the network starts estimating very large poses within the first hundred iterations and never recovers. We also experimented with rendering the features and decoding them, similar to , but found that this resulted in worse alignment performance.

5 Losses

We use three consistency losses to train our model: photometric, depth, and correspondence. The photometric and depth losses are the L1 losses applied between the rendered and input RGB-D frames. Those losses are masked to only apply to valid pixels, as discussed in § 3.4. Additionally, we use the correspondence error calculated in Eq. 4 as our correspondence loss. We weight the photometric and depth losses with a weighting of 1 while the correspondence loss receives a weighting of 0.1.

Experiments

We now empirically evaluate our model on pairwise point cloud registration. Our experiments aim to answer several questions: (1) does unsupervised training provide us with useful features for alignment?; (2) can RGB-D video alleviate the need for the pose supervision required by geometric registration approaches?; (3) how do the different components of the model contribute to its performance?

We address those questions by evaluating our approach on two datasets of indoor scenes: ScanNet and 3DMatch . We find that our approach achieves better registration accuracy than off-the-shelf visual and geometric feature descriptors (§ 4.1). We also find that our approach performs on-par with supervised geometric registration approaches despite using significantly simpler correspondence matching and alignment algorithms; supporting our claim that RGB-D video can alleviate the need for pose supervision. Finally, we analyze our model components through several key ablations (§ 4.2).

Datasets. We evaluate our approach using ScanNet and 3D Match . ScanNet contains RGB-D images and ground-truth camera poses for 1513 scenes, while 3D Match is a much smaller dataset with a total of 101 scenes. We use the official data split of 1045/156/312 scenes for train/val/test for ScanNet. 3D Match only provides a train/test split, so we further divide the train split into train and validation; resulting in 71/11/19 RGB-D sequences for train/val/test split. We generate view pairs by sampling image pairs that are 20 frames apart. We sample the training scenes more densely by sampling all pairs that are 20 frames apart. This results in 1594k/12.6k/26k ScanNet pairs and 122k/1.5k/1.5k 3D Match pairs.

Baselines. We compare our model to several learned and non-learned point cloud registration approaches. Since we are interested in the unsupervised setting, we first compare against methods that do not require pose supervision. Our first set of baselines use off-the-shelf keypoint detectors and descriptors with RANSAC as the robust estimator. For all these baselines, we use Open3D’s RANSAC implementation . Despite being proposed over a decade ago, SIFT features are still used and serve as a strong baseline for a non-learned method. SuperPoint is a recently proposed approach for keypoint detection and description and has achieved state of the art performance in correspondence matching on several benchmarks. Finally, FCGF is a recently proposed geometric feature descriptor that has also achieved state-of-the-art performance on several 3D correspondence benchmarks. Furthermore, FCGF features have been used by several recent approaches for point cloud registration without further fine-tuning .

We also compare against two supervised geometric registration approaches: DGR and 3D MV Registration . Both of these approaches operate on FCGF point cloud embeddings as their input and learn how to extract good correspondences between pairs. There are two salient differences between our approaches: First, our approach is unsupervised, while those approaches rely on pose supervision. Second, our approach operates on RGB-D, while those approaches use the FCGF embeddings of the point cloud without relying on the images. This comparison demonstrates how leveraging the currently ignored RGB modality could alleviate the need for pose supervision and pretrained descriptors. We emphasize that we use the weights provided by the authors which were trained on the 3D Match Geometric Registration benchmark.

Training Details. We train our models with the Adam optimizer with a learning rate of 10−410^{-4} and momentum parameters of (0.9, 0.99). We train each model for 200K iterations. We implement our models in PyTorch , while making extensive use of PyTorch3D and Open3D .

We first evaluate our approach on point cloud registration. Given two RGB-D images, we estimate the 6-DOF pose that would best align the first input image to the second. The transformation is represented by a rotation matrix R\mathbf{R} and translation vector t\mathbf{t}.

Evaluation Metrics. We evaluate pairwise registration by evaluating the pose prediction as well as the chamfer distance between the estimated and ground-truth alignments. We compute the angular and translation errors as follows:

We report the translation error in centimeters and the rotation errors in degrees.

While pose gives us a good measure of performance, some scenes are inherently ambiguous and multiple alignments can explain the scene appearance; \eg, walls, floors, symmetric objects. To address these cases, we compute the chamfer distance between the scene and our reconstruction. Given two point clouds where P\mathcal{P} represents the correct alignment of the scene and Q\mathcal{Q} represents our reconstruction of the scene, we can define the closest pairs between the point clouds as set ΛP,Q={(p,arg min⁡q∈Q∣∣p−q∣∣):p∈P)\Lambda_{\mathcal{P},\mathcal{Q}}=\{(p,\operatorname*{arg\,min}_{q\in\mathcal{Q}}||p-q||):p\in\mathcal{P}). We then compute the chamfer error as follows:

For each of these error metrics, we report the mean and median errors over the dataset as well as the accuracy for different thresholds.

We conduct our experiments on ScanNet and report the results in Table 1. We find that our model learns accurate point cloud registration; outperforming prior feature descriptors and performing on-par with supervised geometric registration approaches. We next analyze our results through the questions posed at the start of this section.

Does unsupervised learning improve over off-the-shelf descriptors? Yes. We evaluate our approach against the traditional pipeline for registration: feature extraction using an off-the-shelf keypoint descriptor and alignment via RANSAC. We show large performance gains over both traditional and learned descriptors. It is important to note that FCGF and SuperPoint currently represent the state-of-the-art for feature descriptors. Furthermore, both methods have been used directly, without further fine-tuning, to achieve the highest performance on image registration benchmarks and geometric registration benchmarks . We also find that our approach learns features that can generalize to similar datasets. As shown in Table 1, our model trained on 3D Match outperforms the off-the-shelf descriptors while being competitive with supervised geometric registration approaches.

Does RGB-D training alleviate the need for pose supervision? Yes. We compare our approach to two recently proposed supervised point cloud registration approaches: DGR and 3D Multi-view Registration . Since their model was trained on 3D Match, we also train our model on 3D match and report the numbers. We find that our model is competitive with supervised approaches when trained on their dataset, and can outperform them when trained on ScanNet. However, a direct comparison is more nuanced since those two classes of methods differ in two key ways: training supervision and input modality.

We argue that the recent rise in RGB-D cameras on both hand-held devices and robotic systems supports our setup. First, the rise in devices suggests a corresponding increase in RGB-D raw data that will not necessarily be annotated with pose information. This increase provides a great opportunity for unsupervised learning to leverage this data stream. Second, while there are cases where depth sensing might be the better or only option (\eg, dark environment or highly reflective surfaces.), there are many cases where one has access to both RGB and depth information. The ability to leverage both can increase the effectiveness and robustness of a registration system. Finally, while we only learn visual features in this work, we note that our approach is easily extensible to learning both geometric and visual features since it is agnostic to how the features are calculated.

2 Ablations

We perform several ablation studies to better understand the model’s performance and its various components. In particular, we are interested in better understanding the impact of the optimization and rendering parameters on the overall model performance. While some ablations can only be applied during training (\eg, rendering choice), ablations that affect the correspondence estimation and fitting can be selectively applied during training, inference, or both. Hence, we consider all the variants.

Joint Rendering. Our first ablation investigates the impact of our rendering choices by rendering the output images from the joint point cloud. In § 3.4, we discuss rendering alternate views to force the model to align the pointclouds to produce accurate renders. As shown in Table 2, we find that naively rendering the joint point cloud results in a significant performance drop. This supports our claim that a joint render would negatively impact the features learned since the model can achieve good photometric consistency even if the pointclouds are not accurately aligned.

Ratio Test. In our approach, we use Lowe’s ratio test to estimate the weight for each correspondence. We ablate this component by instead using the feature distance between the corresponding points to rank the correspondences. Since this ablation can be applied to training or inference independently, we apply it to training, inference, or both. Our results indicate that the ratio test is critical to our model’s performance, as ablating it results in the largest performance drop. This supports our initial claims about the utility of the ratio test as a strong heuristic for filtering correspondences. It is worth noting that Lowe’s ratio test shows incredible efficacy in determining correspondence weights; a function often undertaken by far more complex models in recent work . Our approach is able to perform well using such a simple filtering heuristic since it is also learning the features, not just matching them.

Randomized Subsets. In our model, we estimate tt transformations based on tt randomly sampled subsets. This is inspired by RANSAC as it allows us to better handle outliers. We ablate this module by estimating a single transformation based on all the correspondences. Similar to the ratio test, this ablation can be applied to training or inference independently. As shown in Table 2, ablating this component at test time results in a significant drop in performance. Interestingly, we find that applying it during training and relieving it during testing improves performance. We posit that this ablation acts similarly to DropOut which forces the model to predict using a subset of the features and is only applied during training. As a result, the model is forced to learn better features during training, while gaining the benefits of randomized optimization during inference.

Number of subsets. We find that the number of subsets chosen has a significant impact on both run-time and performance. During training, we sample 10 subsets of 80 correspondences each. During testing, we sample 100 subsets of 80 correspondences each. For this set of experiments, we used the same pretrained weights and only vary the number of subsets used. Each subset still contains 80 correspondences. As shown in Table 3, using a larger number of subsets improves the performance while also increasing the run-time. Additionally, we find that the performance gains saturate at 100 subsets.

Conclusion

We present an unsupervised, end-to-end approach to pairwise RGB-D point cloud registration. We observe that existing approaches to point cloud registration rely on pose supervision for learning geometric point cloud alignment. However, with the increase in cameras with depth sensors, we expect a large stream of unannotated RGB-D data. This provides us with an opportunity to leverage unsupervised learning for more robust RGB-D point cloud registration.

To this end, we propose using view synthesis as a task for unsupervised point cloud registration via differentiable alignment and rendering. At the core of our approach is the notion of achieving geometric alignment through training a model on photometric consistency. Our approach learns to extract features from RGB-D data that allow it to both register and render the input frames. We show that our approach outperforms current state-of-the-art feature descriptors with RANSAC as well as supervised geometric registration approaches. This supports our initial premise of using RGB-D data to alleviate the need for pose supervision.

While our implementation relies solely on features extracted from RGB, our approach does not necessitate this. Specifically, our approach could be extended to learning geometric features for correspondence estimation. Furthermore, while we find that the ratio test allows us to achieve highly accurate registration, it would be interesting to explore whether recently proposed supervised correspondence filtering algorithms can be adapted for unsupervised training as well as how they would compare to the simple ratio test heuristic.

Acknowledgments We would like to thank the anonymous reviewers for their valuable comments and suggestions. We also thank Nilesh Kulkarni, Karan Desai, Richard Higgins, and Max Smith for many helpful discussions and feedback on early drafts of this work.

References