Learning to Orient Surfaces by Self-supervised Spherical CNNs

Riccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva, Samuele Salti, Luigi Di Stefano

Introduction

Humans naturally develop the ability to mentally portray and reason about objects in what we perceive as their neutral, canonical orientation, and this ability is key for correctly recognizing and manipulating objects as well as reasoning about the environment. Indeed, mental rotation abilities have been extensively studied and linked with motor and spatial visualization abilities since the 70s in the experimental psychology literature .

Robotic and computer vision systems similarly require neutralizing variations w.r.t. rotations when processing 3D data and images in many important applications such as grasping, navigation, surface matching, augmented reality, shape classification and detection, among the others. In these domains, two main approaches have been pursued so to define rotation-invariant methods to process 3D data: rotation-invariant operators and canonical orientation estimation. Pioneering works applying deep learning to point clouds, such as PointNet achieved invariance to rotation by means of a transformation network used to predict a canonical orientation to apply direclty to the coordinates of the input point cloud. Despite being trained by sampling the range of all possible rotations at training time through data augmentation, this approach, however, does not generalize to rotations not seen during training. Hence, invariant operators like rotation-invariant convolutions were introduced, allowing to train on a reduced set of rotations (ideally one, the unmodified data) and test on the full spectrum of rotations . Canonical orientation estimation, instead, follows more closely the human path to invariance and exploits the geometry of the surface to estimate an intrinsic 3D reference frame which rotates with the surface.

Transforming the input data by the inverse of the 3D orientation of such reference frame brings the surface in an orientation-neutral, canonical coordinate system wherein rotation invariant processing and reasoning can happen. While humans have a preference for a canonical orientation matching one of the usual orientations in which they encounter an object in everyday life, in machines this paradigm does not need to favour any actual reference orientation over others: as illustrated in Figure 1, an arbitrary one is fine as long as it can be repeatably estimated from the input data.

Despite mental rotation tasks being solved by a set of unconscious abilities that humans learn through experience, and despite the huge successes achieved by deep neural networks in addressing analogous unconscious tasks in vision and robotics, the problem of estimating a canonical orientation is still solved solely by handcrafted proposals . This may be due to convnets, the standard architectures for vision applications, reliance on the convolution operator in Euclidean domains, which possesses only the property of equivariance to translations of the input signal. However, the essential property of a canonical orientation estimation algorithm is equivariance with respect to 3D rotations because, upon a 3D rotation, the 3D reference frame which establishes the canonical orientation of an object should undergo the same rotation as the object. We also point out that, although, in principle, estimation of a canonical reference frame is suitable to pursue orientation neutralization for whole shapes, in past literature it has been intensively studied mainly to achieve rotation-invariant description of local surface patches.

In this work, we explore the feasibility of using deep neural networks to learn to pursue rotation-invariance by estimating the canonical orientation of a 3D surface, be it either a whole shape or a local patch. Purposely, we propose to leverage Spherical CNNs , a recently introduced variant of convnets which possesses the property of equivariance w.r.t. 3D rotations by design, in order to build Compass, a self-supervised methodology that learns to orient 3D shapes. As the proposed method computes feature maps living in SO⁡(3)\operatorname{SO}(3), i.e. feature map coordinates define 3D rotations, and does so by rotation-equivariant operators, any salient element in a feature map, e.g. its arg max⁡\operatorname*{arg\,max}, may readily be used to bring the input point cloud into a canonical reference frame. However, due to discretization artifacts, Spherical CNNs turn out to be not perfectly rotation-equivariant . Moreover, the input data may be noisy and, in case of 2.5D views sensed from 3D scenes, affected by self-occlusions and missing parts. To overcome these issues, we propose a robust end-to-end training pipeline which mimics sensor nuisances by data augmentation and allows the calculation of gradients with respect to feature maps coordinates. The effectiveness and general applicability of Compass is established by achieving state-of-the art results in two challenging applications: robust local reference frame estimation for local surface patches and rotation-invariant global shape classification.

Related Work

The definition of a canonical orientation of a point cloud has been studied mainly in the field of local features descriptors used to establish correspondences between sets of distinctive points, i.e. keypoints. Indeed, the definition of a robust local reference frame, i.e. R(p)={x^(p),y^(p),z^(p)∣y^=z^×x^}\mathcal{R}(p)=\{\hat{\mathbf{x}}(p),\hat{\mathbf{y}}(p),\hat{\mathbf{z}}(p)\mid\hat{\mathbf{y}}=\hat{\mathbf{z}}\times\hat{\mathbf{x}}\}, with respect to which the local neighborhood of a keypoint pp is encoded, is crucial to create rotation-invariant features. Several works define the axes of the local canonical system as eigenvectors of the 3D covariance matrix between points within a spherical region of radius rr centered at pp. As the signs of the eigenvectors are not repeatable, some works focus on the disambiguation of the axes . Alternatively, another family of methods leverages the normal to the surface at pp, i.e. n^(p)\hat{n}(p), to fix the z^\hat{\mathbf{z}} axis, and then exploits geometric attributes of the shape to identify a reference direction on the tangent plane to define the x^\hat{\mathbf{x}} axis . Compass differs sharply from previous methods because it learns the cues necessary to canonically orient a surface without making a priori assumptions on which details of the underlying geometry may be effective to define a repeatable canonical orientation.

On the other hand, PointNets employ a transformation network to predict an affine rigid motion to apply to the input point clouds in order to correctly classify global shapes under rigid transformations. In Esteves et al. prove the limited generalization of PointNet to unseen rotations and define the Spherical convolutions to learn an invariant embedding for mesh classification. In parallel, Cohen et al. use Spherical correlation to map Spherical inputs to SO⁡(3)\operatorname{SO}(3) features then processed with a series of convolutions on SO⁡(3)\operatorname{SO}(3). Similarly, PRIN proposes a network based on Spherical correlations to operate on spherically voxelized point clouds. SFCNN re-defines the convolution operator on a discretized sphere approximated by a regular icosahedral lattice. Differently, in , Zhang et al. adopt low-level rotation invariant geometric features (angles and distances) to design a convolution operator for point cloud processing. Deviating from this line of work on invariant convolutions and operators, we show how rotation-invariant processing can be effectively realized by preliminary transforming the shape to a canonical orientation learned by Compass.

Finally, it is noteworthy that several recent works rely on the notion of canonical orientation to perform category-specific 3D reconstruction from a single or multiple views.

Proposed Method

In this section, we provide a brief overview on Spherical CNNs to make the paper self-contained, followed by a detailed description of our method. For more details, we point readers to .

3D Rotations: 3D rotations live in a three-dimensional manifold, the SO⁡(3)\operatorname{SO}(3) group, which can be parameterized by ZYZ-Euler angles as in . Given a triplet of Euler angles α,β,γ\alpha,\beta,\gamma, the corresponding 3D rotation matrix is given by the product of two rotations about the zz-axis, Rz(⋅)R_{z}(\cdot), and one about the yy-axis, Ry(⋅)R_{y}(\cdot), i.e. R(α,β,γ)=Rz(α)Ry(β)Rz(γ)R(\alpha,\beta,\gamma)=R_{z}(\alpha)R_{y}(\beta)R_{z}(\gamma). Points represented as 3D unit vectors xx can be rotated by using the matrix-vector product RxRx.

where the operator LRL_{R} rotates the function ff by R∈SO⁡(3)R\in\operatorname{SO}(3), by composing its input with R−1R^{-1}, i.e. [LRf](x)=f(R−1x)[L_{R}f](x)=f(R^{-1}x), where x∈S2x\in S^{2}. Although both the input and the filter live in S2S^{2}, the spherical correlation produces an output signal defined on SO⁡(3)\operatorname{SO}(3) .

Spherical and SO(3) correlation equivariance w.r.t. rotations: It can be shown that both correlations in (1) and (2) are equivariant with respect to rotations of the input signal. The feature map obtained by correlation of a filter ψ\psi with an input signal hh rotated by Q∈SO⁡(3)Q\in\operatorname{SO}(3), can be equivalently computed by rotating with the same rotation QQ the feature map obtained by correlation of ψ\psi with the original input signal hh, i.e.:

Signal Flow: In Spherical CNNs, the input signal, e.g. an image or, as it is the case of our settings, a point cloud, is first transformed into a kk-valued spherical signal. Then, the first network layer (S2S^{2} layer) computes feature maps by spherical correlations (⋆\star). As the computed feature maps are SO(3) signals, the successive layers (SO(3) layers) compute deeper feature maps by SO(3) correlations (∗\ast).

2 Methodology

We define the rotated cloud, Vc\mathcal{V}_{c}, in (4) to be the canonical, rotation-neutral version of V\mathcal{V}, i.e. the function gg outputs the inverse of the 3D rotation matrix that brings the points in V\mathcal{V} into their canonical reference frame. (5) states the equivariance property of gg: if the input cloud is rotated, the output of the function should undergo the same rotation. As a result, two rotated versions of the same cloud are brought into the same canonical reference frame by (4).

Due to the equivariance property of Spherical CNNs layers, upon a rotation of the input signal each feature map does rotate accordingly. Moreover, the domain of the feature maps in Spherical CNNs is SO(3), i.e. each value of the feature map is naturally associated with a rotation. This means that one could just track any distinctive feature map value to establish a canonical orientation satisfying (4) and (5). Indeed, defining as Φ\Phi the composition of S2S^{2} and SO⁡(3)\operatorname{SO}(3) correlation layers in our network, if the last layer produces the feature map [Φ(fV)][\Phi(f_{\mathcal{V}})] when processing the spherical signal fVf_{\mathcal{V}} for the cloud V\mathcal{V}, the same network will compute the feature map [LRΦ(fV)]=[Φ(LRfV)]=[Φ(fT)][L_{R}\Phi(f_{\mathcal{V}})]=[\Phi(L_{R}f_{\mathcal{V}})]=[\Phi(f_{\mathcal{T}})] when processing the rotated cloud T=RV\mathcal{T}=R\mathcal{V}, with spherical signal fT=LRfVf_{\mathcal{T}}=L_{R}f_{\mathcal{V}}. Hence, if for instance we select the maximum value of the feature map as the distinctive value to track, and the location of the maximum is at QVmax∈SO⁡(3)Q^{max}_{\mathcal{V}}\in\operatorname{SO}(3) in Φ(fV)\Phi(f_{\mathcal{V}}), the maximum will be found at QTmax=RQVmaxQ^{max}_{\mathcal{T}}=RQ^{max}_{\mathcal{V}} in the rotated feature map. Then, by letting g(V)=QVmaxg(\mathcal{V})=Q^{max}_{\mathcal{V}}, we get g(T)=RQVmaxg(\mathcal{T})=RQ^{max}_{\mathcal{V}}, which satisfies (4) and (5). Therefore, we realize function gg by a Spherical CNN and we utilize the arg max⁡\operatorname*{arg\,max} operator on the feature map computed by the last correlation layer to define its output. In principle, equivariance alone would guarantee to satisfy (4) and (5). Unfortunately, while for continuous functions the network is exactly equivariant, this does not hold for its discretized version, mainly due to feature map rotation, which is exact only for bandlimited functions . Moreover, equivariance to rotations does not hold for altered versions of the same cloud, e.g. when a part of it is occluded due to view-point changes. We tackle these issues using a self-supervised loss computed on the extracted rotations when aligning a pair of point clouds to guide the learning, and an ad-hoc augmentation to increase the robustness to occlusions. Through the use of a soft-argmax layer, we can back-propagate the loss gradient from the estimated rotations to the positions of the maxima we extract from the feature maps and to the filters, which overall lets the network learn a robust gg function.

Training pipeline: An illustration of the Compass training pipeline is shown in Figure 2. During training, our objective is to strengthen the equivariance property of the Spherical CNN, such that the locations selected on the feature maps by the arg max⁡\operatorname*{arg\,max} function vary consistently between rotated versions of the same point cloud. To this end, we train our network with two streams in a Siamese fashion . In particular, given V\mathcal{V}, T∈P\mathcal{T}\in\mathcal{P}, with T=RV\mathcal{T}=R\mathcal{V} and RR a known random rotation matrix, the first branch of the network computes the aligning rotation matrix for V\mathcal{V}, RV=g(V)−1R_{\mathcal{V}}=g(\mathcal{V})^{-1}, while the second branch the aligning rotation matrix for T\mathcal{T}, RT=g(T)−1R_{\mathcal{T}}=g(\mathcal{T})^{-1}. Should the feature maps on which the two maxima are extracted be perfectly equivariant, it would follow that RT=RRV=RT⋆R_{\mathcal{T}}=R{R_{\mathcal{V}}}=R_{\mathcal{T}}^{\star}. For that reason, the degree of misalignment of the maxima locations can be assessed by comparing the actual rotation matrix predicted by the second branch, RTR_{\mathcal{T}}, to the ideal rotation matrix that should be predicted, RT⋆R_{\mathcal{T}}^{\star}. We can thus cast our learning objective as the minimization of a loss measuring the distance between these two rotations. A natural geodesic metric on the SO⁡(3)\operatorname{SO}(3) manifold is given by the angular distance between two rotations . Indeed, any element in SO⁡(3)\operatorname{SO}(3) can be parametrized as a rotation angle around an axis. The angular distance between two rotations parametrized as rotation matrices RR and SS is defined as the angle that parametrizes the rotation SRTSR^{T} and corresponds to the length along the shortest path from RR to SS on the SO⁡(3)\operatorname{SO}(3) manifold . Thus, our loss is given by the angular distance between RTR_{\mathcal{T}} and RT⋆R_{\mathcal{T}}^{\star}:

As our network has to predict a single canonicalizing rotation, we apply the loss once, i.e. only to the output of the last layer of the network.

Soft-argmax: The result of the arg max⁡\operatorname*{arg\,max}{} operation on a discrete SO⁡(3)\operatorname{SO}(3) feature map returns the location i,j,ki,j,k along the α,β,γ\alpha,\beta,\gamma dimensions corresponding to the ZYZ Euler angles, where the maximum correlation value occurs. To optimize the loss in (6), the gradients w.r.t. the i,j,ki,j,k locations of the feature map where the maxima are detected have to be computed. To render the arg max⁡x\operatorname*{arg\,max}{}x operation differentiable we add a soft-argmax operator following the last SO⁡(3)\operatorname{SO}(3) layer of the network. Let us denote as Φ(fV)\Phi(f_{\mathcal{V}}) the last SO⁡(3)\operatorname{SO}(3) feature map computed by the network for a given input point cloud V\mathcal{V}. A straightforward implementation of a soft-argmax layer to get the coordinates CR=(i,j,k)C_{R}=(i,j,k) of the maximum in Φ(fV)\Phi(f_{\mathcal{V}}) is given by

where softmax(⋅)(\cdot) is a 3D spatial softmax. The parameter τ\tau controls the temperature of the resulting probability map and (i,j,k)(i,j,k) iterate over the SO⁡(3)\operatorname{SO}(3) coordinates. A soft-argmax operator computes the location CR=(i,j,k)C_{R}=(i,j,k) as a weighted sum of all the coordinates (i,j,k)(i,j,k) where the weights are given by a softmax of a SO⁡(3)\operatorname{SO}(3) map Φ\Phi. Experimentally, this proved not effective. As a more robust solution, we scale the output of the softmax according to the distance of each (i,j,k)(i,j,k) bin from the feature map arg max⁡\operatorname*{arg\,max}. To let the bins near the arg max⁡\operatorname*{arg\,max}{} contribute more in the final result, we smooth the distances by a Parzen function yielding a maximum value in the bin corresponding to the arg max⁡\operatorname*{arg\,max} and decreasing monotonically to .

Learning to handle occlusions: In real-world settings, rotation of an object or scene (i.e. a viewpoint change) naturally produces occlusions to the viewer. Recalling that the second branch of the network operates on T\mathcal{T}, a randomly rotated version of V\mathcal{V}, it is possible to improve

robustness of the network to real-world occlusions and missing parts by augmenting T\mathcal{T}. A simple way to handle this problem is to randomly select a point from T\mathcal{T} and delete some of its surrounding points. In our implementation, this augmentation happens with an assigned probability. T\mathcal{T} is divided in concentric spherical shells, with the probability for the random point to be selected in a shell increasing with its distance from the center of T\mathcal{T}. Additionally, the number of removed points around the selected point is a bounded random percentage of the total points in the cloud. An example can be seen in Figure 3.

Network Architecture: The network architecture comprises 1 S2S^{2} layer followed by 3 SO⁡(3)\operatorname{SO}(3) layers, with bandwidth B=24B=24 and the respective number of output channels are set to 40, 20, 10, 1. The input spherical signal is computed with K=4K=4 channels.

Applications of Compass

We evaluate Compass on two challenging tasks. The first one is the estimation of a canonical orientation of local surface patches, a key step in creating rotation-invariant local 3D descriptors . In the second task, the canonical orientation provided by Compass is instead used to perform highly effective rotation-invariant shape classification by leveraging a simple PointNet classifier. The source code for training and testing Compass is available at https://github.com/CVLAB-Unibo/compass.

where I(⋅)\text{I}(\cdot) is an indicator function, (⋅)(\cdot) denotes the dot product between two vectors, and ρ\rho is a threshold on the angle between the corresponding axes, 0.970.97 in our experiments. Rep measures the percentage of reference frames which are aligned, i.e. differ only by a small angle along all axes, between the two views. The final value of Rep for a given model is computed by averaging on all the pairs.

Test-time adaptation: Due to the self-supervised nature of Compass, it is possible to use the test set to train the network without incurring in data snooping, since there is no external ground-truth information involved. This test-time training can be carried out very quickly, right before the test, to adapt the network to unseen data and increase its performance, especially in transfer learning scenarios. This is common practice with self-supervised approaches .

Datasets: We conduct experiments on three heterogeneous publicly available datasets: 3DMatch , ETH , and Stanford Views . 3DMatch is the reference benchmark to assess learned local 3D descriptors performance in registration applications . It is a large ensemble of existing indoor datasets. Each fragment is created fusing 50 consecutive depth frames of an RGB-D sensor. It contains 62 scenes, split into 54 for training and 8 for testing. ETH is a collection of outdoor landscapes acquired in different seasons with a laser scanner sensor. Finally, Stanford Views contains real scans of 4 objects, from the Stanford 3D Scanning Repository , acquired with a laser scanner.

Experimental setup: We train Compass on 3DMatch following the standard procedure of the benchmark, with 48 scenes for training and 6 for validation. From each point cloud, we uniformly pick a keypoint every 1010 cm, the points within 3030 cm are used as local surface patch and fed to the network. Once trained, the network is tested on the test split of 3DMatch. The network learned on 3DMatch is tested also on ETH and Stanford Views, using different radii to account for the different sizes of the models in these datasets: respectively 100100 cm and 1.51.5 cm. We also apply test-time adaptation on ETH and Stanford Views: the test set is used for a quick 2-epoch training with a 20% validation split, right before being used to assess the performance of the network. We use Adam as optimizer, with 0.001 as the learning rate when training on 3DMatch and for test-time adaptation on Stanford Views, and 0.0005 for adaptation on ETH. We compare our method with recent and established LRFs proposals: GFrames, TOLDI, a variant of TOLDI recently proposed in that we refer to here as 3DSN, FLARE , and SHOT . For all methods we use the publicly available implementations. However, the implementation provided for GFrames could not process the large point clouds of 3DMatch and ETH due to memory limits, and we can show results for GFrames only on Stanford Views.

Results: The first column of Table 1 reports Rep on the 3DMatch test set. Compass outperforms the most competitive baseline FLARE, with larger gains over the other baselines. Results reported in the second column for ETH and the third column for Stanford Views confirm the advantage of a data-driven model like Compass over hand-crafted proposals: while the relative rank of the baselines changes according to which of the assumptions behind their design fits better the traits of the dataset under test, with SHOT taking the lead on ETH and the recently introduced GFrames on Stanford Views, Compass consistently outperforms them. Remarkably, this already happens when using pure transfer learning for Compass, i.e. the network trained on 3DMatch: in spite of the large differences in acquisition modalities and shapes of the models between training and test time, Compass has learned a robust and general notion of canonical orientation for a local patch. This is also confirmed by the slight improvement achieved with test-time augmentation, which however sets the new state of the art on these datasets. Finally, we point out that Compass extracts the canonical orientation for a patch in 17.85ms.

2 Rotation-invariant Shape Classification

Problem formulation: Object classification is a central task in computer vision applications, and the main nuisance that methods processing 3D point clouds have to withstand is rotation. To show the general applicability of our proposal and further assess its performance, we wrap Compass in a shape classification pipeline. Hence, in this experiment, Compass is used to orient full shapes rather than local patches. To stress the importance of correct rotation neutralization, as shape classifier we rely on a simple PointNet , and Compass is employed at train and test time to canonically orient shapes before sending them through the network.

Datasets: We test our model on the ModelNet40 shape classification benchmark. This dataset has 12,311 CAD models from 40 man-made object categories, split into 9,843 for training and 2,468 for testing. In our trials, we actually use the point clouds sampled from the original CAD models provided by the authors of PointNet. We also performed a qualitative evaluation of the transfer learning performance of Compass by orienting clouds from the ShapeNet dataset.

Experimental setup: We train Compass on ModelNet40 using 8,192 samples for training and 1,648 for validation. Once Compass is trained, we train PointNet following the settings in , disabling t-nets, and rotating the input point clouds to reach the canonical orientation learned by Compass. We followed the protocol described in to assess rotation-invariance of the selected methods: we do not augment the dataset with rotated versions of the input cloud when training PointNet; we then test it with the original test clouds, i.e. in the canonical orientation provided by the dataset, and by arbitrary rotating them. We use Adam as optimizer, with 0.001 as the learning rate.

Results: Results are reported in Table 2. Results for all the baselines come from . PointNet fails when trained without augmenting the training data with random rotations and tested with shapes under arbitrary rotations. Similarly, in these conditions most of the state-of-the-art methods cannot generalize to unseen rotations. If, however, we first neutralize the orientation by Compass and then we run PointNet, it gains almost 60 points and achieves 72.20 accuracy, outperforming the state-of-the-art on the arbitrarily rotated test set. This shows the feasibility and the effectiveness of pursuing rotation-invariant processing by canonical orientation estimation. It is also worth observing how, in the simplified scenario where the input data is always under the same orientation (NR), a plain PointNet performs better than Compass+PointNet. Indeed, as the T-Net is trained end-to-end with PointNet, it can learn that the best orientation in the simplified scenario is the identity matrix. Conversely, Compass performs am unneeded canonicalization step that may only hinder performance due to its errors.

In Figure 4, we present some models from ModelNet40, randomly rotated and then oriented by Compass. The models estimate a very consistent canonical orientation for each object class, despite the large shape variations within the classes.

Finally, to assess the generalization abilities of Compass for full shapes as well, we performed qualitative transfer learning tests on the ShapeNet dataset, reported in Figure 4. Even if there are different geometries, the model trained on ModelNet40 is able to generalize to an unseen dataset and recovers a similar canonical orientation for the same object.

Conclusions

We have presented Compass, a novel self-supervised framework to canonically orient 3D shapes that leverages the equivariance property of Spherical CNNs. Avoiding explicit supervision, we let the network learn to predict the best-suited orientation for the underlying surface geometries. Our approach robustly handles occlusions thanks to an effective data augmentation. Experimental results demonstrate the benefits of our approach for the tasks of definition of a canonical orientation for local surface patches and rotation-invariant shape classification. Compass demonstrates the effectiveness of learning a canonical orientation in order to pursue rotation-invariant shape processing, and we hope it will raise the interest and stimulate further studies about this approach.

While in this work we evaluated invariance to global rotation according to the protocol used in to perform a fair comparison with the state-of-the-art method , it would also be interesting to investigate on the behavior of Compass and the competitors when trained on the full spectrum of SO(3) rotations as done in . This is left as future work.

Broader Impact

In this work we presented a general framework to canonically orient 3D shapes based on deep-learning. The proposed methodology can be especially valuable for the broad spectrum of vision applications that entail reasoning about surfaces. We live in a three-dimensional world: cognitive understanding of 3D structures is pivotal for acting and planning.

Acknowledgments

We would like to thank Injenia srl and UTFPR for partly supporting this research work.

References