Learning SO(3) Equivariant Representations with Spherical CNNs
Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, Kostas Daniilidis
Introduction
One of the reasons for the tremendous success of convolutional neural networks (CNNs) is their equivariance to translations in euclidean spaces and the resulting invariance to local deformations. Invariance with respect to other nuisances has been traditionally addressed with data augmentation while non-euclidean inputs like point-clouds have been approximated by euclidean representations like voxel spaces. Only recently, equivariance has been addressed with respect to other groups and CNNs have been proposed for manifolds or graphs .
Equivariant networks retain information about group actions on the input and on the feature maps throughout the layers of a network. Because of their special structure, feature transformations are directly related to spatial transformations of the input. Such equivariant structures yield a lower network capacity in terms of unknowns than alternatives like the Spatial Transformer where a canonical transformation is learnt and applied to the original input.
In this paper, we are primarily interested in analyzing 3D data for alignment, retrieval or classification. Volumetric and point cloud representations have yielded translation and scale invariant approaches: Normalization of translation and scale can be achieved by setting the object’s origin to its center and constraining its extent to a fixed constant. However, 3D rotations remain a challenge to current approaches (Figure 2 illustrates how classification performance for conventional methods suffers when arbitrary rotations are introduced).
It is natural then to apply pooling in the spectral domain. Spectral pooling has the advantage that it retains equivariance while spatial pooling on the sphere is only approximately equivariant. We also propose a weighted averaging pooling where the weights are proportional to the cell area. The only reason to return to the spatial domain is the rectifying nonlinearity, which is a pointwise operator.
We perform 3D retrieval, classification, and alignment experiments. Our aim is to show that we can achieve near state of the art performance with a much lower network capacity, which we achieve for the SHREC’17 contest and ModelNet40 datasets.
Our main contributions can be summarized as follows:
We propose the first neural network based on spherical convolutions.
We introduce pooling and parameterization of filters in the spectral domain, with enforced spatial localization and capacity independent of the resolution.
Our network has much lower capacity than non-spherical networks applied on 3D data without sacrificing performance.
We start with the related work, then introduce the mathematics of group and in particular sphere convolutions, and details of our network. Last, we perform extensive experiments on retrieval, classification, and alignment.
Related work
We will start describing related work on group equivariance, in particular equivariance on the sphere, then delve into CNN representations for 3D data.
Methods for enabling equivariance in CNNs can be divided in two groups. In the first, equivariance is obtained by constraining filter structure similarly to Lie generator based approaches . Worral et al. use filters derived from the complex harmonics achieving both rotational and translational equivariance. The second group requires the use of a filter orbit which is itself equivariant to obtain group equivariance. Cohen and Welling convolve with the orbit of a learned filter and prove the equivariance of group-convolutions and preservation of rotational equivariance in the presence of rectification and pooling. Dieleman et al. process elements of the image orbit individually and use the set of outputs for classification. Gens and Domingos produce maps of finite-multiparameter groups, Zhou et al. and Marcos et al. use a rotational filter orbit to produce oriented feature maps and rotationally invariant features, and Lenc and Vedaldi propose a transformation layer which acts as a group-convolution by first permuting then transforming by a linear filter.
Recently, a body of work on Graph Convolutional Networks (GCN) has emerged. There are two threads within this space, spectral and spatial . These approaches learn filters on irregular but structured graph representations. These methods differ from ours in that we are looking to explicitly learn equivariant and invariant representations for 3D-data modeled as spherical functions under rotation. While such properties are difficult to construct for general manifolds, we leverage the group action of rotations on the sphere.
Most similar to our approach and developed in parallelthe first version of this work was submitted to CVPR on 11/15/2017, shortly after we became aware of Cohen et al. ICLR submission on 10/27/2017. is , which uses spherical correlation to map spherical inputs to features on (3), then processed with a series of convolutions on (3). The main difference is that we use spherical convolutions, which are potentially one order of magnitude faster, with smaller (one fewer dimension) filters and feature maps. In addition, we enforce smoothness in the spectral domain that results in better localization of the receptive fields on the sphere and we perform pooling in two different ways, either as a low-pass in the spectral domain or as a weighted averaging in the spatial domain. Moreover, our method outperforms in the SHREC’17 benchmark.
Spherical representations for 3D-data are not novel and have been used for retrieval tasks before the deep learning era because of their invariance properties and efficient implementation of spherical correlation . In 3D deep learning, the most natural adaptation of 2D methods was to use a voxel-grid representation of the 3D object and amend the 2D CNN framework to use collections of 3D filters for cascaded processing in the place of conventional 2D filters. Such approaches require a tremendous amount of computation to achieve very basic voxel resolution and need a much higher capacity.
Several attempts have been made to use CNNs to produce discriminative representations from volumetric data. 3D ShapeNets and VoxNet propose a fully-volumetric network with 3D convolutional layers followed by fully-connected layers. Qi et al. observe significant overfitting when attempting to train the aforementioned end-to-end and choose to amend the technique using subvolume classification as an auxiliary task, and also propose an alternate 3D CNN which learns to project the volumetric representation to a 2D representation, then processed using a conventional 2D CNN architecture. Even with these adaptations, Qi et al. are challenged by overfitting and suggest augmentation in the form of orientation pooling as a remedy. Qi et al. also present an attempt to train a neural network that operates directly on point clouds. Currently, the most successful approaches are view-based, operating in rendered views of the 3D object . The high performance of these methods is in part due to the use of large pre-trained 2D CNNs (on ImageNet, for instance).
Preliminaries
Consideration of symmetries, in particular rotational symmetries, naturally evokes notions of the Fourier Transform. In the context of deriving rotationally invariant representations, the Fourier Transform is particularly appealing since it exhibits invariance to rotational deformations up to phase (a truly invariant representation can be achieved through application of the modulus operator).
To leverage this property for 3D shape analysis, it is necessary to construct a rotationally equivariant representation of our 3D input. For a group and function , is said to be equivariant to transformations when
where acts on elements of and is the corresponding group action which transforms elements of . If , . A straightforward example of an equivariant representation is an orbit. For an object , its orbit with respect to the group is defined
Through this example it is possible to develop an intuition into the equivariance of the group convolution; convolution can be viewed as the inner-products of some function with all elements of the orbit of a “flipped” filter . Formally, the group convolution is defined as
The group convolution can be shown to be equivariant. For any ,
2 Spherical harmonics
Following directly the preliminaries above, we can define convolution of spherical signal by a spherical filter with respect to the group of 3D rotations :
where is north pole on the sphere.
To implement (6), it is desirable to sample the sphere with well-distributed and compact cells with transitivity (rotations exist which bring cells into coincidence). Unfortunately, such a discretization does not exist . Neither the familiar sampling by latitude and longitude nor the uniformly distributed sampling according to Platonic solids satisfies all constraints. These issues are compounded with the eventual goal of performing cascaded convolutions on the sphere.
To circumvent these issues, we choose to evaluate the spherical convolution in the spectral domain. This is possible as the machinery of Fourier analysis has extended the well-known convolution theorem to functions on the sphere: the Spherical Fourier transform of a convolution is the pointwise product of Spherical Fourier transforms (see for further details). The Fourier transform and its inverse are defined on the sphere as follows :
To compute the convolution of a signal with a filter , we first expand and into their spherical harmonic basis (8), second compute the pointwise product (9), and finally invert the spherical harmonic expansion (7).
It is important to note that this definition of spherical convolution is unique from spherical correlation which produces an output response on (3). Convolution here can be seen as marginalizing the angle responsible for rotating the filter about its north pole, or equivalently considering zonal filters on the sphere.
3 Practical considerations and optimizations
To evaluate the SFT, we use equiangular samples on the sphere according to the sampling theorem of
3.2 Leveraging symmetry:
Method
Figure 3 shows an overview of our method. We define a block as one spherical convolutional layer, followed by optional pooling, and nonlinearity. A weighted global average pooling is applied at the last layer to obtain an invariant descriptor. This section details the architectural design choices.
In this section, we define the filter parameterization. One possible approach would be to define a compact support around one of the poles and learn the values for each discrete location, setting the rest to zero. The downside of this approach is that there are no guarantees that the filter will be bandlimited. If it is not, the SFT will be implicitly bandlimiting the signal, which causes a discrepancy between the parameters and the actual realization of the filters.
To avoid this problem, we parameterize the filters in the spectral domain. In order to compute the convolution of a function and a filter , only the SFT coefficients of order of are used. In the spatial domain, this implies that for any , there is always a zonal filter (constant value per latitude) , such that . Thus, it only makes sense to learn zonal filters.
The spectral parameterization is also faster because it eliminates the need to compute the filter SFT, since the filters are defined in the spectral domain, which is the same domain where the convolution computed.
A first approach is to parameterize the filters by all SFT coefficients of order . For example, given inputs, the maximum bandwidth is , so there are parameters to be learned (). A downside is that the filters may not be local; however, locality may be learned.
1.2 Localized filters:
From Parseval’s theorem and the derivative rule from Fourier analysis we can show that spectral smoothness corresponds to spatial decay. This is used in the construction of graph-based neural networks , and also applies to the filters spanned by the family of spherical harmonics of order zero ().
2 Pooling
The conventional spatial max pooling used in CNNs has two drawbacks in Spherical CNNs: (1) need an expensive ISFT to convert back to spatial domain, and (2) equivariance is not completely preserved, specially because of unequal cell areas from equiangular sampling. Weighted average pooling (WAP) takes into account the cell areas to mitigate the latter, but is still affected by the former.
We introduce the spectral pooling (SP) for Spherical CNNs. If the input has bandwidth , we remove all coefficients with degree larger or equal than (effectively, a lowpass box filter). Such operation is known to cause ringing artifacts, which can be mitigated by previous smoothing, although we did not find any performance advantage in doing so. Note that spectral pooling was proposed before for conventional CNNs .
We found that spectral pooling is significantly faster, reduces the equivariance error, but also reduces classification accuracy. The choice between SP and WAP is application-dependent. For example, our experiments show SP is more suitable for shape alignment, while WAP is better for classification and retrieval. Table 5 shows the performance for each method.
3 Global pooling
In fully convolutional networks, it is usual to apply a global average pooling at the last layer to obtain a descriptor vector, where each entry is the average of one feature map. We use the same idea; however, the equiangular spherical sampling results in cells of different areas, so we compute a weighted average instead, where a cell’s weight is the sine of its latitude. We denote it Weighted Global Average Pooling (WGAP). Note that the WGAP is invariant to rotation, therefore the descriptor is also invariant. Figure 5 shows such descriptors.
4 Architecture
Our main architecture has two branches, one for distances and one for surface normals. This performs better than having two input channels and slightly better than having two separate voting networks for distance and normals. Each branch has 8 spherical convolutional layers, and channels per layer. Pooling and feature concatenation of one branch into the other is performed when the number of channels increase. WGAP is performed after the last layer, which is then projected into the number of classes.
Experiments
The greatest advantage of our model is inherent equivariance to (3); we focus the experiments in problems that benefit from it; namely, shape classification and retrieval in arbitrary orientations, and shape alignment.
We chose problems related to 3D shapes due to the availability of large datasets and published results on them; our method would also be applicable to any kind of data that can be mapped to the sphere (e.g. panoramas).
3D shapes are usually represented by mesh or voxel grid, which need to be converted to spherical functions. Note that the conversion function itself must be equivariant to rotations; our learned representation will not be equivariant if the input is pre-processed by a non-equivariant function.
Given a mesh or voxel grid, we first find the bounding sphere and its center. Given a desired resolution , we cast equiangular rays from the center, and obtain the intersections between each ray and the mesh/voxel grid. Let be the distance from the center to the farthest point of intersection, for a ray at direction . The function on the sphere is given by .
For mesh inputs, we also compute the angle between the ray and the surface normal at the intersecting face, giving a second channel .
Note that this representation is suitable for star-shaped objects, defined as objects that contain an interior point from where the whole boundary is visible. Moreover, the center of the bounding sphere must be one of such points. In practice, we do not check if these conditions hold – even if the representation is ambiguous or non-invertible, it is still useful.
1.2 Training:
We train using ADAM, for 48 epochs, initial learning rate of , which is divided by 5 on epochs 32 and 40.
We make use of data augmentation for training, performing rotations, anisotropic scaling and mirroring on the meshes, and adding jitter to the bounding sphere center when constructing the spherical function. Note that, even though our learned representation is equivariant to rotations, augmenting the inputs with rotations is still beneficial due to interpolation and sampling effects.
2 3D object classification
This section shows classification performance on ModelNet40 . Three modes are considered: (1) trained and tested with azimuthal rotations (z/z), (2) trained and tested with arbitrary rotations ((3)/(3)), and (3) trained with azimuthal and tested with arbitrary rotations (z/(3)).
Table 1 shows the results. All competing methods suffer a sharp drop in performance when arbitrary rotations are present, even if they are seen during training. Our model is more robust, but there is a noticeable drop for mode 3, attributed to sampling effects. Since we use equiangular sampling, the cell area varies with latitude. Rotations around preserve latitude, so regions at same height are sampled at same resolution during training, but not during test. We believe this can be improved by using equal-area spherical sampling.
We evaluate competing methods using default settings of their published code. The volumetric and point cloud based methods cannot generalize to unseen orientations (z/(3)). The multi-view methods can be seen as a brute force approach to equivariance; and MVCNN generalizes to unseen orientations up to a point. Yet, the Spherical CNN outperforms it, even with orders of magnitude fewer parameters and faster training. Interestingly, RotationNet , which holds the current state-of-the-art on ModelNet40 classification, fails to generalize to unseen rotations, despite being multi-view based.
Equivariance to (3) is unneeded when only azimuthal rotations are present (z/z); the full potential of our model is not exercised in this case.
3 3D object retrieval
We run retrieval experiments on ShapeNet Core55 , following the SHREC’17 3D shape retrieval rules , which includes random (3) perturbations.
The network is trained for classification on the 55 core classes (we do not use the subclasses), with an extra in-batch triplet loss (from ) to encourage descriptors to be close for matching categories and far for non-matching.
The invariant descriptor is used with a cosine distance for retrieval. We first compute a threshold per class that maximizes the training set F-score. For test set retrieval, we return elements whose distances are below their class threshold and include all elements classified as the same class as the query. Table 2 shows the results. Our model matches the state of the art performance (from ), with significantly fewer parameters, smaller input size, and no pre-training.
4 Shape alignment
Our learned equivariant feature maps can be used for shape alignment using spherical correlation. Given two shapes from the same category (not necessarily the same instance), under arbitrary orientations, we run them through the network and collect the feature maps at some layer. We compute the correlation between each pair of corresponding feature maps, and add the results. The maximum value of the correlation function (which takes inputs on ) corresponds to the rotation that aligns both shapes .
Features from deeper layers are richer and carry semantic value, but are at lower resolution. We run an experiment to determine the performance of the shape alignment per layer, while also comparing with the spherical correlation done at the network inputs (not learned).
We select categories from ModelNet10 that do not have rotational symmetry so that the ground truth rotation is unique and the angular error is measurable. These categories are: bed, sofa, toilet, chair. Only entries from the test set are used. Results are in Table 3, while Figure 6 shows some examples. Results show that the learned features are superior to the handcrafted spherical shape representation for this task, and best performance is achieved by using intermediate layers. The resolution at conv4 is , which corresponds to cell dimensions up to , so we cannot expect errors much lower than this.
5 Equivariance error analysis
Even though spherical convolutions are equivariant to (3) for bandlimited inputs, and spectral pooling preserves bandlimit, there are other factors that may introduce equivariance errors. We quantify these effects in this section.
We feed each entry in the test set and one random rotation to the network, then apply the same rotation to the feature maps and measure the average relative error. Table 4 shows the results. The pointwise nonlinearity does not preserve bandlimit, and cause equivariance errors (rows 1, 4). The mesh to sphere map is only approximately equivariant, which can be mitigated with larger input dimensions (input column for rows 1, 5). Error is smaller when the input is bandlimited (rows 1, 7). Spectral pooling is exactly equivariant, while max-pooling introduces higher frequencies and has larger error than WAP (rows 1, 2, 3). Error for an untrained model demonstrates that the equivariance is by design and not learned (row 6). Note that the error is smaller because the learned filters are usually high-pass, which increase the pointwise relative error. A linear model with bandlimited inputs has zero equivariance error, as expected (row 8).
Note that even conventional planar CNNs will exhibit a degree of translational equivariance error introduced by max pooling and discretization.
6 Ablation study
In this section we evaluate numerous variations of our method to determine the sensitivity to design choices. First, we are interested in assessing the effects from our contributions SP, WAP, WGAP, and localized filters. Second, we are interested in understanding how the network size affects performance. Results show that the use of WAP, WGAP, and localized filters significantly improve performance, and also that further performance improvements can be achieved with larger networks. In summary, factors that increase bandwidth (e.g. max-pooling) also increase equivariance error and may reduce accuracy. Global operations in early layers (e.g. non-local filters) escape the receptive field and reduce accuracy.
Conclusion
We presented Spherical CNNs, which leverage spherical convolutions to achieve equivariance to (3) perturbations. The network is applied to 3D object classification, retrieval, and alignment, but has potential applications in spherical images such as panoramas, or any data that can be represented as a spherical function. We show that our model can naturally handle arbitrary input orientations, requiring relatively few parameters and small input sizes.
We are grateful for support through the following grants: NSF-DGE-0966142 (IGERT), NSF-IIP-1439681 (I/UCRC), NSF-IIS-1426840, NSF-IIS-1703319, NSF MRI 1626008, ARL RCTA W911NF-10-2-0016, ONR N00014-17-1-2093, and by Honda Research Institute.