Vector Neurons: A General Framework for SO(3)-Equivariant Networks

Congyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard, Andrea Tagliasacchi, Leonidas Guibas

Introduction

With the proliferation of lower-cost depth sensors, learning on 3D data has seen rapid progress in recent years. Of particular interest are pointcloud networks, such as PointNet or ACNe that fully respect the inherent set symmetry – that point sets are not ordered – by incorporating order-invariant and/or order-equivariant layers. Yet, there are other important symmetries that have been less perfectly addressed in the context of pointcloud processing, with 3D rotations being a prime example. Consider a scenario where one scans an object using their LIDAR-equipped phone to retrieve similar objects. Clearly, the global object pose should not affect the query result. PointNet uses spatial transformer layers , which only attain approximate pose invariance while also requiring extensive augmentation at train time.

To avoid an exhaustive data augmentation with all possible rotations, there is a need for network layers that are equivariant to both order and SO(3) symmetries. Recently, two approaches have been introduced to tackle this setting: Tensor Field Networks and SE(3)-Transformers . While guaranteeing equivariance by construction, both frameworks involve an intricate formulation and are hard to incorporate into existing pipelines as they are restricted to convolutions and rely on relative positions of adjacent points.

In this work, we address these issues by proposing a simple, lightweight framework to build SO(3) equivariant and invariant pointcloud networks. A core ingredient in our framework is a Vector Neuron (VN) representation, extending classical scalar neurons to 3D vectors. Consequently, instead of latent vector representations which can be views as ordered sequences of scalars, we deploy latent matrix representations which can be viewed as (ordered) sequences of 3-vectors. Such a representation supports a direct mapping of rotations applied to the input pointcloud to intermediate layers. This is in contrast to more complex solutions based on Wigner D-matrices . Another appealing property of VN representations is that they remain equivariant to linear layers by construction. The challenge in building a fully-equivariant network lies in the non-linear activations. In particular, standard neuron-wise activation functions such as ReLU will not commute with a rotation operation. A key contribution in this work is a 3D generalization of classical activation functions by implementing them through a learned direction. For example, when applied to a vector neuron, a standard fixed direction ReLU activation would simply truncate the half-plane that points in its opposite direction. Instead, dynamically predicting an activation direction in a linear data-dependent fashion allows us to guarantee equivariance. We further provide an invariant pooling operation as well as normalization layers, which altogether render our framework compatible with various pointcloud network backbones. To demonstrate its versatility and efficiency, we implemented vector neuron versions of two popular architectures: PointNet and DGCNN, and tested them on three different downstream tasks: classification (permutation invariant and rotation invariant), segmentation (permutation equivariant and rotation invariant), and reconstruction (rotation equivariant on the encoder side, and rotation invariant on the decoder side). Despite its simplicity and lightweight architecture, in all tasks, our VN achieved top performance when tested on randomly rotated shapes compared with other equivariant architectures, and markedly improved performance compared to augmentation-induced equivariance approaches.

We propose a new versatile framework for constructing SO(3)-equivariant pointcloud networks.

Our building blocks are lightweight in terms of the number of learnable parameters and can be easily incorporated into existing network architectures.

We support a variety of learning tasks, in particular, we are the first to demonstrate a 3D equivariant network for 3D reconstruction.

When evaluated on classification and segmentation, our VN version of popular non-equivariant architectures achieve state-of-the-art performance.

Related Work

The lack of robustness to rotation of classical deep learning architectures for pointcloud processing like PointNet , PointNet++ , Dynamic Graph CNN (DGCNN) , PCNN , PointCNN (and many others) has driven interest for rotation invariant and equivariant designs. In recent years the field of rotation invariant and equivariant deep learning for geometry processing has been rapidly developing. In what follows, we briefly review methods that achieve invariance and equivariance, as well as overview those that achieve equivariance via pose estimation.

Rotation invariance is a desirable property for tasks like shape classification or segmentation. Many rotation invariant architectures have been proposed to address these issues. For example, introduce cleverly designed rotation invariant operations. GC-Conv relies on multi-scale reference frames based on PCA. RI-Framework and LGR-Net pairs local invariant information with global context. Some works like LGR-Net use surface normals in addition to the points coordinates. SFCNN proposes an approach similar to multi-view by mapping input pointclouds to a sphere and performing operations on the sphere. Other works like rely on more principled approaches borrowing tools from equivariant deep learning.

Rotation equivariant methods

Equivariance via pose estimation

Qi et al. achieved approximate pose equivariance by factoring out SO(3) transformations through object pose estimation. Most works in the literature study instance-level pose estimation, where the ground-truth canonical pose of the 3D CAD models corresponding to the input pointcloud is available . More recently Wang et al. introduced category-level pose estimation, and extension to articulated objects has also been proposed . While both these methods need explicit 2D-to-3D supervision, relaxing supervision is possible by borrowing ideas from Transforming Auto-Encoders . However, while Sun et al. learn category-level as well as multi-category pose estimation in a fully unsupervised fashion, the underlying equivariant backbone is only equivariant by augmentation.

Method

where θ\theta represents learnable parameters.

where we interpret the application of the rotation matrix to the set as VR={ViR}i=1N{\mathcal{V}}R=\{{\bm{V}}_{i}R\}_{i=1}^{N}. To facilitate equivariance in standard pointcloud network architectures, we construct VN layers following traditional designs via a combination of a linear map (Section 3.1) followed by a per-neuron non-linearity (Section 3.2). We additionally introduce equivariant pooling (Section 3.3) and normalization layers (Section 3.4). With these building blocks we are able to assemble a rich variety of complex neural networks in equivariance, including the most basic VN Multi-Layer Perceptron (VN-MLP) as a sequence of alternating linear and non-linear layers.

yielding the desired equivariance property. Note that we omit a bias term as an addition of a constant vector that would interfere with equivariance. Further, note that while this layer is SO(3) equivariant, we can achieve SE(3) equivariance by centering V{\bm{V}} at the origin. Finally, depending on the setting, W{\mathbf{W}} may or may not be shared across the elements V{\bm{V}} of V{\mathcal{V}}.

2 Non-linear layers – Figure 3 and Figure 4

Per-neuron non-linearity is key to the representation power of neural networks. As evident from recent literature, especially useful are functions that split the input domain into two half spaces and map them differently (e.g. ReLU, leaky-ReLU, ELU, etc.). In the case of VN, a 3D version of these non-linearities, V′=fReLU(V){\bm{V}}^{\prime}=f_{\text{ReLU}}({\bm{V}}), is needed. Yet, committing to a fixed frame (i.e., one that does not depend on the input pose) like the standard coordinate system would violate equivariance. Instead, we propose to dynamically predict a direction from the input vector-list feature. We then generalize the classical ReLU by truncating the portion of a vector that points into the negative half-space of the learned direction.

resulting in an output vector-list: fReLU(V)=[v′]c=1Cf_{\text{ReLU}}({\bm{V}})=[{\bm{v}}^{\prime}]_{c=1}^{C} In practice, when computing for the unit direction vector k/∥k∥{\bm{k}}/\|{\bm{k}}\| we implement k/(∥k∥+ε){\bm{k}}/(\|{\bm{k}}\|+\varepsilon) with a small margin ε\varepsilon in the denominator to avoid division by zero at the origin.

As illustrated in Figure 3, q{\bm{q}} can be decomposed into two components: q∥{\bm{q}}_{\parallel} and q⊥{\bm{q}}_{\perp} that are parallel and orthogonal to k{\bm{k}}, respectively. Analogous to the standard scalar ReLU, we apply the nonlinear function to q∥{\bm{q}}_{\parallel} along the direction k{\bm{k}} by clipping q∥{\bm{q}}_{\parallel} to zero, while keeping q⊥{\bm{q}}_{\perp} unchanged. Other types of split-case functions (e.g. leaky-ReLU) follow immediately from this definition. We discuss these and other types of non-linearities in the supplementary material.

It is easy to verify that fReLUf_{\text{ReLU}} is rotation equivariant. In particular, both q{\bm{q}} and k{\bm{k}} are linear maps of V{\bm{V}} and thus commute with a rotation matrix as discussed in (4). Moreover, the inner-product term in the second case would cancel out an orthogonal matrix ⟨qR,kR⟩=⟨q,k⟩\left\langle{\bm{q}}{\bm{R}},{\bm{k}}{\bm{R}}\right\rangle=\left\langle{\bm{q}},{\bm{k}}\right\rangle resulting in a scalar multiplication of a k{\bm{k}}, which is again equivariant.

3 Pooling layers

Pooling is widely used when aggregating local/global neighbourhood information, either spatially (e.g. PointNet++) or by feature similarity (e.g. DGCNN). While mean pooling is a linear operation that respects rotation equivariance, we also define a VN max pooling layer as a counterpart to the classical max pooling on scalars.

and then computing the element of V{\mathcal{V}} that best aligns with K\mathcal{K} and selecting it as our global feature: for each channel c∈[C]c\in[C],

where Vn[c]{\bm{V}}_{n}[c] stands for the vector channel vc∈Vn{\bm{v}}_{c}\in{\bm{V}}_{n}.

Similarly, we can aggregate information locally (local pooling) by grouping kk nearest neighbours in V{\mathcal{V}} and perform the aforementioned pooling seperately for each group.

4 Normalization layers – Figure 5

In contrast to other forms of normalizations, batch normalization aggregates statistics across all batch samples. While technically possible, in the context of rotation equivariant networks, averaging across arbitrarily rotated inputs would not necessarily be meaningful. For example, averaging two input features rotated in opposite directions would zero them out instead of producing that feature in a canonical pose.

We instead apply batch normalization to the invariant component of the vector-list features, by normalizing the 2-norms of the vector-list features.

5 Invariant layers

General invariant architectures are comprised of equivariant layers followed by invariant ones. We now introduce our invariant layer, that can be appended as needed to the output of the equivariant VN layers. Rotation-invariant networks are essential for both classification and segmentation tasks, where the identity of an object or its parts should be invariant to pose.

Note that a specific case of (13) is the inner product of two vectors, in particular the norm of equivariant vector features is rotation invariant.

Finally we define our invariant layer by:

Network Architectures

DGCNN performs a permutation equivariant edge convolution by computing adjacent edge features enm′{\bm{e}}^{\prime}_{nm} followed by a local max pooling:

VN-PointNet

PointNet approximates a permutation symmetric function using

where hh is the same for all xn{\bm{x}}_{n}. Its VN version is written as

Experiments

We evaluate our method on three core tasks in pointcloud processing: classification (Section 5.1), segmentation (Section 5.2), and reconstruction (Section 5.3). In addition to their diversity in the required output, these tasks span different use cases of our proposed equivariant framework: classification and segmentation are rotation-invariant tasks, while reconstruction is rotation-equivariant.

We employed the ModelNet40 and the ShapeNet datasets for evaluation. The ModelNet40 dataset consists of 40 classes with 12,311 CAD models in total. We used 9,843 models for training and the others for testing in the classification task. For the ShapeNet dataset, we followed by using ShapeNet-part for part segmentation, which has 16 shape categories with more than 30,000 models. We also applied the subset of ShapeNet in for shape reconstruction, containing 13 major categories with 50,000 models.

Train/test rotation setup

Network implementations

In classification and segmentation, we implement our VN networks in the identical architectures to their classical counterparts, but with each layer in the shape of ⌊N3⌋×3\lfloor\frac{N}{3}\rfloor\times 3 while the corresponding layer in the scalar network has size NN. This in fact greatly reduces the number of learnable parameters in VN networks, resulting in roughly ⩽2/32=2/9\leqslant 2/3^{2}=2/9 times of parameters compared to the counterpart scalar networks – here the factor 2 in the numerator is because in nonlinearities two components q,k{\bm{q}},{\bm{k}} are both learned (Equation 5). In reconstruction we slightly extend the layer size for the VN encoder. Moreover, in VN-PointNet, we discard the input spatial transformation MLP which learns 3×33\times 3 transformation matrices as our VN network already takes rigid transformations into consideration by construction. In the following experiments, we use mean pooling as aggregation in all networks, which performed better in practice. We will discuss more about the max pooling as well as ablation study on other structures in the supplementary material.

1 Classification – Table 1

2 Part segmentation – Table 2

Table 2 shows our results in ShapeNet part segmentation. Again our method shows consistent results across different rotations and achieves best performance with VN-DGCNN compared with other works, including that uses surface normals in addition to the point coordinates.

3 Neural implicit reconstruction – Table 3

Decoder network

Quantitative results – Table 3

Qualitative results – Figure 6

Conclusions

We have introduced Vector Neurons – a novel framework that facilitates rotation equivariant neural networks by lifting standard neural network representations to 3 space. To that end, we have introduced the vector-neuron counterpart of standard network modules: linear layers, non-linearities, pooling and normalization. Using our framework, we have built a rotation-equivariant version of two leading pointcloud network backbones: PointNet and DGCNN, and evaluated them on 3 tasks: classification, segmentation and reconstruction. Our results demonstrate a consistent advantage to our modified architecture when the input shapes pose is arbitrary, compared to an augmentation based approach.

While our method shines under arbitrary rotation settings, on aligned input shapes and specifically in the task of reconstruction, our VN-OccNet was not able to match the reconstruction quality of vanilla OccNet by a small margin. In future work we plan to investigate this matter.

In this work, we have focused on 3D pointcloud networks, yielding permutation and rotation equivariant architectures. However, it should be clear that our framework has obvious generalizations to higher-dimensional pointclouds in a completely analogous way. We also believed it can find applications in other modalities like meshes, voxel grids, and even in the image domain. Generalization of vector neurons to other transformation groups of interest, such as the full affine group, can also be investigated (the addition of uniform scalings in our framework is quite straightforward).

In summary, by making rotation equivariant modules simple and accessible we hope to alleviate the need to curate and pre-align shapes for supervision and inspire future research on this fascinating topic.

Acknowledgements

We gratefully acknowledge the support of a Vannevar Bush Faculty Fellowship, as well as gifts from the Adobe, Amazon AWS, and Autodesk corporations.

References

Discussions

In this section, we discuss some extensions, alternatives, and explanations to the VN layers in Section 3.

The VN-ReLU defined in Section 3.2 already consists of a built-in linear layer q=WV{\bm{q}}={\mathbf{W}}{\bm{V}} (5) and the the non-linearity is applied to this learned feature q{\bm{q}}. An alternative to this is to construct linear and non-linear layers separately, where the non-linearity is directly applied to each input vector channel v∈V{\bm{v}}\in{\bm{V}} by

Detaching the linear layer from non-linearity allows more flexibility in constructing neural networks and, in practice, gives better results in some cases. However, this also doubles the network depth and can lead to longer training time compared to the entangled linear-ReLU layer in (6). Experimental comparisons will be shown in Section 8.2.

Other non-linearities

Though we only showed how to define VN-ReLU in Section 3.2, a rich library of equivariant non-linearities can be defined in this manner using the input-dependent direction vector k{\bm{k}}. An immediate extension is VN-LeakyReLU, where instead of clipping q∥{\bm{q}}_{\parallel} to zero we contract it by a factor α∈(0,1)\alpha\in(0,1). In the manner of the detached VN-ReLU in (24), the VN-LeakyReLU can be easily expressed as:

2 Local Pooling

For any point x∈X{\bm{x}}\in{\mathcal{X}} with feature V∈V{\bm{V}}\in{\mathcal{V}} we consider its KK nearest neighbours {xk}k=1K\{{\bm{x}}_{k}\}_{k=1}^{K} in the primal space and we denote by Vk∈V{\bm{V}}_{k}\in{\mathcal{V}} the corresponding feature of xk{\bm{x}}_{k}. Similar to global pooling (9), local pooling (in the primal space) is given by:

Feature space locality

3 Batch Normalization

In VN-BatchNorm (10), for each input vector-list feature Vb{\bm{V}}_{b}, all entries in its per-channel 2-norm Nb{\bm{N}}_{b} are non-negative, but after normalizing the distributions, the output “2-norm” Nb′{\bm{N}}^{\prime}_{b} can have negative entries. Geometrically, a negative entry nc′∈Nb′{\bm{n}}^{\prime}_{c}\in{\bm{N}}^{\prime}_{b} means the orientation of its corresponding vector channel is flipped, that is, vc′∈Vb′{\bm{v}}^{\prime}_{c}\in{\bm{V}}^{\prime}_{b} is in the opposite direction of vc∈Vb{\bm{v}}_{c}\in{\bm{V}}_{b}.

To avoid the negative 2-norms, an alternative is to take logarithms on all entries of Nb{\bm{N}}_{b} and then apply the standard BatchNorm to log⁡(Nb)\log({\bm{N}}_{b}). So the VN batch normalization becomes:

where log⁡\log and exp⁡\exp act element-wise. However, taking log⁡\log and exp⁡\exp brings a lot of instability and in practice can cause gradient explosion. Also, logarithms cannot be computed for vectors with zero 2-norms.

Additional Experiments

2 Ablation Studies

Table 6, 7, and 8 show our ablation studies on non-linearity, pooling, and the invariant layer in VN networks on ModelNet40 classification.