Relevance of Rotationally Equivariant Convolutions for Predicting Molecular Properties

Benjamin Kurt Miller, Mario Geiger, Tess E. Smidt, Frank Noé

Introduction

The discovery of novel molecules has been accelerated by advances in computational quantum chemistry and machine learning assisted exploration of chemical space [halicin, robin_molecule_discovery, molecule_net]. The successes have been characterized by designing bespoke neural networks which have relevant properties “baked-in,” such as parameter sharing across calculations on individual atoms, continuous convolutions, invariance to atomic indexing, and invariance to rotation and translation [schnet, ani1]. Meanwhile, there has also been development on neural networks which are equivariant to group action [se3cnn], some with molecules in mind [tensorfieldnetworks, cormorant]. Equivariant neural networks can be seen as a super-set of invariant ones because a group necessarily contains the identity element. The question considered in this paper can loosely be stated as: When doing regression on scalar molecular properties, what is missing when one employs only invariant layers in a neural network as opposed to including equivariant ones?

We explore this question using the QM9 benchmark [gdb17, qm9] by predicting quantum chemical properties of small molecules. While the molecules can rotate and translate, affecting the molecule’s position vectors, the QM9 properties are all scalar and invariant to translation or rotation. Here we compare equivariant neural networks (ENNs) that predict rototranslationally invariant molecular properties but differ by whether their internal features are rotationally invariant (convolutions depend on distances) or equivariant (convolutions depend on distances and angles). We also investigate whether increasing depth in networks with rotationally invariant layers is comparatively effective at reducing test error. The networks are implemented in the PyTorch [pytorch] library e3nn [e3nn] using the SE(3)SE(3) equivariant point modules. QM9 data handling and training routines were borrowed from SchNetPack [schnetpack].

Molecular properties, which depend only on the atomic distance graph, are commonly predicted by kernel methods or Gaussian process regression [mol_kernel_0, mol_kernel_1, mol_kernel_2, mol_kernel_3] or graph neural networks [graphs_0, graph_1], where ENNs are usually employed for predicting physical properties, which depend on atomic displacement vectors [husic2020coarse, geometric_new_paper]. While kernel approaches are more data-efficient, graph neural networks scale to larger amounts of data. Inspiration for our study came from literature on invariant and equivariant ENNs for molecular property prediction. SchNet [schnet, schnetpack] introduced atom-wise features with continuous convolution. Tensor Field Networks [tensorfieldnetworks] and Cormorant [cormorant] generalized the approach to angular-feature based rotation equivariant networks. In parallel, although aimed at voxelized data, se3cnn [se3cnn] developed the gated nonlinearity. The library under consideration, e3nn, represents a superset of Tensor Field Networks, SchNet, and se3cnn. Support for Cormorant’s so-called two-body interaction has also been included in e3nn but is not considered in this experiment.

Although DimeNet [dimenet] is the leading architecture on QM9 regression, it is not considered in our analysis. Their use of directional message passing, while effective, is not trivially compatible with SchNet, Cormorant, or e3nn. Their edge featurization using spherical Fourier-Bessel functions could be incorporated rather simply, but investigation is left for future work.

While QM9 remains the gold standard for most machine learning studies, a new data set called QM7-X [qm7x] contains a wealth of tensor properties suitable for prediction with networks like e3nn or Cormorant. SchNet and DimeNet cannot predict tensor quantities in their current incarnations.

Methods