Learning Gradient Fields for Molecular Conformation Generation

Chence Shi, Shitong Luo, Minkai Xu, Jian Tang

Introduction

Graph-based representations of molecules, where nodes represent atoms and edges represent bonds, have been widely applied to a variety of molecular modeling tasks, ranging from property prediction (Gilmer et al., 2017; Hu et al., 2020) to molecule generation (Jin et al., 2018; You et al., 2018; Shi et al., 2020; Xie et al., 2021). However, in reality a more natural and intrinsic representation for molecules is their three-dimensional structures, where atoms are represented by 3D coordinates. These structures are also widely known as molecular geometry or conformation.

Nevertheless, obtaining valid and stable conformations for a molecule remains a long-standing challenge. Existing approaches to molecular conformation generation mainly rely on molecular dynamics (MD) (De Vivo et al., 2016), where coordinates of atoms are sequentially updated based on the forces acting on each atom. The per-atom forces come either from computationally intensive Density Functional Theory (DFT) (Parr, 1989) or from hand-designed force fields (Rappé et al., 1992; Halgren, 1996) which are often crude approximations of the actual molecular energies (Kanal et al., 2017). Recently, there are growing interests in developing machine learning methods for conformation generation (Mansimov et al., 2019; Simm & Hernandez-Lobato, 2020; Xu et al., 2021). In general, these approaches (Simm & Hernandez-Lobato, 2020; Xu et al., 2021) are divided into two stages, which first predict the molecular distance geometry, i.e., interatomic distances, and then convert the distances to 3D coordinates via post-processing algorithms (Liberti et al., 2014). Such methods effectively model the roto-translation equivariance of molecular conformations and have achieved impressive performance. Nevertheless, such approaches generate conformations in a two-stage fashion, where noise in generated distances may affect 3D coordinate reconstruction, often leading to less accurate or even erroneous structures. Therefore, we are seeking for an approach that is able to generate molecular conformations within a single stage.

Inspired by the traditional force field methods for molecular dynamics simulation (Frenkel & Smit, 1996), in this paper, we propose a novel and principled approach called ConfGF for molecular conformation generation. The basic idea is to directly learn the gradient fields of the log density w.r.t. the atomic coordinates, which can be viewed as pseudo-forces acting on each atom. With the estimated gradient fields, conformations can be generated directly using Langevin dynamics (Welling & Teh, 2011), which amounts to moving atoms toward the high-likelihood regions guided by pseudo-forces. The key challenge with this approach is that such gradient fields of atomic coordinates are roto-translation equivariant, i.e., the gradients rotate together with the molecular system and are invariant under translation.

To tackle this challenge, we observe that the interatomic distances are continuously differentiable w.r.t. atomic coordinates. Therefore, we propose to first estimate the gradient fields of the log density w.r.t. interatomic distances, and show that under a minor assumption, the gradient fields of the log density of atomic coordinates can be calculated from these w.r.t. distances in closed form using chain rule. We also demonstrate that such designed gradient fields satisfy roto-translation equivariance. In contrast to distance-based methods (Simm & Hernandez-Lobato, 2020; Xu et al., 2021), our approach generates conformations in an one-stage fashion, i.e., the sampling process is done directly in the coordinate space, which avoids the errors induced by post-processing and greatly enhances generation quality. In addition, our approach requires no surrogate losses and allows flexible network architectures, making it more capable to generate diverse and realistic conformations.

We conduct extensive experiments on GEOM (Axelrod & Gomez-Bombarelli, 2020) and ISO17 (Simm & Hernandez-Lobato, 2020) benchmarks, and compare ConfGF against previous state-of-the-art neural-methods as well as the empirical method on multiple tasks ranging from conformation generation and distance modeling to property prediction. Numerical results show that ConfGF outperforms previous state-of-the-art baselines by a clear margin.

Related Work

Molecular Conformation Generation Existing works on molecular conformation generation mainly rely on molecular dynamics (MD) (De Vivo et al., 2016). Starting from an initial conformation and a physical model for interatomic potentials (Parr, 1989; Rappé et al., 1992), the conformation is sequentially updated based on the forces acting on each atom. However, these methods are inefficient, especially when molecules are large as they involve computationally expensive quantum mechanical calculations (Parr, 1989; Shim & MacKerell Jr, 2011; Ballard et al., 2015). Another much more efficient but less accurate sort of approach is the empirical method, which fixes distances and angles between atoms in a molecule to idealized values according to a set of rules (Blaney & Dixon, 2007).

Recent advances in deep generative models open the door to data-driven approaches, which have shown great potential to strike a good balance between computational efficiency and accuracy. For example, Mansimov et al. (2019) introduce the Conditional Variational Graph Auto-Encoder (CVGAE) which takes molecular graphs as input and directly generates 3D coordinates for each atom. This method does not manage to model the roto-translation equivariance of molecular conformations, resulting in great differences between generated structures and ground-truth structures. To address the issue, Simm & Hernandez-Lobato (2020) and Xu et al. (2021) propose to model interatomic distances of molecular conformations using VAE and Continuous Flow respectively. These approaches preserve roto-translation equivariance of molecular conformations and consist of two stages — first generating distances between atoms and then feeding the distances to the distance geometry algorithms (Liberti et al., 2014) to reconstruct 3D coordinates. However, the distance geometry algorithms are vulnerable to noise in generated distances, which may lead to inaccurate or even erroneous structures. Another line of research (Gogineni et al., 2020) leverages reinforcement learning for conformation generation by sequentially determining torsional angles, which relies on an additional classical force field for state transition and reward evaluation. Nevertheless, it is incapable of modeling other geometric properties such as bond angles and bond lengths, which distinguishes it from all other works.

Neural Force Fields Neural networks have also been employed to estimate molecular potential energies and force fields (Schütt et al., 2017; Zhang et al., 2018; Hu et al., 2021). The goal of these models is to predict energies and forces as accurate as possible, serving as an alternative to quantum chemical methods such as Density Function Theory (DFT) method. Training these models requires access to molecular dynamics trajectories along with ground-truth energies and forces from expensive quantum mechanical calculations. Different from these methods, our goal is to generate equilibrium conformations within a single stage. To this end, we define gradient fieldsSuch gradient fields, however, are not force fields as they are not energy conserving. analogous to force fields, which serve as pseudo-forces acting on each atom that gradually move atoms toward the high-density areas until they converges to an equilibrium. The only data we need to train the model is equilibrium conformations.

Preliminaries

As the bonded edges in a molecule are not sufficient to characterize a conformation (bonds are rotatable), we extend the original molecular graph by adding auxiliary edges, i.e., virtual bonds, between atoms that are 2 or 3 hops away from each other, which is a widely-used technique in previous work (Simm & Hernandez-Lobato, 2020; Xu et al., 2021) for reducing the degrees of freedom in 3D coordinates. The extra edges between second neighbors help fix angles between atoms, while those between third neighbors fix dihedral angles. Hereafter, we assume all molecular graphs are extended unless stated.

2 Score-Based Generative Modeling

The score function s(x)\mathbf{s}(\mathbf{x}) of a continuously differentiable probability density p(x)p(\mathbf{x}) is defined as ∇xlog⁡p(x)\nabla_{\mathbf{x}}\log p(\mathbf{x}). Score-based generative modeling (Song & Ermon, 2019; Song et al., 2020; Meng et al., 2020; Song & Ermon, 2020) is a class of generative models that approximate the score function of p(x)p(\mathbf{x}) using neural networks and generate new samples with Langevin dynamics (Welling & Teh, 2011). Modeling the score function instead of the density function allows getting rid of the intractable partition function of p(x)p(\mathbf{x}) and avoiding extra computation on getting the higher-order gradients of an energy-based model (Song & Ermon, 2019).

After the noise conditional score networks sθ(x,σi)\mathbf{s}_{\theta}(\mathbf{x},\sigma_{i}) are trained, samples are generated using annealed Langevin dynamics (Song & Ermon, 2019), where samples from each noise level serve as initializations for Langevin dynamics of the next noise level. For detailed design choices of noise levels and sampling hyperparameters, we refer readers to Song & Ermon (2020).

3 Equivariance in Molecular Geometry

Equivariance is ubiquitous in physical systems, e.g., 3D roto-translation equivariance of molecular conformations or point clouds (Thomas et al., 2018; Weiler et al., 2018; Chmiela et al., 2019; Fuchs et al., 2020; Miller et al., 2020; Simm et al., 2021). Endowing model with such inductive biases is critical for better generalization and successful learning (Köhler et al., 2020). For molecular geometry modeling, such equivariance is guaranteed by either taking gradients of a predicted invariant scalar energy (Klicpera et al., 2020; Schütt et al., 2017) or using equivariant networks (Fuchs et al., 2020; Satorras et al., 2021). Formally, a function F:X→Y\mathcal{F}:\mathcal{X}\rightarrow\mathcal{Y} being equivariant can be represented as follows:Strictly speaking, ρ\rho does not have to be the same on both sides of the equation, as long as it is a representation of the same group. Note that invariance is a special case of equivariance.

where ρ\rho is a transformation function, e.g., rotation. Intuitively, Eq. 2 says that applying the ρ\rho on the input has the same effect as applying it to the output.

In this paper, we aim to model the score function of p(R∣G)p(\mathbf{R}\mid G), i.e., s(R)=∇Rlog⁡p(R∣G)\mathbf{s}(\mathbf{R})=\nabla_{\mathbf{R}}\log p(\mathbf{R}\mid G). As log⁡p(R∣G)\log p(\mathbf{R}\mid G) is roto-translation invariant with respect to conformations, the score function s(R)\mathbf{s}(\mathbf{R}) is therefore roto-translation equivariant with respect to conformations. To pursue a generalizable and elegant approach, we explicitly build such equivariance into the model design.

Proposed Method

Our approach extends the denoising score matching (Vincent, 2011; Song & Ermon, 2019) to molecular conformation generation by estimating the gradient fields of the log density of atomic coordinates, i.e., ∇Rlog⁡p(R∣G)\nabla_{\mathbf{R}}\log p(\mathbf{R}\mid G). Directly parameterizing such gradient fields using vanilla Graph Neural Networks (GNNs) (Duvenaud et al., 2015; Kearnes et al., 2016; Kipf & Welling, 2016; Gilmer et al., 2017; Schütt et al., 2017, 2017) is problematic, as these GNNs only leverage the node, edge or distance information, resulting in node representations being roto-translation invariant. However, the gradient fields of atomic coordinates should be roto-translation equivariant (see Section 3.3), i.e., the vector fields should rotate together with the molecular system and be invariant under translation.

where sθ(R)i=∇rilog⁡pθ(R∣G)\mathbf{s}_{\theta}(\mathbf{R})_{i}=\nabla_{\mathbf{r}_{i}}\log p_{\theta}(\mathbf{R}\mid G) and sθ(d)ij=∇dijlog⁡pθ(d∣G)\mathbf{s}_{\theta}(\mathbf{d})_{ij}=\nabla_{d_{ij}}\log p_{\theta}(\mathbf{d}\mid G). Here N(i)N(i) denotes ithi^{th} atom’s neighbors. The second equation of Eq. 3 holds because of the chain rule over the partial derivative of the composite function fG(d)=fG∘gG(R)f_{G}(\mathbf{d})=f_{G}\circ g_{G}(\mathbf{R}).

Under the assumption that we parameterize log⁡pθ(R∣G)\log p_{\theta}(\mathbf{R}\mid G) as a function of interatomic distances d\mathbf{d} and molecular graph representation GG, the score network defined in Eq. 3 is roto-translation equivariant, i.e., the gradients rotate together with the molecule system and are invariant under translation.

The score network of distances sθ(d)\mathbf{s}_{\theta}(\mathbf{d}) is roto-translation invariant as it only depends on interatomic distances, and the vector (ri−rj)/dij(\mathbf{r}_{i}-\mathbf{r}_{j})/d_{ij} rotates together with the molecule system, which endows the score network of atomic coordinates sθ(R)\mathbf{s}_{\theta}(\mathbf{R}) with roto-translation equivariance. See Supplementary material for a formal proof. ∎

After training the score network, based on the estimated gradient fields of atomic coordinates corresponding to different noise levels, conformations are generated by annealed Langevin dynamics (Song & Ermon, 2019), combining information from all noise levels. The overview of ConfGF is illustrated in Figure 1 and Figure 2. Below we describe the design of the noise conditional score network in Section 4.2, introduce the generation procedure in Section 4.3, and show how the designed gradient fields are connected with the classical laws of mechanics in Section 4.4.

2 Noise Conditional Score Networks for Distances

We then adopt a variant of Graph Isomorphism Network (GINs) (Xu et al., 2019; Hu et al., 2020) to compute node embeddings based on graph structures. At each layer, node embeddings are updated by aggregating messages from neighboring nodes and edges:

where N(i)N(i) denotes ithi^{th} atom’s neighbors. After NN rounds of message passing, we compute the final edge embeddings by concatenating the corresponding node embeddings for each edge as follows:

where λ(σi)=σi2\lambda(\sigma_{i})=\sigma_{i}^{2} is the coefficient function weighting losses of different noise levels according to Song & Ermon (2019), and all expectations can be efficiently estimated using Monte Carlo estimation.

3 Conformation Generation via Annealed Langevin Dynamics

Given the molecular graph representation GG and the well-trained noise conditional score networks, molecular conformations are generated using annealed Langevin dynamics, i.e., gradually anneal down the noise level from large to small. We start annealed Langevin dynamics by first sampling an initial conformation R0\mathbf{R}_{0} from some fixed prior distribution, e.g., uniform distribution or Gaussian distribution. Empirically, the choice of the prior distribution is arbitrary as long as the supports of the perturbed distance distributions cover the prior distribution, and we take the Gaussian distribution as prior distribution in our case. Then, we update the conformations by sampling from a series of trained noise conditional score networks {sθ(R,σi)}i=1L\{\mathbf{s}_{\theta}(\mathbf{R},\sigma_{i})\}_{i=1}^{L} sequentially. For each noise level σi\sigma_{i}, starting from the final sample from the last noise level, we run Langevin dynamics for TT steps with a gradually decreasing step size αi=ϵ⋅σi2/σL2\alpha_{i}=\epsilon\cdot\sigma_{i}^{2}/\sigma_{L}^{2} for each noise level. In specific, at each sampling step tt, we first obtain the interatomic distances dt−1\mathbf{d}_{t-1} based on the current conformation Rt−1\mathbf{R}_{t-1}, and calculate the score function sθ(Rt−1,σi)\mathbf{s}_{\theta}(\mathbf{R}_{t-1},\sigma_{i}) via Eq. 3. The conformation is then updated using the gradient information from the score network. The whole sampling algorithm is presented in Algorithm 1.

4 Discussion

Physical Interpretation of Eq. 3 From a physical perspective, sθ(d,σ)\mathbf{s}_{\theta}(\mathbf{d},\sigma) reflects how the distances between connected atoms should change to increase the probability density of the conformation R\mathbf{R}, which can be viewed as a network that predicts the repulsion or attraction forces between atoms. Moreover, as shown in Figure 1, sθ(R,σ)\mathbf{s}_{\theta}(\mathbf{R},\sigma) can be viewed as resultant forces acting on each atom computed from Eq. 3, where (ri−rj)/dij(\mathbf{r}_{i}-\mathbf{r}_{j})/d_{ij} is a unit vector indicating the direction of each component force and sθ(d,σ)ij\mathbf{s}_{\theta}(\mathbf{d},\sigma)_{ij} specifies the strength of each component force.

Handling Stereochemistry Given a molecular graph, our ConfGF generates all possible stereoisomers due to the stochasticity of the Langevin dynamics. Nevertheless, our model is very general and can be extended to handle stereochemistry by predicting the gradients of torsional angles with respect to coordinates, as torsional angles can be calculated in closed-form from 3D coordinates and are also invariant under rotation and translation. Since the stereochemistry study of conformation generation is beyond the scope of this work, we leave it for future work.

Experiments

Following existing works on conformation generation, we evaluate the proposed method using the following three standard tasks:

Conformation Generation tests models’ capacity to learn distribution of conformations by measuring the quality of generated conformations compared with reference ones (Section 5.1).

Distributions Over Distances evaluates the discrepancy with respect to distance geometry between the generated conformations and the ground-truth conformations (Section 5.2).

Property Prediction is first proposed in Simm & Hernandez-Lobato (2020), which estimates chemical properties for molecular graphs based on a set of sampled conformations (Section 5.3). This can be useful in a variety of applications such as drug discovery, where we want to predict molecular properties from microscopic structures.

Below we describe setups that are shared across the tasks. Additional setups are provided in the task-specific sections.

Data Following Xu et al. (2021), we use the GEOM-QM9 and GEOM-Drugs (Axelrod & Gomez-Bombarelli, 2020) datasets for the conformation generation task. We randomly draw 40,000 molecules and select the 5 most likely conformationssorted by energy for each molecule, and draw 200 molecules from the remaining data, resulting in a training set with 200,000 conformations and a test set with 22,408 and 14,324 conformations for GEOM-QM9 and GEOM-Drugs respectively. We evaluate the distance modeling task on the ISO17 dataset (Simm & Hernandez-Lobato, 2020). We use the default split of Simm & Hernandez-Lobato (2020), resulting in a training set with 167 molecules and the test set with 30 molecules. For property prediction task, we randomly draw 30 molecules from the GEOM-QM9 dataset with the training molecules excluded.

Baselines We compare our approach with four state-of-the-art methods for conformation generation. In specific, CVGAE (Mansimov et al., 2019) is built upon conditional variational graph autoencoder, which directly generates atomic coordinates based on molecular graphs. GraphDG (Simm & Hernandez-Lobato, 2020) and CGCF (Xu et al., 2021) are distance-based methods, which rely on post-processing algorithms to convert distances to final conformations. The difference of the two methods lies in how the neural networks are designed, i.e., VAE vs. Continuous Flow. RDKit (Riniker & Landrum, 2015) is a classical Euclidean Distance Geometry-based approach. Results of all baselines across experiments are obtained by running their official codes unless stated.

Model Configuration ConfGF is implemented in Pytorch (Paszke et al., 2017). The GINs is implemented with N=4N=4 layers and the hidden dimension is set as 256 across all modules. For training, we use an exponentially-decayed learning rate starting from 0.001 with a decay rate of 0.95. The model is optimized with Adam (Kingma & Ba, 2014) optimizer on a single Tesla V100 GPU. All hyperparameters related to noise levels as well as annealed Langevin dynamics are selected according to Song & Ermon (2020). See Supplementary material for full details.

Setup The first task is to generate conformations with high diversity and accuracy based on molecular graphs. For each molecular graph in the test set, we sample twice as many conformations as the reference ones from each model. To measure the discrepancy between two conformations, following existing work (Hawkins, 2017; Li et al., 2021), we adopt the Root of Mean Squared Deviations (RMSD) of the heavy atoms between aligned conformations:

where nn is the number of heavy atoms and Φ\Phi is an alignment function that aligns two conformations by rotation and translation. Following Xu et al. (2021), we report the Coverage (COV) and Matching (MAT) score to measure the diversity and accuracy respectively:

where SgS_{g} and SrS_{r} are generated and reference conformation set respectively. δ\delta is a given RMSD threshold. In general, a higher COV score indicates a better diversity, while a lower MAT score indicates a better accuracy.

Results We evaluate the mean and median COV and MAT scores on both GEOM-QM9 and GEOM-Drugs datasets for all baselines, where mean and median are taken over all molecules in the test set. Results in Table 1 show that our ConfGF achieves the state-of-the-art performance on all four metrics. In particular, as a score-based model that directly estimates the gradients of the log density of atomic coordinates, our approach consistently outperforms other neural models by a large margin. By contrast, the performance of GraphDG and CGCF is inferior to ours due to their two-stage generation procedure, which incurs extra errors when converting distances to conformations. The CVGAE seems to struggle with both of the datasets as it fails to model the roto-translation equivariance of conformations.

When compared with the rule-based method, our approach outperforms the RDKit on 7 out of 8 metrics without relying on any post-processing algorithms, e.g., force fields, as done in previous work (Mansimov et al., 2019; Xu et al., 2021), making it appealing in real-world applications. Figure 3 visualizes several conformations that are best-aligned with the reference ones generated by different methods, which demonstrates ConfGF’s strong ability to generate realistic and diverse conformations.

2 Distributions Over Distances

Setup The second task is to match the generated distributions over distances with the ground-truth. As ConfGF is not designed for modeling molecular distance geometry, we treat interatomic distances calculated from generated conformations as generated distance distributions. For each molecular graph in the test set, we sample 1000 conformations across all methods. Following previous work (Simm & Hernandez-Lobato, 2020; Xu et al., 2021), we report the discrepancy between generated distributions and the ground-truth distributions, by calculating maximum mean discrepancy (MMD) (Gretton et al., 2012) using a Gaussian kernel. In specific, for each graph G in the test set, we ignore edges connected with the H atoms and evaluate distributions of all distances p(d∣G)p(\mathbf{d}\mid G) (All), pairwise distances p(dij,duv∣G)p(d_{ij},d_{uv}\mid G) (Pair), and individual distances p(dij∣G)p(d_{ij}\mid G) (Single).

Results As shown in Table 2, the CVGAE suffers the worst performance due to its poor capacity. Despite the superior performance of RDKit on the conformation generation task, its performance falls far short of other neural models. This is because RDKit incorporates an additional molecular force field (Rappé et al., 1992) to find equilibrium states of conformations after the initial coordinates are generated. By contrast, although not tailored for molecular distance geometry, our ConfGF still achieves competitive results on all evaluation metrics. In particular, the ConfGF outperforms the state-of-the-art method CGCF on 4 out of 6 metrics, and achieves comparable results on the other two metrics.

3 Property Prediction

Setup The third task is to predict the ensemble property (Axelrod & Gomez-Bombarelli, 2020) of molecular graphs. Specifically, the ensemble property of a molecular graph is calculated by aggregating the property of different conformations. For each molecular graph in the test set, we first calculate the energy, HOMO and LUMO of all the ground-truth conformations using the quantum chemical calculation package Psi4 (Smith et al., 2020). Then, we calculate ensemble properties including average energy E‾\overline{E}, lowest energy EminE_{\text{min}}, average HOMO-LUMO gap Δϵ‾\overline{\Delta\epsilon}, minimum gap Δϵmin\Delta\epsilon_{\text{min}}, and maximum gap Δϵmax\Delta\epsilon_{\text{max}} based on the conformational properties for each molecular graph. Next, we use ConfGF and other baselines to generate 50 conformations for each molecular graph and calculate the above-mentioned ensemble properties in the same way. We use mean absolute error (MAE) to measure the accuracy of ensemble property prediction. We exclude CVGAE from this analysis due to its poor quality following Simm & Hernandez-Lobato (2020).

Results As shown in Table 3, ConfGF outperforms all the machine learning-based methods by a large margin and achieves better accuracy than RDKit in lowest energy EminE_{\text{min}} and maximum HOMO-LUMO gap Δϵmax\Delta\epsilon_{\text{max}}. Closer inspection on the generated samples shows that, although ConfGF has better generation quality in terms of diversity and accuracy, it sometimes generates a few outliers which have negative impact on the prediction accuracy. We further present the median of absolute errors in Table 4. When measured by median absolute error, ConfGF has the best accuracy in average energy and the accuracy of other properties is also improved, which reveals the impact of outliers.

4 Ablation Study

So far, we have justified the effectiveness of the proposed ConfGF on multiple tasks. However, it remains unclear whether the proposed strategy, i.e., directly estimating the gradient fields of log density of atomic coordinates, contributes to the superior performance of ConfGF. To gain insights into the working behaviours of ConfGF, we conduct the ablation study in this section.

We adopt a variant of ConfGF, called ConfGFDist, which works in a way similar to distance-based methods (Simm & Hernandez-Lobato, 2020; Xu et al., 2021). In specific, ConfGFDist bypasses chain rule and directly generates interatomic distances via annealed Langevin dynamics using estimated gradient fields of interatomic distances, i.e., sθ(d,σ)\mathbf{s}_{\theta}(\mathbf{d},\sigma) (Section 4.2), after which we convert them to conformations by the same post-processing algorithm as used in Xu et al. (2021). Following the setup of Section 5.1, we evaluate two approaches on the conformation generation task and summarize the results in Table 5. As presented in Table 5, we observe that ConfGFDist performs significantly worse than ConfGF on both datasets, and performs on par with the state-of-the-art method CGCF. The results indicate that the two-stage generation procedure of distance-based methods does harm the quality of conformations, and our strategy effectively addresses this issue, leading to much more accurate and diverse conformations.

Conclusion and Future Work

In this paper, we propose a novel principle for molecular conformation generation called ConfGF, where samples are generated using Langevin dynamics within one stage with physically inspired gradient fields of log density of atomic coordinates. We novelly develop an algorithm to effectively estimate these gradients and meanwhile preserve their roto-translation equivariance. Comprehensive experiments over multiple tasks verify that ConfGF outperforms previous state-of-the-art baselines by a significant margin. Our future work will include extending our ConfGF approach to 3D molecular design tasks and many-body particle systems.

Acknowledgments

We would like to thank Yiqing Jin for refining the figures used in this paper. This project is supported by the Natural Sciences and Engineering Research Council (NSERC) Discovery Grant, the Canada CIFAR AI Chair Program, collaboration grants between Microsoft Research and Mila, Samsung Electronics Co., Ldt., Amazon Faculty Research Award, Tencent AI Lab Rhino-Bird Gift Fund and a NRC Collaborative R&D Project (AI4D-CORE-06). This project was also partially funded by IVADO Fundamental Research Project grant PRF-2019-3583139727.

References

Appendix A Proof of Proposition 1

Intuitively, the above equation says that applying any translation γt\gamma_{\mathbf{t}} to the input has no effect on the output, and applying any rotation ρD\rho_{D} to the input has the same effect as applying it to the output, i.e., the gradients rotate together with the molecule system and are invariant under translation.

Denote R^=γt∘ρD(R)\hat{\mathbf{R}}=\gamma_{\mathbf{t}}\circ\rho_{D}(\mathbf{R}). According to the definition of translation and rotation function, we have R^i=r^i=Dri+t\hat{\mathbf{R}}_{i}=\hat{\mathbf{r}}_{i}=D\mathbf{r}_{i}+\mathbf{t}. According to Eq. 3, we have

Here dij=d^ijd_{ij}=\hat{d}_{ij} and d=d^\mathbf{d}=\hat{\mathbf{d}} because rotation and translation will not change interatomic distances. Together with Eq. 11 and Eq. 12, we conclude that the score network defined in Eq. 3 is roto-translation equivariant. ∎

Appendix B Additional Hyperparameters

We summarize additional hyperparameters of our ConfGF in Table 6, including the biggest noise level σ1\sigma_{1}, the smallest noise level σL\sigma_{L}, the number of noise levels LL, the number of sampling steps per noise level TT, the smallest step size ϵ\epsilon, the batch size and the training epochs.

Appendix C Additional Experiments

We present more results of the Coverage (COV) score at different threshold δ\delta for GEOM-QM9 and GEOM-Drugs datasets in Tables 7 and 8 respectively. Results in Tables 7 and 8 indicate that the proposed ConfGF consistently outperforms previous state-of-the-art baselines, including GraphDG and CGCF. In addition, we also report the Mismatch (MIS) score defined as follows:

where δ\delta is the threshold. The MIS score measures the fraction of generated conformations that are not matched by any ground-truth conformation in the reference set given a threshold δ\delta. A lower MIS score indicates less bad samples and better generation quality. As shown in Tables 7 and 8, the MIS scores of ConfGF are consistently lower than the competing baselines, which demonstrates that our ConfGF generates less invalid conformations compared with the existing models.

Appendix D More Generated Samples

We present more visualizations of generated conformations from our ConfGF in Figure 4, including models trained on GEOM-QM9 and GEOM-drugs datasets.