SE(3)-Stochastic Flow Matching for Protein Backbone Generation
Avishek Joey Bose, Tara Akhound-Sadegh, Guillaume Huguet, Kilian Fatras, Jarrid Rector-Brooks, Cheng-Hao Liu, Andrei Cristian Nica, Maksym Korablyov, Michael Bronstein, Alexander Tong
Introduction
Proteins are one of the basic building blocks of life. Their complex geometric structure enables specific inter-molecular interactions that allow for crucial functions within organisms, such as acting as catalysts in chemical reactions, transporters for molecules, and providing immune responses. Normally, such functions arise as a result of evolution. With the emergence of computational techniques, it has become possible to rationally design novel proteins with desired structures that program their functions. Such methods are now seen as the future of drug design and can lead to solutions to long-standing global health challenges. Some recent examples include rationally designed protein binders for receptors related to influenza (Strauch et al. 2017), COVID-19 (Cao et al. 2020a; Gainza et al. 2023), and cancer (Silva et al. 2019).
Background and preliminaries
2 Flow matching on Riemannian manifolds
Riemannian flow matching. Given a probability path that connects to , and its associated flow , we can learn a CNF by directly regressing the vector field with a parametric one . This technique is termed flow matching (Lipman et al. 2022, FM) and leads to a simulation-free training objective as long as satisfies the boundary conditions and . Unfortunately, the vanilla flow matching objective is intractable as we generally do not have access to the closed-form of that generates . Instead, we can opt to regress against a conditional vector field , generating a conditional probability path , and use it to recover the target unconditional path: . The vector field can also be recovered by marginalizing of conditional vector fields: . The Riemannian CFM objective (Chen & Lipman 2023) is then,
As FM and CFM objectives have the same gradients (Tong et al. 2023b), at inference, we can generate by sampling from , and using to propagate the ODE backward in time.
3 Protein Backbone Parametrization
2 FoldFlow-OT
To this end, we propose FoldFlow-OT, a model that accelerates FoldFlow-Base by constructing conditional probability paths using Riemannian optimal transport. The interpolation measure connects and is built from Riemannian OT which solves the Monge optimal transport problem:
3 FoldFlow-SFM
This flow, also known as Doob’s h-transform (Doob 1984), is easy to sample from in a simulation-free manner and correctly maps between arbitrary samplable marginals in expectation. Specifically, we can build a simulation-free bridge by sampling from the conditional probability (Shi et al. 2023; Albergo et al. 2023). See §B for more details.
Modeling Protein Backbones using FoldFlow
To model protein backbones using FoldFlow models we parameterize the velocity prediction as a function that consumes a protein on the conditional path at time and predicts the starting point . Specifically, the predicted velocity is , with . This choice of parameterization has two principal benefits. (1) It allows the usage of specialized architectures specifically designed for structure prediction, and (2) it allows for auxiliary protein-specific losses to be placed directly on the to improve performance.
Architecture. Following by Anand & Achim 2022; Yim et al. 2023b (FrameDiff) we use the structure module of AF2 to model . This begins with a time-dependent node and edge embeddings and , followed by layers of invariant point attention. We use a small MLP head on top of the node embeddings to predict the torsion angle of the oxygen as .
Experiments
2 Protein Backbone Design
We evaluate FoldFlow models in generating valid, diverse, and novel backbones by training on a subset of the Protein Data Bank (PDB) with proteins. We compare FoldFlow to pretrained versions of FrameDiff (Yim et al. 2023b) (FrameDiff-ICML), the improved version on the authors’ GitHub (FrameDiff-Improved), Genie (Lin & AlQuraishi 2023), and RFDiffusion, which is the gold standard (Watson et al. 2023). We also retrain FrameDiff (FrameDiff-Retrained) on our dataset, which contains 10% more admissible structures, while inheriting the majority of the hyperparameters of FoldFlow. We provide a detailed description of all the metrics in §I.6. Figures 3 and 10 visualize generated samples and ESM-refolded structures.
We report our findings in table 2 and observe that FoldFlow outperforms FrameDiff-Retrained on all three metrics. We identify FrameDiff as the most comparable baseline as it is the current SOTA model that does not utilize pre-training while using comparable resources. In contrast to FoldFlow we highlight that RFDiffusion uses a pre-trained backbone and a significantly larger model ( vs. parameters), training set, and compute resources ( vs. GPU days). We also note that Genie is trained on a larger dataset ( vs. ), which hinders rigorous comparisons with FoldFlow. Next, we analyze the performance of FoldFlow on each metric in detail.
Designability. We measure designability using the self-consistency metric with ProteinMPNN (Dauparas et al. 2022) and ESMFold (Lin et al. 2022), counting the fraction of proteins that refold (-RMSD (scRMSD) < and mean scRMSD) over proteins at lengths . In table 2, we find that all FoldFlow models achieve significantly higher Frac. designability score than all FrameDiff models, and appreciably close the gap to RFDiffusion, e.g. for FoldFlow-OT and FrameDiff-Retrained respectively. When retrained on our dataset with 10% more samples, we find that FrameDiff is more designable, but is still below all FoldFlow models. We also note that while FoldFlow-OT creates the highest fraction of designable proteins (excluding RFDiffusion), it has relatively low diversity and novelty. We find that adding stochasticity with FoldFlow-SFM results in a model that beats FrameDiff-Improved on every metric and can dramatically improve novelty at the cost of worse designability (table 2). In fig. 9 in §G.1, we plot designability versus sequence length and observe the largest gains on sequence lengths .
Diversity. We use the average pairwise TM-score of the designable generated samples averaged across lengths as our diversity metric (lower is better). We find an inverse correlation between performance on designability and diversity metrics for FrameDiff models, which interestingly does not hold for FoldFlow models. We note that FoldFlow models have comparable diversity to the baselines with FoldFlow-OT and FoldFlow-SFM being the most diverse.
Novelty. Designing novel but realistic protein structures compared to the training data is also an important goal. Unlike conventional generative modeling problems, e.g. images, the novelty of proteins is particularly important since the entire premise of ML-driven drug discovery requires developing original drugs that may be vastly different than current human knowledge (training data) but also synthesizable (designable) (Marchand et al. 2022; Schneider et al. 2020; Schneider 2018). We measure novelty using two metrics: 1.) the fraction of designable proteins with TM-score as used in Lin & AlQuraishi 2023 (higher is better) and 2.) the average maximum TM-score of designable generated proteins to the training data (lower is better). In fig. 4(a) count the number of designable proteins as a function of the Max TM-score to the training set. We see that FoldFlow-SFM designs the most novel structures against all methods including RFDiffusion and Genie. This substantiates the hypothesis that the stochasticity of the learned SDE is crucial to the improved robustness in high dimensions of FoldFlow-SFM versus FoldFlow-Base, and FoldFlow-OT which allows it to sample designable proteins far outside of the support of the training set.
Ablation study of FoldFlow. Next, we ablate various additions to the FoldFlow-Base model in fig. 4(c) and report the full extended ablation in table 6. We find FoldFlow-OT creates the most designable model, but adding stochasticity helps increase novelty and diversity. We also find that inference annealing is critical to the performance of the model in terms of achieving higher designability.
3 Equilibrium Conformation Generation
Modeling various protein conformations is crucial in determining biological behaviours such as mechanisms of actions or binding affinity to other proteins. Unlike diffusion models, FoldFlow can easily be instantiated from any sampleable source distribution. To test this, we model the equilibrium distribution of a protein given initial predicted structures from pre-trained folding models including OmegaFold (Wu et al. 2022b) and ESMFold (Lin et al. 2022). The training target distribution consists of 200,000 frames at intervals of a molecular dynamics trajectory of the BPTI protein (Shaw et al. 2010); the inference of FoldFlow is tested against 20,000 unseen frames in that trajectory. FoldFlow can successfully model both the general set of conformations, as indicated by ICA of the dihedral angles in fig. 5(b), as well as the highly flexible residues, as seen in the Ramachandran plot in fig. 5(a). We note that our approach can capture all the modes of distribution in contrast to AlphaFold2, which does not model the flexibility well (fig. 5(c)). In table 3, we observe under the 2-Wasserstein () metric for all angles and only residue , FoldFlow with an informed prior outperforms a random prior, and FrameDiff which can only use an uninformed prior which further highlights a key advantage of FoldFlow over diffusion-based approaches.
Related Work
Equivariant generative models. There have been several efforts to incorporate symmetry constraints in generative models. These include building equivariant vector fields for CNFs (Köhler et al. 2020; Katsman et al. 2021; Garcia Satorras et al. 2021; Klein et al. 2023) and finite flows using the affine coupling transform (Dinh et al. 2017; Bose & Kobyzev 2021; Midgley et al. 2023). Applications in theoretical physics have also been impacted by equivariant flows (Boyda et al. 2020; Kanwar et al. 2020; Abbott et al. 2023). Lastly, beyond flows a new genre of models coming into prominence is based on the idea of equivariant score matching (De Bortoli et al. 2022; Brehmer et al. 2023) and diffusion models (Hoogeboom et al. 2022; Xu et al. 2022; Igashov et al. 2022).
Conclusion
Acknowledgements
AJB was supported by the Ivado Phd fellowship. KF is supported by NSERC Discovery grant (RGPIN-2019-06512), CIFAR AI Chairs, and a Samsung grant. CHL is supported by the Vanier scholarship. The authors would like to thank Clément Bonet and Gauthier Gidel for fruitful discussions on Riemannian geometry and optimal transport. We also thank Riashat Islam for providing feedback on early versions of this work. We thank Alexander Stein and the entire DreamFold team for providing a vibrant workspace that enabled this research. The authors would like to thank the restaurant Le Don Donburi for keeping our stomachs full and sustaining this research. Finally, the authors would like to acknowledge Anyscale and Google GCP for providing computational resources for the protein experiments.
References
Appendix A Theoretical Preliminaries
It is a matrix Lie group with the lie algebra given by:
Note that the inner product on Lie groups consumes elements of the Lie algebra and, because the left action is transitive, this inner product is well-defined for all tangent spaces of the group elements.
The distance induced by this metric is given by:
where and are the angle and axis of rotation for .
Similarly, the matrix logarithm can be expressed using the rotation angle:
Although this expression contains an infinite sum, Matthies et al. 1900 has shown that for , it can be approximated by a closed-form equation:
Riemannian Flow Matching is a generalization of Flow Matching on Riemannian manifold. Therefore, the setting as well as the main ideas are similar and are straightforward to adapt to the Euclidean case. This means that the objective is also to regress a conditional vector field built from conditional probability paths. In this section, we described the conditional probability paths and conditional vector fields that were used respectively by Lipman et al. 2022 and Tong et al. 2023b.
The main difference is that the conditional probability path is now a Gaussian conditioned on a latent variable with variance , . The conditional vector field has a closed form derived from the following Theorem:
The unique vector field whose integration map satisfies has the form
We now describe the Flow Matching method Lipman et al. 2022 and OT-CFM Tong et al. 2023b; Tong et al. 2023a; Pooladian et al. 2023b.
which is a probability path from the standard normal distribution () to a Gaussian distribution centered at with standard deviation (). If one sets to be the uniform distribution over the training dataset, the objective introduced by Lipman et al. 2022 is equivalent to the CFM objective (1) for this conditional probability path.
OT-Conditional Flow Matching (Tong et al. 2023b). As explained in the main paper, the probability path used in FM is not the optimal transport probability paths between the distributions and . Therefore, we want to get straighter flows for faster inference and more stable training. To achieve that, we leverage the optimal transport theory and want the probability path to be the Euclidean McCann interpolants defined as . However, as the map is intractable in practice, we rely on the Brenier theorem which makes a connection between the map and the optimal transport plan . Therefore we set the mean of Gaussian conditional probability path as and the latent distribution .
Appendix C Riemannian Optimal transport
Optimal transport in generative models. OT has been used in generative models for several approaches. For GANs, it was used as a loss function (Genevay et al. 2018; Fatras et al. 2021b; Salimans et al. 2018; Arjovsky et al. 2017). More recently, it was used to speeding up training and inference for continuous normalizing flows (Finlay et al. 2020; Liu et al. 2023b; Lipman et al. 2022; Tong et al. 2020; Tong et al. 2023b; Tong et al. 2023a), Schrödinger bridge models (Shi et al. 2023; Liu et al. 2023a; De Bortoli et al. 2021). In this section, we recall its basic definition over a Riemannian manifold. Then we discuss its empirical computation and we finish this section by proving Proposition 1.
Optimal transport on Riemannian manifold was first studied in the seminal work of McCann 2001 and we refer to Villani 2003; Villani 2008 for a review of all results. Recently, optimal transport has also drawn attention from the machine learning community, and we now give a longer introduction on this topic.
The (static) Kantorovich optimal transport problem seeks a mapping from one measure to another that minimizes a displacement cost. Formally, we define the -Wasserstein distance between distributions and on with respect to the cost as:
where denotes the set of all joint probability measures on whose marginals are and . To compute the optimal transport plan, we rely on the POT library Flamary et al. 2021. This problem is a relaxation of the well-known Monge formulation described in the main paper and that we recall now for the sake of readability.
The Monge optimal transport problem is defined as
When is a smooth compact manifold with no boundary and has a density, (McCann 2001, Proposition 9) shows that the map exists and is unique. This is an extension to Riemannian manifold of the well-known Brenier Theorem (Brenier 1991). The optimal transport map and the McCann interpolation have then the following form:
For empirical distributions, the Kantorovich problem is a linear program and can be efficiently solved with the simplex algorithm. We refer to (Peyré & Cuturi 2019, Chapter 3) for a review on how to solve the Kantorovich problem. However, when we deal with large datasets, computing and storing the transport plan for Optimal Transport (OT) can be challenging due to its cubic time and quadratic memory complexity with respect to the number of samples. To address this, a minibatch OT approximation is often employed. While this approach introduces some error compared to the exact OT solution (Fatras et al. 2020), it has been proven effective in various applications such as domain adaptation and generative modeling (Damodaran et al. 2018; Genevay et al. 2018). Specifically, during training, for each source and target minibatch, pairs of points are sampled from the optimal transport plan computed between the pair . We empirically show that the batch size can be small compared to the full dataset size and still give a good performance, which aligns with prior studies (Fatras et al. 2021b; Fatras et al. 2021a). This strategy is also at the heart of the OT-CFM methods (Tong et al. 2023b; Tong et al. 2023a; Pooladian et al. 2023b).
C.2 Proof of Proposition 1
We recall the proposition statement here for convenience and then prove it below.
Let be a connected, complete and smooth () Riemannian manifold without boundary, equipped with its standard volume measure. Let be two compactly supported distributions and set the ground cost with the geodesic distance on . Further, assume that is absolutely continuous with respect to the volume measure on . Then the Kantorovich and Monge problems admit a unique solution that is connected as follows , where is almost uniquely determined everywhere . Furthermore we have that for some -concave function .
Appendix D Stochastic Riemannian flow matching
D.3 Proof of Proposition 2
The proof of this claim follows a similar structure to Chen & Lipman 2023. We proceed by proving eq. 40. Dropping the distributions for conciseness, we have:
Appendix E Extended Figure Information
For the vector-field parametrization, the goal is to create a function that by construction lies on the tangent space of the manifold. For the toy experiments, this is done by using a -layer MLP, and projecting the output of the network to the tangent space of the input. That is, similar to Chen & Lipman 2023, we have:
In this section, we present the qualitative results of our toy experiments. In fig. 7, we can see that all the three models, FoldFlow-Base, FoldFlow-OT FoldFlow-SFM learn to correctly model the modes of the ground-truth distribution with a slight model shrinkage in the FoldFlow-Base.
Appendix G Additional Results and Analysis for the Protein Experiments
Empirical investigation of rotation norms for Inference Annealing.
In fig. 10, we show five proteins generated by FoldFlow-SFM from each backbone length. Here we show the generated structure in green and the best ESM-refolded structure out of eight sequences generated using ProteinMPNN. We can see that FoldFlow-SFM generates diverse folds that refold with diversity in secondary structure and overall 3D conformation.
In fig. 9 we compare the performance of models on designability, diversity, and novelty tasks for different backbone lengths. In particular, we can see that FoldFlow closes the gap between models without pretraining (Genie, FrameDiff, FoldFlow) and RFDiffusion in terms of designability, particularly on shorter sequences ().
We note there is a trade-off between designability and diversity/novelty, both at the short sequence lengths and as sequence length increases. For longer sequences (250, 300), while FoldFlow models are comparable in terms of designability, they generate significantly more diverse and novel structures as compared to all other models, even RFDiffusion (although RFDiffusion still generates significantly more designable proteins at these lengths.
We train our model in Pytorch using distributed data-parallel (DDP) across four NVIDIA A100-80GB GPUs for roughly 2.5 days. We note that this is substantially less than comparable models (table 4). RFDiffusion requires the use of pre-trained weights from RosettaFold which trained for 4 weeks on 64 V100 GPUs (Watson et al. 2023).
G.2 Equilibrium Conformation Generation Experiment
As described in Section 5.3 proteins take on many different physical conformations in the real world. These conformations dictate many important attributes of a protein’s behaviour, e.g., how one protein might bind to another. As a protein’s conformations generally do not deviate greatly from one another, a desirable approach would be to start from a noised version of a known conformation of the protein to generate another conformation. We hypothesize that the flows required to do this are easier to learn than starting from an uninformed source distribution. We find FoldFlow is an ideal candidate for this setting, and we show its efficacy in fig. 5.
Results in Section 5.3 show that FoldFlow generates conformations covering all modes of the true conformation distribution. Moreover, we sample different conformations of BPTI from AlphaFold2 and plot them on the Ramachandran and ICA plots, observing while FoldFlow can capture all modes of the distribution AlphaFold2 only captures one. Further fig. 11 shows that the KL divergence between the distribution of angles generated by FoldFlow is low and uniform, conveying it has learned the distribution of the target. We believe this is an exciting direction meriting larger experiments on more proteins in the future.
To contextualize these results, we compare the performance of these models with various baselines such as 250 samples from a random prior (used as the prior in FoldFlow-Rand and FrameDiff where each residue is sampled from ), 160 conformations sampled from AlphaFold 2, and 250 samples from the training set (Trainset). Results are averaged over 10 seeds for the random prior and the train set. The trainset represents a well-trained model as the Wasserstein distance is not zero even for empirical distributions drawn from the same distribution. The RandomPrior and the AlphaFold2 represent the random and informed priors respectively. All models are significantly better than these two priors, and FoldFlow approaches the performance of samples from the training set.
Figure 12 depicts the Ramachandran plots for residue 56 with scatter plots for FoldFlow FoldFlow-Rand and FrameDiff against a kernel density estimate (KDE) of the test set. We see that FoldFlow with the informed prior is able to model both modes where FoldFlow-Rand and FrameDiff both focus on the mode in the bottom right, centred at .
We have two major findings from this experiment:
An informed prior helps improve performance both overall and on the most flexible residue as seen by comparing the performance of FoldFlow and FoldFlow-Rand in table 3.
FoldFlow (with both informed and random priors) improve over FrameDiff on this task.
The equilibrium conformation generation task, studied here, is an example of a setting where an informed prior may be useful. Recent work has explored other applications of starting from an informed prior, such as protein docking (Somnath et al. 2023; Stärk et al. 2023), single-cell (Tong et al. 2023b) and image-to-image translation (Liu et al. 2023b).
Appendix H Further Discussion of FoldFlow and Related Models
Comparison between flow matching and diffusion approaches. While flow matching and diffusion models bear many similarities they also have key differences which we highlight in this appendix.
Flow matching based approaches enjoy the property of transporting any source distribution to any target distribution. This is in contrast to diffusion where one typically needs a Gaussian-like source distribution.
Flow matching approaches are readily compatible with optimal transport due to the same property of being able to transport and source to any target. Optimal transport which itself has the advantage of providing faster training with a lower variance training objective and reducing the numerical error in inference due to straighter paths. Diffusion models by themselves are not amenable to optimal transport but instead one can do entropic regularized OT. In Euclidean space, this corresponds to a Schrodinger bridge but this is not an optimal transport path is it stochastic.
In general, simulating an ODE is much more efficient than simulating an SDE during inference. Conditional flow-matching and OT-conditional flow matching (Lipman et al. 2022; Tong et al. 2023b) both learn ODEs as the learned flow corresponds to a continuous normalizing flow. Diffusion models on the other hand are SDEs and while being more robust to noise in higher dimensions require more challenging inference.
Comparison to FrameDiff. While our model uses a similar setup to FrameDiff, we introduce a number of improvements that help to stabilize training and improve performance. Indeed our additions lead to improvements on all metrics over the FrameDiff-Improved model released on GitHub which substantially improves on the designability over FrameDiff-ICML. We first recap the improvements made in FrameDiff-Improved over FrameDiff-ICML as detected in the code:
A bug in the score calculation for rotations means that there is a stop gradient in the rotation score calculation and FrameDiff-ICML is not trained to match the rotation score, which makes its performance quite impressive given this limitation. This bug is fixed in the FrameDiff-Improved model which uses a different score calculation.
The dataloader was switched from sampling uniform over proteins in the dataset, to uniform over clusters, then uniform within clusters. As we explore in Section I.5, this changes the distribution of proteins but overall increases diversity as there are many similar proteins in a small number of clusters Figure 13(c).
The rotation loss was changed to use a separate axis and angle component to reduce variance in the loss.
While these items improve the performance of FrameDiff, especially in terms of designability, there are still a few potential areas for improvement.
Appendix I Implementation Details and Experimental Setup
I.2 SDE Training and Inference
I.3 Vector Field Parametrization
Similar to the toy experiment, for the protein modelling case, the architecture is constructed such that the output vector lies on the tangent space.
I.4 Protein Task Hyperparameters
FoldFlow is implemented in Pytorch, and uses the invariant point attention (IPA) implementations from OpenFold (Ahdritz et al. 2022) in the backbone. We use the Adam optimizer with constant learning rate , , . The batch size depends on the length of the protein to maintain roughly constant memory usage. In practice, we set the effective batch size to
for each step. We set and weight the rotation loss with coefficient as compared to the translation loss which has weight .
We also used a trick from FrameDiff-Improved to stabilize the rotation loss. Instead of the loss on the rotation vector, we separate the loss into two components: one on the axis and one on the angle for the rotation vector. This seemed to reduce variance and numerical instability in the training.
I.5 Data and Data Sampling
We use a subset of PDB filtered with the same criteria as FrameDiff, specifically, we filter for monomers of length between 60 and 512 (inclusive) with resolution downloaded from PDB (Berman et al. 2000) on July 20, 2023. After filtering out any proteins with loops we are left with 22248 proteins. To support diversity, we sample uniformly over clusters with similarity of 30% as suggested in FrameDiff-Improved model https://cdn.rcsb.org/resources/sequence/clusters/clusters-by-entity-30.txt. Our model functions most efficiently with batches of proteins of the same length, so we each batch contains proteins of a single length. There are 4268 clusters in our dataset.
To assess the effects of our sampling methods on protein diversity and length distribution, we present three plots. fig. 13(a) illustrates the variability and range of protein lengths in the dataset, giving an overview of available lengths for sampling. fig. 13(b) shows the batch fraction per length, highlighting alterations in sequence length distribution during training due to uniform cluster sampling. Fig. 13(c), which uses a log scale on both axes, unveils the variation in cluster sizes and the skewness in protein distribution across clusters. Uniform cluster sampling enhances batch diversity, aiding model generalization over various protein sequences and structures. However, as observed in fig. 13(b), it slightly modifies the sequence length distribution during training. fig. 13(c) reveals an unevenness in protein distribution across clusters, with two bins containing approximately 14% of proteins.
I.6 Protein Metrics
where is the length of the target sequence, is the length of the common sequence after 3D structural alignment, is the distance (post alignment) of the residues in , and , and is a scaling factor to normalize across protein lengths. The TM-score ranges between with a TM-score of indicating perfectly aligned structure. In general a TM-score are considered roughly similar folds, with TM-score 0.2 corresponding to randomly chosen unrelated proteins.
The RMSD metric. The root-mean-square deviation (RMSD) is a simple metric over paired residues expressed as
where is again the distance between the residues heavy atoms . The RMSD score is length dependent unlike the TM-score, but has been shown to be a more stringent filtering step then TM-score for designability (Watson et al. 2023). In general, as compared to TM-score the RMSD metric is more sensitive local errors and less sensitive to global misalignments.
Designability. A generated protein structure is considered designable if there exists an amino acid sequence which refolds to that structure. We first generate 50 proteins at lengths {100, 150, 200, 250, 300 }, then apply ProteinMPNN with 8 times to generate 8 sequences for every generated structure. Finally we apply default ESMFold and aligned RMSD of the backbone atoms to calculate alignment of each ESMFold-refolded structure with the generated structure. We determine a protein designable if at least one of the 8 refolded structures has an scRMSD score . While a threshold of for designability is standard, this threshold may be unreasonably strict for longer backbones. However, it is unclear how this threshold should decay with increasing sequence length.
Finally, we note the imperfection of the self-consistency designability metric: when ESMFold does not produce the same structure as FoldFlow it does not imply FoldFlow’s structure is wrong, especially for longer sequences where protein folding models are known to perform worse. Both ProteinMPNN and ESMFold are imperfect, and the failure cases of these models has not been well characterized. While the false positive rate of this metric appears to be low, the false negative of this metric has not been quantified.
Diversity. We calculate all pairwise TM-scores for all generated structures that achieve the designability threshold of scRMSD for each length of protein. We then compute the mean over all of these pairwise TM-scores as our diversity metric. For this metric, a lower score is better. We choose to compare diversity on designable proteins as we do not want the designability score to be inflated by models which produce poor, proteins that may be very dissimilar to the space of refoldable proteins at that length.
Novelty. We calculate novelty using two metrics. The first is the minimum TM-score of designable generated proteins to the training data as described in Section I.5. The second metric is motivated by previous research (Lin & AlQuraishi 2023) and is the fraction of proteins that are both designable (scRMSD < 2 Å) and novel (avg. max TM-score < 0.5). We note that all models are not trained on the same dataset: FoldFlow and FrameDiff-Retrained share their dataset andFoldFlow and FrameDiff-ICML use very similar training datasets (only differing in about 10% of structures),. However, Genie and RFdiffusion use substantially larger datasets. Genie is trained on the Swissprot database (Jumper et al. 2021; Varadi et al. 2021) and RFdiffusion is at least pretrained on high-confidence AlphaFold2 structures. These larger training sets may cause novelty to be overestimated for these models as there are structures in their training sets that are far from the training set we use to test novelty against.
Error bounds in Table 2. We also report the standard error of the novelty and designability metrics in table 2. This is calculated by taking the standard error for each metric per sequence length, and then taking the mean over sequence lengths. We note that as the diversity is calculated as the averaged pairwise distances of designable proteins, each estimate of the mean is correlated resulting in an invalid estimate of the standard error.
Appendix J Extended Ablation of experiments
We perform a complete ablation experiment for the four features of the FoldFlow models: stochasticity, optimal transport, auxiliary losses in training and inference annealing. For each of the experiments, we evaluate the performance of the model on the designability, diversity and novelty metrics. These results can be seen in table 6. Overall, we observe the following trends:
Stochasticity improves robustness and the ability of the model to generate novel proteins in settings, as observed in the fraction of novel and designable proteins (Novelty-fraction).
Optimal transport improves the designability of the model by reducing the variance in the training objective in settings.
The auxiliary losses improve the designability of the models in settings.
Inference annealing improves the performance of all FoldFlow models in all metrics.