MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter
Zhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei, Yixin Cao, Kenji Kawaguchi, Xiang Wang, Tat-Seng Chua
Introduction
Language Models (LMs) have demonstrated significant achievements across various domains (Devlin et al., 2019; Zhao et al., 2023). Notably, the wealth of biochemical literature in LMs’ pretraining data has enabled LMs to obtain a high-level understanding of biochemical concepts and molecule properties. This can be reflected by their promising performances in biochemical and medical question-answering benchmarks (Taylor et al., 2022; OpenAI, 2023). Therefore, it becomes increasingly urgent to incorporate these LMs to augment research in chemistry and biology.
For this purpose, we aim to utilize LMs for molecule understanding. As shown in Figure 1a, most existing LMs (Touvron et al., 2023; Zhang et al., 2022; Zeng et al., 2022) represent molecules by their 1D Simplified Molecular Input Line Entry System (SMILES) strings (Weininger, 1988) and process them in a manner similar to texts. While convenient, treating molecules as strings overlooks the molecules’ 2D graph representations, which are crucial to human professionals in comprehending the molecule structures (Wells, 2012). To combat that, recent works (Su et al., 2022; Liu et al., 2022b) represent molecules as graphs and use a Graph Neural Network (GNN; Xu et al., 2019) as the molecular graph encoder. The graph encoder is trained jointly with an LM through cross-modal contrastive learning (Radford et al., 2021; Li et al., 2022), as illustrated in Figure 1b. However, the application scope of cross-modal contrastive learning is limited Alayrac et al. (2022): it is suitable for retrieval tasks, but is insufficient for open-ended molecule-to-text generation tasks, such as molecule captioning (Edwards et al., 2022) and molecule’s IUPAC name prediction (Taylor et al., 2022). This is because molecule-to-text generation is a conditional generation task Keskar et al. (2019); Raffel et al. (2020). It requires the LM to understand 2D graphs as the generation conditions, which contrastive learning cannot achieve. Su et al. (2022) attempt to directly input 2D graphs’ representations into LMs, however showing limited improvement.
To bridge this gap, we devise MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter. MolCA enables the LM to understand 2D graphs as inputs, therefore effectively conditioning the molecule-to-text generation process. To enable the LM to understand 2D graphs, we identify that the key challenge is cross-modal alignment (Li et al., 2023; Merullo et al., 2023; Alayrac et al., 2022): translating the representations of 2D graphs into 1D soft prompts (Li and Liang, 2021) in the text space so that the LM can understand. This translation is facilitated by the cross-modal projector, bridging the gap between the graph encoder’s representation space and the LM’s input space, as illustrated in Figure 1. Specifically, we implement the cross-modal projector as a Q-Former Li et al. (2023) due to its effectiveness in vision-language tasks. With an effective cross-modal projector, we can harness the power of existing large LMs Taylor et al. (2022); Touvron et al. (2023) for molecule-to-text generation. However, given a large LM with billion scale parameters, its efficiency of downstream fine-tuning arises as a new problem. Therefore, we integrate the LM with a uni-modal adapter, i.e., LoRA Hu et al. (2022), to enable its efficient adaptation.
As Figure 2 illustrates, MolCA uses a three-stage training pipeline to integrate its components. The two pretrain stages aim to develop the cross-modal alignment ability of the cross-modal projector. In pretrain stage 1, the projector and the encoder are trained to extract the molecule features that are the most relevant to the text. This stage endows the resulting model with powerful molecule-text retrieval ability. In pretrain stage 2, the cross-modal projector is connected to a frozen LM and trained for molecule captioning. This task forces the cross-modal projector to produce soft prompts that the LM can understand. In the final stage, MolCA is fine-tuned for downstream generation tasks.
Our contributions can be summarized as follows:
We propose MolCA, a pioneering method for molecular language modeling. MolCA enables an LM to perceive 2D molecular graphs, thereby facilitating molecule-to-text generation tasks.
MolCA sets new state-of-the-arts in a variety of benchmarks. It surpasses the baselines by 2.1 and 7.6 BLEU-2 for molecule captioning on CheBI-20 Edwards et al. (2022) and our curated PubChem324k dataset, respectively. Moreover, in predicting IUPAC names, MolCA shows a significant advantage of 10.0 BLEU-2 over the baselines. For molecule-text retrieval, MolCA outperforms the baselines by 20% retrieval accuracy in PubChem324k and achieves the best performances in PCDes Zeng et al. (2022) and MoMu datasets Su et al. (2022).
We conduct ablation studies to show MolCA’s effectiveness of incorporating 2D graphs into LMs for molecule-related tasks. Additionally, our quantitative analysis shows that incorporating 2D graphs helps improve the LM’s ability to count functional groups inside molecules.
Model Architecture
Here we introduce three key components of MolCA’s architecture: 1) a graph encoder for 2D structure understanding, 2) an LM for text generation, and 3) a cross-modal projector to connect the graph encoder and the LM. We describe the uni-modal adapter in Section 3.3.
Graph Encoder. Given the rich structural patterns in molecules, we leverage a GNN-based encoder to encode molecular graphs. Specifically, we employ a five-layer GINE (Hu et al., 2020) that is pretrained on 2 million molecules from the ZINC15 (Sterling and Irwin, 2015) dataset by contrastive learning (You et al., 2020). Given a molecular graph , the graph encoder can generate structure-aware features for every node of :
where denotes the number of nodes in .
Language Model. To achieve effective text generation performance, we employ Galactica (Taylor et al., 2022) as the base LM. Galactica is pretrained on a large collection of scientific literature, which encompasses fields like chemistry, biology, and medicine. Its promising performance in text-based science question-answering benchmarks (Hendrycks et al., 2021; Jin et al., 2019) underscores its understanding of high-level biochemical concepts. Notably, Galactica can process 1D SMILES of molecules, which can potentially benefit our downstream tasks. Galactica is a decoder-only transformer LM based on the OPT (Zhang et al., 2022) architecture.
Cross-Modal Projector. We implement the cross-modal projector as a Querying-Transformer (Q-Former) (Li et al., 2023) to map the graph encoder’s outputs to the LM’s input text space. As shown in Figure 3, Q-former has different procedures for processing 2D molecular graphs and 1D texts. Given text inputs, Q-Former inserts [CLS] tokens at the beginning and processes the texts by N layers of self-attention modules and feed-forward networks. The self-attention modules adopt causal masks (Raffel et al., 2020) when the pretraining task is text generation. On the other hand, given a molecular graph , Q-Former works as a molecule feature extractor. Specifically, it maintains a set of learnable query tokens as inputs. These query tokens can interact with the graph encoder’s output Z through the cross-attention modules (Vaswani et al., 2017) and extract molecule features. The cross-attention modules are added every two layers. Additionally, the query tokens can interact with the text inputs through the same self-attention modules. Note that, the query tokens and text inputs are processed by different feed-forward networks, in order to maintain capacities for processing molecules and texts.
We initialize Q-Former from Sci-BERT (Beltagy et al., 2019), an encoder-only transformer pretrained on scientific publications. Q-Former’s cross-attention modules are randomly initialized.
Training Pipeline
This section delves into the details of MolCA’s three-stage training pipeline (cf. Figure 2). The two pretrain stages leverage a dataset of molecule-text pairs to train the cross-modal projector and the graph encoder. The goal of pretraining is to translate 2D molecular graphs into soft prompts that a frozen LM can understand. The fine-tune stage focuses on efficient adaptation to downstream generation tasks.
In this stage, we aim to optimize the cross-modal projector (i.e., Q-Former) to extract the molecule features most relevant to the text input. This stage serves as a “warmup” training for the cross-modal projector before connecting to the LM. Inspired by BLIP2 Li et al. (2023), we simultaneously apply three cross-modal pretraining tasks that are tailored for Q-Former’s architecture: molecule-text contrasting, molecule-text matching, and molecule captioning. These pretraining tasks endow the Q-Former with a strong molecule-text retrieval ability. Therefore, we save the resulting model from this stage for downstream retrieval tasks. We now elaborate on the three pretraining tasks.
Molecule-Text Contrasting (MTC). We apply cross-modal contrastive learning (Radford et al., 2021) to train the Q-Former to extract text-revelant molecule features. In this task, query tokens and text inputs are fed into the Q-Former separately (left of Figure 3) to obtain Q-Former’s molecule representations and text representations.
where is the temperature-scaled cosine similarity. Temperature is empirically set to .
where is a uniform distribution; and are random negative samples in batch.
Similar to MTC, MTM also computes the similarity between molecule-text pairs. The difference is that MTM can capture more fine-grained similarity between a molecule and a text through the self-attention and cross-attention modules, compared to the simple cosine similarity used by MTC. Therefore, in retrieval experiments, we use MTC to first retrieve the top k samples and use MTM for re-ranking, thereby improving the performance.
Molecule Captioning (MCap). MCap aims to generate the molecule’s text description based on the molecule representations. For this task, we adopt a special masking strategy in self-attention modules to ensure that the queries learn to extract molecule features that correspond to the text descriptions. Specifically, we employ the bi-directional self-attention masks for queries, allowing them to see each other but not the text tokens. Further, we apply causal masks for texts on the same self-attention module to perform autoregressive decoding of text descriptions. Each text token can see the queries and the preceding text, but not the subsequent text tokens. Since the text tokens cannot directly interact with the graph encoder, they must obtain molecule information from the queries, forcing the queries to extract molecule information through the cross-attention modules. Let be the probability of Q-Former generating text for a graph . We use the following loss function:
2 Pretrain Stage 2: Aligning 2D Molecular Graphs to Texts via Language Modeling
In this stage, we aim to align the cross-modal projector’s outputs to the text space of a frozen LM. As Figure 5 illustrates, we feed the cross-modal projector’s representations of 2D molecular graphs to the frozen LM as inputs, and train the model to generate molecules’ text descriptions. This process encourages the cross-modal projector to provide representations that the LM can understand, so as to prompt the text generation. Additionally, we also use a molecule’s 1D SMILES to guide the generation (cf. Figure 5). This is because most LMs (Taylor et al., 2022; Touvron et al., 2023; Zhang et al., 2022) use SMILES during pretraining. Therefore, these LMs have established some correlations between SMILES and their text contexts. Thus, including SMILES can potentially prompt the corresponding biochemical knowledge. On the other hand, incorporating 2D graphs can help capture structural patterns that are hard to learn from 1D SMILES. We will show later in experiments that combining 2D graphs and 1D SMILES can boost performance.
Formally, consider a molecule-text pair and ’s SMILES repsentation , The cross-modal projector representations of are denoted as . We define as the text distribution parameterized by the frozen LM. We optimize the cross-modal projector and the graph encoder by minimizing the following loss function:
3 Fine-tune Stage: Uni-Modal Adapter for Efficient Downstream Adaptation
In this stage, we fine-tune MolCA for downstream generation tasks. As Figure 5 illustrates, we append a text prompt of the task description after the molecule representations. Then, we apply language modeling loss to fine-tune MolCA for generation tasks, such as molecule’s IUPAC name prediction.
where W is kept frozen and the newly added BA is trained during adaptation. Given a small , LoRA can effectively adapt the LM to downstream tasks while requiring little memory overhead for storing gradients.
Experiments
Here we briefly present the experimental settings. More details can be found in Appendix B.
PubChem324k Dataset. We collect PubChem-324k – a dataset containing 324k molecule-text pairs from the PubChem websitehttps://pubchem.ncbi.nlm.nih.gov. Table 1 presents the dataset statistics. Notice that, the dataset includes many uninformative texts, such as “The molecule is a peptide”. Therefore, we sample a high-quality subset of 15k pairs with text longer than 19 words for downstream tasks. This high-quality subset is further randomly divided into the train/valid/test sets. The remaining dataset, which is more noisy, is used for pretraining. Additionally, we filter our pretrain subset to exclude molecules from the valid/test sets of other downstream datasets, including CheBI-20 Edwards et al. (2022), PCDes Zeng et al. (2022), and MoMu Su et al. (2022) datasets. The dataset after filtering includes totally 313k molecule-text pairs.
Baselines. For generation tasks, we compare MolCA with the following baselines: T5 Raffel et al. (2020), MolT5 Edwards et al. (2022), and MoMu Su et al. (2022). For molecule-text retrieval, we also include these methods: MoleculeSTM Liu et al. (2022b), KV-PLM Zeng et al. (2022), and Sci-BERT Beltagy et al. (2019).
2 Molecule Captioning
We evaluate MolCA for molecule captioning on the datasets of PubChem324k and CheBI-20 Edwards et al. (2022). Specifically, we implement MolCA with the base LMs of Galactica, Galactica, and MolT5-Large. We employ full parameter fine-tuning for Galactica and MolT5-Large due to their smaller scales. We fine-tune MolCA and baselines on the dataset’s training set and report the test set performance selected by the valid set. Following Edwards et al. (2022), we adopt BLEU Papineni et al. (2002), ROUGE Lin (2004), and METEOR Banerjee and Lavie (2005) as the evaluation metrics. As shown in Table 2, we observe that:
1. MolCA consistently outperforms the baselines by a large margin. Specifcally, MolCA, achieves the highest performance on all metrics. It outperforms the baselines by 7.6 BLEU-2 on PubChem324k and 2.1 BLEU-2 on CheBI-20.
2. MolCA, outperforms baselines of larger sizes across all metrics, showing that MolCA’s advantage is not limited to model scale.
3 IUPAC Name Prediction
The International Union of Pure and Applied Chemistry (IUPAC) has established a standardized naming system for chemical compounds, known as IUPAC names (Favre and Powell, 2013). Notably, this naming system relies on identifying specific molecule structures, including hydrocarbon chains and double/triple bonds. Therefore, correctly predicting IUPAC names indicates a model’s proficiency to understand molecule structures. We fine-tune MolCA and baselines using the PubChem324k’s training set to generate a molecule’s IUPAC name. As shown in Table 3, MolCA consistently outperforms the baselines by a large margin of 10.0 BLEU-2, highlighting MolCA’s advantage in comprehending molecule structures.
4 Molecule-Text Retrieval
We evaluate MolCA for molecule-text retrieval on the datasets of PubChem324k, PCDes Zeng et al. (2022) and MoMu Su et al. (2022). Specifically, we evaluate MolCA’s checkpoint from pretrain stage 1 without further fine-tuning. For all experiments, MolCA first retrieves the top candidates using MTC, then employs the MTM module for re-ranking. We select Accuracy (Acc) and Recall@20 (R@20) as the evaluation metrics, and report the performances of retrieval in the entire test set. As shown in Table 4, we observe that:
1. MolCA demonstrates superior performance over baselines. Specifically, in PubChem324k, MolCA improves the accuracy by more than 20% over the baselines. In PCDes and MoMu, MolCA also consistently outperforms the baselines, demonstrating its effectiveness for molecule-text retrieval.
2. Incorporating MTM significantly improves MolCA’s performance. This can be attributed to MTM’s ability to model long-range interactions between molecule features and texts, achieved by the cross-attention and self-attention modules.
3. MolCA’s good performances can be partially attributed to our larger pretrain dataset – PubChem324k. As shown in Table 4(a), we compare the performances of MoMu’s original checkpoint (pretrained on 15k molecule-text pairs) with our reproduced MoMu using PubChem324k. The latter improves the retrieval accuracy by over 25%.
5 Ablation Study on Representation Types
Here we ablate the two representations types of molecules: 1D SMILES and 2D graphs. We compare MolCA with its two variants: 1) 1D SMILES: an LM that uses only 1D SMILES for pretraining and fine-tuning. For a fair comparison, we pretrain this variant on PubChem324k’s pretrain subset for molecule captioning before its downstream adaptation; 2) 2D Graph: this variant follows the original MolCA’s training pipeline, except not using 1D SMILES in pretrain stage 2 and fine-tune stage.
End Task Ablation. Table 5 presents the results for molecule-to-text generation and molecule property prediction Hu et al. (2020) tasks. We can observe that combing 2D graphs and 1D SMILES leads to improved performance in all the compared tasks. This demonstrates MolCA’s effectiveness in incorporating molecules’ 2D graph representations.
Counting Functional Groups (FGs). We ablate MolCA’s capability of counting 85 types of FGs inside molecules. An FG is a molecule’s subgraph that exhibits consistent chemical behaviors across different molecules Rong et al. (2020). Correctly counting FGs can help understand a molecule’s properties. As shown in Figure 6, incorporating 2D graphs significantly improves MolCA’s performance in counting FGs, thereby enhancing its ability in understanding molecule structures.
Related Works
Here we briefly review the molecule-related literature. We discuss MolCA’s relations to vision-language pretraining methods in Appendix A.
Molecule Understanding via 1D Language Modeling. Due to the extensive biochemical literature in their training corpus, some open-domain LMs Zhang et al. (2022); Touvron et al. (2023); Chowdhery et al. (2022) have obtained a high-level understanding of molecular and chemical concepts. This is demonstrated through their promising performances in text-related biochemical and medical question-answering benchmarks Hendrycks et al. (2021); Jin et al. (2019). Among these LMs, Galactica Taylor et al. (2022) shows competitive performances for using a corpus that is primarily composed of scientific literature. Focusing on the chemistry domain, KV-PLM (Zeng et al., 2022) models molecules by applying masked language modeling loss on 1D SMILES. Vaucher et al. (2021) propose to predict the chemistry experiment actions by reading chemical reaction equations. MolT5 (Edwards et al., 2022) presents several T5-based Raffel et al. (2020) LMs for SMILES-to-text and text-to-SMILES translations. Further, Christofidellis et al. (2023) propose to fine-tune T5 for chemical reaction prediction and retrosynthesis tasks. MolCA is different from these methods that exclusively utilize 1D SMILES to represent molecules. Instead, MolCA aims to enable LMs to perceive molecules’ 2D graph representations.
Molecule-Text Contrastive Learning. Driven by the demand of a molecule-text retrieval system, Text2Mol (Edwards et al., 2021) employs cross-modal contrastive learning to train a molecular graph encoder of GCNs Kipf and Welling (2017) and a text encoder of Sci-BERT Beltagy et al. (2019). Subsequent works Su et al. (2022); Liu et al. (2022b); Seidl et al. (2023) have proposed enhancements, including the addition of inter-modal contrastive learning loss (Su et al., 2022) and applying the model for text-based molecule editing (Liu et al., 2022b). However, cross-modal contrastive learning is unsuitable for open-ended conditional generation task Alayrac et al. (2022), because of its focus on learning a similarity function. To resolve the problem, we propose MolCA to enable the LM’s understanding of 2D molecular graphs, facilitating MolCA’s capability of open-ended molecule-to-text generation.
Conclusion and Future Works
In this work, we propose MolCA, a novel molecular language modeling method. MolCA aims to enable LMs to perceive 2D graphs for molecule-to-text generation. For this purpose, MolCA features a cross-modal projector to map representations of 2D graphs into the text space of LMs. It also employs a uni-modal adapter for efficient downstream adaptation. MolCA achieves state-of-the-art performances on molecule captioning and molecule-text retrieval benchmarks. Looking forward, we are interested in exploring LMs for 3D molecular modeling and drug discovery tasks.
Limitations
This work focuses on utilizing LMs’ generation ability for molecule-text tasks. Other interesting abilities of LMs, like in-context learning and chain-of-thought reasoning, are beyond the scope of this research. We leave that to future exploration.
While MolCA offers improvements over baselines, we observe that the current performance in molecule captioning is not yet sufficient for practical application. This can be attributed to the scale of pretraining data. To our knowledge, our PubChem324k dataset is the largest dataset of molecule-text pairs. However, compared to the 10M scale dataset (Changpinyo et al., 2021) for vision-language pretraining, our dataset, consists of 324k data points, is comparatively smaller and limits the model’s performance. Remedy solutions may include mining weakly supervised data from biochemical literature.
Broader Impacts
Our work has established new state-of-the-art performances in molecule captioning and molecule-text retrieval. It has broader impacts in two aspects: 1) for chemistry professionals, our method of molecule captioning and molecule-text retrieval could be useful tools, potentially speeding up their research process; 2) for individuals without specialized chemistry knowledge, our method could provide a more affordable way to access the basic chemical information of molecules.
Our model shares the risks of most LMs. It can generate inaccurate information and can potentially be abused to produce biased content. Further, considering the limited scale of our training data, we strongly advise strictly testing our model before applying it in real applications.
Acknowledgement
This research is supported by the National Natural Science Foundation of China (92270114) and the University Synergy Innovation Program of Anhui Province (GXXT-2022-040). This material is based upon work supported by the Google Cloud Research Credit program with the award (6NW8-CF7K-3AG4-1WH1). This research is supported by NExT Research Center.
References
Appendix A Complete Related Works
We present the complete literature review. In addition to the molecule-related literature, as addressed in the main body of the paper, we also discuss MolCA’s relation to vision-language pretraining.
Molecule Understanding via 1D Language Modeling. Due to the extensive biochemical literature in their training corpus, some open-domain LMs Zhang et al. (2022); Touvron et al. (2023); Chowdhery et al. (2022) have obtained a high-level understanding of molecular and chemical concepts. This is demonstrated through their promising performances in text-related biochemical and medical question-answering benchmarks Hendrycks et al. (2021); Jin et al. (2019). Among these LMs, Galactica Taylor et al. (2022) shows competitive performances for using a corpus that is primarily composed of scientific literature. Focusing on the chemistry domain, KV-PLM (Zeng et al., 2022) models molecules by applying masked language modeling loss on 1D SMILES. Vaucher et al. (2021) propose to predict the chemistry experiment actions by reading chemical reaction equations. MolT5 (Edwards et al., 2022) presents several T5-based Raffel et al. (2020) LMs for SMILES-to-text and text-to-SMILES translations. Further, Christofidellis et al. (2023) propose to fine-tune T5 for chemical reaction prediction and retrosynthesis tasks. MolCA is different from these methods that exclusively utilize 1D SMILES to represent molecules. Instead, MolCA aims to enable LMs to perceive molecules’ 2D graph representations.
Molecule-Text Contrastive Learning. Driven by the demand of a molecule-text retrieval system, Text2Mol (Edwards et al., 2021) employs cross-modal contrastive learning to train a molecular graph encoder of GCNs Kipf and Welling (2017) and a text encoder of Sci-BERT Beltagy et al. (2019). Subsequent works Su et al. (2022); Liu et al. (2022b); Seidl et al. (2023) have proposed improvements, including the addition of inter-modal contrastive learning loss (Su et al., 2022) and applying the model for text-based molecule editing (Liu et al., 2022b). However, cross-modal contrastive learning is unsuitable for open-ended conditional generation task Alayrac et al. (2022), because of its focus on learning a similarity function. To resolve the problem, we propose MolCA to enable the LM’s understanding of 2D molecular graphs, facilitating MolCA’s capability of open-ended molecule-to-text generation.
Vision-Language Pretraining (VLP). Both VLP and Molecular Language Modeling aim to bridge the gap between text and another modality. Notably, VLP methods of CLIP Radford et al. (2021) and others Li et al. (2022); Yao et al. (2022) use contrastive learning to connect a visual encoder and a text encoder. These methods can be applied for tasks like image-text retrieval and zero-shot image classification. Recently, a series of VLP works Tsimpoukelli et al. (2021); Merullo et al. (2023); Li et al. (2023); Alayrac et al. (2022) show that visual features can be aligned to the text space of LMs. This cross-modal alignment allows LMs to utilize their language generation and few-shot learning abilities for multi-modal tasks. MolCA draws inspiration from these findings. To the best of our knowledge, we are the first to align 2D molecular graphs to the text space of LMs. Furthermore, we incorporate a uni-modal adapter to improve the adaptation efficiency on downstream tasks.
Appendix B Experimental Settings
Pretrain Settings. MolCA’s pretrain stage 1 has 50 epochs and pretrain stage 2 has 10 epochs. Q-Former has query tokens (). Our optimizer’s configuration follows Li et al. (2023). We use the AdamW optimizer (Loshchilov and Hutter, 2019) with a weight-decay of . The learning rate is scheduled by a combination of linear warmup and cosine decay. The peak learning rate is 1e-4 and the warmup has 1000 steps.
Molecule Captioning. MolCA is fine-tuned for 100 epochs using the same configuration of optimizer and learning rate scheduler. LoRA is implemented using the OpenDelta library Ding et al. (2022) and the PEFT library (Mangrulkar et al., 2022). For the PubChem324k dataset, we set LoRA’s rank to and apply LoRA to Galactica’s modules of [q_proj, v_proj]. This configuration yields a LoRA adapter with 2M parameters, which constitutes 0.12% of the parameters in the Galactica. For the CheBI-20 dataset, we set LoRA’s rank to and apply LoRA to Galactica’s modules of [q_proj, v_proj, out_proj, fc1, fc2]. This configuration yields a LoRA adapter with 12M parameters, which constitutes 0.94% of the parameters in the Galactica.
IUPAC Name Prediction. We collect IUPAC names for molecules in the train/valid/test sets of PubChem324k using the PubChemPy libraryhttps://github.com/mcs07/PubChemPy. The experiment uses the same hyperparameters as the molecule captioning experiment. We append a text prompt “The molecule’s IUPAC name is” after the molecule representations as the task description (cf. Figure 5).
Molecule-Text Retrieval. We use MolCA’s checkpoint from pretrain stage 1 for retrieval without fine-tuning on any other datasets. This is similar to the setting of zero-shot retrieval in Su et al. (2022); Liu et al. (2022b).
Molecule Property Prediction. Following Hu et al. (2020), we fine-tune the models for 100 epochs and report the test performance selected by the valid set. For molecule classification, we attach a linear classifier after the mean pooling of the LM’s hidden states of the last layer. We use the AdamW optimizer with a constant learning rate of 1e-4 and weight decay of . This experiment uses the same LoRA configuration as the molecule captioning experiment in the PubChem324k dataset.
Counting Functional Groups (FGs). We use the molecules in PubChem324k’s train set for fine-tuning and use the molecules in the valid set for evaluation. Following Rong et al. (2020), we use RDkit Landrum (2013) to obtain the ground truth counts of FGs in every molecule. For each FG type, we employ a separate linear classifier to regress its numbers. Our model is trained using the Mean Square Error (MSE) loss function. Other settings, including optimizer and LoRA, are the same as the Molecule Property Prediction experiment.
Galactica. Following the instructions in Taylor et al. (2022), we wrap SMILES sequences with special tokens of [START_I_SMILES] and [END_I_SMILES] before feeding them into Galactica.
PubChem324k Dataset. Our dataset collection process follows the procedures described in Liu et al. (2022b). The resulting dataset is larger due to the frequent updates made to the PubChem database Kim et al. (2021). For each molecule in this website, we use the “description” field in its webpage as the corresponding text description. To avoid information leakage, we replace any common name or IUPAC name of the molecule at the beginning of texts with a text template (i.e., “The molecule”). Detailed statistics of PubChem324k are presented in Table 6.
Appendix C More Experimental Results
Molecule-Text Retrieval. Here we present MolCA’s complete molecule-text retrieval performance on the PubChem324k, PCDes, and MoMu datasets. Following Su et al. (2022), we report the performance of retrieval in a batch of 64 random samples and the performance of retrieval in the entire test set. As shown in Table 7, our conclusions align with those from Section 4.4: 1) MolCA consistently outperforms the baselines for molecule-text retrieval; 2) applying the MTM module for re-ranking is crucial for MolCA’s molecule-text retrieval performances.
Ablating the Pretrain Stages. We conduct ablation studies on MolCA’s two pretrain stages. As shown in Table 8, both the two pretrain stages have significant contributions to MolCA’s molecule captioning performances.
Ablating the Cross-Modal Projector. We compare the performances of our selected cross-modal projector Q-Former and a linear cross-modal projector. For the linear cross-modal projector, we feed the node representations from the graph encoder to the base LM after the linear projector layer. We tune the weights of the graph encoder, linear projector, and the base LM’s LoRA adapter. The experimental setting and hyperparameters are the same as those of MolCA. Table 9 shows the results. We can observe that: 1) Linear cross-modal projector underperforms Q-Former. We conjecture that a linear layer is suboptimal to bridge the modality gap between 2D molecules and 1D texts. This aligns with findings in the MME benchmark Fu et al. (2023), where Q-Former-based methods (e.g., BLIP-2, InstructBLIP Dai et al. (2023), MiniGPT-4 Zhu et al. (2023)) outperform linear cross-modal projector based method (e.g., LLaVA Liu et al. (2023)). 2) Linear cross-modal projector slightly outperforms the SMILES-only baseline. We attribute this improvement to the usage of 2D molecular graphs, but the gains are limited because the linear projector is less effective.
MolCA’s Generation Results. Figure 7 shows MolCA’s molecule-to-text generation results. The two samples of molecule captioning is also presented in Table 10. Specifically, we compare MolCA (i.e., 1D SMILES + 2D Graph) and its variant that is pretrained and fine-tuned using only 1D SMILES. We can observe that using both 1D SMILES and 2D graph leads to more accurate descriptions of molecule structures.
Computational Cost. We present the real-world training time of MolCA’s three training stages in Table 11. All experiments are conducted on two NVIDIA A100 40 GB GPUs. Notably, we observe that the fine-tuning stage is affordable in terms of computational resources.