Does Synthetic Data Generation of LLMs Help Clinical Text Mining?
Ruixiang Tang, Xiaotian Han, Xiaoqian Jiang, Xia Hu
Abstract
Recent advancements in large language models (LLMs) have led to the development of highly potent models like OpenAI’s ChatGPT. These models have exhibited exceptional performance in a variety of tasks, such as question answering, essay composition, and code generation. However, their effectiveness in the healthcare sector remains uncertain. In this study, we seek to investigate the potential of LLMs to aid in clinical text mining by examining their ability to extract structured information from unstructured healthcare texts, with a focus on biological named entity recognition and relation extraction. However, our preliminary results indicate that employing LLMs directly for these tasks resulted in poor performance and raised privacy concerns associated with uploading patients’ information to the LLM API. To overcome these limitations, we propose a new training paradigm that involves generating a vast quantity of high-quality synthetic data with labels utilizing LLMs and fine-tuning a local model for the downstream task. Our method has resulted in significant improvements in the performance of downstream tasks, improving the F1-score from 23.37% to 63.99% for the named entity recognition task and from 75.86% to 83.59% for the relation extraction task. Furthermore, generating data using LLMs can significantly reduce the time and effort required for data collection and labeling, as well as mitigate data privacy concerns. In summary, the proposed framework presents a promising solution to enhance the applicability of LLM models to clinical text mining.
Introduction
Recent advances in large language models (LLM) have dramatically improved the performance of various natural language processing (NLP) tasks, creating new opportunities for automating tasks traditionally performed by humans. OpenAI’s ChatGPT, for instance, has demonstrated its ability to perform as well as humans on MBA exams at the Wharton Business School 1, showcasing its competitiveness with human knowledge and its potential to assist professionals 2; 3. In the healthcare industry, LLMs hold tremendous potential for transforming the field by extracting valuable insights from unstructured data, such as electronic health records and digital medical data. By identifying crucial data points for population health management, clinical trials, and drug discovery, LLMs can help facilitate the development of new drugs and treatment plans. As the development and integration of LLMs in healthcare continue, professionals can expect significant improvements in patient outcomes and overall healthcare delivery.
LLMs possess a distinct advantage in their emergent abilities in Zero-Shot Learning, enabling them to learn and adapt to new tasks through prompt instructions, even if they have never encountered them before 4; 5. For example, by incorporating the instruction prompt ”Translate these sentences from [source Language] to [Target Language]:”, ChatGPT can compete favorably with commercial translation products, such as Google Translate 6. While LLM’s ability to conduct many NLP tasks makes it a valuable tool for users, it also raises significant privacy issues. One major concern is that sensitive information may be inadvertently revealed during the process. This is particularly true in healthcare, where the privacy and confidentiality of patient information are of utmost importance. Therefore, it is important to ensure that there is a mechanism to ensure robust privacy protections that prevent unauthorized access to sensitive information. In addition, reliability and usability are important issues that must be addressed when using LLM. Users must be able to rely on the accuracy and consistency of the model, which requires ongoing refining, testing, and evaluation to ensure that the system is functioning as intended and meeting the needs of users. In this paper, we will focus on improving the reliability of LLM for zero-shot tasks while mitigating the privacy risk.
In order to assess the zero-shot performance of current LLM models for healthcare tasks, we conducted experiments on ChatGPT to investigate its ability to extract structured information from unstructured healthcare texts, specifically for biological named entity recognition (NER) and relation extraction (RE) tasks. Our preliminary findings suggest that using ChatGPT directly only yields poor performance compared to SOTA models trained on the dataset, as indicated in Table 2 and 3. This result highlights the fact that while ChatGPT has demonstrated impressive inference and reasoning abilities in various classic natural language understanding (NLU) tasks, it is not adequate to apply ChatGPT alone to healthcare tasks since it was not specifically trained for this domain 7; 8. In addition to the performance limitations, integrating large language models into hospital systems raises privacy concerns as most LLMs are only available through their APIs, and healthcare providers cannot upload patient information directly to the LLMs’ APIs 9.
To bridge the gap, we propose a novel training paradigm to tackle the challenges of utilizing LLMs for healthcare tasks. Instead of directly applying LLMs in a zero-shot setting, we generate a large volume of synthetic data with labels using LLMs. To improve the quality and diversity of the synthetic data, we use a small amount of human-labeled examples as seeds and create appropriate prompts to guide LLMs in generating a variety of examples with varying sentence structures and linguistic patterns. A post-processing step is employed to eliminate low-quality or duplicated samples produced by LLMs. Finally, we employ the synthetic data to fine-tune a local pre-trained language model. Our experiments on four representative datasets show that the proposed pipeline significantly enhances the performance of the local model compared to LLMs’ zero-shot performance. In addition, the local offline model effectively addresses data privacy concerns by reducing the need for uploading patient data to LLM APIs. In conclusion, our innovative training framework enables the training of a local model with superior performance compared to using LLMs alone. It also mitigates potential privacy concerns and reduces the dependence on costly and time-consuming data collection and labeling.
Preliminaries
Biomedical Named Entity Recognition. Biomedical NER involves identifying and categorizing medical entities, such as diseases, symptoms, drugs, etc., in a medical text. NER uses an IOB (Inside, Outside, Begin) tagging scheme, where each word is assigned a tag indicating whether it is the beginning of a named entity (B), inside a named entity (I), or outside a named entity (O). For example, the sentence ”The symptoms suggest a possible case of rheumatoid arthritis.” would be tagged as ”O O O O O O O B-Disease I-Disease O”. Formally, a sentence in a medical text is denoted as a sequence of words , and the corresponding tags for each word in the sentence are denoted as , where tag is an element of the tag set {B, I, O}.
Biomedical Relation Extraction. Biomedical RE involves identifying and extracting the relationships between medical entities in a text, such as diseases and drugs, symptoms and treatments, etc. Formally, let be a sentence containing two medical entities and , and let be the relation between them. The MRE task can be formulated as a classification problem, where the goal is to learn a function that utilizes the context in the sentence to predict the relation between . The performance of the medical NER and RE models is typically evaluated using standard classification task metrics such as precision, recall, and F1-score.
Zero-shot Learning. Zero-shot learning is an emerging research paradigm that allows LLMs to perform tasks they have not been explicitly trained. This is accomplished by utilizing the LLMs’ capacity to produce coherent text based on a given prompt. The prompt serves as a guide, providing a corpus that describes the task at hand along with a set of potential outputs 5. The LLM then generates the most plausible output based on its acquired knowledge. Recent studies show that LLMs achieve promising zero-shot ability in various traditional NLU tasks 10.
Experimental Dataset. For the NER task, we consider two widely used datasets, including National Center for Biotechnology Information disease corpus (NCBI) and the BioCreative V CDR corpus (BC5CDR) 11. The NCBI dataset contains human-labeled annotations and is used to recognize disease names 12. The BC5CDR corpus collects PubMed articles with annotated chemicals, diseases, and chemical-disease interactions, and it is used in our study to evaluate the chemical and disease recognition task. For the RE task, we adopt two binary relation extraction datasets. The Gene Associations Database (GAD) dataset 13, which is a corpus of gene-disease associations curated from genetic association studies and contains 5,330 annotations. We also consider the EU-ADR corpus, a biomedical relation extraction dataset that contains 100 abstracts with relations between drugs, disorders, and targets 14. However, GAD and EU-ADR have noisy labels and are considered weakly supervised datasets. To ensure an accurate evaluation of the model’s performance, we employed three annotators to manually label 200 data samples from the original test datasets for both GAD and EU-ADR. The ground truth label was determined through majority voting.
Benchmarking LLM on Biomedical NER and RE Tasks
We conducted benchmark experiments on the ChatGPT model. To create prompts that would effectively trigger ChatGPT’s named entity recognition and relation extraction abilities, we take inspiration from the ChatGPT itself by prompting it for advice. Specifically, we asked ChatGPT to provide us with five concise templates that could be used to address biological NER and RE tasks. After testing the generated prompts on a validation set, we selected the most effective ones and added instructions for adapting them to suit downstream datasets, prompts are shown in Table 1.
In Tables 2 and 3, we report the performance of ChatGPT using our prompts and state-of-the-art models trained on the dataset. Results demonstrate that although ChatGPT shows some capability as a generalist model that can perform multiple traditional natural language understanding tasks 15, it performs inferior to SOTA models that have been fine-tuned on specific healthcare tasks. Specifically, ChatGPT performs slightly worse than SOTA in the biological relation extraction task, but there is a substantial performance gap between ChatGPT and SOTA on the biological named entity recognition task. For instance, the average disease recognition F1-score of the SOTA model is 88.60%, whereas ChatGPT achieves only 35.93%. Similarly, the average relation extraction F1-score of the SOTA model is 84.35%, while ChatGPT achieves 78.35%. This outcome is unsurprising, as ChatGPT is trained to tackle general natural language problems and has not been trained specifically for these tasks.
Exploring Synthetic Data Generation of ChatGPT for Clinical Text Mining
Motivation. In Section 3, we demonstrate that ChatGPT achieves only average performance for the biomedical relation extraction task and poor performance for the biomedical named entity recognition task. Additionally, uploading patient data directly would pose significant privacy concerns. Regulations such as GDPR 16 and CCPA 17 prohibit the upload of electronic health record information to ChatGPT, as this could potentially compromise patients’ protected health information (PHI). To leverage ChatGPT’s capabilities in assisting with healthcare-related tasks, we propose utilizing ChatGPT to generate a large volume of training data along with corresponding labels, which can be used to train a local model. This approach requires only a few examples to generate the entire training dataset, thereby resolving the low-resource problem commonly encountered in the healthcare domain. Additionally, the local model can address privacy concerns since the synthetic data contains no patient-sensitive information. Hospitals can use the local model instead of the ChatGPT to perform the downstream tasks while maintaining the privacy of patient data.
Prompt Engineering and Data Generation. To guide ChatGPT in generating the synthetic dataset for our tasks, we designed a suitable instruction prompt inspired by ChatGPT itself. Our approach involved prompting ChatGPT with the following request: “Provide five concise prompts or templates that can be used to generate data samples of [Task Descriptions].” ChatGPT then provided us with five candidate prompts for data generation. We generated 10 data samples using each prompt and manually compared their quality to select the best prompt. As shown in Figure 1, we repeated this process, asking ChatGPT to augment five prompts based on the previous best prompt until we arrived at the optimal prompt for data generation after three rounds of testing.
We present the optimal prompts for data generation in Table 6. To ensure the synthetic data’s high quality, we adopt the following measures: (1) To ensure the generated data is in the same distribution as the target dataset, we instruct ChatGPT to generate sentences that mimic the style of PubMed Journal articles, which is the data source of the NCBI and GAD dataset. (2) To prevent ChatGPT from generating duplicated examples, we added different seeds in the prompt for each round of data generation. For named entity recognition, we used different entity seeds, e.g., ”familial adenomatous polyposis,” and asked ChatGPT to generate sentences containing the seed entity. For the relation extraction task, we use examples from the original dataset combined in a format of sentence label as seeds. As we previously mentioned, since the GAD and EU-ADR datasets are highly noisy, we manually selected 50 positive and 50 negative samples from the original training dataset. For each round, we randomly sampled three positive examples and three negative examples as the seed samples. The generated examples, as shown in Table 4 and 5, are fluent and similar to sentences sampled from scientific articles.
Named Entity Recognition
In this section, we evaluate the effectiveness of our proposed synthetic data generation approach for the named entity recognition (NER) task, following the methodology outlined in Section 4. We first extract the seed entities from the training set and use them to generate synthetic sentences with annotations for the target entity type. Specifically, for each seed entity, we generate sentences with the corresponding entity annotations, in our experiments, we set . We then use the synthetic dataset to fine-tune three pre-trained language models. To evaluate the performance of the baseline models and our proposed methods, we use a subset of the test set for all datasets due to the computational limitations of ChatGPT. We report the precision, recall, and F1 scores for each model. The models evaluated in our experiments include three settings: (1) zero-shot, where the models were not fine-tuned on any dataset, and the model is ChatGPT. (2) models fine-tuned on synthetic data generated by our approach, and (3) models fine-tuned on the original training set. The pre-trained language models used in our experiments are BERT 18, RoBERTa 19, and BioBERT 20.
We present the comparison of the baseline methods and our methods in Table 7. From the experimental results, we observed that fine-tuning the models on the synthetic data generated using our approach leads to significant improvements compared to the zero-shot scenario in all the evaluated metrics. The average performance of BERT fine-tuned on synthetic data improved more than on Precision, on Recall, and on F1 than ChatGPT. Moreover, in some cases, our proposed method even achieves comparable performance to the models fine-tuned on the original training set. For example, for the BC5CDR Chemical dataset, the Recall of the BERT model fine-tuned on synthetic data obtained , improved from in the zero-shot scenario, which is comparable to the recall when fine-tuned on the original training set. The results demonstrate the effectiveness of our synthetic data generation approach in improving the performance of these models.
The Effect of the Number of the Generated Sentences.
To investigate the impact of the number of synthetic sentences generated on the effectiveness of our proposed method, we conducted experiments with varying numbers of synthetic sentences and ratios of seed entities. As mentioned previously, we generated sentences with annotations for seed entities. In the first experiment, we used seed entities for synthetic data generation, while in the second experiment, we generated $$ sentences for each entity. The results are presented in Figure 2. Our findings showed that increasing the number of synthetic sentences can improve model performance up to a certain point, beyond which the improvement becomes marginal. Similarly, adjusting the ratio of synthetic to real entities in the training dataset can enhance model performance, especially for under-represented entities.
1 Relation Extraction
In this section, we aim to assess the efficacy of our synthetic data generation approach for relation extraction. We follow the methodology outlined in Section 4 and randomly sample three positive and three negative examples from the 100 training manually labeled dataset to use as seed examples. For each round of generation, we generate three positive sentences and three negative sentences. In this way, we collect 6437 and 6424 synthetic examples for the GAD and EU-ADR, respectively. Given that the GAD and EU-ADR datasets are noisy, we manually label 200 test samples from the original test dataset to evaluate our models. We measure models’ performance based on their precision, recall, and F1 scores. Our experiments cover three model settings: (1) zero-shot, where we directly leverage ChatGPT with the prompt shown in Table 1 for inference, (2) models fine-tuned on synthetic data generated by our approach, and (3) models fine-tuned on the original training set. Similarly to the NER task, we adopt BERT, RoBERTa, and BioBERT as the backbone models for our experiments.
We present the comparison of the baseline methods and our methods in Table 8. From the experimental results, we observed that fine-tuning the models on the synthetic data generated using our approach leads to notable improvements compared to the zero-shot scenario in all the evaluated metrics. The average performance on Precision, Recall, and F1 improve more than 6%, 10%, and 8% percentage than ChatGPT. The model trained on the synthetic data achieves comparable performance to the models fine-tuned on the original training set. Notably, for the GAD dataset, the model trained on the synthetic dataset achieved slightly better results than the original dataset. The results demonstrate the effectiveness of our synthetic data generation approach for the relation extraction task.
The Effect of the Number of the Generated Sentences.
To investigate the impact of the number of synthetic sentences on the effectiveness of our proposed method, we conducted experiments with varying numbers of synthetic sentences and ratios of seed examples. As mentioned previously, we collected 6437 and 6424 synthetic examples for the GAD and EU-ADR, respectively. In the first experiment, we used a range of synthetic data to train our local model, while in the second experiment, we varied the pool size of our seed examples from 0 to 100. The results of these experiments are presented in Figure 3. Our findings indicate that increasing the number of synthetic sentences can improve model performance up to a certain point, beyond which the improvement becomes marginal. Specifically, we found that 3500 synthetic sentences are sufficient for obtaining optimal results. Additionally, using a larger number of seed examples can increase the quality and diversity of the generated data. Our experiments demonstrated that 80 seed examples are sufficient for both relation extraction tasks. It is worth noting that not providing seed examples can lead to ChatGPT generating duplicated examples, resulting in a significant drop in model performance.
Analysis of Generated Texts
Our previous research has demonstrated that ChatGPT is capable of producing high-quality synthetic data. However, a potential concern is that ChatGPT has been trained on a publicly available dataset, which means that it may have already encountered the dataset used in our experiments. This raises the possibility of ChatGPT inadvertently leaking information from the original dataset. To address this issue, we utilized the sentence transformer to obtain embeddings for both the original and synthetic data, and then projected them using T-SNE. The resulting distribution of sentences, as illustrated in Figure 4, revealed distinct patterns between the synthetic and original data, indicating that ChatGPT did not simply memorize and reproduce the dataset. This distribution shift can also explain the observed performance gap between models fine-tuned on synthetic versus original data. Our future work aims to explore methods for producing synthetic data with a similar distribution as the original data.
Related Work
In this section, we review the literature related to the topic of our paper, including previous research on Large Language Models and NLP for Biomedical applications..
Large Language Models. Recently, Large Language Models (LLMs) have attracted increasing attention due to their high performance and capability to understand natural language. Consequently, researchers and practitioners are exploring the use of LLMs to assist experts in a variety of domains, including education 21, healthcare 22; 23, and content creation 24. OpenAI GPT-3 7 series has been a breakthrough for LLMs, which is trained by generating the next word in a sequence given the preceding words. These models have been widely used for tasks such as text generation and language modeling, and have also achieved impressive results on many benchmarks. The following works, such as BLOOM 25, PaLM 26 and LLaMA 27, are optimized for specific tasks such as code generation and document ranking. Recently, the instruction-tuned version of GPT-3, ChatGPT, has emerged as a game-changer in LLMs. ChatGPT is capable of generating coherent text from scratch. In this work, we use ChatGPT as the zero-shot baseline and we use it to generate syntectic data for us.
NLP for Biomedical. The Natural Language Processing (NLP) technique is widely applied in the biomedical domain, as evidenced by numerous studies 28; 29. NLP for Biomedical has various applications, including the analysis of electronic health records (EHRs)30; 31; 32; 33, drug discovery34; 35, and medical chatbots 36; 37. The use of LLMs for Biomedical is gaining traction among both industry and academic researchers. Previous work 20; 38; 39 has explored the application of NER and RE to biomedical tasks. Biomedical NER and RE has diverse usage in the healthcare domain, including analyzing EHRs 40; 41; 42, extracting clinical trials 43; 44, and drug development 45; 46; 47. The major challenges facing NLP in Biomedical include developing accurate models for biomedical text analysis and ensuring patient data privacy. In this work, we propose the use of synthetic data to fine-tune offline models, which can not only improve prediction accuracy but also protect patient privacy.
Conclusion
In this study, we set out to explore the potential of ChatGPT to assist with clinical text mining tasks, with a particular focus on named entity recognition and relation extraction. However, our initial attempts to use ChatGPT directly for these tasks yielded unsatisfactory results and raised privacy concerns. Therefore, we developed a new framework that involved generating high-quality synthetic data with ChatGPT and fine-tuning a local offline model for downstream tasks. The use of synthetic data resulted in significant improvements in the performance of these downstream tasks, while also reducing the time and effort required for data collection and labeling, and addressing data privacy concerns as well. In the future, we aim to further refine our framework to enhance the data quality and extend its application to other clinical tasks.