Evaluation of ChatGPT on Biomedical Tasks: A Zero-Shot Comparison with Fine-Tuned Generative Transformers
Israt Jahan, Md Tahmid Rahman Laskar, Chun Peng, Jimmy Huang
Introduction
The rapid growth of language models Rogers et al. (2021); Zhou et al. (2023) in the field of Natural Language Processing (NLP) in recent years has led to significant advancements in various domains, including the biomedical domain Kalyan et al. (2022). Although specialized models (e.g., BioBERT Lee et al. (2020), BioBART Yuan et al. (2022a), BioGPT Luo et al. (2022), etc.) have shown promising results in the biomedical domain, they require fine-tuning using domain-specific datasets. This fine-tuning process can be time-consuming due to the requirement of task-specific large annotated datasets. In contrast, zero-shot learning enables models to perform tasks without the need for fine-tuning on task-specific datasets. ChatGPT, a large language model, has demonstrated impressive zero-shot performance across various tasks Laskar et al. (2023). However, its performance in the biomedical domain remains to be thoroughly investigated. In this regard, this paper presents a comprehensive evaluation of ChatGPT on four key biomedical tasks: relation extraction, question answering, document classification, and summarization.
In this paper, our primary objective is to explore the extent to which ChatGPT can perform these tasks without fine-tuning and assess its performance by comparing with state-of-the-art generative fine-tuned models, BioGPT and BioBART. To our best knowledge, this is the first work that evaluated ChatGPT on benchmark biomedical datasets. Our evaluation of ChatGPT can have a profound impact on the biomedical domain that lacks domain-specific datasets by exploring its zero-shot learning capabilities. To ensure the reproducibility of our evaluation and to help facilitate future research, we will release all the ChatGPT-generated responses along with our evaluation code here: https://github.com/tahmedge/chatgpt-eval-biomed.
Related Work
The effective utilization of transformer-based Vaswani et al. (2017) NLP models like BERT Devlin et al. (2019) have also led to significant progress in the biomedical domain Lee et al. (2020); Alsentzer et al. (2019); Beltagy et al. (2019); Gu et al. (2020); Peng et al. (2019) in recent years. BERT leverages the encoder of the transformer architecture, while GPT leverages the decoder of the transformer. In addition to these models, sequence-to-sequence models like BART Lewis et al. (2019) that leverage both the encoder and the decoder of transformer have also emerged as a powerful approach in various text generation tasks.
It has been observed that domain-specific pre-training of these models on the biomedical text corpora followed by fine-tuning on task-specific biomedical datasets have helped these models to achieve state-of-the-art performance in a variety of BioNLP tasks Gu et al. (2021). However, one major limitation of using such fine-tuned models is that they require task-specific large annotated datasets, which is significantly less available in the BioNLP domain in comparison to the general NLP domain. In this regard, having a strong zero-shot model could potentially alleviate the need for large annotated datasets, as it could enable the model to perform well on tasks that it was not trained on.
Recently, large autoregressive language models like GPT-3 Brown et al. (2020) have demonstrated impressive few-shot learning capability. More recently, a new variant of GPT-3, called the InstructGPT model Ouyang et al. (2022) has been proposed that leverages the reinforcement learning from human feedback (RLHF) mechanism. The resulting InstructGPT models (in other words, GPT-3.5) are much better at following instructions than the original GPT-3 model, resulting in an impressive zero-shot performance across various tasks. ChatGPT, a very recent addition to the GPT-3.5 series, has been trained using dialog-based instructional data alongside its regular training phase. Though ChatGPT has demonstrated strong zero-shot performance across various NLP tasks Laskar et al. (2023); Qin et al. (2023); Bang et al. (2023); Yang et al. (2023), it is yet to be investigated in the biomedical domain. To this end, this paper aims to evaluate ChatGPT in the biomedical domain.
Our Methodology
For a given test sample , we prepare a task instruction and concatenate the text in the test sample with the task instruction to construct the prompt . Then the prompt is given as input to ChatGPT (gpt-3.5-turbo) to generate the response . In this paper, we evaluate ChatGPT on 4 biomedical tasks across 11 benchmark datasets. Below, we describe these tasks, the datasets we use for evaluation, and the prompt that we construct for each task depending on the respective dataset.
Given a text sequence , the biomedical relation extraction task aims to extract relations between entities mentioned in the text by identifying all possible relation triplets. In this paper, we evaluate drug-target-interaction in the KD-DTI dataset Hou et al. (2022), chemical-disease-relation in the BC5CDR dataset Li et al. (2016), and drug-drug-interaction in the DDI dataset Herrero-Zazo et al. (2013). Our prompts for these datasets are demonstrated in Table 1.
(ii) Document Classification:
Given a text document , the goal is to classify the type of the document. For this task, we use the HoC (the Hallmarks of Cancers corpus) dataset Baker et al. (2016) that consists of 1580 PubMed abstracts. This dataset was annotated at the sentence level by human experts among ten currently known hallmarks of cancer. Our prompt is shown in Table 1.
(iii) Question Answering:
For the question-answering task, we evaluate the performance of ChatGPT on the PubMedQA dataset Jin et al. (2019). Here, the objective is to determine whether the answer to a given question can be inferred from the reference context. We give the question, the reference context, and the answer as input to ChatGPT to determine whether the answer to the given question can be inferred from the given reference context, with ChatGPT being prompted to reply either as yes, no, or maybe (see Table 1 for details).
(iv) Abstractive Summarization:
Given a text sequence , the goal is to generate a concise abstractive summary of . To this end, we evaluate ChatGPT on various biomedical summarization tasks, such as healthcare question summarization (we used MeQSum Abacha and Demner-Fushman (2019) and MEDIQA-QS Abacha et al. (2021) datasets), medical answer summarization (we used MEDIQA-ANS Savery et al. (2020) and MEDIQA-MAS Abacha et al. (2021) datasets), and dialogue summarization (we used the iCliniq and HealthCareMagic datasets Zeng et al. (2020); Mrini et al. (2021) for doctor-patient dialogue summarization to generate short queries for healthcare forums describing patient’s medical conditions). We show our prompts for this task in Table 2.
Experiments
Since ChatGPT is a generative model, we consider two state-of-the-art generative transformers as our baselines. Below, we first present these baselines, followed by presenting the results.
The backbone of BioGPT Luo et al. (2022) is GPT-2 Radford et al. (2019), which is a decoder of the transformer. The BioGPT model was trained over PubMed titles and abstracts via leveraging the standard language modeling task. We compare zero-shot ChatGPT with BioGPT models fine-tuned on relation extraction, document classification, and question-answering tasks.
BioBART:
BioBART is a sequence-to-sequence model that was pre-trained over PubMed abstracts Yuan et al. (2022a). The pre-training process involves reconstructing corrupted input sequences. We compare the zero-shot ChatGPT with BioBART fine-tuned on abstractive summarization datasets.
2 Results & Discussion
We first compare the performance of ChatGPT with BioGPT on relation extraction, document classification, and the question-answering task (see Table 3). Then we compare its performance with BioBART on summarization datasets (see Table 4). More evaluation details are given in Appendix A.1.
We observe that in the BC5CDR and KD-DTI datasets for relation extraction, ChatGPT led to higher recall scores but much lower precision scores compared to the fine-tuned BioGPT model. This is because ChatGPT tends to generate long and descriptive responses, leading to many inaccurate relation extractions. Though in terms of F1, it outperforms fine-tuned BioGPT in the BC5CDR dataset, it fails to outperform in the KD-DTI dataset. More importantly, it outperforms BioGPT in the DDI dataset in all metrics: Precision, Recall, and F1.
While analyzing the results in different datasets, we observe that in both BC5CDR and DDI datasets where ChatGPT outperforms BioGPT, the training set is small, only 500 and 664 instances, respectively. On the other hand, in the KD-DTI dataset where ChatGPT fails to outperform BioGPT, the training set contains 12000 instances. This gives us a strong indication that even in the biomedical domain, zero-shot ChatGPT can outperform fine-tuned biomedical models in smaller-sized datasets.
We also observe that more descriptive prompts may help ChatGPT to obtain better Precision scores. Contrary to the KD-DTI dataset, we describe the definition of each interaction type in the DDI dataset (see Table 1) where ChatGPT performs the best. To further investigate the effect of prompts in relation extraction, we evaluate the performance in BC5CDR with a new prompt: i. Identify the chemical-disease interactions in the passage given below: [PASSAGE]. We observe that the Precision, Recall, and F1 scores are decreased by 16.07%, 10.3%, and 14.29%, respectively, with this prompt variation.
Document Classification Evaluation:
We observe that in the HoC dataset, the zero-shot ChatGPT achieves an F1 score of 59.14, in comparison to its counterpart fine-tuned BioGPT which achieves an F1 score of 85.12. We also investigate the effect of prompt tuning by evaluating with two new prompts that are less descriptive (see Appendix A.2 for more details): i. Prompting without explicitly mentioning the name of 10 HoC classes, drops F1 to 38.20. ii. Prompting with the name of each HoC class is given without providing the definition of each class, drops the F1 score to 46.93.
Question Answering Evaluation:
We observe that in the PubMedQA dataset, the zero-shot ChatGPT achieves much lower accuracy than BioGPT (51.60 by ChatGPT in comparison to 78.20 by BioGPT). However, the BioGPT model was fine-tuned on about 270K QA-pairs in various versions of the PubMedQA dataset for this task. While ChatGPT achieves more than 50% accuracy even without any few-shot examples in the prompt.
Summarization Evaluation:
We observe that in terms of all ROUGE scores Lin (2004), ChatGPT performs much worse than BioBART in datasets that have dedicated training sets, such as iCliniq, HealthCareMagic, and MeQSum. Meanwhile, it performs on par with BioBART in the MEDIQA-QS dataset. More importantly, it outperforms BioBART in both MEDIQA-ANS and MEDIQA-MAS datasets. Note that MEDIQA-ANS, MEDIQA-MAS, and MEDIQA-QS datasets do not have any dedicated training data and ChatGPT achieves comparable or even better performance in these datasets compared to the BioBART model fine-tuned on other related datasets Yuan et al. (2022a). This further confirms that zero-shot ChatGPT is more useful than domain-specific fine-tuned models in biomedical datasets that lack large training data.
Conclusions and Future Work
In this paper, we evaluate ChatGPT on 4 benchmark biomedical tasks to observe that in datasets that have large training data, ChatGPT performs quite poorly in comparison to the fine-tuned models (BioGPT and BioBART), whereas it outperforms fine-tuned models on datasets where the training data size is small. These findings suggest that ChatGPT can be useful in low-resource biomedical tasks. We also observe that ChatGPT is sensitive to prompts, as variations in prompts led to a noticeable difference in results.
Though in this paper, we mostly evaluate ChatGPT on tasks that require it to generate responses by only analyzing the input text, in the future, we will investigate the performance of ChatGPT on more challenging tasks, such as named entity recognition and entity linking Yadav and Bethard (2018); Yan et al. (2021); Yuan et al. (2022b); Laskar et al. (2022a, b, c), as well as problems in information retrieval Huang et al. (2005); Huang and Hu (2009); Yin et al. (2010); Laskar et al. (2020, 2022d). We will also explore the ethical implications (bias or privacy concerns) of using ChatGPT in the biomedical domain.
Limitations
Ethics Statement
The paper evaluates ChatGPT on 4 benchmark biomedical tasks that require ChatGPT to generate a response based on the information provided in the input text. Thus, no data or prompt was provided as input that could lead to ChatGPT generating any responses that pose any ethical or privacy concerns. This evaluation is only done in some academic datasets that already have gold labels available and so it does not create any concerns like humans relying on ChatGPT responses for sensitive issues, such as disease diagnosis. Since this paper only evaluates the performance of ChatGPT and investigates its effectiveness and limitations, conducting this evaluation does not lead to any unwanted biases. Only the publicly available academic datasets are used that did not require any licensing. Thus, no personally identifiable information has been used.
Acknowledgements
We would like to thank all the anonymous reviewers for their detailed review comments. This work is done at York University and supported by the Natural Sciences and Engineering Research Council (NSERC) of Canada and the York Research Chairs (YRC) program.
References
Appendix A Appendix
Since ChatGPT generated responses can be lengthy, and may sometimes contain unnecessary information while not in a specific format, especially in tasks that may have multiple answers (e.g., Relation Extraction), it could be quite difficult to automatically evaluate its performance in such tasks by comparing with the gold labels by using just an evaluation script. Thus, for some datasets and tasks, we manually evaluate the ChatGPT generated responses by ourselves and compare them with the gold labels. Below we describe our evaluation approach for different tasks:
Relation Extraction: The authors manually evaluated the ChatGPT generated responses in this task by comparing them with the gold labels. To ensure the reproducibility of our evaluation, we will release the ChatGPT generated responses.
Document Classification: We created an evaluation script and identifies if the gold label (one of the 10 HoC classes) is present in the ChatGPT generated response. For fair evaluation, we lowercase each character in both the gold label and the ChatGPT generated response. Our evaluation script will be made publicly available to ensure reproducibility of our findings.
Question Answering: Similar to Document Classification, we also evaluated using an evaluation script that compares the gold label and the ChatGPT generated response (here, we also convert each character to lowercase). The evaluation script will also be made public.
Abstractive Summarization: We used the HuggingFace’s Evaluatehttps://huggingface.co/docs/evaluate/index library Wolf et al. (2020) to calculate the ROUGE scores and the BERTScore for the Abstractive Summarization task evaluation.
A.2 Effects of Prompt Variation
We investigate the effects of prompt tuning in the HoC dataset by evaluating the performance of ChatGPT based on the following prompt variations:
Prompting with explicitly defining the 10 HoC classes achieves an F1 score of 59.14 (see Row 1 in Table 5).
Prompting without explicitly mentioning the name of 10 HoC classes, drops F1 to 38.20 (see Row 2 in Table 5).
Prompting with the name of each HoC class is given without providing the definition of each class, drops the F1 score to 46.93 (see Row 3 in Table 5).
Our findings demonstrate that more descriptive prompts yield better results.
A.3 Sample ChatGPT Generated Responses
Some sample prompts with the ChatGPT generated responses for Relation Extraction, Document Classification, and Question Answering tasks are given in Table 6 and for the Abstractive Summarization task are given in Table 7.