A Survey of Large Language Models in Medicine: Progress, Application, and Challenge
Hongjian Zhou, Fenglin Liu, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Peilin Zhou, Junling Liu, Yining Hua, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, David A. Clifton
Introduction
The recently emerged general large language models (LLMs) , such as PaLM , LLaMA , GPT-series , and ChatGLM , have advanced the state-of-the-art in various natural language processing (NLP) tasks, including text generation, text summarization, and question answering. Inspired by these successes, several endeavors have been made to adapt general LLMs to the medicine domain, leading to the emergence of medical LLMs . For example, based on PaLM , MedPaLM and MedPaLM-2 have achieved a competitive accuracy of 86.5 compared to human experts (87.0 ) in the United States Medical Licensing Examination (USMLE) . Based on publicly available general LLMs (e.g. LLaMA ), several medical LLMs, including ChatDoctor , MedAlpaca , PMC-LLaMA , BenTsao , and Clinical Camel , have been introduced. As a result, medical LLMs have gained growing research interests in assisting medical professionals to improve patient care .
Although existing medical LLMs have achieved promising results, there are some key issues in their development and application that need to be addressed. First, many of these models primarily focus on biomedical NLP tasks, such as dialogue and question answering, but their practical utility in clinical practice is often overlooked . Recent research has begun to explore the potential of medical LLMs in different clinical scenarios, including Electronic Health Records (EHRs) , discharge summary generation , health education , and care planning . However, they mainly perform case studies and invite clinicians to perform the human evaluation on a small number of samples, thus lacking evaluation datasets for assessing model performance in clinical scenarios. Second, most existing medical LLMs report their biomedical NLP performances mainly on answering medical questions, neglecting other biomedical tasks, such as text summarization, relation extraction, information retrieval, and text generation. These research gaps motivate this survey that offers a comprehensive review of the development of LLMs and their applications in medicine. We aim to cover topics on existing medical LLMs, various biomedical tasks, clinical applications, and arising challenges.
As shown in Figure 1, this review seeks to answer the following questions:
What are LLMs? How can medical LLMs be effectively built? (Section 2)
How are the current medical LLMs evaluated? What capabilities do medical LLMs offer beyond traditional models? (Section 3)
How should medical LLMs be applied in clinical settings? (Section 4)
What challenges should be addressed when implementing medical LLMs in clinical practice? (Section 5)
How can we optimize the construction of medical LLMs to enhance their applicability in clinical settings, ultimately contributing to medicine and creating a positive societal impact? (Section 6)
For the first question, we summarize the principles of existing medical LLMs, detailing their basic structures, the number of parameters, and the data used for model development. Additionally, we provide insight into the construction process of these models. This information is valuable for researchers and medical practitioners looking to build their own medical LLMs tailored to specific needs, such as computational limits, private data, and local knowledge bases.
For the second question, we have conducted an extensive survey on the performances of existing medical LLMs across ten biomedical NLP (discriminative and generative) tasks. This analysis allows us to understand how medical LLMs outperform traditional medical AI models in different aspects. By showcasing their abilities, we aim to clarify the strengths that medical LLMs bring to the table when deployed in clinical settings.
The third question focuses on the practical application of medical LLMs in clinical settings. We provide guidelines for seven clinical application scenarios, offering specific implementations of medical LLMs and highlighting which abilities are used for each scenario.
The fourth question emphasizes the challenges that must be overcome when deploying medical LLMs in clinical practice. These challenges include hallucination (i.e. generation of coherent and contextually relevant but factually incorrect outputs) , explainability , ethical, legal, and safety concerns . We also advocate a broader evaluation of medical LLMs, such as trustworthiness for ensuring their responsible and effective use in clinical settings.
For the final question, we look into future directions for developing medical LLMs, providing a guide for researchers and practitioners looking to advance this field and maximize the potential of medical LLMs in medicine.
By addressing these questions and providing a holistic perspective on medical LLMs, we hope to foster deeper understanding, broader collaboration, and faster advancement in the field of medicine AI. In summary, this review makes several contributions:
We provide a comprehensive review of large language models in medicine and summarize their evaluations in ten biomedical downstream tasks.
We highlight the clinical applications of medical LLMs and offer practical guidelines for their deployment in various clinical settings.
We identify and discuss the challenges of applying medical LLMs in clinical practice, aiming to inspire further research and development in this area.
The overall structure of the review is as follows: Section 2 reviews existing research on LLMs and medical LLMs, emphasizing how to efficiently construct medical LLMs. Section 3 summarizes the performance of existing medical LLMs on representative biomedical AI tasks. Section 4 details how medical LLMs are applied in medicine. Section 5 delves into the challenge of existing medical LLMs. Section 6 introduces several potential opportunities to improve medical LLMs in terms of development and deployment. The conclusion of the paper is given in Section 7. Finally, Appendix A provides a detailed background, including technical details and model development, of the general large language models. Appendices B and C illustrate the detailed performances of LLMs on the downstream biomedical AI tasks.
The Principles of Medical Large Language Models
For clarity, this section focuses on summarizing the principles of medical large language models, leaving the detailed introduction to the background of general LLMs in Appendix A. Table 1 summarizes the currently available medical LLMs according to their model development. Existing medical LLMs are mainly pre-trained from scratch, fine-tuned from existing general LLMs, or directly obtained through prompting to align the general LLMs to the medical domain. Therefore, we introduce the principles of medical LLMs in terms of these three methods: pre-training, fine-tuning, and prompting. Meanwhile, we further summarize the medical LLMs according to their model architectures in Figure 2.
Pre-training typically involves training an LLM on a large corpus of medical texts, including both structured and unstructured text, to learn the rich medical knowledge. The corpus may include electronic health records (EHRs) , clinical notes , DNA sequence , and medical literature . In particular, PubMed , MIMIC-III clinical notes , and PubMed Central (PMC) literature , are three widely used medical corpora for medical LLM pre-training. A single corpus or a combination of corpora may be used for pre-training. For example, PubMedBERT and ClinicalBERT are pre-trained on PubMed and MIMIC-III, respectively. In contrast, BlueBERT combines both corpora for pre-training; BioBERT is pre-trained on both PubMed and PMC. The University of Florida (UF) Health EHRs are further introduced in pre-training GatorTron and GatorTronGPT . MEDITRON is pre-trained on Clinical Practice Guidelines (CPGs). The CPGs are used to guide healthcare practitioners and patients in making evidence-based decisions about diagnosis, treatment, and management.
To meet the needs of the medical domain, pre-training medical LLMs typically involve refining the following commonly used training objectives in general LLMs: masked language modeling, next sentence prediction, and next token predictionPlease refer to Section A.1.2 for an introduction of these three pre-training objectives.. For example. BERT-series models, e.g., BioBERT , PubMedBERT , ClinicalBERT , and GatorTron , mainly adopt the masked language modeling and the next sentence prediction for pre-training; GPT-series models, e.g., BioGPT , and GatorTronGPT , mainly adopt the next token prediction for pre-training.
After pre-training, medical LLMs are typically fine-tuned and then evaluated on different biomedical AI tasks to assess their capabilities for understanding and generating relevant text. In detail, once pre-trained, the LLM will have learned rich general language representations that can be leveraged for various downstream tasks. To achieve strong performance on these tasks, the LLM can be fine-tuned (i.e. further trained) on a small, task-specific dataset. This allows the model to adapt its general language representations to the specific requirements of the target task. The combination of large-scale pre-training and fine-tuning has been shown to be effective in achieving state-of-the-art performance.
Currently, the widely-used downstream tasks for evaluating the medical LLMs are the question answering (QA) and named entity extraction (NER) . The former task (i.e. QA) requires the model to generate responses/answers to questions using the medical knowledge it has learned, which is crucial for applications such as diagnostic support and medical research. The latter task (i.e. NER) involves identifying medical entities such as diseases, treatments, and medications from the text. Two widely used benchmarks for evaluating model performance on these tasks are Biomedical Language Understanding Evaluation (BLUE) , and Biomedical Language Understanding & Reasoning Benchmark (BLURB) . The BLUE benchmark includes ten public datasets for evaluating the performance on NER, relation extraction, document classification, sentence similarity, and inference. The BLURB is a more comprehensive benchmark than BLUE, consisting of thirteen datasets for the above tasks and an additional QA task.
2 Fine-tuning
It is high-cost and time-consuming to train a medical LLM from scratch, due to its requirement of substantial (e.g. several days or even weeks) computational power and manual labor. One solution is to fine-tune the general LLMs with medical data, and researchers have proposed different fine-tuning methods for learning domain-specific medical knowledge and obtain medical LLMs.
Current fine-tuning methods include Supervised Fine-Tuning (SFT), Instruction Fine-Tuning (IFT), and Parameter-Efficient Tuning. The resulting fine-tuned medical LLMs are summarized in Table 1.
aims to leverage high-quality medical corpus, which can be physician-patient conversations , medical question-answering , and knowledge graphs . The constructed SFT data serves as a continuation of the pre-training data to further pre-train the general LLMs with the same training objectives, e.g. next token prediction. SFT provides an additional pre-training phase that allows the general LLMs to learn rich medical knowledge and align with the medical domain, thus transforming them into specialized medical LLMs.
The diversity of SFT enables the development of diverse medical LLMs by training on different types of medical corpus. For example, DoctorGLM and ChatDoctor are obtained by fine-tuning the general LLMs ChatGLM and LLaMA on the physician-patient dialogue data, respectively. MedAlpaca based on the general LLM Alpaca is fine-tuned using over 160,000 medical QA pairs sourced from diverse medical corpora. Clinicalcamel combines physician-patient conversations, clinical literature, and medical QA pairs to refine the LLaMA-2 model . In particular, Qilin-Med and Zhongjing are obtained by incorporating the knowledge graph to perform fine-tuning on the Baichuan and LLaMA , respectively.
In summary, existing studies have demonstrated the efficacy of SFT in adapting general LLMs to the medical domain. They show that SFT improves not only the model’s capability for understanding and generating medical text, but also its ability to provide accurate clinical decision support .
constructs instruction-based training datasets , which typically comprise instruction-input-output triples, e.g. instruction-question-answer. The primary goal of IFT is to enhance the model’s ability to follow various human/task instructions, align their outputs with the medical domain, and thereby produce a specialized medical LLM.
Thus, the main difference between SFT and IFT is that the former focuses primarily on injecting medical knowledge into a general LLM through continued pre-training, thus improving its ability to understand the medical text and accurately predict the next token. In contrast, IFT aims to improve the model’s instruction following ability and adjust its outputs to match the given instructions, rather than accurately predicting the next token as in SFT . As a result, SFT emphasizes the quantity of training data, while IFT emphasizes their quality and diversity. Since IFT and SFT are both capable of improving model performance, there have been some recent works attempting to combine them for obtaining robust medical LLMs.
In other words, to enhance the performance of LLMs through IFT, it is essential to ensure that the training data for IFT are of high quality and encompass a wide range of medical instructions and medical scenarios. To this end, MedPaLM-2 invited qualified medical professionals to develop the instruction data for fine-tuning the general PaLM. BenTsao and ChatGLM-Med constructed the knowledge-based instruction data from the knowledge graph. Zhongjing further incorporated the multi-turn dialogue as the instruction data to perform IFT. MedAlpaca simultaneously incorporated the medical dialogues and medical QA pairs for instruction fine-tuning.
After instruction fine-tuning, the performance of the resulting medical LLMs (e.g., MedPaLM 2 and Clinical Camel ) is typically evaluated on multiple QA datasets (e.g., MedQA (USMLE) , MedMCQA , PubMedQA , and MMLU ). These studies demonstrated the ability of IFT to improve downstream performance.
aims to substantially reduce computational and memory requirements for fine-tuning general LLMs. The main idea is to keep most of the parameters in pre-trained LLMs unchanged, by fine-tuning only the smallest subset of parameters (or additional parameters) in these LLMs. Commonly used parameter-efficient tuning techniques include Low-Rank Adaptation (LoRA) , Prefix Tuning , and Adapter Tuning .
In contrast to fine-tuning full-rank weight matrices, 1) LoRA preserves the parameters of the original LLMs and only adds trainable low-rank matrices into the self-attention module of each Transformer layer . Therefore, LoRA can significantly reduce the number of trainable parameters and improve the efficiency of fine-tuning, while still enabling the fine-tuned LLM to capture effectively the characteristics of the downstream tasks. 2) Prefix Tuning takes a different approach from LoRA by adding a small set of continuous task-specific vectors (i.e. “prefixes”) to the input of each Transformer layer . These prefixes serve as the additional context to guide the generation of the model without changing the original pre-trained parameter weights. 3) Adapter Tuning involves introducing small neural network modules, known as adapters, into each Transformer layer of the pre-trained LLMs . These adapters are fine-tuned while keeping the original model parameters frozen , thus allowing for flexible and efficient fine-tuning. The number of trainable parameters introduced by adapters is relatively small, yet they enable the LLMs to adapt to downstream scenarios or tasks effectively.
In general, parameter-efficient tuning is valuable for developing LLMs that meet unique needs in specific (e.g. medical) domains, due to its ability to reduce computational demands while maintaining the model performance. For example, medical LLMs DoctorGLM , MedAlpaca , Baize-Healthcare , Zhongjing , CPLLM , and Clinical Camel adopted the LoRA to perform parameter-efficient fine-tuning to efficiently align the general LLMs to the medical domain.
3 Prompting
Fine-tuning considerably reduces computational costs compared to pre-training, but it requires further model training and collections of high-quality datasets for fine-tuning, thus still consuming some computational resources and manual labor. In contrast, the “prompting” methods efficiently align general LLMs (e.g. PaLM ) to the medical domain (e.g., MedPaLM ), without training any model parameters. Popular prompting methods include Zero/Few-shot prompting, Chain-of-Thought (CoT) prompting, Self-consistency prompting, and Prompt Tuning.
aims to directly give an instruction to prompt the LLM to efficiently perform a task following the given instruction. Zero-shot prompting does not provide an example. Few-shot prompting presents the LLMs with a small number of examples or task demonstrations before requiring them to perform a task. It allows the LLMs to learn from these examples or demonstrations to accurately perform the downstream task and follow the given examples to give corresponding answers . Therefore, few-shot prompting allows LLMs to accurately understand and respond to medical queries. For example, MedPaLM significantly improves the downstream performance by providing the general LLM, PaLM , with a small number of downstream examples such as medical QA pairs.
further improves the accuracy and logic of model output, compared with Zero/Few-shot Prompting. Specifically, through prompting words, CoT aims to prompt the model to generate intermediate steps or paths of reasoning when dealing with downstream (complex) problems . Moreover, CoT can be combined with few-shot prompting by giving reasoning examples, thus enabling medical LLMs to give reasoning processes when generating responses. For tasks involving complex reasoning, such as medical QA, CoT has been show to effectively improve model performance . Medical LLMs, such as DeID-GPT , MedPaLM , and MedPrompt , use CoT prompting to assist them in simulating a diagnostic thought process, thus providing more transparent and interpretable predictions or diagnoses. In particular, MedPrompt directly prompts a general LLM, GPT-4 , to outperform the fine-tuned medical LLMs on medical QA without training any model parameters.
is built on CoT to further enhance the robustness of the response . It encourages the model to perform multiple attempts at generating multiple answers to the same question, and then select the most consistent answer across different attempts, thus improving model performance even when CoT is ineffective. This can be particularly useful in the medical domain , where consistency in diagnosis or treatment recommendation is crucial.
aims to improve the downstream model performance by employing both prompting and fine-tuning techniques . The prompt tuning method introduces learnable prompts, i.e. trainable continuous vectors, which can be optimized or adjusted during the fine-tuning process to better adapt to different downstream scenarios and tasks. Therefore, they provide a more flexible way of prompting LLMs than the “prompting alone” methods that use discrete and fixed prompts, as described above. In contrast to traditional fine-tuning methods that train all model parameters, prompt tuning only tunes a very small set of parameters associated with the prompts themselves, instead of extensively training the model parameters. Thus, prompt tuning effectively and accurately responds to medical problems , with minimal incurring computational cost.
Existing medical LLMs that employ the prompting techniques are listed in Table 1. Recently, MedPaLM and MedPaLM-2 propose to combine all the above prompting methods, resulting in Instruction Prompt Tuning, to achieve strong performances on various medical question-answering datasets. In particular, using the MedQA dataset for the US Medical Licensing Examination (USMLE), MedPaLM-2 achieves a competitive overall accuracy of 86.5% compared to human experts (87.0%), surpassing previous state-of-the-art method MedPaLM by a large margin (19%).
Biomedical NLP Tasks
In this section, we will introduce two popular types of downstream tasks: generative and discriminative tasks, including ten representative downstream tasks that further build up clinical applications. We first describe the downstream tasks and their widely-used evaluation datasets and then discuss LLMs suitable for these tasks and compare their performance. Table 2 summarizes the details of widely used evaluation datasets for each downstream task. Figure 3 illustrates the performance comparisons between different LLMs. For clarity, we will only cover a general discussion of those downstream tasks. The detailed definition of the downstream task and the performance comparisons can be found in Appendices B and C.
Discriminative tasks are for categorizing or differentiating data into specific classes or categories based on given input data. They involve making distinctions between different types of data, often to categorize, classify, or extract relevant information from structured text or unstructured text. The representative tasks (see Table 2) include Question Answering, Entity Extraction, Relation Extraction, Text Classification, Natural Language Inference, Semantic Textual Similarity, and Information Retrieval.
The typical input for discriminative tasks can be medical questions, clinical notes, medical documents, research papers, and patient EHRs. The output can be labels, categories, extracted entities, relationships, or answers to specific questions, which are often structured and categorized information derived from the input text. In existing LLMs, the discriminative tasks are widely studied, and used to make predictions and extract information from input text.
For example, based on medical knowledge, medical literature, or patient EHRs, the medical question answering (QA) task can provide precise answers to clinical questions, such as symptoms, treatment options, and drug interactions. This can help clinicians with making more efficient and accurate diagnoses . Entity extraction can automatically identify and categorize critical information (i.e. entities) such as symptoms, medications, diseases, diagnoses, and lab results from patient EHRs, thus assisting in organizing and managing patient data . The following entity linking task aims to link the identified entities in a structured knowledge base or a standardized terminology system, e.g., SNOMED CT , UMLS , or ICD codes . This task is critical in clinical decision support or management systems, for better diagnosis, treatment planning, and patient care.
2 Generative Tasks
Different from discriminative tasks that focus on understanding and categorizing the input text, generative tasks require a model to accurately generate fluent and appropriate new text based on given inputs. These tasks include medical text summarization , medical text generation , and text simplification .
For medical text summarization, the input and output are typically long and detailed medical text (e.g. “Findings” in radiology reports), and a concise summarized text (e.g., the “Impression” in radiology reports). Such text contains important medical information that enables clinicians and patients to efficiently capture the key points without going through the entire text. It can also help medical professionals to draft clinical notes by summarizing patient information or medical histories.
In medical text generation, e.g. discharge instruction generation , the input can be medical conditions, symptoms, patient demographics, or even a set of medical notes or test results. The output can be a diagnosis recommendation of a medical condition, personalized instructional information, or health advice for the patient to manage their condition outside the hospital.
Medical text simplification aims to generate a simplified version of the complex medical text by, for example, clarifying and explaining medical terms. Text simplification can be easily confused with text summarization . While text summarization concentrates on giving shortened text while maintaining most of the original text meanings, text simplification focuses more on the readability part, hence there is no extractive approach for text simplification. In particular, complicated or opaque words will be replaced; complex syntactic structures will be improved; and rare concepts will be explained . It is possible that in extreme cases, text simplification may increase the length of a text for readability improvement. Thus, one example application is to generate easy-to-understand educational materials for patients from complex medical texts. It is useful for making medical information accessible and understandable to a general audience, without altering the essential meaning of the text.
3 Performance Comparisons
Figure 3 shows that some existing general LLMs (e.g. GPT-3.5-turbo and GPT-4 ) have achieved strong performance on existing downstream tasks. This is most noticeable for the QA task where GPT-4 (shown in the blue line in Figure 3) consistently outperforms existing task-specific fine-tuned models, and is even comparable to human experts (shown in purple line). The QA datasets of evaluation include MedQA (USMLE) , PubMedQA , and MedMCQA . To better understand the QA performance of existing medical LLMs, in Figure 4, we further demonstrate the QA performance of medical LLMs on the MedQA dataset over time in different model development types. It also clearly shows that current works, e.g., MedPrompt , have successfully proposed several prompting methods to enable LLMs to outperform human experts.
However, on the non-QA tasks, as shown in Figure 3, the existing general LLMs perform worse than the task-specific fine-tuned models. For example, on the entity extraction task using the NCBI disease dataset , the state-of-the-art task-specific fine-tuned model BioBERT achieves an F1 score of 89.36, substantially exceeding the F1 score of 56.73 by GPT-4. We hypothesize that the reason for the strong QA capability of the current general LLMs is that the QA task is close-ended; i.e. the correct answer is already provided in multiple candidates. In contrast, most non-QA tasks are open-ended where the model has to predict the correct answer from a large pool of possible candidates, or even without any candidates provided.
Overall, the comparison proves that the current general LLMs have a strong question answering capability, however, the capability on other tasks still needs to be improved. Therefore, we advocate that the evaluation of medical LLMs should be extended to a broad range of tasks including non-QA tasks, instead of being limited mainly to medical QA tasks. Hereafter, we will discuss specific clinical applications of LLMs, followed by their challenges and future directions.
Clinical Applications
This section discusses the clinical applications of LLMs, each subsection consisting of a specific application and a discussion on how LLMs perform this task.
Medical diagnosis involves the medical practitioner using objective medical data from tests and self-described subjective symptoms to conclude the most likely health problem occurring in the patient . It is important to diagnose a patient in an accurate and timely manner because the effectiveness of treatment for most diseases is extremely time-sensitive. A missed or wrong diagnosis often has negative consequences, ranging in severity from a minor inconvenience to death. For example, breast cancer has the highest mortality rate in the world because many communities lack trained personnel to perform proper checks in a timely manner. It may increase the accessibility of professional healthcare to incorporate LLMs into the medical diagnosis pipeline.
For example, a recent application of LLMs for medical diagnosis is through a graph model that returns the top knowledge paths (i.e., sub-graphs including the most relevant nodes and edges) regarding the pathology of diseases. Dr. Knows , a graph-based model that selects the top diagnosis cases with knowledge paths trained on a real-world hospital dataset. The knowledge paths come from the unified medical language system (UMLS) knowledge graph and can be used as diagnostic paths to support the diagnostic predictions made by the LLMs. The path encoder then generates a path representation, and the path ranker assesses the paths created for logical association with the input, generating a ranked list of probable disease diagnostics. Assessing the CUI-F score, a clinical metric that is a combination of CUI (concept identifier) recall and CUI precision, it was shown that when used in addition to existing models, this method improves the aforementioned CUI-F score by 8 percent to 18 percent depending on the base model chosen .
One distinct limitation of using LLMs as the sole tool for medical diagnosis is that it is heavily reliant on the subjective text inputs from the patient. Since LLMs are solely text-based, they lack the inherent capability to analyze medical diagnostic imagery. Given that objective medical diagnoses frequently depend on visual images, LLMs are often unable to directly conduct diagnostic assessments as they lack concrete, visual evidence to support disease diagnosis . However, they can help with diagnosis as a logical reasoning tool for improving accuracy in other vision-based models. One such example is ChatCAD , where images are first fed into an existing computer-aided diagnosis (CAD) model to obtain tensor outputs. These outputs are translated into natural language, which is subsequently fed into ChatCAD to summarize results and formulate diagnoses. ChatCAD achieves a recall score of 0.781, higher than that (0.382) of the state-of-the-art task-specific model R2GenCMN . We note that although the more recent GPT-4V(ision) is capable of interpreting images in the general domain, its extension to the medical image domain is yet to materialize. Nevertheless, all aforementioned methods of implementing LLMs rely on image transformation into text beforehand. There are also concerns over patient privacy, algorithm accountability, and the potential for bias , when using LLM for medicine.
2 Formatting and ICD-Coding
The international classification of diseases (ICD) is a method of standardizing diagnostic and procedural (e.g. surgeries) information of a clinical session. These ICD codes are recorded in the individual’s electronic health records (EHR) every time a patient visits a doctor. They are also used for tracking health metrics, outcomes of treatments, and billing. There is a strong need to automate the ICD labeling process, because it is time-intensive and currently often manually entered by doctors themselves.
LLMs can help to automate ICD coding by isolating medical terms from clinical notes and assigning corresponding ICD codes . For example, PLM-ICD is a multi-class classification model fine-tuned for automatic ICD coding. The base model used in PLM-ICD is domain-specific with medicine-specific knowledge to enhance the ability to understand medical terms. PLM-ICD uses segment pooling, the algorithm that divides long input texts into shorter representations using LLMs, when the input surpasses the maximum allowable length. Lastly, it relates the encoding to the augmented labels to output ICD codes for each clinical input. PLM-ICD produced a 92.6 macro AUC score and a 98.9 micro AUC scoreMacro AUC computes the AUC for each class separately and then averages them, giving equal weight to each class, while Micro AUC aggregates the contributions of all classes to compute the AUC, focusing on the overall performance across all samples. when implemented on the MIMIC-III full dataset , higher than scores of existing state-of-the-art models .
One important issue with LLMs is their potential biases and hallucinations. In particular, traditional multi-label classification models can easily constrain their outputs to a predefined list of (usually >1000) ICD candidate codes, through a classification neural network. In contrast, generative LLMs struggle to effectively handle this extensive list through prompting techniques. As a result, the LLM may assign an ICD code that is not in the candidate list, or even a non-existent ICD code, to the input medical text. It leads to confusion among healthcare professionals when interpreting medical records for diagnoses and medical procedures . It is therefore crucial to establish a proactive mechanism to detect and rectify errors before they enter patient Electronic Health Records (EHRs).
3 Clinical Report Generation
Clinical reports, such as radiology reports , discharge summaries , discharge instruction , and patient clinic letters , refer to standardized documentation that healthcare workers must complete after each patient visit . It is closely linked to medical diagnosis as a large portion of the report is often diagnostic results. It is typically tedious and time-consuming for potentially overworked clinicians to write clinical reports, and thus they are often incomplete or error-prone. Adopting LLMs for clinical report generation can provide an objective means of increasing their completeness and reducing clinical workload. The generated report is intended to be a document that the clinician can review, modify, and approve as necessary, rather than taking human “out of the loop” .
LLMs can be used intuitively as a summarization tool to help with clinical report generation. Given a diagnosis as input, an LLM can use its capabilities of text summarization (described in previous sections) to produce a clear and concise final conclusion. In this instance, LLMs do not directly contribute to improving the accuracy of the conclusion. Instead, they act as a tool for the tedious work that otherwise would have to be undertaken by doctors.
Another popular utilization of LLMs for generating clinical reports relies on a vision-based model or manual input from a doctor that acts as a precursor in the pipeline . This precursor step analyzes the input medical image, and then feeds the LLM with some type of annotation. The LLM uses the annotation alongside some other text prompt (e.g. report format, and input by medical personnel) to generate an accurate and fluent report that follows the requested format. This has been shown to greatly reduce the workload on doctors .
Recent existing medical LLMs for clinical report generation focus on ChatCAD , a scheme that combines the vision-based Computer-Aided Diagnosis (CAD) with the text-based LLMs. ChatCAD has been shown to improve the diagnostic performance F1 score of the state-of-the-art report generation methods R2GenCMN by 16.42 percent . In this scheme, the CAD first generates some rudimentary text-based prompts based on the input medical diagnostic images, which are subsequently fed into an LLM to further interpretation. The LLM then combines the inputs from CAD and other inputs (such as report format) to generate a formal report.
Although LLMs have generated more complete and accurate clinical reports than the human counterparts , they still have the issue of hallucinations, as well as a tendency to approach inputs with a literal view instead of an assumption-based perspective that is often taken by human doctors. There are also observations that human-written reports are generally more concise than reports generated by LLMs .
4 Medical Education
The importance of the healthcare profession requires no explanation , and training for specific roles in this field is critical. Medical education applies to both professionals and the general public, which is arguably equally as important . LLMs can be incorporated into the medical education system in different ways, including answering questions, helping students prepare for medical exams, and acting as a Socratic tutor .
Karabacak et al. have proposed several benefits of incorporating LLMs into the medical education system, specifically for preparing medical students for medical exams and subsequently scenarios in the real world. They suggest that medical education can be augmented by generating scenarios, problems, and corresponding answers by an LLM. Students will benefit from being exposed to a larger variety of problems than what is in their textbook. Since LLMs can generate novel content, it will ensure that students always have new problems to practice on . LLMs can also generate feedback on the students’ responses to practical problems, allowing students to know their areas of weakness in real time. Inherently, these will better prepare the medical students for the real world since they would have been exposed to more scenarios than before .
Another use of LLMs in the medical field is providing information to the public. Medical dialogues are often complex and difficult to understand for the average patient. LLMs can tune the textual output of prompts to use varying degrees of medical terminology for different audiences. This will make medical information easy to understand for the average person, while ensuring medical professionals have access to the most credible information .
Potential downsides of using LLMs in medical education include the current lack of ethical training, and biases in training datasets (e.g. under-representing some groups) . In addition, misinformation (e.g. hallucination) presents a challenge in utilizing LLMs for medical education.
5 Medical Robotics
Medical robots can be used in many facets of medicine, including surgery , transporting patients, assisting nurses , and medical rehabilitation . They can also address the shortage of medical staff and perform tasks beyond the human’s physical capabilities.
Robotics need environmental information to function. They acquire input data via sensors, analyze these data, plan routes, and execute the planned route to perform the required action. Route planning is a crucial stage among the above. Graph-based Robotic Instruction Decomposer utilizes LLMs in route planning, in which scene graphs (instead of image recognition) are used to receive environmental information and plan tasks in each execution stage for a given natural language instruction. It can also predict upcoming tasks and plan pre-defined robotic movements in the scene graph. LLMs will then take the instruction, scene graph, and robot graph as inputs to generate the planned route in text form as their output. This method has been shown to outperform GPT-4 by over 25.4% accuracy in simultaneously predicting correct action and object, and 43.6% accuracy in correctly predicting instruction tasks .
Using LLMs in medical robotics can also improve human-computer interactions. Robots with improved interactivity may recognize human emotions and requests through natural language inputs. This makes patient communication with robots less intimidating and more user-friendly than the current practice .
Some challenges with implementing medical robotics are similar to those with collaborative robots (cobots), as both cases involve robots operating alongside humans that require trusting robots to always do the right thing . Different from cobots, when implementing LLM into the algorithm for route planning and robotic motion, there is a concern with errors in judgment due to the effects of bias and hallucinations. This concern can be mitigated by the better ability of traditional robots to inhibit damage than cobots. The low prediction accuracy score (< 0.5) of current LLM-powered medical robotics renders them still unsuitable for implementation in real clinical settings , which also indicates research opportunities.
6 Medical Language Translation
There are two main areas of medical language translation. One is the translation of medical terminology from one language to another . The other is the translation of professional medical dialogue into expressions that are easy to understand by non-professional personnel . Both areas are important for making communication convenient, through different languages or between different groups of people.
The current language barrier to global collaboration in both research and medical techniques can be largely reduced with the help of LLMs. Machine translation has been proven seven percent more accurate than traditional human services . Language translation also improves accuracy in education resources and research articles, making knowledge accessible worldwide .
For the second area, LLMs can improve medical education by identifying the skill levels of specific students, and thus providing corresponding terminology and structures they will understand. This application will also help patients, especially the elderly and the less knowledgeable, to understand professional medical speech .
One ethical consideration of using LLM to perform medical translation is the potential for discriminatory verbiage to be inserted inadvertently into the output. Such verbiage is difficult to prevent due to the nature of the pipeline, and thus may cause miscommunications and even legal consequences. Additionally, potential misinformation caused by translation errors may confuse patients and, in the worst case, may cause them to take the wrong medical advice and execute it, inflicting harm to themselves .
7 Mental Health Support
Mental health support involves both diagnosis and treatment. Depression, a common mental health problem, is treated through a variety of therapies, including cognitive behavior therapy, interpersonal psychotherapy, psychodynamic therapy, etc. . Many of these techniques are primarily dominated by patient-doctor conversations, but psychological consulting and subsequent treatments can be cost-prohibitive for many. The ability of LLMs to serve as conversation partners and companions may significantly lower the barrier to entry for patients with financial or physical constraints , thus increasing the accessibility to mental health treatment resources . There have been various research works on the effects of incorporating LLMs or chatbots into the treatment plan .
The level of self-disclosure has a heavy impact on the effectiveness of mental health diagnosis and treatment. The more a patient is willing to share, the more accurate the diagnosis and, therefore, the more accurate the treatment plan. Studies have shown that patient willingness to discuss mental health-related topics with a robot is high . Alongside the convenience and lower financial stakes, mental health support by LLMs/chatbots has the potential to be more effective than human counterparts in many scenarios.
One major difficulty in employing LLMs for mental health support is the difference between written and spoken communications. Hill et al. found that people answered questions differently when asked to write the answer down, compared with when verbally expressing their answers. This may be a barrier that LLMs have to break in order to mimic a therapist to a high degree . Future studies could include longer-term studies to analyze how social penetration (i.e., the development of interpersonal relationships) over time affects information disclosure .
Challenges
Despite their potential for personalized medicine and improved patient care, LLMs face several challenges in their applications to medicine. The large scale of these models requires substantial computational resources, which can pose a limitation to their applications. LLMs are susceptible to “hallucination”, where they generate incorrect or misleading information . Issues surrounding patient privacy and data bias present significant hurdles that must be addressed to ensure the ethical and equitable use of LLMs in medicine .
Despite these challenges, the future of LLMs in medicine remains promising , with ongoing research and technological advances. In this section, we address these challenges and discuss solutions to the implementation of LLMs in an array of healthcare applications.
Hallucination of LLMs refers to the phenomenon where the generated output contains inaccurate or nonfactual information. It can be categorized into intrinsic and extrinsic hallucinations . Intrinsic hallucination generates outputs logically-contradicting factual information, such as wrong calculations of mathematical formulas . Extrinsic hallucination happens when the generated output cannot be verified, typical examples include LLMs ‘faking’ citations that do not exist or ‘dodging’ the question. When integrating LLMs into the medical domain, fluent but nonfactual LLM hallucinations can lead to the dissemination of incorrect medical information, causing misdiagnoses, inappropriate treatments, and harmful patient education. It is therefore vital to ensure the accuracy of LLM outputs in the medical domain.
Current solutions to mitigate LLM hallucination can be categorized into training-time correction, generation-time correction, and retrieval-augmented correction. The first (i.e. training-time correction) adjusts model parameter weights, thus reducing the probability of generating hallucinated outputs. Its examples include factually consistent reinforcement learning and contrastive learning. The second (i.e. generation-time correction) adds a ‘reasoning’ process to the LLM inference to ensure reliability, using drawing multiple samples or a confidence score to identify hallucination before the final generation. The third approach (i.e. retrieval-augmented correction) utilizes external resources to mitigate hallucination, for example, using factual documents as prompts or chain-of-retrieval prompting technique .
2 Lack of Evaluation Benchmarks and Metrics
Current benchmarks and metrics often fail to evaluate LLM’s overall capabilities, especially in the medical domain. For example, MedQA (USMLE) and MedMCQA offer extensive coverage on QA tasks but fail to evaluate important LLM-specific metrics, including trustworthiness, helpfulness, explainability, and faithfulness . It is therefore imperative to develop domain and LLM-specific benchmarks and metrics.
Singhal et al. proposed HealthSearchQA consisting of commonly searched health queries, offering a more human-aligned benchmark for evaluating LLM’s capabilities in the medical domain. Benchmarks such as TruthfulQA and HaluEval evaluate more LLM-specific metrics, such as truthfulness, but do not cover the medical domain. Future research is necessary to meet the need for more medical and LLM-specific benchmarks and metrics than what are currently available.
3 Domain Data Limitations
Current datasets in the medical domain (Table 1) remain relatively small compared to datasets for training general-purpose LLMs (Table 3). These limited small datasets only cover a small space of the vase domain of medical knowledge. This results in LLMs exhibiting extraordinary performance on open benchmarks with extensive data coverage, yet falling short on real-life tasks such as differential diagnosis and personalized treatment planning .
Although the volume of medical and health data is large, most require extensive ethical, legal, and privacy procedures to be accessed. In addition, these data are often unlabeled, and solutions to leverage these data, such as human labeling and unsupervised learning , face challenges due to the lack of human expert resources and small margins of error.
Current state-of-the-art approaches , , typically fine-tune the LLMs on smaller open-sourced datasets to improve their domain-specific performance. Another solution is to generate high-quality synthetic datasets using LLMs to broaden the knowledge coverage ; however, it has been discovered that training on generated datasets causes models to forget . Future research is needed to validate the effectiveness of using synthetic data for LLMs in the medical field.
4 New Knowledge Adaptation
LLMs are trained on extensive data to learn knowledge. Once trained, it is expensive and inefficient to inject new knowledge into an LLM through re-training. However, it is sometimes necessary to update the knowledge of the LLM, for example, on a new adverse effect of a medication or a novel disease. Two problems occur during such knowledge updates. The first problem is how to make LLMs ‘forget’ the old knowledge, as it is almost impossible to remove all ‘old knowledge’ from the training data, and the discrepancy between new and old knowledge can cause unintended association and bias . The second problem is the timeliness of the additional knowledge - how do we ensure the model is updated in real-time ? Both problems pose significant barriers to using LLMs in medical fields, where accurate and timely updates of medical knowledge are crucial in real-world implementations.
Current solutions to knowledge adaptation can be categorized into model editing and retrieval-augmented generation. Model editing alters the knowledge of the model by modifying its parameters. However, this method does not generalize well, with their effectiveness varying across different model architectures. In contrast, retrieval-augmented generation provides external knowledge sources as prompts during model inference; for example, Lewis et al. enabled model knowledge updates by updating the model’s external knowledge memory.
5 Behavior Alignment
Behavior alignment refers to the process of ensuring that the LLM’s behaviors align with the objectives of its task. Development efforts have been spent on aligning LLMs with general human behavior, but the behavior discrepancy between general humans and medical professionals remains challenging for adopting LLMs in the medical domain. For example, ChatGPT is well aligned with general human behavior, but their answers to medical consultations are not as concise and professional as those by human experts . In addition, misalignment in the medical domain introduces unnecessary harm and ethical concerns that lead to undesirable consequences.
Current solutions include instruction fine-tuning, reinforcement learning from human feedback (RLHF) , and prompt tuning . Instruction fine-tuning refers to improving the performance of LLMs on specific tasks based on explicit instructions. For example, Ouyang et al. used it to help LLMs generate less toxic and more suitable outputs. RLHF uses human feedback to evaluate and align the outputs of LLMs. It has been shown to be effective in multiple tasks, including becoming helpful chatbots and decision-making agents . Prompt tuning can also align LLMs to the expected output format. For example, Liu et al. uses a prompting strategy, chain of hindsight, to enable the model to detect and correct its errors, thus aligning the generated output with human expectations.
6 Ethical, Legal and Safety Concerns
Concerns have been raised regarding using LLMs (e.g. ChatGPT) in the medical domain , with a focus on ethics, accountability, and safety. For example, the scientific community has disapproved of using ChatGPT in writing biomedical research papers due to ethical concerns. The accountability of using LLMs as assistants to practice medicine is challenging . Li et al. and Shen et al. found that prompt injection can cause the LLM to leak personally identifiable information (PII), e.g. email addresses, from its training data, which is a significant vulnerability when implementing LLM in the medical domain.
With no immediate solutions available, we have nevertheless observed research efforts on understanding the cause of these ethical and legal concerns. For example, Wei et al. propose that PII leakage is attributed to the mismatched generalization between safety and capability objectives (i.e., the pre-training of LLMs utilizes a larger and more varied dataset compared to the dataset used for safety training, resulting in many of the model’s capabilities are not covered by safety training). Moreover, governments and large corporations are currently seeking to regularize and monitor the use of AI in various fields, including healthcare and medicine.
Future Directions
Although LLMs have already made a significant impact on people’s life through chatbots and search engines, their integration into medicine is still in the infant stage. Numerous new avenues of medical LLMs await researchers and practitioners to explore how to better serve the general public and patients. These avenues include introducing new benchmarks, establishing interdisciplinary collaborations, developing multimodal LLMs, and applying LLMs to less established medicine fields.
Recent studies have underscored the shortcomings of existing benchmarks in evaluating Large Language Models (LLMs) for clinical applications . Traditional benchmarks, which primarily gauge accuracy in medical question-answering, inadequately capture the full spectrum of clinical skills necessary for LLMs . Criticisms have been leveled against the use of human-centric standardized medical exams for LLM evaluation, arguing that passing these tests does not necessarily reflect an LLM’s proficiency in the nuanced expertise required in real-world clinical settings .
In response, there is an emerging consensus on the need for more comprehensive benchmarks. These should include capabilities like sourcing from authoritative medical references, adapting to the evolving landscape of medical knowledge, and clearly communicating uncertainties . Additionally, considering the sensitive nature of healthcare, these benchmarks should also assess factors such as fairness, ethics, and equity, which, though crucial, pose quantification challenges . The aim is to create benchmarks that more effectively mirror actual clinical scenarios, thus providing a more accurate measure of LLMs’ suitability for medical advisory roles.
2 Multimodal LLM Integrated with Time-Series, Visual, and Audio Data
Multimodal LLMs (MLLMs), or Large Multimodal Models (LMMs), are LLM-based models designed to perform multimodal (e.g. involving both visual and textual) tasks . While LLMs primarily address NLP tasks, MLLMs support a broader range of tasks, such as comprehending the underlying meaning of a meme and generating website codes from images . This versatility suggests promising applications of MLLMs in medicine. Several MLLM-based frameworks integrating vision and language, e.g. MedPaLM M , LLaVA-Med , Visual Med-Alpaca , Med-Flamingo , and Qilin-Med-VL , have been proposed to adopt the medical image-text pairs for fine-tuning, thus enabling the medical LLMs to efficiently understand the input medical (e.g. radiology) images.
A recent study proposes to integrate vision, audio, and language inputs for automated diagnosis in dentistry. However, there exist only very few medical LLMs that can process time series data, such as electrocardiograms (ECGs) and sphygmomanometers (PPGs) , despite such data being important for medical diagnosis and monitoring.
Similar to LLMs, MLLMs are associated with data privacy and quality challenges. Furthermore, the multimodal nature of MLLM introduces unique issues, including limited perception capabilities , fragile reasoning chains , sub-optimal instruction-following ability , and object hallucination . More research is needed to address these issues to ensure safe and effective applications of MLLM in medicine.
3 Medical Agents
LLM-based agents (i.e., intelligent systems or software designed to perform specific medical tasks or functions) have achieved significant progress in solving complex tasks (e.g. software design, molecular dynamics simulation) through human-like behaviors, such as role-playing and communication . However, integrating these agents effectively within the medical domain remains a challenge. The medical field involves numerous roles and decision-making processes, especially in disease diagnosis that often requires a series of investigations involving CT scans, ultrasounds, electrocardiograms, and blood tests. The idea of utilizing LLMs to model each of these roles, thereby creating collaborative medical agents, presents a promising direction. These agents could mimic the roles of radiologists, cardiologists, pathologists, etc., each specializing in interpreting specific types of medical data. For example, a radiologist agent could analyze CT scans, while a pathologist agent could focus on blood test results. The collaboration among these specialized agents could lead to a more holistic and accurate diagnosis. By leveraging the comprehensive knowledge base and contextual understanding capabilities of LLMs, these agents not only interpret individual medical reports, but also integrate these interpretations to form a cohesive medical opinion.
This multi-agent approach could significantly enhance diagnostic accuracy, reduce the time taken for diagnosis, and alleviate the workload on healthcare professionals. Furthermore, incorporating feedback loops within this system can enable continuous learning and improvement. As these Medical Agents interact with real-world medical data and cases, they can refine their decision-making algorithms and adapt to emerging medical trends and novel diseases. However, this approach also raises several challenges and considerations. First, ensuring the privacy and security of patient data is paramount, as these systems would handle sensitive medical information. Second, the reliability and accuracy of the interpretations by the agents need rigorous validation to meet medical standards. Lastly, the ethical implications of AI in healthcare, especially in decision-making roles, must be carefully examined. Overall, collaborative medical agents not only promise to improve healthcare delivery, but also open up new avenues for research and development in AI-assisted medical decision-making.
4 LLMs in Medical Sub-domains
Current LLM research in medicine has largely focused on general medicine, likely due to the greater availability of data in this area . This has resulted in the under-representation of LLM applications in specialized fields like ‘rehabilitation therapy’ and ‘sports medicine’. The latter, in particular, holds significant potential, given the global health challenges posed by physical inactivity. The World Health Organization identifies physical inactivity as a major risk factor for non-communicable diseases (NCDs), impacting over a quarter of the global adult population . Despite initiatives to incorporate physical activity (PA) into healthcare systems, implementation remains challenging, particularly in developing countries with limited PA education among healthcare providers . LLMs could play a pivotal role in these settings by disseminating accurate PA knowledge and aiding in the creation of personalized PA programs . Such applications could significantly enhance PA levels, improving global health outcomes, especially in resource-constrained environments.
5 Interdisciplinary Collaborations
Just as interdisciplinary collaborations are crucial in safety-critical areas like nuclear energy production, collaborations between the medical and technology communities for developing medical LLMs are essential to ensure AI safety and efficacy in medicine. The medical community has primarily LLMs provided by technology companies without questioning their data training, which is a sub-optimal situation. Medical professionals are therefore encouraged to actively participate in creating and deploying medical LLMs by providing relevant training data, defining the desired benefits of LLMs, and conducting tests in real-world scenarios to evaluate these benefits . Such assessments would help to determine the legal and medical risks associated with LLM use in medicine and inform strategies to mitigate LLM hallucination .
Conclusion
Large language models (LLMs) have made tremendous progress in natural language processing in recent years, opening up new opportunities for their application in medicine. This comprehensive review of existing medical LLMs provides details on their model architecture, parameter size, pre-training data, fine-tuning data, evaluation benchmarks, etc. We also summarize their performance across diverse biomedical NLP tasks.
Our analysis reveals that while LLMs have achieved promising results on benchmarks, significant gaps remain between benchmark performance and real-world clinical utility. We therefore further explore the potential of LLMs in various clinical applications such as diagnosis, clinical note generation, medical education, and other scenarios. We also discuss potential solutions to the challenges of applying LLM in medical applications, such as hallucination, lack of explainability, data shortage, and evaluation limitations.
As medical LLM applications are still in their infancy, to fully realize the benefits of LLMs in medicine, future research and development needs to focus on: i) developing new evaluation benchmarks with medical-specific metrics like trustworthiness, safety, fairness, etc; ii) strengthening interdisciplinary collaboration between medical and AI communities; iii) building multimodal LLMs to integrate time series, visual, and audio data; iv) and applying LLMs to more medical sub-domains.
In summary, this review provides a comprehensive overview of the principles, applications, and challenges of LLMs in medicine, intended to promote further research and exploration in this interdisciplinary field. With the rapid development of foundation models, the LLMs could significantly improve future clinical practice and medical discoveries for the benefit of society. It requires sustained interdisciplinary collaboration between clinicians and AI researchers, doctor-in-the-loop, and human-centered design, to address existing challenges with applying LLMs in the medical domain. The co-development of the collaboration for appropriate training data, benchmarks, metrics, and deployment strategies could enable faster and more responsible implementation of medical large language models than the current practice.
Inclusion and Ethics
This study does not require the recruitment of human participants. We used public data and literature for analysis.
Data availability
The data used in our work are publicly available https://github.com/AI-in-Health/MedLLMsPracticalGuide.
Code availability
The code for this paper is available at https://github.com/AI-in-Health/MedLLMsPracticalGuide.
References
Author Contributions
FL, ZL, JL, and DC conceived the project. HZ and FL conceived and designed the study. HZ, FL, BG, XZ, and JH conducted the literature review, performed data analysis, and drafted the manuscript. All authors contributed to the interpretation and final manuscript preparation. All authors read and approved the final manuscript.
Competing Interests
The authors declare no competing interests.
Appendix A Appendix: Background
In this section, we describe the background from the 1) Formulation of Large Language Model, 2) General Large Language Model. The details of existing LLMs are shown in Table 3.
The impressive performance of LLMs can be attributed to their Transformer architecture, large-scale pre-training, and scaling laws. Please refer to for details.
The language model Transformer is first used in machine translation and is then successfully applied to achieve state-of-the-art results in multiple NLP tasks. The natural strength of the Transformer lies in its fully-attentive mechanism, in which no recurrence is required. It is based solely on attention mechanisms and eliminates recurrence and convolutions entirely. This not only enables efficient understanding and modeling of long-text , but also allows highly paralleled training , thus reducing training and inference costs.
These characteristics make the Transformer highly scalable, and therefore it is relatively easy and efficient to obtain LLMs through large-scale pre-training strategy, such as the encoder-only LLM BERT , the encoder-decoder LLM T5 , and the decoder-only LLM GPT .
A.1.2 Large-scale Pre-training
The success of LLMs relies on the large-scale pre-training, during which an LLM is trained on massive corpora of open-domain unlabeled text (e.g., CommonCrawl, Wiki, and Books ) in an unsupervised or self-supervised learning manner. The common training objectives are masked language modeling, next sentence prediction, and next token prediction.
In masked language modeling, a portion of the input text is masked, and the model is tasked with predicting the masked words based on the remaining unmasked context. This encourages the model to learn contextual representations, capturing the semantic and syntactic relationships between words .
In next sentence prediction training, the model is given two sentences, and is tasked with predicting the second sentence that logically follows the first. The model learns the relationships between sentences and improves its ability to capture the overall coherence of the text .
Next token prediction is another common training objective, where the model is required to predict the next token (e.g., word) in a sequence given the previous tokens. In this way, it helps the model to understand the given context and reason for the next token, developing predictive capabilities . This training objective is often used in existing popular LLMs, e.g., GPT-series models , PaLM , LLaMA , and LLaMA-2 .
Large-scale pre-training allows LLMs to learn a wide range of linguistic knowledge and a broad understanding of language patterns and concepts. It enables LLMs to perform well on “zero-shot” and “few-shot” learning scenarios. In these scenarios, they can accurately perform the downstream tasks without human-labeled data for training. This is highly desirable in low-resource (e.g. marginalized language or medical) applications, where a large amount of labeled data for training is usually unavailable.
A.1.3 Scaling Laws
LLMs are essentially scaled-up versions of Transformer architecture with increased numbers of Transformer layers, model parameters, and the volume of pre-training data. The “scaling laws” predict how much improvement can be expected in a model’s performance as its size increases (in terms of parameters, layers, data, or the amount of training computed). The scaling laws proposed by OpenAI show that to achieve optimal model performance, the budget allocation for model size should be larger than the data. In contrast, the scaling laws proposed by Google DeepMind show that both model and data sizes should be increased in equal scales.
The scaling laws has been instrumental in developing LLMs, guiding researchers and practitioners to efficiently allocate resources and anticipate the benefits of scaling their models. Multiple LLMs have been proposed as a result, advancing the development of natural language understanding and generation.
A.2 General Large Language Models
In this section, we briefly introduce existing general LLMs . General LLMs can be divided into three categories based on their architecture (Table 3 ): encoder-only LLMs, encoder-decoder LLMs, and decoder-only LLMs. Please refer to for details.
Encoder-only LLMs, typically consisting of a stack of Transformer encoder layers, are designed to comprehend input sequences and produce dense context-aware representations, which aim to capture the semantic and syntactic properties of the input sequence. These models typically employ a bidirectional training strategy that allows them to integrate context from both the left and the right of a given token in the input sequence. This bi-directionality enables the models to achieve a deep understanding of the input sentences . Therefore, encoder-only LLMs are particularly suitable for language understanding tasks that require a comprehensive understanding of the input text, such as sentiment analysis , document classification , named entity recognition , and other tasks where the full context of the input is essential for accurate predictions. The strong performance of encoder-only LLMs in language understanding tasks has attracted significant research interest. A large number of encoder-only LLMs have been proposed, among which BERT is the representative one. The remaining includes DeBERTa, ALBERT, and RoBERTa. ELECTRA , and ERNIE .
Encoder-only LLMs represent a vital development in the field of natural language processing, with their bidirectional training and deep contextual understanding setting new benchmarks for a range of downstream tasks.
A.2.2 Decoder-only LLM
Decoder-only LLMs, which utilize a stack of Transformer decoder layers, are characterized by their uni-directional (left-to-right) processing of text, enabling them to generate language in a sequential manner. Unlike encoder-only LLMs, decoder-only LLMs are not designed for bidirectional context understanding, but are instead good at language generation tasks. During training, this architecture is trained unidirectionally using the next token prediction training objective to predict the next word in a sequence, given all the previous words. This process aligns naturally with language generation tasks such as text completion, storytelling , dialogue , and structured generation task - code generation . During inference, the decoder-only LLMs can directly generate sequences autoregressively (i.e. word-by-word). The prominent examples of decoder-only LLMs are the GPT (Generative Pre-Training Transformer) series developed by OpenAI and the LLaMA (Large Language Model Meta AI) series developed by Meta . Both have been employed successfully in language generation.
As a result of the open-source of LLaMA, a large number of improved LLM based on LLAMA have been proposed, e.g. Alpaca and Vicuna . The remaining popular decoder-only LLMs include PaLM , Bard , and GPT-4, as shown in Table 3.
A.2.3 Encoder-decoder LLM
Encoder-decoder LLMs are designed to simultaneously process input sequences and generate output sequences. They typically consist of a stack of bidirectional Transformer encoder layers followed by a stack of unidirectional Transformer decoder layers. The encoder processes and understands the input sequences, acquiring context-aware representations, while the decoder generates the output sequences based on the encoded representations . Combining the benefits of both the encoder and the decoder, these LLMs are suitable for tasks such as 1) machine translation, where the encoder processes the source language text, and the decoder generates the translation in the target language , 2) summarization , where the encoder reads the full-length document and the decoder produces a concise summary, and 3) even non-language tasks, such as protein structure prediction . Examples of encoder-decoder LLMs include Flan-T5 , and ChatGLM .
Appendix B Appendix: Discriminative Tasks
Question Answering (QA) aims to give answers to the given queries, by generating multiple-choice or free-text responses. A multiple-choice response happens when the model is asked a question with the material for answering the question included. For example, the QA for ’Is hypertension a risk factor for cardiovascular disease, yes or no?’ is multiple-choice-orientated. In contrast, an open question without potential answers provided gives rise to a free-text response, an example QA being ’What are the common symptoms of influenza?’.
Table 2 shows the commonly used biomedical QA datasets: MedQA (USMLE) , PubMedQA , and MedMCQA . Such datasets are relatively small in size compared to general datasets, thus the model performance can be readily improved by pre-training on general data and then finetuning on the biomedical data . Pergola et al. proposed a biomedical-specific masking method: instead of masking tokens randomly, the model identifies biomedical-related tokens and masks them to focus on in-domain learning. Wang et al. defined the self-questioning prompting (SQP) and utilize it on the BioASQ dataset. To extract useful information for specific tasks, SQP lets the GPT model ask questions about the given text and then answer these questions. However, hallucination remains a threat to the quality of outputs in QA. One solution is to use biomedical search systems such as Almanac . Another approach is to use extra datasets as an augmentation for the QA task. There are many chatbots based on LLM QA, such as Clinical Camel , DoctorGLM , ChatDoctor , HuaTuo , HuaTuoGPT , and MedAlpaca . Except for a few , most chatbots are a black box the the consumers, and thus providing insight into them is a potential direction for further studies.
Existing works use accuracy as the metric for model performance in QA. We compare the performance of task-specific (fine-tuned) BERT variations , Med-PaLM-2 , FLAN-PaLM , GPT-4 , GPT-3.5 , Clinical Camel , Galactica , BioMedLM , BioGPT , and PMC-LLaMA on the datasets described in Table 2. The citations in this paragraph correspond to the sources that provide data for assessing the model performance. Table 4 details their performances on five widely-used datasets, indicating that GPT-series models perform much better than BERT-series models. For GPT-series models, the observation that the few-shot setting outperforms the zero-shot setting still holds. Med-PaLM-2 has the highest accuracy among all models. These results indicate that it is feasible to fine-tune general LLMs on medical data and improve their performance.
B.2 Entity Extraction
Entity extraction, or named entity recognition (NER) aims to identify named entities mentioned in unstructured text, and classify them into predefined categories such as the names of persons, organizations, locations, medical codes, time expressions, quantities, monetary values, percentages, etc. Existing task-specific fine-tuned models leverage large-scale pre-trained language representations, and then fine-tune them on NER datasets to achieve state-of-the-art performances. The self-attention mechanisms are used for efficiently capturing complex patterns and dependencies.
In the healthcare domain, medical NER aims to extract medical entities such as disease names, medication, dosage, and procedures from clinical narratives, electronic health records (EHRs), and chemicals and proteins from scientific literature. This enables NER to provide a solid basis for clinical decision support systems, automated patient monitoring, etc.
Table 2 shows commonly used biomedical NER datasets, such as NCBI Disease , JNLPBA , and BC5CDR . To perform the NER task, for each input token, the existing models output a dense representation, which not only embeds the tokens but also includes its relation with other tokens in the text. Therefore, with the dense representations of input tokens extracted as vectors, additional layers are applied to the last Transformer layer to fit the downstream entity extraction task. The widely-used additional layers are softmax, BiLSTM , CRF , and their combinations . Gu et al. proposed a distillation method that uses few-shot GPT-3.5 to extract the correct entities and create the training set for the student model. This method shows that GPT models can be used to label the dataset first to achieve unsupervised learning. Wang et al. defined a new prompting method self-questioning prompting (SQP) for NER. Wang et al. further evaluated the performances of BARD, GPT-3.5, and GPT-4 with SQP and achieved some state-of-the-art results.
The entity-level F1 score is widely used to evaluate the models’ performance. Recent studies on the model performance of LLMs on biomedical NER tasks include GPT-3 , GPT-3.5 , GPT-4 , and ChatGPT . The citations in this paragraph correspond to the sources that provide data for assessing the model performance.
We summarize the performances of existing LLMs on the NER task in Table 5. It clearly shows that both the encoder-only BERT-series models (e.g. BioBERT , SciBERT , PubMedBERT , BlueBERT , and ClinicalBERT ) significantly outperform most decoder-only LLM GPT-series models (e.g., GPT-3 and GPT-4 ), under both the zero-shot and few-shot settings. Recent work has investigated the performance of GPT-3 , GPT-3.5 , and GPT-4 on biomedical NER tasks. These GPT models still need further development to outperform traditional task-specific fine-tuned models .
B.3 Relation Extraction
Different from entity extraction, relation extraction (RE) aims to find the relation between entities in a text. It is highly related to entity extraction since identifying entities first can improve the model’s performance. A relation can be ordered or unordered. For example, in the sentence ’I work in London’ the relation ’workspace’ is order-sensitive and the pair (I, London) is of such a relation, but in the sentence ’Alice and Bob are friends’ the relation ’friend’ is not order-sensitive. In the biomedical area, relation extraction is often used to extract relations in a MedAbstrct text for further applications like text summarization.
For relation extraction, Table 2 shows representative biomedical datasets such as BC5CDR, BioRED, and DDI. In practice, most datasets indicate the potential entity for the relation instead of asking the model to find it, so RE is usually considered as a classification task. One common technique for task-specific fine-tuned models is to use the classification token (e.g., [CLS] token in Transformer ), which learns the information of all input tokens, and is further used to classify the relation. Some studies use only the classification token and apply softmax in the last layer . Others use not only the classification token, but also all other token representations. It has been shown that using all token representations outperforms using classification token only . Specifically, Wang et al. defined the self-questioning prompting (SQP) as discussed in Sec B.1 and utilize it in the DDI RE task.
Existing studies on RE model performance typically use the F1-score, since the model performs a straightforward classification task. Table 6 compares the performance of BioBERT , SciBERT , PubMedBERT , BioGPT , GPT-3 , GPT-3.5 , GPT-4 , and ChatGPT on the datasets introduced above. It clearly shows that at present task-specific fine-tuned models outperform GPT models. For GPT models, the F1-score of the best GPT model (GPT-4) is below 70%, while that of all the task-specific fine-tuned models based on BERT is approximately 80%. There is no appreciable difference in F1-score between zero-shot and few-shot .
B.4 Text Classification
Unlike entity extraction or relation extraction at the token level, text classification is a sentence or document-level task. It is a classic classification problem of assigning predefined labels to a text, where it is common for a text to have multiple labels. The task is to predict all correct labels given the input medical text, and can be used as the preprocessing of medical text simplification and summarization.
Table 2 shows commonly used datasets (e.g. HoC and OHSUMED) for text classification. Since this is a text-level task, the [CLS] vector is augmented when using task-specific fine-tuned (Transformer-based) models. Another approach for distilling the overall information is to use the weighted sum of final attention layer outputs. Recent work has shown that adding a custom attention layer after the original model (BERT) improves the model performance . McCreery et al. double-fine-tuned the BERT model, in the sense that they first fine-tuned the model on a general dataset and then fine-tuned it on a medical-specific dataset. Another approach is to combine graph-based models with general LLMs, as in ChatGraph that used ChatGPT to extract text information and apply it to a graph-based model that outperforms GPT models. CohortGPT used Chain-of-Thought prompting and knowledge graph to outperform few-shot ChatGPT and GPT-4.
Same as above, the F1-score is used for evaluating the model performance. Table 7 compares the performance of PubMedBERT , BioGPT , GPT-4 , ChatGPT , and ChatGraph on four benchmark datasets. It shows that GPT models perform better on the i2b2 dataset than on the OHSUMED dataset, possibly because its few-shot setting is not yet sufficiently good. Overall, few-shot GPT models outperform zero-shot ones, and the F1 scores tend to increase with more parameters.
B.5 Natural Language Inference
Natural Language Inference (NLI) is a sentence-level task, involving two sentences: hypothesis () and premise (). The task to to determine whether can be inferred from . The outcome is one of the following three labels: (i) Entailment: the hypothesis can be inferred from the premise, or logically, . (ii) Contradiction: the negation of the hypothesis can be inferred from the premise, or logically, . (iii) Neutral: all other cases, or logically, .
The commonly used datasets, such as MedNLI and BioNLI , are shown in Table 2 . Due to the scarcity of biomedical datasets, researchers also use general datasets (e.g. SNLI and MultiNLI ). In task-specific fine-tuned models, similar to RE, [CLS] vector is often added to the text to show the overall information. The difference is that here we have two (instead of one) sentences, and so a [SEP] vector is also applied to indicate the separation of premise and hypothesis, yielding an overall input that looks like ‘[CLS], premise, [SEP], hypothesis’. For BERT-based models, a classification head is applied to the final [CLS] vector to predict the label. Kanakarajan et al. first pre-trained BioBERT on MIMIC-III and then fine-tuned the model on MedNLI . Cengiz et al. used the so-called two-stage sequential transfer learning method. They trained the BioBERT model on SNLI and MultiNLI first then fine-tuned it on the MedNLI dataset . The majority vote method is used to combine the prediction output of different trained models. Gu et al. combined GPT-3.5 and PubMedBERT using knowledge distillation. They first fed GPT-3.5 original texts and asked it to distill the data, and the distilled data was subsequently used to further train the PubMedBERT model. Wang et al. defined the self-questioning prompting (as discussed in Sec B.1) and utilize it in the MedNLI task.
The F1-score is used in the same manner as above. Table 8 compares the performance of task-specific fine-tuned models (e.g. BioBERT, ClinicalBERT, BlueBERT, and SciBER) on the benchmark MedNLI dataset , together with the performances of GPT-3.5 and GPT-3.5-Distillation. It shows that BioBERT has the best performance, and the current general LLMs still need further improvement on natural language inference.
B.6 Semantic Textual Similarity
Semantic Textual Similarity (STS) is similar to natural language inference (NLI). Their difference is that STS is a quantitative task, while NLI is a qualitative task. STS aims to produce a numerical value in a range (such as ) for any pair of sentences that indicates the degree of similarity, with the interpretation that means two sentences are completely independent and means two sentences are completely correlated or equivalent.
As shown in Table 2, the ‘2019 n2c2/OHNLP’ , BIOSSES , and MedSTS are widely-used benchmark datasets for STS. Since STS is a text-level task, [CLS] vector is naturally added to extract the overall information. Mutinda et al. added a fully connected layer after the [CLS] vector as the architecture. Yang et al. combined the representation of different models and added a fully connected layer. Wang et al. dropped the concept of [CLS] and used the Hierarchical Convolution (HConv) layer as the added last layer. Wang et al. defined the self-questioning prompting (SQP) and utilize it on the BIOSSES dataset. Xiong et al. used the concatenation of character level, sentence level, and entity level representation (CSE-concate) to extract the information that is fed to the further added MLP layer.
For further training on a pre-trained model, there is a common understanding that the new dataset should have a similar distribution to that of the dataset for pre-training to improve the model performance. The iterative training is introduced for this purpose , where the phrase "iterative" refers to the following two iterative steps: a) It first freezes the model and computes the outputs from the dataset, then it chooses a subset of the dataset so the outputs of the subset have a similar distribution to that of the pre-training dataset. b) The model is further trained using the obtained subset obtained in Step a).
The Pearson correlation coefficient is used for evaluation purposes. For pairs of data , we define Here, represents the actual similarity and represents the predicted similarity. Table 9 compares the performance of BERT , ClinicalBERT , XLNet , RoBERTa , HConv-BERT , BARD , GPT-3.5 , GPT-4 , BERT (CSE-concate) , and ClinicalBERT (iterative training) on BIOSSES and ‘2019 n2c2/OHNLP’ datasets.
It shows that the task-specific fine-tuned model RoBERTa achieves the best results on the ‘2019 n2c2/OHNLP’ dataset. Although GPT-4 has the highest Pearson correlation coefficient of 91.60%, we note that, with significantly fewer model parameters, the task-specific fine-tuned model ClinicalBERT achieves a competitive result compared to GPT-4. The overall results indicate that further work on improving the current general LLMs is needed.
B.7 Information Retrieval
Information retrieval (IR) plays an important role in the clinical area. It is the process of retrieving relevant knowledge or information related to the query from a number of unstructured data. It satisfies one’s need for searching information, such as article recommendations and literature searches. IR in the biomedical domain contains multiple tasks, among which text summarization and text simplification are discussed in Sec C.1 and Sec C.2, respectively. Question answering is also an important task included in IR, previously discussed in Sec B.1. In this subsection, we will concentrate on the query-article relevance task, also known as the ranking task. Specifically, for a query and a dataset of articles , we aim to find the most relevant articles where the degree of relevance is defined task-specifically.
Table 2 shows commonly used datasets, such as TREC-COVID , NFCorpus , and BioASQ in the biomedical area. In addition, a general dataset, BEIR (Benchmarking IR) , is often included in the training step of biomedical IR. Jin et al.
developed a BERT-based model BioCPT that encodes the query and articles for the ranking task. They split the method into training, inference, and evaluation steps. In the training step, they introduce query-to-document loss and document-to-query loss to train the encoders for the query and articles. In the inference step, they concatenate the encoding of the query and its best-fit article together with non-relevant articles found by maximum inner product search (MIPS) to further train the model to rank at the top. In the evaluation step, for each input query , the model evaluates over the whole dataset to find the best relevant articles by MIPS and used the model from the inference step to rank those articles.
Sun et al. applied zero-shot ChatGPT to rank the most relevant documents without abnormal prompting. They also further used GPT-4 to re-rank the top 30 documents retrieved by ChatGPT. Abonizio et al. introduced two LLM-based data augmentation methods, namely InPars and Promptagator, for IR. For InPars, they used GPT-3 and GPT-J to generate a new query for a randomly selected document. They used few-shot prompting which provides the model with good query examples or bad query examples. For Promptagator, the major difference is that a more dataset-specific prompting is applied. Ateia and Kruschwitz proposed a query expansion technique that expands the current query into a more comprehensive query, which consistently improves the performance of any successive tasks. It is done purely by the GPT model with regular instructional prompting. Similarly, Wang et al. used ChatGPT to generate more refined Boolean queries for systematic reviews. They showed that ChatGPT is able to generate or refine queries with higher precision.
The Normalized Discounted cumulative gain at (NDCG@) is used for evaluation. Based on the notation defined above, we further denote as the relevance of article to query . If we have the best ideal articles in order and articles retrieved by the model , we can define as: \text{NDCG@}k=\left({\sum_{i=1}^{k}\frac{\text{rel}(q,d_{i}^{\text{model}})}{\log(i+1)}}\right)\bigg{/}\left({\sum_{i=1}^{k}\frac{\text{rel}(q,d_{i}^{\text{ideal}})}{\log(i+1)}}\right)\,.
We compare the performance of Google’s Generalizable T5-based dense Retrievers (GTR) , OpenAI’s cpt-text , BioCPT , ChatGPT , and ChatGPT+GPT-4 on the datasets we introduced before.
Table 10 shows that ChatGPT and GPT-4 have the best performance. We observe that increasing parameters does not necessarily improve the model performance; for example, GTR-XXL has a lower NDCG@10 score than those of other GTR models, but performs better. We also note that the performance of all models appears unstable across different data; for example, ChatGPT+GPT-4 has only a 38.5% score for NFCorpus, but has a score of 85.5% for TREC-COVID.
Appendix C Appendix: Generative Tasks
There are two types of text summarization: extractive and abstractive summarizations. Extractive summarization aims to find the most important sentences in the text, while omitting the redundant or irrelevant sentences. In contrast, abstractive summarization generates brand-new text that summarizes the given text. Therefore, the extractive approach is more related to token-level tasks, while the abstractive approach is more relevant to high-level tasks (e.g. text generation).
Table 2 shows the commonly used biomedical datasets PubMed and MentSum . Some general datasets, such as GigaWord , CL-SciSumm , and S2ORC , are also commonly used for text summarization.
For extractive summarization, the key task is to define some scoring system that scores all sentences and hence finds the most important ones. Moradi et al. used a clustering-based method to summarize medical text. They vectorize the tokens of the text by BERT and cluster the sentence vector into clusters, and then they define an informativeness score that chooses one sentence from each cluster to form a summary. Moradi et al. used the graph-based model for summarization, where sentences are treated as nodes and relations as edges. The relations are measured by the cosine similarity of vectors representing the sentences. They then used different graph ranking algorithms to choose important sentences as a summary.
Du et al. used purely the Transformer-based model as the scoring system. They i) tokenized the whole text with [CLS] and [SEP] augmented, ii) further augmented the corresponding sentence and token positions to each token, and iii) fed the whole vector into the model. A sigmoid layer is added to the model so the output is between zero and one, and the output is considered the score for each sentence. Since the Transformer automatically extracts relations of one sentence to others, there is no need for manually designing a score. Chen et al. also relied only on the Transformer to score the sentences. Different from Du et al. , they proposed the AlphaBERT that takes the character-level tokens as input,and the training is split into pre-training and fine-tuning.
McInerney et al. combined the summarization model with a query to generate specific summaries. Pang et al. proposed a principled inference framework with top-down and bottom-up inference techniques to improve summarization models.
There are also studies on summarizing more than one piece of text at a time. They mainly relied on graph-based or Transformer-based models to extract relations between different texts . Zhang et al. applied one-shot GPT-3.5, using dialogues and summaries from the same category as prompts to generate abstractive summarization. More studies are needed to qualitatively analyze the performance of GPT models on biomedical text summarization .
The standard Recall-Oriented Understudy for Gisting Evaluation (ROUGE) data, including ROUGE-1 and ROUGE-2, is used for evaluation. Table 11 compares the performance of BioBERT variations using different training sets , rankings (PageRank and PPF) , and added last layers on PubMed. It shows that models fine-tuned on Pubmed or PMC data achieve improved performance. Differences in the ranking strategy do not correspond to appreciable differences in the model performance. Overall, current BERT models are far below the expectation of true text summarization, measured by the ROUGE-2 scores.
C.2 Text Simplification
Text simplification is to generate a new text that recovers almost all the information of the original text while improving its readability. The task has potential in biomedical education, since one major characteristic of medical education is its opaque vocabulary .
Table 2 shows the commonly used text simplification datasets. Patel et al. used NER to identify the medical terms. Lexical substitution is then applied to reduce the complexity of the text. Jeblick et al. tested the performance of ChatGPT on simplifying self-collected radiology reports. They tried different prompting texts and found that the prompt ‘Explain this medical report to a child using simple language’ performs the best. However, by evaluating from a) factual correctness, b) completeness, and c) harmfulness, they concluded that GPT models may generate harmful texts which is unacceptable in the medical domain. Joseph et al. evaluated the performance of zero-shot GPT-3 and Flan-T5 on MultiCochrane. They also fine-tuned mT5 and Flan-T5 on this dataset. Yang et al. introduced a data augmentation method for text simplification based on LLMs. For text without its simplified counterpart, they used GPT-3 to generate multiple choices for simplified text, and further trained a BERT model as the score to choose the best one as the simplification. Van et al. transferred the simplification task into a prediction task. They assume they have the original text and a unfinished simplified text . The task is to predict the next token . In particular, at each time, the model receive as input and predict . They used different models like BERT, RoBERTa, XLNet, and GPT-2. They also tried to combine the predicted token of all four models to outperform any of them.
For evaluation, we use the bilingual evaluation understudy (BLEU) score . We compare the performance of GPT-3 , Flan-T5 , and mT5 on MultiCochrane (English) (please see Table C.1). For AutoMeTS dataset , we use the accuracy of the next token as we discussed and we compare the performance of BERT, RoBERTa , XLNet , GPT-2 , and their combination (please see Table C.1). The citations in this paragraph correspond to the sources that provide data on the performance of the models. For MultiCochrane, T5 models have much better performance than other baselines. For AutoMeTS, a combination of different models actually outperforms any of them, but overall, none of those models have a high accuracy or BLEU score.
C.3 Text Generation
Text generation is a broad task that includes or relates to multiple specific tasks such as question answering (Sec B.1) and text summarization (Sec C.1). Here we focus on data-to-text tasks with open (instead of well-defined) answers, which involves taking structured data (e.g. a table) and producing text as output that describes these data. For example, generating patient clinic letters, radiology reports, and medical notes .
Yermakov et al. introduced a new dataset BioLeaflets and evaluated the performance of multiple LLMs on the data-to-text generation task. They found that T5 and BARD achieve the best performance for this task. However, multiple questions remain. Current LLMs may generate texts with typos, hallucinations, and repetitious words. Also, LLMs are not mature enough to produce coherent long text so far. Ranjit et al. studied the task of generating reports for chest X-ray images. They also investigated and discussed the hallucinations that occurred in the GPT-generated texts. There are concerns about using model-generated pseudo-text for attacking since humans without expert knowledge cannot easily see the factual errors in the generated text. Rodriguez et al. works on preventing attacks in the biomedical domain.