Evaluating Large Language Models for Radiology Natural Language Processing

Zhengliang Liu, Tianyang Zhong, Yiwei Li, Yutong Zhang, Yi Pan, Zihao Zhao, Peixin Dong, Chao Cao, Yuxiao Liu, Peng Shu, Yaonai Wei, Zihao Wu, Chong Ma, Jiaqi Wang, Sheng Wang, Mengyue Zhou, Zuowei Jiang, Chunlin Li, Jason Holmes, Shaochen Xu, Lu Zhang, Haixing Dai, Kai Zhang, Lin Zhao, Yuanhao Chen, Xu Liu, Peilong Wang, Junhao Chen, Pingkun Yan, Jun Liu, Bao Ge, Lichao Sun, Dajiang Zhu, Xiang Li, Wei Liu, Xiaoyan Cai, Xintao Hu, Xi Jiang, Shu Zhang, Xin Zhang, Tuo Zhang, Shijie Zhao, Quanzheng Li, Hongtu Zhu, Dinggang Shen, Tianming Liu

Introduction

In recent years, large language models (LLMs) have emerged as prominent tools in the realm of natural language processing (NLP). Compared with traditional NLP models, LLMs are trained on expansive datasets and demonstrate impressive capabilities ranging from language translation to creative content generation, and problem-solving. For example, OpenAI’s dialogue model, ChatGPThttps://openai.com/blog/chatgpt, has garnered widespread attention due to its outstanding performance, sparking a trend in the development of LLMs that has had profound effects on the growth of the entire AI community.

ChatGPT was developed based on GPT-3.5 version , released in November 2022. It is trained on a massive dataset of text and code and can generate text, translate languages, write different kinds of creative content, and answer your questions in an informative way. It is currently widely used in many areas, such as intelligent customer service, summary generation, and more, offering greater possibilities for applications of language models . The multimodal model GPT-4https://openai.com/research/gpt-4 with 1.8 trillion parameters was released in March 2023, with overall performance and accuracy surpassing the previous version. Also, the introduction of ChatGPT plugin functionality has directly endowed ChatGPT with the ability to use other tools and connect to the internet, breaking the constraints of the model’s data. The popular rise of ChatGPT has led to a surge in the development of new LLMs. There are now several hundred open-source LLMs available, such as Hugging Face’s BLOOM , which is trained with 176 billion parameters across 46 natural languages and 13 programming languages while most LLMs are not publicly released and are mainly based on Latin languages with English as the main language. Also, starting from Meta’s open source LLaMA series of models, researchers from Stanford University and other institutions have successively open-sourced LLaMA-based lightweight classes such as Alpacahttps://crfm.stanford.edu/2023/03/13/alpaca.html, Koalahttps://bair.berkeley.edu/blog/2023/04/03/koala/, and Vicunahttps://github.com/lm-sys/FastChat, etc. emerged. The research and application threshold of this type of model is greatly reduced, and the training, and reasoning costs have been repeatedly reduced.

Similarly, LLM ecology in other countries such as China has also begun to take shape. At present, the LLMs in China can basically be divided into three tracks: companies, institutions, and universities, such as Baidu’s ERNIE Bothttps://yiyan.baidu.com/, Huawei’s Pangu models series, IDEA’s Ziya-LLaMAhttps://github.com/IDEA-CCNL/Fengshenbang-LM series, Fudan University’s MOSShttps://github.com/OpenLMLab/MOSS. Many corresponding models are based on LLaMA, chatGLMhttps://github.com/THUDM/ChatGLM-6B, BLOOM, and other models that do not follow the transformer approach, such as baichuanhttps://modelscope.cn/models/baichuan-inc/Baichuan-13B-Base/summary, RWKV . These models are being used for various tasks, including natural language understanding, natural language generation, machine translation, and question answering.

The versatility of these models extends into the medical field . For example, LLMs can be used to generate personalized medical reports , facilitate online medical consultations, remote medical diagnosis, and guidance , and aid in medical data mining , among other applications. In radiology and medical image analysis, , which are fields that have been continuously intertwined with developments in AI , one significant application lies in the interpretation of images. Generative AI can help automate the process of preliminary diagnosis, potentially saving physicians’ time. It could be particularly useful in situations where there is a shortage of trained radiologists. Moreover, physicians may no longer need to manually enter data into the patient’s electronic medical record. In addition, these models are capable of aiding in clinical decision-making. By processing and analyzing a patient’s radiological data along with other relevant medical information, LLMs can generate patient-specific reports and provide possible diagnoses, treatment recommendations, or potential risks. Furthermore, advances in LLM have led to the development of many specialized biomedical LLMs, such as HuatuoGPT , an open-source model from the Chinese University of Hong Kong with the BLOOMZ as backbone, uses the data distilled from ChatGPT and the real data of doctors in the supervised fine-tuning stage. And XrayGLM,https://github.com/WangRongsheng/XrayGLM, the first Chinese multimodal large model dedicated to the diagnosis of chest X-rays, which is based on VisualGLM-6Bhttps://github.com/THUDM/VisualGLM-6B and then fine-tuned on two open chest X-rays datasets. Others include QiZhenGPThttps://github.com/CMKRG/QiZhenGPT, BioMedLMhttp://github.com/standford-crfm/BioMedLM.html, BioGPT , PMC-LLaMA , Med-PaLM , etc., which demonstrate significant potential of LLMs in the medical field.

Despite the escalating ubiquity of LLMs in various sectors, a comprehensive understanding and evaluation of their performance, particularly in the specialized field of radiology NLP, remains noticeably absent. This paucity of knowledge is even more stark when we consider the emerging LLMs developed in other countries such as China, a significant portion of which boast robust bilingual capabilities in both English and Chinese. Often untapped and under-evaluated, these models could offer unique advantages in processing and understanding multilingual medical data. The scarcity of in-depth, scientific performance evaluation studies on these models in the medical and radiology domains signals a significant knowledge gap that needs addressing. Given this backdrop, we believe it is paramount to undertake a rigorous and systematic exploration and analysis of these world-wide LLMs. This would not only provide a better understanding of their capabilities and limitations but also position them within the global landscape of LLMs. By comparing them with established international contenders, we aim to shed light on their relative strengths and weaknesses, providing a more nuanced understanding of the application of LLMs in the field of radiology. This, in turn, would potentially contribute to the optimization and development of more efficient and effective NLP/LLM tools for radiology.

Our study is focused on the crucial aspect of radiology NLP, namely, interpreting radiology reports and deriving impressions from radiologic findings. We rigorously evaluate the selected models using a robust dataset of radiology reports, benchmarking their performance against a variety of metrics.

Initial findings reveal that distinct differences were observed in the models’ respective strengths and weaknesses. The implications of these findings, as well as their potential impact on the application of LLMs in radiology NLP, are discussed in detail.

In the grand scheme, this study serves as a pivotal step towards the wider adoption and fine-tuning of LLMs in radiology NLP. Our observations and conclusions are aimed at spurring further research, as we firmly believe that these LLMs can be harnessed as invaluable tools for radiologists and the broader medical community.

Related work

The rapid development of LLMs has been revolutionizing the field of natural language processing and domains that benefit from NLP . These powerful models have shown significant performance in many NLP tasks, like natural language generation (NLG) and even Artificial General Intelligence (AGI). However, utilizing these models effectively and efficiently requires a practical understanding of their capabilities and limitations, and overall performance, so evaluating these models is of paramount importance.

To compare the capabilities of different LLMs, researchers usually test with benchmark datasets in various fields (such as literature, chemistry, biology, etc.), and then evaluate their performance according to traditional indicators (such as correct answer rate, recall rate, and F1 value). The most recent study from OpenAI includes the pioneering research study that assesses the performance of large language models (i.e. GPT-4) on academic and professional exams specifically crafted for educated individuals. The findings demonstrate exceptional performance of GPT-4 across a diverse array of subjects, encompassing the Uniform Bar Exam and GRE. Furthermore, an independent study conducted by Microsoft reveals that GPT-4 outperforms the USMLE, the comprehensive medical residents’ professional examination, by a significant margin . Holmes et al. explore the utilization of LLMs in addressing radiation oncology physics inquiries, offering insights into the scientific and medical realms. This research serves as a valuable benchmark for evaluating the performance in radiation oncology physics scenarios of LLMs.

Unlike the studies above that used traditional assessment research, Zhuang et al. introduced a novel cognitive science-based methodology for evaluating LLMs. Specifically, inspired by computerized adaptive testing (CAT) in psychometrics, they proposed an adaptive testing framework for evaluating LLMs that adjusts the characteristics of test items, such as difficulty level, based on the performance of individual models. They performed fine-grained diagnosis on the latest 6 instruction-tuned LLMs (i.e. ChatGPT (OpenAI), GPT-4 (OpenAI), Bard (Google), ERNIEBot (Baidu), QianWen (Alibaba), Spark (iFlytek)) and ranked them from three aspects of Subject Knowledge, Mathematical Reasoning, and Programming. The findings demonstrate a noteworthy superiority of GPT-4 over alternative models, achieving a cognitive proficiency level comparable to that of middle-level students. Similarly, traditional evaluation approaches are also not suitable for code generation tasks. Zheng et al. presented a multilingual model with 13 billion parameters for code generation, and building upon HumanEval (Python only), and they developed the HumanEval-X benchmark for evaluating multilingual models by hand-writing the solutions in C++, Java, JavaScript, and Go.

Nevertheless, a conspicuous dearth of assessment pertaining to substantial models in the realm of NLP within the domain of radiology persists. Consequently, this investigation endeavors to furnish an analytical appraisal of substantial models operating in the field of radiology. This inquiry represents the pioneering endeavor to encompass an exhaustive evaluation of large-scale language models (LLMs) within the purview of radiology, thereby serving as a catalyst for future investigations aimed at appraising the efficacy of LLMs within intricately specialized facets of medical practice.

2 Large Language Model Development in Other Countries

Many teams in other countries have conducted numerous attempts and research on LLMs. Examples include MOSS, developed by the OpenLMLab team at Fudan University, the BaiChuan series models developed by the Baidu team, the Sun and Moon model developed by SenseTime, PanGu developed by Huawei, ChatGLM-med developed by HIT, and YuLan-Chat developed by RUC. An introduction to these models will follow.

MOSS is an open-source series of large language models in Chinese and English, consisting of multiple versions. The base model, moss-moon-003-base, is the foundation of all MOSS versions and is pre-trained on high-quality Chinese and English corpora, encompassing 700 billion words. To adapt MOSS to dialogue scenarios, the model moss-moon-003-sft is fine-tuned on over 1.1 million rounds of dialogue data using the base model. It possesses the capabilities of instruction-following, multi-turn dialogue comprehension, and avoidance of harmful requests. Furthermore, in addition to providing consultation and question-answering functionalities, the moss-moon-003-sft-plugin model is trained not only on over 1.1 million dialogue data but also on more than 300,000 enhanced dialogue data with plugins. It extends the capabilities of the moss-moon-003-sft model to include plugin functionalities such as using search engines, generating images from text, performing calculations, and solving equations. Both moss-moon-003-sft and moss-moon-003-sft-plugin have versions with 4-bit quantization and 8-bit quantization, making them suitable for lower resource environments.

BaiChuan, an open-source large-scale pretraining language model developed by BaiChuan Intelligence, is based on the Transformers architecture. It has been trained on 1.2 trillion tokens and contains 7 billion parameters, supporting both Chinese and English languages. During training, a context window length of 4096 tokens was used. In actual testing, the model can also scale to over 5000 tokens. It achieved the best performance among models of the same size on standard Chinese and English language benchmarks (C-Eval /MMLU ).

Chat-GLM-6B is an open bilingual language model based on the General Language Model(GLM) framework, with 6.2 billion parameters. The model is trained for about 1 trillion tokens of Chinese and English corpus, supplemented by supervised fine-tuning, feedback bootstrap, and reinforcement learning with human feedback. It is optimized for Chinese QA and dialogue. Furthermore, the model is capable of generating answers that align with human preference.Chat-GLM-med is a variation of Chat-GLM-6B that is fine-tuned specifically for Chinese medical instructions. This Chinese medical instruction dataset is constructed using a medical knowledge graph and the GPT3.5 API. The fine-tuning process was performed on top of the existing Chat-GLM-6B model.

YuLan is developed by the GSAI team at the Renmin University of China. It utilizes the LLaMA base model and is fine-tuned on a high-quality dataset of Chinese and English instructions. The dataset construction involves three stages: Open-source Instruction Deduplication, Instruction Diversification based on Topic Control, and Instruction Complexification. These stages aim to enhance the diversity of the instruction learning dataset.

There are various types of LLMs available at the current stage, and most of them are based on a base model which is a pre-trained language model following the Transformer architecture. These models are further fine-tuned using domain-specific or high-quality data constructed on top of the base model. By leveraging high-quality, domain-specific data, a model with specialized knowledge can be derived from the base model.

3 Applications in the Medical Field

The development of LLMs could result in many potential applications in the medical field. In general, they could be used in the following four areas: clinical documentation, clinical decision support, knowledge-based medical information retrieval and generation, and medical research.

There is a large amount of clinical writing to be done by physicians and clinical professionals every day. Some can be quite laborious and time-consuming. LLMs can be possibly applied to assist in documenting patient information and symptoms , generating accurate and comprehensive clinical notes and test reports, and thus effectively reducing the writing load of physicians and clinical professionals. For example, an LLM summarizing the Impression from the radiology report was developed, presenting a paradigm in applications in similar domains .

LLMs can also be applied to provide clinical decision support through recommending medicine usage , identifying appropriate imaging services from clinical presentations , or determining the cause of disease from numerous clinical notes and reports. When integrated with other modalities, like imaging, it can generate comprehensive information, assist physicians in disease diagnosis . In addition, from cases of patients with similar symptoms, LLMs can generate patient disease outcomes, giving a prediction of what the treatment may look like, and supporting the physicians and patients make decisions on treatment options.

Knowledge-based application is another place where LLMs can play a big role. For example, it could be useful to have an LLM application developed to answer health-related questions from patients . With the training of data in specific domains, LLM applications could provide physicians and health professionals with relevant medical information from a vast amount of scientific literature, research papers, and clinical guideline, enabling quick access to up-to-date information on disease, treatments, drug interactions, and more . Knowledge-based LLMs can help educate medical trainees and patients by answering generic or specific questions . By integrating with a patient’s medical record, LLMs can provide personalized information and explanation of drug usage, ongoing treatment, or any relevant questions patients may have.

LLMs can be of great benefit to the medical research community and public health . For instance, the privacy of patient medical records is a big concern in clinics. The removal of identification information is mandatory before medical records used for research and results are released to the public. An LLM application help remove the identification information from medical records and could be widely utilized and beneficial to medical research . Training of clinical NLP models may suffer from a lack of medical text data; augmentation of medical text data by LLMs could provide additional samples profiting the NLP model training . Moreover, LLMs could conduct data collection, processing, and analysis about specific diseases, providing quantified metrics and valuable insight to researchers .

Methodology

This section will discuss our testing methods for LLMs. We will begin by introducing the datasets MIMIC and OpenI, which we use for evaluation. Our testing approach involves employing a fixed set of prompts and parameters to assess the performance of LLMs in the field of radiology, specifically focusing on deriving impression-based performance from findings. To ensure consistency, we set several hyperparameters of the LLMs, namely the temperature to 0.9, the top_k to 40, and the top_p to 0.9. To evaluate the model’s zero-shot and few-shot performance, we utilize zero-shot, one-shot, and five-shot examples as prompts. The experimental results and their detailed analysis are presented in the results section.

Our testing approach involves utilizing a fixed set of prompts and parameters to evaluate the LLMs. The model’s inference parameters, namely the temperature, top_k, and top_p, are fixed at 0.9, 40, and 0.9, respectively, to ensure consistency. We engage zero-shot, one-shot, and five-shot prompts to examine the model’s zero-shot and few-shot performance. A zero-shot prompt involves presenting the model with a new task, with no prior examples provided. A one-shot prompt involves providing the model with one prior example, while a five-shot prompt provides the model with five prior examples. This variation in prompts offers a nuanced understanding of how the LLMs operate under different conditions and degrees of prior exposure.

2 Model Selection

Considering both resource constraints and the need for uniformity in model comparison, our evaluation specifically focuses on Large Language Models (LLMs) with approximately 7 billion parameters. The choice of this parameter count is based on two primary considerations. First, models of this size strike a balance between computational efficiency and model performance. They allow for faster inference, making it feasible to thoroughly evaluate the models over the complete testing dataset in a practical timeframe. Second, this parameter count is well-represented across different types of LLMs, allowing for a broad and diverse range of models to be included in the study.

For open-source models, we procure the necessary code and model parameters directly from their official GitHub repositories. These repositories provide comprehensive documentation and community support, ensuring that the models are implemented and evaluated correctly.

For commercially available models, such as Sensenova, ChatGPT, GPT-4, PaLM2, and Anthropic Claude2, we utilize their respective Application Programming Interfaces (APIs). These APIs offer a structured and standardized way of interacting with the models, enabling us to input our pre-determined prompts and parameters and receive the model outputs in a consistent and reliable manner.

HuatuoGPT is a language model developed by the Shenzhen Research Institute of Big Data from the Chinese University of Hong Kong, Shenzhen. HuatuoGPT-7B is trained on the Baichuan-7B corpus, while HuatuoGPT-13B is based on Ziya-LLaMA-13B-Pretrain-v1. The advantage of HuatuoGPT is in its integration of real-world medical data and the information-rich base of ChatGPT. This allows HuatuoGPT to provide detailed diagnoses and advice in medical consultation scenarios, similar to a doctor’s approach . HuatuoGPT has two versions: HuatuoGPT-7B and HuatuoGPT-13B. In our experiments, we used the HuatuoGPT-7B version.

Luotuo is a Chinese language model exploited and maintained by the researchers Qiyuan Chen, Lulu Li, and Zihang Leng. Luotuo is fine-tuned by the LLaMA on Chinese corpus utilizing LoRA technique and does well in Chinese infering . Luotuo has three versions: Luotuo-lora-7b-0.1, Luotuo-lora-7b-0.3, and luotuo-lora-7b-0.9. Luotuo-lora-7b-0.3 was used in the experiments.

Ziya-LLaMA denotes bilingual pre-trained language models based on LLaMA. It is a member of the open-source general large model series and is introduced by the Center for Cognitive Computing and Natural Language Research (CCNL) at the IDEA Research Institute. Ziya-LLaMA boasts remarkable versatility, demonstrating proficiency across a wide array of tasks including translation, programming, text classification, information extraction, summarization, copywriting, common sense Q&A, and mathematical calculation. Its comprehensive training process comprises three stages: large-scale continual pre-training, multi-task supervised fine-tuning, and human feedback learning. Ziya has four version, Ziya-LLaMA-13B-v1.1, Ziya-LLaMA-13B-v1, Ziya-LLaMA-7B-Reward, and Ziya-LLaMA-13B-Pretrain-v1. In this study, we investigated the Ziya-LLaMA-13B-v1.

YuYan-Dialogue YuYan-Dialogue is a Chinese language dialogue model by fine-tuning the YuYan-11b on a large multi-turn dialogue dataset of high quality and developed by Fuxi AI lab, Netease.Inc. It is trained on a large Chinese novel dataset of high quality and has very strong conversation generation capabilities. YuYan-Dialogue has only one version that is YuYan-Dialogue. Therefore, we used it in our experiments.

BenTsao BenTsao is a medical language model based on LLaMA-7B model developed by SCIR Lab in Harbin Institution of Technology. It has undergone Chinese medical instruction fine-tuning and instruction tuning. They built a Chinese medical instruction dataset through the medical Knowledge graph and GPT3.5 API, based on which, they further fine-tuned the model, improving the question-and-answer effect of LLaMA in the medical field. BenTsao has four versions, LLaMA-med, LLaMA-literature, Alpaca-med, Alpaca-all-data. Here, we used the LLaMA-med (BenTsao) for comparison.

XrayGLM Xray-GLM is a vision-language model developed by Macao Polytechnic University. It is based on the VisualGLM-6B and fintuned on the translated Chinese version MIMIC-CXR, OpenI dataset. It has strong ability on chest Xray VQA. Here, we used the newest version of the Xray-GLM for comparison.

ChatGLM-Med ChatGLM-Med is a language model developed by SCIR Lab in Harbin Institution of Technology. It is based on the ChatGLM-6b and has undergone Chinese medical instruction fine-tuning and instruction tuning. They built a Chinese medical instruction dataset through the medical Knowledge graph and GPT3.5 API, and on this basis, and fine-tuned the model based on the instructions of ChatGLM-6B, improving the question-and-answer effect of ChatGLM in the medical field. Here, we chose the newest version of ChatGLM-Med model for comparison.

ChatGPT/GPT4 ChatGPT and GPT4 are both highly influential large language models developed by OpenAI. The full name of ChatGPT is gpt-3.5-turbo, which is developed on the basis of gpt2 and gpt3.The training process of ChatGPT mainly refers to instructGPT , ChatGPT is an improved instructionGPT. The main difference from GPT-3 . is that the new addition is called RLHF (Reinforcement Learning from Human Feedback, human feedback reinforcement learning) . This training paradigm enhances human conditioning of the model output and enables a more comprehensible ranking of the results. ChatGPT has strong language understanding ability and can handle various language expressions and queries. ChatGPT has an extensive knowledge base that can answer various frequently asked questions and provide useful information. GPT-4 is a successor to GPT-3, so it may be more capable in some ways. In our experiments, we used the ChatGPT and GPT4.

ChatGLM2/ChatGLM ChatGLM2 is a large language model developed by Tsinghua University, developed on the basis of the ChatGLM using the GLM framework . ChatGLM2 has more powerful performance, which can handle longer contexts and perform more efficient reasoning with a more open protocol. What’s more, ChatGLM2 is an excellent bilingual pre-trained model . There are many versions of ChatGLM2 depending on the size of the pattern instruction set. This work mainly tested ChatGLM2-6B and ChatGLM-6B.

QiZhenGPT QiZhenGPT is a model developed by Zhejiang University. It uses the Chinese medical instruction data set constructed by QiZhen Medical Knowledge Base, and based on this, performs instruction fine-tuning on the Chinese-LLaMA-Plus-7B, CaMA-13B, and ChatGLM-6B models. QiZhenGPT has an excellent effect in Chinese medical scenarios, and it is more accurate in answering questions than ChatGLM-6B. According to different model objects fine-tuned by instructions, QizhenGPT has three types : QiZhen-Chinese-LLaMA-7B, QiZhen-ChatGLM-6B, and QiZhen-CaMA-13B. In this work, we tested mainly on QiZhen-Chinese-LLaMA-7B.

MOSS MOSS-MOON-003 is the third version of the open-sourced plugin-augmented bilingual (i.e. Chinese and English) conversational language model MOSS, specifically from the MOSS-MOON-001 to MOSS-MOON-003, developed by the OpenLMLab from Fudan University . The MOSS-MOON-003-sft is fine-tuned with supervision on approximately 1.1M multi-turn conversational data to the base model, MOSS-MOON-003-base. The advantage of MOSS-MOON-003 is it can follow bilingual multi-turn dialogues, refuse inappropriate requests and utilize different plugins due to its base model (i.e. MOSS-MOON-003-base was pre-trained on 700B English, Chinese, and code tokens), fine-tuning on multi-turn plugin-augmented conversational data, and further preference-aware training. There are 10 versions available: MOSS-MOON-003-base, MOSS-MOON-003-sft, MOSS-MOON-003-sft-plugin, MOSS-MOON-003-sft-int4, MOSS-MOON-003-sft-int8, MOSS-MOON-003-sft-plugin-int4, MOSS-MOON-003-sft-plugin-int8, MOSS-MOON-003-pm, MOSS-MOON-003, and MOSS-MOON-003-plugin. In our experiments, we used the MOSS-MOON-003-sft version.

ChatFlow ChatFlow is a fully-parameterized training model developed by the Linly project team, built upon the foundations of LLaMa and Falcon and based on the TencentPretrain pre-training framework and a large-scale Chinese scientific literature dataset . By utilizing both Chinese and Chinese-English parallel incremental pre-training, it transfers its language capabilities from English to Chinese. The key advantage of ChatFLow is that it addresses the issue of weaker Chinese language understanding and generation abilities found in the open-source models Falcon and LLaMa. It significantly improves the encoding and generation efficiency of Chinese texts. ChatFlow comes in two versions, namely ChatFlow-7B and ChatFlow-13B. For our experiments, we utilized the ChatFlow-7B version.

CPM-Bee CPM-Bee is a large model system ecology based on OpenBMB, and it is a self-developed model of the Facing Wall team. It is a completely open source, commercially available Chinese-English bilingual basic model, and it is also the second milestone achieved through the CPM-Live training process. CPM-Bee uses the Transformer autoregressive architecture, with a parameter capacity of tens of billions, pre-training on a massive corpus of trillions of tokens, and has excellent basic capabilities. There are four versions of CPM-Bee: CPM-Bee-1B, CPM-Bee-2B, CPM-Bee-5B, CPM-Bee-10B. In this experiment, we tested the performance of CPM-Bee-5B (CPM-Bee).

PULSE The PULSE model is a large-scale language model developed on the OpenMEDLab platform. It is based on the OpenChina LLaMA 13B model, which is further fine-tuned using approximately 4,000,000 SFT data from the medical and general domains. PULSE supports a variety of natural language processing tasks in the medical field, including health education, physician exam questions, report interpretation, medical record structuring, and simulated diagnosis and treatment. PULSE has two versions, PULSE_7b and PULSE_14b. In this experiment, we tested the version of PULSE_7b.

Baichuan Baichuan, developed by Baichuan Intelligence, is a large pre-trained model based on the Transformer architecture. The baichuan-7B model, comprising 7 billion parameters, was trained on approximately 12 trillion tokens, utilizing the same model design as LLaMa. Subsequently, they further developed the baichuan-13B model, which is even larger in size and trained on a greater amount of data . The key advantage of the Baichuan model lies in its use of an automated learning-based data weighting strategy to adjust the data distribution during training, resulting in a language model that supports both Chinese and English. It has demonstrated robust language capabilities and logical reasoning skills across various datasets. Two versions of the Baichuan model are developed: baichuan-7B and baichuan-13B. For our experiments, we utilized the baichuan-7B version.

AtomGPT AtomGPT , developed by Atom Echo, is a large language model based on the model architecture of LLaMA . AtomGPT uses a large amount of Chinese and English data and codes for training, including a large number of public and non-public data sets. Developers use this method to improve model performance. AtomGPT currently has four versions: AtomGPT_8k, AtomGPT_14k, AtomGPT_28k, AtomGPT_56k. In this experiment, we chose AtomGPT_8k for testing.

ChatYuan ChatYuan large v2 is an open-source large language model for dialogue, supports both Chinese and English languages, and in ChatGPT style. It is published by ClueAI. ChatYuan large v2 can achieve high-quality results on simple devices that allows users to operate on consumer graphics cards, PCs, and even cell phones. It got optimized for fine-tuning data, human feedback reinforcement learning, and thought chain. Also, comparing with its previous version, the model is optimized in many language abilities, like better at both Chinese and English, generating codes and so on. ChatYuan has three versions: ChatYuan-7B, ChatYuan-large-v1, ChatYuan-large-v2. In our experiments, we tested the ChatYuan-large-v2.

Bianque-2.0 Bianque is a large model of healthcare conversations fine-tuned by a combination of directives and multiple rounds of questioning conversations. Based on BianQueCorpus, South China University of Technology chose ChatGLM-6B as the initialization model and obtained BianQue after the instruction fine-tuning training. BianQue-2.0 expands the data such as drug instruction instruction, medical encyclopedic knowledge instruction, and ChatGPT distillation instruction, which strengthens the model’s suggestion and knowledge query ability. By using Chain of Questioning, the model can relate more closely to life and to improve questioning skills, which is different from most language model. It has two versions: Bianque-1.0 and Bianque-2.0. In our experiments, we tested Bianque-2.0.

AquilaChat AquilaChat is a language model developed by the Beijing Academy of Artificial Intelligence. AquilaChat is an SFT model based on Aquila for fine tuning and Reinforcement learning. The AquilaChat dialogue model supports smooth text dialogue and multiple language class generation tasks. By defining extensible special instruction specifications, AquilaChat can call other models and tools, and is easy to expand its functions. AquilaChat has two versions: AquilaChat-7B and AquilaChat-33B. In our experiments, we used the AquilaChat-7B version.

Aquila Aquila is a language model developed by the Beijing Academy of Artificial Intelligence. Aquila-7B is a basic model with 7 billion parameters. The Aquila basic model inherits the architectural design advantages of GPT-3, LLaMA, etc. in terms of technology, replaces a batch of more efficient low-level operator implementations, redesigns and implements the Chinese English bilingual tokenizer, upgrades the BMTrain parallel training method, and achieves nearly 8 times the training efficiency compared to Magtron+DeepSpeed Zero-2. Aquila has two versions: Aquila-7B and Aquila-33B. In our experiments, we used the Aquila -7B version.

Chinese-Alpaca-Plus Chinese-Alpaca-Plus is a language model developed by Yiming Cui etc. Chinese-Alpaca-Plus is a language model based on LLaMA. Chinese-Alpaca-Plus has improved its coding efficiency and semantic understanding of Chinese by adding 20000 Chinese tags to the existing Glossary of LLaMA . Chinese-Alpaca-Plus has three versions: Chinese-Alpaca-Plus-7B, Chinese-Alpaca-Plus-13B, and Chinese-Alpaca-Plus-33B. In our experiments, we used the Chinese-Alpaca-Plus-7B version.

TigerBot Tigerbot-7b-sft-v1 is a language model developed by the Tigerbot Company. TigerBot-7b-sft-v1 is a large-scale language model with multiple languages and tasks. Tigerbot-7b-sft-v1 is an MVP version that has undergone 3 months of closed development and over 3000 experimental iterations.Functionally, Tigerbot-7b-sft-v1 already includes the ability to generate and understand most of the classes, specifically including several major parts: content generation, image generation, open-ended Q&A, and long text interpretation.Tigerbot-7b-sft has two versions: Tigerbot-7b-sft-v1 and Tigerbot-7b-sft-v2. In our experiments, we used the tigerbot-7b-sft-v1 version.

XrayPULSE XrayPULSE is an extension of PULSE and made by OpenMEDLab. OpenMEDLab utilize MedCLIP as visual encoder and Q-former (BLIP2) following a simple linear transformation as the adapter to inject the image to PULSE. For aligning the frozen visual encoder and the LLM by the adapter, OpenMEDLab generate Chinese-version Xray-Report paired data from radiology. By extending PULSE, XrayPULSE is fine-tuned on Chinese-version Xray-Report paired datasets and aims to work as a biomedical multi-modal conversational assistant. The basic model is PULSE and we did the tests on XrayPULSE by modifying the Checkpoint file.

DoctorGLM DoctorGLM is the first chinese diagnosis large language model (released at 3rd april 2023) that developed by ShanghaiTech University . It is fine-tuned on ChatGLM-6B using real-world online diagnosis dialogue. DoctorGLM has several updates and two different parameter-efficient finetune setting (p-tuning and LoRA). In our experiments, we used the DoctorGLM-5-22 p-tuning version.

Robin-7B-medical Robin-medical (LMFlow) is a toolkit providing a complete fine-tuning workflow for a large foundation model to support personalized training with limited computing resources. It is developed by Diao et al. from the Hong Kong University of Science and Technology. They provide a series of LoRA models based on the LLama model called Robin-medical, which are specially fine-tuned on the PubMedQA and MedMCQA datasets. The advantage of LMFlow is that it introduces an extensible and lightweight toolkit to simplify the fine-tuning and inference of general large foundation models. This allows people to fine-tune foundation models to mitigate the current status that most existing models exhibit a major deficiency in specialized-task applications. Robin-medical has 7B, 13B, 33B and 65B versions. We tested the 7B version in our experiments.

PaLM2 PaLM2 is a large language model developed by Google. PaLM 2 is a language model based on a tree structure, which makes use of the context and grammatical rules in the language to make the model’s understanding of text information more refined, accurate and comprehensive. Different from traditional sequence-based models (such as GPT), PaLM2 uses some new methods that are more popular than traditional methods, such as Tree-LSTM , Bert , etc. Compare to PaLM, PaLM2 excels at advanced reasoning tasks including code and math, classification and question answering, translation and multilingualism It excels at advanced reasoning tasks including code and math, classification and question answering, translation and multilingualism. It’s also being used in other state-of-the-art models, like Med-PaLM2 and Sec-PaLM. We tested the PaLM2 version in our experiments.

SenseNova SenseNova is a large language model developed by SenseTime. Through the trinity flywheel of data, model training and deployment, it can provide various large models and capabilities such as natural language, content generation, automatic data annotation, and custom model training. Based on the previous accumulation of NLP work by SenseTime, SenseNova is still good in the domestic large language model. Based on the "SenseNova" large-scale model system, SenseTime has also developed a series of generative AI models and applications including Miahua SenseMirage, Ronin SenseAvatar, Qiongyu SenseSpace, and Gewu SenseThings. We mainly tested SenseNova in this work.

Anthropic Claude2 Claude2 is a large language model developed by Anthropic, which is characterized by helpful and trustworthy. It is developed on the basis of Claude1.3. Anthropic uses a technical framework they call Constitute AI to achieve harmless processing of language models. Claude2 has a more powerful text processing function than GPT4, can handle larger-scale text, and has stronger context understanding ability and Chinese understanding ability. Claude is currently available in two versions, the powerful Claude, which excels at a wide range of tasks from complex dialogue and creative content generation to detailed instruction following, and the faster and more affordable Claude Instant, which also Can handle casual conversations, text analysis, summarization, and document question answering. We tested the latest version of Anthropic Claude2 for this work.

BayLing Bayling is an instruction-following large language model equipped with advanced language alignment. It is a product from Natural Language Processing Group, Institute of Computing Technology, Chinese Academy of Science. BayLing can be effortlessly deployed on a consumer-grade GPU. It shows superior capability in English/Chinese generation, instruction following and multi-turn interaction. Bayling has three versions: BayLing-7B-v1.0, BayLing-13B-v1.0, BayLing-13B-v1.1. In our experiments, we tested BayLing-7B.

3 Uniform Testing Prompts

For a fair and equitable comparison across different LLMs, we adopt a uniform approach in the selection and use of testing prompts. The same prompts are used across all models and conditions, regardless of whether they are zero-shot, one-shot, or five-shot scenarios.

In a zero-shot evaluation, the models are presented with a new task, with no prior examples given. For the one-shot scenario, we provide the model with one prior example. Meanwhile, in the five-shot scenario, the model is given five examples to learn from. These scenarios aim to mimic real-world usage conditions where models are given a limited number of examples and are expected to generalize from them.

4 Datasets

Our study utilizes two comprehensive and publicly available datasets, the MIMIC-CXR and the OpenI datasets. These datasets were utilized to test the performance and efficacy of various LLMs in generating radiology text reports.

In our study, we used these datasets to evaluate the capabilities of the LLMs. We focused on the "Findings" and "Impression" sections of each report as they provide comprehensive and detailed textual information about the imaging findings and the radiologists’ interpretations.

The MIMIC-CXR dataset is a substantial repository of de-identified chest radiographs (CXRs) that are complemented with their corresponding radiology reports. The dataset contains medical data from over 60,000 patients who were admitted to the Beth Israel Deaconess Medical Center between 2001 and 2012. The radiology reports in the MIMIC-CXR dataset typically consist of two sections: "Findings" and "Impression". The "Findings" section details observations from radiology images, while the "Impression" section provides summarized interpretations of these observations.

4.2 OpenI Dataset

The OpenI dataset is another essential resource that was used in our study. It is a freely available repository that consists of radiology images paired with their respective reports. This dataset provided an independent external platform to validate the performance and generalizability of our LLMs across different data sources.

We followed an existing literature approach to randomly divide the dataset into separate segments for testing purposes. This division resulted in a subset of 2400, 292, and 576 reports for various testing scenarios.

Results

This section presents the evaluation results of various large language models (LLMs) on two extensive datasets, OpenI and MIMIC-CXR. The performance of the models was assessed under three distinct shot settings: zero-shot, one-shot, and five-shot. Model performance was evaluated using three key metrics: Recall@1 (R-1), Recall@2 (R-2), and Recall@L (R-L).

On the OpenI dataset, Anthropic Claude2 excelled in the zero-shot setting, achieving an R-1 score of 0.2372, an R-2 score of 0.1259, and an R-L score of 0.2193. These results notably surpassed those of other models under the same setting. In the one-shot scenario, the model achieving the highest R-1 score was BayLing-7B with 0.1268, followed closely by Luotuo-lora-7B-0.3 and Ziya-LLaMA-13B-v1 with scores of 0.152 and 0.1502, respectively. However, BayLing-7B was the standout performer in the five-shot setting, registering the highest scores across all metrics with an R-1 score of 0.4506, an R-2 score of 0.3452, and an R-L score of 0.4436.

2 MIMIC-CXR Dataset Results

The evaluation on the MIMIC-CXR dataset showed that the Anthropic Claude2 model retained its superior performance in the zero-shot setting, achieving an R-1 score of 0.3177, an R-2 score of 0.153, and an R-L score of 0.256. PaLM2 emerged as the leading model in the one-shot setting, delivering an R-1 score of 0.2711, an R-2 score of 0.1446, and an R-L score of 0.2251. In the five-shot scenario, the BayLing-7B model continued to outperform other models with the highest R-1 score of 0.2901, R-2 score of 0.1722, and R-L score of 0.2747.

However, some models like AtomGPT_8k registered considerably lower performance across all shot settings and both datasets. For example, AtomGPT_8k scored remarkably low in the OpenI zero-shot setting, with an R-1 score of 0.0287. It continued to score low across other shot settings and in the MIMIC-CXR dataset.

In conclusion, this evaluation underscores the significant diversity in the capabilities of different LLMs, emphasizing the need for careful model selection for specific tasks. The performance variance across different shot conditions has important implications for task-specific LLM selection in future research and applications.

Discussion

The present study has conducted one of the most exhaustive assessments of world-wide LLMs, focusing primarily on their utilization within the domain of radiology. The meticulous evaluation of these models, using extensive radiology report datasets and juxtaposing them with established global leading models, provides significant insights into their capabilities, limitations, and potential roles within the healthcare sector.

Our findings underscore that multiple LLMs perform comparably in interpreting radiology reports. This alignment points to their advanced natural language understanding skills and highlights their potential utility in enhancing radiology practice, where they can aid in automating radiological image interpretation, assisting in preliminary diagnosis, and thereby freeing up time for healthcare professionals. This is particularly beneficial in regions with limited access to radiologists or in healthcare scenarios where high volumes and time constraints pose significant challenges.

2 Inter-model Differences and Implications

While the performance of the world-wide models showed broad alignment, our results also spotlighted some disparities between the different models. This variance in strengths and weaknesses indicates that the choice of an LLM for a specific application should depend on the particular requirements of that task. Hence, a more profound understanding of these models, to which our study contributes, is critically essential for their effective deployment in the field.

3 Implications of Evaluation Metrics

The evaluation metric adopted in our study is Rouge Score, an N-gram-based method that inherently measures how well models conform to set answers. GPT-4, a universally recognized powerful model, did not outperform its counterpart, ChatGPT, nor did it surpass other models in the Rouge Score. This discrepancy invites a questioning of the significance of Rouge Score as a measure of radiology knowledge. The BayLing model, for instance, tended to produce succinct answers which, despite their brevity, may be of high quality and accuracy. On the contrary, GPT-4 may be more verbose and consider issues more comprehensively, showing some level of distrust in the input. The difference in results highlights the need to carefully interpret the evaluation scores, taking into account the unique characteristics of each model.

4 Model Size and Performance

Our analysis reveals that to achieve high performance in this specific task, there is no strict need for large models. Models with 7B parameters can produce impressive results, suggesting that we might be on the verge of a fourth industrial revolution driven by these more accessible, lightweight models. This prompts a reconsideration of the belief that model performance is strongly correlated with the size of the model. In fact, smaller models also demonstrated strong capabilities, raising the question of whether intelligence truly arises from the number of parameters and data accumulation.

5 Multimodal LLMs: The Next Frontier

The advent of multimodal LLMs, capable of managing multiple forms of input such as text and images, creates fascinating prospects for future research. Evaluating these models’ aptitude to directly interpret radiological images, in addition to textual reports, could revolutionize radiology practice. These multimodal models could find uses in areas like disease detection and diagnosis, treatment planning, and patient monitoring.

Conclusion

In this comprehensive study, we rigorously evaluated the performance of 32 significant world-wide LLMs in the healthcare and radiology sector, comprising both global leading models such as ChatGPT, GPT-4, PaLM2, Claude2 and a robust suite of LLMs developed in other countries such as China. The overarching goal of this exploration was to benchmark these models in the context of interpreting radiology reports, enabling a nuanced understanding of their diverse capabilities, strengths, and weaknesses. Our findings affirm the competitive performance of many Chinese LLMs against their global counterparts, emphasizing their untapped potential in healthcare applications, particularly within radiology. This suggests a trajectory towards a future where these multilingual and diverse LLMs contribute to an enhanced global healthcare delivery system.

Looking ahead, our large-scale study’s insights offer a compelling foundation for further exploratory research. There is immense scope for expanding these LLMs into different medical specialties and developing multimodal LLMs, the latter of which could handle complex and diverse data types to provide a more comprehensive understanding of patient health. However, as we navigate this evolving landscape of LLMs, it is imperative to give due consideration to their effective application and ethical deployment. In conclusion, our study hopes to catalyze further exploration and discussion, envisioning an era where LLMs significantly aid in healthcare provision and contribute to an enhanced standard of global patient care.

References