Radiology-Llama2: Best-in-Class Large Language Model for Radiology

Zhengliang Liu, Yiwei Li, Peng Shu, Aoxiao Zhong, Longtao Yang, Chao Ju, Zihao Wu, Chong Ma, Jie Luo, Cheng Chen, Sekeun Kim, Jiang Hu, Haixing Dai, Lin Zhao, Dajiang Zhu, Jun Liu, Wei Liu, Dinggang Shen, Tianming Liu, Quanzheng Li, Xiang Li

Introduction

Transformer-based large language models (LLMs) such as ChatGPT and GPT-4 have shown impressive capabilities in natural language processing . The development in transformer-based NLP models has also spurred advancements in developing and applying transformer-based models in computer vision and other modalities . Since November 2022, inspired by the versatile capabilities and wide popularity of ChatGPT, LLMs have been applied in clinical studies , pharmacy , radiology , Alzheimer’s disease , agriculture and brain science research .

However, their application in specialized domains like healthcare has been limited.

First, localized large language models are a must for real world healthcare, since hospitals cannot share data or upload data to commercial models such as ChatGPT or GPT-4 due to privacy regulations .

In addition, LLMs trained on general domain, such as ChatGPT , GPT-4 and PaLM 2 , lack medical knowledge in specialized domains such as radiology, and it is necessary to design a model that is properly trained on domain data that is clinically meaningful.

Moreover, our Radiology-Llama2 perfectly imitates the style or radiologists, yet models like ChatGPT generate comprehensive but Wikipedia-like responses unlike the concise and simple language style of real radiologists that facilitates quick information exchange.

Finally, this work opens the door for personalized radiological assistants that are tailored to the style of individual physicians .

This work addresses this gap through Radiology-Llama2, an LLM tailored for radiology through instruction tuning to generate radiology impressions from findings. Evaluations show it surpasses general LLMs in coherence, conciseness and clinical utility of generated impressions.

State-of-the-Art Performance: Outperforms any other language models in deriving clinical impressions , setting a new benchmark on MIMIC-CXR and OpenI datasets.

Flexibility and Dynamism: Unlike its BERT-based counterparts , Radiology-Llama2 is not tied to a specific input structure, allowing for a broader range of inputs and adaptability to different tasks within radiology, including complex reasoning.

Clinical Usability with Conversational Capabilities: Generative LLMs offers inherent conversational functionality , enabling it to provide contextual insights and responses in a human-like manner. This makes Radiology-Llama2 particularly useful for medical professionals in a clinical setting, enhancing both diagnosis and reporting.

Related work

Recent developments in NLP are marked by the emergence of LLMs such as GPT-3 , GPT-4 , PaLM , and PaLM-2 . Contrasting the earlier pre-training and fine-tuning approach observed in BERT , GPT , GPT-2 , and their variants , these new LLMs exhibit few-shot and zero-shot learning capabilities using in-context learning. Furthermore, open-source models like LLaMA and Bloom have entered the scene, promoting broader accessibility.

There’s also an increasing interest in instruction-tuned models such as Alpaca , StableLM , and Dolly .

2 Domain-Specific Language Models (DSLMs)

DSLMs, such as AgriBERT , are tailored to specific domains, aimed at optimal performance in related tasks. Specifically, AgriBERT is trained on agricultural texts, making it suitable for tasks in agriculture. SciEdBERT is designed for the educational sector and focuses on middle school chemistry and physics, offering insights into evaluating students’ responses. ClinicalRadioBERT , in the healthcare sector, is adept at radiation oncology and emphasizes its training from clinical notes and related literature. These DSLMs highlight the potential and adaptability of specialized models across various sectors .

Methodology

Radiology-Llama2 is trained on a large radiology dataset using instruction tuning to generate radiology impressions from findings. Evaluations by radiologists show it surpasses general LLMs in coherence, conciseness and clinical utility of generated impressions.

MIMIC-CXR is a large chest radiographs dataset which consists of 227,835 imaging studying based on 65,379 patients presenting to the Beth Israel Deaconess Medical Center Emergency Department between 2011–2016 . There are 377,110 available images in the dataset where each imaging studying contains one or more images (typically a frontal view and a lateral view). This dataset also has the corresonding free-text radiology reports and has been de-identified to ensure the US Health Insurance Portability and Accountability Act of 1996 (HIPAA) Safe Harbor requirements. Widely application of MIMIC-CXR dataset has been implemented in computer vision, natural language processing and decision support etc.

1.2 OpenI Dataset

OpenI dataset is a publicly available dataset aiming to make clinical documents for secondary use in the region of research and education . This dataset collects 8121 images form the hospitals’ picture archiving systems accompanied with 3996 corresponding radiology reports from the Indiana Network. Manual coding has been added into the radiologist reports in order to increase the relevancy of the retrieved clinical documents. Similar to MIMIC-CXR dataset, OpenI uses manually verification after the automatic method to achieve de-identification.

2 Instruction Tuning

Instruction tuning is a foundational component of the Radiology-Llama2 framework. Instruction tuning addresses the fundamental disconnect between the traditional training objectives of LLMs and the user-specific goals of instruction following. This technique involves additional training using pairs of human-specified instructions and corresponding desired outputs. It serves to align the model with task-specific user objectives, enhance model controllability, and allow for rapid domain-specific adaptation, all while maintaining computational efficiency.

To bolster learning, instructions are formatted specifically. For instance, the "Findings" text is supplied with a succinct instruction, such as "Derive the impression from findings in the radiology report", while the "Impression" text from the same report serves as the target output. This approach calibrates the model in alignment with the desired task, thus yielding an instruction-adhering language model optimized for radiology reports.

2.2 Domain-Specific Knowledge Acquisition

Through training on domain-specific data, the model is also adept at assimilating domain-specific knowledge that is quintessential to radiology. Consequently, Radiology-Llama2 is proficient in capturing language patterns, terminologies, and logical reasoning essential for interpreting radiology reports.

The initial instruction tuning is centered on the "Findings" to "Impression" conversion, which holds significant clinical value. Recognizing the potential of diverse instruction pairs in radiology, ongoing engagements with radiologists are aimed at formulating a diverse set of clinically pertinent instruction tuning pairs to augment the capabilities of Radiology-Llama2.

3 Experimental Setting

The Radiology-Llama2 model employs an advanced training regimen. For reference, when training Radiology-GPT with the Llama2-7b-chat base model, the training was facilitated by Low-Rank Approximations (LoRA) . The choice of LoRA was motivated by its compact size and portability which are conducive to model sharing and deployment.

The experimental setting for the training comprised the following configurations:

lora_r (rank of low-rank factorization): 8

lora_alpha (scaling factor for the rank): 16

The target modules for LoRA were set to "q_proj" and "v_proj", corresponding to the query and value matrices in the self-attention mechanism of the transformer architecture. The training was conducted on a server equipped with 4 Nvidia A100 80GB GPUs.

Results

The present study evaluated the performance of various large language models on two key datasets pertinent to radiology: MIMIC-CXR and OpenI. The assessment employed Rouge-1, Rouge-2, and Rouge-L as the primary metrics, given their widespread acceptance for evaluating the quality of generated text.

Radiology-Llama2 significantly outperforms all comparison models across all ROUGE metrics: ROUGE-1, ROUGE-2, and ROUGE-L, on both MIMIC-CXR and OpenI datasets. Results can seen in Table 1 and Table 2.

For the MIMIC-CXR dataset, Radiology-Llama2 achieves scores of 0.4834 in ROUGE-1, 0.324 in ROUGE-2, and 0.4427 in ROUGE-L. These scores are markedly higher than those of the second-best performing model, Anthropic Claude2, which manages 0.3177 in ROUGE-1 and 0.153 in ROUGE-2. This demonstrates that Radiology-Llama2 not only captures a higher proportion of overlapping unigrams between the generated and reference summaries but also maintains this superiority in capturing bigrams and maintaining a longer sequence of content overlap.

Similarly, in the OpenI dataset, Radiology-Llama2 sustains its exemplary performance, recording scores of 0.4185 in ROUGE-1, 0.2569 in ROUGE-2, and 0.4087 in ROUGE-L. In comparison, the second-best model, Anthropic Claude2, scores 0.2372 in ROUGE-1 and 0.1259 in ROUGE-2. The substantial gap between the two models across all metrics underscores Radiology-Llama2’s robustness and generalizability across datasets.

On the lower end of the performance spectrum, Baichuan-7B exhibits exceptionally low scores on both datasets. For instance, its ROUGE-2 score is a meager 0.0057 on the MIMIC-CXR dataset, and its ROUGE-L score is only 0.0029. This emphasizes the limitations of such models in capturing even the basic elements of content overlap.

It is also noteworthy that Radiology-Llama2 maintains its superiority not only in terms of single-term overlap but also in capturing a longer chain of content, as reflected in its ROUGE-L scores.

This paper has provided two specific examples. The first example can be seen in Figure 2 and Figure 3, the finding is “The lungs are hyperexpanded. Heart size normal. No mass or focal opacities seen. Stable degenerative changes of the thoracic spine.” Six LLMs are required to derive the impression from the findings in the radiology report. It can be seen that LLAMA2, ChatGPT and Alpaca all understand the content of the radiology report, however, their answers are too trivial and fail to catch the most important points, which will make their given impression hard to understand and increase the difficulty to be used clinically. In contrast, although StableLM, Dolly and LLAMA basically understand the intention of the text, their answers are not satisfactory. Their answers are not only redundant but also contain a lot of irrelevant information. Radiology-GPT, which has been performed a fine-tuning on a specific data set, has the answer which is obviously more capable and concise, but the gap with the impression given by Radiologist is not small. Therefore, in comparison, it can be clearly seen that the excellent performance of Radiology-LLAMA2, the answer of Radiology-LLAMA2 is not only accurate and concise, but also has the report style which is closest to Radiologist. This intuitively proves the performance of Radiology-LLAMA2 and proves its strong clinical application potential.

In summary, Radiology-Llama2 consistently demonstrates superior performance in generating clinically relevant and coherent radiology reports, as evidenced by its high ROUGE scores across multiple datasets.

2 Expert Evaluation

To supplement the quantitative evaluations, we conducted an expert-based assessment of the models. We randomly selected 10 records each from MIMIC-CXR and OpenI datasets and had two experienced radiologists manually evaluate the generated radiology impressions based on five key criteria: Understandability, Coherence, Relevance, Conciseness, and Clinical Utility.

Understandability: Radiology-Llama2 and Radiology-GPT both stood out with a score of 48.5, suggesting that their generated impressions are highly understandable to radiologists. ChatGPT closely followed with a score of 47. On the other end of the spectrum, Llama scored the lowest at 14, indicating significant limitations in generating understandable content.

Coherence: Radiology-Llama2 once again led the cohort with a score of 47.5, closely followed by Radiology-GPT at 48 and ChatGPT at 46.5. Llama lagged considerably, scoring only 13.5, which raises questions about its ability to produce coherent clinical text.

Relevance: In terms of the relevance of the generated text to the radiology findings, ChatGPT surprisingly took the lead with a score of 49.5. Radiology-Llama2 was close behind with a 46.5, while Radiology-GPT scored 43.5. Llama and StableLM were the least relevant models, scoring 14 and 25, respectively.

Conciseness: Radiology-Llama2 excelled in generating concise impressions, scoring the highest at 49. Llama, StableLM, and Dolly demonstrated shortcomings in this regard, scoring 14, 24.5, and 27.5 respectively.

Clinical Utility: Radiology-Llama2 emerged as the most clinically useful model, garnering a top score of 50. In contrast, Llama scored the lowest in this metric as well, with a score of 13, highlighting its limited utility for clinical applications.

In addition, we evaluated the Rouge scores of a few select models on these 10 samples to cross-verify with the experts’ evaluation. Please see 1 for more details.

Overall, Radiology-Llama2 consistently demonstrated superior performance across all five criteria, affirming its status as a highly effective tool for generating radiology impressions. While other models like Radiology-GPT and ChatGPT showed competence in certain areas, they were unable to match the all-around excellence of Radiology-Llama2. Models like Llama and StableLM displayed significant limitations, performing poorly across multiple criteria.

Discussion

The high Rouge scores attained by Radiology-LLAMA2 suggest that this language model has the potential for significant impact in clinical radiology diagnosis. Its ability to swiftly generate coherent and clinically relevant reports could be transformative, particularly in busy radiology departments where timely and accurate reporting is of the essence. By automating certain aspects of the reporting process, the model serves as a valuable assistive tool for radiologists, allowing them to focus more on complex cases that require nuanced human expertise.

2 Need for Diverse Training Data

While Radiology-LLAMA2 performs impressively on generating findings and impressions, its utility could be further broadened by incorporating more diverse forms of training data. For instance, it could be trained on instructional text for medical procedures, summaries of patient histories, or physician’s notes on differential diagnoses. By diversifying the data sources, the model would be better equipped to assist in various facets of radiological practice, such as recommending further diagnostic tests or suggesting possible treatment paths.

3 Multimodality: Adding Image Capabilities

The addition of image analysis capabilities would elevate Radiology-LLAMA2 from a text-generation model to a truly multimodal diagnostic tool. Future iterations could integrate machine learning algorithms for image recognition, enabling the model to make direct observations from X-rays, MRIs, or CT scans. This would potentially create a more holistic diagnostic process where textual and visual data are analyzed in tandem for more accurate and comprehensive diagnoses.

4 Conversational Assistant to Radiologists

Beyond report generation, Radiology-LLAMA2 could be developed into a conversational assistant that helps radiologists in real-time. This would enable a more dynamic interaction, where the model could assist in tasks ranging from quick data retrieval to offering second opinions on diagnoses. Such a system could act as a "second pair of eyes," providing immediate feedback and thus serving as a valuable safeguard against diagnostic errors.

Conclusion

Radiology-Llama2 demonstrates localized LLMs can transform radiology when designed appropriately. With proper oversight, it has great potential for clinical decision support and other applications. This work pave the way for specialized LLMs in other medical domains.

In conclusion, Radiology-Llama2 represents an important advance in applying LLMs to healthcare. With continued research into model design and evaluation, such specialized LLMs can enable breakthroughs in medical AI.

References