RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models
Peng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu, Yun Li, Gang Li, Linjun Zhang, Huaxiu Yao
Introduction
Artificial Intelligence (AI) has showcased its potential in medical diagnosis, including disease identification, treatment planning, and recommendations Tăuţan et al. (2021); Wang et al. (2019); Ye et al. (2021); Xia et al. (2024b); Li et al. (2024). In particular, the recent development of Medical Large Vision Language Models (Med-LVLMs) has introduced more accurate and customized solutions to clinical applications Li et al. (2023); Moor et al. (2023); Zhang et al. (2023); Wu et al. (2023). While Med-LVLMs have demonstrated promising performance, they remain prone to generating responses that deviate from factual information, potentially resulting in inaccurate medical diagnoses. This susceptibility to hallucination underscores the need for enhanced mechanisms to ensure factual alignment in critical medical applications (see an example in Figure 1(a)) Royer et al. (2024); Xia et al. (2024a)). Such errors pose a significant risk to clinical decision-making processes and can lead to adverse outcomes.
Recently, Retrieval-Augmented Generation (RAG) Gao et al. (2023) has emerged as a promising method for enhancing the factual accuracy of responses from Med-LVLMs. By integrating external, reliable data sources, RAG guides the model in producing factual medical responses, enriching its knowledge base with supplementary information. For example, RAG has been used in tasks such as visual question answering (VQA) Yuan et al. (2023) and report generation Kumar and Marttinen (2024); Tao et al. (2024). However, as illustrated in Figure 1(b) and Figure 1(c), directly applying RAG strategy to Med-LVLMs presents two significant challenges: (1) A small number of retrieved contexts may not cover the reference knowledge required for the question, thus limiting the model’s factual accuracy. Conversely, a large number of retrieved contexts may include low-relevance and inaccurate references, which can interfere with the model’s generation; (2) Med-LVLMs may overly rely on the retrieved information. In this situation, the model might correctly answer on its own, but incorporating the retrieved contexts could lead to incorrect responses.
To tackle these challenges, we propose the Reliable mUltimodaL RAG called RULE for MEd-LVLMs. First, RULE introduces a provable strategy for factuality risk control through calibrated selection of the number of retrieved contexts , ensuring that Med-LVLMs provably achieve high accuracy without the need for additional training Angelopoulos et al. (2021). Specifically, this strategy modifies the Med-LVLM through a post-processing step that performs hypothesis testing for each to determine whether the risk can be maintained above an acceptable threshold. This process begins by calculating the -value for each . Fixed sequence testing is then used to determine which values can be accepted. Second, to mitigate over-reliance on retrieved knowledge, we introduce a knowledge balanced preference fine-tuning strategy. This strategy harmonizes the model’s internal knowledge with retrieved contexts during medical response generation. Here, we identify samples where the model initially responds correctly but gives incorrect answers after incorporating retrieved contexts as dispreferred samples, indicating retrieval over-dependence. Conversely, ground-truth responses are considered as preferred samples. The curated preference data is then utilized for fine-tuning the preferences in Med-LVLMs.
Our primary contributions of this paper is RULE, which introduces an innovative approach to enhance retrieval-based Med-LVLMs. RULE not only controls factual risk by calibrating the selection of reference contexts but also balances the model’s knowledge and retrieved contexts through preference fine-tuning using a curated preference dataset. Across three medical Visual Question Answering (VQA) benchmarks, including radiology and ophthalmology, our empirical results demonstrate that RULE effectively improves the factual accuracy of Med-LVLMs, achieving a 8.06% improvement over the best prior methods for mitigating hallucination. In addition, empirically verify the effectiveness of the proposed components and demonstrate the compatibility of RULE.
Preliminaries
In this section, we will provide a brief overview of Med-LVLMs and preference optimization.
Medical Large Vision Language Models. Med-LVLMs connects the LLMs and medical visual modules, enabling the model to use medical images and clinical queries as inputs . This allows the model to autoregressively predict the probability distribution of the next token. The text output of Med-LVLMs is denoted as .
Preference Optimization. Preference optimization has achieved remarkable results in efficiently fine-tuning LLMs, significantly aligning their behavior with the goals. Typically, give an input , a language model policy can produce a conditional distribution with as the output text response. The recently popular DPO Rafailov et al. (2023) utilizes preference data achieve objective alignment in LLMs. The preference data is defined as , where and represent preferred and dispreferred responses given an input prompt . The probably of obtaining each preference pair is where is the sigmoid function. In DPO, the optimization can be formulated as classification loss over the preference data as:
where represents the reference policy, which is the LLM fine-tuned through supervised learning.
Methodology
In this section, as illustrated in Figure 2, we will introduce RULE as an efficient solution for improving factuality of Med-LVLMs. Specifically, our approach consists of three main modules that work together to optimize the model’s performance. First, we apply the retrieval strategy to Med-LVLMs, enhancing the model’s ability to leverage retrieved information. Second, we implement a statistical method to control the factuality risk through calibrated selection of retrieved contexts. Third, we develop a preference optimization method to balance the model’s reliance on its own knowledge and the retrieved contexts. Next, we will detail these three key modules in detail as follows:
Med-LVLMs often generate non-factual responses when dealing with complex medical images. RAG can provide the model with external knowledge as a reference, thereby effectively enhancing the factual accuracy. In the multimodal knowledge retrieval stage, RULE retrieves textual descriptions/reports that are most similar to the features of the target medical images. These references contain a wealth of image-based medical facts and serve to guide the generation of responses for the medical image.
2 Factuality Risk Control Through Calibrated Retrieved Context Selection
For the RAG strategy, the top-3/5 result is typically used as a reference Gao et al. (2023). However, it sometimes fails to encompass all relevant retrieved contexts, especially when facing the fine-grained features of medical images. Additionally, an excessive amount of retrieved contexts may introduce low-relevance and inaccurate references, which can interfere with the model’s generation. Thus, an algorithm that can automatically determine the optimal number of retrieved contexts, based on the risk of factual errors, is particularly crucial.
where is the Kullback-Leibler divergence between two Bernoulli distributions and denotes risk upper bound. representing the probability that, in a binomial distribution with parameters and , denoted by , the observed value is less than or equal to . Then, the minimum of these two probabilities is taken. Finally, we use any family-wise error rat (FWER)-controlling procedure, such as Bonferroni correction Van der Vaart (2000) or sequential graphical testing Bretz et al. (2009), to choose . For example, for Bonferroni correction, if is less than or equal to , where denotes tolerance level, then is added to the set . The proposed strategy calculates the model’s factuality risk under different values, computes the corresponding probabilities using two approaches, and selects those values that meet the risk tolerance to control the overall factuality risk.
We have the following result that ensures with probability at least , the factuality risk produced is controlled by .
Let . If the training dataset is and the output of the above algorithm , then
In practice, we calibrate the selection of on the validation sets of each dataset to minimize factuality risk. Consequently, the optimal calibrated by this algorithm can be directly used on the test sets.
3 Knowledge Balanced Preference Tuning
In addition to selecting the optimal number of retrieved contexts, it is likely that these contents often fail to fully capture the details of every lesion or normal area in medical images. Therefore, when the retrieved contexts is inaccurate, a reliable Med-LVLM is expected to remain unaffected by the unreliable information and independently use its own knowledge to answer medical questions. However, empirically, as illustrated in Table 1, approximately half of all incorrect responses by the retrieval-augmented Med-LVLM are due to an over-reliance on retrieved contexts. This significantly affects the application of the retrieval augmented generation strategy to Med-LVLMs.
To address this issue, we propose a Knowledge-Balanced Preference Tuning (KBPT) strategy to mitigate over-reliance on retrieved contexts and enhance factuality in medical content generation. Specifically, we select samples from the a separate set with samples are not used to fine-tune the retriever in Section 3.1, where denotes input medical image, ground-truth answer and question, respectively. We identify responses where the model originally answers (i.e., ) correctly but gives incorrect answers after incorporating retrieved contexts as dispreferred responses, as they indicate over-dependence on the retrieval. Conversely, ground-truth answers are considered preferred responses. We denote the preference dataset as , where , are represented as preferred and dispreferred responses, respectively.
Based on the curated preference data, we fine-tune the Med-LVLM using direct preference optimization. Following Eqn. (1), the loss is calculated as follows:
Experiment
In this section, we evaluate the performance of RULE, aiming to answer the following questions: (1) Can RULE effectively improve the factuality of Med-LVLMs compared to other baselines and open-sourced Med-LVLMs? (2) Do all proposed components boost the performance? (3) How does RULE change attention weights of retrieved contexts to balance model knowledge and retrieved contexts? (4) How do different types of data or models influence DPO fine-tuning?
Implementation Details. We utilize LLaVA-Med-1.5 7B Li et al. (2023) as the backbone model. During the preference optimization process, we adapt LoRA fine-tuning Hu et al. (2021). For the training of retriever, the vision encoder is a ResNet-50 He et al. (2016), and the text encoder is a bio-BioClinicalBERT Alsentzer et al. (2019). We use the AdamW optimizer with a learning rate of , weight decay of and a batch size of 32. The model is trained for 360 epochs. For more detailed information on training hyperparameters and training data, please see Appendix A and C. Baselines. We compare RULE with LVLM hallucination mitigation methods that have already shown promising results in natural images, including Greedy Decoding, Beam Search Sutskever et al. (2014), DoLa Chuang et al. (2023), OPERA Huang et al. (2023), VCD Leng et al. (2023). These methods manipulate the logits of the model’s output tokens to enhance factual accuracy. Furthermore, we compare the performance with other open-source Med-LVLMs, including Med-Flamingo Moor et al. (2023), MedVInT Zhang et al. (2023), RadFM Wu et al. (2023).
Evaluation Datasets. To ensure that the retrieved report content is relevant to the visual question-answering content and to facilitate experimentation, we utilize three medical vision-language datasets, i.e., MIMIC-CXR Johnson et al. (2019), IU-Xray Demner-Fushman et al. (2016), and Harvard-FairVLMed Luo et al. (2024), encompassing radiology and ophthalmology. The training set is split into two parts: one part is used to train the retriever (Section 3.1), and the other part is used to construct the preference dataset for KBPT (Section 3.3).
Additionally, we construct VQA pairs for KBPT and evaluation. Specifically, the reports from training set for preference dataset and reports from original test set are input into GPT-4 OpenAI (2023) to create closed-ended VQA data with yes or no answers, e.g., "Is there any pulmonary nodule?". By sampling segments from a medical report, we can generate a sequence of concise, closed-ended questions posed to the model, each with accurate answers. The questions are in yes/no format, making it easier to analyze errors caused by over-reliance on retrieved contexts compared to open-ended questions. The detailed construction process and dataset statistics are provided in the Appendix A.
Evaluation Metrics. We use Accuracy as the primary metric and, for detailed comparisons, we also adopt Precision, Recall, and F1 Score.
2 Results
In this section, we provide comprehensive comparison results with different baseline methods and other open-sourced Med-LVLMs.
Comparison with Baseline Methods. We present the results of a comparison between RULE and various hallucination reduction methods in Table 2. According to these results, RULE demonstrates the best overall performance, effectively and accurately diagnosing diseases with an average accuracy improvement of 20.8% across all datasets. We also observe that RULE performs notably better on the IU-Xray and Harvard-FairVLMed compared to MIMIC-CXR. This difference is attributed to the excessive length of the reports available for retrieval in MIMIC-CXR, where overly long references tend to confuse the Med-LVLM. In addition, even when dealing with the relatively niche ophthalmology data (i.e., Harvard-FairVLMed), RULE demonstrates superior results, significantly enhancing the factual accuracy of the Med-LVLM. In contrast, the performance of decoding methods is quite unstable, showing significant rates of missed or incorrect diagnoses across different datasets, as indicated by the precision and recall values.
Comparison with Other Med-LVLMs. In Table 3, we present the comparison with different open-sourced Med-LVLMs. RULE demonstrates state-of-the-art (SOTA) performance across all datasets. Although the second-best model, MedVInT, outperforms other models, RULE achieves an average accuracy improvement of 47.4% over it. Whether in radiology or ophthalmology, RULE demonstrates remarkable performance, significantly surpassing other open-source Med-LVLMs. This indicates that RULE is generally applicable and effective in the medical multimodal diagnosis, providing consistent improvements across various medical image modalities.
3 How Does RULE Improve the Performance?
In this section, we conduct a set of analyses demonstrate how different components contribute to the performance and illustrate how RULE enhances overall performance, which are details as follows:
Ablation Studies. To further illustrate the effectiveness of the components of RULE, we conduct ablation experiments on three datasets. The results are shown in Table 4. We find that the basic RAG strategy ("R") slightly improves factual accuracy on two datasets but decreases it on MIMIC-CXR. The limited retrieved contexts can not cover the fine-grained features of medical images, resulting in unstable factual accuracy improvements. With the aid of the factuality risk control strategy ("FRC"), retrieval performance see a stable increase, outperforming the original Med-LVLM. Considering the model’s over-reliance on retrieved contexts, the knowledge balanced preference tuning ("KBPT") further enhances the model’s reliability and significantly improves its performance. Ultimately, by combining these two strategies, RULE achieves optimal performance.
How does RULE Mitigate the Issue of Over-Reliance on Retrieved Contexts? To better understand how RULE mitigates the Med-LVLM’s over-reliance on retrieved contexts, we measure the Med-LVLM’s error and over-reliance ratios, and visualize the text and image attention maps of the models before and after fine-tuning using a randomly selected case, as shown in Figure 3. The quantitative results in Figure 3(a) demonstrate the significant positive impact of RULE in mitigating the model’s over-reliance on retrieved contexts, with the error rate and over-reliance rate decreasing by an average of 42.9% and 47.3%, respectively. Attention maps Figure 3(b) illustrate the model’s attention scores for text and image tokens. We find that, on the text side, the model with knowledge balanced preference tuning shows a significantly reduced focus on retrieved contexts, effectively mitigating over-reliance on such information. The model focuses more on the question and leverages its own knowledge to answer, rather than relying solely on the retrieved contexts, effectively enhancing factual accuracy.
Analyzing Preference Data Type in KBPT. We further conduct a thorough analysis of the data types used in constructing preference data for KBPT. Three formats are considered: medical image captioning (prompted as “Please describe this medical image"), visual question-answering (VQA), and a mixture of both. The selected data are samples where the model makes errors due to over-reliance on retrieved contexts. The results are shown in Table 5. We observe that models fine-tuned using VQA data perform the best across all three datasets. This indicates that when retrieved contexts are incorporated into VQA questions, the Med-LVLM, through KBPT, can learn this paradigm of integrating and balancing its own knowledge with retrieved context to maximize factual accuracy. However, when the data is in the form of captioning, it may enhance the model’s ability to describe medical facts, but it merely distances the model’s answers from the retrieved contexts. The model fails to understand how to balance retrieval content with its own knowledge.
4 Compatibility Analysis
To demonstrate the compatibility of RULE, we conduct KBPT on LLaVA-Med-1.0 as well. The experimental results on three datasets are shown in Figure 4. We find that our knowledge balanced preference tuning method demonstrates good compatibility across different models, significantly improving factual accuracy across multiple datasets. Based on LLaVA-Med-1.0, RULE increases accuracy by an average of 16.7%. This indicates that RULE has a noticeable positive effect on mitigating over-reliance on retrieved contexts, thereby enhancing the Med-LVLM’s factual accuracy.
5 Case Study
Figure 5 presents two representative case results, demonstrating that RULE can effectively enhance the factual accuracy of med-LVLMs. In case 1, LLaVA-Med provides a factually incorrect answer. After applying the RAG strategy, the model still exhibits factual issues, whereas our method effectively addresses this and improves accuracy. In case 2, LLaVA-Med initially provides a correct answer, but due to the model’s over-reliance on retrieved contexts, it subsequently produces an incorrect response. RULE balances the weight of inherent knowledge and retrieved contexts, enhancing factual accuracy.
Related Work
Factuality in Med-LVLMs. The rapid development of Large Vision and Language Models (LVLMs) Liu et al. (2023b, a); Zhu et al. (2023); Alayrac et al. (2022); Zhou et al. (2024a, b) has begun to impact medical diagnosis. A series of Med-LVLMs Li et al. (2023); Moor et al. (2023); Wu et al. (2023); Zhang et al. (2023), represented by LLaVA-Med, have emerged, demonstrating impressive performance across various medical image modalities. However, Med-LVLMs still exhibit significant factual errors, producing medical responses that conflict with the visual medical information. This could potentially lead to misdiagnoses or missed diagnoses. Recently, several benchmarks Royer et al. (2024); Xia et al. (2024a) have been established to evaluate the accuracy of Med-LVLMs in tasks such as VQA or report generation. Beyond evaluating factuality, improving the factual accuracy of Med-LVLMs remains an underexplored area.
Retrieval Augmented Generation. RAG has recently been recognized as a promising solution Gao et al. (2023). It enhances the model’s ability to generate accurate facts by incorporating contextual information from external datasets. In medical multimodal analysis, the RAG approach has been applied to various tasks such as medical VQA Yuan et al. (2023) and report generation Kumar and Marttinen (2024); Tao et al. (2024); He et al. (2024). However, in Med-LVLMs, applying RAG-based approaches overlook two critical issues: the number of retrieved contexts and whether the model overly relies on these reference. These factors can significantly affect the model’s performance and may even degrade it. In RULE, we systematically address these challenges and enhance the factuality of Med-LVLMs.
Conclusion
In this work, we aim to enhance the factuality of Med-LVLM by addressing two key challenges in medical RAG. Specifically, we first introduce a provably effective strategy for controlling factuality risk through the calibrated selection of retrieved contexts. Second, we develop a preference optimization strategy that addresses errors stemming from the model’s excessive dependence on retrieved contexts, aiming to balance its intrinsic knowledge and the retrieved information. Experiments on three medical imaging analysis datasets demonstrate the effectiveness of RULE.
Limitations
This work explores a reliable multimodal RAG method for Med-LVLMs to enhance factual accuracy. Our primary focus is on factual accuracy. Future research can explore other issues related to deploying Med-LVLMs in clinical settings, such as safety, fairness, robustness, and privacy.
Acknowledgement
This research was supported by Cisco Faculty Research Award.
References
Appendix A Data
The quantities of all the data used are shown in Table 6 and Table 7. It is notable to note that for training the retriever, this refers to the number of image-text pairs; for fine-tuning, it refers to the number of QA items. “All" represents the total quantity used to construct the preference dataset, where only the samples with correct original answers that become incorrect after adding retrieved contexts are included in the training of knowledge balanced preference tuning (“KBPT").
A.2 Instructions
We convert the medical reports into a series of closed-ended questions with yes or no answers. To ensure the quality of the VQA data, we perform a round of self-checks using GPT-4 OpenAI (2023). Finally, we conduct an round of manual filtering to remove questions with obvious issues or those related to multiple images or patient histories. The prompt templates used are shown in Table 8.
A.3 Involved Datasets
We utilize three open-source medical vision-language datasets, i.e., MIMIC-CXR Johnson et al. (2019), IU-Xray Demner-Fushman et al. (2016), Harvard-FairVLMed Luo et al. (2024).
MIMIC-CXR Johnson et al. (2019) is a large publicly available dataset of chest X-ray images in DICOM format with associated radiology reports.
IU-Xray Demner-Fushman et al. (2016) is a dataset that includes chest X-ray images and corresponding diagnostic reports.
Harvard-FairVLMed Luo et al. (2024) focuses on fairness in multimodal fundus images, containing image and text data from various sources. It aims to evaluate bias in AI models on this multimodal data comprising different demographics.
Appendix B Evaluated Models
We evaluate four open-source Med-LVLMs, i.e., LLaVA-Med Li et al. (2023), Med-Flamingo Moor et al. (2023), MedVInT Zhang et al. (2023), RadFM Wu et al. (2023). The selected models are all at the 7B level.
LLaVA-Med Li et al. (2023) is a vision-language conversational assistant, adapting the general-domain LLaVA Liu et al. (2023b) model for the biomedical field. The model is fine-tuned using a novel curriculum learning method, which includes two stages: aligning biomedical vocabulary with figure-caption pairs and mastering open-ended conversational semantics. It demonstrates excellent multimodal conversational capabilities.
Med-Flamingo Moor et al. (2023) is a multimodal few-shot learner designed for the medical domain. It builds upon the OpenFlamingo Alayrac et al. (2022) model, continuing pre-training with medical image-text data from publications and textbooks. This model aims to facilitate few-shot generative medical visual question answering, enhancing clinical applications by generating relevant responses and rationales from minimal data inputs.
RadFM Wu et al. (2023) serve as a versatile generalist model in radiology, distinguished by its capability to adeptly process both 2D and 3D medical scans for a wide array of clinical tasks. It integrates ViT as visual encoder and a Perceiver module, alongside the MedLLaMA Wu et al. (2024) language model, to generate sophisticated medical insights for a variety of tasks. This design allows RadFM to not just recognize images but also to understand and generate human-like explanations.
MedVInT Zhang et al. (2023), which stands for Medical Visual Instruction Tuning, is designed to interpret medical images by answering clinically relevant questions. This model features two variants to align visual and language understanding Wu et al. (2024): MedVInT-TE and MedVInT-TD. Both MedVInT variants connect a pre-trained vision encoder ResNet-50 adopted from PMC-CLIP Lin et al. (2023), which processes visual information from images. It is an advanced model that leverages a novel approach to align visual and language understanding.
Appendix C Implementation Details
Following the settings of CLIP Radford et al. (2021), we adopt the same architecture and hyperparameters for the vision and text encoders. The vision encoder is a ResNet-50 He et al. (2016), and the text encoder is a bio-bert-based model Alsentzer et al. (2019). We use the AdamW optimizer with a learning rate of , weight decay of and a batch size of 32. The model is trained for 360 epochs. The reports available for retrieval are from the training set of the corresponding dataset. In our experiments, we apply cross-validation to tune all hyperparameters with grid search. All the experiments are implemented on PyTorch 2.1.2 using four NVIDIA RTX A6000 GPUs. It takes roughly 2.5 and 4 hours for fine-tuning CLIP and LLaVA-Med-1.5 7B, respectively.
Appendix D Proofs
Proof of Proposition 1: According to the definition, denotes the Med-LVLM. denotes the top retrieved contexts. The dataset is , where is the target image, is the ground-truth answer, is the target question. By the definition of ,
Therefore, can be written as the average value of a function evaluated at each data point in . Then, by combining Theorem 1, Proposition 1 and Proposition 2 of Angelopoulos et al. (2021), we finish the proof.