GSCo: Towards Generalizable AI in Medicine via Generalist-Specialist Collaboration
Sunan He, Yuxiang Nie, Hongmei Wang, Shu Yang, Yihui Wang, Zhiyuan Cai, Zhixuan Chen, Yingxue Xu, Luyang Luo, Huiling Xiang, Xi Lin, Mingxiang Wu, Yifan Peng, George Shih, Ziyang Xu, Xian Wu, Qiong Wang, Ronald Cheong Kin Chan, Varut Vardhanabhuti, Winnie Chiu Wing Chu, Yefeng Zheng, Pranav Rajpurkar, Kang Zhang, Hao Chen
Introduction
Promoted by the rapid development in Large-Language Models (LLMs) , large-scale vision-language models (LVLMs) have demonstrated impressive capabilities in various multi-modal tasks, like visual question answering and image captioning . In the field of healthcare, although the large-scale vision language models , especially generalist foundation models , achieved considerable development in multiple medical tasks, their performances are still far from the specialist counterparts. A significant challenge arises from the insufficient availability of high-quality datasets for instruction tuning, which is essential for training a generalist foundation model. To bridge this gap, previous works proposed to build the instruction-tuning dataset based on large-scale image-text pairs collected from PubMed Central. Although they have successfully generated a multitude of instruction-tuning data, there are still two constraints. Firstly, the data generation process relies exclusively on textual information (i.e., image captions) without considering the associated images, potentially leading to inconsistencies between the generated instruction data and the corresponding medical images. Secondly, the data may encounter the long-tail problem, which refers to a scenario where some rare diseases occur significantly less frequently than more common diseases. Recognizing the abundance of medical image datasets with label-level annotations such as classification and detection , we observe the potential for generating high-quality data beyond the label’s expression. In this work, we introduce a diagnosis-guided bootstrapping method. In contrast to prior works that exclusively rely on textual information for data generation, our method incorporates both image and textual information. Concretely, we utilize the well pre-trained LVLM to generate detailed medical reports consisting of findings and conclusions based on the diagnostic information. Findings enumerate observations, while conclusions summarize the final diagnosis. By implementing this method, the generated data not only ensures accuracy but also significantly enhances the richness of information within the text. Building upon the diagnosis-guided data generation approach, in this work, we develop MedDr, a large-scale, generalist foundation model for healthcare. In contrast to previous work , as illustrated in Figure 1, MedDr accommodates a broader range of medical image modalities, encompassing radiology, pathology, dermatology, retinography, and endoscopy. Comprehensive experiments indicate that MedDr achieves state-of-the-art performance across various medical downstream tasks, such as visual question answering, medical image diagnosis, and medical report generation.
Moreover, to further enhance the reliability and accuracy of the model’s responses (especially on rare or unseen diseases), we propose a simple but effective Retrieval-Augmented Medical Diagnosis strategy for the medical generalist model. Both quantitative and qualitative experiments demonstrate the effectiveness of the proposed strategy and also validate the generalization ability of MedDr, hence showcasing the significant potential for applying Retrieval-Augmented Generation (RAG) to the medical generalist model.
To summarize, our key contributions are as follows:
Diagnosis-Guided Bootstrapping. We introduce a novel approach to generate the data guided by diagnoses. Our method incorporates label-level information to effectively enhance textual content while preserving the overall accuracy of the generated text.
MedDr. We develop a generalist foundation model for healthcare, capable of handling diverse medical data modalities, including radiology, pathology, dermatology, retinography, and endoscopy. It achieves state-of-the-art performance on various downstream tasks.
Retrieval-Augmented Medical Diagnosis. We propose a simple but effective retrieval-based strategy for medical generalist models, which not only enhances the prediction accuracy of the model but also boosts its generalization performance. To our knowledge, we are the first to apply the RAG strategy to the medical vision-language generalist model.
Related Work
The success of large language models , such as GPT-4 , LLaMA-2 , and PaLM-2 , has generated interest in building vision-language models. This interest has resulted in a considerable amount of work in the general domain , including GPT-4V , PaLM-E , MiniGPT-4 , LLaVA , and InternVL . However, the development of large-scale vision-language models (generalist foundation models) in medicine remains under-explored. Currently, there are two branches of research in this area. The first branch focuses on combining a large language model with several specific vision models in medicine , such as Visual Med-Alpaca , OphGLM , and ChatCAD . These models integrate a large language model with a specialized vision model to address medical tasks beyond the language modality. The other branch combines the vision and language parts as a whole, and the two modules can be trained jointly to acquire the ability to handle broader medical scenarios . Models such as Med PaLM M , Med-Flamingo , and LLaVA-Med target training a large-scale vision language model for various medical applications across different modalities, while RadFM focuses on the medical diagnosis of 2D and 3D radiology images. Different from previous generalist foundation models, the proposed MedDr can tackle more medical image modalities while maintaining strong performances.
Data Construction in Medicine
In contrast to the general domain, which boasts a plethora of large-scale vision-language datasets such as LAION and Conceptual 12M that facilitate the development of large-scale vision-language models, the medical domain is characterized by a paucity of datasets. Predominant medical vision-language datasets, namely MIMIC-CXR , PadChest , and CheXpert , are primarily utilized for chest x-ray image diagnostics and typically comprise fewer than one million images. Other datasets, such as PMC-OA and PMC-Inline , despite their large scale, mostly derive from academic publications within PMC Central, which may not reflect well-established medical knowledge. Additionally, datasets generated by large language models, for instance, PMC-VQA , PMC-CaseReport , and the instruction-tuning dataset developed for LLaVA-Med are heavily dependent on text associated with images, potentially giving rise to discrepancies between the images and the accompanying instructions. In contrast, in this work, we propose a diagnosis-guided bootstrapping strategy that aims to harness the capabilities of vision-language models to curate a high-quality multi-modal medical dataset, thereby enhancing the training of large-scale vision-language models in medicine.
Retrieval Augmentation in Medical Vision Language Models
Retrieval augmentation involves augmenting a language model with relevant retrieved information . The use of retrieval augmentation in vision-language models has gained attention in recent years. In the medical domain, retrieval augmentation has been applied to medical VQA and report generation . RAMM retrieved medical image captions from a database sourced from PubMed Central to help improve the performance of vision question answering. eCLIP retrieved related medical reports via the given test image to enhance a large language model for report generation. MCSAM retrieved information from a cross-modal memory bank to improve the quality of report generation. However, there is a lack of research exploring the potential of retrieval augmentation in enhancing the performance of medical generalist foundation models. To address this gap, this study leverages retrieval augmentation to improve the performance of the proposed MedDr model.
Method
As depicted in Figure 2, the proposed methodology can be categorized into three distinct components: Diagnosis-Guided Bootstrapping, Medical Instruction Tuning, and Retrieval-Augmented Medical Diagnosis. The subsequent sections provide a comprehensive elaboration of these components and their respective details.
Different from previous works , which constructed the instruction tuning dataset based on image-text pairs crawled from the internet, in this work, we aim to build the dataset based on the high-quality medical image classification dataset.
As shown in Figure 3 (a), we observe that LVLMs in the general domain exhibit a comprehensive understanding of disease-related information and their associated symptoms, owing to the LLM’s extensive training on diverse corpora. However, the model encounters challenges when it comes to correlating the knowledge with concrete medical images, leading to erroneous diagnostic predictions.
Meanwhile, if we provide the model with specific disease and modality information alongside the given image, the LVLM in the general domain is also capable of generating high-quality medical reports. Figure 3 (b) shows a case of “ulcerative colitis”, where the findings enumerate the observations in the image and the impression encapsulates the conclusion.
Based on our observations, we propose a diagnosis-guided bootstrapping strategy that leverages both visual and textual information to construct the instruction tuning dataset. Specifically, we format the instruction as follows:
You are a helpful medical assistant, and your task is report generation. You are given {Modality} image and the diagnosis is {Disease}. You need to provide a medical report consisting of findings and impressions.
In contrast to previous works , which generated data from textual information only, our approach offers a distinct advantage. It facilitates the utilization of numerous label-level annotated datasets in medicine and guarantees that the generated information remains pertinent to the accompanying images. Following this method, we construct the medical report dataset encompassing diverse medical image modalities and integrate them into the training set. As illustrated in Figure 3 (c), MedDr can generate detailed medical reports without relying on explicit inputs such as modality and diagnostic information.
2 Medical Instruction Tuning
To prepare the training data, we integrate the generated data described in Section 3.1 with pre-existing data from a range of medical tasks, such as medical image diagnosis, medical report generation, and medical visual question answering. We then create specific instructions for each medical task in the model’s training process, which are detailed in Appendix A. The language modeling loss is utilized as the loss function to train the model.
3 Retrieval-Augmented Medical Diagnosis
LLMs have demonstrated significant capabilities but still encounter challenges such as out-of-domain knowledge . At inference time, we propose a retrieval-augmented medical diagnosis strategy to enhance the model’s generalization ability.
Figure 4 illustrates the proposed method, where the database is built on the training data across multiple medical tasks and modalities. Concretely, given an image-text pair, we encode the image by the vision encoder of MedDr and take the visual embedding as the key while the text is the value.
When conducting the retrieval, the visual embedding of the query image is encoded and then taken to calculate the cosine similarity between the query and keys in the constructed database. The Top- most similar items from the database are retrieved, and then their meta information is incorporated into the instruction as additional textual clues to help the model make medical decisions.
This design presents two notable advantages. Firstly, it effectively eliminates the requirement for an additional embedding module, thus mitigating the associated computational costs. By incorporating the vision encoder of MedDr as the embedding module, we can utilize the intermediate results directly as embeddings for query during inference without incurring any additional overhead. Secondly, this design facilitates the expansion of MedDr on out-of-domain data without retraining. This expansion capability enhances the model’s generalization ability, enabling it to perform effectively on data beyond its original training domain.
4 Implementation Detail
In this work, we employ InternVL , a state-of-the-art large-scale vision-language model in the general domain, as our foundation model, which contains about 40B parameters, consisting of a 6B vision encoder and a 34B Large Language model. The model is fine-tuned on both collected and generated data. The number of training samples is about 2M. Please refer to Appendix A for detailed information about the training dataset and instruction prompt. The fine-tuning recipe follows the suggestions provided by InternVL. We fix all parameters except for the LoRA component, which is composed of approximately 0.1B parameters, accounting for 0.4% of the total parameters. Meanwhile, we also leverage DeepSpeed ZeRO Stage 3 to optimize the training procedure. The model is trained on 8 NVIDIA H800 GPUs for two epochs.
Experiments
To evaluate the model’s performance more comprehensively, we select open-source large-scale vision language models in both the general and medical domains as our baseline models.
RadFM mainly focuses on the radiology modality. It consists of a 3D ViT as the vision backbone and PMC-LLaMA-13B as the LLM.
LLaVA-Med is built on LLaVA . It is fine-tuned on about 600K concept alignment samples and 60K instruction tuning samples in one day with 8 A100 GPUs.
Med-Flamingo is developed based on OpenFlamingo-9B , which can handles multiple images interleaving with texts.
InternVL is one of the most powerful open-source large-scale vision-language models in the general domain. It surpasses GPT-4V and Gemini on several multi-modal tasks.
We reproduce the above models based on their open-source checkpoint and evaluate the model using the same test data. The testing prompt is following their official implementation.
2 Visual Question Answering
Visual Question Answering (VQA) task requires the model to answer the question based on the image provided, which needs a comprehensive understanding of both image and text. We conduct experiments of the visual question answering task on four benchmark datasets, e.g., VQA-RAD , Slake-VQA , Path-VQA and PMC-VQA . The RAD-VQA and Slake-VQA datasets primarily focus on radiology data (e.g., CT, MRI, and X-ray). Following the official split, their test sets consist of 451 questions and 1061 questions, respectively. Path-VQA is a pathology VQA dataset consisting of 32,799 question-answer pairs of 7 categories, generated from 4,998 images. We utilize the official split and take 6761 questions as the test dataset. Compared with previous datasets, the PMC-VQA dataset covers a broader medical scope. It is built on the materials from PubMed and consists of more than 227K questions. Following the split of RadFM , we utilize about 90K questions as the test split. Following MultiMedEval , we evaluate the results using NLG and classification metrics.
The comparison results are shown in Table 1. RadFM incorporates RAD-VQA, Slake-VQA, and PMC-VQA in its training dataset, achieving satisfying performance on the radiology dataset. However, when faced with an out-of-domain dataset Path-VQA, the model struggles to answer relevant questions. LLaVA-Med is fine-tuned on specific datasets separately and obtains higher performance than unified models like RadFM and Med-Flamingo . On the PMC-VQA dataset, however, without fine-tuning, the model’s performance is far from satisfactory. The performance of the Med-Flamingo model is relatively lower due to the absence of relevant VQA data in the training set. As an LVLM in the general domain, InternVL achieves impressive performance on medical tasks, demonstrating the model’s generalization ability. Overall, MedDr exhibits superior performance across all datasets, even outperforming the fine-tuned LLaVA-Med model on certain evaluation metrics.
3 Medical Report Generation
The Medical Report Generation (MRG) task requires the model to list all the observations and provide a diagnosis, which is indeed a challenge for the model’s ability to capture details. We conduct experiments of medical report generation tasks on two benchmark datasets, e.g., MIMIC-CXR and IU-Xray . Following R2Gen , the test sets consist of 3858 samples and 1180 samples, respectively. Besides, we utilize both NLG metrics and model-based metrics in MultiMedEval to evaluate the models.
As shown in Table 2, it should be noted that apart from RadFM and MedDr, other models are not trained on relevant datasets. Among these “layman” models , InternVL achieves the best performance. This is because InternVL can probably understand the instructions better and generate more reasonable responses. Compared with RadFM , MedDr achieves better performance on almost all metrics, which underscores the exceptional capabilities of MedDr in comparison to the radiology specialist model.
4 Medical Image Diagnosis
Medical Image Diagnosis is one of the most important tasks in the medical domain, which requires the model to diagnose the given image within a predefined label set. We conduct experiments on nine benchmark datasets across diverse medical image modalities. VinDr-PCXR is a pediatric chest X-ray dataset comprising 9,125 samples with 15 diagnosis categories. VinDr-SpineXR is a spinal lesions detection and classification dataset consisting of 10,469 images with 13 types of abnormalities. HAM10000 is a dermatology dataset of common pigmented skin lesions, which consists of 10,015 images categorized as seven different diseases. PneumoniaMNIST is based on a prior dataset of 5,856 pediatric chest X-ray images. The task is a binary-class classification of pneumonia against normal. OCTMNIST is developed from a dataset consisting of 109,309 optical coherence tomography (OCT) images and comprises 4 diagnosis categories. ChestMNIST is based on the NIH-ChestXray14 dataset, comprising 112,120 chest X-ray images with the text-mined 14 disease labels. BreastMNIST is a breast ultrasound dataset and the sample is categorized into normal, benign, or malignant. OrganAMNIST is based on 3D CT benchmark LiTS . The images are from the center slices of the 3D bounding boxes in axial views and are classified into 11 body organs. WCE is a colon disease dataset curated from other datasets , which contains 4 kinds of diagnoses. For VinDr-PCXR and VinDr-SpineXR datasets, we follow the split of RadFM . For HAM10000 , MedMNIST , and WCE datasets, we follow the official split. Following MultiMedEval , we utilize accuracy and macro-F1 score to evaluate the results.
Experiment results are provided in Table 3. For some methods, obtaining reasonable responses for unseen modalities can be challenging. Therefore, we can not evaluate their metrics and denote the result as “N/A” in the table. VinDr-PCXR and VinDr-SpineXR are quite challenging. Even incorporating them in the training set, the performances of RadFM and MedDr are still not deemed satisfactory. MedDr achieves a better Macro-F1 score because the model’s prediction is more diverse, while RadFM always predicts frequent labels and neglects the long-tail labels. It is worth noting that InternVL achieves the second-best performance on most datasets, just lower than MedDr. This indicates that even without training on domain-specific tasks, scaling the model can enhance its performance in general. Meanwhile, the utilization of a diagnosis-based data generation strategy has enabled us to train our model on more diverse medical modalities. Consequently, MedDr exhibits superior performance in dermatology and endoscopy, surpassing other existing models by a large margin, validating the efficacy of our data generation method.
5 Retrieval-Augmented Medical Diagnosis
In this section, we explore the efficacy of our proposed RAG strategy. We conduct experiments on MedDr and Med-Flamingo , which also has the capability to handle multiple image inputs. The database is constructed based on the corresponding training set. For medical image diagnosis tasks, we retrieve the labels of the top five most similar images. For visual question answering and medical report generation tasks, we retrieve the most similar images along with their corresponding annotations. Table 4 shows the results. The “voting” column represents the results obtained through voting on the retrieved samples. For medical report generation task, we take the top- retrieved report as the prediction. Even though we only used the simple similarity-based retrieval approach, the metrics obtained through “voting” are already higher than previous results on several datasets, which shows the high quality of retrieved results.
As shown in other columns, both Med-Flamingo and MedDr have shown significant improvements in various metrics with retrieval augmentation, demonstrating the effectiveness of our proposed method. Especially, in the BreastMNIST dataset, MedDr achieves 87.8% accuracy, which is even higher than the specialist model (86.3%) in MedMNIST . It should also be noted that the BloodMNIST dataset is not covered in the training dataset. Therefore, the performances of both Med-Flamingo and MedDr without RAG are unsatisfactory. However, powered by retrieval augmentation, MedDr achieves 95.5% accuracy, which not only demonstrates the effectiveness of RAG but also validates the generalization ability of our model. Compared with “voting”, MedDr has inferior performance on some datasets. It could be due to the bias of some labels during the training process, leading to inconsistent results between the retrieved samples and the model’s decision. As for Med-Flamingo, its performance is upper-bounded by “voting”, which means it heavily relies on the retrieved results to make the final diagnosis. For example, in the medical report generation task, we find that Med-Flamingo tends to directly rephrase or even copy the retrieved samples without any modification. Overall, with retrieval-augmented medical diagnosis strategy, MedDr obtains significant performance improvement on various tasks, demonstrating the effectiveness of our method.
6 Qualitative Evaluation
In this section, we present some qualitative results to demonstrate the effectiveness of our proposed method and the superiority of MedDr. Figure 5 shows a collection of retrieved items from various datasets and modalities. We find that the majority of the retrieved items have the same label as the query image, indicating that they can be effectively considered as references for the model. However, it should be noted that a small portion of the retrieved items have labels that are not consistent with the query, which highlights the need for a comprehensive understanding of both contextual information and medical images.
Table 5 presents illustrative examples of the medical report generation tasks. On benchmark dataset MIMIC-CXR , we compare MedDr with RadFM , which is a generalist foundation model with a specialization in radiology. We find that most findings generated by RadFM only describe the normal (healthy) status of the patient, while the abnormal information is more critical for this task. In contrast, MedDr lists both normal and abnormal findings of the patient. Moreover, we also conduct the report generation on retinography images to demonstrate the generalization ability of MedDr.
Conclusion
In this work, to relieve the limitations faced by large-scale vision-language models in medicine due to the scarcity of high-quality image-text data, we propose a novel approach that leverages both image and label information to generate diagnosis-based datasets for vision-language tasks. Our approach enables the development of MedDr, a generalist foundation model for healthcare, which is capable of effectively handling diverse medical image modalities and various medical downstream tasks. Furthermore, to enhance the model’s generalization ability during inference, we introduce a simple but effective retrieval-augmented medical diagnosis strategy. Through extensive experiments on visual question answering, medical report generation, and medical image diagnosis, we demonstrate the superiority of our proposed model and strategy. Our results highlight the potential of large-scale vision-language models in revolutionizing healthcare by facilitating accurate and comprehensive analysis of medical data across various modalities.
References
Appendix A Training Dataset and Instruction Prompt
SLAKE is a bilingual radiology VQA dataset comprising 642 images and 14K questions. We only use the English part of the training split, which consists of 4,919 question-answer pairs.
VQA-RAD is a manually constructed dataset where clinicians asked naturally occurring questions of radiology images and provided reference answers. Following the official split, we use 3,064 question-answer pairs of the training set.
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images, where each question is manually checked to ensure correctness. Following the official split, we use 19,755 question-answer pairs of the training set.
PMC-VQA is a large-scale medical visual question-answering dataset built from image-text pairs from PubMed Central, covering broader medical image modalities. Following the official split, we use 152,603 question-answer pairs of the training set.
PMC-CaseReport is an auto-generated visual question-answering dataset based on the cases report papers in the PMC-Inline dataset. Following the official split, we use 254,105 question-answer pairs of the training set.
The instruction prompt for visual question answering is:
User: You are a helpful medical assistant. You are required to answer the question based on the medical image. The question is {Question}.
A.2 Medical Report Generation
MIMIC-CXR presents 371,920 chest X-rays associated with 227,943 imaging studies from 65,079 patients. Following RadFM and R2Gen , we use 337,292 cases for training.
IU-Xray is a set of chest X-ray images paired with their corresponding diagnostic reports. The dataset contains 7,470 pairs of images and reports. Following R2Gen , we use 4,730 cases from the training split.
The instruction prompt for medical report generation is:
User: You are a helpful medical assistant. Your task is report generation. You are given a chest x-ray image, and you are required to generate a summary report about the image. MedDr: {Medical Report}.
A.3 Medical Image Diagnosis
VinDr-SpineXR is a large annotated medical image dataset for spinal lesions detection and classification from radiographs. Following RadFM , we use 6,129 samples for training.
VinDr-PCXR is an open-source large-scale pediatric chest X-ray dataset for the interpretation of common thoracic diseases. Following RadFM , we use 4,585 samples for training.
VinDr-Mammo is a large-scale benchmark dataset for computer-aided detection and diagnosis in full-field digital mammography. Following RadFM , we use 6,047 samples for training.
VinDr-CXR is an open large-scale dataset of chest X-rays with radiologist’s annotations. The training set contains 15,000 scans, and 3 radiologists independently label each image. Following the official split, we use 45,000 samples for training.
CheXpert is a large public dataset for chest radiograph interpretation, consisting of 224,316 chest radiographs of 65,240 patients. Following the official split, we use 223,414 samples for training.
ChestX-ray14 is a medical imaging dataset which comprises 112,120 frontal-view X-ray images of 30,805 patients with the text-mined fourteen common disease labels. Following the official split, we use 86,524 samples for training.
PCam200 is a public pathological H&E image dataset from Patch Camelyon in 200 microns by 512 px made in the same manner from Camelyon2016 challenge dataset . Following the official split, we use 28,539 samples for training.
PAD-UFES-20 is a dermatology classification dataset consisting of 2,298 images for six different diagnostics. We use all of the 2,298 samples for training.
Dermnet consists of dermatology images of 23 types of skin diseases taken from Dermnet. Following the official split, we use 15,557 samples for training.
HAM10000 is a large collection of multi-source dermatoscopic images of pigmented lesions. Following the official split, we use 10,015 samples for training.
ISIC2020 is dataset of the SIIM-ISIC Melanoma Classification Challenge 2020. The dataset contains 33,126 dermoscopic training images of unique benign and malignant skin lesions from over 2,000 patients. Following the official split, we use 33,126 samples for training.
Kvasir is a multi-class image dataset for computer-aided gastrointestinal disease detection. Following the official split, we use 8,000 samples for training.
Kvasir Capsule is an endoscopy dataset consisting of 47,238 images with anatomical landmarks and pathological and normal findings. Following the official split, we use 47,248 samples for training.
WCE is a curated colon disease dataset based on Kvasir and ETIS-Larib-Polyp DB Dataset . Following the official split, we use 3,200 samples for training.
GastroVision is a multi-center open-access gastrointestinal (GI) endoscopy dataset that includes different anatomical landmarks, pathological abnormalities, polyp removal cases, and normal findings from the GI tract. We use all of the 8,000 samples for training.
ODIR is a structured ophthalmic database of 5,000 patients with age, color fundus photographs from left and right eyes and doctors’ diagnostic keywords from doctors. Following the official split, we use 6,392 samples for training.
Fundus1000 contains 1,000 fundus images with 39 categories. We use all of the 1,000 samples for training.
RFMiD2.0 is a multi-label dataset including around 860 retinal fundus images annotated by three eye specialists. Following the official split, we use 455 samples for training.
Retinal OCT-C8 is a large-scale dataset for ophthalmic research containing 24,000 optical coherence tomography (OCT) images that are organized into eight categories. Following the official split, we use 18,000 samples for training.
UltraBreast is a private breast ultrasound dataset that contains 45896 cases that are labeled benign or malignant.
The instruction prompt for medical image diagnosis is:
User: You are a helpful medical assistant. Your task is disease diagnosis. You are given a {Modality} image. The possible diagnoses are:{Label Set}. MedDr: {Label}.
A.4 Synthetic Datasets and Prompts
As introduced in Section 3.1, we construct a large-scale medical report dataset across diverse medical modalities. Concretely, we generate 196,760 samples in total based on the VinDr-SpineXR , VinDr-PCXR , VinDr-Mammo , VinDr-CXR , ChestX-ray14 , PAD-UFES-20 , Dermnet , Kvasir , WCE , Kvasir Capsule , ODIR , Fundus1000 and RFMiD2.0 datasets.
The instruction prompt for the diagnosis-guided dataset is:
User: You are a helpful medical assistant. Your task is report generation. You are given a {Modality} image. You need to provide a medical report consisting of findings and impressions. Findings lists the observations and impression outlines the final diagnosis. MedDr: {Medical Report}.
OpenI-Based Dataset
To augment the diversity of our training data, we also collect 245,371 image-based case studies from OpenI . We summarize the title of the case and the image caption based on the image by InternVL and obtain high-quality images and corresponding text summaries.
The instruction prompt for the OpenI-based dataset is:
User: You are a helpful medical assistant. You are given a medical image. You are required to generate a detailed description and analysis from a medical perspective based on the image. MedDr: {Description}.