Advancing Multimodal Medical Capabilities of Gemini
Lin Yang, Shawn Xu, Andrew Sellergren, Timo Kohlberger, Yuchen Zhou, Ira Ktena, Atilla Kiraly, Faruk Ahmed, Farhad Hormozdiari, Tiam Jaroensri, Eric Wang, Ellery Wulczyn, Fayaz Jamil, Theo Guidroz, Chuck Lau, Siyuan Qiao, Yun Liu, Akshay Goel, Kendall Park, Arnav Agharwal, Nick George, Yang Wang, Ryutaro Tanno, David G. T. Barrett, Wei-Hung Weng, S. Sara Mahdavi, Khaled Saab, Tao Tu, Sreenivasa Raju Kalidindi, Mozziyar Etemadi, Jorge Cuadros, Gregory Sorensen, Yossi Matias, Katherine Chou, Greg Corrado, Joelle Barral, Shravya Shetty, David Fleet, S. M. Ali Eslami, Daniel Tse, Shruthi Prabhakara, Cory McLean, Dave Steiner, Rory Pilgrim, Christopher Kelly, Shekoofeh Azizi, Daniel Golden
Introduction
Medical data from diverse sources like biobanks, electronic health records, medical imaging, wearables, biosensors, and genomic sequencing are enabling the development of multimodal AI solutions that can better capture the complexity of human health and disease (Acosta et al., 2022). While AI in medicine has primarily focused on narrow tasks with single input and output types (Rajpurkar et al., 2022), recent advances in generative AI show promise in addressing multimodal, multi-task challenges in medical settings (Moor et al., 2023a, b).
The emergence of large language models (LLMs) and large multimodal models (LMMs) such as Flamingo (Alayrac et al., 2022), PaLI (Chen et al., 2022), GPT-4 (Achiam et al., 2023), GPT-4v (OpenAI, 2023), PaLM (Anil et al., 2023; Chowdhery et al., 2023), LLaMA (Touvron et al., 2023), LLaVa (Liu et al., 2023, 2024a), and Mistral 7B (Jiang et al., 2023) that promise significantly enhanced context length and improved multimodal capabilities suggests that the realization of highly complex multimodal reasoning across various medical data will soon be achievable. These advancements have catalyzed the expansion of LLMs specifically designed for medical domains, such as Med-PaLM and its successor Med-PaLM 2 (Singhal et al., 2023a, b), Clinical Camel (Toma et al., 2023), MedAlpaca (Han et al., 2023), BioMistral (Labrak et al., 2024), sc-GPT (Cui et al., 2024), and others. Going beyond text alone, recent works have extended the capabilities of these base multimodal models by building models that cover various medical imaging modalities like Med-PaLM M (Tu et al., 2024), Med-Flamingo (Moor et al., 2023b) as well as those that focus on a specific imaging domain, such as radiology (Tanno et al., 2024; Thawkar et al., 2023; Hyland et al., 2023; Xu et al., 2023; Hamamci et al., 2024) and histopathology (Sun et al., 2024; Ikezogwo et al., 2024; Lu et al., 2024).
The release of the Gemini models (Gemini Team, Google, 2023; Google, 2024), with their advanced multimodal capabilities and breakthroughs in long-context understanding, marked a significant step forward in multimodal reasoning. Given its inherent human focus, medicine is a field in which advanced multimodal systems like Gemini are expected to be transformative (Acosta et al., 2022). Evaluations have already started to evaluate the base performance of these newer multimodal models (Pal and Sankarasubbu, 2024). However, the true potential of multimodal foundation models in the medical field remains largely underexplored due to the complexity of optimizing for problems in this field (Moor et al., 2023a; Rajpurkar et al., 2022) and a lack of diverse and meaningful evaluations that are grounded in clinical use cases (Royer et al., 2024; Zhang et al., 2023a; Fleming et al., 2023). To better understand the nuances of model capabilities and limitations, it is necessary to optimize multimodal models for a diversity of relevant clinical applications and rigorously evaluate them on appropriate clinical datasets.
This report details our efforts in exploring Gemini’s capabilities across a range of challenging multimodal medical tasks. Our evaluation benchmarks include 2D and 3D radiology images, histopathology patches, ophthalmology images, dermatology images, and genetic risk scoring. Our benchmark suite includes both open benchmark datasets and our own curated datasets. Open benchmark datasets have the advantage of being established and enabling direct comparison to others’ work, but they are often limited or methodologically flawed, leading to results that can overstate performance. For the custom benchmarks that we introduce, we have prioritized high quality metrics that are closely correlated to clinical utility. In particular, we focused on expert human evaluations for quantifying performance on CXR and CT report generation and on open visual question answering (VQA) questions from VQA-Rad. Additionally, we compared Med-Gemini to previous work or to the non-medically tuned version of Gemini where possible.
Where we believed it was helpful, we have proactively improved the quality of certain open benchmarks. This includes updating and correcting erroneous labels (such as MIMIC-CXR-JPG classification labels), extending the task scope of datasets (such as introducing VQA question/answer pairs for MIMIC-CXR), and refining data splits to remove train-test contamination (such as PAD-UFES-20 and VQA-Rad). We hope to release these improvements publicly soon.
In this report, we expand the fine-tuned family of models, Med-Gemini, specifically focusing on medical imaging and genomics. The models described here were tuned on a dataset of 7 million samples obtained from 3.7 million medical images and cases, spanning medical image classification, VQA, report generation, and genomic risk prediction, detailed in Section 2. Importantly, this dataset includes mostly free text paired with medical data, which eliminates the need for expensive expert labeling of the training data. We intentionally explore both medical image-based tasks and also the non-image-based task of polygenic risk prediction in order to evaluate the potential of Med-Gemini beyond imaging and in the crucial medical domain of long term risk prediction. Our findings demonstrate that LMMs have significantly advanced over the past year and are able to perform an increasing range of challenging tasks. Our key contributions are summarized as follows:
Med-Gemini: A family of generalist medical AI models fine-tuned from Gemini, capable of performing a diverse set of medical tasks including medical image classification, VQA, report generation, and genomic risk prediction. Med-Gemini extends Gemini’s capabilities to include interpretation of diverse medical data, including both genomics and 2D and 3D medical images. Additional capabilities of Med-Gemini are described in “Capabilities of Gemini Models in Medicine” by Saab et al. (2024).
Clinically-relevant benchmarking: We evaluate Gemini and Med-Gemini on a comprehensive set of clinically relevant benchmarks including 22 datasets across five different tasks and six distinct medical image modalities. Our evaluation suite includes eight out-of-distribution datasets to assess generalization capabilities of this new family of models. Our assessments primarily consist of automated metrics but we rely on expert human evaluation for tasks where expert human judgment is critical, namely chest X-ray and computed tomography (CT) report generation and radiology VQA on open questions in the VQA-Rad dataset.
Promising or best-in-class performance in several clinically relevant tasks: Med-Gemini demonstrates best in class performance on chest X-ray and CT report generation and chest X-ray classification. Med-Gemini can also be used to predict disease and mortality risk more accurately than a standard linear polygenic risk score (PRS) based approach. Med-Gemini approaches the performance of models trained using orders of magnitude more training examples on dermatology, histopathology, and ophthalmology image classification and demonstrates competitive performance across several VQA tasks across pathology and radiology.
Datasets
Many different public and private datasets were used in the training and evaluation of Med-Gemini. All datasets were de-identified. Open datasets were used in accordance with their existing licenses and private datasets were used with permission and appropriate licenses.
Datasets were split into train, validation, and test sets by patient identifier when available. When patient identifiers were not available for a dataset, we ensured that there was no case or image overlap between splits.
MIMIC-CXR contains 377,110 images from 65,379 patients, with de-identified free-text reports describing the images (Johnson et al., 2019a, c; Goldberger et al., 2000). This dataset is the largest public chest X-ray dataset, acquired in the emergency department of Beth Israel Deaconess Medical Center in the US. For each patient, there are multiple views and a corresponding report labeled for 13 common radiological conditions using the CheXpert labeler (Irvin et al., 2019) or with “no finding” if no condition is present. Available labels include atelectasis, cardiomegaly, consolidation, edema, enlarged cardiomediastinum, fracture, lung lesion, lung opacity, pleural effusion, pleural other, pneumonia, pneumothorax, support devices, and no finding. We used the MIMIC-CXR training set (237,912 images) to fine-tune Gemini as described in Section 3 and detailed in Table 1. We further employed the test cases of MIMIC-CXR as a benchmark for multiple evaluation tasks including classification, report generation and VQA. For the report generation task, we used the chest X-ray image corresponding to the frontal view (anterior-posterior or posterior-anterior) to generate the Findings and Impression sections, similar to prior works (Tanno et al., 2024). For cases where no frontal view was available, we excluded them from our evaluation. For the VQA task, we utilized the condition-dependent VQA dataset (e.g. pleural effusion presence/location/severity) introduced in Xu et al. (2023), which we will make publicly available soon. In addition, we have used radiologist-adjudicated updated labels for findings that are also planned to be released soon (Park et al., 2024).
This dataset consists of digital X-ray images of knee joints which were collected from hospitals and diagnostic centers. Original images are 8-bit grayscale. Each knee X-ray image was manually annotated by two medical experts following the Kellgren and Lawrence system for classification of osteoarthritis (Gornale and Patravali, 2020). There are a total of 1,633 unique images in this dataset, and we utilized 1,469 images in our training.
The dataset consists of 1,373 patients, 1,641 skin lesions, and 2,298 images. Skin lesion images were collected from various smartphones and exhibit variations in resolution, size, and lighting conditions. The dataset was acquired in collaboration with the Dermatological and Surgical Assistance Program (PAD) at the Federal University of Espírito Santo (UFES-Brazil) (Pacheco et al., 2020). PAD is a non-profit program offering free skin lesion treatment, particularly to those who cannot afford private care. The dataset includes images of six different skin lesion types and diagnostics: three skin diseases and three skin cancers. These include basal cell carcinoma (BCC), squamous cell carcinoma (SCC), actinic keratosis (ACK), seborrheic keratosis (SEK), Bowen’s disease (BOD), melanoma (MEL), and nevus (NEV). We randomly split the dataset into train (corresponding to 90% of total samples) and test (10% of total samples). We utilized 2,047 training samples from this dataset in our training corpus when fine-tuning Gemini as described in Section 3 and detailed in Table 1. We intend to publicly release our dataset split soon.
We utilized CT volumes from the validation subset of the NLST dataset (NLST, 2014) and processed them through a lung cancer screening system (Kiraly et al., 2024) to create a dataset of the most salient 2D slices from CT volumes. Captions were assigned based on the scr_group label in the NLST participant dictionary: values of 1 and 2 were considered as having nodules whereas a value of 3 were considered as not having a nodule. We then selected a total of 2,199 slices consisting of 1,324 studies with nodules and 875 without nodules for further analysis from the previously defined validation set (Ardila et al., 2019). The lung cancer screening system generated captions for each slice based on the first detected region. Slices without nodules received simpler descriptions, for example: “An axial CT slice of the middle lungs with no nodules.” For slices containing nodules, the captions included: location of the nodule (left or right lung), suspicion level for malignancy, and estimated size in millimeters based on the screening system’s output. We then split the slices into training and validation evenly based on the presence of nodules, allocating 80% for training, including 1,759 image and caption pairs, and 20% for validation, including 440 image and caption pairs. All 2D slice images were set to a window of .
This dataset is a large bilingual (English and Chinese) VQA dataset meticulously annotated by experienced physicians (Liu et al., 2021). It offers 642 images with 14,028 question-answer pairs in three imaging modalities (i.e. CXR, CT, MRI). Slake-VQA includes various areas of radiology, covering human body regions like the brain, neck, chest, abdomen, and pelvic cavity. The dataset comprises 9,849 VQA samples for training, 2,109 for validation, and 2,070 for testing. Questions are diverse, including both open-ended (free-form) and closed-ended (yes/no) formats. They probe various image aspects such as plane, quality, position, organ, abnormality, size, color, shape, and related medical knowledge. We used only English-language examples from the official splits, which included 4,919 training, 1,053 validation, and 1,061 test examples.
This is a dataset of question-answer pairs on pathology images (He et al., 2020). The dataset includes both open-ended questions and closed-ended (yes/no) questions and is built with automated methods using two publicly-available pathology textbooks and a publicly-available digital library. The dataset includes 32,632 question-answer pairs on 4,289 images. The official training, validation, and test splits contain 19,654, 6,259, and 6,719 QA pairs. We leveraged the official train and test sets for training our model and evaluating its performance, respectively.
The VQA-Med-2019 dataset offers a collection of medical images and associated question-answer (QA) pairs for model training and evaluation (Ben Abacha et al., 2021). It includes a training set with 3,200 medical images and 12,792 QA pairs, a validation set with 500 medical images and 2,000 QA pairs, and a test set containing 500 medical images and 500 questions. For the purpose of this report, we removed all images that overlapped with VQA-Rad (Lau et al., 2018) to avoid contamination. This resulted in 12,664 QA pairs used in training.
Genetic factors play a significant role in an individual’s risk of developing various diseases. In this work, we used UK Biobank (Bycroft et al., 2018), a resource of nearly 500,000 de-identified individuals with genetic, lifestyle, and health information, to develop a task that takes as input an embedding of an individual’s genomic data and uses it to predict an individual’s status for various broad health outcomes. We extracted a set of 432,090 samples of European genetically inferred ancestry with genomic data passing quality control thresholds and split it randomly into train, validation, and test splits containing 60%, 20%, and 20% of the samples, respectively. Following best practices for polygenic risk prediction, we avoided including individuals who were genetically similar in two different data splits (Choi et al., 2020).
PMC-OA is a medical dataset with image-caption pairs collected from PubMedCentral’s OpenAccess subset. Using the method described in Zhang et al. (2023b), we retrieved 3,110,109 scientific papers containing 15,505,259 image-caption pairs. To ensure meaningful analysis, we filtered for image-caption pairs containing at least one photographic image (e.g., excluding images corresponding to data figures), resulting in a final dataset of 2,246,656 image-caption pairs.
1.2 Private datasets
Pathological examination of tissue samples is crucial for effective diagnosis and treatment planning. Data from nine tasks across six tissue types from prior work (Lai et al., 2023) were used in our training set (Table A.8). Multi-class annotation masks were used for both sampling image patches from whole-slide images as well as generating captions. Patches of size 256 256 were sampled from whole slide images in a class-balanced manner. For each of the nine tasks, up to 10,000 image patches were sampled for three different magnification levels (2, 1, and 0.5 microns-per-pixel), resulting in 207,603 unique patches. Patch-level captions were created via prompting of a large language model (Gemini Pro) with inputs including structured slide-level metadata as well as patch-level annotation labels. Multiple captions per class for each task were generated and then manually reviewed to ensure an appropriate level of detail and accuracy, resulting in 5–7 captions per class across tasks. Combining the sampled patches with our curated captions resulted in 1,550,976 image-text pairs for fine-tuning. For examples of some of our curated captions corresponding to the annotation labels, see Table A.7.
Diabetic retinopathy is the leading cause of blindness in the working-age population of the developed world. We used the de-identified dataset from EyePACS Inc. (Cuadros and Bresnick, 2009) and converted diabetic lesion-level presence labels to captions. The lesions considered were microaneurysms, hemorrhages, hard exudates, panretinal photocoagulation (PRP) scars, neovascularization of the disc and neovascularization elsewhere. For caption conversion, if a given image has lesion presence, for example, microaneurysm and hemorrhage, the associated generated caption was “microaneurysm is present, hemorrhage is present.” For healthy eyes, we use “no diabetic retinopathy related lesion” as the caption. 12,976 images with lesions and 3,000 healthy eye images were used to construct the dataset.
A comprehensive dataset comprising 753,247 CT studies with associated radiology reports from 615,384 patients was obtained from three major hospital regions in the United States. These CT studies included head/neck, chest, heart, abdominal, spine, and extremity regions imaged with and without contrast. To ensure robust evaluation, we employed a patient-level random split for training, validation, and testing. The data was divided into 70% for training, 15% for validation, and 15% for testing on the patient level. After an ingestion process this resulted in 657,719 training volumes and a total of 23,649 validation volumes. Due to the reliance on expert evaluation for report generation, a subset of 92 non-contrast head/neck CT volumes from unique patients in the test set was used for model assessment. Volumes were prepared as described in Section 2.3. Only axial image volumes containing more than 10 slices were included in the prepared data and the volume within the study with the most axial slices was selected for inference.
We also carefully processed the existing dataset to create an extra 2D CT slice dataset specifically tailored for training our 2D model. This involved filtering radiology reports for specific series and image numbers, selecting the correct images, and windowing them to a window of . To ensure that the text pertained directly to the CT slice in question and was a comprehensive description of it, captions were generated by combining the sentence of the report referencing the image along with the following sentence. This process resulted in a dataset of 4,009 images consisting of 3,207 training and 802 validation examples, primarily focused on CT studies of the abdomen and pelvis.
The CXR-US2 dataset corresponds to the training set of US1 in Xu et al. (2023). This dataset consists of 132,680 frontal chest X-ray images from 12,988 patients taken at an academic medical center in Illinois, USA. Further descriptive statistics can be found in Xu et al. (2023).
2 Held-out datasets for evaluation and benchmarking
Beyond the test sets associated with our training datasets (described above), we also utilized multiple held-out and out-of-distribution (OOD) datasets.
The CheXpert dataset is similar to the MIMIC-CXR dataset and consists of 224,316 chest X-ray images (both frontal and lateral views) from 65,240 patients (Irvin et al., 2019). It labels 14 distinct thoracic conditions, including “No Finding”. The original CheXpert dataset contains positive, negative, uncertain and unmentioned labels. During evaluation we considered the “unmentioned” label as negative and included only chest X-rays depicting frontal views.
The VQA-Rad dataset (Lau et al., 2018) comprises 315 radiology images sourced from CT, MRI, and X-ray scans, and it encompasses three anatomical regions including the head, abdomen, and chest. This dataset includes a wide array of question types, spanning 11 distinct categories, such as modality, plane, organ system, abnormality, and more, where 58% of the question-answer pairs are designed to be closed-ended (yes/no or limited choices), while the remaining 42% are open-ended.
The standard and official splits of the dataset feature 1,797 QA pairs for training and 451 for testing purposes. However, due to contamination of images included in both training/test in the original dataset release, we constructed a new, non-overlapping test and tuning split for the subset of chest X-ray images and associated question-answer pairs, first described in Xu et al. (2023). In this study, we went one step further and created new image-disjoint splits of train, validation and test sets for all three image types. We ensured that the previous X-ray-only validation and test sets were subsets of the new validation and test sets, respectively, thus enabling comparisons with the ELIXR model (Xu et al., 2023) on the new test set, as no former validation examples were included in the new test set. Aside from these constraints, we sampled in a manner that roughly equalizes both the ratio of closed to open question-answer pairs for each anatomical region, see Table A.5, as well as the distribution across the 11 different question types across the three splits, see Table A.6 in Section A.1.3. Henceforth, we refer to this new three-way split as the “balanced split”. In total, the balanced VQA-RAD dataset split, which we will make publicly available soon, comprises 2,248 pairs of questions and answers, encompassing 1,299 closed-ended questions and 949 open-ended questions.
The ChestX-ray14 dataset (Summers, 2019) is a comprehensive medical imaging dataset containing 112,120 frontal-view chest X-ray images from 30,805 unique patients. ChestX-ray14 builds upon the ChestX-ray8 dataset (Wang et al., 2017), expanding the number of labeled diseases to fourteen common thoracic pathologies, including Atelectasis, Consolidation, Infiltration, Pneumothorax, Edema, Emphysema, Fibrosis, Effusion, Pneumonia, Pleural thickening, Cardiomegaly, Nodule, Mass, and Hernia. Because these labels are automatically derived using NLP techniques and therefore contain inherent uncertainty, we restricted our evaluation to a subset of 1,962 cases focusing on three radiologist-adjudicated conditions (Majkowska et al., 2020), namely lung opacity, pneumothorax, and fracture.
This dataset utilizes histopathology images from The Cancer Genome Atlas (TCGA), for which different study types correspond to different cancer types with additional information via portal.gdc.cancer.gov. The patches from this dataset are sampled from 2,952 training slides, 1,466 validation slides, and 1,489 test slides across ten (10) distinct TCGA study types: BLCA, BRCA, COAD, HNSC, KIRC, LIHC, LUAD, LUSC, OV, and STAD (Lai et al., 2023). We used the test set as an out-of-distribution dataset to evaluate our model’s ability to generalize to different histopathology-related tasks.
2.2 Private datasets
This is a private research dataset of a similar scale as MIMIC-CXR, which we refer to as IND1 (Nabulsi et al., 2021). This dataset comprises 263,021 de-identified frontal chest X-rays (digital and scanned) along with their corresponding reports. The X-rays were collected from five regional centers (Bangalore, Bhubaneswar, Chennai, Hyderabad, and New Delhi) across a large hospital group in India between November 2010 and January 2018 (Ahn et al., 2022). We used the same test set as (Tanno et al., 2024), and following their framework, 300 of those cases are used for human evaluation.
The TTH tissue type dataset, introduced by Weng et al. (2019) and Lai et al. (2023), represents a patch-level tissue type classification task. This internal dataset comprises 17,319 training slides, 6,488 validation slides, and 6,719 test slides, encompassing a total of 16 distinct tissue types. These tissue types include Appendix, Breast, Cervix, Colon and Rectum, Fallopian Tube, Gallbladder, Liver, Lymph Node, Ovary, Placenta, Prostate, Skin, Thyroid, Upper GI, Uterus, and Vas Deferens. We used the test set to evaluate generalization of our model to different histopathology tasks.
3 Data preprocessing
When available, images acquired in DICOM format were used to directly create examples for training and inference. In the case of X-rays, raw pixel data were extracted from the DICOM image pixel data, and the look up table (LUT, part of the DICOM metadata) was subsequently applied. If a DICOM file contained multiple LUTs, we used the first LUT entry. If the window width and window center were defined, these were also used for preprocessing. The final pixel data were re-scaled to the full range of $$ for the 16-bit PNG format. X-ray images in a preprocessed format were taken as is.
All 3D CT volumes were derived from DICOM images. Only axial slices were used to establish a standardized anatomical perspective. Slices were sorted based on the Image Position (Patient) attribute and used to compute slice spacing and reconstruct volumes. Subsequently, images were clipped with a Hounsfield Unit (HU) range of to cover a full spectrum of densities (e.g. the typical window/level values of brain, soft tissues) and then scaled to . Finally, tricubic interpolation was used to resample all images to a voxel spacing of 0.7mm 0.7mm 1.4mm, ensuring consistent uniform resolution for accurate comparative analysis.
A genomic featurization for an individual consists of polygenic risk scores (PRSs) for 7,415 traits. Each PRS estimates the genetic risk of the individual for a particular disease or trait, calculated by aggregating the estimated effects of many common variants associated with the condition. Each PRS was computed using genome-wide association study summary statistics computed by the Pan-UKB Consortium (Pan-UKB team, 2020). These genomic features were then converted to images by projecting the PRSs into patch-aligned squares of 8 8 pixels with values between $$. The 3 RGB channels of the images were used to stack 3 different p-value thresholds of the projections. The PRS features were obtained from the genetic information of 314,540 individuals of European genetically inferred ancestry from the UK Biobank (Sudlow et al., 2015; Bycroft et al., 2018).
To create training and evaluation labels, we selected eight in-distribution health outcomes which have strong heritability (i.e. genetic information plays an important role in influencing susceptibility (Visscher et al., 2008)), span multiple organ systems, and are challenging to predict from polygenic risk scores alone: coronary artery disease, stroke, type 2 diabetes, glaucoma, chronic obstructive pulmonary disease (COPD), rheumatoid arthritis, major depression, and all-cause mortality. Additionally, to assess model generalization, we selected six out-of-distribution (OOD) health outcomes that share genetic correlation with one or more of the in-distribution health outcomes: hypertension, hypercholesterolemia, atrial fibrillation, diabetic retinopathy, asthma, and pneumonia (Table A.10).
Patches with initial size of 256 256 pixels were sampled from whole slide images using multi-class annotation masks in a class-balanced manner across three different magnification levels (2, 1, and 0.5 microns-per-pixel).
Images from all 2D datasets were uniformly resized to 768 768 pixels, preserving aspect ratio with padding, with pixel intensities scaled to $$. This ensured image resolution would be high enough for the fine-grained detail of medical images. For text, we used the native Gemini SentencePiece tokenizer (Kudo and Richardson, 2018; Gemini Team, Google, 2023) without modification.
Modeling Methodology
Gemini builds upon the robust foundation of Transformer decoders (Vaswani et al., 2017; Parmar et al., 2018), offering significant architectural and optimization enhancements for efficient, stable large-scale training (Gemini Team, Google, 2023; Barham et al., 2022). This equips Gemini with exceptional natural language understanding and text generation capabilities. Of particular interest for medical data processing, Gemini’s multimodal design draws inspiration from foundational Google research on Flamingo (Alayrac et al., 2022), CoCa (Yu et al., 2022; Yan et al., 2022) and PaLI (Chen et al., 2022), enabling enhanced multimodal understanding and reasoning.
Gemini handles video understanding by encoding frames as a sequence within its large context window (Gemini Team, Google, 2023; Google, 2024). This allows seamless integration of video frames, multi-slice images, text, or audio inputs. The model even supports variable input resolutions, enabling it to prioritize computational resources for tasks requiring high-resolution analysis. Gemini 1.5 specifically is a mid-size model with a context window of up to 1 million tokens and performance on par with the largest Gemini model, 1.0 Ultra. Given this exceptional efficiency, we chose to finetune Med-Gemini from Gemini 1.5.
2 Multimodal fine-tuning
Three custom versions of the Gemini 1.5 Pro vision encoder were trained for 2D modalities, 3D modalities, and genomics. In our initial experiments, we found that custom vision encoders for each type of data format performed better than a single vision encoder for all data formats. Furthermore, fine-tuning the vision encoder as well as the language component in Gemini led to significantly better visual understanding in comparison to a model that used the native vision encoder of Gemini 1.5 Pro models. From these three custom vision encoders, we trained three specific variants of Med-Gemini which we refer to as Med-Gemini-2D, Med-Gemini-3D, and Med-Gemini-Polygenic. Notably, Med-Gemini-2D includes all conventional medical images that are encoded in 2D (e.g. chest X-ray, CT slices, pathology patches), Med-Gemini-3D was built on top of Med-Gemini-2D and handles 3D medical data (e.g. CT), and Med-Gemini-Polygenic was trained for a novel image encoding derived from non-image features (e.g. genomics). For all three model variants, fine-tuning was framed as a captioning or VQA task.
All 2D modalities were fine-tuned together using the training mix described in Section 2 and Table 1 to create Med-Gemini-2D. The 2D modalities used for fine-tuning included the described radiology, pathology, dermatology, and ophthalmology images.
To interpret 3D medical data, we leveraged the video understanding capabilities of Gemini (Google, 2024). Use of the Gemini video encoder allows Med-Gemini-3D to process multiple 2D slices, replacing the time axis with the depth dimension, with computed tomography (CT) as our example modality. This 3D fine-tuned model can then synthesize information across a series of 2D slices to generate radiology reports. Use of this video encoding capability will permit analysis of other volumetric and time-series medical data (e.g. MRI, ultrasound) in the future.
Genomics “images” (polygenic risk scores (PRS) projected into 2D, see Section 2.3) were included in the mixture of datasets used to fine-tune the Med-Gemini-Polygenic vision encoder, and were trained to predict eight broad health outcomes (coronary artery disease, stroke, type 2 diabetes, glaucoma, chronic obstructive pulmonary disease, rheumatoid arthritis, major depression, and all-cause mortality) in a captioning task.
To optimize the instruction-following capabilities of the fine-tuned Med-Gemini even further, we subsequently employed an instruction-tuning phase. In this phase, we fine-tuned Gemini 1.5 Pro on a curated collection of multimodal data consisting of carefully crafted instruction and response pairs. By exposing the model to these examples, we refined its ability to not only understand the content of medical images and signals, but also to follow nuanced instructions and generate tailored outputs.
3 Model training and inference infrastructure
Like its predecessor Gemini 1.5 Pro and all other Gemini models, Med-Gemini was trained on large-scale Google TPUv4 accelerator pods spread across multiple data-centers. This training setup significantly scales up from our previous flagship PaLM family (Chowdhery et al., 2023). The Gemini architecture ensures efficient serving on TPU accelerators at scale. For detailed information on training and serving Gemini models, see (Gemini Team, Google, 2023; Google, 2024).
Evaluation and Results
The following sections explore in detail how Gemini and Med-Gemini perform across various modalities and tasks in the medical field. Due to restrictions in our data licenses, our evaluation was limited to internal models. Our evaluation leveraged a robust dataset suite encompassing 22 datasets across four different clinically relevant tasks (report generation, VQA, classification, risk prediction). This dataset includes eight out-of-distribution datasets to assess generalization and spanned seven distinct medical image modalities. An overview of the evaluation datasets is provided in Table 2. The total number of evaluation samples across these datasets exceeded 40,000.
To rigorously evaluate Med-Gemini’s in-distribution and out-of-distribution performance, we employed a comprehensive medical image classification benchmark. This benchmark encompassed diverse modalities: skin lesion classification, chest X-ray classification, histopathology patch classification, and fundus image classification. We approached classification as a generative multi-choice task for zero-shot classification (no supporting example in prompt) and linear probing for label-efficient setups. This design allowed for a thorough assessment of Med-Gemini’s robustness and adaptability across various medical imaging domains.
Our chest X-ray image classification evaluation focused on two key classification scenarios. First, we considered multi-label classification for the presence of each of five types of frequently occurring conditions: atelectasis, cardiomegaly, consolidation, pulmonary edema, and pleural effusion. This follows the suggestions from Tanno et al. (2024); Irvin et al. (2019); Azizi et al. (2021, 2023). Second, we performed binary classification for all images as either normal or abnormal, based on the CheXpert “no finding” label for frontal chest X-rays (Irvin et al., 2019). These two scenarios are used consistently across MIMIC-CXR and our out-of-distribution dataset CheXpert (Irvin et al., 2019). In addition, for ChestX-ray14 (Wang et al., 2017; Summers, 2019) we focused the evaluation of our model on three specific conditions including lung opacity, pneumothorax, and fracture. For all test tests, we only included images that were frontal view (i.e. view position “AP” or “PA”). In addition, for MIMIC-CXR it is required the original report to contain a “Findings” section that could be extracted via regular expression matching.
We manually explored the validation set, prompting each model either with a multi-select prompt for all labels (i.e.,5 conditions plus normal/abnormal) at once, or multiple binary Yes/No prompts for each label separately, and found binary prompts to yield better Macro F1 results for Med-Gemini for all labels, and slightly better results for Gemini Ultra except for predicting normal/abnormal. The prompts used for evaluation are listed in Section A.1.2. Answers were generated using nucleus sampling with a temperature of 0.0, a top_p of 0.75 and an output token limit of 200. Generated answers were normalized and matched against “yes”/“no” ground truth strings, which directly corresponded to 1.0 and 0.0 label values for all evaluation data sets. For multi-label, multi-class scenarios, we evaluated the average accuracy using the class-weighted F1 score. Details of the metrics used can be found in Section A.3. The MIMIC-CXR labels were revised based on a selective review of flagged reports by board-certified radiologists. See Section A.1.1 for more details about the revised MIMIC-CXR labels. For the MIMIC CXR evaluations, we excluded case/condition combinations with an “uncertain” (-1.0) or no label (blank), except for the “No Findings” condition, where all cases with a non-positive (1.0) label were considered negative. Classification results using data-efficient learning are described separately in Section A.2.1.
Table 3 shows the comparison of the performance on the chest X-ray classification task between Med-Gemini and Gemini Ultra for in- and out-of-distribution datasets. Our medically tuned model outperformed Gemini Ultra across most labels on the in-distribution MIMIC-CXR dataset. Notably, we demonstrated significantly stronger performance on the normal/abnormal classification despite using a multi-select prompt for Gemini, which specified all the “abnormal” conditions (and yielded better results on the validation set than a dedicated binary prompt), versus a very short normal/abnormal binary prompt for Med-Gemini. However, on the more challenging out-of-distribution datasets (CheXpert and ChestX-ray14), performance is varied. Med-Gemini excels in some tasks such as cardiomegaly or pleural effusion detection on CheXpert, while lagging in others like fracture detection in ChestX-ray14, which is a strong minority class there. These results suggest room for improvement in handling significant domain shifts.
We evaluated the image embeddings of Med-Gemini-2D via linear probing on the 11 tasks from Lai et al. (2023) and summarized in Table A.8. Med-Gemini was fine-tuned on data corresponding to 9 of these tasks (in-distribution), while 2 tasks were held-out (out-of-distribution). Together, these evaluation tasks cover a total of 17 tissue types, 12 cancer types, and several different types of classification tasks (e.g.,tumor identification, grading, subtyping) across 3 magnifications. Linear probing was done as in Lai et al. (2023): briefly, a logistic regression model with L2-regularization was fit for each task, and task-specific regularization weights and magnifications were selected using the validation sets. Linear probe metrics on the test sets were calculated using 5,000 patches with logistic regression models trained on embeddings from the 10,000 train set plus 5,000 validation set patches. Confidence intervals for macro-averaged AUCs were computed via blocked bootstrap (blocking on slides) with 10,000 replicates. For comparison, we also report performance with embeddings from an ImageNet21k-based ViT-S/16 model trained using the AugReg method (Steiner et al., 2021), embeddings produced by the vision encoder in Gemini Ultra, and embeddings from a histopathology-specialized model trained via self-supervision (Lai et al., 2023) (PathSSL). Results are reported in Figure 2. While the PathSSL embedding model is specialized to the histopathology domain, the image embeddings in Med-Gemini-2D achieved comparable performance while also demonstrating strong results across multiple other clinical domains.
Med-Gemini-2D achieves competitive classification accuracy using just dermatological images alone as input, and does not rely on metadata (e.g. patient demographics, lesion symptoms, living conditions). Such metadata, while provided in PAD-UFES-20, are not always readily available in clinical settings. We note that our evaluation is not directly comparable to Med-PaLM M (Tu et al., 2024) since (a) Med-PaLM M inputs an additional 14 clinical attributes and (b) we created different train and test splits to remove patient overlap between splits in the original Med-PaLM M work (we hope to publicly release these updated splits soon). To establish context for model performance, we compared Gemini performance with Derm Foundation (Google, 2024) which is a specialized dermatology model developed by Google. We trained linear probing classifiers on top of the Derm Foundation embeddings in this comparison.
We utilized the following three metrics for evaluation. (1) Weighted-AUC (by class prevalence): We first extracted the embedding outputs from Med-Gemini-2D’s image encoder, Gemini Ultra’s image encoder and Derm Foundation. Then we individually trained linear probes on top of the embeddings on the entire PAD-UFES-20 train split to classify 6 skin lesion types, and computes their weighted-AUC on the test split. (2) Weighted-F1 (by class prevalence): For Med-Gemini-2D and Gemini Ultra, we extracted classification prediction based on string matching from the model output. For Derm Foundation, we took the argmax of the linear probe prediction as the predicted class. (3) Accuracy: We computed the classification prediction follows the same method as in weighted-F1.
Table 4 shows the performance comparison. From the AUC linear probing results, we can see that both Gemini Ultra and Med-Gemini-2D produced robust embeddings for skin lesion classification and had on-par performance when compared with the specialized Derm Foundation model. From F1 and accuracy, we can see that the fine-tuning in Med-Gemini-2D improved the LLM’s understanding of the embedding space for skin lesions and achieved performance close to the Derm Foundation model.
We evaluated the performance of Med-Gemini-2D on four ophthalmology classification tasks. First, we approached the identification of three common Diabetic Retinopathy (DR) lesions, more specifically hard exudates, hemorrhage, and panretinal photocoagulation (PRP) scars, as a multi-label classification challenge. Then, we treated anomaly detection as a binary classification problem, where the model was tasked with determining the presence or absence of DR lesions in a fundus image. EyePACS dataset (Cuadros and Bresnick, 2009) was used for all tasks, with fundus images balanced for equal distribution of positive and negative labels.
We benchmarked Med-Gemini-2D against Gemini Ultra, which used both the fundus image and a multiple-choice-like prompt for prediction. To extract the prediction labels, we searched for specific markers that the LLM was constrained to output (e.g. (G) if no DR lesion is present). In contrast, Med-Gemini-2D relied only on the image, with prediction labels extracted by searching for keywords, like “hemorrhage.” For the anomaly detection task, we compare to a third model similar to (Krause et al., 2018) that has been trained using supervised learning to detect different grades of DR (none, mild, moderate, severe, proliferative). If the predicted DR grade is “none,” the fundus is considered normal (no DR lesion), otherwise, the image is predicted to have DR lesions present. We note that this supervised model was carefully trained on a much larger dataset containing more than 3 million fundus images from diverse manufactures/data sources/geography, and could be considered as an “upper bound” of this task.
Table 5 displays the performance of Med-Gemini-2D and Gemini Ultra on the classification of hard exudates, hemorrhages and PRP Scars, utilizing accuracy, sensitivity, specificity, and F1 score as evaluation metrics. It also compares the performance of DR lesions detection between, Med-Gemini-2D, Gemini Ultra and the strong supervised model. Results demonstrate that Med-Gemini-2D consistently outperformed Gemini Ultra on both multi-label and binary classification tasks. Notably, Med-Gemini-2D achieved significantly higher specificity in hard exudate classification (96.4% vs. Gemini Ultra ’s 39.0%) and hemorrhage classification (81.1% vs. Gemini Ultra ’s 28.1%). These results highlight the benefits of task-specific fine-tuning for this highly-specialized medical domain. Med-Gemini-2D underperformed on the anomaly detection task compared to the strong supervised model, but it is important to acknowledge the strong supervised model’s significant advantage in the volume of labeled data used during its training process.
For three other attempted classification tasks, namely detection of microaneurysms, neovascularization of the optic disc and neovascularization elsewhere, Med-Gemini-2D appeared to be miscalibrated, predicting most cases as negative in the LLM text output. While this is likely related to the training dataset distribution and overall data mixing ratio, further work is needed to improve question-answering based classification and calibration.
2 Visual question answering (VQA)
We assessed Med-Gemini-2D’s performance on VQA tasks across a range of diverse medical specialties, including radiology, dermatology, and pathology, and spanning a wide range of open-ended and closed-ended questions. Table 6 summarizes the overall VQA results. Prompt templates were manually optimized for each model and VQA dataset on the validation splits, and are listed in Table A.2. Model answers were generated using the same method and parameters as for CXR classification, see Section 4.1. That is, the generative sampling was not constrained by a given test vocabulary in any manner, as it was in related work (Li et al., 2023b, a; Zhang et al., 2023a), typically only to the test set’s ground truth answers, for the reasons described e.g. in Tu et al. (2024); Van Sonsbeek et al. (2023). In other words, answers were generated in a truly generative, open-ended, and zero-shot manner.
For close-ended questions, we measured accuracy based on exact matches of normalized model vs ground truth answers, and compare against SOTA based on vocabulary-constrained answer generations, since the freely generated answers mostly matched the overall set of ground truth answers. For open-ended questions, we report the average token-wise F1 score (Tu et al., 2024) between the normalized answers of the model and the ground truth, and only compare against SOTA results where answers were generated in the same zero-shot manner. In addition, for Med-Gemini-2D results on VQA-Rad, one board-certified radiologist scored the answers using the 3-point scoring rubric introduced in Xu et al. (2023), in order to compare against results of the prior ELIXR model.
We assessed Med-Gemini-2D’s VQA capabilities in radiology using three datasets from distinct domains. First, we evaluated on the MIMIC-CXR VQA test set, an in-distribution benchmark containing 226 question-answer pairs for 48 chest X-ray images suggested by Xu et al. (2023). Second, we used the English-only 1,061 question and answer pairs in the test split of Slake VQA, a large bilingual (English and Chinese) VQA dataset. Finally, we employed the VQA-Rad dataset, leveraging the new three-way balanced split detailed in Section 2. As discussed previously, to evaluate out-of-distribution performance and facilitate a head-to-head comparison with the previous best-in-class model (ELIXR), we did not fine-tune our model on either the VQA-Rad training images or questions. This approach ensures both ELIXR and our model are tested on the same, larger VQA-Rad test set (Xu et al., 2023). Moreover we evaluated the performance of our model on both chest X-ray only and all modality (CT, MRI, and X-ray).
As Table 6 demonstrates, Med-Gemini-2D outperformed many previous results and Gemini across different subsets and metrics. Specifically, in the chest X-ray only (ELIXR split) subset, our model achieved a remarkable expert-evaluated accuracy score of 71.9 and an accuracy of 78.8 in closed-ended questions, improving the best-in-class number by 14 and 11.7, respectively. In the chest X-ray-only balanced split subset, our model maintained strong performance with an expert-evaluated accuracy of 71.8 improving best-in-class number by 16 and an accuracy of 78.1 in closed-ended questions. Moreover, across all modalities in the balanced split, our model achieved competitive results and improved over Gemini, demonstrating its versatility. In the MIMIC-CXR VQA dataset, it achieved an accuracy of 78.6% on Yes/No questions and a tokenized F1 score of 52.5 overall. In the Slake VQA dataset, our model significantly outperformed Gemini and achieves performance close to state-of-the-art with 84.8 accuracy on close-ended questions, showcasing its capability across different domains. Its mean tokenized F1-score across all English questions at 75.8 is lower than for the MedPaLM-M model (Tu et al., 2024) at 89.3, which might be partially attributable to it being prompted in a zero-shot manner, versus a one-shot text-only prompt for the latter. Contrary to MedPaLM-M, Med-Gemini-2D was not fine-tuned with one-shot examples, hence this prompting technique would yield worse results during inference.
To evaluate Med-Gemini-2D’s VQA capabilities in pathology, we utilized the PathVQA dataset (He et al., 2020). For this dataset, our model achieved an accuracy of 83.3 at Yes/No questions, and a tokenized F1-score of 58.7 over all questions which improves over Gemini. Its overall zero-shot results are slightly below those of the MedPaLM-M model (Tu et al., 2024), which reports an overall tokenized F1-score of 62.7, albeit employing a text-only one-shot prompting technique here as well. This technique involved an additional exemplar question-and-answer pair along with an image placeholder string provided as a one-shot example during evaluation. However, while these results can provide a general sense of VQA capabilities for images comprising both histopathology and general anatomic pathology photographs and diagrams, given the known issues with QA pairs and image quality in this auto-generated dataset (Lu et al., 2024), we suggest cautious interpretation.
We also conducted a qualitative review of model behavior for histopathology and radiology VQA tasks. Examples are shown in Figure 6, and Figure 7 .
3 Report generation for chest X-rays
In clinical practice, the role of the radiologist extends far beyond narrow interpretation of radiology images. Radiologists are tasked with conveying nuanced findings within a broader clinical context, synthesizing information, and providing recommendations for patient care. Expert radiologists use natural language to articulate this synthesis of imaging findings, overall impressions, and recommendations in written reports. Unlike some prior work, our model was tuned for the difficult task of generating both the ‘FINDINGS’ and ‘IMPRESSION’ sections of chest X-ray reports for frontal view chest radiographs (anterior-posterior or posterior-anterior), covering comprehensively both the observations and inferences typically made by radiologists during a study.
Table 8 presents the performance comparison of various models in generating radiology reports for chest X-rays using the publicly available MIMIC-CXR dataset. The “Sections” column indicates whether the model generates the ‘FINDINGS’ (‘F’) or ‘IMPRESSIONS’ (‘I’) section of the report, with metrics drawn from published research. Higher values in all metrics indicate superior performance. Notably, our model undertakes the more challenging task of generating both sections (F + I) for frontal chest X-rays, aiming to capture the radiologist’s holistic interpretation of the study.
Following common practice, we leveraged the established n-gram based methods such as ROUGE-L, BLEU-4 to evaluate the generated reports quality against the ground-truth. Additionally, we measured the RadGraph F1-score, which is the F1 score between the entities extracted from the reference report and generated one using RadGraph (Jain et al., 2021). RadGraph accounts for not only the absence or presence of findings in the report, but also their relationships to image features. Med-Gemini achieved a RadGraph F1-score of 24.4%, marking a notable improvement of 3.9% compared to the previous top-performing model.
For the IND1 dataset we did not evaluate automated metrics, as automated metrics such as RadGraph F1-score are specifically trained on MIMIC-CXR to measure performance of US-style chest X-ray report and are not capable of handling the out-of-distribution format of IND-1 dataset reports obtained in an India-based clinical setting.
For report generation we devised a novel evaluation rubric, expanding on those used in Flamingo-CXR (Tanno et al., 2024) and Med-PaLM M (Tu et al., 2024), to understand potential impact on clinical management of patients. The evaluation rubric consists of six categories that compare two reports. It provides an improved granularity around patient impact and was used for both the CXR and CT generated reports. Table 9 defines the labels for comparing the AI and original radiologist reports for the same study. This rubric along with training materials and examples were provided to radiologist labelers as training material, prior to any labeling. In each example seen by labelers, the origin of the reports (AI vs. original) was masked and the reports were shown in random order to avoid bias.
Five India-based board-certified radiologists, one India-based thoracic specialist, and one US-based academic thoracic radiologist evaluated a total of 606 cases: 306 from the MIMIC test dataset and 300 from IND1. After the study completion, readers were compared using their mean Quadratic Kappa (Sim and Wright, 2005) to the two thoracic specialists. Two readers falling below 0.2, i.e. “none to slight agreement” were eliminated from the final results. The results were computed based on the total sum of categories for the selected reports after elimination of scores of the X category. The percentage of cases within each category were then plotted sequentially along a horizontal plot for all, abnormal, and normal cases as shown in Figure 3 and summaries are shown in Table 7
In examining cases that fell into the A1 and B1 categories, i.e. where one report captures clinical findings but both would result in the same patient management, similar reasoning was given in both categories. These included missing less critical findings and descriptiveness of findings. Examples of missed findings include: mild cardiomegaly, calcified granulomas, and old fractures. In terms of descriptiveness, examples include: better descriptions of bulla, proper identification of devices, and clearly discerning mass versus pneumonia and other less explicit diagnoses. Reports falling into categories A2 and B2 missed key findings, including: failures in assessing tube positions, missed nodules, and missed pneumothraces.
4 Report generation for head/neck CT volumes
3D imaging modalities often involve more complex data preparation and longer radiologist interpretation time in comparison to 2D images such as X-rays, making the paired image-text data required for generative AI modeling scarcer and more expensive. Additionally, radiology reports tend to be much longer and imaging features much sparser for 3D images than for 2D images. Given this relative data scarcity and information complexity (with correspondingly increased memory requirements), end-to-end modeling to convert 3D radiology images to text reports has previously been infeasible. With its increased computational capacity and extensive domain-specific pretraining, Med-Gemini-3D, building on other recent generative AI work such as Hamamci et al. (2024), is the first LLM-based generative AI model able to interpret a 3D medical imaging modality end to end from the CT volume to text.
Using the same human evaluation rubrics introduced in Section 4.3, we evaluated a total of 92 non-contrast head/neck CT studies consisting of 27 Normal-labeled cases without findings and 65 Abnormal-labeled cases that contained findings, including both acute findings such as cerebrovascular accidents as well as findings that are common consequences of aging, such as atrophy. Studies were initially divided into normal and abnormal candidates based on the length of the impression sections of the reports. A random subset within each was selected and then manually classified by a board-certified radiologist into the normal or abnormal category based on the full radiology report. Studies classified as normal contained no findings.
In reviewing the reports, a single academic board-certified examined the study and all series using a web-based Picture Archiving and Communication System (PACS) viewer. The radiologist graded the two reports using the same rubric presented for evaluating CXR reports. For each rating, the radiologist also recorded a comment describing why the rating was given. The model generated the report based on a single series with the most slices and did not have access to any of the other series. The model was given the patient history in the form of text during inference.
Results are shown in Figure 4 and Table 10. We found that 45% of AI reports on normal studies and 57% of AI reports on abnormal studies would have resulted in the correct clinical management of the patient, though some of those AI reports included errors that would not directly affect management. We did find, however, that only 17% of AI reports were considered to be of equivalent or better quality than the original radiologist reports. In examining the notes on errors from the generated reports, i.e.,those that were scored B2, roughly half involved missed findings while the other half involved hallucinations such as identified subdural hematomas or cysts. In terms of B1 category reports, comments about the generated report mention it either incorrectly estimates or under-characterizes white matter changes.
While our early results presented here leave significant room for future improvement, the potential opportunity for AI in volumetric imaging is vast. This difficulty in reporting on volumetric data can result in concerning diagnostic delays (NHS, 2024). The ability to safely triage, expedite, and quality check existing reports could be highly beneficial in health systems around the world.
5 Disease prediction from genetic information
Personalized medicine can benefit greatly from genetics, as disease risks depend heavily on an individual’s genetic makeup. To leverage this powerful information, we expanded our model’s ability to process genetic information in the form of an RGB image by featurizing the genome into polygenic risk scores (PRS) as explained in Section 2.
To assess the disease risk prediction capability of Med-Gemini-Polygenic, we created benchmarks by training linear models on all PRS featurizations plus demographics (“Ensemble of PRSs and demographics”) which is the current best practice for using PRSs for disease prediction (Albiñana et al., 2023; Truong et al., 2024). For in-distribution health outcomes (see Section 2) used in Med-Gemini-Polygenic training, we directly applied the trained “Ensemble of PRSs and demographics” models as benchmarks. For health outcomes that were never used in the training process (out-of-distribution or OOD, see Section 2) but share some genetic correlation with the in-distribution outcomes, we first calculated their phenotypic correlations with all the in-distribution outcomes, and used the model trained to predict the most correlated in-distribution outcome to generate a maximally strong performance benchmark.
We evaluated Med-Gemini-Polygenic performance on case/control balanced datasets sampled from the test split (200 cases, 200 controls per outcome) for computational efficiency (Section A.2.2). We obtained a disease probability score from Med-Gemini-Polygenic by prompting it to predict the status of a given health outcome using a text prompt and the genetic risk “image” (Table A.9), and computed the probability as the ratio of the likelihoods of the model generating a positive and negative prediction. Med-Gemini-Polygenic achieved higher AUCs than the PRS linear model benchmarks for all in-distribution health outcomes except glaucoma (Figure 5). To evaluate zero-shot generalization ability, we prompted Med-Gemini-Polygenic to predict disease status for the six out-of-distribution health outcomes. Med-Gemini-Polygenic achieved similar performance to benchmarks trained on the most correlated in-distribution outcome (Table A.11) despite never being instructed about the associations between in-distribution and out-of-distribution outcomes (Figure 5).
Additionally, we compared the performance of linear probes of the Med-Gemini-Polygenic embeddings and directly prompting Med-Gemini-Polygenic in the evaluation sets of 400 individuals. Comparisons of the AUCs show that while Med-Gemini-Polygenic performs similarly to the linear probe when each is only given demographic information, it often outperforms the linear probe when incorporating both PRSs and demographics (Figure A.5). This performance increase is largely attributable to Med-Gemini-Polygenic modeling non-linear interactions between genomic information and demographics (Table A.12).
We caution that the AUC values reported here represent an upper bound on model performance since the GWASs used to create the PRS features were performed within the UK Biobank. However, the relative performance of different models that all operate on this in-sample data is the measure of interest for these analyses.
Qualitative Results
In this section we provide a few examples showcasing our model’s capability in medical dialogue for diverse set of medical modalities including chest X-ray, CT, fundus, dermatology, pathology, depicted in Figures 6, 7 as well as 2D (Figure 8) and 3D (Figure 9) radiology report generation. As highlighted in these examples, Med-Gemini is able to provide accurate and reasonable multimodal dialogue and interpretation capabilities across a variety of medical imaging domains. At the same time, expert review of these examples highlights areas for improvement regarding the phrasing, accuracy, appropriate level of detail, and completeness of generated responses.
In addition to understanding automated report generation capabilities, it is important to consider plausible real world assistive use cases. As a proof of concept, we experimented with directing the model’s attention to a specific region/organ within the CXR (Figure 10).
Lastly, as shown in the above examples, even though Med-Gemini was only fine-tuned with data directly related to image interpretation (e.g. there were no question-answer pairs related to treatments or symptoms in the fine-tuning set), Med-Gemini can still leverage the medical knowledge from Gemini pretraining to give simple but reasonable answers to those questions. While we emphasize that real-world medical diagnosis, prognosis, and treatment information is much more complicated and nuanced than the examples provided here, these examples serve as a proof of concept for combining large model pretraining with domain specialization, an active area for further improvements.
Related Work
Large language models (LLMs) built on Transformer architectures (Parmar et al., 2018; Vaswani et al., 2017) have seen rapid advancement, driving significant progress in natural language processing and multimodal modeling. Pathway scaling methods (Barham et al., 2022) have been crucial in enabling the development of ever-larger models like the PaLM family including PaLM, PaLM 2, and PaLM-E (Anil et al., 2023; Chowdhery et al., 2023; Driess et al., 2023). Other significant LLMs include BERT (Devlin et al., 2018), GPT family (Radford et al., 2019; Brown et al., 2020; Achiam et al., 2023), T5 (Raffel et al., 2020), and LLaMA (Touvron et al., 2023), Hyena (Poli et al., 2023), Mistral 7B (Jiang et al., 2023). LLMs are often refined through techniques like Chain of Thought (CoT) prompting (Wei et al., 2022) or fine-tuning (FLAN) (Wei et al., 2022).
These advancements have catalyzed an expansion of LLMs specifically designed for medical domains, such as PubMedGPT (Bolton et al., 2022), BioGPT (Luo et al., 2022), Med-PaLM (Singhal et al., 2023a) and its successor Med-PaLM 2 (Singhal et al., 2023b), Clinical Camel (Toma et al., 2023), MedAlpaca (Han et al., 2023), BioMistral (Labrak et al., 2024), LLMs for clinical trial recruitment (Wornow et al., 2024), and others. Language models can handle omic information, as demonstrated by models such as HyenaDNA (Nguyen et al., 2024), BioT5 (Pei et al., 2023), sc-GPT (Cui et al., 2024), and ProtLLM (Zhuo et al., 2024).
Beyond language and text alone, multimodal models like Flamingo (Alayrac et al., 2022), PaLI (Chen et al., 2022), GPT-4 (Achiam et al., 2023), GPT-4v (OpenAI, 2023), and LLaVa (Liu et al., 2023, 2024a) have demonstrated remarkable capability in processing both text and images. Gemini (Gemini Team, Google, 2023; Google, 2024) introduced further advancement in multimodal capabilities, exhibiting a distinct ability to reason across text, images, and other modalities such as video and audio.
Building upon these capable generic multimodal models, for medical applications specifically, recent works include vision-language models that span multiple medical imaging modalities as well as those that focus on a specific imaging domain, such as radiology or histopathology. Efforts such as Med-Flamingo (Moor et al., 2023b), BiomedCLIP (Zhang et al., 2023a), Med-PaLM M (Tu et al., 2024), BiomedGPT (Zhang et al., 2023a), Flamingo-CXR (Tanno et al., 2024), LLaVa-Med (Li et al., 2024), PMC-VQA (Zhang et al., 2023b), RadFM (Wu et al., 2023), ELIXR (Xu et al., 2023), XrayGPT (Thawkar et al., 2023), MAIRA-1 (Hyland et al., 2023), HeLM (Belyaeva et al., 2023), M-REGLE (Zhou et al., 2024), CONCH (Lu et al., 2024), PLIP (Huang et al., 2023), PathAsst (Sun et al., 2024), QuiltNet-B-32 (Ikezogwo et al., 2024) and many others specifically explore the potential of multimodal models for medical applications, signaling a growing interest in this area. These methods cover a range from generalist to specialist approaches. Models such as MAIRA-1 (Hyland et al., 2023), XrayGPT (Thawkar et al., 2023), Radiology-GPT (Liu et al., 2024b), and CT2Rep (Hamamci et al., 2024) focus on radiology report generation, and among modalities choose only chest X-ray or chest CT report generation. Some of these approaches broaden their capabilities to cover multiple types of modalities but focus on only one task, such as methods that aim for VQA capabilities like LLaVA-Med (Li et al., 2024), and PMC-VQA (Zhang et al., 2023b), aiming to build assistants for medical question answering.
While specialized VLMs demonstrate particular strengths, generalist models capable of handling a wide range of tasks and modalities, such as Med-PaLM M, are gaining prominence. The field of medical AI is witnessing the emergence of comprehensive ‘Generalist Medical AI’ models (Moor et al., 2023b, a; Tu et al., 2024; Zhang et al., 2023a) and the orchestration of AI tools for medical tasks using LLMs (Ferber et al., 2024). These models aspire to provide robust interaction with medical information in a manner similar to what general-purpose LLMs have done for broader domains. Pioneering efforts like these offer important initial insights into the potential for large multimodal models to provide assistance across various medical tasks using a unified platform. This inconsistency underscores the urgent need for a unified benchmark to enable meaningful evaluation in this rapidly evolving field.
Evaluation of medical VLMs suffers from a lack of consistency and standardization, creating a new landscape for works proposing new benchmarks to fill this gap. Multiple recent works demonstrate this inconsistency with varying tasks, datasets, and completely distinct sets of metrics, hindering direct comparison even for a same dataset. Along these lines, multiple recent works (Royer et al., 2024; Tu et al., 2024; Wu et al., 2023; Moor et al., 2023b; Fleming et al., 2023) suggest multimodal benchmarks such MultiMedEval, MultiMedBench, RadBench, RadMD, and MedMD to evaluate these generalist and multimodal models in a more systematic fashion.
Discussion
In this study, we present three new models within the Med-Gemini family, based upon Gemini 1.5, across various medical modalities. We show promising performance across a number of tasks, including classification, VQA, and report generation. Our Med-Gemini models are able to process complex medical data types, including 2D and 3D radiology images, histopathology patches, ophthalmology images, dermatology images, and genetic risk scores. Importantly, our models were fine-tuned using predominantly medical data and paired free text descriptive reports. These reports are ubiquitous in healthcare and our ability to use them as a training objective reduces the need for further expensive expert labelling.
The results in this study show early potential across a number of different tasks and individual modalities. We believe that the combination of tasks and multiple modalities in future work will enable AI models to address a far wider range of applications than has been previously possible. Longer context windows and improved reasoning abilities will enable decision-making that incorporates historical context, more closely reflecting how human specialists operate.
The opportunity for LMMs to analyze complex medical types including 3D radiology and large pathology images presents an exciting range of potential downstream applications. This work showcases our early explorations in CT, a three-dimensional modality that has been challenging to integrate with LMMs to date. This is due to a combination of vast data size, architectural limitations, and the jump in clinical task complexity of interpreting 3D imaging modalities (vs. 2D). While our results are currently a proof of concept, and do not yet reach performance required for clinical use, we expect architectures to rapidly improve. We look forward to exploring other similar complex modalities in future work.
While our findings in this study are promising and provide a glimpse into the potential of LMMs in medicine, it is important to thoroughly test them beyond traditional academic benchmarks. This is necessary to ensure they are safe and reliable before considering use in real-world situations, especially in safety critical areas like healthcare. In this work, we have tried to go deeper into the nuance of medical evaluation through the use of panels of specialists to assess and rate the performance of models on tasks such as report generation and question answering. We believe that an increasingly diverse range of healthcare professionals need to be deeply involved in future iterations of this technology, helping to guide the models towards capabilities that have valuable real world utility. There are a number of areas on which future evaluations should focus before models like these are considered safe and effective for clinical use:
Despite the potential of machine learning in healthcare, there is growing concern about the reliability of algorithm validation methods. In medical image analysis, improvement on simple benchmark performance metrics may not translate to improved outcomes in clinical settings, leading to a disconnect between expectations and real-world usefulness. Benchmark datasets are an important step towards developing clinically useful models, but given their limitations in size, scope, and reflection of real world distributions, they are not themselves a proxy for real-world performance. The potential for generative AI lies foremost in assisting, rather than replacing human specialists in the diagnosis and management of disease; evaluations should shift from static benchmarks to realistic clinical scenarios that assess AI-human collaboration and its impact on patient outcomes.
LLMs and LMMs trained on vast datasets risk inheriting biases and errors from their source data. This can lead to misdiagnoses and amplification of systemic bias. Before models like these are used in real world settings, careful evaluations that address safety and bias risks should be performed and any discovered risks should be mitigated (Weng et al., 2024). End users should also carefully validate model performance for their specific use cases and patient populations.
While LLMs exhibit impressive zero-shot generalization, it’s important to note that their massive training datasets increase the potential for data contamination, which may result in overestimation of their true generalization abilities. Large models like Gemini might have inadvertently “seen” examples related to the task during training, even if those examples were not explicitly labeled. This hidden exposure compromises our understanding of models’ true ability to generalize to completely novel concepts when evaluating on open datasets. Researchers are actively investigating the impact of data contamination to ensure we accurately gauge capabilities of such large models (Vogel et al., 2022; Udandarao et al., 2024). Prospective studies, while typically more expensive and time-consuming to execute than retrospective studies, are another option for mitigating this risk.
Conclusion
Multimodal generative AI, exemplified by powerful models like Gemini, holds great potential to revolutionize healthcare. While medicine is a rapidly growing use case for these new models, general purpose models may not naturally perform well in the medical domain due to its highly specialized data.
To explore the potential for models like Gemini in medicine, we developed several models within the new Med-Gemini family, a series of models built upon the multimodal foundation of Gemini and fine-tuned on a diverse range of medical data including radiology, histopathology, ophthalmology, dermatology and genomics. We assessed our Med-Gemini models’ performance using a comprehensive medical benchmarking suite, including both established benchmarks and custom benchmarks designed to reflect clinical relevance. Notably, some benchmarks involved evaluations by medical experts for tasks such as generating CXR and CT reports and radiology VQA.
Med-Gemini-2D sets a new standard for expert-evaluated chest X-ray report generation, outperforming previous models, and Med-Gemini-3D showcases the first LMM-based report generation for 3D CT. Beyond report generation, Med-Gemini-2D demonstrates exceptional performance in VQA and classification across various medical imaging modalities. Beyond imaging, Med-Gemini-Polygenic outperforms conventional polygenic risk score methods in predicting disease risk. These results demonstrate the potential of the Gemini foundation and the fine-tuned Med-Gemini family in the medical domain. Nonetheless, the results also underscore the need for further rigorous research to ensure safe and effective implementation in real-world clinical settings.
While advanced capabilities on individual medical tasks are useful in their own right, we envision a future in which all of these capabilities are integrated together into comprehensive systems to perform a range of complex multidisciplinary clinical tasks, working alongside humans to maximize clinical efficacy and improve patient outcomes. The results presented in this study represent a step towards realizing this vision.
Contributions and Acknowledgments
Authors are listed here associated with their primary workstreams. Many authors contributed to additional workstreams beyond the one under which they are listed.
Google Research and Google DeepMind Leadership
7EyePACS, Inc and Meredith Morgan University Eye Center, University of California at Berkeley
Acknowledgements
This project was an extensive collaboration between many teams at Google Research and Google DeepMind. We thank Kevin Swersky and Mike Schaekermannn for their feedback and insight, which significantly contributed to the enhancement of this report. We also thank Sami Lachgar, Lauren Winer, Maggie Shiels, Jessica Valdez, Jon Small, Aaron Abood, Rishad Patel, Christian Wright, Annisah Um’rani, Jean-baptiste Alayrac, Aishwarya Kamath, Viorica Patraucean, Rory Sayres, Abbi Ward, Louis Blankemeier, Olga Kanzheleva, Taedong Yun, Ksenia Konyushkova, Christos Kaplanis, Juanma Zambrano Chaves, Alan Karthikesalingam, Vivek Natarajan, and Can Kirmizi for their valuable insights, technical support and feedback during our research. We thank Kimberly Kanada and Ilana Traynis for their review of the qualitative examples shown in this manuscript. We are grateful to Jonathon Shlens, Dale Webster and Oriol Vinyals for their support during the course of this project. We also thank Michael Colligan and Brittany Stein from DeepHealth/RadNet for their support with data curation.
This research was conducted using the UK Biobank Resource under application number 65275. The results shown here are in part based upon data generated by the TCGA Research Network. The authors thank the National Cancer Institute for access to NCI’s data collected by the National Lung Screening Trial (NLST). The statements contained herein are solely those of the authors and do not represent or imply concurrence or endorsement by NCI.
Data Availability
Except IND1, CXR-US2, and CT-US1, Eyepacs, and TTH, which are private datasets, the rest of the datasets utilized for developing, benchmarking, and evaluation of Gemini and Med-Gemini in this report are publicly accessible with appropriate permissions. We intend to publicly release our updated classification labels and custom VQA question and answer pairs for the MIMIC-CXR dataset, our splits for the PAD-UFES-20 and VQA-Rad datasets, and several suggested replacement question and answer pairs for the VQA-Rad dataset which were recommended by our reading radiologist. This text will be updated when that data is available.
Code Availability
We will not open-source the model code and weights because of the safety concerns associated with unmonitored use in medical settings. To ensure responsible innovation, we will collaborate with our research partners and healthcare providers to validate and explore safe applications of the Gemini and Med-Gemini through Google Cloud APIs.
Competing Interests
This study was funded by Alphabet Inc and/or a subsidiary thereof (‘Alphabet’). Authors who are affiliated with Google Research, Google DeepMind, and Verily Life Sciences are employees of Alphabet and may own stock as part of the standard compensation package.
Use of AI in Manuscript Preparation
This manuscript was written manually, with a small number of copy edits performed using Gemini. The authors take all responsibility for the contents.
References
Appendix A.1 Additional data details
One of the limitation of the MIMIC-CXR dataset is the lack of ground-truth labels. MIMIC-CXR JPG (Johnson et al., 2019b) extracted structured labels from 277,827 radiology reports using CheXpert (Irvin et al., 2019), a natural language processing (NLP) tool to extract observations from radiology reports. In order to improve upon these labels on the test subset, we utilized Med-PaLM 2 (Singhal et al., 2023b) coupled with US-based board certified radiologists to refine those labels. This work is further adjudication of the labels used in Xu et al. (2023). We first used a keyword search to identify reports containing text associated with the a given finding (e.g.,“Cardiomegaly”). Next, Med-PaLM 2 was applied to the flagged radiology reports on a per-label basis using two queries shown in Table A.1 for a total of 23,824 queries. All identified positive and negative labels that disagreed with the original labels were flagged for human verification.
Three US-based board certified radiologists reviewed the 1,378 flagged labels and a fourth academic US-based board certified thoracic radiologist adjudicated the responses of reviewer disagreements. For each finding and report, radiologists selected one of four possible labels defined by MIMIC-CXR JPG: positive, negative, uncertain, and not mentioned. Zero-round adjudication was performed on the reviewers’ annotations. There was strong inter-rater agreement (Fleiss’ = 0.71); reviewers were unanimous for 77% of the labels. In cases of disagreement between reviewers, majority vote was used (21%), and when all three reviewers disagreed (2%) a senior academic thoracic radiologist provided the final determination.
In the final analysis of the flagged reports and findings, Med-PaLM 2’s label matched the ground truth 66% of the time while the original labels were correct in 19% of the cases. The labels are in preparation to be released (Park et al., 2024).
A.1.2 Prompts for VQA and CXR classification evaluations
We explored both binary question prompting for each of the evaluated top 5 conditions for MIMIC-CXR including the abnormal/normal class, as well as multi-select prompts for all of them at once on the validation set. Since binary prompts overall yielded better macro F1 scores for Med-Gemini, these were used at evaluation. For Gemini Ultra binary question were used for the top-5 conditions and a multi-select prompt for the normal/abnormal condition, see Tables A.3 and A.4. Each prompt template was crafted for each model in aiming to optimize its performance, which yielded very short prompt templates for both classification and VQA for Med-Gemini, since it was fine-tuned with clinical questions.
A.1.3 New balanced splits for VQA-Rad dataset
The official train/test split of the VQA-Rad (Lau et al., 2018) dataset comprises 1,797 QA pairs for training (i.e. dataset field QID_para {’freeform’, ’para’}) for 313 different MedPIX®images (per field IMAGEID), and 451 QA pairs for 203 different images in the test set (i.e. field QID_para {’test_freeform’, ’test_para’}). Although images were sampled such that each is not only for a different case, but also a different patient, see (Lau et al., 2018), 202 of the test IMAGEIDs and also match the train set IMAGEIDs. Hence most of the test images also appear in the train set, only the questions and answers differ. For some VQAs even the latter is not completely true, since VQA-Rad contains paraphrased questions which share the same answer.
To remove this train/test contamination issue, in Xu et al. (2023) we proposed a different validation/test split, which is based on the IMAGEIDs in order to ensure disjoint images and thus patients. In this work we split this relatively large test set further into a new test and train set, which are roughly equal-sized, and assign a few remaining former test set IMAGEIDs and corresponding VQAs to the existing validation set, such that all three new splits not only are roughly equal-sized, but their ratio of open to closed questions (as determined by field A_TYPE) are approximately equal within each of the three depicted anatomical regions (chest, head and abdomen), see Table A.5. We chose to equalize this ratio since open-ended questions are more difficult for AI models, and the level of difficulty ought to be similar for each new split. Similarly, while swapping individual IMAGEIDs and associated question and answers, we approximately equalized the distribution of question types (field Q_TYPE) within each split, in order to gain similar ones, see Table A.6.
A.1.4 Polygenic risk prediction
We crafted prompts for predicting the status of various health outcomes using both an individual’s PRS image and their demographic information. An example prompt for predicting coronary artery disease is shown in Table A.9.
For linear probes of out-of-distribution outcomes, we used data of a related in-distribution outcome to train the linear probe and then evaluated the predictions on the out-of-distribution outcome. For example, in order to evaluate diabetic retinopathy, we train the linear probe to predict type 2 diabetes and evaluate the type 2 diabetes predictions on diabetic retinopathy data. In general, the most related in-distribution outcome is defined as the outcome with the highest Matthew’s correlation coefficient with the out-of-distribution outcome across individuals in our training set (Table A.11). To evaluate Med-Gemini-Polygenic, we directly prompted Med-Gemini-Polygenic to predict the out-of-distribution outcome without providing any information about correlations between out-of-distribution and in-distribution outcomes.
Appendix A.2 Additional results
We performed data-efficient classification for Chest X-ray classification task focusing on examples across 8 different findings (atelectasis, cardiomegaly, airspace opacity, consolidation, fracture, pneumothorax, pleural effusion, and pulmonary edema). We also deploy two out-of-distribution datasets including ChestX-ray14 and CheXpert for this purpose. Our data-efficient classification follows the protocol from (Xu et al., 2023) except that instead of training a Multilayer Perceptron (MLP) as a nonlinear classifier, we train a linear probe on top of the frozen image encoder. Following the ELEVATER(Li et al., 2022) method, we initialize the weights of the final linear layer with the text embeddings for the class label. Training parameters includes a learning rate of 0.2, a batch size of 512, and 300 epochs utilizing the Layer-wise Adaptive Rate Scaling (LARS) optimizer.
In alignment with previous best-in-class method, ELIXR (Xu et al., 2023), the linear classifiers were trained on 5 different varying sample sizes including 0.01% to 100% subsets of the training data to facilitate direct comparability of results to Xu et al. (2023). The smallest sample size includes 64 samples. Figure A.1 shows aggregated results of Med-Gemini vs. ELIXR on ChestX-ray14 and CheXpert (Xu et al., 2023) for 5 and 6 various runs, respectively. Comparison between data-efficient classification results of Med-Gemini vs. ELIXR reveals that linear probes trained on top of visual embeddings from Med-Gemini exhibit robust performance in data-efficient classification, although approximately one order of magnitude inferior than ELIXR at the sample size as low as 64 samples.
A.2.2 Polygenic risk prediction
Beyond evaluating Med-Gemini-Polygenic, we also compared linear probes of the Med-Gemini-Polygenic embeddings to linear probes of the demographics only and the ensemble of PRSs and demographics. Figure A.2 uses the same balanced sets of 400 individuals as used in Figure 5, and Figure A.3 uses larger balanced sets containing all the positive cases per health outcome and an equal number of controls. The AUC metrics are relatively consistent between both evaluation sets. Furthermore, we computed Med-Gemini-Polygenic performance on coronary artery disease and COPD in 4000-sample evaluations (Figure A.4), and observed stable results. Taken together, these results suggesting that our evaluation set of 400 individuals is representative of overall model performance.
In addition, we demonstrated that using the Med-Gemini-Polygenic framework likely results in better predictive performance than linear models trained with all PRS featurizations plus demographics regardless of future sample sizes available by conducting sample size ablation tests on the PRS ensemble models. We observed performance plateaus for the linear model with at most samples (Figure A.6).
Finally, we investigated the relative contributions of the genomic embedding and modeling non-linear interactions between genomic representations and demographic information by comparing the performance of Med-Gemini-Polygenic to two other non-linear models: a gradient-boosted decision tree (GBDT) of the “Embeddings and demographics” (“Embeddings”) and a GBDT of the most correlated individual PRS at each of the three significance thresholds and demographics (“Best PRSs”). Med-Gemini-Polygenic and the GBDT of “Embeddings and demographics” yield comparable performance across all traits, and consistently outperform the GBDT of “Best PRS” for in-distribution outcomes, confirming the importance of both multi-PRS predictors and accurately modeling non-linear interactions between genetic contributors and demographic information (Table A.12).
A.2.3 MIMIC-CXR classification
Table A.13 shows the comparison between performance of Med-Gemini and Gemini Ultra measure by F1-score for the original label and revised label as explained in Section A.1.1. Revised MIMIC-CXR labels significantly improve chest X-ray classification performance measured by F1-score (%). Our results demonstrate the impact of accurate ground truth on model evaluation.
A.2.4 Histopathology classification
Table A.14 details the linear probing results for the histopathology patch-classification task, reporting 1-vs-rest AUC (%) with 95% confidence intervals. The confidence intervals were obtained using blocked bootstrap resampling over test set slides with 10,000 replicates. Our model’s image embeddings match the performance of the histopathology-specialized model (PathSSL) on 6 out of 9 in-distribution tasks, with room for improvement on the remaining tasks. While Gemini and Med-Gemini-2D perform similarly overall, Med-Gemini shows a trend towards higher mean AUC on most in-distribution tasks and both out-of-distribution tasks.
Appendix A.3 Evaluation metrics
Beyond human and expert evaluation, we leverage a range of automated metrics tailored to specific tasks. For classification tasks, this may include basic accuracy and AUC (Area Under the ROC Curve) metrics. For tasks like report generation, where the fidelity and informativeness of the generated text are crucial, we employ wide variety of metrics such as BLEU, Rouge-L or RadGraph F1-score to probe the quality of our models.
Used for image classification and close-ended VQA inference tasks. Measures the percentage of correct predictions vs. the ground truth.
AUC is a performance metric for classification models that indicates how well a model distinguishes between different classes. AUC is calculated by plotting the True Positive Rate (TPR) against the False Positive Rate (FPR) at various classification thresholds. The TPR measures the proportion of correctly identified positive instances, while the FPR measures the proportion of incorrectly identified negative instances. The area under this curve represents the model’s overall ability to separate classes. An AUC of 1.0 indicates a perfect classifier, while an AUC of 0.5 implies the model has no better discriminative power than random guessing.
The F1 score is a valuable metric for evaluating classification models, especially when dealing with imbalanced datasets. F1 score is calculated as the harmonic mean of precision (the proportion of true positive out of all predicted positives) and recall (the proportion of true positives correctly identified). The Weighted F1 Score which is used for VQA, averaging F1 scores across classes based on their frequency. Macro-F1 score used for image classification averaging F1 scores across classes without considering imbalances.
Used for image classification in ophthalmology related tasks. Measures the percentage of correctly identified positive cases out of all actual positive cases. A model with high sensitivity minimizes false negatives.
Used for image classification in ophthalmology related tasks. Measures the percentage of correctly identified negative cases out of all actual negative cases. A model with high specificity minimizes false positives.
Tokenized F1-score provides a granular evaluation of language models by calculating precision, recall, and F1-score at the individual token level. This means it rewards partial matches, recognizing the model’s ability to identify elements within a sequence even if they’re not perfectly aligned. For this purpose True positives and false positives are determined as the number of correctly generated tokens and tokens generated but not present in the ground truth, respectively. False negatives are tokens present in the ground truth but missed by the model.
Rouge-L measures evaluates the quality of generated text and text summarization by comparing the longest common subsequence (LCS) between generated and reference text (Lin, 2004). Higher scores indicate better content and better salient point capturing. Rouge-L assesses the similarity between generated and reference text by measuring the overlap of their LCS and calculating recall based on the LCS length relative to the reference text This metric takes into account the order of words in the text, which makes it particularly suitable for evaluating summaries or text generation tasks where the order of words matters. The higher the Rouge-L score, the better the quality of the generated text compared to the reference text. ROUGE-L relies heavily on LCS and exact matches limiting the contextual understating of the generated text and increasing the sensitivity to sentence length. A high ROUGE-L score doesn’t necessarily ensure that the generated text is grammatically correct, well-structured, or reads naturally.
CIDEr (Consensus-based Image Description Evaluation) is a metric specifically designed to assess the quality of captions generated for images and short text passages. It goes beyond simple word overlap by considering both the n-gram matches (sequences of consecutive words) and the importance of those n-grams (Vedantam et al., 2015). In the preprocessing, both generated and reference texts are converted to lowercase and common stop words (“the”, “a”, “an”) are removed. Words are also stemmed, reducing them to their root form (e.g., “running” becomes “run”). Then every generated text is broken down into a series of n-grams which are sequences of ‘n’ consecutive words. A weight is assigned to each n-gram based on its Term Frequency-Inverse Document Frequency (TF-IDF). This means common n-grams across all texts receive lower weights, while those that are more informative and distinctive get higher weights. The cosine similarity is calculated between the TF-IDF weighted n-gram vectors of the generated text and each reference. The individual similarity scores are averaged to produce the final CIDEr score. CIDEr can struggle to recognize texts that are semantically similar but use different synonyms and suffer from limited contextual understanding.
The BLEU (Bilingual Evaluation Understudy) score is a widely used metric for evaluating the quality of AI generated text. It essentially compares a generated text to a set of human-written reference, providing a score that indicates how similar they are (Papineni et al., 2002). BLEU focuses on n-gram precision, meaning it checks how often sequences of n consecutive words in the generated text appear in any of the reference. It also considers a brevity penalty to discourage generations that are significantly shorter than the reference text. Higher BLEU scores indicate better translation quality, with a perfect score of 1.0 signifying a perfect match between the generated text and the reference. BLEU score has limitations including lack of penalization for grammatical correctness, fluency, or semantic equivalence. Additionally, the quality of the reference and ground truth can impact the BLEU score.
RadGraph F1-score (Jain et al., 2021) is a performance metric specifically designed to evaluate the accuracy of models that extract structured medical information from radiology reports. Unlike standard F1-scores, RadGraph F1-score considers not only whether a finding is correctly identified but also the accuracy of its relationships with other findings within the report. This is crucial because radiology reports often describe complex relationships between abnormalities, locations, and other attributes. Although RadGraph F1-score has its shortcomings, in comparison to other automated NLG metrics provides a more holistic assessment of a model’s ability to understand the nuanced information present in free-text radiology reports.
While RadGraph F1-score offers a more nuanced evaluation than standard F1-scores for radiology report analysis, it has potential limitations. First, it relies on accurate RadGraph creation from the original text. Errors in entity extraction or relation identification during this pre-processing stage could cascade into the RadGraph F1-score calculation. Secondly, it might be overly strict for partial matches and slight discrepancies in relationships or minor variations in wording could significantly penalize the score. Finally, it may not fully account for the clinical relevance of certain errors, treating all mismatches equally despite the potential for varying real-world impact.
To compute the RadGraph F1-score, the model’s predictions on chest X-ray images are compared against ground-truth report made by radiologists or other experts. To increase robustness of our calculation to slight format changes, before passing the ground-truth and the generated report through the RadGraph F1-score package (Yu et al., 2023), we normalize both free-form text to lowercase. The F1-score takes into account both false positives (cases where the model incorrectly identifies an abnormality) and false negatives (cases where the model fails to detect a true abnormality). By considering both precision (the ratio of true positives to the total number of predicted positives) and recall (the ratio of true positives to the total number of actual positives), the F1-score provides a balanced assessment of the model’s performance. A higher RadGraph F1-score indicates better performance in accurately identifying abnormalities in medical images, which is crucial for assisting radiologists in diagnosis and treatment planning.
Appendix A.4 Supplementary Table for Performance Summary
Table A.15 presents the aggregate performance of Med-Gemini compared to the previous state-of-the-art (SoTA), or a strong baseline where available. Figure 1 illustrates the relative improvement gained by using one of our Med-Gemini models over the SoTA or strong baseline, using Gemini as a reference point when no SoTA is available. For pathology classification, we averaged AUC performance across all sub-datasets. For report generation, we calculated the micro average performance across normal and abnormal cases, expert identified “AI generated report is superior or similar to original report" (see Table 7)