ChatGPT Makes Medicine Easy to Swallow: An Exploratory Case Study on Simplified Radiology Reports
Katharina Jeblick, Balthasar Schachtner, Jakob Dexl, Andreas Mittermeier, Anna Theresa Stüber, Johanna Topalis, Tobias Weber, Philipp Wesp, Bastian Sabel, Jens Ricke, Michael Ingrisch
Introduction
"ChatGPT, am I dying? What does this medical report mean? Can you explain it to me like I’m five?" With the latest release of OpenAI’s Large Language Model (LLM) ChatGPT , algorithmic language modeling has reached a new milestone in generating human-like responses to user text inputs. Just five days after its release, ChatGPT had already attracted over a million users and gained a significant amount of media attention . Given the level of popularity and widespread access to the public, the question as to how people will use such models arises and which opportunities and challenges are associated with them. In our experience, the text output of ChatGPT was astonishingly convincing for a variety of tasks, such that we expect disruptive change across numerous domains and industries in a very short time.
Among a myriad of potential downstream applications, LLMs are able to simplify complex text and subsequently make it more accessible to a broader audience. There is a huge need for simplification in domains where expert knowledge is necessary to put text into context and interpret it. For example, legal or medical documents are written by trained experts in highly specialized language, but often have immediate consequences for clients or patients, respectively, who are usually non-experts. Simplifying these documents promises to enable individuals to actively and autonomously oversee them. One domain where text simplification might be particularly useful is radiology. Radiological findings are typically communicated in a free-text report in specialized medical jargon, targeting a clinician or doctor as the recipient. For readers without a medical background, these reports are often inaccessible and potentially misleading. For example, a study showed, that only 4% of all analyzed radiology reports were readable by the average US adult. They concluded to use simpler and more structured language to improve patient-centered care.
Ideally, radiological findings are communicated in a timely, personalized conversation between doctor and patient. However, in resource-constrained health systems, this conversation is often delayed. Therefore, patients might try to make their radiology reports accessible upfront. For instance, patients might research isolated medical terms on the internet without being able to place them in the correct medical context, potentially leading to misinterpretations and confusion. In this context, developed an online system that augments radiology reports with lay-language definitions. Alternatively, patients might use free simplification serviceshttps://washabich.de (visited on 12/28/2022), which often lack popularity, accessibility, or scalability. We expect, that in the very near future, publicly accessible LLMs such as ChatGPT will increasingly be applied by interested patients to simplify their own medical reports filling the gap.
However, ChatGPT was not explicitly trained for medical text simplification and is not intended to be used for this critical task. While it can generate plausible-sounding text, the content does not need to be true, as shown in Figure 1. This raises the question of whether LLMs such as ChatGPT are able to simplify radiology reports such that the output is factually correct, complete, and not potentially harmful to the patient. Patients, i.e., potential LLM users, are not able to answer this question. The quality of these simplified reports needs to be thoroughly assessed by expert radiologists.
In this work, we conducted an exploratory case study to investigate the phenomenon that emerging LLMs such as ChatGPT may be used or misused by patients to simplify their radiology reports. This unintended use may raise chances and challenges for patient-centered care. Therefore, we asked 15 radiologists to rate the quality of three radiology reports simplified by ChatGPT regarding their factual correctness, completeness, and potential harm using a questionnaire. Descriptive statistics and inductive categories were used to evaluate the questionnaire. Based on our findings, we elaborated opportunities and challenges of using ChatGPT-like LLMs for simplifying radiology reports.
Background
In this section, we introduce LLMs, as well as the technical specifics of ChatGPT. We shortly discuss techniques for summarization and simplification of radiology reports and mention known limitations of LMMs. This will serve as a basis for discussing chances and potential challenges for a scenario where these models are used to simplify radiology reports by patients.
Since introducing the concept of attention in deep learning models , the transformer is an established architecture with dominance in nearly all natural language processing (NLP) benchmarks, including question answering, translation, and text classification . While the base-model architecture remained relatively unchanged throughout the years, significant progress was made by scaling the number of layers and internal dimensions resulting in so-called Large Language Models (LLMs) with billions of parameters, which lead to increased model capacity and abilities. Popular architectures grew in parameter size starting, e.g., from BERT (345M; ) over MegatronLM (8.3B; ) and T-5 (11B; ) up to BLOOM (176B; ) and PaLM (540B; ). A significant success factor for fitting LLMs is an enormous training dataset, e.g., the Pile , which contains documents from Arxiv, PubMed, Stack Exchange, Wikipedia, as well as a subset of Common Crawlhttp://commoncrawl.org/, and GitHub, among others. For these kinds of LLMs, introduced the terminology of foundation models, which defines training on a very large data basis and the ability to adapt to a variety of downstream tasks.
2 ChatGPT
ChatGPT is an LLM developed by OpenAI that was first released on November 30th, 2022. The user can directly prompt the model via an API in a conversational way, e.g., allowing for follow-up questions or admission of mistakes . The backbone of ChatGPT is based on the generative pre-trained transformer series (GPT; ). Despite the success and capacity of the third GPT iteration (GPT-3) with 175B parameters, the challenge of engineering text prompts for achieving the desired generative output remained. This is due to the autoregressive training procedure, which tasks the model to predict a token based on the previous text and thus is optimized for text completion and not dialogues. To improve the dialogue capabilities of the model as well as to reduce bias and general toxicity, developed InstructGPT, a fine-tuned GPT-3, that leverages a novel training procedure in order to follow user instructions more persistently. The basic idea for InstructGPT is reinforcement learning from human feedback (RLHF, ), which is composed of three phases: First, human labelers manually create responses for randomly sampled prompts from a prompt database. This handcrafted data served to fine-tune GPT-3 in a supervised fashion (SFT model), i.e., through training on the manually created output for the respective prompt. In the second step, SFT is tasked to create multiple outputs for a given prompt, which are subsequently ranked by a human labeler from best to worst. This new ranking dataset serves to train a reward model, which assigns a scalar reward as a measure of quality for the given SFT outputs. In the final step, concepts from reinforcement learning and proximal policy optimization were applied to fine-tune SFT using a previously trained reward model as an approximate reward function. This allows evaluation of the quality of the online-generated outputs using the reward model.
The setups of training InstructGPT and ChatGPT are nearly identical with minor differences concerning the data collection . The major difference lies in the backbone itself. InstructGPT is based on GPT-3. ChatGPT uses GPT-3.5, a newer, iterated version of the original GPT-3. Additionally, OpenAI embedded various non-disclosed safety mechanisms into ChatGPT to promote a secure and non-toxic environment.
3 Simplification and Summarization of Radiology Reports
Despite being closely related, simplification and summarization express two different concepts. Text summarization describes the task of creating a short version of a text including all important aspects. A supervised recurrent model for this domain was proposed by with increased performance in comparison to previous non-neural approaches. noted a lack of medical terminology in BERT and therefore proposed fine-tuning the model on chest x-ray reports before summarizing impression sections. A domain adaption was done by for a German version of BERT with the goal of summarizing German chest x-ray reports. Further approaches for abstractive summarization employed, e.g., a pointer-generator network or a reinforcement-learning inspired procedure to solve this task. For extractive summarization of medical data, BioBERT was fine-tuned by and a sentence-ranking framework with random forests in the backbone was proposed by . Overall, argued that the continuously growing availability of biomedical text data has led to increased attention to the research field of text summarization, where meanwhile the majority of methods focus on machine learning.
In contrast to summarization, simplifying a text does not necessarily imply shortening, but describes a transformation to make it more readable and understandable . For instance, identified a considerate proportion of recurring standard phrases in their radiology report dataset and argued that these could be easily and consistently replaced with simpler terminology, while more varied descriptions of unique pathological findings were less easy to simplify automatically. A similar procedure was applied in , where human evaluators report a simplification of the medical reports after the substitution. The replacement method was also adapted and extended by a semantic network in . The open-access consumer health vocabulary (CHV) was used by for simplifying text. They came to the conclusion that even some CHV terms may not be appropriate for laypersons.
In addition to this technical work on how to generate simplified reports, studies were conducted to investigate how well patients can understand conventional radiology reports, and if they prefer simplified reports. Using readability tests, showed that the required reading level for their set of lumbar spine reports exceeded 12th grade. They argued that the patients’ understanding of complex texts should be considered since patients increasingly read reports themselves. As mentioned in the introduction, concluded that only 4% of more than 100,000 consecutive radiology reports were at a reading level below the reading level of an average US adult. This finding was further investigated in a randomized controlled trial , where the majority of patients showed a preference for patient-friendly radiology reports over traditional radiology reports.
There is a clear need for patients to be able to better understand radiology reports . While machine learning has made significant progress in the area of radiology report summarization, to the best of our knowledge, automated report simplification has not yet received adequate attention within the NLP community. ChatGPT and other foundation models present promising opportunities for simplifying radiology reports, but it is important to consider their limitations, particularly when used in the medical field.
4 Known Limitations of LLMs
The GPT-3 publication discussed the potential for misuse since the generated outputs become close to indistinguishable from human language. Furthermore, they highlighted model-inherent issues with biases and fairness. Access to GPT is only granted over an API and not to the original model. With an effort to democratize artificial intelligence, EleutherAI published GPT-J as an open-source alternative for GPT-3 . also elaborated on intrinsic biases, which are inherent to foundation models. In addition, they brought up the concept of extrinsic harms that occur by adaption to a downstream task. Moreover, the aspect of centralization of power was mentioned as a problem, where only a specific part of society is able to access tools like an LLM. critically discussed the aspect of training data being one source of harmful and biased behavior of LLMs and depicted the financial as well as environmental impacts of large model training. Additionally, were able to extract potentially sensitive training data from LLMs, indicating that larger models tend to be more vulnerable to these attacks due to increased memorization. Another important issue is that LLMs can produce a text which sounds plausible to the reader but is incorrect or nonsensical. In a benchmark to measure truthful answers, larger models tended to score lower than their smaller versions with fewer parameters . Recently, Meta launched the Galactica LLM to support scientific writing. However, the API was taken offline three days after its release due to a backlash: The model produced seemingly correct-sounding but factually incorrect articles.
When using ChatGPT to simplify radiology reports, we expect the output to sound plausible, while its factual correctness and completeness are not guaranteed. In the next chapter, we present our approach for assessing the quality of ChatGPT-generated simplified reports.
Methods
To investigate the capabilities of ChatGPT for simplification of radiology reports, we conducted an exploratory case study (Figure 2) based on the research question: What is the radiologists’ opinion on the quality of simplified radiology reports generated with ChatGPT? We designed a questionnaire that included three fictitious original radiology reports written by an experienced radiologist and a unique simplified version of each generated by ChatGPT. We asked 15 radiologists to rate the simplified reports and analyzed the results of the questionnaire.
A radiologist with 10 years of experience wrote three artificial radiology reports. The artificial reports did not contain any sensitive personalized information. Each report contains multiple findings, which are associated with one another, e.g. tumor with edema, meniscus lesion with cruciate ligament lesion, and systemic metastasis. All three reports are of moderate complexity, based on a definition, where a report is of low complexity in case it contains no pathological or at most one explicit pathological finding, and of high complexity in case of many and especially unclear findings. The fictitious reports mimic real cases in clinical routine, i.e., the reports include previous medical information, describe the findings on the image, and contain a conclusion.
The first report Knee MRI (A.1.0) describes a case in musculoskeletal radiology. The case of a neuroradiological MRI of the brain is the subject of the second report Head MRI (A.2.0). It describes a follow-up examination and mentions comparisons to the previous examinations. The third report (A.3.0), referred to as Oncol. CT, describes a fictitious oncological imaging event, reporting a follow-up whole-body CT scan. To give an impression of the complex wordings in radiology reports, some excerpts of the original reports are shown in Figure 3.
2 Simplifying Reports using ChatGPT
The original radiology reports were simplified by prompting the ChatGPT (version December 15th, 2022) online interface with the request "Explain this medical report to a child using simple language:" followed by the original radiology report in text format.
This prompt was derived in a heuristic fashion using a separate fictitious radiology report that was different from the original reports described in Section 3.1. First, we tested different prompt versions of different lengths and compositions, including "Simplify", "Explain like I’m five", "Explain this medical report like I’m five", "Explain to a child using simple language", "Explain this medical report using simple language", and "Explain this medical report to a child using simple language". Based on several trials, we gained the impression that the latter prompt produced the best and most stable results and we perceived the language of the generated reports to be significantly simpler compared to the original report. We assume that this might be traced back to its detailed and explicit form, containing a concise call to action ("Explain"), specifying a narrow domain ("medical report"), and asking for a specific level of language complexity ("to a child using simple language"). During our tests, including phrases like "to a child" or "like I’m five" in the prompt on top of "simplify" or "using simple language" appeared to be beneficial for language simplification. While our final prompt choice "Explain this medical report to a child using simple language" worked well for our radiology report examples in this study, we do not claim that this version is necessarily the best option for the task of simplifying radiology report with ChatGPT.
The output of ChatGPT is not deterministic by default. In principle, the model outputs probabilities for the next token in the autoregressive generation procedure. However, the current settings of the ChatGPT interface do not allow tuning the temperature parameter responsible for handling the predicted token probabilities. To account for the variability in the text output of ChatGPT and to achieve good coverage of its generative capability, we prompted the model 15 times for each of the three original reports, respectively, i.e., we generated 15 different simplified reports per original report. Before each repeat of the prompt, we restarted the chat session to ensure that no tokens in the cache could distort the generated response. All 45 ChatGPT-generated simplified reports can be found in Appendix A.
3 Questionnaire
We designed a questionnaire to query radiologists on the quality of the simplified reports generated with ChatGPT. We defined the quality of radiology reports as the combination of (i) factual correctness, (ii) completeness, and (iii) harmfulness. Factual correctness was not further specified. By our definition, a complete report includes all key medical information, relevant to the patient. We specify harm as the potential that a subject could draw wrong conclusions from a simplification, which might result in any physical or psychological harm or lead to unwanted change in therapy or compliance.
The anonymized questionnaire "Questionnaire - Quality of Simplified Radiological Reports" (Appendix B.1) consisted of the three original radiology reports (section 3.1) and three unique simplified reports created with ChatGPT, respectively (section 3.2). On the front page, the participating radiologists were explicitly informed that the simplified radiology reports were generated with "the machine learning language model ChatGPT" and received a description of the setup of the questionnaire and an instruction on how to answer. Further, we asked for consent to participate and the years of experience, starting from the first year of residency.
Each questionnaire contained three blocks: (i) the original report, (ii) a single, randomly selected, unique simplified version of the original report, referred to as a "simplified report", and (iii) a series of questions to access the quality of the simplified reports. We asked the participants to rate their level of agreement with each criterion on a five-point Likert scale (formulated as a statement, respectively). Additionally, each question was accompanied by a follow-up question in which we asked the radiologists to provide text evidence for their assessment.
Factual Correctness: "The simplified radiological report is factually correct." Follow-up: "Highlight all incorrect text passages (if applicable) of the simplified report with a text marker".
Completeness: "Relevant medical information for the patient is included in the simplified radiological report." Follow-up: "List all relevant medical information, which is missing in the simplified report (if applicable)."
Potential Harm: "The simplified report leads patients to draw wrong conclusions, which might result in physical and/or psychological harm." Follow-up: "List all potentially harmful conclusions, which might be drawn from the simplified report (if applicable)."
15 radiologists with varying levels of experience from our clinic (Department of Radiology, University Hospital, LMU Munich) were asked to answer our questionnaire independently. Each radiologist received the same three original reports and different, unique simplified reports.
4 Evaluation
The questionnaires were collected and checked for consent and completeness. For each participant, the years of experience were recorded, respectively. The radiologists’ ratings on the Likert Scales for factual correctness, completeness, and potential harm were evaluated for each of the three cases (Knee MRI, Brain MRI, Oncol. CT). Statistical parameters for the ordinal scales were calculated: median, 25%-quantile (), 75%-quantile (), interquartile range (IQR), minimum, maximum, mean, and standard deviation (SD). The word count for the original reports and each simplified report was noted. For each report, all passages of the highlighted text, as well as answers in the free-text fields were transcribed manually to a spreadsheet (Table LABEL:app:tab:questionnaire_answers). Additionally, the percentage of free-text questions and text highlights where text evidence was provided by the participants was calculated. Finally, the free-text answers were inductively categorized by content.
Results
In this section, we present our evaluation of the questionnaires, which were answered by 15 radiologists with a median (IQR) experience of 5 (9) years. We analyzed their ratings on the Likert Scale, followed by the free-text analysis and the word count of the reports.
We first evaluated the radiologists’ ratings for all 45 simplified reports (see Table 1 and Figure 4(a)). The participants generally agreed (median = 2) with the statements that the simplified reports are factually correct and complete, respectively. For both quality criteria, 75% of all ratings were given for "Agree" or "Fully agree" (), while "Fully disagree" was not selected at all. In the case of completeness, we found a slight tendency towards more positive ratings compared to factual correctness. 25% of all completeness ratings were "Fully agree" (), more than for factual correctness. Further, "Neutral" and "Disagree" were only selected four times, but were selected ten times for factual correctness. Accordingly, the mean is slightly lower (closer to 1 = Fully agree) for completeness (mean = 1.8) compared to factual correctness (mean = 2.2). In line with these findings, the participants disagreed (median = 4) on the potential of wrong conclusions drawn from the simplified reports resulting in physical and/or psychological harm.
Compared to factual correctness (SD = 0.9) and completeness (SD = 0.7), the radiologists’ ratings of potential harm are more broadly distributed (SD = 1.0) with a considerable proportion of ratings for "Neutral" and "Agree". "Strongly Disagree" was not chosen at all.
In the second step, we also evaluated the radiologists’ ratings for each of the three cases (Knee MRI, Brain MRI, and Oncol. CT) individually. The results of the three reports are illustrated in Figure 4(b) and summarized in Table 2. The median of the ratings showed no relevant differences in factual correctness, completeness, and potential harm. However, there is a tendency that the Head MRI and the Oncol. CT reports are judged to have a higher risk of patients drawing wrong conclusions with potentially harmful consequences.
2 Free-text Analysis
In the following, we summarize the results of the follow-up questions. For each follow-up question, we applied inductive categorization of the free-text data by content. We present the identified categories along with a few selected examples for each. 51% of the participants highlighted incorrect passages, while only 22% and 36% of the radiologists listed missing relevant information and potentially harmful conclusions, respectively (Table 3).
We asked the participants to mark incorrect passages in the text (see highlighted passages in the simplified texts in Appendix A), which unveiled a variety of weak points in the simplified reports of ChatGPT.
We observed that in the simplified text generated by ChatGPT medical terminology was misinterpreted in several cases. The abbreviation "DD" (differential diagnosis) was often mistaken as a final diagnosis (A.2.6, A.2.8), e.g., the simplified report contains "The conclusion of the report is that the mass on the right side of the head is a type of cancer called a glioblastoma " for the original statement " DD distant GBM manifestation". Furthermore, "thyroid struma" was described by ChatGPT in one instance as "infection in their thyroid gland" (A.3.13) and in another as "extra thyroid gland" (A.3.7). The finding "No pneumothorax" was described as "no holes in the walls between the lungs and chest" (A.3.3). The lateral compartment of the knee was described as being on the "part of your body that’s on the outside" (A.1.15). The term "retroperitoneal" (referring to lymph nodes in need of further control) was wrongly interpreted as "in the person’s back" (A.3.9). Lastly, a growing mass of "currently max. 22 mm" was incorrectly simplified as a "small, abnormal growth that has gotten bigger" (A.2.3).
In some cases we observed that the automated simplification led to passages of imprecise language: The medial compartment of the knee was described as the "middle part of your leg" (A.1.5), the brain as the "head" (A.2.7), the abdomen as the "stomach" (A.3.7) and thyroid struma as "thyroid problems" (A.3.12). Partial regression was described as "gotten smaller and is not spreading as much as before" (A.3.12), intercondylar area as "the area between the two bones that make up the lower part of the leg" (A.1.4) and cartilage was described as "hard, smooth substance" (A.1.9). Metastases were imprecisely simplified as "spots" (A.3.5). A case of ambiguous simplification of the original text is found where the simple report ended with " is no evidence of the cancer spreading to other parts of the body." (A.3.13). In the original report pulmonary metastases were described.
Some findings were not contained in the original report, i.e. they were made up in the simplified reports, for instance, "no signs of cancer in the thyroid gland" (A.3.3) or "brain does not seem to be damaged" (A.2.9).
We observed odd and unsuited language in several simplified reports. The idiom "wear and tear" (A.1.4) was used to describe the degeneration of tendons and a radiological examination was circumscribed as "to see how the cancer is doing" (A.3.8). In report A.2.6 contrast media was paraphrased with the word "dye" and tumor as a "group of abnormal cells".
Grammatical errors such as "CT scan" instead of "CT scanner" (A.3.7) were the exception.
2.2 Missing Key Medical Information
Furthermore, the simplified reports often missed key medical findings mentioned in the original report.
For Knee MRI (A.1.0), a participant noted that the information on cartilage damage on the inner side of the knee was missing in the simplified report (A.1.5). In the Head MRI report (A.2.0) the information on "no intermediate or recent ischemia" (A.2.2) was missing in one instance. Another missed finding was the growth of the lesion on the parietooccipital side in the conclusion of the simplified report (A.2.4). For Oncol. CT (A.3.0) one participant reported that mentioning "spots" in contrast to metastases removes medical information. Furthermore the fact, that the solid portions of the pulmonary metastases are decreasing (A.3.5) - consistent with therapy response - is not included in the simplified report. In another simplified report (A.3.13), where the thyroid struma is misinterpreted as infection, the description of the enlargement is missing.
We observed missing or unspecific location information in the simplified reports. In the simplified report A.1.15 the location of the described finding (knee) was not mentioned. A participant mentioned that "the back" (A.3.14) is not precise enough to describe the affected retroperitoneum. Non-specific descriptions of the location resulted in problems with the identification of the present tumor and the excised tumor (A.2.11).
2.3 Potentially Harmful
Finally, we asked the participants to list all potentially harmful conclusions, which might be drawn from the simplified report. The given answers corresponded mostly to already highlighted passages of incorrect text or missing key medical information and corroborated the detected weaknesses. Odd language and grammatical errors were not considered potentially harmful.
Participants listed the misinterpretation of differential diagnosis ("DD") for final diagnosis as potentially harmful for the patient, e.g., "GBM is one (likely) DD which implies that other DDs exist (e.g. radionecrosis)" (A.2.7). Lymph nodes were presented in the simplified report like "they might have cancer"(A.3.9). This was considered harmful since the original report stated that there was "no evidence of recurrence or new lymph node metastases". Furthermore, the wording "small growth" (A.2.3) was deemed a harmful conclusion, as the original report Head MRI (A.2.0) states a progression in size, which "almost doubled".
Participants assessed the hallucination "brain does not seem to be damaged" (A.2.9) as potentially harmful to the patient, as the original report describes a growing mass.
We observed that imprecise language highlighted as incorrect was sometimes also graded as potentially harmful to the patient. One participant noted that "there is no (new) spread of cancer" in the simplified oncological report (A.3.13). Additionally, participants listed these potentially harmful conclusions: "not clear that spots are pulmonal metastasis" (A.3.5), "changes happened recently (how recently?)", and "some extra fluid sign. incr. edema" (A.2.13).
We found that potentially harmful conclusions can originate from missing information in the simplified report. The sentence "the parts that are staying the same size are changing" (A.3.5) was found to be an unclear statement missing interpretation. One participant stated that "information about the lesion showing growth" is missing from the conclusion of the simplified report (A.2.4).
Cases, where ChatGPT misses the correct information about a location of a disease, can lead to potentially harmful conclusions. Examples were: "misunderstanding which lesion is stable and which one is in progress can lead the patient to some wrong expectations" (A.2.11) and "patient might be misled to look at cancer spots on their back" (A.3.14).
3 Word Count
The median word count for the simplified Head MRI and Oncol. CT reports were comparable to their respective original reports (Table 4). The simplified Knee MRI report with a median (IQR) word count of 414 (60) is generally longer compared to the original report (word count = 222).
Discussion
In this exploratory study, we evaluated the opinion of radiologists on the quality of radiology reports which were simplified by ChatGPT. Participating radiologists rated the simplified radiology reports according to (i) factual correctness, (ii) completeness, and (iii) potential harmfulness. In the following, we discuss the implication of our results presented in Section 4.
We prompted ChatGPT to generate 45 simplified radiology reports for this study. According to the authors of ChatGPT, the model is in principle able to state to not know the answer . However, in no instance, the model refused to give an answer or indicated that the response could be incorrect while generating consistent plausible-sounding outputs.
Although ChatGPT was not explicitly trained for the purpose of simplifying radiology reports, most radiologists agreed that the generated simplified reports are factually correct (median = 2, Section 4). However, some radiologists also identified statements that were not considered factually correct (Figure 4). In particular, we identified the following error categories: misinterpretation of medical terms, imprecise language, hallucination, odd language, and grammatical errors.
On the one hand, we speculate that the problem of misinterpretation of medical terms, imprecise language, odd language, and grammatical errors may be solved by the adoption of a ChatGPT-like model for the medical domain, as well as by optimizing the prompt. On the other hand, we also observed hallucination, which is a type of behavior where the model adds statements to the simplified report that cannot be derived from the original report (Section 4.2.1). This is an intrinsic problem in generative models like LLMs .
1.2 Completeness
Prompting ChatGPT for simplification is likely to cause responses, where content is removed or rephrased in oversimplified terms. Therefore, we tasked radiologists to verify the completeness of the simplified reports. Overall, the radiologists agreed that the simplified reports contained the medical information relevant for the patients (median = 2, Section 4.1). This indicates that ChatGPT is able to identify the most important aspects of the complex medical content of radiology reports. Nevertheless, the participants also detected missing key medical information, which we summarized in the two categories missed findings and unspecific location (Section 4). These findings suggest that paraphrasing and explaining complex medical terminology in more simple terms can have the side effect of losing medical context and professional preciseness.
It is an open question if the problem of missing key medical information could be solved by adapting the prompt or if this is an unreachable goal for the current version of ChatGPT.
1.3 Potential Harm
Non-maleficence, i.e., to not harm individual patients, is a major principle of medical ethics, also in the context of AI applications . Thus, we asked the radiologists if they think that "The simplified report leads patients to draw wrong conclusions, which might result in physical and/or psychological harm.”
Overall, radiologists disagreed that potentially harmful conclusions will be drawn from the ChatGPT-generated reports (median = 4, Table 1). This finding may encourage patients to use ChatGPT or similar LLMs as a tool to simplify radiology reports. However, as shown in Section 4.2.1 and 4.2.2, simplified reports can contain errors and key medical findings might be missed. About half of the radiologists found critical statements in the simplified reports with a high potential of leading patients to wrong conclusions, which bears the great risk of causing physical or psychological harm (Section 4.2.3). For example, the medical finding that certain lymph nodes "might have cancer" is an overstatement in the simplified report A.3.9. The original report clearly states that there is "No evidence of recurrence or new lymph node metastases" and only suggests to "further control" the "Unchanged prominent retroperitoneal lymph nodes". This might cause psychological harm since the patient could conclude that the CT shows hints of new lymph node metastases. Another example can be found in the simplified report A.2.3. Here, only a "small growth" of the contrast-enhancing mass is described, whereas the original report (A.2.0) states a progression in size, which "almost doubled". This is a clear understatement, which could lead the patient to underestimate the progression of the disease.
While, overall, radiologists do not see a high risk for potential harm, they still point out that single errors occur which might be potentially harmful. Users of ChatGPT should be aware of this problem.
2 Opportunities and Challenges
With openly available LLMs like ChatGPT, patients might use these tools to simplify their radiology reports in the near future. This scenario raises a number of opportunities and challenges.
Improved accessibility of personal medical information may be beneficial to the patient in multiple ways. Simplified reports give patients the chance to better comprehend their health situation instead of making them feel overwhelmed by the medical terminology found in the original report and might help patients to better prepare themselves for future doctor-patient interactions. For instance, the simplified report might raise questions patients could then ask the medical practitioner. This might lead to an improved understanding of their own health condition which, in turn, empowers the patients, strengthens their position throughout the medical treatment, and facilitates informed patient decisions. Simplified radiology reports are particularly valuable for patients who do not speak the native language of the country in which they are treated. Even though ChatGPT works best for English, other languages are supported with varying quality. In the case of poorly supported languages, reports could be translated into English using established tools prior to simplification. Overall, easily accessible simplified radiology reports may help to increase patient autonomy and allow them to take on a more active role throughout their medical treatment process facilitating patient-centered care.
However, when patients simplify their reports in advance of a medical consultation, patients will not be assisted by an expert while reading and attempting to interpret them. Without medical advice, patients could misinterpret certain findings or even learn about a life-threatening diagnosis from their simplified report, but the information might not be accompanied by additional context, e.g., therapy options. This could put the patient in a state of mental stress. Additionally, there might be an increased risk of patients making their own clinical decisions, such as delaying or even omitting further doctor appointments or terminating a therapy without professional medical consultation. Even though the language conveying the information might be simplified, the content of radiology reports and the process of clinical decision-making remains complex.
In this study, we found that ChatGPT already performs surprisingly well at simplifying radiology reports out of the box. It produces fairly accurate and complete simplified reports. If this technology continues to become even more easily accessible for laypersons, it has the potential to be applied by a wide range of users in the near future. To fully leverage this tool as a radiology report simplifier with maximum benefit for patients, several technical hurdles need to be overcome.
Among others, we see the following challenges: ChatGPT was not developed for the specific task of simplifying radiology reports. A domain adaption of ChatGPT to the medical field could further improve the quality of generated simplified radiology reports, e.g., leading to a more correct interpretation of medical terms. The task of simplification could be enforced by using a reward model trained to estimate the quality of simplified radiology reports in the RLHF training procedure of ChatGPT. Like other LLMs, ChatGPT might have intrinsic biases due to imbalanced training data . We hypothesize that for the task of simplifying radiology reports, rare pathologies might be handled less accurately than more common ones. Furthermore, medicine is an active and dynamic field of science, which is subject to progress and change. However, ChatGPT - like most LLMs - is trained on data with a certain timestamp . New research findings in medicine are thus not reflected in the model’s output. Another problematic property of ChatGPT is the non-deterministic output. ChatGPT generates text by predicting tokens on a probabilistic basis, which does not guarantee the most likely response. While this fosters dynamic dialogue, it could also lead to high variance with respect to output reliability. As of now, there is no way for users to control ChatGPT, e.g., with a temperature parameter or seed. This makes it difficult, if not almost impossible, to reproduce a certain output, even when applying the same input multiple times. Finally, there are privacy concerns when sensitive data is uploaded to a proprietary service.
Simplified radiology reports have the potential to increase patient autonomy during the treatment process. With the release of ChatGPT, patients now have unprecedented possibilities to automatically generate such simplified reports themselves. However, there is an unavoidable risk for patients to draw potentially harmful conclusions from those reports, when no medical oversight is provided.
We envision a future integration of ChatGPT-like LLMs directly in the clinic or radiology centers. In this scenario, a simplified radiology report would always be automatically generated with an LLM alongside the original report, proofread by a radiologist, and corrected where necessary. Both reports would then be issued to the patient. We imagine the application of an appropriate LLM that is adapted to the medical domain, properly certified, and hosted in a privacy-preserving way, e.g., on-premise. In our opinion, this would be a cost-efficient way to leverage state-of-the-art technology for improving patient-centered care.
3 Limitations of the Study
The number of original radiology reports (n = 3) and experts (n = 15) for assessing the quality of the simplified reports is small. Further, we used fictitious radiology reports written by an experienced radiologist instead of original reports of patients to protect patient privacy. The radiologist who generated the simplified reports is German, i.e., not a native English speaker. A German abbreviation ("Z.n." for "Zustand nach", which can be paraphrased as "following"), which does not exist in English, was found in the original reports. The original reports used in this study are at a medium complexity level. It is unclear how the quality of less or more complex reports would be rated. Also, the language simplicity of ChatGPT-generated radiology reports was analyzed only qualitatively during the design process of the input prompt ("Explain this medical report to a child using simple language"). We did not further investigate the ability of ChatGPT to simplify language with respect to established guidelines for simple language. We do not claim that our prompt for triggering a simplification of radiology reports through ChatGPT, is optimal for this task. Our prompt is the result of a heuristic evaluation of experimenting with different inputs. Finally, the quality of the simplified reports generated with ChatGPT was only evaluated from a radiologist’s point of view. The patients’ perspective on the readability, accessibility, and added value of those simplified reports was not part of this study.
Conclusion
In this case study, we asked radiologists to assess the quality of simplified radiology reports generated with ChatGPT. We found that most radiologists agree that the simplified reports are factually correct, complete, and not potentially harmful to patients. This positive vote illustrates the great potential that LLMs could have for medical text simplification. In particular, we argue that simplified reports generated with LLMs like ChatGPT would strengthen patients’ autonomy. However, the radiologists also identified factual incorrect statements, missing key medical information, and text passages, which might lead patients to draw potentially harmful conclusions from the simplified reports, if physicians are not kept in the loop. We hypothesize that LLMs like ChatGPT may be used by patients to simplify radiology reports that contain critical medical information, even though the models were not developed for this specific medical use case. Thus, we see a need for technical improvements and fine-tuning of ChatGPT or similar LLMs to the specific task of simplifying radiology reports.
Our exploratory case study can only give first insights into the field of automated simplification of radiology reports with ChatGPT, especially due to the limited number of participants and radiology report cases. Given the great potential of LLMs for simplifying radiology reports, there is a demand for future research to validate the findings of our case study and to further explore the many possibilities of this new technology in the medical domain. This requires quantitative studies on the quality of simplified radiology reports generated with ChatGPT and similar LLMs. Also, the opinion of patients regarding the accessibility of those simplified reports and their overall added value within the treatment process should be investigated. We propose to include the automatic generation of simplified radiology reports in the clinical process based on a domain-adapted model. These should be approved by experts and issued to patients alongside their original report. Overall, we see great potential in using LLMs like ChatGPT to improve patient-centered care in radiology and other medical domains.
Acknowledgments
We thank all radiologists who participated in this study. This work has been partially funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) as part of BERD@NFDI - grant number 460037581.
References
Appendix A Appendix
In Appendix A, we summarize the original reports and the simplified reports generated by ChatGPT. The incorrect text passages (question 1a) identified by participants were highlighted in the simplified reports.