Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

Jungo Kasai, Yuhei Kasai, Keisuke Sakaguchi, Yutaro Yamada, Dragomir Radev

Introduction

Much recent work builds approaches to various language tasks on the impressive abilities of large language models (e.g., GPT-3 and 4, Brown et al., 2020; OpenAI, 2023; BLOOM, Scao et al., 2022). Large language models (LLMs) can perform downstream tasks by conditioning generation on a small number of task demonstrations or task instructions, without the need for any parameter updates Brown et al. (2020); Wei et al. (2022a); Su et al. (2023). Recent work (Kung et al., 2022; Choi et al., 2023; Kung et al., 2022; Nori et al., 2023; Bubeck et al., 2023, inter alia) started to evaluate the performance of these LLM-based approaches across diverse specialty areas and disciplines, beyond conventional natural language processing benchmarks, such as the GLUE tasks Wang et al. (2018, 2019), semantic parsing Zelle and Mooney (1996); Yu et al. (2018), and grammatical error correction Napoles et al. (2017); Coyne and Sakaguchi (2023). These evaluation results suggest that LLMs have the potential to transform many applications and industries in the future.

Meanwhile, many of these diverse benchmarks are limited to the English language. While the training data for LLMs are often English-centric Brown et al. (2020); Wang and Komatsuzaki (2021); Zhang et al. (2022); Touvron et al. (2023), many LLMs exhibit multilinguality. For example, ChatGPT and GPT-4 have shown competitive performance on the WMT machine translation benchmark for some language pairs Jiao et al. (2023); Barrault et al. (2020). LLM-based services, such as ChatGPT, are now used by non-English speakers every day. We thus argue that it has become increasingly critical to benchmark LLMs on diverse specialized domains in non-English languages. This work makes a first step towards this goal. Specifically, we evaluate LLMs (GPT-3 and 4 and ChatGPT) on Japanese medical lincensing examinations from the past five years (2018-2023), including the current year, and release the data as the Igaku QA (医学 QA) benchmark. The exam takes place every year for final-year medical school students in Japan. It consists of 400 multiple-choice questionsA couple of arithmetic questions directly ask for numbers (e.g., daily urine protein excretion) rather than correct choices. More detail and the passing criteria are discussed in §2. and covers a wide range of topics in medicine and public health. Final-year students who pass the exam are given the Japanese medical license and enter a two-year residency program as doctors.

Importantly, our Igaku QA benchmark does not rely on any translation of English resources and comes solely from the Japanese medical licensing examinations. Our project is led by native Japanese-speaking NLP researchers and a practicing cardiologist based in Japan. Many previous benchmarks in non-English languages are created by translating exisitng English datasets Conneau et al. (2018); Artetxe and Schwenk (2019); Artetxe et al. (2020); Lewis et al. (2020); Longpre et al. (2021); Ponti et al. (2020). Evaluation on such datasets can introduce serious problems. The content of these translation-based datasets often becomes English-centric and substantially diverges from actual use cases by native speakers of the target language Mohammad et al. (2016); Clark et al. (2020); Asai et al. (2021, 2022). This divergence becomes particularly severe in medical applications. For instance, medical practice should follow the rules and law in the country. As illustrated by an exam problem in Fig. 1 (the original problem and the ChatGPT output were in Japanese but translated by the authors for readability), euthanasia is illegal in Japan, and doctors should not suggest it in their clinical practice in Japan (Choice b in Fig. 1). However, we observed that ChatGPT chose this option. Indeed, in the Japanese medical licensing exam, 20+ choices are considered prohibited (禁忌), and test takers who choose four or more prohibited choices will fail regardless of their total scores (see more detail for evaluation criteria in §2.2). Medical practice also requires knowledge about specific statistics or systems in the country (e.g., what are the responsibilities and duties of a public health center (保健所) in Japan?). Our approach avoids these potential pitfalls in designing evaluation for non-English languages and provides useful evaluation data to the research community.

Our experiments (§3) show that unlike the previous language models, GPT-4 can successfully pass the Japanese medical lincensing examinations over the past five years, including the current year. This result suggests the potential of non-English AI applications in medical support, education, and assessment as LLMs continue to improve in the future. Nonetheless, GPT-4 still substantially underperforms the majority-vote performance among the medical school students. Moreover, though the results on Japanese are as promising as the recent findings on the United States Medical Licensing Examination (USMLE; Kung et al., 2022), there are significant limitations in Japanese (and similarly distant languages): increased API costs and smaller context window sizes due to tokenization and lack of customization specific to the country (§3.3) We hope that our evaluation results and Igaku QA benchmark will spur further research on clinical applications of LLMs, especially in non-English languages in the world.

Background

In this section, we briefly discuss the medical licensure process in Japan and its difference from the US system (§2.1). We then describe the exam structure, evaluation criteria, and topics that are covered (§2.2), as well as our Igaku QA collection process (§2.3). More example problems will be presented in §3.3.

Fig. 2 illustrates the standard timeline of the medical licensing process in Japan in comparison to the United States. Students in Japan typically take the National Medical Practitioners Qualifying Examination (NMPQE) in their final year of the six-year medical school education. The exam covers a wide range of topics and assesses students’ knowledge about clinical and social medicine and public health. Note that hands-on clinical exposure typically happens after passing the exam and obtaining the license. This differs from the US system, where the licensing process consists of three steps (Step 1, foundational sciences; Step 2 CK, clinical knowledge; Step 3, generalist medical practice) and students enter a residency program during this process. For more comprehensive discussion and the historical context of the difference between the United States and Japan, see Kuwabara et al. (2015).

2 Details of NMPQE

Since 2018,The exams from 2017 or earlier had Parts A-I with more problems in total. the Japanese medical licensing exam is structured in the following format: it consists of Parts A-F, each comprising 50-75 multiple-choice questions with five answer choices, totaling to 400 questions in all. There are some problems that require selecting two or three choices, in which case all choices need to be correctly selected to earn the point(s). Note that there are a few exceptions: a small number of problems are arithmetic questions that ask for numbers directly or contain more than five choices. In 2022, a total of 10,061 people took the exam, and 91.7% of them passed.https://www.mhlw.go.jp/stf/shingi2/0000197611_00004.html. See Appendix §B for more exam statistics.

In the multiple-choice questions, 25+ choices are marked as prohibited choices (禁忌肢). These are choices that correspond to decisions that should be strictly avoided in medical practice in Japan. For example, euthanasia is illegal in Japan and doctors are not allowed to suggest it in their medical practice (see Choice b in Fig. 1). Similarly, when a patient desires to have children in the future and there is a viable alternative, a total hysterectomy is considered as a prohibited choice.

Evaluation Criteria

Parts A-F split into two sections: required (Parts B and E) and general sections (others). Regarding the required section, one point is awarded for each general question, and three points for each clinical practical question. For the general section, one point is awarded for each question.Although the total number of problems remains the same from year to year, a few questions are often disregarded due to their difficulty or ambiguity.

A Student passes the exam if and only if all the following three criteria are satisfied:

The score on the required section is 80% of the total score or higher.

The score on the general section is 70% of the total score or higher.

Only up to three prohibited choices (禁忌肢) may be selected.

Categories

Fig. 3 plots the numbers of problems broken down by category from 2022. The categorization is based on exam preparation books that are widely used by Japanese medical students 国試対策問題編集委員会 (2022), and it has been confirmed by the second author of this paper, who is a native Japanese speaker and a practicing cardiologist in Japan. The exam problems span 28 categories that cover a wide range of topics in medicine, including public health, cardiology, psychiatry, and obstetrics.

Problems with Images

Naturally, some problems (∼\sim25%) contain images (e.g., X-ray photographs in a clinical case problem), though not all of them strictly require images to answer. Most of the large language model APIs currently available do not take as input images (including GPT-3/4 APIs), and thus some problems cannot be answered by design without using images. We still include these problems in our benchmark and encourage researchers and practitioners to develop LLMs that can work in multimodal settings.

3 Benchmark Collection

We collect the exam problems and their answers in the past five years (from 2018 through 2023), including the current year, from the official website of the Ministry of Health, Labour and Welfare in Japan.https://www.mhlw.go.jp/kouseiroudoushou/shikaku_shiken/ishi/. We also collect additional metadata, such as the percentage of test takers who selected each choice, as well as the average accuracy of the test takers, based on the exam preparation books 国試対策問題編集委員会 (2018, 2019, 2020, 2021, 2022).We have yet to get access to the 2023 metadata. We open-source all problems with their answers as a benchmark.https://github.com/jungokasai/IgakuQA.

Notice that we do not rely on any translation of sources from other languages (e.g., English) or countries, and the benchmark comes solely from resources that are originally written in Japanese. This avoids potential problems that many translation-based datasets have. First, the content of the benchmark is aligned with actual usage in the target language and country; this helps us better understand model behaviors or failures in a more realistic way. Moreover, since all problems are originally written in Japanese, we mitigate the risk of the translationese effect: translated text differs from naturally-occurring text lexically, syntactically and stylistically, resulting in divering evaluations Baker (1993); Lembersky et al. (2011); Graham et al. (2020).

Experiments and Analysis

We benchmark popular LLM APIs that are currently available as of March 31, 2023 on our Igaku QA dataset. For simplicity, all of the models are used in a closed-book setting where no external resources are provided. We leave it to future work to extend baselines to open-book settings.

We experiment with three LLM APIs: GPT-3 (text-davici-003; Brown et al., 2020), ChatGPT (gpt-3.5-turbo),Concurrent work manually runs and evaluates ChatGPT on the 2023 exam Kaneda et al. (2023). and GPT-4 OpenAI (2023). Details of their training data and architecture are not well documented, but it is believed that they are autoregressive language models built upon the transformer architecture Vaswani et al. (2017). Recent work has begun to evaluate these LLM APIs on diverse benchmarks in English beyond standard natural language processing datasets (Choi et al., 2023; Nori et al., 2023, inter alia). We follow these efforts but focus on the Japanese language, which is typologically distant from English (script and word order etc.). While they are designed primarily for applications in the English language, recent studies have shown that they can be used in non-English languages as well Bang et al. (2023); Jiao et al. (2023).

Prompting and Output Formatting

LLM APIs are used with prompts for various downstream tasks, and previous work demonstrated that different prompts result in different downstream performances Liu et al. (2021). By default, we use a simple prompting method. Fig. 4 (GPT-3) and Fig. 5 (ChatGPT and GPT-4) illustrate our prompts. We use three in-context examples randomly sampled from the Japanese medical licensing exam in 2006. See Appendix §C for more detail about the three in-context examples.

In addition to these simple prompting baselines, we explore the following alternative with ChatGPT.ChatGPT’s API cost is significantly lower than GPT-3 and GPT-4, so we used it to explore differen methods. ChatGPT-EN first translates the problems and answer choices into English, followed by inference in the English language. We also explored a prompting method with intermediate steps that have proven successful in various tasks, including multi-hop question answering Wei et al. (2022b); Press et al. (2023). However, similar to the findings in law school exams where Chain-of-Thought prompting did not improve performance Choi et al. (2023), we did not find any improvements from adding intermediate steps of explanations on ChatGPT.We release all model outputs in our experiments and leave it to future work to explore better methods to produce intermediate reasoning steps. Lastly, we also provide the Student Majority baseline that picks the choice(s) selected by the highest percentage of test takers.

Evaluation Methods

We perform automatic evaluations by exact matching. As discussed in §2.2, almost all problems are multiple-choice questions with a few exceptions that require numbers. Exact matching is a reliable metric on Igaku QA since there are no free-form answers, contrasting with open-ended generation tasks that often require human evaluations or advanced metrics Kasai et al. (2022a, b); Khashabi et al. (2022); Hu et al. (2023). There were a small number of cases where an LLM fails to follow the format specified by the in-context examples (e.g., outputting text, instead of choosing an option). Note that this formatting issue was limited in our case, but there are ways to force a strict answer format, which later work can explore Nori et al. (2023)

2 Results

Seen in Table 3.2 are the results from the Japanese medical licensing examinations from the past five years (2018-2023). We see a consistent trend over the five years: GPT-4 achieves the best performance, followed by ChatGPT/ChatGPT-EN/ChatGPT-Exp and GPT-3. Moreover, GPT-4 and ChatGPT-EN are the only ones that do not select more than three prohibited choices over the five years. GPT-4 manages to pass the exam in all six years but substantially underperforms the student majority baseline.

ChatGPT-EN outperforms ChatGPT to a certain degree in the majority of cases, suggesting limitations of LLMs’ multilinguality when translation is not done explicitly.

Prohibited Choices

As shown in Table 3.2, unlike the student majority baseline, the LLMs sometimes select prohibited choices (§2.2) that should be strictly avoided in medical practice. Fig. 6 shows one of those problems. ChatGPT chooses Choice e “use oral hypoglycemic agents if dietary therapy is ineffective,” but this is considered as a prohibited choice; there are significant concerns about suggesting the use of oral hypoglycemic agents during pregnancy due to the potential dangers, including fetal teratogenesis, hypoglycemia, hyperbilirubinemia, and polycythemia Sutherland et al. (1974); Langer et al. (2000); Kavitha et al. (2013). This result demonstrates critical challenges when LLMs are applied to specialized, high-stakes applications, such as medicine, finance, and law.

Geographic and Temporal Context

We also found several problems require geographic and/or temporal context, departing from conventional question answering datasets. For instance, the problem in Fig. 7 requires Japan-specific knowledge. Open-book approaches or retrieval augmentation can be used to further improve the performance on these problems Kasai et al. (2022c). While evaluation on geographic or temporal context is not the main focus of the Igaku QA benchmark, it is one of the challenges that large-scale question answering systems face in real-world applications Zhang and Choi (2021); Jang et al. (2022a, b); Liška et al. (2022); Kasai et al. (2022c).

GPT vs. Medical Students

Fig. 8 compares the student accuracy (the ratio of the students who select the correct choice(s)) and the GPT-4 result ( green: correct; red: wrong) for each problem from 2022. We find correlation between the student accuracy and the likelihood of the correct prediction, suggesting that GPT-4 struggles on questions that are also difficult for humans. We see similar patters from other models (Appendix §C).

We evaluated GPT LLM APIs on the Japanese national medical licensing examinations that students take at the end of their six-year medical school education. Here we discuss the connections to well-studied clinical natural language processing (for non-English languages in particular), multilingual language modeling, and open-domain question answering.

Similar to many other applications of natural langauge processing (NLP), English is by far the most resource-rich language in clinical NLP Névéol et al. (2018). For example, many advanced NLP tools, such as part-of-speech taggers (Smith et al., 2005; Tsuruoka et al., 2005; Divita et al., 2006, inter alia), are developed for biomedical applications in the English language. Some efforts in clinical NLP for non-English languages include: core NLP models and pipelines (e.g., parsing Nishimoto et al. (2008), abbreviation/vocabulary expansion Shinohara et al. (2013); Ahltorp et al. (2016), question answering Ito et al. (2016), and pretrained transformers Wada et al. (2020); Kawazoe et al. (2021) for Japanese biomedical text); datasets and resources Rebholz-Schuhmann et al. (2013); Neveol et al. (2014); Aramaki et al. (2014); Kors et al. (2015); and crosslingual transfer Deléger et al. (2009); Papaioannou et al. (2022). As LLMs and generative models become increasingly powerful and popular among English speakers and speakers of other languages like Japanese, evaluations of these models should be diversified accordingly. Benchmarks that have been developed to assess the qualifications and skills for human experts, such as bar or medical licensing examinations, can be useful in this regard. For a more comprehensive survey on clinical NLP in languages other than English, see Névéol et al. (2018).

Multilingual Language Models

Much recent work on multilingual NLP hypothesized that although each language is unique, different languages manifest similar characteristics (e.g., morphological, lexical, syntactic) which can be exploited by training a single, polyglot model with data from multiple languages Ammar (2016). This polyglot approach has proven successful in various NLP tasks, including syntactic dependency parsing Ammar et al. (2016), semantic role labeling Mulcaire et al. (2018), named entity recognition Xie et al. (2018), and language modeling for phonetic sequences Tsvetkov et al. (2016) and for speech recognition Ragni et al. (2016). More recently, researchers developed multilingual pretrained language models Mulcaire et al. (2019b, a); Xue et al. (2021); Liu et al. (2020) that can be used for machine translation or crosslingual transfer in downstream tasks. Though there are variants that use crosslingual supervision (e.g., Lample and Conneau, 2019), many of these polyglot models can benefit from joint training of different languages without any explicit supervision. We suspect that similar polyglot language modeling is happenning in LLMs, such as ChatGPT and GPT-4, which we tested on our Igaku QA benchmark in Japanese, a language typologically distant from English.

Open-Domain and Multilingual Question Answering

Much prior work proposed datasets for open-domain QA for English and beyond Clark et al. (2020); Asai et al. (2021, 2022); Longpre et al. (2021); Zhang et al. (2021). Several works pointed out the problem of translation-based question answering evaluations Clark et al. (2020); Asai et al. (2021): questions raised mainly by English speakers can diverge from information needs from speakers of other languages. For instance, these translation-based benchmarks can overly represent English-centric topics, such as American politics, sports, and culture. To mitigate this English-centric problem, some datasets only sample questions from native speakers of each langauge Clark et al. (2020). Consistent with such data creation methods, our Igaku QA consists of problems that are written by native Japanese speakers to evaluate the qualifications and skills for medical practice in the country.

We presented our evaluations of the GPT APIs on the Japanese medical licensing examinations from 2018 to 2022. The newest model, GPT-4, outperforms the others and manages to pass the examinations. Through our benchmark, we highlighted several important limitations of the current LLM APIs when they are applied to a specialized domain in Japanese, a language typologically distant from English. We open-source our benchmark as Igaku QA, as well as the model outputs and meta information for future research.

This work evaluates large language models on Japanese medical licensing examinations. We highlight several core limitations of our evaluations: reproducibility and potential data leakage, language coverage, and scope of evaluation.

First, as our experiments are performed using black-box LLM APIs, our results are not fully reproducibile, and the results may change with updates in the APIs. Further, since the language model training data and setups are not well documented, there are potential risks of data leakage that overestimates the performance of LLMs. To mitigate these issues, we release all model outputs and experimental settings as well as the Igaku QA benchmark. This way if there is any update in the APIs, we can easily update our results and analyze changes in behaviors after the update. We have also included results from the current year (March 2023), which we believe is after the training of GPT-4, to address potential data leakage. We observed consistent performance with the previous years.

Clearly, our benchmark is limited to the Japanese language and Japanese medical licensing examination. It is an important research avenue to explore evaluations in more languages and domains. Nonetheless, evaluation in the medical domain requires expertise, including knowledge specific to the country and its medical system and standard medical practice. As discussed in this paper, there are potential risks if benchmarks are simply translated to various languages. The second author of this work is a doctor in a Japanese hospital, and such interdisciplinary efforts are necessary.

Lastly, we note limitations in the scope of our evaluations. For example, we did not use image information during evaluations because the current OpenAI LLM APIs do not support image input. While some problems with images can be solved based solely on the problem text, many problems with images (and, of course, medical practice in general) need multimodal reasoning. We leave it to future work to test models in multimodal settings.

Despite these challenges, we believe that it is important to benchmark black-box LLMs; they are increasingly used by people around the world across various disciplines. We hope that our evaluations and Igaku QA benchmark will contribute to a better understanding of their behaviors, failures, and potential risks and benefits in diverse areas.

The day after I completed this manuscript, I was eagerly awaiting your usual email and feedback. To my great shock, I received the unexpected and sad news of your sudden passing. I went to your office at Yale to leave white lilies, which I believe symbolize the purity of your life-long commitment to mentorship, education, and research.

LILY (Language, Information, and Learning at Yale) is also the name of your NLP lab at Yale University. I am extremely fortunate to be part of the LILY lab since the beginning in Spring 2017. I still vividly remember the day I visited your office for the first time. At the time, I was in my senior year with almost no prior research experience. Despite this, you kindly offered to mentor me on my research project. After graduation, I sought your advice as to what I should do next. Soon after, you and Professor Bob Frank from Yale Linguistics very kindly secured funding and offered me a research assistant position. This experience became the foundation of my NLP research career. During my Ph.D. at the University of Washington, we continued to meet regularly and collaborate on many exciting projects.

Among many other things, your attitude towards research has always struck me as passionate and open-minded. The field of NLP has experienced many changes since I started. We used to talk a lot about building core NLP models using LSTMs. Now, we are seeing tremendous progress from large language models. You always showed great enthusiasm for the latest advancements and how they are transforming the way we approach NLP problems. Your passion for this field was contagious. You consistently encouraged me to be open to new ideas, even when I was skeptical or anxious about new directions. As you led by example, no matter how the field changes, researchers have a responsibility to demonstrate limitations and potentials of new technologies for society.

As one of the researchers who were extremely lucky to have you as an advisor, I feel obliged to pay it forward by continuing to support younger generations of scholars. Thank you so much for the amazing six years. May you find eternal rest and peace. ご冥福をお祈りいたします。

We thank Noriyuki Kojima and Koji Shiono for their helpful feedback on this work.

Appendix A Cateogries

Our categorization of the exam problems is based on books widely used among Japanese medical students 国試対策問題編集委員会 (2018, 2019, 2020, 2021, 2022). Table 2 presents the Japanese-to-English translations of the category names. We have 28 categories, ranging from public health to anesthesiology

Appendix B Additional Exam Information

Fig. 9 plots the passing rates of the Japanese medical licensing examination in the past five years. The exam is typically taken by final-year medical students, and they obtain the Japanese medical license after passing the exam. As shown in the figure, the passing rate slightly varies from year to year but are generally high (around 90%).

Breakdown by Category

Figs. 10-13 plot the numbers of the problems over the 28 categories from year 2018 to 2021. The four years all have a catgory distribution similar to the exam in 2022 (Fig. 3). Different from the United States Medical Licensing Examination (USMLE), where three steps are taken over years, the Japanese medical licensing examination is usually taken once by final-year medical students. See also Fig. 2 for the standard timelines.

Appendix C Additional Settings and Results

Similar to Fig. 8, Figs. 14-16 compare the student accuracy (the ratio of the students who selected the correct choice(s)) and the ChatGPT/ChatGPT-EN/GPT-3 results (green: correct; red: wrong) for each problem from 2022. We find correlation between the student accuracy and the likelihood of the correct prediction, suggesting that the LLMs struggle on questions that are also difficult for humans.

In-Context Examples

Table 3 our three in-context examples (translated into English for readability) for GPT-3 and ChatGPT/GPT-4. All of the three in-context examples were sampled from the Japanese medical licensing exam in 2006, which is also available on the official website of the Ministry of Health, Labour and Welfare of Japan (厚生労働省).https://www.mhlw.go.jp/topics/2006/04/tp0419-1.html. All our prompt templates are available online.https://github.com/jungokasai/IgakuQA..