Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset

Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, Michael Lingzhi Li

Introduction

Recent advancements brought by large language models (LLMs) such as T5 (Raffel et al., 2020) and GPT-4 (OpenAI, 2023) have revolutionized natural language processing (NLP). However, evaluating LLMs in the medical field poses significant challenges due to the paucity of standardized and comprehensive datasets compiled from reliable and unbiased sources (Li et al., 2023). Most existing medical datasets (Hendrycks et al., 2020; Abacha et al., 2019b; Li et al., 2023; Zhou et al., 2022) for language model evaluation have limitations that hinder comprehensive assessment of LLM performance (Nori et al., 2023). Many datasets are insufficient in terms of size and diversity, preventing a thorough evaluation of LLM capabilities. Furthermore, most datasets primarily focus on text generation tasks rather than utilizing clear choice evaluations, impeding objective and quantitative measurement of LLM performance. Additionally, a majority of these datasets (Li et al., 2023; Pal et al., 2022; Zhu et al., 2020) are sourced from online forums and consumer feedback, which could suffer from significant bias and error. These challenges are particularly amplified in non-English languages, such as Chinese, due to the pervasive inequality in language resources that exists in the NLP field (Bird, 2020; Zeng et al., 2022; Fang et al., 2023). Overall, due to the lack of qualified evaluation datasets, the strengths and weaknesses of LLMs in the medical field have not been fully studied.

In response, we present a novel dataset called CMExam to overcome these challenges and benchmark LLM performance. CMExam is sourced from authentic medical licensing exams. It contains more than 60K questions and utilizes the multiple-choice question format to allow standardized and objective evaluations. Questions in CMExam have corresponding solution explanations that can be used to test LLM’s reasoning ability in an open-ended manner. To offer diverse perspectives for measuring LLM performance in the medical field, we created five additional question-wise annotation dimensions based on authenticated resources and objective metrics. To reduce the substantial time and labor costs associated with annotating large-scale datasets, we propose an innovative strategy called GPT-Assisted Annotation. This approach harnessed the power of GPT-4 to automate the initial annotation process. Subsequently, the annotated data underwent a meticulous review and manual verification conducted by two medical professionals. Figure 1 shows an example question from CMExam and the annotation process.

Furthermore, we benchmark the performance of general domain LLMs and medical domain LLMs on answer prediction (multiple-choice) and answer reasoning (open-ended) tasks of CMExam. This comprehensive assessment aims to highlight the strengths and weaknesses of various approaches in Chinese medical QA, with a focus on LLMs. The main findings of this benchmark are as follows:

GPT-4 (OpenAI, 2023) demonstrates impressive zero-shot performance on the answer prediction task compared to other models, though still significantly lagging behind human performance.

GPT-3.5 (Brown et al., 2020) and GPT-4 generated reasonable answers on the answer reasoning task despite low BLEU and ROUGE scores. This is because they tended to generate short answers with reasonable quality.

Existing medical domain LLMs, such as Huatuo (Li et al., 2023) and DoctorGLM (Xiong et al., 2023), exhibit poor zero-shot performance on both tasks, indicating their limited coverage of medical knowledge and substantial room for improvement.

Lightweight LLMs (e.g., ChatGLM (Du et al., 2022)) fine-tuned on CMExam with supervision chieve performance close to GPT-3.5 on the answer prediction task. They also significantly outperform GPT-3.5 and GPT-4 on the reasoning task while having only 3% of the parameters of GPT-3.5.

In summary, this study provides valuable insights into the performance of LLMs in medical contexts from multiple perspectives, benefiting both the artificial intelligence research community and the medical research community. Our findings contribute to a deeper understanding of the capabilities and limitations of LLMs in the medical domain. Additionally, the CMExam dataset and benchmark introduced in this study serve as valuable resources to inspire researchers to explore more effective ways of integrating medical knowledge into LLMs, ultimately enhancing their performance in medical applications.

Related Work

Medical Question-Answering Datasets Table 1 presents a summary of medical QA datasets published after 2017. In particular, we focus on categorizing the data source and question types of the different datasets. Most existing medical QA datasets adopt an open-ended format, primarily because they were constructed directly from consumer questions and answers from doctors. However, multiple-choice and fill-in-the-blank questions provide a more standardized and objective evaluation, and only a small portion of medical QA datasets have adopted these formats. Notable examples include CliCR (Šuster and Daelemans, 2018), MEDQA (Jin et al., 2021), MMLU (Hendrycks et al., 2020), MLEC-QA (Zeng et al., 2023a), and MedMCQA (Pal et al., 2022). Note that the multiple-choice questions in MultiMedQA (Singhal et al., 2022) come from MEDQA, MedMCQA, and MMLU.

Data source types generally determine the reliability of a dataset. Consumer questions collected from web sources require human review to ensure the correctness of the answers. As datasets grow in size, quality control becomes increasingly challenging (Li et al., 2023). In contrast, datasets built from case reports (e.g., CliCR), research literature (e.g., BioAsq (Krithara et al., 2023)), medical books, exams, and related practices (e.g., MMLU and MedMCQA) are often more reliable.

From Table 1, we observe that there are few datasets based on multiple-choice questions from authoritative sources. This characteristic distinguishes CMExam from the MLEC-QA dataset, which is also derived from the Chinese National Medical Licensing Examination. In essence, CMExam has been meticulously crafted as a foundational benchmark dataset. It introduces question explanations for reasoning ability inspection, incorporates expansive annotation facets with authoritative references, and includes question-wise medical competencies and difficulty ratings calculated from human performance. These features make CMExam an indispensable resource for authoritative LLM performance assessment and meaningful human-machine comparisons. Table 2 presents a list of innovations and characteristics of CMExam, which are discussed in detail in the following sections.

Other Benchmark Datasets of Large Language Models The assessment of LLMs has witnessed significant progress, with the introduction of diverse benchmarks that evaluate different dimensions across multiple languages and models. Many datasets focus on assessing natural language understanding and reasoning capabilities of LLMs. RACE (Lai et al., 2017) includes English exams for Chinese middle and high school students. TriviaQA (Joshi et al., 2017) consists of question-answer pairs authored by trivia enthusiasts. DROP (Dua et al., 2019) evaluates reading comprehension with discrete reasoning and arithmetic components. GLUE (Wang et al., 2018) encompasses four existing NLU tasks, while SuperGLUE (Wang et al., 2019) extends it with a more challenging benchmark of eight language understanding tasks. Other datasets, such as HellaSwag (Zellers et al., 2019) and WinoGrande (Sakaguchi et al., 2021), focus on commonsense reasoning. TruthfulQA (Lin et al., 2021) includes health, law, finance, and politics, to assess LLMs’ ability to mimic human falsehoods, while MMCU (Zeng, 2023) covers medical, legal, psychology, and education to evaluate multitask Chinese understanding. In addition to language understanding and reasoning, several datasets focus on specific subjects and topics, such as Python coding tasks (Chen et al., 2021) and middle school mathematics questions (Cobbe et al., 2021). Notably, both C-Eval (Huang et al., 2023) and M3KE (Liu et al., 2023) serve as multi-level multi-subject evaluation benchmarks, making them particularly suitable for assessing the capabilities of LLMs across multiple domains.

The CMExam Dataset

Data Collection and Pre-processing CMExam comprises authentic past licensed physician exams in the Chinese National Medical Licensing Examination (CNMLE) collected from the Internet. The CNMLE, also known as the Physician Qualification Examination, is a standardized exam that assesses applicants’ medical knowledge and skills in China. It includes a written test with multiple-choice questions covering various medical subjects and a clinical skills assessment simulating patient diagnosis and treatment. We excluded questions that rely on non-textual information, including questions with external information such as images and tables, and questions with keywords ”graph” and ”table”. Duplicate questions were removed from the dataset. In total, 96,161 questions, 68,119 of which were retained after pre-processing. The dataset was then randomly split into training/development/test sets with a ratio of 8:1:1. Each question in the dataset is associated with an ID, five candidate answers, and a correct answer. 85.24% of questions have brief solution explanations and questions in the test set contain additional annotations.

Data Annotation CMExam provides a comprehensive analysis of LLM performance through five additional annotation dimensions. The first dimension involves disease groups based on the 11th revision of the International Classification of Diseases (ICD-11) (World Health Organization (WHO), 2021). ICD-11 is a globally recognized standard classification system for documenting and categorizing health conditions, consisting of 27 major disease groups. The second dimension comprises 36 clinical departments derived from the Directory of Medical Institution Diagnostic and Therapeutic Categories (DMIDTC) http://www.nhc.gov.cn/fzs/s3576/201808/345269bd570b47e7aef9a60f5d17db97.shtml, published by the National Health Commission of China. DMIDTC is an authoritative guide used for categorizing and naming diagnostic and therapeutic subjects within healthcare institutes. In cases where the question cannot be successfully classified by ICD-11 or DMIDTC, the annotation is marked as ”N/A”. The third dimension refers to medical disciplines, which are categorized based on the List of Graduate Education Disciplinary Majors (2022) published by the Ministry of Education of the People’s Republic of China http://www.moe.gov.cn/srcsite/A22/moe_833/202209/t20220914_660828.html. This dimension encompasses seven categories representing study majors used in universities. The fourth dimension was created by two medical professionals within the team to assess the primary medical competency tested by each associated question. It consists of four categories. The fifth dimension represents five potential difficulty levels of each question, determined by analyzing the correctness rate observed in human performance data collected alongside the questions. For detailed information on these additional annotations including their potential values, please refer to Table 9, 12, 10, 11. And our proposed GPT-Assisted Annotation strategy is shown in supplementary materials.

Dataset Characteristics The CMExam dataset has several advantages over previous medical QA datasets regarding: 1)Reliability and Authenticity: CMExam is sourced exclusively from the CNMLE that undergoes rigorous review and validation processes, ensuring its accuracy and adherence to established medical standards. 2) Standardization and Comprehensiveness: CMExam includes both multiple-choice questions that ensure fair and objective evaluations of models’ performance and question-wise open-ended reasoning that allows in-depth analysis and assessment of model reasoning abilities and comprehension. Despite the inherent absence of explanations within the CNMLE, we cross-referenced exam questions with solutions offered by diverse online medical examination preparation platforms, effectively enhancing the dataset’s informational depth. CMExam reflects the comprehensive coverage of medical knowledge and reasoning required in clinical practice, as it is sourced from carefully designed national medical exams. The inclusion of five additional annotation dimensions enhances the dataset’s rigor and offers valuable insights for in-depth evaluation and analysis. 3) Scale: CMExam consists of over 60K high-quality questions, providing a large and reliable dataset.

Data Statistics The dataset has a total of 68,119 questions, with 65,950 answers being single-choice and 2,169 being multiple-choice, with a maximum of five answer choices. Among all questions, 85.24% have associated solution explanations https://www.yikaobang.com.cn/, http://www.jinyingjie.com/, https://www.lanjiyin.com.cn/. Figure 2 shows additional statistics visualization and more basic statistics of CMExam can be seen in supplementary materials. Within the test set, 4,493 questions (65.97%) have corresponding disease group annotations. The most prevalent disease group is Traditional Medicine Disease Patterns (TMDP), followed by Digestive System Diseases, Certain Infectious (Digest) and Parasitic Diseases (InfDis), Endocrine, Nutritional, or Metabolic Diseases (Endo), and Circulatory System Diseases (Circ). For the associated clinical department annotations, 4,965 questions (72.90%) have been assigned values. The two most frequently represented clinical departments are Internal Medicine (IM) and Traditional Chinese Medicine (TCM), with Dentistry (Dent) and Surgery (Surg) following closely. Every question in the test set has been labeled with a discipline, where Clinical Medicine (ClinMed) comprises the largest proportion. Additionally, each question has been categorized into a competency area, with Medical Fundamentals (MedFund) being the predominant category. The difficulty levels of the questions align with common exam patterns, with a greater number of easy questions and a smaller number of hard questions.

Benchmarks

Model Selection The LLMs we benchmarked on the CMExam can be divided into two groups based on domains: 1) General Domain LLMs: This group comprises GPT3.5/4 (Brown et al., 2020; OpenAI, 2023), ChatGLM (Du et al., 2022; Zeng et al., 2023b), LLaMA (Touvron et al., 2023), Alpaca (Taori et al., 2023), and Vicuna (Chiang et al., 2023). These models are general-purpose language models trained on a massive amount of general-purpose corpora; 2) Medical Domain LLMs: This group can be further divided into two subgroups. The first subgroup consists of representative LLMs specifically designed for the medical domain, including DoctorGLM (Xiong et al., 2023) and Huatuo (Wang et al., 2023). DoctorGLM is a healthcare-specific language model initialized with ChatGLM-6B parameters and further fine-tuned on Chinese medical dialogues extracted from ChatGPT. Huatuo, on the other hand, is a knowledge-enhanced model, which builds upon the LLaMA architecture and is additionally supervised-fine-tuned with knowledge-based instruction data harvested from the Chinese medical knowledge graph (CMeKG). The second subgroup comprises medical LLMs that were constructed through supervised fine-tuning of LLMs using the CMExam training set. This subgroup includes models fine-tuned on BERT (Devlin et al., 2019), RoBERTa (Liu et al., 2019), PromptCLUE (Zhang and Xu, 2022) (T5-based), BART (Shao et al., 2021), Huatuo, ChatGLM, LLaMA, Alpaca, and Vicuna.

Human Performance To effectively gauge the medical proficiency of LLMs, incorporating a measure of human performance into the benchmarking process is of paramount importance. Therefore, during data collection, we preserved the accuracy of human responses for each question. Human performance is estimated by computing a weighted average of response accuracy within each dimension, with weights determined by the number of respondents. This design ensures a robust comparison of LLMs’ performance relative to human capabilities, particularly when larger respondent samples contribute to a question’s accuracy.

Experimental Setting For GPT models, we leveraged OPENAI’s API to access the GPT-3.5-turbo and GPT-4-0314 models, given that their open-source variants are currently unavailable. The LLaMA, Alpaca, and Vicuna models were used in their respective 7B versions, while ChatGLM was evaluated using its publicly accessible 6B version. Additionally, we performed fine-tuning on open-source models using the CMExam dataset. We used P-tuning V2 (Liu et al., 2021) for ChatGLM-6B, with the length of prefix tokens set to 128, and the learning rate set to 2e-2, LoRA (Hu et al., 2021) for LLaMA, Alpaca, Vicuna, and Huatuo models, with the rank set to 8, alpha set to 16, and dropout at 0.05. For BERT models, we followed the fine-tuning methods outlined in (Devlin et al., 2019), with batch size set to 16, learning rate set to 2e-4, hidden dropout probability set to 0.4, and maximum input length set to 192. The fine-tuning processes for all models except BERT involved a batch size of 64, a maximum input length, and a target length of 256. All fine-tuning was performed using NVIDIA V100 GPUs for 10 epochs.

Metrics We assess model performance on multiple choice questions using accuracy and weighted F1 score. These metrics are commonly employed in information retrieval and question-answering tasks to evaluate model performance. For the open-ended solution explanations of CMExam, BLEU (Papineni et al., 2002) and ROUGE (Lin and Hovy, 2003) were used to evaluate the discrepancy between model-generated explanations and ground truth.

2 Results and Analysis

Overall Comparison We first assessed the performance of general domain LLMs and medical domain LLMs for answer prediction and reasoning tasks. The results are displayed in Table 3. For the answer prediction task, GPT-4 significantly outperforms other methods, demonstrating a zero-shot performance with an accuracy of 61.6% and an F1 score of 0.617. While a performance gap still exists when compared to human performance (which stands at 71.6% accuracy), it’s noteworthy that this gap has been greatly reduced from what was observed with GPT-3.5. Among lightweight, general domain LLMs, ChatGLM outperforms LLaMA, Alpaca, and Vicuna, likely attributable to their limited coverage of the Chinese corpus. This restriction seemingly hampers their ability to provide accurate responses to CMExam queries. Furthermore, a noticeable deficiency in zero-shot performance is evident in lightweight medical domain LLMs such as Huatuo, owing to their restricted medical corpus diversity, which hampers the acquisition of broad medical knowledge and accurate interpretation of CMExam questions. Our findings suggest that finetuning models with CMExam enhance their performance. For instance, with an accuracy of 45.3%, ChatGLM-CMExam is comparable to GPT-3.5’s performance, despite utilizing only about 3% of the parameters employed by GPT-3.5. It is noteworthy that encoder-only LLMs, such as BERT and RoBERTa, remain a robust baseline for answer prediction tasks. Their performance can par with, or even exceed, that of certain decoder-only LLMs, such as LLaMA-CMExam and Alpaca-CMExam, despite having fewer parameters.

For the solution explanation task, we observe that GPT models performed poorly on the BLEU metric, likely due to their tendency to generate short explanations. However, they exhibited an advantage on the ROUGE metric. As DoctorGLM is unable to return answer options according to the prompt, we only report its performance in the solution explanation task. Through finetuning, LLM was able to generate more reasonable explanations. For instance, ChatGLM-CMExam achieved scores of 31.10 and 18.94 on BLEU-1 and BLEU-4, respectively, and scores of 43.94, 31.48, and 29.39 on the ROUGE metrics.

Results by Disease Groups Drawing upon ICD-11 annotations (26 categories), we conducted an analysis of the performance of several LLMs across various categories. To mitigate the potential impact of random variability resulting from the number of questions, we limited our analysis to categories containing more than 100 questions. According to Table 5, LLMs have uneven performance and significant gaps in knowledge. GPT-4’s accuracy ranges from 74.4% for Neo to 44.3% for TCMDP, GPT-3.5’s accuracy ranges from 63.9% for Neo to 31.0% for TCMDP and ChatGLM-CMExam’s accuracy ranges from 54.7% for Psy to 42.9% for RESP.

Results by Clinical Departments To compare model performance regarding the clinical department dimension (36 categories), we only analyzed categories with more than 50 questions to ensure result representativeness. Results presented in Table 5 highlight that the models show relatively high accuracy on questions associated with commonly encountered departments, such as Emergency Medicine (EM), Internal Medicine (IM) and Surgery (Surg). Their accuracy on questions associated with rarer departments, such as Traditional Chinese Medicine (TCM). There is a marked discrepancy in the average accuracy among different departments, with the highest being 50.9% and the lowest being only 13.9%. This observation suggests there are notable variations in medical knowledge and reasoning approaches among different departments. Consequently, it may be necessary to examine specific optimization strategies for different departments.

Results by Medical Disciplines Then, we evaluated LLM performance across seven medical disciplines. As depicted in Table 7, the performance of LLMs across disciplines such as Traditional Chinese Medicine (TCM), Traditional Chinese Pharmacy (TCPharm), and Pharmacy (Pharm) was notably subpar, with all accuracy rates falling below 42%. This pattern suggests a potential deficiency in the exposure of these models to data within these categories. Conversely, disciplines such as ClinMed and Ph&PM demonstrated higher accuracy rates, likely due to the abundance of relevant data. The observed variability in performance across different disciplines underscores the distinctiveness of data characteristics and complexities inherent to each field, thereby advocating for discipline-specific model optimizations and enhancements.

Results by Competencies Evaluations based on medical competency areas aimed at a higher-level understanding of model capability in solving medical problems. As indicated in Table 7, the lowest average accuracy across LLMs was observed within the domain of mastering Medical Fundamentals (MedFund), with a meager average score of 42.1%. This result demonstrates that these models, predominantly trained on general textual data, have inadequate exposure to medical-specific data. While fine-tuning did provide some improvement, these models could benefit from additional medical scenario data to further augment their performance. It is worth highlighting that the average accuracy in the domain of Public Health Laws and Ethics (PHL) was reasonably high, notably achieving an average of 47.6%. In addition, the LLMs showcased their proficiency in accurate disease diagnosis.

Results by Question Difficulty To evaluate model performance in tackling questions of varying levels of difficulty, we conducted experiments regarding the question difficulty dimension, which was calculated based on human exam-taker performance. As shown in Table 8, there’s an evident trend where model accuracies decrease as question complexity rises. This pattern suggests that more sophisticated questions demand an extensive knowledge base and complex reasoning, which are challenging for the LLMs, thus reflecting patterns observed in human performance.

Results by Question Length Finally, to investigate if model performance is associated with input lengths, we compared their performance regarding question lengths. Figure 3 illustrates that Large Language Models (LLMs) generally show higher accuracy with problem lengths between 60 and 90. However, their performance seems to falter with problems that are either too short or overly long. Additionally, we noticed that the effect of question length varies across different LLMs. For instance, GPT models tend to incrementally improve as the problem length expands, performing optimally within the 50 to 90 range. Conversely, ChatGLM-CMExam’s performance fluctuates noticeably with varying lengths, and it tends to fall short compared to GPT models when addressing longer problems.

Conclusion and Discussions

In this work, we developed CMExam, a dataset sourced from the stringent Chinese National Medical Licensing Examination, featuring 60,000+ multiple-choice questions, with detailed explanations. CMExam ensures reliability, validity, and adherence to medical standards. It also demonstrates the practicality of employing GPT-4 to automate the annotation process, which strikes a harmonious balance between efficiency and cost-effectiveness while maintaining the desired level of accuracy and reliability of the annotation. Utilizing this large and reliable corpus, we tested several LLMs for answer selection and reasoning tasks. A performance gap was observed between LLMs and human experts, signaling the need for additional LLM research. CMExam’s standardization and comprehensiveness also ensure objective evaluations of models while enabling in-depth analysis of their reasoning capabilities. The questions cover a wide spectrum of medical knowledge, augmented with five additional annotation dimensions for rigorous evaluation. This study aims to spur further exploration of LLMs in medicine by providing a comprehensive benchmark for their evaluation. We anticipate CMExam to contribute significantly to future advancements of LLMs, particularly in handling medical question-answering tasks.

Limitations Firstly, while CMExam is derived from meticulously designed medical examinations, our process of excluding questions requiring non-textual information may inadvertently affect the balance of the remaining questions, potentially introducing unexpected biases. It is critical to acknowledge this aspect while interpreting any findings or analyses conducted using this dataset. Furthermore, the current BLEU and ROUGE metrics primarily evaluate the explanation task, but these measures are insufficient for assessing the reasonableness of the answer. In future work, we will incorporate human evaluation to provide a more comprehensive assessment of the models.

Ethics CMExam is a dataset derived from the Chinese National Medical Licensing Examination, which aligns with numerous datasets containing similar National Medical Licensing Examinations (Zeng et al., 2023a; Hendrycks et al., 2020; Jin et al., 2021; Pal et al., 2022; Singhal et al., 2022). We have ensured adherence to applicable legal and ethical guidelines during data collection and use. The authenticity and accuracy of the exam questions have been thoroughly verified, providing a reliable basis for evaluating LLMs. Please note that the CMExam dataset is intended for academic and research purposes only. Any commercial use or other misuse that deviates from this purpose is expressly prohibited. We urge all users to respect this stipulation in the interest of maintaining the integrity and ethical use of this valuable resource.

Societal Impacts While CMExam aims to enhance LLM evaluations in the medical field, it should not be misused for assessing individual medical competence or for patient diagnosis. Conclusions drawn from models trained on this dataset should acknowledge its limitations, especially given its single source and the specific context of the CNMLE. The use of this dataset should strictly be limited to research purposes to avoid potential misuse.

References

Appendix A Appendix

This section presents four tables of additional annotations that contain translation. It showcases abbreviations, full English names, and Chinese names for each group in each annotation dimension. Table 9 showcases all disease groups included in the 11th revision of the International Classification of Diseases (ICD-11). We present the disease group in the same order found on the official website. Table 12 offers a classification of 36 clinical departments derived from the Directory of Medical Institution Diagnostic and Therapeutic Categories. Table 10 presents a breakdown of medical disciplines based on the List of Graduate Education Disciplinary Majors published by the Ministry of Education of the People’s Republic of China. This categorization comprises seven study majors used in universities. Table 11 provides all groups of areas of medical competency assessed in Chinese medical licensing exams.

A.2 Instructions for Pre-annotation

In this section, we present instructions used to pre-annotate CMExam test set data using GPT4. As shown in Figure 4,5,6,7, we first constrained the output from GPT4 to return only specific categories. We then annotated each of the five additional annotation dimensions relevant to this study with all the category information for each dimension. Next, we provided specific prompt information and finally, we performed filtering on the GPT4 output to improve the effectiveness of pre-annotation. During the actual annotation process, specific categories and prompt information should be filled in the grey background areas.

A.3 Analysis of Model Generation Ability

In Figure 8, we present partial explanations generated by various models for a medical question from the CMExam dataset. Notably, GPT-4 and GPT-3.5 produce concise and sensible explanations, which may account for the lower BLUE scores. Conversely, models like Vicuna, LLaMA, and Huotuo demonstrate a more prominent repetition phenomenon, while Alpaca simply duplicates the provided options without providing an explanation.

Fine-tuning models on the CMExam dataset significantly reduces the repetition phenomenon and improves the overall reasonableness of the explanations. For instance, the ChatGLM-CMExam model analyzes each option in a similar manner to the solution explanation. However, some models still generate unreasonable explanations, as observed in LLaMA-CMExam, Alpaca-CMExam, and Vicuna-CMExam. This could be attributed to their training on generic data and lack of specific knowledge in the medical domain. This underscores the significance of training large language models with a focus on the medical domain.

A.4 Analysis of Model Generation Correctness

To assess the accuracy of model-generated explanations, we conducted a study using a randomly selected sample of 50 cases in which the Language Models (LLMs) correctly predicted the answers. Medical experts were then invited to manually verify the correctness of the explanations, focusing not only on the accuracy of the answer predictions but also on the quality of the accompanying explanations.

Our investigation revealed that despite the correct answer predictions by the models, certain samples exhibited errors in their corresponding explanations. These errors were categorized by the experts into three groups: explanations that were irrelevant, repeated, or inaccurate. The statistics presented in Figure 9 demonstrate that the number of samples with accurate explanations generated by the GPT models exceeded 45, accounting for over 90% of the total. However, it is important to note that both the ChatGLM and ChatGLM-CMExam models may produce some erroneous explanations, primarily consisting of inaccuracies and irrelevance. We have included examples of these incorrect explanations in Figure 10.

A.5 Analysis of Few-Shot and Chain-of-Thought Prompts

In our research, we designed few-shot and chain-of-thought prompts for the answer prediction and reasoning tasks and conducted experiments on the GPT models. As shown in Table 13, our results demonstrate that while the use of few-shot or chain-of-thought prompts did not yield significant improvements in the prediction task, but there was a notable enhancement in the reasoning task.

Specifically, for the GPT-4 model, the utilization of few-shot prompts increased the BLUE-1 from 0.17 to 5.95, and the BLUE-4 from 0.06 to 2.25. Furthermore, incorporating chain-of-thought prompts further increased the BLUE-1 to 7.29. Similarly, positive effects were observed on the GPT-3.5 model, where few-shot prompts improved the BLUE-1 and BLUE-4 to 14.62 and 4.80, respectively. Additionally, the ROUGE-1, ROUGE-2, and ROUGE-L increased to 38.08, 18.35, and 18.37.

These improvements can be attributed to the fact that few-shot prompts provide examples that GPT models can reference when generating detailed explanations for each option during the reasoning process. Similarly, chain-of-thought prompts can achieve similar effects, aiding in the enhancement of model performance.

A.6 Data statistics

Questions in CMExam have a median length of 17 (Q1: 12, Q3: 32). Regarding solution explanations, the median length is 146 tokens (Q1: 69, Q3: 247). Table 14 shows more basic statistics of CMExam,

A.7 Guidelines for Expert-Annotation

During the annotation phase, we invited one expert physician from the Second Affiliated Hospital of Zhejiang University and one senior doctoral student from Zhejiang University School of Medicine to carry out the annotations. The expert physician has over two years of clinical experience. The annotation guidelines have the following sections:

Comprehensive Question Understanding: Prior to initiating the annotation process, meticulously comprehend the medical question, ensuring a holistic grasp of its context and significance.

Subject Categorization: Identify the precise subject or medical field that the question pertains to, such as cardiology, pediatrics, or pathology.

Principal Symptoms or Medical Conditions: Ascertain and pinpoint the primary symptoms or medical conditions expounded in the question.

Examination of Pertinent Factors: Scrutinize the question for any associated factors that might be present, including the severity of the ailment, its etiology, and patient history given in the question.

Examination of Pertinent Factors: Scrutinize the question for any associated factors that might be present, including the severity of the ailment, its etiology, and patient history given in the question.

Appropriate Classification System Usage: Use the accurate classification system for annotation in alignment with the determined subject and symptoms. Suitable systems could encompass the 11th revision of the International Classification of Diseases (ICD-11), the Directory of Medical Institution Diagnostic and Therapeutic Categories (DMIDTC), and others.

Addressing Multiple Annotations: In scenarios where the question encompasses multiple symptoms or medical conditions, opt for the most related classification for annotation.

Ensuring High-Quality Annotations: Adhere to the guidelines and definitions within the chosen classification system. This diligence helps avert subjectivity and ambiguity, fostering precision in the annotations.

Navigating Queries and Uncertainties: Should any doubts or uncertainties emerge during the annotation process, consult the official documents and glossaries of the chosen classification system. Engaging in discussions with professionals is also advised to achieve clarity.

Resolving Discrepancies: When disagreements emerge between annotators, a collaborative discussion shall be initiated. The objective is to reach a consensus and unify the annotation decision.

A.8 Prompt strategies for Pre-Annotation

During the experimental process, we indeed tried different prompts to enable GPT to better understand and complete the annotation task. The specific strategies were as follows:

Without ICD-11 Category Instructions: We did not provide detailed ICD-11 category information as instruction but directly supplied the question information and asked GPT to respond. Under this setup, a significant portion of the categories returned by GPT did not strictly belong to ICD-11 classifications, yielding unsatisfactory results.

Batch Processing for Cost Efficiency: Initially, we concatenated multiple questions and, through a single dialogue, had GPT return annotations for multiple questions. Under this setup, expert validation showed that the accuracy of GPT’s annotations was relatively low.

Consistency in Formatting: When no format guidance was given, GPT’s return format was inconsistent, resulting in a higher parsing cost. Hence, after multiple trials, we eventually opted for more rigorous format guidance.

Our annotation process was carried out in two stages: First, GPT conducted an initial round of pre-annotation. Subsequently, we invited an expert physician from the Second Affiliated Hospital of Zhejiang University and a doctoral student from Zhejiang University School of Medicine to annotate. The expert physician had over two years of clinical experience. In instances where there were disagreements in annotations, both annotators would discuss and eventually arrive at a consensus.