Can large language models reason about medical questions?

Valentin Liévin, Christoffer Egeberg Hother, Andreas Geert Motzfeldt, Ole Winther

Introduction

Self-supervised pre-training promises to turn vast quantity of raw data (e.g., text, images, audio) into general-purpose models. Language representations have transformed the field of natural language processing, from simple word vectors (Mikolov et al., 2013; Pennington et al., 2014) to deep contextualized representations (Peters et al., 2018; Vaswani et al., 2017; Devlin et al., 2018; Radford et al., 2018), language models are now ubiquitous in natural language processing, notably, thanks to the Transformer architecture (Vaswani et al., 2017) and its compatibility with massively parallel computation hardware.

In recent years, tremendous resources have been allocated to scale Transformer-based language models (Brown et al., 2020; Rae et al., 2021; Chowdhery et al., 2022; Thoppilan et al., 2022; Hoffmann et al., 2022; Smith et al., 2022; Zhang et al., 2022; Lieber et al., 2021; Fedus et al., 2021; Laurençon et al., 2023) to using hundreds of billions of parameters and to training on gigabytes of text. This so far translated in sustained gains (Kaplan et al., 2020) and enabled new ways to interact with language models. This progress made many of the past benchmarks obsolete and sparked a general interest for designing difficult enough benchmarks (e.g., BIG-bench; Srivastava et al. (2022)). Pre-train, prompt and predict (Liu et al., 2021) is an emerging paradigm for applying LLMs to new problems, without fine-tuning the weights on the task. Prompt-based learning consists in augmenting the problem with instructions such that the model’s completion of the prompt will correspond to a solution. This allows for LLMs to learn from a few examples (coined shots) which are simply incorporated into the prompts (Brown et al., 2020).

Initially, scaling language models up appeared to benefit more knowledge-intensive tasks than the reasoning-heavy ones (Rae et al., 2021). Nevertheless, Wei et al. (2022) demonstrated that LLMs could be applied to System 2 problems by prompting the model to generate step-by-step solutions, coined “Chain-of-Thought” (CoT). CoT prompting led to substantial improvements on many reasoning-intensive tasks (Wei et al., 2022; Zhou et al., 2022; Drozdov et al., 2022; Nye et al., 2021), allowing to bridge the gap with human-level performances for most of the hard BIG-bench tasks (Suzgun et al., 2022). As an alternative to writing reference step-by-step solutions, zero-shot CoT (Kojima et al., 2022) allows generating CoTs using single and domain-agnostic cue: “Let’s think step by step” (see example in Figure 1). The CoTs that result from that prompt not only appear to expose valid reasoning but also translate into superior zero-shot performances (see example in Figure 1).

Applying LLMs to real-life scenarios will require implementing additional safeguards. Language models may amplify the social biases present in the training data, may hallucinate incorrect facts and may lack or robustness (Bender et al., 2021), for instance to adversarial attacks (Wang et al., 2021). Therefore, deploying LLMs into sensitive areas such as healthcare must be operated with great care (Korngiebel & Mooney, 2021; Sezgin et al., 2022). Nonetheless, large language models are powerful tools and therefore have the potential to transform the field of machine intelligence. At the dawn of this research work, although LLMs had been tested on large benchmarks (MMLU Hendrycks et al. (2020), BIG-bench Srivastava et al. (2022)), studies applied to the medical domain were still needed. Specialized datasets such as the MedQA-USMLE (Jin et al., 2020) enable assessing the capabilities of LLMs in realistic clinical scenarios requiring specialized medical knowledge, advanced reasoning capabilities and human-level reading comprehension skills.

This article – written in three stages (v1: July 2022; v2: December 2022; v3: September 2023) – evolved along with the remaining of the field. December 2022 was a turning point in machine learning history; new records were achieved on medical benchmarks by the domain-specific Med-PaLM (Singhal et al., 2022; 2023b), ChatGPTChatGPT was released to the public on November 30, 2022 – chat.openai.com and GPT-4 (Nori et al., 2023). ChatGPT sparked the interest of the public and the research community, which hastened to benchmark it against USMLE questions Gilson et al. (2023); Kung et al. (2023), turning to self-curated data instead of the peer-reviewed MedQA benchmark.USMLE steps 1,2 and 3 were evaluated separately whereas the MedQA aggregates all steps. Similarly to our work, Singhal et al. (2022) and Kung et al. (2023) involved human experts to evaluate the generated explanations on USMLE questions. Concurrently, significant progress happened on the open-source world (Llama-2; Touvron et al. (2023)). Recently, Chen et al. (2023) investigated both generalist and finetuned open-source LLMs applied to medical benchmarks. CoT prompting and ensemble methods are now commonplace in the literature (Singhal et al., 2022; 2023b; Nori et al., 2023; Chen et al., 2023) whereas retrieval-augmentation (grounding) remains less common (Wang et al., 2023; Liévin et al., 2022).

This paper investigates the performances, interpretability and limitations of CoT prompting for medical question answering. We utilized the GPT-3.5 series (InstructGPT and Codex). This research was conducted in three rounds; first, using InstructGPT, we investigated variations of zero-shot CoT prompting for medical reasoning (domain-specific CoT cues, retrieval augmentation), looking both at the answering performances and the limitations based on an expert evaluation. In the second round, thanks to the Codex beta program, we investigated how scaling inference-time compute could be applied to challenge both the human baseline and to quantify uncertainty. Last, we benchmarked a range of open-source models. Our contributions are:

We assess how GPT-3.5 perform on multiple-choice medical board exam question datasets (MedQA-USMLE and MedMCQA) and a medical reading comprehension dataset (PubMedQA) using prompt engineering. We explore zero-/few-shot, direct/CoT, domain-specific CoT cues and retrieval augmentation.

We propose an evaluation protocol for evaluating generated CoTs (three main categories: reasoning, knowledge and reading comprehension). A medical expert annotated subset of CoTs generated by zero-shot InstructGPT and supports that InstructGPT, in many cases, can reason and exploit memorized expert knowledge.

We demonstrate that scaling inference-time compute enables Codex 5-shot CoT to be well-calibrated and to reach the passing score on the three medical datasets.

We benchmark open-source LLMs on the MedQA-USMLE and MedMCQA.

This article has evolved over three distinct versions, each exploring different facets of LLMs:

v1 - July 2022: Investigated InstructGPT (expert evaluation & benchmarking prompting strategies).

v2 - December 2022: Scaled experiments and passed the MedQA-USMLE using Codex.

v3 - September 2023: Evaluated open-source models Llama-2, Vicuna, Guanaco, Falcon, etc.

Method

This paper explores variations of prompt engineering for medical question answering. The prompt templates are summarized in Figure 2.

We studied two classes of prompts: the direct prompt and zero-shot CoT. The direct prompt triggers the model to generate the answer using a single completion step (i.e., “The answer is”) whereas, when applying the zero-shot CoT framework, we use a two-steps prompting scheme: first an initial reasoning prompt with a CoT cue (e.g., “Let’s think step by step”) which completion is the CoT, second an extractive prompt which completion is the answer (e.g., “Therefore the answer is”). In the zero-shot CoT setting, this corresponds to the setup described in Kojima et al. (2022), the direct setting corresponds to Brown et al. (2020).

We experimented with inserting examplars (or shots) of question-answer pairs and question-explanation-answers triplets in the prompts. We built each shot using the zero-shot template, replacing the output with the reference explanations and answers. In the few-shot CoT setting, our setup matches the one from Wei et al. (2022).

We denote x{\mathbf{x}} the answer string, y{\mathbf{y}} a prompt and z{\mathbf{z}} a completion generated from an LLM denoted pθp_{\theta}. In the zero-shot setting, sampling z^∼pθ(z∣y)\hat{{\mathbf{z}}}\sim p_{\theta}({\mathbf{z}}|{\mathbf{y}}) is a two-steps process (first generate the CoT, then extract the answer) pictured in Table LABEL:tab:prompt-design. Using a sampling temperature τ\tau, kk completions z^1,…,z^k\hat{{\mathbf{z}}}_{1},\ldots,\hat{{\mathbf{z}}}_{k} can be sampled from the generative LLMs. Following Wang et al. (2022), we aggregate the completions and estimate the marginal answer likelihood as (Figure 3)

LLMs memorise part of the knowledge embedded into the training data, nonetheless, models might fail to re-use this knowledge effectively during prediction. Conditioning the predictions on a knowledge base is an alternative research direction for improving language models (Lewis et al., 2020; Borgeaud et al., 2021; Lazaridou et al., 2022).

We investigated whether grounding the model with additional context could improve the answering accuracy. We experimented with a simple BM25 retriever and used Wikipedia as a knowledge base. Read more details in Appendix G.

Experiments

This section is separated into three parts: (i) introducing the datasets and the GPT-3.5 models, (ii) investigating zero-shot medical reasoning with InstructGPT and (iii) scaling inference-time compute with Codex (using longer few-shot prompts and sampling many completions per question).

Our source code is available on Github.github.com/vlievin/medical-reasoning – DOI: 10.5281/zenodo.10301874 A collection of generated CoTs, reusable for downstream tasks, are accessible through ToughtSource (Ott et al., 2023).github.com/OpenBioLink/ThoughtSource All our benchmark results are summarized in Appendix A, Table S2.

This study is centered around three medical multiple-choice question answering datasets: USMLE (Jin et al., 2020) which includes difficult real-world medical questions targeting medical professionals, the MedMCQA (Pal et al., 2022) which gathers questions from medical school entrance exams and the PubMedQA (Jin et al., 2019) which includes reading comprehension questions about PubMed abstracts. The three datasets are summarized in Table 3. For each dataset, we gathered questions with explanations (long answer) which we used as reference CoTs in few-shot learning scenarios. We present the three datasets in further details in Appendix C. Furthermore, we compare the MedQA-USMLE with the MMLU-USMLE dataset (Hendrycks et al., 2020) in Appendix D, we found the MedQA questions to be more challenging than the MMLU ones.

We study a collection of closed- and open-source models. The 175B parameter GPT-3.5 series (Brown et al., 2020) the human-aligned GPT-3 (InstructGPT, text-davinci-002, Ouyang et al. (2022)), the code-finetuned GPT-3 (Codex, code-davinci-002, Chen et al. (2021)). A collection of open-source models ranging from 7B to 70B parameters: Llama-2 (Touvron et al., 2023), Vicuna (Zheng et al., 2023), Guanaco (Dettmers et al., 2023), Falcon (Almazrouei et al., 2023), MPT (Team, 2023) and GPT-NeoX (Black et al., 2022). We used greedy decoding (temperature τ=0\tau=0) with k=1k=1 sample unless specified (e.g., ensemble methods).

In Appendix E, we report the test USMLE accuracy for four GPT-3 versions: a small GPT-3, the largest GPT-3 trained without human-alignment, InstructGPT and Codex. The smaller model text-curie-002 delivered close to random performances, with a maximum accuracy of 27.9%. The non-aligned largest GPT-3 text-davinci-001 scored 40.2%, whereas the largest code pre-trained Codex scored 52.9% and the code pretrained and human-aligned InstructGPT scored 47.1%.

2 Investigating zero-shot reasoning with InstructGPT

In this section, we investigate whether the good generative capabilities of LLMs can be applied to answer medical questions in a zero-shot setting. We investigate variations of the zero-shot CoT framework: using domain-specific CoT cues and augmenting the prompt with Wikipedia passages.

In addition to the original zero-shot CoT cue “Let’s think step by step” we tested 29 other domain-specific variations such as “Let’s think step by step like a medical expert”. The study is available in Appendix B. We selected five CoT cues displayed in Table 3. In Appendix I, we display CoT samples for more exotic cues such as “Let’s follow a Bayesian step by step approach” and “Let’s work by elimination” and “Let’s reflect on each answer option”.

Zero-shot benchmark

In Table 4, we report the performances of InstructGPT for the direct prompt and the aggregated performances for the five domain-specific CoT cues (Table 3). We explored augmenting the prompts with retrieved Wikipedia passages (grounding) and report the performances of an ensemble model with majority voting, akin to Wang et al. (2022).

InstructGPT outperformed the domain-specific and finetuned BERT baselines on the three datasets. Without BM25 grounding, InstructGPT scored +1.4% on the USMLE questions, +1.0% on the MedMCQA exam questions and +1.1% on PubMedQA over the best BERT methods.

Without BM25 grounding, the direct prompt remained, on average, a better alternative to the CoT prompts. Performances were lower for each of the considered CoT cues, except in the case of the USMLE dataset, for which half of the CoT prompts resulted in small improvements over the direct prompt (+1.1% using the CoT prompt #1 vs. using the direct prompt). Nonetheless, the domain-specific CoT prompts #2–5 did not significantly outperform the original CoT prompt #1.

In an attempt to exploit the good reading comprehension skills of InstructGPT, we explored conditioning the completions on Wikipedia passages. When using the direct prompt, we recorded gains on the USMLE (+1.3%) and on the MedMCQA (+2.7%) datasets, suggesting that retrieval augmentation might be beneficial.

Combining the predictions of multiple prompts outperformed the single-prompt predictions, except in the case of the PubMedQA dataset, for which the direct prompt performed exceptionally well. The best performances on the USMLE and MedMCQA datasets were obtained by combining retrieval-augmented prompts, setting a maximum of 53.1% accuracy on the USMLE dataset and 48.8% valid. accuracy on the MedMCQA dataset.

Expert evaluation of the generated CoTs

InstructGPT delivered strong performances using zero-shot CoT prompting. In this section, we investigate whether the CoTs are sound and seek to understand better how the model fails and succeeds. We considered three general skills that we expect are required to be mastered to answer medical questions: (i) performing non-trivial reasoning steps, (ii) recalling knowledge that is not provided in the context and (iii) the ability to comprehend the question and the context. Based on the three skills, we defined three success patterns (A, B, C) and three failure patterns (D, E, F).

A subset of 50 CoTs generated based on USMLE questions was annotated by a medical expert (C.E.H.) using the six categories. For each category and each CoT, we reported a match if the pattern could be observed at least once. This means that a CoT can be labelled with both a correct and an incorrect pattern for the same skill. We showcase thirty annotated CoTs (three in Figure 9, 27 in Appendix I).

We report the frequencies of occurrence for the six patterns in Table 5. We found that most of the questions answered incorrectly triggered generating CoTs that contained reasoning errors (pattern D, 86%), and that exhibited a lack of knowledge (pattern E, 74%). Misunderstanding of the questions or the context was less frequently observed (Pattern F, 50%). We observed that CoTs leading to questions answered correctly could still show failure patterns but we also observed that the CoTs leading to incorrect answers were not entirely incorrect, as 59% contained at least one correct reasoning step, 65% showed proper recall of knowledge. Furthermore, inspecting the CoTs leading to incorrect answers more closely, we found that 47% of those were inconclusive:Labelling questions as inconclusive or not was also performed by C.E.H. the model could not narrow down the prediction to a single answer.

Answering bias

In Figure 4, we report the frequencies of the USMLE answers and the frequencies of predicted labels (zero-shot InstructGPT) for the direct and CoT prompts. Both prompting schemes led to biased predictive frequencies. Direct prompting led to over-estimating the labels C and D while under-estimate the label A. CoT prompting led to under-estimating B and C while over-estimating the label D. We repeat the experiment using randomly permuted labels and observed similar patterns, see Appendix F.

3 Scaling inference-time compute with Codex

In the second round of experiments, we investigated whether using more inference-time compute, thanks to the Codex beta program, could be utilized to obtain better performances and more interpretable outputs. Codex enables using longer prompts, we used five-shot prompts and experimented with sampling k=100k=100 completions with temperature τ=0.5\tau=0.5 for each question. We report question answering performances and results on uncertainty quantification.

In Figure 5, we report the performances of Codex 5-shot CoT given subsets of k′<kk^{\prime}<k CoTs. We report the best finetuned models and the human baseline. In line Wang et al. (2022), increasing the budget of samples yields better results. Using an ensemble of the kk samples, Codex 5-shot CoT reaches the passing score on the three tasks (see Table 1): the USMLE dataset (60.2% ≥\geq 60%), the MedMCQA dataset (62.7% ≥\geq 50%) and on the PubMedQA dataset (78.2% ≥\geq 78%). Additional results, including performances in zero-shot settings, are available in Table S2, Appendix A. Although Codex performed exceptionally well with 5 shots, Codex yield feeble performances with zero-shot CoT; inspecting the generated CoTs revealed lesser-quality samples (Appendix I).

Uncertainty quantification

We investigate the answering likelihood Equation 1 given by Codex 5-shot CoT with k=100k=100 samples. In Figure 6, we report the maximum probability assigned by the model for correctly vs. incorrectly answered questions along with the calibration plots for the three datasets. Codex 5-shot CoT appears to be overall calibrated, although the calibration is worse for the PubMedQA dataset.

4 Benchmarking Open-Source Models

In the rapidly evolving landscape of LLMs, a prevalent question is the performance gap between open-source and closed-source models. Our study focused on the capabilities of InstructGPT and Codex. Given a budget of 2.000 A100 hours, we benchmarked a range of open-source LLMs, with parameter sizes ranging from 7 to 70 billion, against the 175-billion-parameter Codex. In Figure 7, we report the predictive performances, calibration plot and bias for Llama-2, Vicuna 1.5 and Codex using up to k=100k=100 CoT samples. We provided additional results in Figure 8 in Appendix H (zero- and 5-shot, MedQA-USMLE and MedMCQA).

Discussion

Zero-shot InstructGPT and Codex outperformed finetuned BERT models on three challenging question-answering datasets (section 3.2 and Appendix A). In the case of the USMLE and the MedMCQA datasets, the retrieval-augmented BERT baselines were outperformed by several LLMs, regardless of augmenting the prompts with Wikipedia passages. This suggests that LLMs, without finetuning, can mobilize medical knowledge and problem-solving skills.

For both InstructGPT and Codex, single-sample CoT prompting was not found to be competitive with direct prompting (section 3.2 and Appendix A). Nevertheless, CoTs are human-readable and therefore interpretable. Our expert evaluation (section 3.2) revealed that CoTs are often sound: even InstructGPT still does mistakes, it was often able to reason, recall medical knowledge and comprehend the given problem. In section 3.2 and Appendix B, we explored domain-specific CoTs cues such as “ Let’s think step by step like a medical expert”. Although such prompts, taken separately, did not outperform the original zero-shot CoT prompt (see Table S2 in Appendix A), more specific prompts appeared to trigger alternative strategies such as working by elimination or manipulating equations (see Appendices B and I). Investigating whether a task-specific prompt could help solve specific tasks will be left for future research. A collection of generated CoT samples are presented in Appendix I, many more samples are available on our GitHub page.

The expert evaluation of the generated CoTs (section 3.2) and the good results obtained on the medical exam questions (see Table S2, Appendix A) suggest that GPT-3.5 memorizes domain knowledge. Nevertheless, despite the simplicity of the BM25 retriever and the small number of retrieved documents prepended in each prompt, grounding InstructGPT resulted in slight improvements (see Table 4). This suggests that InstructGPT is not omniscient and so (i) using stronger retrievers such as commercial search engines (Lazaridou et al., 2022) or dense retrievers (Karpukhin et al., 2020), (ii) using a more complete knowledge base Borgeaud et al. (2021), or (iii) leveraging inference-time compute by retrieving, re-ranking and processing more passages (Lazaridou et al., 2022) might improve performances. Nevertheless, how to best combine 5-shot CoT prompting with retrieval augmentation remains a promising research direction.

In section 3.2, we exposed the biases induced by the use of direct and CoT prompts. In the case of the direct prompt, the answer D was most often selected, which might be due to its proximity to the generated answer. In the case of the CoT prompts, the labels A and D were selected more often, which might be a result of often beginning CoTs with content related to option A. Based on an inspection of the CoTs, we speculate that GPT-3 defaults to this behaviour when it cannot answer but still attempts to complete the prompt with a default answer (D or A). Shuffling the answer options might be one way to overcome this limitation, however, other forms of biases might still be present.

CoTs can be combined and/or filtered using human or automated feedback (Wang et al., 2022; Cobbe et al., 2021). In section 3.3, we showed that sampling and combining up to k=100k=100 completions using Codex or Llama-2 with 5-shot CoT prompts was sufficient to reach both the MedMCQA and the challenging USMLE, although a large gap remains between our models and the human experts.

In section 3.3 and 3.4, we looked at the probability assigned to correct and incorrect predictions using the ensemble model from Equation 1. We found Codex and Llama-2 to be close to well-calibrated, corroborating the results of Kadavath et al. (2022) that “language models (mostly) know what they know”.

In Appendix E, we compared multiple GPT-3 models in the zero-shot setting. Best performances are obtained using Codex, outperforming the human-aligned InstructGPT, which is a finetuned version of Codex. Human alignment might impair performances; Codex (without alignment) was not as robust as InstructGPT (with alignment) in zero-shot CoT setting (see performances in Table S2 in Appendix A, see CoT samples in Appendix I). Nevertheless, 5-shot prompting allowed us to bypass the zero-shot limitations of Codex. We observed a similar pattern when comparing the versions of LLama-2 70b: the base version outperformed the chat version (Appendix H). Instruction-finetuned models might lose in-context learning abilities.

Open-source models, despite having fewer parameters, are approaching the performance of proprietary ones (Figure 7, 8). For instance, Llama-2 outperforms Codex with just half the parameters.

Instruction-finetuned LLMs like Guanaco (Dettmers et al., 2023) and Vicuna (Zheng et al., 2023) performed exceptionally well (Figure 8). Surprisingly, Vicuna 1.5 13B’s superior performance to both Llama-2 versions underscores the significance of high-quality datasets for instruction-based fine-tuning (Zhou et al., 2023).

Conclusion

We applied zero-shot, few-shot direct and CoT prompting to medical question answering with and without retrieval augmentation. Zero-shot InstructGPT significantly outperformed the finetuned BERT baselines. CoT prompting proved to be a powerful tool leading to better performances and more interpretable predictions. Our expert evaluation suggests that, LLMs can mostly comprehend complex medical questions, can often recall expert-domain knowledge and can often perform non-trivial reasoning steps.

Although InstructGPT and Codex still make mistakes, we found that scaling inference-time compute by sampling many CoTs per question could overcome part of these limitations. With 100 samples, Codex 5-shot CoT delivered unprecedented performances on the three datasets, bridging the gap with human-level performances and virtually passing the USMLE by 0.2% points. Our exploration into open-source LLMs indicated their competitive stance in medical benchmarks. Llama-2 outperformed Codex by 2 points on the USMLE in spite of a much smaller parameter footprint.

However, deploying LLMs in real-life clinical scenarios will require the development of more robust techniques. We exposed one form of bias (ordering of the answer options affects the predictions) but many more might affect predictions, including those hidden in the training data (e.g., gender, race, …). Nevertheless, a lack of knowledge might be more easily compensated, our experiment with BM25, albeit limited, suggests that augmenting the prompt with factual data improves performances.

Since the completion of version 2 of this work, both GPT-4 and MedPalm 2 have achieved performance on USMLE around 85% Nori et al. (2023); Singhal et al. (2023a). This is not unexpected given the evolution the LLM field has witnessed recently. Although benchmark contamination in training sets for both proprietary and open-source LLMs is a valid concern, these results indicate that both open and closed sourced LLMs hold great potential for assisting human decision-making in medicine and beyond.

We thank OpenAI for granting access to the Codex beta program. We acknowledge EuroHPC Joint Undertaking for awarding us access to MeluXina at LuxProvide, Luxembourg. V.L.’s work was funded in part by Google DeepMind through a PhD grant. OW’s work was funded in part by the Novo Nordisk Foundation through the Center for Basic Machine Learning Research in Life Science (NNF20OC0062606). V.L., A.G.M. and O.W. acknowledge support from the Pioneer Centre for AI, DNRF grant number P1.

Author Contributions

Conceptualization, V.L., C.E.H. and O.W.; Methodology, V.L. and O.W.; Software, V.L. and A.G.M.; Investigation, V.L. and A.G.M.; Writing – Original Draft, V.L.; Writing – Review & Editing, all authors.; Data Curation, C.E.H. ; Funding Acquisition, O.W.; Supervision, O.W. and V.L.

References

Appendix A Summary of the Results

In Table S2, we summarize the performances of InstructGPT and Codex on the MMLU-USMLE, MedQA-USMLE, MedMCQA and PubMedQA datasets in zero-shot, few-shot, with and without grounding. All our results on the validation set of the MedMCQA are estimated using 1k samples. The results of the MedMCQA test set require submitting an official submission. We used a sampling temperature of τ=0\tau=0 for all experiments except when drawing k>0k>0 samples and using majority voting (MV). For the majority voting model, we used k=100k=100 samples and τ=0.5\tau=0.5 for Codex, τ=0.9\tau=0.9 for Vicuna.

We composed an initial set of 30 zero-shot CoT prompt variations. In Table S1, we report the accuracy for each of the 30 prompts based on a subset of 100 USMLE validation questions. Given an estimated accuracy uncertainty of 5% (see the paragraph “uncertainty estimation” below), we concluded that the first half of the results are all reasonable candidates for the study.

For the remaining of this paper, we selected 5 prompts: the original “Let’s think step by step”, the medical variation “Let’s think step by step like a medical expert” and the top-three CoT cues reported in Table S1.

In Figure S1, we report the agreement rate for all 30 prompts on the 100 validation questions. Whereas most of the prompts followed a rather consistent pattern, with an agreement rate superior to 50%, a minority of the prompts seemed to agree less with the majority of the prompts, such as “Let’s reflect on each answer option step by step”, “Let’s follow a Bayesian step by step approach” or “Let’s work by elimination”. In Figure S4, we showcase four chain-of-thoughts selected to highlight the diversity of the completions and the ability of InstructGPT to adopt diverse problem-solving strategies. Yet, strategies are not always executed correctly: in Figure S4, example 2, GPT-3 ultimately finds the correct answer (Missense mutation) but identified the wrong diagnostic (the 6-year-old boy suffers from sickle cell disease).

Appendix C Datasets

Jin et al. (2020) gathers historical questions from the United States Medical Licensing Examination (USMLE), which targets trained medical professionals. The questions are notorious for being challenging as they often require strong problem-solving skills coupled with comprehensive medical knowledge. Each question features a description of a medical case and a question that emulates the real clinical setting. The more recent MMLU dataset (Hendrycks et al., 2020) has 31 validation and 272 test USMLE questions (around 105 words/question). In Appendix D, we benchmark both USMLE datasets and found the MedQA USMLE dataset to be more difficult. The MedQA-USMLE data does not come with explanations. Instead, we use the MMLU-USMLE CoTs from Chung et al. (2022) that are available from https://github.com/jasonwei20/flan-2.

Pal et al. (2022) is a large-scale multiple-choice question answering collected from Indian medical school entrance exams (AIIMS and NEET-PG). The MedMCQA covers a broad range of medical topics (dentistry, psychiatry, surgery, …) and requires being able to follow a variety of reasoning types (logic, factual, comparison, …). However, questions are often more knowledge-centred than the USMLE questions, which tend to focus more on problem-solving skills.

Jin et al. (2019) is a collection of expert-annotated yes/no/maybe research questions derived from PubMed abstracts. Whereas the questions from the USMLE and the MedMCQA datasets are self-contained and might be answered using general medical knowledge and methodology, each PubMedQA question is contextualized on a provided abstract. Therefore PubMedQA primarily focuses on evaluating reading comprehension skills.

Appendix D MedQA-USMLE versus MMLU-USMLE

In Table S3, we report the performances of the three medical question answering datasets as well as the professional medicine subset of the MMLU dataset (Hendrycks et al., 2020), which was also explored in recent related work (Chung et al., 2022).

Based on Codex performances, the MedQA-USMLE dataset appears to be more challenging than the MMLU-USMLE counterpart. Codex (Chen et al., 2021) in a 5-shot setting (Direct and CoT prompting, τ\tau=0), scores around 13.2% lower accuracy on the MedQA-USMLE (∼\sim56.4%) than on the MMLU-USMLE (∼\sim69.6%). Succeeding the USMLE requires a score of around 60%.

We report the test USMLE accuracy for multiple GPT version in Table S4 for the direct and CoT #1 prompts. Note that Codex (code-davinci-002) is a large GPT-3 model pre-trained on text and code; InstructGPT (text-davinci-002) is a version of Codex finetuned based human-feedback to “follow the user’s instructions helpfully and safely”.https://beta.openai.com/docs/model-index-for-researchers

The smallest model performed only slightly better than at random, with an accuracy of maximum 27.8% for the curie model, whereas the largest model non-aligned text model text-davinci-001 scored a maximum of 40.2% for all prompts. The best performances are obtained with the text and code pre-trained model code-davinici-002 (52.9%). Human-alignment appears to damage answering performances: text-davinci-002 scored a maximum of 47.1%. This suggests that advanced medical reasoning capabilities only emerge in the largest of the GPT-3 models, and that code pre-trained is highly effective. In this experiment, human-alignment led to a decrease of accuracy, although we found InstructGPT to overall produce more readable samples than Codex in zero-shot CoT settings (Appendix I).

In Table S5, we report the frequencies of answers and of predicted labels with and without label permutation. We report the frequencies for InstructGPT as well as Codex.

Querying InstructGPT using the CoT prompts resulted in a more faithful predictive distribution of the labels. Nonetheless, a bias towards the labels A and D and a tendency to avoid predicting labels B and C could still be observed. To confirm whether this bias originates from the data or the model, we permuted the labels and repeated the experiment for prompts number 0 and 1 and observed the same trend. Codex exhibits similar trends, although few-shot learning seems to yield more faithful predictive distributions.

In all cases, models tend to default to the label D. In Figure S7, we present two CoT leading to mispredicted label D. In both cases, GPT-3 fails to narrow down to one answer options and defaults to option D.

Figure S2 presents some of the results from Table S5 for the USMLE and extend it with the frequencies observed in the two other datasets for Codex 5-shot CoT (k=100k=100 samples). The bias appears less important for the MedMCQA and PubMedQA datasets than for the USMLE dataset.

Wikipedia articles were converted into overlapping passages of size 100 words and indexed along with their respective article titles. Given a question q{\mathbf{q}}, an answer choice a{\mathbf{a}}, and weights β1=1,β2=1,β3=0.5\beta_{1}=1,\beta_{2}=1,\beta_{3}=0.5. The weights were chosen based on a qualitative assessment of the retrieved passages on a few questions. we retrieved passages d{\mathbf{d}} based on a composite BM25 score defined as

We assessed the performances of open-source models (Vicuna, Guanaco, GPT-NeoX, MPT-instruct, Falcon and Llama-2) on the MedQA-USMLE and MedMCQA datasets in zero-shot and 5-shot settings, all using greedy decoding (τ=0\tau=0). We report results in Figure S3.

In Figure S4, we report four selected CoTs generated from the prompt variations studied in Appendix B

In Figure S5, we display CoTs generated by Codex. Codex appears to yield CoTs of lower quality than InstructGPT (frequent repetitions, less verbosity).

We provided nine more expert-labelled chain-of-thoughts in Figures S6, S7, S8, S9, S10, S11, S12, S13 and S14. Note that patterns reported in Table 5 cannot always be matched to text segments, as one highlighted text segment does not always correspond to a single category (reasoning and knowledge patterns are often entangled).