Conformal Prediction with Large Language Models for Multi-Choice Question Answering

Bhawesh Kumar, Charlie Lu, Gauri Gupta, Anil Palepu, David Bellamy, Ramesh Raskar, Andrew Beam

Introduction

Large language models (LLMs) have recently achieved impressive performance on a number of NLP tasks, such as machine translation, text summarization, and code generation. However, lingering concerns of trust and bias still limit their widespread application for critical decision-making domains such as healthcare.

One well-known issue with current LLMs is their tendency to “hallucinate” false information with seemingly high confidence. These hallucinations can occur when the model generates outputs not grounded in any factual basis or when the prompt is highly unusual or ambiguous. This behavior of LLMs may also result from how these models are trained — using statistical sampling for next-token prediction — which can progressively increase the likelihood of factual errors as the length of generated tokens increases (LeCun, 2023). Factually incorrect outputs may confuse and deceive users into drawing wrong conclusions, ultimately decreasing the overall system’s trustworthiness. Decisions based on unpredictable or biased model behavior could have significant negative and socially harmful consequences in high-stakes domains such as healthcare and law.

Therefore, we seek to explore principled uncertainty quantification (UQ) techniques for LLMs that can provide guaranteed error rates of model predictions. Ideally, these UQ techniques should be model agnostic and easy to implement without requiring model retraining due to the intensive computing costs and limited API access associated with many LLMs. To this end, we investigate conformal prediction, a distribution-free UQ framework, to provide LLMs for the task of multiple-choice question-answering (MCQA).

Based on our experiments, we find the uncertainty, as provided by conformal prediction, to be strongly correlated with accuracy, enabling applications such as filtering out low-quality predictions to prevent a degraded user experience. We also verify the importance of the exchangeability assumption in conformal prediction (see section 2) for guaranteeing a user-specified level of errors.

To summarize, our contributions are the following:

we adapt conformal prediction for MCQA tasks to provide distribution-free uncertainty quantification in LLMs,

show how the uncertainty provided by conformal prediction can be useful for downstream tasks such as selective classification,

and assess the performance of conformal prediction when the exchangeability assumption is violated for in-context learning in LLMs.

Conformal Prediction

Uncertainty quantification (UQ) techniques are critical to deploying machine learning in domains such as healthcare (Bhatt et al., 2021; Kompa et al., 2021b, a). Conformal prediction (Gammerman et al., 2013; Vovk et al., 2022) is a flexible and statistically robust approach to uncertainty quantification. Informally, the central intuition behind conformal prediction is to output a set of predictions containing the correct output with a user-specified probability.

By providing a more nuanced understanding of the model’s confidence and a statistically robust coverage guarantee, conformal prediction paves the way for improved and more reliable applications of machine learning models across various domains (Kumar et al., 2022).

Prediction sets. Formally, let C:X→2Y\mathcal{C}:\mathcal{X}\rightarrow 2^{\mathcal{Y}} be a set-valued function that generates a prediction sets over the powerset of YY given an input XX. This prediction set naturally encodes the model’s uncertainty about any particular input by the size of the prediction set.

Expressing uncertainty as the set size is an intuitive output that can be helpful in decision-making contexts (Babbar et al., 2022). For example, in medical diagnosis, the concept of prediction set is similar to a differential diagnosis, where only likely and plausible conditions are considered given the observed symptoms of a patient (Lu et al., 2022c). Indeed, conformal prediction has been utilized for uncertainty quantification in healthcare applications such as medical imaging analysis (Lu et al., 2022a, b; Lu & Kalpathy-Cramer, 2022).

Coverage guarantee. Conformal methods generate prediction sets that ensure a certain user-specified probability of containing the actual label, regardless of the underlying model or distribution. This guarantee is achieved without direct access or modification to the model’s training process and only requires a held-out calibration and inference dataset. This makes conformal prediction well-suited to LLM applications when retraining is costly and direct model access is unavailable through third-party or commercial APIs.

The coverage guarantee states that the prediction sets obtained by conformal prediction should contain the true answer on average at a user-specified level, α\alpha. This property is called coverage, and the corresponding coverage guarantee is defined as:

where α∈(0,1)\alpha\in(0,1) is the desired error rate, and C\mathcal{C} is the calibrated prediction set introduced above. (Xtest,Ytest)∼Dcalibration{\left(X_{\text{test}},Y_{\text{test}}\right)\sim\mathcal{D_{\text{calibration}}}} is an unseen test point that is drawn from the same distribution as the data used to calibrate the prediction sets.

Conformal Calibration Procedure. As previously mentioned, conformal prediction only needs the scores of a model to calibrate and construct the prediction sets. We now describe how to calibrate the prediction sets for a specific score function.

Let f:X→Δ∣Y∣f:\mathcal{X}\rightarrow\Delta^{\lvert\mathcal{Y}\rvert} be a classifier with a softmax score, where Δ\Delta is a ∣Y∣\lvert\mathcal{Y}\rvert-dimensional probability simplex. A common choice for the score function, least ambiguous set-valued classifiers (LAC) (Sadinle et al., 2019), is defined as

where [f(X)]Y\left[f(X)\right]_{Y} is the softmax score at the index of the true class.

To calibrate the prediction sets to our desired level of coverage, we need to estimate a threshold q^α\hat{q}_{\alpha} that is the 1−α1-\alpha quantile of the calibration scores

where {s1,…,sn}\{s_{1},\ldots,s_{n}\} are the LAC scores of the calibration set.

At inference time, prediction sets can be constructed in the following manner:

Exchangeability assumption. Conformal prediction assumes that the data used to calibrate the prediction sets is exchangeable with the test data at inference time. If this assumption holds, the coverage guarantee, as stated in Equation 1, will hold, and the resulting prediction sets will have the desired error rate.

Exchangeability can be viewed as weaker than the independent and identically distributed (IID) assumption (Bernardo, 1996). This assumption is often made in machine learning with regard to the training, validation, and test sets. The threshold used to determine the size of the prediction set is estimated on a held-out calibration data set that is assumed to be exchangeable with the test distribution.

Prompt Engineering

In this paper, we focus on the task of multiple-choice question answering (MCQA) and frame MCQA as a supervised classification task, where the objective is to predict the correct answer choice out of four possible options. We wish to quantify the model uncertainty over the predicted output using conformal prediction. We condition each option choice (A, B, C, and D) on the prompt and question and use the LLaMA-13B model (Touvron et al., 2023) to generate the logit corresponding to each multiple-choice answer. We normalize the four logits using the softmax to obtain valid probabilities for each option.

One-shot prompting. LLMs are very sensitive to the exact input prompt, which has motivated a whole field of in-context learning and prompt engineering or prompt tuning (Zhou et al., 2023; Wei et al., 2023). Context learning refers to the ability of LLMs to understand and make predictions based on the context in which the input data is presented without updating the model weights. Prompt engineering methods vary significantly among tasks and require heavy experimentation and reliance on hand-crafted heuristics. For the current setup, model performance on classification tasks is often sensitive to the prompts used. Thus, we experiment with several prompting strategies before finalizing our prompts.

We use one-shot prompting by including one context example. For each subject, we use a slightly different prompt. For example, we prompt the model to assume it is the “world’s best expert in college chemistry” when generating predictions for college chemistry subjects.

We also use ten different prompts for each subject to generate ten softmax probability outputs to reduce variance. We obtain the final probability outputs for a question by averaging the softmax outputs corresponding to these ten prompts. The ten prompts for a given subject only vary in terms of the one-shot question. A sample prompt for high school biology is provided below:

GPT-4 generated examples. We explore two approaches for the one-shot example in the prompts: (1) One-shot example is one of the questions in the MMLU dataset for that subject. We then exclude this specific question for generating predictions with the resulting prompt. (2) We use GPT-4 to generate multiple-choice questions for each subject. We then cross-check the questions and answers produced by GPT-4 for correctness and select ten correct question-answer pairs.

We use the following prompt to generate MCQs for clinical knowledge from GPT-4: “Give me 15 multiple choice questions on clinical knowledge with answers”. Specific questions and answers generated by the GPT-4 are available from our code (refer to Section 4.4.) We have also included a subset of sample GPT-4 generated questions and answers as well as MMLU-based questions and answers in the Appendix (A.1 )

We generate MCQs for other subjects using similar prompts. GPT-4-based one-shot questions produce more accurate answers than MMLU-based questions, as shown in Figure 1.

After controlling for the size of the prompts (limited to 700 tokens), we find that MMLU-based and GPT-4 based one-shot questions produce similar accuracy on the sixteen subjects we evaluate. We conduct all the following experiments on prompts that use GPT-4-based one-shot questions since they are shorter on average and achieve similar performance.

Experiments

We use the LLaMA-13B model (Touvron et al., 2023) to generate predictions for MCQA. LLaMA-13B is an open-source 13 billion parameter model trained on 1 trillion tokens and has been shown to achieve good zero-shot performance on various question-answering benchmarks. For our dataset, we use the MMLU benchmark (Hendrycks et al., 2021), which contains MCQA questions from 57 domains covering subjects such as STEM, humanities, and medicine.

For our experiments, we considered the following subset of MMLU: computer security, high school computer science, college computer science, machine learning, formal logic, high school biology, anatomy, clinical knowledge, college medicine, professional medicine, college chemistry, marketing, public relations, management, business ethics, and professional accounting. We group these domains into three broad categories: “business”, “medicine”, and “computer science”. These 16 subjects represent diverse domains and have sufficient samples (each with at least 100 questions).

We perform classification by obtaining logit scores corresponding to option choices ‘A’, ‘B’, ‘C’, and ‘D’ conditioned on the one-shot prompt and the question. For example, for the sample prompt and question pair described in section 3, we find the logit score corresponding to the next tokens corresponding to each of the four options. We then take the softmax over the logit scores corresponding to the options choices to obtain probability scores. The softmax scores corresponding to ten different prompts (that vary in terms of one-shot questions) are averaged to obtain final probability scores for each question-option pair.

2 Setup

We randomly split the data into equal-sized calibration and evaluation sets for each subject and averaged results over 100 random trials for our conformal prediction experiments. For each trial, we randomly sample 50%50\% of data for calibration and 50%50\% to evaluate coverage and set size. Thus, we have at least 50 samples for calibration. While the theoretical guarantee of conformal prediction holds on average for even such a small number of calibration samples, the individual 100 random trials may not always have exact coverage. A higher calibration size can reduce variance in coverage associated with the different random trials (Angelopoulos & Bates, 2021b).

3 Results

Naive Calibration in LLMs. Previous works have studied calibration in context to LLMs. (Si et al., 2022) looked at the limitation of traditional calibration metrics like Expected Calibration Error (ECE) and Maximum Calibration Error (MCE). (Jiang et al., 2021) looked at T5, BART, and GPT-2 language models and found that the models are not calibrated for question-answering tasks. More recently, (Kadavath et al., 2022) found that large language models are well calibrated for various MCQA tasks. In current work, we examine the calibration error in the softmax probability output for the MCQA task for the LLaMA-13B language model. To this end, we calculate the Expected Calibration Error (ECE) and Maximum Calibration Error (MCE), metrics that measure the average and maximum discrepancy between the confidence of the model’s predictions and their accuracy. We find that the naive softmax output of the model is reasonably well calibrated across subjects on average, with ECE varying between a minimum of 1% for high school biology to a maximum of 7% for marketing (refer figure 9 in the appendix.) This aligns with previous findings on calibration error in LLMs (Kadavath et al., 2022). Nonetheless, MCE is significant for most subjects, indicating that the model is under-confident or over-confident at specific confidence levels. Additionally, there are no formal guarantees in terms of calibration errors.

Difference in coverage and set sizes between subjects. We next implement the conformal prediction procedure and compare coverage and prediction set size between subjects in Figure 3 and Figure 4 at the error rate α=0.1\alpha=0.1. The coverage guarantee of conformal prediction holds across all subjects (Figure 3). Comparing Figure 2 and Figure 4, we see that for each of the three categories, uncertainty — as measured by prediction set sizes — is, in general, significant for subjects with low top-1 accuracy and low for subjects with high top-1 accuracy.

For example, more challenging subjects such as formal logic and college chemistry have the most uncertainty on average, while “easier” subjects such as marketing have the lower average uncertainty. We show more results for different α\alpha values in Table 1.

Selective classification with conformal prediction. Conformal Prediction framework can also be used for selective classification (Angelopoulos et al., 2022b; Angelopoulos & Bates, 2021a). In Figure 5, we analyze the correlation between uncertainty (as measured by conformal prediction) and top-1 accuracy performance. Specifically, we look at top-1 accuracy across subjects stratified by the size of the prediction set outputted by conformal prediction. We find a robust negative correlation between set size and top-1 accuracy for all subjects. This is intuitive as models with low confidence scores should correspond to less accurate predictions.

The accuracy for prediction sets with only one prediction is significantly higher than naive top-1 accuracy, as shown in Figure 7 (refer k=1k=1 accuracy). Thus, our results demonstrate that the set size obtained from conformal prediction procedure can filter low-quality predictions in downstream applications for LLMs. For example, highly uncertain predictions in a disease screening application should be flagged for manual review and not shown to the user.

Size-stratified coverage and comparison with naive top-kk prediction sets. Size-stratified coverage measures error-rate guarantee across prediction sets of different sizes (Angelopoulos et al., 2022a). This experiment shows that coverage is not trivially satisfied by naively forming prediction sets by simply taking the top-kk highest softmax probabilities. In Figure 7, we show the coverage when all prediction sets have a fixed set size and find that coverage decreases sharply with size. This is in contrast to prediction sets formed by conformal prediction in Figure 6, where we find that even prediction sets of size one have close to the desired level of coverage (90%90\% when α=0.1\alpha=0.1) across most subjects. Indeed, we found that coverage is consistent over all set sizes for conformal prediction.

Conformal prediction can be thought of as outputting “adaptive” prediction sets that try to attain the proper level of coverage (depending on the chosen error rate α\alpha) instead of “fixed” prediction sets of size kk.

Exchangeability assumption across subjects. In Figure 8, we test the exchangeability assumptions between subjects by calibrating on one subject and evaluating coverage on a different subject, grouped into three categories of subjects. Recall that the exchangeability assumption is needed for the coverage guarantee of Equation 1 to hold.

On the main diagonal, where the prediction sets are calibrated and evaluated on the same subject, we observed little deviation from the desired coverage rate of 90%90\%. For example, prediction sets calibrated and evaluated on the same subject had close to the desired error rate of 10%10\% when α=0.1\alpha=0.1. On the off-diagonal, we can see significant disparities between some subjects. For example, when prediction sets are calibrated on MCQA data from “high school computer science” and evaluated on “business ethics”, coverage is only around 83%83\%, less than the desired 90%90\% coverage. However, for subjects from similar domains and accuracy, such as “clinical knowledge”, “anatomy”, and “high school biology”, we find relatively more minor deviations from the targeted coverage rate when calibrated on out-of-subject data. This may result from good generalization capabilities and relatively calibrated softmax probability (Kadavath et al., 2022) outputted by the LLMs.

4 Code Availability

We release the code at this Github repository. The code repository also contains the question-answer pairs generated by GPT-4 for our prompts.

Discussion

As Large Language Models (LLMs) become increasingly powerful and are deployed in mission-critical systems, obtaining formal uncertainty guarantees for these models is crucial.

In this work, we investigated uncertainty quantification in LLMs in the context of multiple-choice questions using conformal prediction, a statistical framework, for generating prediction sets with coverage guarantees.

We found that naive softmax outputs of LLMs are relatively well calibrated on average but can suffer from under-confidence and over-confidence, and the extent of miscalibration varies across different subjects. To have a formal guarantee on the error rate of the model prediction, we implemented the conformal prediction procedure on the naive softmax output of the LLM.

The conformal prediction framework produces valid prediction sets with error rate guarantees when calibration and evaluation sets come from the same distribution. We also explored the application of conformal prediction procedures for selective classification tasks. We found that conformal prediction can be used to discard predictions with unusual and low-quality outputs where the model is not confident, as indicated by the size of its prediction sets.

Developers of LLM systems should provide estimates of uncertainty to improve trustworthiness in their outputs to users.

Uncertainty quantification can be useful for downstream applications such as filtering biased, unusual, or low-quality outputs.

Conformal prediction is one approach to uncertainty quantification where a user-specified error rate can be statistically guaranteed when the calibration data is exchangeable with the test data.

For our specific dataset (MMLU) and LLM (LLaMA-13B), we find that softmax outputs obtained as described in section 4.1 are reasonably calibrated on average. Nonetheless, models suffer from under-confidence and overconfidence, especially at the tail ends of probability distribution (refer figure 9 in the Appendix.)

Our work has some limitations. Our findings were limited to the MCQA task on the MMLU dataset using the LLaMA-13B model. Future works could extend our findings to multiple models and data sets. Further, it would be interesting to extend the conformal prediction framework to more general settings like free-form text generation to control for inaccurate, biased, and harmful outputs from LLMs. It would also be interesting to explore exchangeability conditions in LLMs further when calibration and evaluation data sets are from different distributions (i.e., not just from MMLU), which is a more realistic scenario.

Despite these limitations, our work represents, to our knowledge, the first exploration of conformal prediction for LLMs in classification tasks. Our results contribute to the growing body of research on uncertainty estimation and generalization capabilities of LLMs and serve as a step forward in developing more robust and reliable uncertainty measures for increasingly capable large language models. Such measures are essential for ensuring LLMs’ safe and responsible deployment in mission-critical applications.

Acknowledgement

We thank Prof. Yoon Kim, Abbas Zeitoun, and Anastasios Angelopoulos for helpful discussions and feedback on this work.

References

Appendix A Appendix

A.1.2 Professional Accounting

A.1.3 Clinical knowledge