Large Language Models Are Not Robust Multiple Choice Selectors

Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, Minlie Huang

Introduction

Multiple choice question (MCQ) is a prevalent input format of large language models (LLMs). An MCQ typically encompasses a question accompanied by multiple candidate options, from which the model is tasked to select the most suitable answer, as exemplified in Figure 1. Current LLM-centric scenarios have widely utilized the task format of MCQ, for instance, within benchmarks targeted at assessing LLMs (Hendrycks et al., 2020; Zhong et al., 2023; Huang et al., 2023) and in LLM-based automatic evaluation frameworks (Chiang et al., 2023; Zheng et al., 2023). In any scenario, we always expect LLMs to robustly select reliable answers in MCQs.

Unfortunately, we observe that LLMs are vulnerable to option position changes in MCQs. We show in Table 2 that, in the 0-shot MMLU evaluation (Hendrycks et al., 2020), a simple “answer-moving attack” by always moving the golden answers to a specific position causes LLMs’ dramatic performance fluctuations. For instance, moving the golden answers to position D degrades the accuracy of gpt-3.5-turbo by 6.3 (from 67.2 to 60.9). When moving to A, llama-30b is boosted by 15.2 and surpasses gpt-3.5-turbo (68.2 vs. 65.3), which starkly contrasts with their original performance (53.1 vs. 67.2).

LLMs’ poor robustness to option position changes results from their biased behavior: they prefer to select specific option IDs as answers (like “Option A”), which we term as selection bias. As a simple verification, we randomly sampled 1,000 MMLU test samples, where we controlled the number of correct answers being A/B/C/D as 250 each. Among these samples, llama-30B selects A/B/C/D 34.6% / 27.3% / 22.3% / 15.8% of the time, while gpt-3.5-turbo for 22.5% / 25.6% / 32.3% / 19.6%, respectively (averaged over 10 runs). These proportions are statistically nonuniform (χ2\chi^{2} Test, pp-value ≪10−4\ll 10^{-4}) and align well with the performance fluctuations in Table 2.

Through extensive empirical evaluation (§2.3), with 20 LLMs, on three benchmarks, and with varying option numbers (from two to five), we show that selection bias is prevalent across various LLMs and cannot be well mitigated by simple prompting strategies (§2.6), like Chain-of-Thought prompting (Wei et al., 2022; Kojima et al., 2022). It varies with models but manifests a cross-domain similarity within the same model (§2.3). With careful ablation analyses (§2.4), we find that, contrary to the common view in previous work (Wang et al., 2023a; Pezeshkpour & Hruschka, 2023), selection bias arises less from LLMs’ position bias, where they are deemed to favor options presented at specific ordering positions (like first or second). In contrast, we pinpoint one more salient intrinsic cause of selection bias as the model’s token bias when predicting answers from the option IDs given the standard MCQ prompt, where the model a priori assigns more probabilistic mass to specific ID tokens (e.g., A/B/C/D).

To efficiently mitigate selection bias, we propose a method called PriDe (§3), referring to Debiasing with Prior estimation. In PriDe, we assume the model’s prior bias for option IDs can be separated from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples (e.g., 5%), and then applies it to debias the subsequent samples. The whole debiasing procedure needs no sample labels, takes place during the inference time, and requires only negligible extra computational costs. We demonstrate that PriDe achieves superior debiasing effectiveness to strong baselines, especially in the low-cost scenario (§4.1). Furthermore, the prior estimated by PriDe provides a good interpretation for selection bias (§4.1) and can generalize well across different domains (§4.2), which highlights its practical potential in broader scenarios.

(1) We identify the ubiquitous selection bias in LLMs and provide extensive empirical analyses (with 20 LLMs, on three MCQ benchmarks) and valuable insights on this problem. (2) We pinpoint LLMs’ token bias as one primary intrinsic cause of selection bias. (3) We propose a label-free, inference-time debiasing method PriDe, which demonstrates notable effectiveness and efficiency, interpretability, and cross-domain generalization. We hope that this work can inspire future research on the bias and robustness of LLMs.

Investigation on Selection Bias

Our study focuses on the causal, decoder-only LLMs since this architecture has become the dominant choice for modern LLMs. We experiment with 20 LLMs from popular LLM families across various sizes: llama-7/13/30/65B (Touvron et al., 2023a), llama-2(-chat)-7/13/70B (Touvron et al., 2023b), vicuna-v1.3-7/13/33B, vicuna-v1.5-7/13B (Chiang et al., 2023), falcon(-inst)-7/40B (Almazrouei et al., 2023), and gpt-3.5-turbo-0613 (OpenAI, 2022). The models except gpt-3.5-turbo are all open source on the HuggingFace website, and we can access their output probabilities. gpt-3.5-turbo is the commercial API of ChatGPT. It accepts textual prompts and returns generated texts without providing access to output probabilities.

Benchmarks

We conduct experiments on MMLU (Hendrycks et al., 2020), ARC-Challenge (Clark et al., 2018), and CommonsenseQA (CSQA) (Talmor et al., 2019), which all are MCQ benchmarks widely used for LLM evaluation. Our selection of benchmarks takes into account the diversity of tasks and domains. Specifically, MMLU and ARC consist of 4-option MCQs, while CSQA consists of 5-option ones, and MMLU covers tests from 4 domain categories spanning 57 subjects. The diverse domains facilitate us to derive general observations and enable us to explore cross-domain generalization. See Table 3 in Appendix A for detailed data statistics.

Evaluation

Our evaluation protocol follows the mainstream LLM evaluation frameworks, such as HuggingFace LLM Leaderboard, EleutherAI lm-harness, the original MMLU implementation, and OpenAI Evals (see Appendix B for their project URLs). Specifically, for open-source models, we access the output probabilities of option ID tokens A/B/C/D/E and use the maximal one as the model prediction. For gpt-3.5-turbo, we compare the golden answer with the first generated token, with the decoding temperature set to 0. See Figure 6 and 7 in Appendix C for the input formats.

Our evaluation mainly considers the 0-shot setting, which excludes biases introduced by in-context examples, but we also conduct 5-shot experiments. The in-context examples come from the development sets and are shared across all the test samples within the same task.

2 Measurement of Selection Bias

In our study, selection bias is defined as the model’s behavioral bias to select specific option IDs as answers. To measure selection bias, one naive way is based on the counting for model predictions, which, however, is susceptible to label imbalance. We instead propose to measure selection bias based on the balance of recalls of different option IDs and use the standard deviation of recalls (RStd) as a quantitative metric. This measurement is intuitive that greater recall imbalance indicates more pronounced selection bias and is not as susceptible to label imbalance as the counting-based measurement. More importantly, recall balance well reflects the model’s robustness to option position changes, as illustrated in Figure 2. Hence, we reasonably expect that reducing selection bias (measured by recall balance) will improve LLMs’ robustness to option position changes in MCQs.

Note that measuring selection bias with recall balance implies the premise that the golden answers should be randomly placed. We show in Figure 18 in Appendix F that randomly shuffling the options does not obviously change selection bias, validating the above premise.

3 Key Observations

We first conduct an extensive evaluation of LLMs on various benchmarks to gain a preliminary understanding of selection bias. We show partial results in Figure 3 for a brief presentation and put the full results in Appendix D. We draw the following main observations and insights:

Intuitively, selection bias is likely to originate from LLMs’ training data, where some answers (e.g., C) may occur more frequently than others. However, we do not observe consistent patterns of selection bias within the same model family where the models are trained with the same training data (e.g., llama-7/13/30/65B, llama-2-7/13/70B). We speculate that selection bias arises as a product of complex interactions between training data composition and ordering, model capacity (number of parameters), and other factors like hyperparameters.

Selection bias within the same LLM displays a moderate similarity across different domains.

For instance, under the 0-shot setting, llama-30B consistently prefers A/B on various benchmarks, while gpt-3.5-turbo favors C/B more. While the preference ranking may not strictly persist across tasks or domains, there is an overarching tendency for each model to lean towards certain option IDs (e.g., A and B) and away from others (e.g., C and D). It suggests that selection bias is an inherent behavioral bias of LLMs that is less impacted by tasks or domains.

In-context examples can reduce but may meanwhile alter selection bias.

As exemplified in Figure 3, llama-30B disfavors C under the 0-shot setting but becomes biased towards it under the 5-shot setting. We find that this alteration still does not display noticeable patterns within the same model family. It indicates that in-context examples can introduce new biases that will be intertwined with the inherent selection bias, making the latter complex and less regular.

4 What Causes Selection Bias?

Given the ubiquity of selection bias in various LLMs, we now seek to figure out the intrinsic causes resulting in this behavioral bias. We propose two hypotheses: (1) Token bias. In the standard MCQ prompt (Figure 1), when selecting answers from the option IDs, the model may a priori assign more probabilistic mass to specific ID tokens (such as A or C). (2) Position bias. The model may favor options presented at specific ordering positions (such as the first or second one). Note that in a recent work, Wang et al. (2023a) similarly found that GPT-4 exhibits the bias towards the responses from “Assistant 1”. However, it is still unclear whether it is because the preferred responses are “selected via the ID token 1” or because they are “presented first”.

One challenge here is that option IDs are bound with options’ ordering positions, e.g., the ID B is naturally tied with the second-presented option. To distinguish the impacts of the two hypothesized causes, we conduct two ablation experiments. (1) Shuffling option IDs. We randomly shuffle the default ID ordering A/B/C/D, for instance, into B/D/A/C or C/A/D/B, etc. In this way, B can denote the option presented at any ordering position, thus eliminating the impact of position bias and leaving only token bias. But this ablation will obviously impair the naturalness and quality of the MCQ prompt, and may consequently degrade model performance (as shown in Table 2.4). (2) Removing option IDs and asking the model to directly select option contents. In this way, the change of selection bias would indicate the impact of token bias, while the remaining part corresponds to position bias. When evaluating LLMs without option IDs, we require gpt-3.5-turbo to generate the whole selected option, which is then compared with the golden answer. For open-source models, we compute the likelihoods of options, normalized by their lengths, and use the maximum one as the model prediction. See Figure 8 and 9 in Appendix C for the input formats.

As shown in Figure 3 and Table 2.4, the removal of option IDs notably reduces selection bias (RStd decreases), while RStd is little changed by shuffling option IDs. The former observation, in most cases, holds for various LLMs, on different benchmarks, and with varying option numbers (from two to five, where 2/3-option settings are constructed from the original data), see Appendix D and Table 4 in Appendix F for detailed results. We also try to replace the default ID symbols A/B/C/D with several reasonable alternatives, including a/b/c/d, 1/2/3/4, and (A)/(B)/(C)/(D), but observe no remarkable reduction in RStd from the default one, as shown in Table 2.4. These results confirm that the model’s token bias is one primary intrinsic cause of selection bias.

However, with option IDs removed, the remaining selection bias (corresponding to the impact of position bias) varies with models and tasks, see Appendix D and Table 4 in Appendix F. For instance, the remaining selection bias of llama-13B, vicuna-v1.3-7B, and gpt-3.5-turbo is only marginal, while that of llama-30B, vicuna-v1.3-33B, and falcon-40B is still pronounced (although having been reduced much). The selection bias of llama-2-13/70B even slightly increases in MMLU and ARC after option IDs being removed while still decreasing in CSQA, implying the potential counteraction between token bias and position bias. These results suggest that the model’s position bias is somewhat present but quite irregular, largely depending on models and tasks.

5 Can We Debias LLMs by Removing Option IDs?

Despite the notably reduced selection bias, we find that removing option IDs usually degrades model performance (except in a few cases under the 5-shot setting), see Table 4 and 5 in Appendix F. This performance degradation results from the way we leverage LLMs to answer MCQs without option IDs, i.e., calculating and comparing the likelihoods of options, which is referred to as the “cloze prompt” format in Robinson & Wingate (2022). Their study demonstrates that asking LLMs to predict option IDs forms a better MCQ prompt than the “cloze prompt”, which is consistent with our observation. Besides, selecting answers by calculating and comparing the likelihoods of options is not as convenient and straightforward to implement as directly predicting option IDs. We thus suggest that removing option IDs is not a practical method to mitigate selection bias.

6 Can Simple Prompting Strategies Mitigate Selection Bias?

As a preliminary debiasing attempt, we apply two simple prompting strategies to gpt-3.5-turbo: (1) Explicit debiasing instruction: We append an explicit debiasing instruction in the system message of gpt-3.5-turbo (“Please note that the provided options have been randomly shuffled, so it is essential to consider them fairly and without bias.”). (2) Chain-of-Thought prompting (Wei et al., 2022; Kojima et al., 2022): gpt-3.5-turbo is first prompted with “Let’s think step by step:” to generate its thought process and then produces the final answer. We follow the implementation in OpenAI Evals, see Figure 10 in Appendix C for details. As shown in Table 2.4, the two prompting strategies cannot mitigate selection bias well. It suggests that selection bias is an inherent behavioral bias of LLMs that cannot be addressed by simple prompt engineering.

Methodology

Before proposing our debiasing method, we first introduce a strong permutation-based debiasing baseline that our method builds upon. It averages the model’s prediction distributions under various option permutations (Wang et al., 2023a; Zheng et al., 2023), which intuitively cancels out both the model’s token bias and position bias.

Formally, we use qq to denote the MCQ question. Suppose the nn default-ordered option IDs (e.g., A/B/C/D) are did_{i} and the default-ordered option contents are oi,i∈{1,2,…,n}o_{i},i\in\{1,2,\dots,n\}. We use II to denote a permutation of {1,2,…,n}\{1,2,\dots,n\}, I\mathcal{I} to a set of possible IIs. We use gI(i)g_{I}(i) to denote the index of ii in II, and xIx^{I} to the concatenation of the default-ordered option IDs and the II-permuted option contents, so that oio_{i} is tied with dgI(i)d_{g_{I}(i)} in xIx^{I}. The permutation-based debiasing baseline can be formulated as:

2 Prediction Probability Decomposition

3 Debiasing with Prior Estimation

Taking the logarithm of both sides of Equation 3 and summing over all I∈II\in\mathcal{I}, we can obtain:

Recall our observation in §2.3 that selection bias within the same LLM displays a moderate cross-domain similarity. This implies that the prior for option IDs is likely to generalize across different samples and domains, which motivates us to compute the priors of partial test samples and use them as an approximation for the remaining samples. It can largely improve debiasing efficiency since no more computational overhead is needed for subsequent samples once the prior is estimated.

When K≪∣D∣K\ll|\mathcal{D}|, the overhead for prior estimation will be negligible compared to the whole inference cost. The overall procedure of PriDe is summarized as Algorithm 1.

Experiments

Figure 4 presents the debiasing results (averaged over all the models) versus the computational costs under the 0-shot setting, see detailed breakdowns in Table 4 and 5 in Appendix F. PriDe achieves superior debiasing effectiveness and performance improvement to Full/Cyclic Perm, especially in the low-cost scenario. This also holds under the 5-shot setting, see Figure 19 in Appendix F. In Figure 17 in Appendix F, we show that the estimated prior manifests a clear correlation with the empirical selection bias (i.e., the recalls of different option positions before debiasing). It suggests that PriDe can provide a good interpretation for the model’s selection bias. Furthermore, we observe that the priors are stable when estimated with different sizes of test samples (from 2% to 20%). It suggests that we are able to obtain a reliable estimate of prior even with a limited computational budget. It also implies that the model’s prior bias for option IDs exhibits a similar pattern across different samples, confirming our design motivation of PriDe in §3.3.

Note that one may propose to debias with fewer permutations (such as 2 or 3 random permutations, with ×2\times 2 or ×3\times 3 costs) as a low-cost alternative to Cyclic (×n\times n) or Full (×n!\times n!) Perm. In Figure 21 and 22 in Appendix F, we show that PriDe can still be combined with these methods and notably boost debiasing effectiveness and efficiency.

2 Generalization Analysis

In practical scenarios, we may not guarantee that the test samples always come from the same or similar domains. We hope that the prior estimated by PriDe can be generalizable: Once the prior is estimated using a small number of samples from domain X, it can be used to debias not only other samples from domain X but also samples from domain Y. To this end, we evaluate the cross-domain generalization of estimated priors on the four category domains of MMLU and ARC (5 domains in total, all are 4-option MCQs). We first use PriDe to estimate the prior with α\alpha test samples from a source domain X, and then apply it to debiasing the test samples from a target domain Y. As shown in Figure 5, the estimated priors exhibit reasonable generalization across different domains.

However, we find that although the generalization in terms of debiasing is promising, there may be a slight degradation in model performance when the domain gap is large (e.g., from STEM/Humanities to ARC). While PriDe is not designed to improve model performance but rather to mitigate selection bias and thereby improve LLMs’ robustness, we suggest updating the estimated prior using new samples when there are predictable domain shifts in test samples, whose overhead is still negligible compared to the whole inference cost.

3 How Debiasing Affects Model Predictions?

We notice that although it is not our initial intention, the debiasing methods (PriDe and Cyclic/Full Perm) usually improve the model performance. With PriDe (α=5%\alpha=5\%) and Cyclic Perm as examples, we seek further insights on how debiasing affects model predictions and consequently improves model performance. We are especially interested in how PriDe works with the estimated prior and how it works differently from Cyclic Perm.

Related Work

The realm of large language models (LLMs) has been undergoing rapid and significant developments since the launch of GPT-3 (Brown et al., 2020). Prominent examples, such as ChatGPT (OpenAI, 2022), LLaMA (Touvron et al., 2023a; b), Alpaca (Taori et al., 2023), and Vicuna (Chiang et al., 2023), have emerged successively over the past year. These models usually boast billions of parameters and are trained to comprehend natural language and follow human instructions (Ouyang et al., 2022; Wang et al., 2023b).

Multiple Choice Questions (MCQs)

As a concise task format, MCQs are widely adopted in LLM-centric scenarios. For instance, in the automatic evaluation framework of Chiang et al. (2023); Zheng et al. (2023), GPT-4 (OpenAI, 2023) is presented with a question and two model answers and is tasked to determine which one is better. Numerous standard language model benchmarks also employ the task format of MCQs, such as MMLU (Hendrycks et al., 2020), ARC (Hendrycks et al., 2020), AGIEval (Zhong et al., 2023), C-Eval (Huang et al., 2023). It thus piques our research interest in the robustness of LLMs in the context of MCQs.

Bias and Robustness of LLMs

In our study, we use “bias” to refer to the systematic error within LLMs rather than prejudice in culture, gender, etc. The bias in language models has always been an important research area closely related to model robustness. For instance, Zhao et al. (2021) showed that GPT-3 is sensitive to task instructions and in-context examples, which arises from its bias towards certain answers. Chen et al. (2022); Si et al. (2023); Pan et al. (2023) explored how LLMs’ few-shot learning is influenced by the construction of in-context examples. Wang et al. (2023a); Zheng et al. (2023) found that GPT-4 leans towards the first presented answers and may produce unfair evaluation results.

Contemporaneous with our work, Pezeshkpour & Hruschka (2023) similarly observed that LLMs are sensitive to option position changes in MCQs and verified position bias as one cause of sensitivity. However, their study does not ablate the impact of option IDs, making their investigation on “position bias” less convincing. In fact, our study finds that position bias is less regular and largely depends on models and tasks, while token bias is a more salient intrinsic cause of LLMs’ selection bias and, consequently, poor robustness. Their study is also conducted on very limited models (only GPT-4 and InstructGPT without open-source LLMs, somewhat hindering reproducibility) and tasks (only three MMLU subtasks, one Big-Bench subtask (Srivastava et al., 2023), and CSQA). In contrast, we draw more general observations through extensive cross-model and cross-task empirical evaluation, which further inspires our proposal of the computation-efficient, interpretable, and generalizable debiasing method PriDe.

Conclusion

This work studies the inherent selection bias of large language models (LLMs), which makes them vulnerable to option position changes in multiple choice questions (MCQs). Through extensive empirical analyses, we pinpoint that this behavioral bias stems primarily from token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens when predicting answers from the option IDs, and partially from position bias, where the model favors options presented at specific ordering positions. In particular, token bias is a more salient intrinsic cause of selection bias, while position bias is less regular and depends on models and tasks. Our proposed debiasing method PriDe estimates the model’s prior bias for option IDs and separates it from the overall prediction distribution. It remarkably mitigates selection bias, with no need for sample labels and only negligible computational overhead. We especially highlight PriDe’s high efficiency, interpretability, and cross-domain generalization. We hope that the empirical analyses in this work and our debiasing method can inspire future research on the bias and robustness of LLMs.

References

Appendix A Data Statistics

Appendix B Reference Projects

https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard

EleutherAI lm-harness

https://github.com/EleutherAI/lm-evaluation-harness

Original MMLU implementation

OpenAI Evals

Appendix C Prompts Used in Experiments

Appendix D Evaluation Results of Selection Bias

Appendix E Proof for Permutation-based Debiasing Baseline

Our probability decomposition assumption (Equation 3) can also give a theoretical proof of the soundness of the permutation-based debiasing baseline (Equation 1). Note that gIg_{I} is the inverse mapping of fIf_{I}: gI=fI−1g_{I}=f_{I}^{-1}, so we can rewrite Equation 3 as:

By taking the logarithm and summing over I∈II\in\mathcal{I}, we can similarly derive:

Appendix F Supplementary Experimental Results