Debating with More Persuasive LLMs Leads to More Truthful Answers

Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, Ethan Perez

Introduction

Most existing approaches to align LLMs rely on the availability of labelled data (Ouyang et al., 2022; Menick et al., 2022). However, faced with models that can answer questions in increasingly broad context, obtaining such data requires domain expertise (OpenAI, 2023; Gemini Team et al., 2023). As these systems continue to advance, they will surpass expert knowledge. Consequently, there will be no ground truth to rely on, rendering most alignment approaches unusable. We need mechanisms that provide scalable oversight (Amodei et al., 2016; Christiano et al., 2018; Irving et al., 2018; Bowman et al., 2022): alignment methods that scale with model capability.

A promising paradigm to overcome the need for ground-truth labels is using less capable models to align stronger models (Cotra, 2021; Bowman et al., 2022; Burns et al., 2023). Fundamental to these approaches is the assumption that it is easier to identify or critique the correct answer than it is to generate it (Goodfellow et al., 2014; Bai et al., 2022; Saunders et al., 2022). However, evaluating the correctness of a critique can be difficult, in which case a critique of a critique can help. In similar spirit, Irving et al. (2018) propose debate as a method to align superhuman AI. The idea is that a human or a weaker model can judge the correctness of an answer more accurately by using the adversarial critiques that are generated by a debate.

To address the challenge of evaluating models without ground truth, we investigate the question: can debateNote that the form of debate here is taking inspiration from Irving et al. (2018), but is operationalised differently. aid weaker judges in evaluating stronger models? To simulate stronger and weaker models, we use information-asymmetric debates in a reading comprehension task (Michael et al., 2023; Radhakrishnan, 2023). We give the debaters full access to the underlying text while judges have no access to the text. This setup renders the debaters experts (i.e. stronger) and the judges non-experts (i.e. weaker) in the task. Experts are provided with a quote tool, which allows them to present externally verified quotes from the text. The judges use the arguments from the debate to answer the reading comprehension question. We test this setup for both human and LLM judges and evaluate against a non-adversarial baseline called consultancy, in which a single expert model argues for one particular answer.

To investigate how debate will scale with increased model capabilities, we introduce a metric called persuasiveness. Persuasiveness is measured by judge approval, meaning it does not require ground-truth labels. We optimise model outputs for persuasion with inference-time methods, generating a range of expert models. The resulting models are evaluated by unseen LLM judges whose preferences have not been used in optimisation. Finally, we evaluate debate and consultancy with human judges, conducting a large-scale study with over a thousand hours of annotator time spent providing judgements. Our findings are as follows:

Weak judges can supervise strong debaters. The result holds both when using LLMs and when using humans to judge outputs. Specifically, for the most persuasive models we find that non-expert human judges achieve 88% and non-expert LLM judges achieve 76% accuracy with debate, where naive performance is 60% and 48% respectively. Debate also outperforms the single-model baseline consultancy, with which human and LLM judges achieve 78% and 54%, respectively.

Optimising debaters for persuasiveness improves a judge’s ability to identify truth in debates. Using inference-time methods such as best-of-NN and critique-and-refinement, we find that models optimised for judge approval become better at arguing for the correct answer relative to arguing for the incorrect answer. In particular, using persuasive debaters leads to higher judge accuracy. By contrast, judge accuracy decreases as consultants are more persuasive. We find that this effect generalises to unseen judges whose preferences have not been optimised for.

Human judges are well calibrated and achieve a lower error rate with debate. Based on confidence ratings of human judges, we find that judges are better calibrated with debate than with consultancy. Furthermore, debate achieves higher judge accuracy than consultancy across all rejection thresholds.

Although greater access to information is only one way in which future models may be stronger than their supervisors, our results pave the way for further research on adversarial oversight methods. We provide empirical evidence in one domain indicating that as models get more capable, debate enables scalable oversight by both human and model judges.

Methods

We are concerned with protocols that enable non-experts to elicit truth from experts. Here, we discuss the protocols we investigated and the task setting in which they are evaluated, as illustrated in Figure 2. Furthermore, we define novel metrics for evaluating the strength of the debaters that do not rely on labelled data.

Debate — We first introduce debate, a protocol in which two expert models (debaters) argue for opposing answers to a question. Debate runs for a pre-determined number of rounds NN, during which a transcript of the debaters’ arguments is kept. In each round, debaters see the arguments from previous rounds and simultaneously generate their arguments for the next round. After NN rounds, a judge reads the transcript and attempts to choose the correct answer. Each debater tries to convince the judge to pick their answer, and the judge is tasked with picking the correct answer. The adversarial nature of the protocol stems from the conflicting incentives between the debaters, as each debater strategically presents arguments to explain why their opponent’s claims are false. At the start of a round, debaters receive nearly-identical prompts explaining the game, their assigned answer, and the current transcript.

Interactive debate — Similar to debate, interactive debate uses two experts arguing for opposing answers. In addition to the mechanisms above, after each round, the judge may make a statement. This enables the judge to highlight sub-problems, ask for clarifications, and direct questions to either debater. The statement is appended to the transcript and is visible to all players for the rest of the debate.

Consultancy (Baseline) — For comparison with debate, we use the consultancy baseline established by (Michael et al., 2023). In consultancy, a single expert model (the consultant) is assigned a specific answer and aims to persuade the judge that their answer is correct. The judge aims to elicit the correct answer, asking the consultant probing questions. Consultancy runs for a pre-determined number of rounds NN (fixed to be the same as debate), in which the consultant and judge sequentially make statements. At the end of consultancy, the judge decides which answer to choose. At the start of each round, the consultant is provided with a prompt containing the rules of the game, their assigned answer and the current transcript. Consultancy builds up a transcript of a dialogue between the consultant and judge as rounds continue. In all our evaluations, we run consultancy for both the correct and incorrect answers; producing the same 50/50 prior as debate.

2 Experimental Details

To operationalise the oversight problem, we conduct information-asymmetric experiments (Michael et al., 2023) in which non-experts do not have access to the necessary information to answer questions, but the experts do.

Task — We evaluate the ability of non-expert judges to answer questions from the reading comprehension dataset “Question Answering with Long Input Texts, Yes!” (QuALITY; Pang et al., 2022). Judges are not allowed access to the original comprehension text, restricting their ability to answer questions and inducing their non-expert status. We motivate this setting and provide implementation recommendations in Appendix C. We use texts from the Project Gutenberg science-fiction story subset, which have a median length of approximately 7000 tokens. From these questions, we select the hard subset, where all untimed annotators chose the correct answer and rated the answer as unambiguous. For each question, we provide two answers: the correct answer and the best distractor. We use two data splits for LLM judge experiments: TLT_{L} (400 train set questions) and DLD_{L} (291 development set questions). For human experiments, where a story can only appear once, we use THT_{H} (153 drawn from both sets) and DHD_{H} (47 drawn from both sets). For question filtering details, see Appendix D.1.

Quote verification tool — In each protocol, expert models have access to a quotation tool to present quotes from the text in tags. Similar to Radhakrishnan (2023), if the quote directly matches a portion of the text, the tool generates verified else, it generates unverified tags (see Appendix D.2). The judge is instructed to trust only verified quotes (see Appendix D.8).

Number of rounds and word limit — We run protocols for three rounds. For analysis of the number of rounds, see Appendix A.8. To control for the quantity of information presented to the judge across protocols and mitigate the LLM judge verbosity bias, we restrict transcripts to 900 words in total, limiting consultants to 300 words per argument and debaters to 150 words (see Appendix D.6).

Language models — We evaluate four large language models as debaters and consultants: GPT-4-Turbo, GPT-3.5-Turbo, Claude 2.1 and Claude 1.3. GPT models have been fine-tuned via reinforcement learning from human feedback (Ouyang et al., 2022), and Claude models have been fine-tuned via Constitutional AI (Bai et al., 2022). Debaters are presented with an egocentric view of the transcript, in which their arguments appear first. We present all models with a scratchpad (Wei et al., 2022; Nye et al., 2022), which allows the model to output a set of tokens enclosed in tags that other players will not see (see Appendix G.1).

To generate a large set of different debater capability levels for our LLM judge experiments, we apply inference-time optimisation using best-of-NN (boNN) sampling. With boNN, models are sampled NN times, and a preference model is used to select the most persuasive arguments. The preference model is given a prompt to optimise for judge feedback. For a subset of models, we additionally use critique-and-refinement (cNN) (Saunders et al., 2022), in which another model is used to produce NN critiques; a preference model then grades these and the most helpful is used to refine the answer. For all preference models, we use the log probabilities from GPT-4-Turbo. We document prompt iteration in Appendix D.3 and learning with examples in Appendix A.9.

For LLM judging, we used GPT-4-Turbo, GPT-3.5-Turbo and Claude 2.1. Prompts are provided in Appendix G.2. Similar to Zheng et al. (2023), we found strong biases with LLM judges. To mitigate positional biases, we run judgements over each debate twice - once with the original transcript and once with the debater order swapped. To mitigate verbosity biases, we apply strict word limits. We document judge biases in Appendix A.7 and mitigation techniques in Appendix D.6.

Cross-play tournament — To investigate how debate scales with model capabilities, we require a way to compare debaters. We run cross-play debates, in which different debaters argue against one another. To compare a large set of debaters, we use a tournament to generate informative matches. For nn debaters, it is not feasible to run all possible matches as this requires O(n2)O(n^{2}) matches. Instead, we use a Swiss-style tournament to decide matches, which produces rankings in O(nlog⁡n)O(n\log n) matches (see Appendix D.4).

Human judges — We recruit 30 human judges via the referral-based annotator platform Pareto.https://pareto.ai/ Annotators judge debates via a web interface in which they are presented with the question, answers, and debate transcript. The 30 annotators go through a training phase and are filtered down to 21 based on performance in a test phase. To avoid any optimisation based on ground truth labels (Perez et al., 2021), we filter annotators based on a combination of inter-annotator agreement and the quality of explanations they provide (see Appendix E.2).

Additional base- and toplines — We introduce a naive judge and an expert judge to evaluate the worst and best-case performance. For the naive judge, we present only the question and possible answers to a non-expert model; this helps identify if there is any contamination of the dataset in the non-experts’ training. For the expert judge, we present the full text to demonstrate a reasonable upper bound for what a non-expert judge can hope to achieve. For human experiments, we did not run an expert judge baseline, but use Pang et al. (2022) who report annotators with access to the text achieve 89% accuracy on the hard subset. We filter for 100% untimed accuracy, and as such, expect 89% to be a lower bound for expert human performance.

3 Measuring Debater and Judge Performance

We are interested in oversight protocols with no access to labels, but existing work uses metrics that assume access to the ground truth. In this section, we introduce metrics to measure debater capabilities without this assumption. Given a dataset D\mathcal{D} consisting of NN questions (qq) and candidate answers (a1,a2a_{1},a_{2}), D={qi,ai1,ai2}i=1N\mathcal{D}=\{q_{i},a_{i1},a_{i2}\}_{i=1}^{N}, a debate match is defined over a dataset by (D1,D2,J)(D_{1},D_{2},J), comprising two models D1D_{1} and D2D_{2}, and a Judge JJ. The first argument refers to which model is assigned a1a_{1} to support (in this case D1D_{1}), and the second argument refers to which model is assigned a2a_{2}. A consultancy match with a consultant DD is expressed as (D,−,J)(D,-,J) or (−,D,J)(-,D,J).

Win rate — We define the win rate as the frequency with which a judge picks a specific debater’s answer. For the match (D1,D2,J)(D_{1},D_{2},J), the win rate, ω1\omega_{1}, for Debater D1D_{1} is given by:

In practice, this win rate depends on the assignment of the answer choice, which gives an unfair advantage, as some assignments may be easier to argue for than others (e.g. some answers are easier to defend than others). To mitigate this, we flip assignments such that D1D_{1} and D2D_{2} argue for the opposite assignments and then take the average, ωˉ1\bar{\omega}_{1}. Assuming debaters perform equally well across all questions, we say D1D_{1} is more persuasive than D2D_{2} if:

Elo rating — To capture a general notion of persuasiveness, we evaluate aggregate win rates for each debater across a population of co-debaters. We parameterise win rates by a latent skill, using the Elo ranking metric (Elo, 1978). We calculate ratings by minimising error between predicted and observed win rates (see Appendix D.5).

Aggregate rating is defined as the latent skill for a debater independent of assignment. Given two debaters, D1D_{1} and D2D_{2}, and a judge JJ, we define each player’s aggregate rating, E1E_{1} and E2E_{2}, as the solution to the following equality:

Win rate and aggregate rating are unsupervised metrics and can be used to optimise oversight protocols without ground truth labels.

To understand truthfulness in debates, we also consider the win rate ωC\omega_{C} conditioned on being assigned all correct answers (assuming ground truth labels). We calculate a players’ correct rating ECE^{C} and incorrect rating EIE^{I} as:

Judge accuracy — In line with the original motivation for debate, we calculate judge accuracy, α\alpha, as the accuracy over a match between copies of the same model; α(D,J)=α(D,D,J)\alpha(D,J)=\alpha(D,D,J). As the consultant’s intention is unknown to a judge, we take accuracy over both potential matches for consultancy α(D1,J)=12(α(D1,−,J)+α(−,D1,J))\alpha(D_{1},J)=\frac{1}{2}(\alpha(D_{1},-,J)+\alpha(-,D_{1},J)).

Results with Language Model Judges

Insight 1: Debate helps LLM judges answer hard questions better than consultancy. We perform hyper-parameter and prompt optimisation for debaters by optimising for win rate (see Appendix D.7). We evaluate each protocol on test sets (THT_{H}, TLT_{L}), using GPT-4-Turbo and best-of-16 sampling for both debaters and consultants. To evaluate performance, we report judge accuracy along with standard error of the mean. We find that all protocols help non-experts answer questions better than the naive judge baseline (see Figure 1 and 9). Debate outperforms consultancy and, with stronger judges, reduces the gap with the expert judge (75.6% vs 92.5%), demonstrating promising performance at eliciting the truth. We provide sample transcripts in Appendix H.

Insight 2: Arguing for the correct answer provides an advantage to debaters, and this advantage increases with persuasiveness. In a cross-play tournament, we evaluate 20 different debaters against distinct opponents (e.g. no model plays against itself). Each debater is one of four different LLMs, augmented by a combination of critique-and-refinement or best-of-NN. In each augmentation, we use a preference model without access to the text (like the judge model), thereby simulating optimising against the judge. Debaters play four matches each, with GPT-4-Turbo as a judge. Cross-play matches were generated on DLD_{L}, obtaining 80 win rates, which were used to generate Elo ratings (outcomes listed in Appendix F).

We present Elo ratings for each model in Figure 3-left and -middle. For a given model, we observe that models are more persuasive (obtain higher Elo) when assigned to the correct answer, and the advantage increases for models with higher aggregate Elo. This suggests that, for debates in QuALITY, telling the truth is easier than lying (a central assumption of Irving et al.). We find this result holds across different judges (see Figure 5) and that the strongest indicator of persuasiveness is base LLM, with GPT-4-Turbo outperforming others. Moreover, models are less persuasive when using critique-and-refinement than when using best-of-NN, suggesting generating useful critiques for refinement is more difficult than identifying good arguments.

Insight 3: Higher aggregate Elo rating leads to higher judge accuracy. Given a ranking of debaters, we next evaluate how aggregate Elo affects judge accuracy. We take each debater from the tournament and generate self-play debates (in which copies of the same model play against each other); we use GPT-4-Turbo to judge these debates. To ensure that our results generalise, we use separate questions for cross-play (DLD_{L}) and self-play (TLT_{L}) matches.

We report judge accuracy against aggregate rating, noting a strong correlation between these two values (Figure 3-right). As debaters are optimised against the unsupervised metric of win rate (judge preference), judge accuracy improves, suggesting that training models to win debates may produce more truthful outcomes. This provides weak evidence that under optimal play debate produces truthful information (Irving et al., 2018), and suggests that with even stronger debaters, non-experts using debate can obtain high accuracy. We analyse debater characteristics in Appendix A.4.

Insight 4: Higher consultant skill leads to worse accuracy. We generate a large set of consultants using best-of-NN and critique-and-refinement augmentations. We find that consultants under optimisation are more persuasive, demonstrated by a higher win rate (see Figure 4-left). Since models are fine-tuned using RLHF, they tend to perform well when advocating correct answers and poorly for incorrect answers. Optimisation might disproportionately improve consultants’ ability to advocate for incorrect answers as it provides an opportunity to explore deceptive approaches. This results in a degradation of judge performance, as the judge a priori does not know if a consultant is arguing the correct or incorrect answer (see Figure 4-middle and -right). We directly compare to debate in Appendix A.3.

Insight 5: More capable judges are better at identifying correct arguments across all debater strengths. To explore how different judge models affect debate performance, we re-run the same cross-play matches with Claude 2.1 and GPT-3.5-Turbo judges. Each judge produces different win rates, aggregate ratings and judge accuracy for each debate (see Figure 5). Strong judges generate a larger range of aggregate debater ratings than weak judges; they can distinguish between good arguments more easily, leading to higher accuracy across the full range of debater Elos. We find that even when preference and judge models are different LLMs, strong debaters improve debate accuracy. See approaches to improve judge performance in Appendix A.5.

Results with Human Judges

We are interested in understanding whether our LLM judge results generalise to human judges, whether human judges perform better than LLM judges, and any differences in how the protocols should be implemented when using humans. To control for the amount of human interaction required, we investigate two variants of each protocol (see Table 1). In static protocols, there is a fixed transcript with just a single human judge interaction at the end to choose the answer. In interactive protocols, the human judge interacts with the models throughout, e.g. by asking clarifying questions.

We conduct five weeks of iteration during which judges receive seven batches of questions (see Appendix 29). For the test phase, we release a batch of 47 questions per judge, which is completed over a period of four weeks. Answer position and assignment are randomised (see Appendix E). To reduce learning effects, after initial training, judges are not given feedback. We collect 953 human judgements in the test phase, 2362 in the iteration phase, and 6476 in total.

Insight 6: Debate protocols produce higher accuracy than consultancy protocols for human judges. We report judge accuracy over each protocol (Figure 6-left). We find that in both static and interactive settings, debate outperforms consultancy significantly (pp=0.001, pp=0.048, respectively). In consultancy, incorrect experts cause low judge accuracy, suggesting that it is easier to be misled in consultancy. Across all protocols, humans have higher judge accuracy than LLM judges (Figure 1).

Insight 7: More non-expert interaction does not improve accuracy. We find identical judge accuracy between static and interactive debate. This suggests that adding non-expert interactions does not help in information-asymmetric debates. This is surprising, as interaction allows judges to direct debates towards their key uncertainties.

Insight 8: Highly confident debate judgements are reliable. Human judges give their answers as confidences between 5% and 95%. To assess the performance of human judges at varying confidence thresholds, we consider classification with rejection. For example, in static debate, rejecting judgements with confidences below 75% retains 65% of questions and achieves a judge accuracy of 94%. We find that across all rejection thresholds, debate protocols achieve a higher judge accuracy than consultancy protocols (Figure 6-middle). Human judges are also better calibrated using debate protocols than in consultancy protocols (Figure 6-right). We explore additional methods to improve human judge reliability in Appendix 27.

Insight 9: LLM debaters are the bottleneck for higher judge accuracy. We conduct error analysis by reading transcripts and judge explanations from all cases where a human judge chooses the incorrect answer. For debate, the most common error is weak arguments from the correct debater, usually due to poor choice of quotes. This suggests that our debaters are far from optimal. For consultancy, the most common error is the judge failed to elicit key quotes from the incorrect consultant (see Appendix B.1).

Related Work

Many approaches exist that attempt to supervise strong models (Christiano et al., 2018; Bowman et al., 2022). Similar to debate, methods attempt to exploit the fact that it is easier to identify a correct solution than it is to generate a correct solution (Christiano et al., 2017; Stiennon et al., 2020; Saunders et al., 2022). Other approaches encourage models to decompose their reasoning (Nye et al., 2022; Wei et al., 2022; Radhakrishnan et al., 2023; Yao et al., 2023), similar to how debate generates a transcript. Alternatively, we can develop inductive biases that allow stronger models to be supervised directly by weaker models (Burns et al., 2023).

Other approaches exist to augment human decision-making with human-AI teams. For example, combined teams can improve reasoning in credit risk prediction (Chromik et al., 2021). In comparison, we leverage more general LLMs, which can be applied over a series of tasks, e.g. learning how to generate jailbreaks (Nikola, 2023). Human-AI teams have been shown to be overly confident in their suggestions (Bansal et al., 2021), whereas we find human judges in debate to consistently be underconfident.

Irving et al. (2018) originally proposed the ‘debate game’ as a mechanism for training safe AI systems. Since then, a body of work has focused on verifying the usefulness for oversight (Barnes, 2020; Parrish et al., 2022b, a; Michael et al., 2023). These studies are all conducted with human debaters, while in our work the debaters are LLMs. Using LLM debaters ensures that we can control for debater skill and investigate self-play debates.

There is much previous work with LLM debaters (Perez et al., 2019; Michael et al., 2023; Radhakrishnan, 2023; Du et al., 2023). In Perez et al. (2019), debates are conducted over comprehension, but debaters are limited to extracting relevant statements from a source text, not generating their own arguments. Michael et al. (2023) introduces the information-asymmetric debate setting for QuALITY but found no positive results when using LLM debaters. Their focus was primarily human debaters and, therefore, they did not consider cross-play win rates for improving debater capabilities. Concurrent to our work, Radhakrishnan (2023) conducts debaters in QuALITY; by contrast, their protocol is only a single turn of debate, focuses on training debaters via reinforcement learning, and does not use human judges.

Discussion, Limitations & Conclusion

In this work, we explore debate as a method to elicit more truthful answers from LLMs. We demonstrate that by allowing non-experts to judge a transcript between two experts, we can identify the correct answers to questions. Additionally, we show that this oversight mechanism can be automated with LLM judges. Although the original debate protocol that Irving et al. (2018) propose involves a stricter protocol in which only a sub-component has to be judged to validate the entire debate, our results show that judging over full debate transcripts is already useful for producing expert labels for data using only non-experts and untrustworthy experts. Our findings generalise to different base LLMs for both the expert debaters and non-expert judges, as well as to human judges. This indicates that the debate protocol is robust to variation in judge skill, which is important as models advance.

Our work has important limitations. In our setup, the difference between strong and weak is only in access to information. In the future, stronger models may differ in reasoning ability or another skill. Furthermore, we evaluate models that have been fine-tuned with RLHF, which have a propensity for honesty; it is unclear if debate will be a suitable technique for deceptive models (Greenblatt et al., 2023; Hubinger et al., 2024).

Finally, our results are limited to setups where the debaters can provide verified evidence to the judge (provided by the debater quote tool in our case). Without such a system, a debater arguing for the incorrect answer could simply create an alternative narrative in which their answer is correct (the judge, without access to the underlying story, would have no means to discover this). Other domains have different notions of evidence, and debater tool-use may differ or be unnecessary. For example, arguments in mathematical debates may require access to simulators, while physics debates can be grounded in experimental data. We posit that such tool-use capabilities will help judges to decide debates more quickly and accurately.

In our domain, we observe that debate becomes more truth-seeking with increased model persuasiveness. This finding is explained by models becoming relatively better at arguing for the correct answer compared to the incorrect answer when their outputs are optimised for judge approval. This indicates that optimising for persuasiveness can lead to more truthful models, paving the way for future research in fine-tuning LLMs via debate. Furthermore, we show how debate can be used to augment human judgements and generate accurate labels to questions beyond their knowledge. Overall, these results demonstrate that debate is a promising approach for scalable oversight.

Impact Statement

This work focuses on methods to allow humans to supervise models, particularly focusing on the open problem of supervision for superhuman models. We believe this work to be useful in this endeavour; however, the incorrect application of supervision can make detecting malicious behaviour harder (Hubinger et al., 2024), and as such, we recommend caution in its application.

Contributions

AK led the project, and was involved in all components. JH ran LLM experiments and led infrastructure. DV ran human experiments and developed infrastructure. KS developed infrastructure. LR, AR, EG, SB, TR and EP advised on research. LR led writing with EG, SB, EP, TR as editors. EP generated a tremendous amount of experiment ideas.

Acknowledgements

We would like to thank Jack Hopkins, Julian Michael and David Rein for their valuable technical insights and discussions. We thank Rob Kirk, Henry Sleight, Timon Willi, Maxime Beau, Nat McAleese and Jesse Mu for their comments on early versions of the paper. We thank Rebecca Ward-Diorio for help writing the paper. We thank Phoebe Yao, Des Balisa, and Ignacio Mackenney for their help running the human trials, as well as the annotators who participated. AK is grateful to Nat McAleese, Richard Ngo and Beth Barnes for many formative discussions over the years. AK was supported by the EPSRC Grant EP/S021566/1, the Effective Ventures Fund and the UCL International Scholar Award for Doctoral Training Centres. JH and DV were funded by the MATS Program https://www.matsprogram.org/ for part of the project. We also thank the OpenAI API Academic Access Program, the Claude Research Access Program, Anthropic and FAR AI for additional funding.

References

Appendix A Further Results with Language Model Judges

Leveraging the outcomes from the cross-play tournament depicted in Figure 3-Right and the consultancy findings shown in Figure 4-Right, we present a side-by-side comparison of judge accuracy and expert win rate for both protocols in Figure 7. The data presented include only Claude 2.1 and GPT-4-Turbo experts employing the best-of-N augmentation strategy. For simplicity, results from the critique-and-refinement protocol and GPT-3.5-Turbo have been excluded.

The analysis reveals a direct correlation between the persuasiveness of debaters, as enhanced by best-of-N augmentation, and the improvement of judge accuracy. Conversely, in the consultancy protocol, an increase in persuasiveness leads to a reciprocal effect on judge accuracy.

A.2 Comparing protocol accuracy with more samples

Figure 9 shows the performance of LLM judges across all our protocols with dataset TLT_{L}. This question set is approximately eight times larger than the question set THT_{H} used to compare protocols in Figure 1, which was tractable to run when not using human judges. In addition, we ran interactive debate with LLM interactive judges rather than humans, as previously presented in Figure 6. Figure 9 shows two main insights:

We find that debate is still a stronger oversight protocol over a larger set of questions than consultancy for each non-expert judge.

Interactive debate does not improve the LLM judge accuracy when the model engages with the debaters. This is consistent across each judge. In contrast, for humans, interactive debate had the same accuracy as static debate.

To analyse these results further, we provide the Performance Gap Recovered (PGR) metric (Burns et al., 2023) in Figure 9. To calculate this, we use the following formula:

The PGR metric measures the percentage of accuracy each oversight protocol recovers compared to the expert baseline, where the judge has access to the text. We find that debate recovers the most performance, and the advantage over other protocols increases as the judge’s capability increases. Conversely, in the consultancy protocol, we observe a decline in PGR performance correlating with enhancements in the base model’s capability.

A.3 Best-of-N versus critique-and-refinement

Figure 10 shows equivalent data as presented in Figures 3 and 4 and highlights the impact of augmenting the experts with best-of-NN (boNN) or critique-and-refinement for debate and consultancy.

The consultancy aggregate win rate and debate aggregate Elo can both be considered as the skill of the expert in each protocol with respect to persuading the judge. The aggregate win rate is calculated after combining the win rate across the incorrect and correct assignments in consultancy. The aggregate Elo is calculated after 20 debaters competed against each other in a cross-play tournament (Appendix 2).

Judge accuracy in debate improves with rising debater skill, yet diminishes in consultancy settings as consultant proficiency increases.

Critique-and-refinement improves skill compared to the vanilla model with no augmentation (Bo1). However, boNN leads to more skilled debaters, which continue to improve as N increases.

We find that the skill plateaus as you go beyond Bo16 for GPT-4-Turbo. Bo32 uses a higher temperature of 1.0 compared to 0.8, which may have pushed the diversity of arguments too far to see further gains.

Combining both boNN and critique-and-refinement is a competitive strategy but not as effective as boNN alone when controlling for the number of inference calls to the expert models. Applying boNN to critique-and-refinement involves sampling the initial arguments and the refinements boNN times, leading to the number of inference calls being equal to 2*boNN + cNN. Therefore, Bo4 Co8, uses 16 calls which is equivalent to Bo16.

A.4 Analysis of debater argument statistics

After running the cross-play tournament across the 20 debaters denoted in Appendix 2, we analysed the quote usage and word counts of a subset of these debaters. On the left of Figure 11, it shows the average number of quotes used by the debaters in their transcript across self-play runs on TLT_{L}. This is further broken down into verified, unverified and no quote tags. Verified refers to quotes that are direct matches, while unverified denotes those that are not. ”No quote tags” is when the debater does not use tags around their quotes, which means they also count as unverified by the judge. We find that GPT-4-Turbo debaters use predominantly verified quotes, which is likely one reason they have a higher debater skill than Claude 2.1 and GPT-3.5-Turbo. Furthermore, debaters enhanced by best-of-NN or critique-and-refinement use more verified quotes, on average, across all base models. Claude models tend to use more quotes in their arguments. However, they are shorter, as shown by Figure 12. As well as stronger debaters using more long verified quotes, we show a link with authenticity since GPT-4-Turbo uses fewer fake quotes compared to Claude and GPT-3.5-Turbo (Figure 12). Our heuristic for finding fake quotes is by normalising the quotes and the text, finding the n-gram similarity and choosing the quotes that have a similarity of 0 while being over 7 words long.

The distribution of the number of words per argument (Figure 11-Right), shows the impact of applying our strict word filter. This filter only allows arguments between 70 and 150 words. We use this to mitigate the judge verbosity bias with the methodology explained in Appendix D.6. We also observe that GPT-4-Turbo has a tighter distribution around the target word limit of 100 than other base models.

A.5 Improving LLM judge performance

Majority voting improves judge performance with debate. When running multiple judgements with a GPT-4-Turbo judge, aggregating the outcomes and taking the majority vote leads to a 4% absolute gain in accuracy (19% relative reduction in error). We sweep the temperature and show in Figure 13 that using a higher temperature of 0.8 and a judge prompt that uses Chain of Thought (Wei et al., 2022) leads to a stable performance when using as few as 3 votes. Performance does not increase when using more votes. See Figure 27 for a similar analysis with human judges.

Majority voting degrades the judge’s accuracy with consultancy. The same experiment was performed with consultancy, and Figure 13 shows a small degradation of up to 0.5% when increasing the number of votes. This shows more evidence that as the judge gets stronger, performance with consultancy decreases.

Fine-tuning improves the judge but not if using Chain of Thought. We fine-tuned GPT-3.5-turbo on 569 samples in the human-to-human debate dataset provided by Michael et al. (2023). We removed the final interactive judge question to stop bias towards the correct answer (since often the human would say, “I’m leaning towards A”). The transcripts were inserted into an identical judge prompt, and an assistant message was added with the correct letter “A” or “B”.

We found gains in judge accuracy when fine-tuning to predict only the ground truth label (e.g. A or B) in the LLM’s completion (denoted Gold Label in Figure 14). For example, the training data would be of the form: User: , Assistant: A. When we fine-tuned the judge to imitate the human chain of thought reasoning of the answer, we saw a degradation in accuracy. We believe this degradation is because deliberation is difficult for LLMs, and learning to imitate human reasoning prevents the model from learning deliberation efficiently.

Few-shot learning with 4 examples improves the judge. We collected data in the same way described for fine-tuning but additionally filtered the data by only keeping debates with 3 rounds and less than 1600 tokens in the transcript. We ensured an equal split between swap and non-swap assignments of answers to A and B (see how we define swapping answers in Appendix D.6). These examples are randomly sampled from and inserted into the start of the judge prompt. We ended up with a list of messages alternating between user and assistant messages, where the assistant messages have the ground truth answer letter. We also removed all interactive judge messages.

We swept the number of few-shot samples used by the judge in Figure 14-Right and observed that performance improves for stronger judges when using 2-4 few-shot examples (6% absolute boost for GPT-4-Turbo which is similar compared to majority voting in Figure 13). GPT-3.5-Turbo (16k context window) does not benefit from any few-shot samples, as shown by its accuracy degrading as we added more. Furthermore, the performance across all judges either plateaued or degraded with more samples depending on the data split used.

Lastly, we hypothesised that few shot learning would be higher quality when using only samples which human judges judged incorrectly. In Figure 14-Right, we observe that using those samples leads to equivalent performance up to 4-shots. Beyond this point, the accuracy of few-shot learning using the ”human incorrect” starts to degrade for GPT-4-Turbo.

A.6 Impact on using and not using the quote system

We conduct an ablation on the usefulness of the quote verification system for our debate setting. We consider three variants of the QuALITY setting, in which models can 1) only use quotes, 2) can only generate arguments, and 3) use both quotes and arguments. We run both debate and consultancy with the expert as GPT-4-Turbo with best-of-16 optimisation and non-expert as GPT-4-Turbo. We found that across both protocols, arguments-only settings result in lower judge accuracy, as experts arguing for the incorrect answer can lie more easily (see Figure 15). Conversely, using quotes only strongly benefits the expert arguing the correct answer, demonstrated by the much higher accuracy over quotes and arguments. We choose to use both arguments as this is a natural setting for human readability.

A.7 Understanding LLM judge biases

We found four types of bias in our LLM judges:

Verbosity — judges favour longer arguments, even lower quality and less persuasive arguments.

Positional — judges favour answers in certain positions over others. For example, in the debate case, this manifests as choosing A instead of B, even when arguments are swapped.

Overconfidence — judges assign miscalibrated probabilities to their chosen answers and tend to be overconfident in their judgements.

Sycophancy — judges sometimes believe unverified quotes and give them credence when they are told they should not trust them.

Self-enhancement bias is when an LLM favours completions generated by itself compared to other models. Other work has shown that this is another bias that LMs as judges have (Zheng et al., 2023). This is not an issue for our self-play debates where the same instance of the model acts as both the incorrect and correct debater. Using different judges for our cross-play tournament does alter the final ranking of some of the 20 debaters (see Appendix F). However, there is no clear sign that the judges rank debaters of the same base model more highly.

See how we address the verbosity and positional bias in Appendix D.6. See how we iterated on our prompts to reduce the impact of sycophancy bias in Appendix D.8.

We find a self-defeating bias in our debaters and consultants, where they will often concede when arguing for the incorrect answer. We use the next best ”distractor” answer (see Appendix D.1), which alleviates issues that arise from the self-defeating bias since the story can be ambiguous enough that LLMs tend not to refuse to argue for the wrong side. See our prompt iteration method to reduce the bias further in Appendix D.7.

Verbosity bias — In initial experiments, we faced challenges in determining whether enhancements in protocols or prompts were genuine or merely a result of the judge’s preference for lengthier arguments. For instance, the critique-and-refinement strategy often resulted in significantly longer arguments after integrating the critic’s feedback. Despite showing a higher win rate than the baseline, we remained uncertain if this win rate was due to the actual strength of the arguments.

We investigated verbosity bias by computing the Elo ratings of debaters from a Swiss tournament where debaters were instructed to use arguments of 100, 200, and 300 words (featuring Claude 2.1 and GPT-4-Turbo) . As depicted in Figure 16, a notable rise in aggregate Elo was observed with higher word counts, indicating a judging preference by GPT-4-Turbo towards lengthier arguments over shorter, more truthful responses.

In addition, Figure 16 shows the word count distribution when you ask the model to produce an argument with a target number of words. The distribution is often centred on the words specified but has a significant variance. Sometimes, the Claude API will write 70 words greater than what was asked for in the prompt. The distribution differs between Claude and GPT-4-Turbo and shows that GPT-4-Turbo is more accurate with a smaller variance. Based on these distributions, we chose to ask the models for 100 words and reject any that lie outside 70 and 150.

Positional bias — LLMs, as judges, have a positional bias as shown by the error bars in Figure 18. This means that judges’ accuracy can change dramatically when judges are provided with the same set of arguments but the answers corresponding to A and B are swapped (illustrated in Table 7). GPT-4-Turbo and humans have a similar positional bias but it is higher for humans in the naive case.

We observe that positional bias is significantly more pronounced in less capable models like GPT-3.5-Turbo, which lack deliberation capabilities, suggesting that these models might resort to guesswork. We hypothesised that GPT-4-Turbo’s confidence in its answers, as extracted from the log probabilities, would hover around 0.5 when the judgement outcome is altered by swapping the positions of answer choices. This was partially confirmed by the reduced average confidence for different outcomes, as illustrated in Figure 18. We noted 25% of the total cases where the model remained over 98% confident even when it reversed its decision due to the positioning of answers.

We extracted log probabilities of the chosen answer by running GPT-4-Turbo with the best-of-NN preference prompt provided in Appendix G.4. We then converted these to probabilities by taking the softmax of log probabilities of the tokens for A and B. Using GPT-4-Base would have provided more calibrated confidence. However, we found that the positional bias was so strong that the model would always predict choice A.

Overconfidence bias — We used the same method to find overconfidence bias as when finding the positional bias confidence. We extracted GPT-4-Turbo log probabilities for the predicted answer choice given the transcripts in each protocol used in the human trials: static consultancy, interactive consultancy, debate and interactive debate. Figure 19 plots calibration curves for each protocol and the selective accuracy curve in the same way as presented for human judges in Figure 6.

The calibration plots highlight that LLMs are consistently overconfident which is a difference from human judges. There is little difference between static and interactive consultancy. Interactive debate becomes more calibrated like for the human judges but this is likely because the LLM can read the human’s interactive messages.

A.8 Impact of number of rounds on accuracy

Using GPT-4-Turbo debaters with best-of-16, we run debate and consultancy for up to 6 rounds (see Figure 21). We observe that when using the debate protocol, weaker judges such as GPT-3.5-Turbo suffer from the inability to effectively deliberate as rounds continue, as shown by a significant decrease in judge accuracy after 2 rounds. Interestingly, we get the best LLM judge accuracy after only 1 or 2 rounds. We used 3 rounds for all experiments in this work because, qualitatively, it was the minimum we needed for a reasonable amount of interaction between debaters. For instance, they can critique their opponent in round 2 and defend in round 3. When considering interactive protocols, we decided it was appropriate to have two human judge interactions, which is possible in a 3 round debate. Furthermore, Michael et al. (2023) find that human judges have a failure mode of ending the debate or consultancy too early, which we wanted to avoid.

Because more rounds do not increase judge accuracy (Figure 21), it suggests the incorrect debater becomes more persuasive to LLM judges as rounds continue. As discussed in Appendix B.1, we find a common source of error was poor quote selection from the correct debater, and we believe that improving the recall of the best quotes for the correct answer would mitigate the judge accuracy decreasing with a number of rounds.

For consultancy, on the other hand, the accuracy increases as rounds continue. We hypothesise that this is due to the judge becoming less sycophantic towards the consultant as the interactive judge asks more questions. This behaviour can be clearly seen in Figure 21 where the incorrect and correct consultant both become less persuasive to the judge as rounds continue.

A.9 Improving the debaters with fine-tuning and few-shot learning

Fine-tuning helps but not significantly due to context length limitations. — Using the methodology in Appendix D.9, we fine-tune GPT-3.5-Turbo on arguments provided in the human-human debate dataset (Michael et al., 2023). This method involves truncating the Gutenberg story to fit in the context window, and running cross-play matches against Claude 2.0 on Slate questions in the QuALITY dataset to find the optimal model. An example Slate questions is “Why does the author tell a story about his vehicle?” and the correct answer choice is “To make a point about what has the most impact on the economy”. Unlike Gutenberg stories, slate questions were used since they fit into the context window.

Figure 22 shows that the fine-tuned model gets an Elo that sits between GPT-3.5-Turbo and Claude 2.0. This shows fine-tuning helps but does not boost debater skill as much as we would have liked. In future work, we’re excited to apply fine-tuning to stronger base models with longer context windows, allowing us to run debates on Gutenberg stories and fine-tune using stories that are not truncated.

Few shot learning does very little to increase debater skill. — We construct few-shot samples by adding the transcript between tags. To indicate which behaviour to imitate, we mark the name of the debater who wins in the transcript as DebaterWinner and the other as DebaterLoser. These example transcripts are added at the start of the first user message in the debater prompt found in G.1. We balance examples of incorrect and correct debaters winning and also the swap position.

Next, we ran a Swiss tournament with 4 Claude 2.1 debaters (vanilla, 8-shot, 32-shot and best-of-8). The aggregate Elo indicates that few-shot learning does little to help debater skill, and increasing to 32 samples leads to degradation in skill compared to the vanilla baseline (see plot on the right of Figure 22). Furthermore, best-of-8 significantly outperforms few-shot learning, so we did not use few-shot strategies in our larger Swiss tournament presented in the main paper.

A.10 Analysis of different preference models

Using a different preference model than the judge decreases accuracy with best-of-NN. — To understand the impact of the preference model on accuracy, we ran the same set of best-of-NN and critique-and-refinement strategies used in the cross-play tournament (listed in Appendix 2) but using GPT-4-Base. As shown in Figure 23, the self-play accuracy with a GPT-4-Turbo judge decreases. This can be explained due to the different preferences of these two models. The ranking of the arguments generated with best-of-NN will persuade the two models differently. Therefore, we match the preference and judge models to maximise debate accuracy. Having a disjoint preference model and judge does not invalidate the claim made that judge accuracy increases with debater strength because Figure 5 shows increasing linear trends for 3 different judges.

The preference model’s selection probability for the incorrect debater decreases with rounds. — Figure 24 indicates that it becomes harder for the incorrect debater to convince the GPT-4-Turbo judge as the rounds go on and the opposite for the correct debater. The variance in the probabilities is much higher for the incorrect debater, too, as shown by the absolute difference between arguments rated the highest and the lowest in the best-of-16 debater. The larger variance suggests there is a variety of poor answers that the incorrect debater can use, and best-of-NN does the job of discarding the unpersuasive ones.

GPT-4-Base has a big shift in selection probability distribution between the correct and incorrect debaters. From Figure 24, you can observe a big difference in distribution between GPT-4-Turbo and GPT-4-Base. GPT-4-Base has a big positional bias, which may partially explain the shift in the distribution between correct and incorrect.

Appendix B Further Results with Human Judges

Human judges chose the wrong answer 161 times in our test set, out of 953 total annotations. Judges chose the wrong answer in 22% of consultancies and 12% of debates. We investigate each of these judgements and label them based on the errors which contributed to the wrong answer. Each incorrect judgement can be caused by multiple errors. To perform this investigation, we read the transcripts, judge explanations, and relevant portions of the underlying texts.

Errors in Debate — Most incorrect judgements in debate were attributable to poor performance by the correct debater, primarily due to poor quote selection. Often the story would contain a quote that either clearly demonstrated the correct answer, or clearly disproved a claim made by the incorrect debater, and the correct debater would fail to choose this quote. Another common problem was repeating quotes that had already been used in earlier rounds, rather than presenting new evidence.

Errors in Consultancy — By far the most common error in consultancy is an inability of the judge to extract key evidence from the incorrect consultant. Often, the incorrect consultant was able to strategically quote from the text to present evidence that seemed to plausibly support its’ answer, while avoiding important quotes that would prove it wrong. The judge attempts to elicit this contrary evidence with questions, but is often unsuccessful. It is then very difficult for the judge to know if no contrary evidence is shown because they are dealing with an honest consultant and no such evidence exists, or because they are dealing with a dishonest consultant who is hiding it.

B.2 Judge Variation

We used 21 human judges, who were chosen based on their inter-annotator agreement and quality of their explanations (see Appendix E.2). Figure 25 shows the variance in accuracy, confidence and calibration between judges in consultancy and debate. Calibration is measured with a Brier score. Each judge performed an approximately equal number of annotations in each protocol. In consultancy protocols, answer assignment was randomised per-annotation. This means there is no guarantee that each judge received an equal number of correct and incorrect consultancies. For this reason, we calculate the combined consultancy accuracy for each judge as 12(correct_mean_accuracy+incorrect_mean_accuracy)\frac{1}{2}(\text{correct\_mean\_accuracy}+\text{incorrect\_mean\_accuracy}). We combine static and interactive protocols for this analysis, as we do not observe significant differences between them. Debate has a lower variance in accuracy and calibration compared to consultancy. This is due to the very high variance in incorrect consultancy. Correct consultancy is quite similar to debate. Judges were also slightly more confident in their answers when judging debates.

B.3 Question Variation

We use 47 questions for the human test set. Question selection is described in Appendix D.1. Each question received an approximately equal number of annotations in each protocol. Figure 26 shows the variance in accuracy, confidence and judge calibration between questions in consultancy and debate. As with judges, the ratio of correct to incorrect consultancies per-question is not guaranteed to be equal. We combine static and interactive protocols for this analysis. The variance difference between debate and consultancy is less pronounced for questions than for judges. Correct consultancies are still much lower variance than incorrect consultancies.

B.4 Human Judge Ensembles

We explore a variety of ensemble methods in an attempt to boost the accuracy of our oversight protocols. In our human experiments, each question was answered by several different judges under each protocol. We report mean accuracy in Figure 6, but here we consider the accuracy of different ensemble methods if the judgements for a question are pooled together for each protocol.

We consider the following ensemble methods:

Most confident — From NN judgements, choose the one where the judge gave the highest confidence. If several are tied for most confident, choose randomly among them.

Strict majority — Use the answer that the majority of judges chose. If there is a tie, choose randomly.

Confidence-weighted majority — Sum the confidences assigned to each answer and choose the answer with the highest aggregate confidence.

Squared confidence-weighted majority — Similar to confidence-weighted majority, but square the confidences first. This gives much more weight to highly confident judgements.

For each ensemble method, we are interested in seeing how accuracy changes as more annotations are used. Given that we have already collected NN judgements for a particular question and protocol, we calculate the accuracy of an ensemble function ff for k<=Nk<=N annotations as the mean accuracy when ff is applied to all combinations CiC_{i} of NN choose kk.

The number of annotations collected for each question/protocol combination varies from 4 to 8. In cases where k>Nk>N, we use the value for f(N)f(N) instead.

Figure 27 shows that all ensemble methods increase accuracy for all protocols. The confidence-weighted protocols are the most successful, boosting debate accuracy from 88% to 95-98%.

An open question is whether the benefits of ensembling are mainly due to variations in judge performance (i.e. reducing the impact of judge mistakes), or due to variations in debater skill (i.e. reducing the impact of honest debater mistakes).

B.5 Human Judge Iteration

Before our main human experiment on the DHD_{H} test set, we spent five weeks iterating on our experimental setup with humans, conducting seven smaller experiments. The overall goal of this iteration phase was to validate and refine our human experiment protocol, and to validate new debater models that seemed promising in our LLM-judged experiments. The experimental variations are described in the list below, and results are shown in Figure 28. Note that each experiment uses a different set of questions, because human judges cannot be given the same question multiple times. All experiments use the question selection criteria described in Appendix D.1, but there may still be variance in the difficulty of different question batches. Figure 29 shows the same plots as Figure 6, but using combined results from the test set and all iteration experiments.

Human judge iteration experiments - experiments were run between Nov 9th 2023 and Jan 26th 2024 (start dates given below):

Experiment 1 (Nov 9th) — 32 questions, 567 judgements. Static and interactive protocols. GPT-4-32k models, no inference-time optimisation (boNN or cNN). Motivation: Initial test of our experimental setup with humans.

Experiment 2 (Nov 15th) — 16 questions, 152 judgements. Static only. Changed model to GPT-4-Turbo. Using bo4 with a GPT-3.5-Turbo-Instruct preference model. Motivation: Test the impact of GPT-4-Turbo and inference-time optimisation in humans.

Experiment 3 (Nov 16th) — 16 questions, 170 judgements. Static only. bo4. Using a new prompt that encourages more conflict between debaters. Motivation: Evaluate new prompt variant.

Experiment 4 (Nov 24th) — 16 questions, 336 judgements. Static only. Only debates, no consultancies. First test of Critique & Refinement in humans. bo1, c8. Motivation: Test Critique and Refinement in humans.

Experiment 5 (Nov 30th) — 32 questions, 578 judgements. Static only. bo4. Change preference model to GPT-4-base. New prompts. Two experimental conditions - short and long. Short enforces 80-110 words per argument. Long enforces 300-350 words per argument. Motivation: Test the effect of argument length in humans (we had not previously enforced any word limits).

Experiment 6 (Dec 7th) — 16 questions, 318 judgements. Static only. Controlling words per round for the first time by giving consultants twice the word count of debaters. Debater arguments limited to 150 words, consultants limited to 300. bo4, c8. New prompt for debaters. Motivation: Test the effect of controlling for transcript length between protocols - rule out hypothesis that debate wins because transcripts are longer.

Experiment 7 (Dec 12th) — 16 questions, 241 judgements. Interactive only. New consultant prompt. bo16 for consultants, debaters use bo4 c6. Motivation: Test interactive protocols with our optimised models.

Experiment 8 (Dec 20th) — 47 questions, 953 judgements. Static and interactive protocols. bo16 for all models. Changed preference model to GPT-4-Turbo. Motivation: Get statistically significant results across all protocols using our most optimised models.

Experiment 9 (Jan 10th) — 24 questions, 480 judgements. Naive judge baseline in which humans were not provided with any transcript.

Experiment 10 (Jan 12th) — 32 questions, 519 judgements. Low Elo experiment using Claude-2.1 and GPT-3.5 models.

B.6 Low Elo debaters

We investigated how low-Elo models perform when judged by humans (Figure 30). Our test set experiments used GPT-4-Turbo-bo16, which had an Elo rating of 141 across all of our LLM judges. For the low-Elo experiment, we use Claude-2.1-bo4 (Elo 79) and GPT-3.5-bo16 (Elo -60). Elo ratings use Claude-2.1-bo1 as the 0-point. In all cases we use GPT-4-Turbo as the preference model, and the intermediary judge in static consultancy uses the same model as the consultant.

We see from the consultancy results that both low-Elo models are significantly less persuasive to humans than GPT-4-Turbo-bo16. We also see that this leads to significantly worse performance in debate (compare with Figure 6). Interestingly, the results from both low-Elo models are very similar, despite the Elo gap between them being larger than the gap between Claude-2.1-bo4 and GPT-4-Turbo-bo16. This suggests that the Elo rankings for human judges are different than those for LLM judges, indicating that humans and LLMs differ in what they are persuaded by.

B.7 Human judge feedback

After Experiment 4, we collected annotator feedback via a Google form. 16 out of 21 annotators completed it. Key insights are given here, and the full data will be released alongside the paper. The full data has numerous responses to qualitative questions about the protocols, efficacy of the models, and judging strategies.

75% of judges said the found it easier to judge debates than consultancies, while 19% said they were the same.

56% said that it is easier to judge longer debate arguments, while 38% said they’re the same.

No judges said that they usually decide their answer after the first round.

44% said it would be useful if debates had more than 3 rounds.

81% said that they mostly rely on quotes when making their decision.

Appendix C Recommendations for working on debate

This section explains three different task setups for debate where information or capability differ between the debaters and judges. With current model capabilities up to GPT-4, we then motivate why information asymmetry and the QuALITY dataset is a good test bed for debate. We then provide iteration and implementation recommendations to others who want to work on debate in the future.

We considered three different regimes for debate:

Capability asymmetric — the debaters are stronger than the judge with respect to base model capability. For example, using GPT-4-Turbo debaters and a GPT-3.5-Turbo judge. Here, the baseline to compare to is the GPT-3.5-Turbo judge’s few-shot or chain of thought performance.

Information asymmetric — the debaters have access to privileged information about the question, which the judges do not. For example, all agents are GPT-4-Turbo, but the judge cannot access the necessary text to answer the question. The baseline to compare to in this setting is consultancy, where the judge engages with only one copy of the model.

Symmetric — the debaters and judges have equal capability and information. In practice, this means all agents in the debate game use the same base LLM, such as GPT-4-Turbo. The baseline to compare here is the accuracy of GPT-4-Turbo by using few-shot examples or chain-of-thought reasoning.

We call a protocol an inference time protocol if uses models which are not trained or fine-tuned for the task. Augmentation occurs when the protocol runs in deployment using methods such as few-shot learning, chain of thought or best-of-NN (where the most persuasive argument is chosen from many completions by a preference model). Protocols may be used for self-improvement if multiple copies of the same model are used to improve accuracy on a task. An example is self-refine (Madaan et al., 2023), where models give feedback on completions before further refining their answers.

We discuss in the next section why debate does not perform well for self-improvement at inference time in the symmetric regime for current models.

C.2 Task setup findings

Inference-time debate is ineffective for self-improvement in symmetric regimes with current LLMs. In our prior research, we examined the efficacy of debate as a self-improvement protocol in symmetric regimes. This approach involved leveraging debate in an attempt to enhance an LLM’s reasoning and factual accuracy. We used GPT-4 to argue between two potential answers to various questions in datasets such as MMLU (Hendrycks et al., 2020) before scoring the accuracy of a GPT-4 judge that deliberated between debater arguments. Our findings indicate that models using debate produce lower test accuracy compared to 0-shot or chain-of-thought approaches, as evidenced by the accuracy observed across 10 tasks in Figure 31. Notably, only in TruthfulQA (Lin et al., 2022) does Debate outperform zero-shot, 5-shot, and chain of thought methodologies.

We hypothesised that debate might perform better on datasets where the judge model has low 0-shot accuracy (all of our prior datasets had a 0-shot accuracy above 70%). GPQA (Rein et al., 2023) is a benchmark written by biology, physics, and chemistry domain experts. It has questions that are hard to find the answer to on the internet, and GPT-4 has an accuracy of less than 40% with few-shot chain of thought. Anecdotally, when applying debate we found debaters lacked the knowledge to build a persuasive argument, thereby not assisting the judge in choosing the correct answer.

In conclusion, debate in our inference time setup does not lead to truth-seeking behaviour in situations where the judge has the same information and capability as the debaters. However, we think that work on self-improvement with debate is still promising, especially when training models (which we did not do) and using harder datasets.

Inference-time debate is ineffective for regimes with capability asymmetry for model capabilities up to Claude-2. In a sandwiching setup, we use a less capable judge, Claude 2, with more proficient debaters like GPT-4. Capability asymmetry arises due to the different reasoning capabilities of the models. Despite this asymmetry, we found no improvement to judge’s performance when running our debate protocol. For instance, Claude’s performance on tasks such as MMLU did not surpass zero-shot despite access to arguments by duelling GPT-4 debaters.

Our hypothesis was that the weaker judge’s deliberation ability was the bottleneck for getting this setting to work. Once the next generation of models is released, the sandwiching setup is worth revisiting, especially on difficult domains such as GPQA.

Inference-time debate is effective at increasing judge accuracy in a task that has information asymmetry. Our findings indicate that information asymmetry is important when testing the effectiveness of debate with LLM capability up to GPT-4-Turbo. Information asymmetry forces the judge to rely on the debaters for problem resolution. In the QuALITY setup, the asymmetry stems from the judge’s lack of access to the story necessary for answering the reading comprehension question. By contrast, in knowledge-symmetric scenarios, the judge can draw upon their own expertise to address the problem and disregard the debate transcript entirely. Another advantage of information asymmetry over capability asymmetry is being able to use the most capable models as the judge. At the time of writing, this is GPT-4-Turbo, and we find the improved deliberation capability crucial for reliably iterating on debate protocols since the aggregate ELOs are spread out over a larger range (as shown by Figure 25).

C.3 Dataset recommendations

We recommend QuALITY as a dataset to test debate. The QuALITY dataset is an excellent resource for testing debate, primarily due to the ability to establish information asymmetry and the rarity of self-defeating behaviour in LLM debaters. The latter is largely attributed to its questions, which are inherently easier to argue for either side due to being able to select quotes from lengthy narratives.

QuALITY’s stories average 7k tokens in length, and are used to create the asymmetry between the debaters and judges. Each story provides plenty of material for debaters to analyse and discuss over three debate rounds. This breadth of context enables the strategic use of quotes, potentially in a deceptive manner.

Extensive metadata accompanying each question is available, such as the proportion of dataset annotators agreeing that the question is answerable. Filtering for high answerability ensures the reliability of the question set for debate experiments, safeguarding against debates on unresolvable topics.

An option to use a highly plausible but incorrect answer option known as the “best distractor”. Using the best distractor reduces the likelihood that the debater assigned to that answer uses self-defeating behaviour (such as conceding to the opposing side). One reason for the reduced likelihood is that there are many quotes can make it seem like the best distractor answer is true, so we avoid the issue of RLHF trained models refusing to argue for a falsehood.

Many other datasets such as MMLU and AQuA (Goswami et al., 2023) are less suitable to test debate with current models because it can be very challenging to craft a proof for the incorrect answer in a persuasive way. If unable to formulate a convincing argument, the LLM debater might often concede or acknowledge subtle errors introduced in their solution to align with the assigned answer. Self-defeating behaviour is not a surprise, given that RLHF training promotes honesty. RLHF training moderately promotes answers that look good to humans which can often be linked with correctness, but in debate, we need to generate plausible arguments for incorrect answers. Our attempts to mitigate self-defeating behaviour via prompt engineering were unsuccessful. We believe helpful-only models would fare better as debaters in these datasets. We fine-tuned GPT-3.5-Turbo for debate, but found that it could not produce compelling debates or adhere to intricate prompts. We’re excited for future work fine-tuning GPT-4-Turbo for debate.

C.4 Iteration recommendations

Optimise for Elo rating rather than judge accuracy. Ultimately, we want debate to improve the judge accuracy on a task. However, it is important not to optimise for this directly when improving debaters. In competitive debate settings, rewarding debaters for persuasiveness to the judge is crucial for developing the skills required to win, making Elo a great metric. The calculated Elo should consider the aggregate win rate when the debater gets a chance to argue for both sides of the argument. Therefore, the debater with the best Elo will be the most persuasive when arguing for correct and incorrect answers. Elo also has the advantage of being an unsupervised metric. An example of where optimising for accuracy can go wrong is when debaters have self-defeating behaviour. Incorrect debaters are much more likely to concede, leading to inflated accuracy since the judge can easily choose the correct answer.

Test prompt improvements against each base model family. Testing each prompt is tedious, but we found small prompt changes can lead to unexpected behaviour. For instance, we reworded a small portion of our consultant prompt, and an extra 3% of the total questions became affected by concessions. Furthermore, it is important to test with multiple LLMs since we found that some prompt changes did not transfer from the family of Claude models to GPT models.

Try out interactive judging yourself while iterating on protocols. We used the same annotation platform (see Appendix E.1) provided in our human trial as a tool to read transcripts and judge debate questions interactively. We learned a lot about our debaters’ argument quality when judging questions. Acting as the judge without knowing the answer allows you to learn how persuasive the incorrect debater is compared to the correct one. In addition, it also allows you to find self-defeating behaviour and subsequently change the prompt to stop it from happening. Duplicate quoting, lack of engagement with the opponent or interactive judge and inability to use quote tags correctly were other examples of failure modes we quickly picked up on by reading lots of transcripts.

C.5 Implementation recommendations

Use a Swiss-style tournament when calculating Elos for multiple debaters. Running cross-play tournaments is compute-intensive, especially if running a round-robin tournament where each debater variant plays each other. Implementing a Swiss-style tournament was vital for fast iteration and reducing costs. A Swiss tournament assumes you start with an initial seed ranking of your players, then make pairings based on the nearest neighbour in the list. This list is then sorted after each round based on the total number of matchups won. This means it has O(nlog⁡n)O(n\log n) complexity rather than O(N2)O(N^{2}). A helpful framing is that the tournament implements an adaptive sorting of the initial list to ultimately find the true final ranking of the players. Once the tournament is complete, we use the fine-grain win rates over the question set the debaters used for each match-up and use the method in Appendix D.5 to find Elo ratings for each player.

Judge debates with both answer permutations to mitigate LLM positional bias Positional bias is evident when LLM judges achieve different accuracy levels if answers A and B are swapped. It is especially common in less capable models such as GPT-3.5-Turbo. This bias is less pronounced but still present in more sophisticated models like GPT-4-Turbo. To mitigate this, we encourage running judgements with both permutations of answer position for every debate transcript (see Table 7).

Apply strict word limit filters or truncation to mitigate verbosity bias in LLM judges. Verbosity bias is when the judge favours longer arguments, which we found to be the case with a significant rise in Elo rating as the argument word count increased. Decoupling this bias and argument quality is difficult, and we found qualitatively that this preference can falsely attribute skill to lengthier arguments. To mitigate this, we recommend using a strict word limit with rejection sampling as in Appendix D.6.

LLMs are poor at deliberation, so use at least GPT-4-Turbo capability for a judging model. Through our experience of protocol iteration, we often found confusing results after judging with GPT-3.5-Turbo and Claude, particularly with a smaller sample size. Weaker judges are more prone to being misled by trivial aspects of one side of an argument. GPT-4’s superiority as a judge is evidenced by a notably wider performance gap recovered, as shown in Figure 9, and it also has a larger area under the curve in Figure 5.

Use the same prompts for both the correct and incorrect debater. It might be tempting to use a different prompt for the incorrect debater, perhaps explaining deceptive strategies to improve performance. There are several reasons why this is ill-advised:

Admitting to the model that it is arguing the incorrect answer or it should use deceptive strategies is more likely to lead to self-defeating behaviour in current RLHF-based models, thus invalidating the debate. If a model gives up every time on the incorrect side, the judge accuracy is 100%, but the protocol is useless when you don’t have labels and the model does not know the correct answer.

The behaviour between debaters could be subtly different, leading to information leakage to the judge.

In a reinforcement learning setup where debaters are rewarded for convincing the judge, the copies of the model should operate with identical prompts.

At inference time, you may not have labels for the correctness of the answer choices, so you would not know which answer to assign to which debater.

Appendix D Implementation Details

Dataset — QuALITY (Pang et al., 2022) is a multiple-choice Q&A dataset for long-document comprehension. It contains documents from a number of sources, such as Slate articles or project Gutenberg short stories. Each document has a set of comprehension questions (with 4 possible answers) written by crowdworkers who have first read the entire document. Different sets of crowdworkers then attempt to answer the questions under 2 possible conditions: 1) Untimed annotators can take as much time as they want to read the document when answering the question; 2) Speed annotators are only allowed to read the document for 45 seconds before answering. The annotators also provide feedback on each question, giving ratings on 1) if the question is answerable and unambiguous 2) How much context from the passage is required to answer 3) Which unchosen answer is the best ”distractor” answer (question writers were encouraged to write difficult distractor answers).

Question Selection — We use only questions from the project Gutenberg short science fiction stories to ensure that judges have no prior real-world knowledge to influence their answers. Most of the stories are from the 1950s-1970s, making it unlikely that our human judges have read them. We wanted to select questions that were difficult enough to generate non-trivial debates, while still having clear and unambiguous correct answers. To this end, we applied the following filtering to the QuALITY dataset:

100% of untimed annotators chose the correct answer

Less than 50% of timed annotators chose the correct answer

All untimed annotators agree that the question is answerable and unambiguous

Average ”context required” rating from untimed annotators is at least 1.5

Writer label matches the gold label (the answer voted as the correct answer by annotators matches what the question writer labelled as the correct answer)

This is essentially a stricter subset of the HARD QuALITY questions.

QuALITY questions have 4 answers, however for our debates we used only 2 answers. For each question, we used the correct answer and the best ”distractor” answer. For the human experiments, we manually reviewed each question and pair of answers to ensure they made sense. We removed questions that were incompatible with our 2-answer requirement, e.g. questions where one answer was ”all of the above”, ”none of the above”, etc.

Human vs LLM judge considerations — For our LLM judges, we could use the same set of questions for each experiment. However, since humans remember past experiments, unlike LLM judges, we had to apply the following additional criteria when selecting questions for each human experiment: 1) This question has not been seen by human judges before 2) No other question in this story has been seen by human judges before (to avoid leaking any prior knowledge). This left us with a total of 200 usable questions for the human judge experiments.

D.2 Quote Verification System

Having access to the text, unlike the judge, the debater’s best strategy for winning a debate is to choose compelling evidence from the story that supports their assigned answer. Debaters are told in their prompt that any quotes they use should be wrapped in XML tags. Here is a snippet from their prompt (see Appendix G.1) explaining the quote system to the debater.

Our quote verification system is a simple direct string match after normalisation has been applied to both the quote and the story text. Normalisation involves stripping punctuation and making the text lowercase.

There are three possible cases after a debater uses a quote:

If it is a direct string match, then the quote is wrapped in v_quote XML tags to indicate it has been verified.

If it fails the string match, then it is wrapped in u_quote XML tags to indicate it is unverified.

If the debater does not put quote tags around the quote, then it will not be processed by the quote verification system and will remain as is.

D.3 Optimising Debaters

Best-of-N — At a high level, best-of-NN (boNN) is a mechanism that allows an LLM to generate many completions, followed by having a preference model choose one response that does well on a specific set of criteria. The preference model used is often another LLM prompted to assign a reward based on the criteria provided. The temperature of the LLM can be raised as the number of completions increases to encourage diverse samples.

In our case, we use boNN to select the most persuasive argument a debater could use during the debate round, according to a judge. See how boNN works alongside critique-and-refinement in Algorithm 1.

The first stage of boNN is to generate NN completions from the debater in the current round of the debate, followed by extracting the argument in tags (since we do not want the judge preference model to see the hidden scratchpad). Next, we apply the quote verification system as explained in Appendix D.2.

The second stage is getting a rating from the preference model (which does not have access to the the story like the judge) on how persuasive the argument is. As we use simultaneous debate, the correct and incorrect debaters concurrently generate responses for a given round. Therefore, we provide a dummy opponent response when appending the candidate arguments to the transcript. We set this to "My answer is the best choice, and my opponent is wrong.".

The preference model has a similar prompt to the judge but specific instructions to output A and B only (see Appendix G.4). We sample a single token from GPT-4-Turbo and extract the top 5 log probabilities to provide a score for tokens A and B. If the debater is the correct debater and defending assignment A, then the score used will be the log probability for A. On the rare occasion that the log probability for A or B is not present in the top 5, we assign it a value of -100, which will reject that sample compared to others.

Once all N arguments are rated, the one with the greatest log probability is chosen to use in the transcript, while the others are discarded. See an example in Table 3 that shows how the judge sees a potential argument alongside the dummy answer and the log probability given for A being chosen by the preference model judge.

We found that using the same preference model as the judge model led to the largest gains in aggregate Elo and judge accuracy (as shown in Figure 23).

If using a completion model instead of chat, it can be helpful to append an assistant message such as “Assistant: Based on the provided arguments, the correct answer is most likely to be (”. This trick ensures that the model is biased towards giving logits for A or B.

Critique and Refinement Critique-and-refinement is a technique to improve the initial output of an LLM on a specific task. Other work, such as Self-Refine (Madaan et al., 2023), uses iterative feedback and refinement steps using the same model to improve its initial output. This can be extended by generating many critiques and choosing the most helpful one (Saunders et al., 2022).

In our debate protocol, we ask a critic model that has access to the story to produce NCN_{C} critiques of the argument after it is extracted and processed by the quote verification system. A preference model then gives a rating to each of the critiques based on which is most helpful for the debater to improve their argument. Helpful is described as being contextually relevant and directly applicable to strengthening the debater’s assigned answer. The critique preference model is asked on the final line of the prompt if the critique is helpful Y/N, and the critique rating is assigned to the log probability corresponding to “Y”. Only a single token needs to be sampled.

In a follow-up message to their original response, the debater is provided with the critique with the highest log probability from the preference model and asked to refine their argument. See prompts in Appendix G.5. The critic and preference model prompts differ slightly between debate and consultancy, but the refinement prompt is identical.

If best-of-NN sampling is also being used, then the NAN_{A} refinements will be generated and chosen by the judge preference model. When NC=1N_{C}=1, this protocol is equivalent to one turn of self-refine. This is illustrated in Algorithm 1.

Common pitfalls in the arguments generated by the LLM debaters include the repetition of quotes, the use of unverified quotes, inadvertently supporting the opposition’s stance, and overlooking quotes that could bolster the argument. The critic addresses these issues by providing targeted, constructive feedback. This feedback is specifically designed to identify and rectify weaknesses in the arguments, such as pointing out where quotes may be repetitive or unverified or where they inadvertently lend credence to the opposing side. A crucial aspect of the critic’s role is ensuring that quotes are direct string matches and are properly encapsulated within tags. By highlighting these areas and suggesting more persuasive and accurately sourced quotes, the critic plays a pivotal role in refining the debater’s presentation. Ensuring it is manageable for the debater to implement the feedback within their assigned word limit, the critique length is limited to 150 words and two bullet points (and double that for consultancy). We found when testing with five bullet points that debaters would implement everything in their refinement and stop listening to the word limit provided.

Concessions occur more often when using critique-and-refinement since the critic can provide tough or impossible feedback for the incorrect debater to implement. Lots of iteration was needed to reduce the frequency of concessions, especially in Consultancy, where minor edits sometimes lead to a 2-3% increase of the total questions being impacted. We found the following to help reduce concessions:

Write the answer they should be supporting again in the refinement prompt message.

Tell the debater they don’t have to include all the items in the critique if doing so would compromise the argument.

Tell the critic they should provide constructive criticism to help the debater continue arguing for their assigned side. Without supporting quotes, the critic is told to help the debater construct a strategic argument without relying on quotes.

Other issues included ensuring the debater did not mention that it had implemented feedback from a critique. We prompted the model not to include the phrase “The critique suggests”. Empirically, we noticed that the incorrect debater was more likely to mention the critique, so we decided to also use rejection sampling if ”critique” was mentioned in the response. Occasionally, Claude 2.1 refuses to refine the argument for ethical reasons. When this occurred, we used the original argument before refinement. An area for improvement is ensuring the critic does not misquote the story when suggesting quotes.

D.4 Cross-Play Tournament

Tournament Details We use a Swiss-style tournament system for our cross-play tournament, allowing us to evaluate a large number of players (NN) efficiently. Running a round-robin tournament where each player plays everyone would require O(N2)O(N^{2}) matches. The Swiss tournament has O(nlog⁡n)O(n\log n) complexity, allowing each player ample opportunities to compete against similarly ranked players, leading to accurate final rankings.

Players should be seeded initially so the best pairings happen between similarly matched players in early rounds. The number of rounds is determined by the ceiling of the log⁡2\log 2 of the total number of players, ensuring a balanced and manageable competition structure.

Pairings for each round are generated based on the closest neighbour that has not played yet. We avoid repeat matchups. In cases with an odd number of participants, a player who does not find a match is granted a bye. This process ensures equitable distribution of opponents that are closest in strength.

Each match is conducted, and the winner is determined based on the player’s number of wins in both correct and incorrect debater permutations. Following each match, scores are updated: winners receive a point, while losers do not. This scoring system allows us to dynamically re-ranking players after each tournament round, reflecting their performance in the tournament.

After the tournament, participants are ranked based on their total score. We then use game history to calculate aggregate Elos based on the win rates in each match. The pseudo-code for the Swiss tournament is given in Algorithm 2.

We use 291 questions from the quality development set (denoted DLD_{L}) for all cross-play matches and use 400 questions from the train set (denoted TLT_{L} when running self-play matches to get judge accuracy. The filtering of these sets is equivalent to the method described in Appendix D.1, and we allow a maximum of 5 questions per story.

Models and Hyper-Parameters We use 20 players in our Swiss tournament and choose various best-of-NN and critique variations across four base models to get an extensive range of skill levels. The initial ranking for each player is given in Table 4 and corresponding hyperparameters in Table 5 and Table 6. The initial ranking was determined by preliminary experiments where we ran smaller round-robin tournaments between Claude and GPT models. Combining best-of-NN with critique-and-refinement in smaller round-robin tournaments boosted aggregate Elo more for Claude than GPT-4-Turbo models. It was assumed that all GPT-4-Turbo models would beat Claude variations, even though that was not true after running the tournament. For best-of-NN, we did temperature tuning using smaller tournaments of 4 players, finding 0.8 led to robust improvements across all NN.

We finished the tournament after 4 rounds rather than 5 because we found the resulting Elos were representative, and running a 5th round would not be a good use of API credits.

Comprehensive results for debater win rates across all 40 matches and with 3 judges are provided in Appendix F. The final ranking order for each judge is also provided.

D.5 Calculating Elo Ranking

Elo ratings, initially developed for chess, offer a simple yet robust method for estimating the relative skill levels of players in competitor-versus-competitor games. The algorithm we employ for Elo rating calculation is based on the assumption that the performance of each player is a normally distributed random variable. The expected score of a player is modelled as a logistic function of the difference in ratings between the player and their opponent.

Expected Win Rate: The expected win rate for a player P1P_{1} against another player P2P_{2} with Elo ratings E1E_{1} and E2E_{2}, respectively, is given by the formula:

This function yields a value between 0 and 1, representing the expected probability of P1P_{1} winning against P2P_{2}.

Cost Function for Elo Rating: The cost function used in Elo rating optimisation is the squared error between the predicted win rate, based on Elo ratings, and the actual win rate observed in historical data. Given a set of player pairs with observed win rates and match counts, the cost function is defined as:

where ωPi,Pj\omega_{P_{i},P_{j}} is the actual win rate of player PiP_{i} against player PjP_{j} in the dataset.

Optimisation: To compute the optimal Elo ratings for a given dataset, we estimate the first-order gradients and apply a simple optimisation technique. The player ratings are first initialised at the zero references of a specific policy (in most cases, we choose the Claude 2.1 model). We estimated ratings using Broyden–Fletcher–Goldfarb–Shanno (BFGS) algorithm. To estimate confidence intervals we use statistical bootstrapping over 1000 seeds.

D.6 Mitigating LLM bias

Strict word limit — We use a strict word limit with rejection sampling to mitigate the verbosity bias of LLM judges (see the impact in Appendix A.7. Our methodology is provided in Algorithm 3, which illustrates generating a pool of candidate responses from the LLM and subsequently filtering them based on criteria in an is_valid. We specify the LLM response should contain an tag, and the argument within those tags must be within the word limit boundaries in Table 6. We generate 3 candidates per completion required for boNN, which, empirically, we find maximises the chance of obtaining a sufficient number of valid responses.

When we don’t have enough responses after rejection sampling, we pad the valid responses with invalid responses. This strategy guarantees that the total number of responses aligns with the requirement of best-of-NN. On the occasions where pad the responses, we truncate the argument if they were over the specific word limit. Our truncation function carefully adjusts responses, particularly paying attention to not disrupt structural elements like quotes. If truncation slices in the middle of a quote, we add a closing tag at the end so a partial quote can still be sent for verification. If we truncate, we append ... to the end of the argument.

Swapping answers — In order to counteract the positional bias of the LLM judge, we maintained a balanced collection of transcripts with the answers swapped between positions A and B (see Table 7). Without this modification, the LLM judge’s accuracy might be artificially inflated, leading to misleading performance assessments, particularly when contrasted with human judges, who exhibit minimal positional bias as demonstrated in our experiments detailed in Appendix A.7.

In conducting all LLM experiments, we use several techniques to allow us to easily swap the arguments when running judging, without needing to run the debate with reversed positions. The prompts for our debaters are designed without any reference to positions A and B, nor do they include the opponent’s name, ensuring that the debaters never incorporate these identifiers in their arguments. If the debaters were to use names and reference the letter corresponding to the answer, swapping the transcript accurately would become more complex, especially when considering name-associated genders, as would be the case with names like Alice and Bob (e.g. his and her). It is worth noting that the prompts for interactive versions of our protocol include these identifiers since humans often refer to the debaters by name or by answer choice letter. In this case, the debate has to be run twice in both positions, which is what we do for human trials.

Furthermore, for an effective judge prompt, it was important to clarify which debater was arguing what. We clearly state that Debater A argues for answer A, and Debater B argues for answer B. The transcript uses the debaters’ names in each round, and Debater A always goes first.

D.7 Debater Prompt Iteration

We iteratively produce debater and consultant prompts using a fixed LLM judge. Importantly, a stronger consultant/debater prompt might have the same average accuracy compared to a weaker one. Thus, it is important to optimise the win rate for consultants and aggregate Elo for debaters. This ensures that we choose prompts that help our experts become more persuasive on both the correct and incorrect side.

A frequent failure mode with LLM debates and consultants is that models tend to concede, in which they acquiesce their assigned answer. Concessions happen at a much higher rate for the incorrect side, so improve protocol accuracy because they clarify to the judge which answers to choose. Thus, this is a reason why optimising for Elo is important. Since the LLMs we used are trained with RLHF, when debaters argue an incorrect answer, they often concede even when explicitly told not to due to their self-defeating bias, so careful prompt iteration is required.

We provide the same prompt (apart from the assigned argument to defend) to both debaters, independent of whether they are arguing the correct or incorrect answer. This is the same for consultancy. We experimented with having separate prompts, for example, giving the incorrect debater tips on how to be deceptive due to being given the incorrect answer. We did see better protocol accuracy; however, we discovered this was usually the case due to self-defeating behaviour where the debater would admit they were defending the incorrect answer. This is another reason why optimising Elo is better than accuracy because an increase in self-defeating behaviour would decrease the Elo. Furthermore, using identical prompts avoids information leakage to the judge and keeps the method applicable for test inference where you don’t know the labels.

Figure 32 provides an example ablation of debater prompts, where we ran a Swiss tournament with 4 debater prompts. The base prompt has the rubric (telling the debaters how the judge will score the transcript), the thinking steps and the extra assistant message removed. We then replace each of these progressively in the other prompts and show that each item improves debater Elo on Claude 2.1. The changes make less of an impact with GPT-4-Turbo, with the largest increase from adding the thinking stage.

We found that whilst iteration worked well by running Swiss tournaments with different prompts, it did not generalise well between different base LMs. For example, prompt iterations were originally tested with Claude 2, but sometimes, changes did not hold up for transfer to GPT-4-Turbo. We ended up iterating over both base models simultaneously in order to not overfit.

Debater Prompt Highlights — Here, we list key features of our prompts after iterating using the aforementioned method.

After the task description in the system prompt, the user provides the question and assigned answer. We found it helpful to reduce concessions by adding an assistant message echoing back that they understand and which answer they are assigned.

Include a detailed description of the task, quote system and judge criteria in the system prompt. We found it beneficial to do this in the system prompt and put question-related content in the user prompt.

Using a scratchpad within XML thinking tags is vital for several reasons. If done well, crafting the advice impacts the win rate significantly compared to other prompt changes.

We provide round-dependent thinking advice, which helps the debaters engage with each other in later rounds. The debaters are told to give an opening statement with a clear context for someone who has not read the story in round 1. They are told to critique their opponent’s argument in round 2 and provide a rebuttal in round 3. Without this specific advice, debaters often repeat their arguments or build on their side rather than pick apart their opponent.

Debaters are advised to list out potential quotes to use in their argument, which helps increase quote usage.

Consultants, it is crucial to get them to repeat back the side they are arguing for during their thinking stage and also state the importance of not conceding.

XML tags mark where parts of the prompt start and finish clearly. E.g. , , , , tags are used. We use and tags in the output to make parsing easy.

Stopping the debater using “In summary” was useful since, qualitatively, we found it to be a waste of the word limit when judging debates.

Reminders of structure and word limit are essential to repeat just before generation to get consistent argument lengths, even though it is in the system prompt.

D.8 Judge Prompt Iteration

For the project, we periodically iterated on judge prompts after improving debaters.

Without careful instructions, LLMs tend to believe whatever they are told, which would lead to poor accuracy if debaters decided to use fake quotes to strengthen their arguments. We found it very important to explain the quote system clearly and to find a prompt where the judge cares about whether quotes are verified or unverified.

Our first set of prompts would lead to a judge that believed a set of unverified quotes just as much as a set of verified ones (as shown by the consistent level of accuracy by the ‘Not strict on verified quotes’ in Figure 33). After changing the prompt to include more explicit instructions about the quote system and extra advice in the thinking stage (as shown in Appendix G.1), we see that the accuracy is high when the incorrect debater arguments are transformed into unverified quotes and low when the correct debater arguments are transformed. This trend shows that the judge down-weights the arguments as it should if they are unverified, meaning fake quotes or lousy use of the quote system will be penalised.

D.9 Fine-tuning debaters methodology

Fine-tuning our debaters presented a challenge, as the extensive length of the Gutenberg QuALITY stories far exceeded the context window capacity of GPT-3.5-Turbo in the OpenAI fine-tuning API. To address this, we truncate the stories and only keep context around the quotes used by the debaters in the human-human debate dataset (Michael et al., 2023).

To effectively truncate and split the story for inclusion in our dataset, we use a methodology that focuses on preserving the quotes used by the human debaters while discarding surrounding text. This approach is useful for shortening the story text to fit in a 4k context window without sacrificing the context necessary for understanding the debate. These are the steps we follow:

Quote Normalisation: Each quote is normalised by replacing specific punctuation marks and converting the text to lowercase. This standardisation aids in accurately comparing and detecting quotes within the story.

Sliding Window: We implement a sliding window to scan through the story. The window size and overlap are parameters, varied to optimise quote retention and story length to be in the range of 500 and 3200 tokens. We sweep window sizes ranging from 200 to 5 tokens, with overlaps between 20 to 2 tokens.

Quote Detection and Retention: Within each window, we check for the presence of normalised quotes. The entire window is marked for retention if a quote is detected (determined by an n-gram similarity threshold).

Combination: After processing the entire story, we compile the retained windows into a truncated story. This approach reduced the average story size from 7k to 2.8k tokens.

For each debate round, the prompt construction included the transcript history to that point, and the fine-tuning completion target was the next argument of the human debater. This process yielded 1063 samples for fine-tuning, selectively using arguments from the winning debaters.

Due to context window limitations, we tested the fine-tuned debaters on Slate questions in the QuALITY dataset rather than Gutenberg. These are news articles and are significantly shorter. We swept the number of epochs, finding that 4 epochs led to the best-performing model in terms of win rate against Claude 2.0.

D.10 LLM experiment dates

We used the following model ids in the time periods below: gpt-3.5-turbo, gpt-4-1106-preview, claude-2.1 and claude-v1.3.

Figure 1 — LLM results on the human dataset. Jan 10th 2024 for debate, naive and expert. Jan 26th 2024 for interactive debate and consultancy.

Figure 3 — debate under optimisation pressure. Dec 18th 2023 to Dec 29th 2023 for cross-play and self-play runs.

Figure 4 — consultancy under optimisation pressure. Jan 3rd 2024 to Jan 5th 2024 for all runs.

Figure 5 — different judges for debate tournament. Dec 30th 2023 to Jan 2nd 2024 for all judges.

Appendix — further LLM results. All runs between Jan 3rd 2024 and Jan 26th 2024.

Appendix E Human Trial Details

We built an annotation platform for annotators to submit judgements for their assigned debates.

Question assignment — Questions were grouped into experiments. Each experiment consisted of one or more protocols to be compared, where a protocol includes the method (debate vs consultancy), whether judge interaction is allowed, prompts, and parameters such as word limit, boNN number, temperature etc. Each experiment was assigned to all judges. For each judge, the mapping of questions ↔\leftrightarrow protocols within an experiment was varied with a Latin squares approach, and the sequence of questions was randomised. The debate or consultancy transcript was generated independently each time, so that even if two judges are assigned the same protocol for the same question, they will see unique transcripts. Judges moved through the experiments in sequence and answered questions within each Experiment in their assigned order. For each transcript, the order of answers is randomised. In consultancy protocols, the answer the consultant is assigned to argue for is also randomised. This results in a potential imbalance of correct/incorrect consultancies assigned to each judge, which makes it difficult to compare individual judge skill in consultancy, but the large number of questions in the whole experiment means the overall ratio of correct/incorrect consultants is close to 50/50. This approach was used to avoid potentially leaking information to judges - if we always split correct/incorrect 50/50 per judge, then a judge who has agreed with the first several consultants in an experiment should be biased towards disagreeing with the remaining consultants. To account for this potential imbalance, we calculate accuracy in consultancy as 12(mean_acc_correct+mean_acc_incorrect)\frac{1}{2}(\text{mean\_acc\_correct}+\text{mean\_acc\_incorrect}).

Judgement submission — The main interface of the annotation platform shows the judge their currently-assigned question and the transcript of the debate or consultancy that has been created for that question. In static protocols, the entire 3-round transcript is pre-generated, including any LLM judge interactions in consultancy. For interactive protocols, only the first round is generated, because subsequent rounds rely on the interactions provided by the human judge. Interactive protocols provide a text box for judges to enter their interactions before each subsequent round is generated. The time to generate an additional round is typically 30-60 seconds. After reading the transcript, judges submit their decision by assigning probabilities to answers A and B, indicating their confidence in each answer. The answer they assign the higher confidence to is considered their chosen answer. Confidences are assigned in 5% increments, with a maximum of 95%. Assigning 50% confidence is not possible. Judges must also provide a free-form text explanation of why they made their decision before submitting each judgement. Judges are not shown whether their chosen answer was correct (except during the training phase).

E.2 Human Judge Recruitment and Onboarding

Recruitment — We recruited 30 workers through the crowdworker platform Pareto. Our recruitment criteria were for the workers to be fluent English speakers with high general reasoning skills (assessed based on referrals and vetting by Pareto). All workers were paid $35/hr throughout training, testing, and experiments.

Training — We gave workers a 1-week training phase where they used the annotation platform to answer 24 questions, divided equally among our four protocols: static debate, static consultancy, interactive debate, and interactive consultancy. During training, workers could see the correct label after submitting each answer, allowing them to learn from their mistakes. All training questions were sourced from the Easy subset of Slate questions, while the main experiments would use Hard Gutenberg. We did this because we care about oversight methods that can generalise from easy labelled data to hard unlabelled data. For this reason, it was important that workers never see labels from the hard subset.

Filtering — After training, workers went through a 1-week filtering phase. This involved 24 questions each, split equally among the four protocols. Filtering questions were sourced from the Hard Gutenberg subset, and workers could not see the correct labels after answering.

At the end of the filtering phase, we reduced the pool of workers from 30 down to 21. As discussed above, we did not want to rely on using labels from the Hard subset during the judge selection process, so rather than assessing workers based on their accuracy, we used the following criteria:

Inter-rater agreement: the degree of alignment of each worker’s responses with the consensus opinion across questions.

Explanation quality: the collection of explanations submitted by each worker with their answers was read and assigned an overall high/medium/low-quality score based on writing quality.

We removed all workers who had inter-rater agreement scores in the bottom quartile, and all workers who had been rated ”low” on explanation quality. We found that these two metrics had significant agreement in their ranking of workers. After filtering, we checked the accuracy scores and found that accuracy and inter-rater agreement resulted in very similar worker rankings.

Calibration Training — During training, in addition to showing workers the correct labels for their submitted answers, we also showed them an overall score, which was calculated based on quadratic scoring. This was intended to encourage judges to give well-calibrated confidence because quadratic scoring is a proper scoring rule. However, in practice, judges found this score confusing, and many believed that the way to maximise the score was to always give a maximally confident answer to every question.

To address this issue, after the filtering phase, we removed the overall score from the interface and gave judges instructions that they should aim to give well-calibrated confidence scores. To help them do this, we had them all complete the Metaculus calibration training gamehttps://www.metaculus.com/tutorials/calibration/ before moving on to the main experiments. This training took 1-2 hours and was paid at the normal rate.

Appendix F Comprehensive Tournament Results

Appendix G Prompts

G.2 Judge Prompts

G.3 Additional Baselines

G.4 Best-of-N Prompts

G.5 Critique-and-refinement Prompts

Appendix H Samples

Here, we present a selection of transcripts from our human experiments. Verified quotes are highlighted in green, while unverified quotes are highlighted in yellow. The entire human dataset is released alongside this work.

This static debate is a good example of how QuALITY questions require careful reading but have clear correct answers. With just a single consultant, it would be easy to present convincing quotes for either side, but with two debaters the correct answer is very clear.

Question: Who is Teena and what role does she play in Chapter one and chapter two? A (correct): Teena is Eddie’s friend and neighbor. She accompanies him on a prospecting hike, where they don’t find any trace of radioactivity but still enjoy a lunch together. (Debater A) B: Teena is Eddie’s friend and neighbor. She accompanies Eddie on a hike through the hills behind the college, where he teaches her all about isotopes. (Debater B)

H.2 Successful correct static consultancy

In static consultancy we use an LLM judge to question the consultant. Question: How does the CPA prevent crime? A (correct): The CPA prevents crime using constant surveillance, subliminal messaging, public shaming, and various psychological treatments, including lobotomy. (Consultant) B: The CPA prevents crime with brain implants that suppress the criminal nature.

H.3 Successful incorrect interactive consultancy

In interactive consultancies the human judge asks the questions. Anecdotally, human judges tend to ask more pointed questions to quickly hone in on key weaknesses in the argument.

Question: Why is Conners upset with Bridges? A (correct): Conners received a report that Bridges was behaving unprofessionally. B: Conners was chewed out by a Senator because Bridges was trying to get information. (Consultant)

H.4 Failed debate - suboptimal evidence use

The most common cause of judges choosing the wrong answer in debate is for the correct debater to make poor choices in quote selection. In this debate, both debaters are arguing over who said which quote. The short word limit makes prevents quoting the entire relevant section of the story, but it should still be possible for the correct debater to strategically select the right portion to make a very convincing case. But in the transcript we see the correct debater make a poor quote choice and keep repeating it in later rounds.

”What’re you doing here on Earth?” I asked him. ”Researching for one of our meat trusts. We’re protein eaters. Humanoid flesh is considered quite a delicacy. How about you?” ”Scouting the place for thrill tourists. My job is to go around to these backward cultures and help stir up inter-tribal, or international, conflicts—all according to how advanced they are. Then our tourists come in—well shielded, of course—and get their kicks watching it.” Paul frowned. ”That sort of practice could spoil an awful lot of good meat.”

Question: What is Paul doing in Tangier? A (correct): He is on a business trip to find a source of protein. (Debater A) B: He is on a business trip scouting locations for thrill-seeking tourists. (Debater B)

H.5 Failed consultancy - cannot extract key evidence

Consultancy is difficult because if the consultant is arguing for the incorrect answer, they will simply not choose any quotes that support the correct answer. The judge can try to elicit this evidence, but it is difficult. In this story there are many possible quotes that would support answer A, the most obvious being ””We’ve been tripping for centuries, visiting every habitable star system from our home near Canopus. But Haron is too demanding.”.

Question: Why does no one on town know who Haron Gorka is? A (correct): He’s not a real resident, but an interstellar visitor. B: He’s not a real resident. He’s using a fake name while he stays in town. (Consultant)