Improving Factuality and Reasoning in Language Models through Multiagent Debate

Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, Igor Mordatch

Introduction

Large language models (LLMs) have demonstrated remarkable language generation, understanding, and few-shot learning capabilities in recent years. These methods are trained on a massive corpus of text on the internet, where the quality and accuracy of extracted natural language may not be ensured. Thus, current models may suffer from confidently hallucinating facts or making implausible jumps in chains of reasoning. An extensive body of recent work has focused on improving factual accuracy and reasoning in language models. These range from prompting models with few or zero-shot chain-of-thought demonstrations, use of verification, self-consistency, or intermediate scratchpads.

We note that these techniques are applied over a single model instance. Instead, we propose a complementary approach inspired by The Society of Mind and multi-agent settings, where multiple language model instances (or agents) individually propose and jointly debate their responses and reasoning processes to arrive at a single common answer. More specifically, given a query, multiple instances of a language model first generate individual candidate answers to a query. Then each individual model instance reads and critiques the responses of all other models and uses this content to update its own answer. This step is then repeated over several rounds. This process induces models to construct answers that are consistent with both their internal critic as well as sensible in light of the responses of other agents. The resulting quorum of models can hold and maintain multiple chains of reasoning and possible answers simultaneously before proposing the final answer.

We find that our debate approach outperforms single model baselines such as zero-shot chain of thought and reflection on a variety of six reasoning, factuality, and question-answering tasks. Using both multiple model agents and multiple rounds of debate are important to achieve the best performance. Given an initial query, we find that individual model instances propose a diverse range of answers despite being the same model class (although we also investigate the case of mixing different model types, such as chatGPT and Bard ). After debating and examining the responses of other model instances, we find that the population almost always converges on a single and more accurate common answer. Debate results are also less likely to include false facts that models are internally uncertain of. This is because as the debate progresses, individual model instances tend to disagree on uncertain facts and omit them from the answer (Figure 7). Lastly, we find that debate does not just act to amplify one correct answer in a model quorum - we find many cases where all the models initially make incorrect predictions, but then arrive at the correct answer as debate progresses (Figure 4,11).

We use the same methodology and prompt templates for all our tasks and require only black-box access to language model generations – no model-internal information such as likelihoods or gradients is needed. This allows our method to be used with common public models serving interfaces. The method is also orthogonal to other model generation improvements such as retrieval or prompt engineering (in fact, we combine our debate method with zero-shot chain of thought). While the debate process is more costly, requiring multiple model instances and rounds, it arrives at significantly improved answers and may be used to generate additional model training data, effectively creating a model self-improvement loop.

To help evaluate the effect of our approach on factual accuracy, we introduce a new benchmark and dataset evaluating factual accuracy of famous computer scientist biographies. We find that contemporary language models have an especially high tendency to hallucinate factually incorrect biographies, often misrepresenting the relevant institutions and dates. Moreover, these facts often inconsistent across different language model instances. By asking models to come to a consensus across their answers, such inconsistent facts may be either removed or corrected.

In summary, our work contributes the following. First, we present a novel approach to improving factual correctness and reasoning accuracy in contemporary language models, leveraging a multi-agent debate process between models. Second, we introduce a new benchmark of factual correctness which contemporary language models struggle with. Finally, we evaluate the performance of our debate procedure in language generation, both in terms of the number of agents, the underlying rounds of debate, and the prompts that elicit such behavior across a set of six different reasoning and factual accuracy tasks.

Language Generation through Multiagent Debate

We present an approach to generate language responses through multiagent debate. We provide an overview of our approach in Section 2.1. We further discuss convergence to consensus in the debate process in Section 2.2. The overall overview of our approach is shown in Figure 2.

Consider your work process when solving the following math question on an exam: “What is the area of a triangle with side lengths of 3, 4, 5?". In one thread of work, you may recognize that the triangle side-lengths directly correspond to a right triangle, and thus directly compute the area as 0.5×3×4=640.5\times 3\times 4=64. To make sure that you have the right answer, you may then try to solve the problem differently by estimating an angle θ\theta in the triangle using the Law of Cosines, and then obtain the area by using the formula 0.5×3×4×sin⁡(θ)0.5\times 3\times 4\times\sin(\theta), arriving at another answer to the given exam problem.

When these lines of work give the same answer, your confidence about the answer increases. In contrast, when these answers are different, individual lines of work may engage in a mental “debate" procedure, where you closely cross-examine the reasoning and assumptions of each line of work and refine solutions until a consistent answer.

Similarly, consider writing a biography of a historical figure. To ensure the factuality of the biography, you may consult multiple different sources on each fact. Facts that are consistent in each source increase your confidence about the fact. In contrast, facts that are inconsistent require careful cross-examination between sources to determine the final consistent data.

To mimic the above multi-threaded reasoning process and multi-source factuality checking processes, we propose to generate answers subject to a multi-agent debate procedure between multiple instances of large language models. Given a question, multiple agents represented as copies of a large language model, generate answers to the question. Each response serves as a possible thought process or source of information which agents may re-examine to find consistent final answers.

After initial responses are generated from different agents, we initiate a round of debate between agents. Individual responses from other agents are concatenated and given as context to each agent, with each agent instructed to construct a new response based on such responses. Each language agent is thus responsible for both verifying the collection of responses given by other agents, and refining its own response based on other agents’ responses. We iteratively repeat this debate procedure over multiple rounds for improved performance.

Concretely, we first prompt each agent to independently solve the given problem or task. After each agent generates a response, we feed each agent a consensus prompt, illustrated in Figure 3, where each agent is instructed to update their responses based on the responses of other agents. This resultant consensus prompt may then be repeatedly given, using the updated responses of each agent. We illustrate an overview of this multiagent debate procedure in Figure 2.

Note that our proposed approach operates in an orthogonal manner to existing approaches to prompt language models. Given a question, we may apply additional techniques for prompting language models to further improve our debate procedure by eliciting additional more detailed responses from language models. We illustrate the synergy of our approach with existing approaches to prompting language models in Figure 6 and directly apply zero-shot chain-of-thought reasoning in our evaluations.

2 Consensus in Debates

Given multiple rounds of debate, how can we ensure that a set of language model agents will converge to a final consensus answer? In general, debate can be seen as a multi-agent game, where convergence is not guaranteed. Empirically, however, we find that language models are able to converge on a single shared answer after multiple rounds of debate (Figure 4).

We found that we could control the duration of debates by how changing how much a language model trusts its own outputs over those generated by other models through different prompts. We illustrate two prompts below in Figure 3, which we use to induce different debate durations between language models, and illustrate the effect of such prompts in Figure 12. In general, we found that prompts that encouraged models to be more “stubborn’ based on their own solutions led to longer debates and better final solutions. Overall, we observed that language model agents were relatively "agreeable", perhaps as a result of instruction tuning or reinforcement learning based on human feedback .

Experiments

In our experiments, we evaluate our multiagent debate procedure and answer the following questions: (1) To what extent does multiagent debate improve reasoning? (2) To what extent does multiagent debate improve factual validity? (3) What design choices enable multiagent debate to improve language generation performance?

We first evaluate the extent to which multiagent debate improves the underlying reasoning process in language models.

We evaluate our approach on three reasoning tasks of increasing difficulty:

Arithmetic. We first evaluate the ability of models to correctly evaluate an arithmetic expression (containing addition, multiplication, and subtraction) consisting of six different two-digit numbers. For example: What is the result of 12+15*21+0-3*27?

GSM8K. Next, we consider harder mathematical reasoning tasks. Using the GSM8K dataset , the models must correctly solve grade school mathematical reasoning tasks.

Chess Move Prediction. Finally, we consider the strategic reasoning of the ability of models, and ask models to predict the best next move in a game of chess, given the first 14 moves of a chess game between two chess grand-masters described in PGN notation .

We report the accuracy of final answers in arithmetic and GSM8K tasks and report the pawn score (advantage) of predicted moves, as estimated by Stockfish in the Chess move prediction tasks. Additional details may be found in the Appendix.

Baselines.

We compare our approach to three alternative approaches to generate responses for reasoning problems. First, we ask agents to directly generate responses (single agent). Next, we consider asking language models to generate and then "self-reflect" on the responses generated . Finally, we consider generating responses using multiple agents and performing majority voting . As the focus of our experiments is to verify the effectiveness of multiagent agent debate, we run both baselines and our approach, using the identical starting prompt and language model across all evaluations. We evaluate models in a zero-shot setting, with prompts found in the Appendix of the paper. We use chatGPT-based language model in all our experiments except those in Figure 11 where we compare multiple language models.

Due to computational expense, we evaluate our approach across benchmarks mainly using three agents with two rounds of debates, although we found further gains with both more agents and rounds of debate (Figure 10). Additional evaluation details are found in the Appendix.

Quantitative Results.

In Table 1, we report the results of each approach on arithmetic, grade school math, and chess reasoning task. In each task, we observe that utilizing multiple different agents to generate solutions improves performance over using a single language model agent to generate a solution. Simultaneously, we also see that reflection, where a language model is asked to critique its early generation, generally gives a modest boost in performance. Multiagent debate, which may be seen as a combination of both reflection and multiagent generation, gives a substantial boost in reasoning across each of the tasks.

Qualitative Results.

In Figure 4, 5, we provide qualitative illustrations of the debate procedure between models. Interestingly, we find cases in which all models initially give an incorrect response, yet the result of debate still obtains the correct answer as agents critique each others’ reasoning. Thus, the purpose of our debate isn’t just to amplify a correct answer – all models can initially be wrong but arrive at the correct answer through the debate process.

Compatibility with other reasoning methods.

Our multiagent generation procedure operates orthogonally approach to other prompting methods which focus on single-agent generation. In Figure 6, we illustrate the performance of multi-agent debate with and without zero-shot chain-of-thought prompting on GSM8K. In both settings, multiagent generation is beneficial.

2 Extracting Factual Information from Multiagent Debate

We next evaluate the extent to which multiagent debate improves the underlying factuality in language models.

We evaluate the factuality of language models in three different settings:

Biographies. To evaluate the factuality of language models, we introduce a novel task of accurately generating historical biographies of people. In preliminary testing, we found that existing language models had a tendency to hallucinate many facts on this task. We constructed ground truth bullet point biographies of 524 well-known computer scientists. We then asked language models to generate bullet point biographies for each person, and evaluated the accuracy at which each ground truth bullet point agreed with generated bullets. We report additional evaluation details in the Appendix.

MMLU. Next, we assess the factuality of language models in responding to different factual knowledge questions typically learned and assessed in different exams. We utilize the existing MMLU dataset to benchmark the accuracy of responses.

Chess Move Validity. Lastly, we study the hallucinations in language models when planning under to the given rules of an existing environment or game. Specifically, we measure the validity of possible moves in a game of Chess given by BIG-Bench Chess-State Tracking Benchmark task of chess-move prediction. In this task, an agent is given a set of next moves, and must make a valid next move of a piece on a board.

Baselines.

We use the same baselines as in Section 3.1. The multiagent (majority) is not directly applicable in this setting as individual responses are not easily comparable, and so we omit baseline comparison with the majority voting in this setting.

Results.

We analyze the performance of each method in Table 2. We found that approaches based on reflection led to poor performance in the factuality setting. In contrast, debate gives the best performance in this setting also, and significantly outperforms each baseline. We illustrate a debate between agents on the biography task in Figure 7 and on MMLU in Figure 8. We found that multiagent debate improved and settled on bullets that were more consistent across agents.

We found that different language agents tended to give different answers when the underlying language model was uncertain about the question. However, directly asking each agent about their confidence of the answer led to high confidence assessments on each answer. However, when these different language agents were asked to communicate with each other, each agent would quickly change their opinion to a consensus answer which was more accurate. We illustrate this in Figure 9. Interestingly, we found that on facts that the language model was confident in (i.e. many instances of the same model all gave the same answer), it was very difficult to convince an agent to change their opinion, suggesting that “ease of persuasion” may be a method to assess factual confidence.

3 Analysis: Understanding Multiagent Debate

Finally, we analyze how multiagent debate improves the underlying language generation procedure in language models.

First, we analyze the impact of agents number in debate. In Figure 10(a), we increase the number of agents used in debate, while fixing the debate length to be two. On arithmetic, performance monotonically increases with the increased number of agents. For larger number of agents, we first summarize all agent responses with chatGPT instead of directly concatenating responses due to context length error.

Rounds of Debate

Next, we analyze the impact of the number of rounds of debate in multiagent debate. In Figure 10(b), we increase the debate length between agents, while fixing the number of agents to three. We find that on the arithmetic task, the performance also monotonically increases with debate length. However, we found that additional debate rounds above four led to a similar final performance to 4 rounds of debate.

Effect of Debate Length on Accuracy

As discussed in Section 2.2, the underlying convergence time in the debate between agents can be controlled by the extent to which agents are encouraged to maintain their opinions. In Figure 12, we consider the effect of short and long-form prompts discussed in Figure 3. We find that debates using longer prompts lead to slower convergence to correct answers, but also lead to a better final consensus on the correct answer. We provide an analysis of consensus between agents in Figure 14.

Using Different Initialization Prompts

In our experiments we use the same prompts for all agents. We also consider the effect of using different questions, where we first instruct each language model to behave like a different persona (professor, doctor, mathematician) on the MMLU dataset. We found that improved performance on MMLU from 71.1 to 74.2 with different agents, suggesting further gains can be obtained with different initialization prompts.

Summarization.

While in the majority of experiments in the paper we directly concatenate the responses of other agents as context for an agent to generate a new response, this is expensive when the number of agents involved in debate gets large. We may alternatively first summarize the responses from all other agents into a single response that we provide to agent at each round for more efficient debate. We apply this strategy in Figure 10 to enable the use of five or more agents in debate. In Figure 13, we analyze the effect compared to directly concatenating the responses of other agents. We find this improves the performance of debate, suggesting that summarization is another tool that can further improve multiagent debate.

Utilizing Different Language Models

Our existing debate results are reported using multiple instances of a chatGPT language model. We further assess the impact of using two different language models, where we ask chatGPT and Bard language models to debate with each other on a set of 20 GSM8K math problems. In this set, we find that multi-agent debate improves the performance of both agents, with Bard solving 11 problems, chatGPT solving 14 problems, and joint multi-agent debate solving 17 problems. We qualitatively illustrate a debate between agents in Figure 11. While both agents initially provide incorrect answers to the problem, chatGPT is able to utilize the incorrect response given by Bard to generate the final correct answer.

Related Work

A wide range of work has explored how to enable reasoning and factuality in language models. To improve reasoning, approaches have relied on prompting techniques such as scratchpads , verification , chain-of-thought demonstrations , and intermediate self-reflection and finetuning . To improve factuality, approaches have relied on training techniques such as RLHF , pruning truthful datasets , external knowledge retrieval and training-free methods based off likelihood estimation .

Our work provides an alternative way to obtain reasoning and factuality in language models using multiagent debates, which only requires black-box access to a language generator. Prior work also has explored how to take the majority vote across different models while in this work, we use the power of a language model to combine different answers. Most similar to our work, Irving et al. also proposes a debate procedure to verify the accuracy and safety of powerful AI agents. In contrast to our approach, in their work, agents are asked to alternatively provide proof of a input, and humans are tasked with assessing these debates and determining safety.

Compositional Generation.

Our work is also related to existing works that focus on text generation by combining different models . Most similar to our work, propose to combine multiple different large pretrained models together for multimodal reasoning. In contrast, in our work, we aim to use communication between different language models to enable more effective reasoning and factuality in language models.

Limitations and Discussion

In this paper, we present an orthogonal approach to improve the performance of language models using multi-agent debate. We find that the approach is simple and effective across a wide set of different reasoning and validity language modeling tasks.

In comparison to other prompting techniques, our multiagent debate procedure is more computationally expensive, as it requires both multiple language generations, and an underlying debate procedure. However, we believe that this approach may be seen as a method to generate additional data that may be distilled back to self-improve the original base model.

Further, we observed that as debates became longer in duration, current language models sometimes struggled to fully process the entire debate input, and typically only focused on the most recent generations. We believe that this performance will be alleviated with longer-context and improved language models or by summarizing early portions of the debate.

Finally, we found that while debates typically converged into single final answers, these answers were not necessarily correct. Despite answers being incorrect, language models would confidently affirm that their answer is correct and consistent with all other agent responses. We believe this result is in part due to the fact that LMs do not correctly express their uncertainty when generating responses, and believe that other orthogonal approaches to improve this performance would improve the results of multiagent debate.

References

Appendix A Appendix

In this appendix, we provide additional analysis and visualizations of the debates used in the main paper in Section A.1. We further provide detailed experimental details on each dataset in Section A.2.

In Figure 14, we illustrate the consensus between agents using either short or long consensus prompts discussed in Figure 3. The use of debate prompts that encourage agents to adapt more to the opinions of other agents improves consensus.

Additional Qualitative Visualizations.

We added additional qualitative visualizations of the debate process. In Figure 16, Figure 17, Figure 18, Figure 19, Figure 20, we illustrate debates between agents in the GSM8K dataset which result in the correct answer. In Figure 21, Figure 22, Figure 23, we further illustrate debates in GSM8K which lead to the incorrect answer. We further provide an example illustration of debate in arithmetic in Figure 24, arithmetic with summarization of individual responses of agents in Figure 25, MMLU in Figure 26, a debate with the full contents biographies in Figure 27, and debate in chess in Figure 28. In general, we found that debate improved the performance of final generated answers, though sometimes answers would converge to the incorrect value.

A.2 Evaluation Details

We provided detailed evaluation details for each setting in the paper. We run all experiments using the gpt-3.5-turbo-0301 model. We provide a table listing the prompts used to prompt models and initialize debate in Table 15.

To evaluate the arithmetic task, we generated six random integers for each task between 0 and 30. We then evaluated the extent to which the correct integer answer was correctly obtained. We evaluated models on one hundred generated arithmetic tasks.

Grade School Math.

To evaluate the GSM8K task, we evaluated the accuracy at which models were able to obtain the final correct answer, as extracted from a box. We evaluated models on one hundred grade school math problems.

Chess.

To evaluate the chess reasoning task, we used chess games from https://www.pgnmentor.com/players/Adams.zip. We asked chatGPT to predict the next move for white to move at turn 14 and reported the relative Stockfish pawn score with search depth 20 after executing the suggested move from chatGPT. We evaluated models on three hundred selected chess games.

Biographies.

To evaluate the biographies task, we compare each generated bullet point biography for a person with a ground truth set of facts about the person extracted from Wikipedia. We iteratively loop through each ground truth fact, and validate the extent to which the generated biography matches a particular bullet by prompting chatGPT with the prompt: Consider the following biography of : Is the above biography above consistent with the fact below? Give a single-word answer, yes, no, or uncertain. We then evaluate and report the percentage of ground bullets that chatGPT returns either yes or no on. We ignored ground truth bullets that chatGPT returns returned uncertain.

We found this evaluation metric provided a fast way to evaluate how relatively correct a generated bullet point biography is. However, we found that generated facts could contain incorrect information that was not captured in the ground truth bullet and thus could not be validated through this metric. Nevertheless, we believe this evaluation scheme estimates the relative accuracy of a generated biography.

MMLU.

To evaluate MMLU, we measured the accuracy in which models were able to select the correct multiple-choice answer in each problem. We evaluated models on one hundred selected MMLU questions randomly distributed across each of the subject areas.

Chess Validity.

To evaluate chess validity, we consider the BIG-Bench Chess-State Tracking Benchmark , where we used the hardest reported task in the benchmark synthetic_short. Each generated answer was deemed correct as long as it was one of the valid answers in the sequence. We evaluated models of one hundred selected chess validity tasks.