ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
Justin Chih-Yao Chen, Swarnadeep Saha, Mohit Bansal
Introduction
A large body of work has focused on improving the reasoning capabilities of Large Language Models (LLMs) by imitating various human cognitive processes. These include phenomena like reflecting on and critiquing one’s own predictions, being receptive to feedback, and learning from feedback. Of note, self-reflection is an introspective process that allows the model to improve its outputs by generating feedback from the model itself (Madaan et al., 2023; Shinn et al., 2023). However, self-reflection suffers from Degeneration-of-Thought – when the model is overly confident in its answer, it is unable to generate novel thoughts even after multiple rounds of feedback (Liang et al., 2023). Moreover, it is difficult for the model to refine knowledge that it already does not contain.
In order to promote more diverse thoughts, past work has drawn inspiration from the society of minds in multi-agent systems (Minsky, 1988; Zhuge et al., 2023). Communication between multiple agents plays a vital role in complex decision-making. This has prompted recent developments of multi-agent debating frameworks (Liang et al., 2023; Du et al., 2023), in which multiple agents participate in a multi-round debate to arrive at a common final answer. Despite the increased reasoning diversity obtained through the process of a debate, multiple agents have typically been limited to different instances of the same underlying model, ChatGPT (OpenAI, 2022).In this work, we refer to multi-agent as multiple instances of the same underlying model (e.g., ChatGPT), whereas multi-model model-agent refers to different models (e.g., ChatGPT, Bard and Claude2) as agents. This results in an inherent model bias, a restricted knowledge scope, and a lack of external feedback from other models due to identical pre-training data and model architectures across all agents. Relatedly, ensemble methods like self-consistency generate the most consistent answer via sampling diverse reasoning paths from the same model (Wang et al., 2023b) but do not incorporate any internal or external feedback. In general, when multiple agents propose diverse solutions to a problem, the success of such a multi-agent system is fundamentally reliant on the ability to estimate each agent’s confidence and accordingly, convince other agents (with explanations) to reach a better consensus. This puts forward the following question: if multiple diverse LLMs are asked to collaboratively solve a task, are they capable of discussing their solutions with each other such that a better consensus is reached?
We aim to solve complex reasoning problems by learning from diverse insights and external feedback, originating from agents that belong to different model families. Collaborative processes such as brainstorming, group meetings, and discussions play a pivotal role in reaching a consensus and arriving at more refined solutions to complex problems (Li et al., 2022b). Effective discussion also entails the selection of stances, voting, convincing, exchange of information, and a diversity of opinions. This leads us to propose \method, a novel method of round-table conference for consensus among diverse LLM agents. \method consists of multiple discussion rounds between diverse LLM agents who try to convinceWhen we say that an agent tries to convince another agent, we mean that it learns (based on corrective explanations) to defend or argue for its stance while still being receptive to the other agent’s argument such that a better consensus is reached. each other to either rectify their answers or become more confident of their initial correct answers (see Fig. 1 for a broad overview). The central motivation of \method stems from the fact that in a collaborative environment, all participants engaging in a discussion hold their own opinions at the beginning, and a consensus within is achieved through various communicative aspects, including convincing others, voting for a decision, and the adjustment of positions along with their associated confidence levels.
Given a reasoning problem, \method begins with each agent first generating an answer, its associated uncertainty, and a corresponding explanation (as a Chain-of-Thought (Wei et al., 2022)) for the answer. Then all agents enter a multi-round discussion phase. Each discussion round comprises of all agents generating a revised explanation and answer based on all other agents’ explanations and answers from the previous round. The goal of the revised response is to convince other agents to reach a better consensus. In particular, \method initiates a discussion by designing a discussion prompt for each agent, that lets it condition on (1) grouped answers from all agents, (2) corresponding explanations generated in the previous round, and (3) demonstrations of samples with human explanations (that rectify an agent’s initial incorrect answer) that can convince other agents. We leverage them in an in-context learning framework to teach models to generate their own convincing explanations (see Fig. 3). Even in cases where an agent initially offers an incorrect answer and explanation, it can consider another agent’s convincing explanation and amend its response accordingly. In each round of the discussion, we estimate an agent’s uncertainty via a confidence-estimation prompt (Tian et al., 2023; Xiong et al., 2023a). Once all agents converge to the same answer (i.e., a consensus has been reached), we employ these confidences to compute a weighted vote as the final answer.
We primarily develop \method with three state-of-the-art LLMs, Bard (Anil et al., 2023), Claude2 (Anthropic, 2023), and ChatGPT (OpenAI, 2022). We show our method’s efficacy on multiple commonsense reasoning (StrategyQA (Geva et al., 2021), ECQA (Aggarwal et al., 2021)) and mathematical reasoning (AQuA (Ling et al., 2017) and GSM8K (Cobbe et al., 2021)) benchmarks. Our first result demonstrates that across all four datasets, \method outperforms prior single-agent (e.g., Self-Refine (Madaan et al., 2023) and Self-consistency (Wang et al., 2023b)) and multi-agent baselines (Debate (Du et al., 2023) and Judge (Liang et al., 2023)). On the commonsense reasoning benchmarks, \method exhibits especially strong results, outperforming a much stronger model like GPT-4 by up to 3.4%. We find that \method not only improves the overall team performance, but also leads to significant gains for each agent individually. We also conduct detailed analyses of the individual components of \method and demonstrate that leveraging diverse LLM agents and the usage of convincing samples lead to maximum improvements. In particular, convincing samples lead to a 4% improvement as compared to general human explanations without our novel answer-rectifying selection criterion. Convincing samples also generally benefit prior methods like the multi-agent debate (Du et al., 2023). We also analyze the accuracy at the end of each discussion round and show that \method not only improves performance after each round compared to multi-agent debate but also reaches faster consensus (i.e., in a lesser number of rounds), thus pointing to its efficiency.
Finally, as an initial investigation, we implement an alternative version of \method, wherein ChatGPT is replaced with GPT-4 as one of the agents. Note that GPT-4 is a much stronger model in terms of its performance compared to the other agents in consideration here (Zheng et al., 2023; OpenAI, 2023) and is also substantially more expensive. However, our results demonstrate that even when one agent (GPT-4) possesses considerably greater strength than others (Bard and Claude2), collaborative discussions facilitated by our framework individually benefit all agents, even improving GPT-4’s initial accuracy by large margins (e.g., an absolute 10.0% on StrategyQA).
In summary, our primary contributions are as follows:
We propose \method, a novel method for improving reasoning with diverse Large Language Models involved in a Round Table Conference.
We study the role of confidence estimation and discussion in multi-agent systems and the ability of an agent to convince other agents (by learning from corrective explanations) to reach a better consensus.
We conduct extensive experiments on multiple reasoning datasets, involving math and commonsense and show that \method improves upon prior single-agent and multi-agent baselines and also outperforms GPT-4 on some benchmarks. We also experiment with a version of \method using GPT-4 as an agent and show that mutual discussion among diverse agents significantly improves GPT-4’s initial accuracy.
We analyze the effect of individual components in \method and observe that the usage of diverse LLM agents and convincing samples leads to significant gains. \method also improves the efficiency of discussion, reaching a faster and better consensus compared to a multi-agent debate baseline.
Motivation and Problem Setup
When confronted with complex reasoning tasks or open-ended questions, humans often resort to collective brainstorming, discussions, and leveraging the power of group intelligence, also referred to as the wisdom of the crowd or the society of minds (Minsky, 1988). Taking inspiration from this, we propose \method that harnesses multiple Large Language Models (LLMs) in a multi-agent discussion procedure with confidence estimation and convincing explanations to improve the overall reasoning capabilities.
We assume that we are given a test problem and there are agents participating in a round table discussion. Each agent is a distinct LLM, potentially trained with different pre-training data and model architectures. All agents are capable of generating an answer and a corresponding explanation (as a Chain-of-Thought (Wei et al., 2022)) for the test problem. For each agent , we utilize a small number of demonstrations of convincing samples . Each convincing sample for an agent is an instance of a question , gold answer , and a human explanation that helps rectify an agent’s initial incorrect answer (see more details in Sec 3.2). The objective of \method is to improve the team performance on a given task by holding multiple rounds of discussion between the agents, quantifying the uncertainty associated with each agent, and convincing the other agents to reach a better consensus. Note that convincing samples serve as an additional performance enhancer: even when the dataset lacks human explanations or one opts not to utilize them, our method can still yield performance improvements independently of this technique (more details in Sec. 5.3).
\method: A Group-Discuss-And-Convince Framework
operates in three phases. In phase 1, all agents generate their initial responses. In phase 2, \method initiates a multi-round discussion between the agents. Once the discussion terminates, in phase 3, \method generates the final answer. The overview of our method is demonstrated in Fig. 2, and the workflow is illustrated in Algorithm 1.
operates with each agent initially generating an answer , an explanation , and an associated confidence for the generated answer. Each agent conditions on a zero-shot prompt that instructs it to reason about the problem ‘step-by-step’. See ‘Phase 1’ in Fig. 2 and the prompt is shown in Fig. 5.
2 Phase 2: Multi-round Discussion
then enters a discussion phase, consisting of rounds (see ‘Phase 2’ in Fig. 2). Discussion for round proceeds as follows. For each agent , \method develops a discussion prompt (as shown in Fig. 5), consisting of the following three components.
Grouped responses of all agents from the previous round. consists of the answers and explanations of all agents from the previous round . In order to foster better discussions, \method summarizes this information by grouping the responses into distinct answer categories and appends all plausible explanations for each answer, as illustrated on the left side of Fig. 2 and Fig. 5.
Confidence associated with the answers. All agents are not equally confident of their answers. Hence, an effective discussion should also take each agent’s uncertainty into account. Since our agents are black-box models, we estimate each agent’s confidence in round by directly prompting the agent to verbally quantify its uncertainty (Xiong et al., 2023b). This is also shown in our prompt in Fig. 5.
Convincing samples from all other agents. Finally, the prompt also contains convincing samples for all other agents . When an agent tries to reassess its reasoning in light of the reasoning provided by other agents, we hypothesize that it should benefit from conditioning on demonstrations that can convince other agents. In order to obtain such convincing samples for an agent , we choose a small number of samples (4 in our experiments) from the training set for which the agent’s initial answer is wrong but conditioning on the corresponding human explanation, rectifies the answer.We experiment with datasets that are annotated with high-quality human explanations in prior work. Additionally, these human explanations are background facts (for commonsense tasks) or some intermediate steps leading to the final answer (for mathematical tasks). Hence, the explanation will never give away the answer directly and the model still needs to reason over it to derive the correct answer. See Fig. 3 for an illustration of the process.
Based on the above components, the prompt is formally defined as:
Each agent in discussion round then conditions on the corresponding discussion prompt to generate an updated answer , explanation , and associated confidence , to be used in the next round. In each round, an agent can review the stances and opinions on the table, reassess all agents’ reasoning processes and solutions, and subsequently provide an updated response. Demonstrations of convincing explanations enable the agent to generate explanations that are more likely to convince other agents to reach a better consensus in the following round.
3 Phase 3: Final Answer Generation
For each data point, \method continues the discussion for a maximum of rounds or terminates it as soon as a consensus is reached (i.e., all agents agree on the same answer). At the end of any round , \method generates the final answer for that round using a weighted voting scheme (see the right side of Fig. 2). In particular, it converts the model’s confidence into a weight and employs this weight in a weighted voting scheme to determine the final answer. Directly using confidence scores as the voting weights is less effective due to the overconfidence problem of LLMs (Xiong et al., 2023b; Tian et al., 2023; Mielke et al., 2022). Specifically, LLMs tend to produce consistently high confidence scores, which can make it challenging to discern subtle distinctions in confidence levels across different outputs. To address this issue, we employ the following simple yet effective rescaling technique to adjust the confidence scores , facilitating better differentiation of confidence levels.
where is the original confidence for agent in round and is the corresponding adjusted score. As we will show later in the experiments, this simple recalibration method, used as a weighing scheme for obtaining the final answer, works well in practice across multiple datasets. Fig. 9 in the Appendix also shows that it helps reduce the Expected Calibration Error (ECE), a popular calibration metric (Naeini et al., 2015). While we note that recalibration can also be achieved through a learned model (e.g., Platt Scaling (Platt et al., 1999)), we refrain from using such models because \method is primarily designed as a few-shot method, and developing a recalibration model would necessitate access to a substantial number of annotated samples.
We use to perform a weighted vote to generate the final answer as follows.
where is a distinct answer generated by any of the agents.
Experimental Setup
We primarily implement \method with three state-of-the-art LLMs: ChatGPT, Bard, and Claude2, engaging them in up to three rounds of discussion. Later, in Section 5.1, we also develop and experiment with a version of \method that replaces ChatGPT with GPT-4 as one of the agents. Henceforth, we will refer to the initial predictions made by the agents as ‘Round 0’ of discussion. During decoding, we set the temperature to 0.7 for ChatGPT and Bard and use the default setting for Claude2. All implementations involving ChatGPT are using gpt-3.5-turbo-0613 from Azure OpenAI.https://oai.azure.com/ We retrieve results from Claude2 by posting requests to their webpagehttps://claude.ai/chats, and for Bard, we use chat-bison-001 from PaLM2 APIhttps://developers.generativeai.google/products/palm to obtain responses. For each agent, we use four demonstrations of convincing samples.
2 Tasks and Metrics
We evaluate \method on two commonsense reasoning and two math reasoning tasks. These include (1) StrategyQA (Geva et al., 2021), a task of implicit reasoning for multi-hop questions, (2) ECQA (Aggarwal et al., 2021), a commonsense reasoning dataset, (3) GSM8K (Cobbe et al., 2021), a benchmark of math word problems, and (4) AQuA (Ling et al., 2017), a dataset of algebraic word problems. Owing to the costly nature of conducting experiments with black-box models and the limit imposed on the number of API calls, we follow many prior works (Du et al., 2023; Bian et al., 2023; Besta et al., 2023; Yao et al., 2023) and experiment with a subset of 100 samples (from the validation set for StrategyQA and the test set for all other datasets). We report accuracy and its associated standard deviation for all tasks. For each experiment, we conduct at least three runs on the same test samples with the same prompts, primarily accounting for the variance caused due to the decoding strategy.
Results and Analysis
Our first experiment evaluates the overall reasoning capabilities of \method. Initially, we focus on the version of \method with ChatGPT, Bard, and Claude2 as the three agents. Then, later in this section, we also report our findings when using a stronger GPT-4 model as an agent. We compare \method to prior works that can be broadly grouped into three categories:
Vanilla single-agent methods. Our first set of baselines includes zero-shot Chain-of-Thought prompting with GPT-4, ChatGPT, Bard, and Claude2 that instructs the model to answer the question ‘step-by-step’ (Kojima et al., 2022).
Advanced single-agent methods. Next, we compare with (1) a Self-Refine (SR) baseline that iteratively generates feedback leveraging the model itself and then uses that feedback to refine the output (Madaan et al., 2023), (2) a Self-Consistency (SC) baseline that samples multiple reasoning paths and generates the most consistent answer (Wang et al., 2023b), and (3) their combination, SR+SC, that first conducts multiple iterations of refinement, followed by a majority vote of the refined answers. We implement these baselines on top of ChatGPT.
Multi-agent methods with a single backbone model. Our final baselines are two recently proposed multi-agent debating methods. In particular, we compare with Du et al. (2023), who propose a multi-agent debate between multiple instances of ChatGPT and Liang et al. (2023), who additionally include a judge to monitor the debate process.
For fair comparisons, all iterative methods (either involving refinement, debate, or discussion) go through 3 rounds of iteration and all multi-agent methods are implemented with three agents. We report our results in Table 1. Our primary observation is that across all four datasets, \method, developed with ChatGPT, Bard, and Claude2 as the agents, improves upon all single-agent and multi-agent baselines that are also built on top of these agents (see last row). On commonsense reasoning tasks like StrategyQA and ECQA, our method also outperforms GPT-4 (without using it as an agent). Note that between all single agents, GPT-4 exhibits significantly better performance on all four benchmarks. Therefore, \method’s ability to match or surpass it while leveraging the three comparatively weaker agents (ChatGPT, Bard, and Claude2) shows the promise of our framework. On the math reasoning tasks (GSM8K and AQuA), \method matches or closes the gap to GPT-4, which is also the state-of-the-art LLM on GSM8K. GPT-4’s especially strong results on GSM8K could be attributed in part to the inclusion of some of GSM8K’s training samples in GPT-4’s pre-training data (OpenAI, 2023).
As also shown in prior works, advanced single-agent methods are better than their vanilla counterparts (see ‘Method Category’ column in Table 1) (Wang et al., 2023b; Madaan et al., 2023). Multi-agent debate with ChatGPT (Du et al., 2023) improves results further, especially on the math datasets. On the other hand, debate with multiple instances of Bard and Claude2 is not effective. We hypothesize that the feedback, originating out of multiple instances of the same underlying model of Bard and Claude2 is not diverse enough. While multi-agent debate with Bard and Claude2 is not effective, when they team up with ChatGPT in a multi-round discussion, \method outperforms debate frameworks that are built on top of these agents. It obtains maximum improvements of 7.7% accuracy on the commonsense reasoning tasks compared to the strongest baseline, Multi-agent debate with Claude2. Improvements in the math reasoning tasks are relatively moderate, because of ChatGPT’s initial strong performance on these tasks.
So far, we have demonstrated that the final team performance of the agents improves through the discussion process. Next, we investigate the round-wise accuracy of each individual agent on the StrategyQA dataset in Table 2. In addition to the accuracy obtained by each agent individually, we also report the team performance, using three different voting mechanisms. These are: (1) our proposed weighted vote, (2) simple majority vote, and (3) choosing the agent with the maximum confidence. We observe that after the initial response generation, both individual and the team accuracy increase for at least two rounds when using weighted voting and majority voting mechanisms. Note that simply choosing the most confident agent proves ineffective. Finally, as discussion progresses further to round 3, each agent’s performance tends to saturate (see more details about \method’s faster and better consensus in Sec. 5.4).
Using GPT-4 as an agent in \method.
In the above section, we showed the effectiveness of \method using ChatGPT, Bard, and Claude2 as the three agents to even outperform GPT-4 in some cases. Based on our results in Table 1 and prior work (OpenAI, 2023; Zheng et al., 2023), GPT-4 is likely the strongest (and also, the most expensive) LLM out of all the models we experiment with. Next, as an initial investigation, we also study the potential of GPT-4 to participate in a multi-round discussion with comparatively weaker agents. To this end, we implement \method with GPT-4, Bard, and Claude2 (i.e., replacing ChatGPT with GPT-4). In Table 3, we report the accuracy obtained by each agent at the end of each discussion round. In addition, we provide the zero-shot performance of each agent as an additional baseline. Note that the zero-shot results are different from \method’s Round 0 results because of the differences in their respective prompts: the latter incorporates convincing samples. With increasing rounds, the accuracy of each agent improves, showing that all models benefit from mutual discussions. GPT-4’s absolute improvement by 10% is particularly encouraging because it is the strongest participant, and our result highlights the potential for a stronger agent to obtain useful external feedback from comparatively weaker agents, and thereby augmenting its own capability. To further validate that this improvement is indeed due to the discussion process of \method with other agents, we compare GPT-4’s final accuracy () with 3 rounds of Debate and Self-Refine baselines (both also implemented with GPT-4). We observe both of these baselines yield significantly lower accuracies, at and respectively. In summary, our \method method holds the potential to involve agents with diverse capabilities in round-table discussions, such that all agents improve individually. Note that the weighted voting scheme becomes less effective in such scenarios and tends to converge towards the dominant agent (GPT-4, in this case). This is why we primarily focus on studying agents with similar capabilities (e.g., ChatGPT, Bard, and Claude2) in this paper and all our other analyses in the following sections are also with this setup.
2 Ablations of \method: All Components are Beneficial
In Table 4, we evaluate the effect of individual components of \method on the StrategyQA dataset. In particular, we compare \method with four of its variants: (1) w/o Multiple Models: Instead of using different models as different agents, we use ChatGPT as the backbone for all three agents, (2) w/o Grouping: We simply concatenate the generated responses from different agents without summarizing and grouping their answers, (3) w/o Convincingness: We remove convincing samples from the initial prompt and the discussion prompt, and (4) w/o Confidence Estimation: We do not use any confidence estimates during the discussion and compute majority vote as the final answer. We show that each component has a positive impact on \method with varying capacities. The effect of using different models as different agents is particularly significant and we observe a 6.8% improvement compared to only using ChatGPT as all three agents in \method. This reinforces our hypothesis that diverse LLMs have complementary strengths and when put together in a round table discussion, they can learn from diverse external feedback and refine their responses to reach a better consensus. Next, grouping answers is beneficial too, demonstrating that summarizing the stances of each agent fosters better discussion. We also show that using convincing samples leads to a 4.5% improvement in accuracy. We analyze the role of convincing samples in more detail in Sec. 5.3. Finally, estimating the confidence of each agent and using it to weigh the agents’ answers outperforms majority voting.
3 Convincing Samples Improve Both \method and Multi-agent Debate
In this section, we conduct a comprehensive analysis of the role of convincing samples that are, essentially, demonstrations of answer-rectifying human explanations. Recall that \method selects a sample as convincing if the corresponding human explanation helps rectify an agent’s initially incorrect answer. Based on this, Table 4 showed that at the cost of collecting a small number of human explanations (four in our case), we can obtain significant improvements (\method row versus ‘w/o Convincingness’ row). Next, we consider a scenario where no human explanations are present. Table 5 shows that \method, even without access to any convincing samples, outperforms the multi-agent debate baseline by absolute 7.8 points (first versus second row, 74.5% v.s. 66.7%). If random human explanations (i.e., general explanations that may not necessarily ensure answer rectification, as per Fig. 3) are available (third row), we obtain some small improvements; but our convincing samples that are selected based on our novel answer-rectification criterion (last row) improve the results substantially. In Appendix A.2 and A.3, we show two illustrative examples of the discussion process without and with convincing samples respectively.
Next, in Table 6, we show that convincing samples, in general, can boost other multi-agent frameworks as well. On top of Multi-agent Debate (with ChatGPT agents), the inclusion of convincing samples leads to improved results compared to the original setup and one with random human explanations. To summarize, being able to convince another agent is a generic concept that can be applied to other multi-agent systems.
4 Analysis per Discussion Round: \method Reaches Faster and Better Consensus
terminates discussion as soon as a consensus is reached i.e., all agents have converged to the same answer. Extending the number of discussion rounds will be costlier due to the increased API calls to black-box models. Hence, achieving faster consensus while maintaining comparable accuracy improvements is more efficient. To study this, in Fig. 4, we plot the accuracy trends at the end of each discussion round; in Fig. 4, we plot the fraction of samples for which consensus has been reached after each discussion round; and finally, in Fig. 4, we analyze accuracy as a function of consensus. From the first plot, we make two important observations: (1) \method improves reasoning performance for two discussion rounds, following which the accuracy saturates, (2) Compared to the debate baselines, \method is not only superior after every round but also peaks at a highest accuracy of 79.0% versus 71.3% for the baselines. Next, from Fig. 4, our observations are also two-fold: (1) In the initial rounds (round 0 and 1), \method’s consensus percentage is lower because the discussion takes place between diverse LLM agents. Diverse agents lead to more differences in opinions initially. (2) However, as the discussion proceeds, \method establishes consensus for all samples by round 3, while in the debate baseline, 13% of the samples do not converge even after round 4. Finally, Fig. 4 shows that for the fraction of samples that enter the multi-round discussion phase (i.e., their initial answers did not have a consensus), accuracy is positively correlated with consensus percentage. In other words, as a greater number of samples reach a consensus, accuracy proportionally improves, effectively pointing to better consensus. In summary, we demonstrate that \method reaches faster and better consensus compared to prior multi-agent baselines, in spite of starting with more diverse responses from different models.
Related Work
Progress in Large Language Models has led to the development of advanced prompting and fine-tuning techniques for reasoning. Representative methods include Chain-of-Thought (CoT) (Kojima et al., 2022; Wei et al., 2022; Wang et al., 2023a) and Tree-of-Thought prompting (Yao et al., 2023), self-consistency (Wang et al., 2023b), meta-reasoning over multiple paths (Yoran et al., 2023), use of scratchpads (Nye et al., 2021), training verifiers (Cobbe et al., 2021), self-reflection (Shinn et al., 2023; Madaan et al., 2023; Wang & Zhao, 2023), and fine-tuning via bootstrapping models (Zelikman et al., 2022; Lewkowycz et al., 2022). Eliciting reasoning from a single agent, while promising, is prone to model bias, degeneration-of-thought, and is fundamentally limited by a lack of diverse insights about the problem due to the absence of external feedback (Liang et al., 2023).
2 Reasoning in Multi-Agent Systems
A recent line of work has explored student-teacher frameworks with the goal of distilling reasoning capabilities from a stronger teacher to a weaker student (Magister et al., 2023; Fu et al., 2023; Ho et al., 2023; Saha et al., 2023; Mukherjee et al., 2023). As opposed to a teacher teaching weaker agents, we are interested in developing a multi-agent system where different LLM agents have their unique strengths and try to collaboratively improve the performance on a reasoning task by convincing each other (using corrective human explanations for in-context learning) to reach a better consensus. Among notable prior works, researchers have proposed multi-agent debating frameworks (Du et al., 2023; Liang et al., 2023; Chan et al., 2023; Xiong et al., 2023a) but such efforts are still largely limited to multiple instances of the same underlying language model. We argue that relying on a single model limits the potential of complementary benefits from different model families and the advantage of ensemble learning. Different models possess varied strengths and weaknesses and consequently, combining the contributions of each model holds the promise of improved robustness and overall accuracy. Moreover, estimating the confidence of each agent and being able to defend or improve one’s opinions become more prominent components in such multi-model multi-agent systems because of the individual differences. Overall, Table 7 summarizes \method’s key differences compared to prior single-agent and multi-agent reasoning methods.
3 Ensembling Large Pretrained Models
Large pretrained models, by virtue of being trained on different data and with architectural variations, exhibit distinct capabilities. This has led to the development of ensembles (Sagi & Rokach, 2018) in multimodal learning (Zeng et al., 2023; Li et al., 2022a). Mixture of Experts, a popular ensemble learning technique, trains multiple smaller specialized models to improve robustness and overall accuracy (Jacobs et al., 1991; Shazeer et al., 2017; Du et al., 2022). Specific to language models, Self-Consistency (Wang et al., 2023b) generates diverse reasoning paths using CoT and chooses the most consistent answer as the final output. Jiang et al. (2023) propose LLM-Blender, a method to rank and fuse generations from different models. Different from these, we study communication via explanations between distinct LLM agents and their ability to discuss and convince each other in order to improve collective reasoning.
Conclusion
We presented \method, a multi-agent framework for improving reasoning with diverse LLM agents, engaged in multiple rounds of discussion via confidence estimation and generating explanations that can convince other agents. \method demonstrated strong results on multiple commonsense and mathematical reasoning benchmarks, consistently outperforming prior single-agent and multi-agent baselines and even improving upon GPT-4 on some benchmarks. Moreover, when GPT-4 was used as one of the agents, \method improved its initial accuracy by 10 absolute points. We also showed that compared to a multi-agent debate baseline, \method helps establish better and faster consensus between agents. \method shows the promise of leveraging diverse language agents in a collaborative setup to discuss and accomplish complex tasks.
Limitations
Given that the current best open-source models often face difficulties with lengthy instructions and prompts (Zheng et al., 2023), our framework employs three prominent API-based models as agents. However, we note that we lack complete knowledge of the data that these models have been exposed to, their scales in terms of parameters, and due to their API access, we also do not possess complete control over their behavior. Depending on API-based models also necessitates the need to prompt these models to obtain their confidence estimates. While this approach proves effective as evidenced by our results, we note that these estimates remain post-hoc in nature. Nevertheless, it is worth highlighting that this limitation could potentially be mitigated in the future should a new state-of-the-art open-sourced model emerge, demonstrating robust capabilities in adhering to instructions, handling extensive prompts, and adapting from feedback. Moreover, we are also making our code, prompts, and result logs publicly available to enable replication of our findings.
Acknowledgments
We thank Peter Hase and Elias Stengel-Eskin for useful feedback and suggestions regarding experiments. This work was supported by NSF-CAREER Award 1846185, NSF-AI Engage Institute DRL-2112635, DARPA MCS Grant N66001-19-2-4031, Accelerate Foundation Models Research program, and a Google PhD Fellowship. The views contained in this article are those of the authors and not of the funding agency.
References
Appendix A Appendix
We show the prompts used in \method in Fig. 5. The initial prompt encompasses (1) the convincing samples that demonstrate how to convince other agents, (2) the test question, and (3) a requirement for ‘step-by-step’ reasoning. The prompt also instructs the agent to express their confidence level, ranging from 0.0 to 1.0, indicating the likelihood of their answer being correct. The discussion prompt is an extension of the initial prompt, instructing the agent to review and express agreement or disagreement with other agents’ solutions. To facilitate discussions, we design a grouping scheme that aggregates information based on the current opinions at the table. For instance, if two agents affirm that the answer to a given question is ‘yes’ while the third agent disagrees with a ‘no’, the designed grouping mechanism in discussion prompt consolidates this information rather than simply concatenating all responses.
A.2 Demonstration of \method without Convincing Samples
We notice that when \method operates in the absence of convincing samples, the agents tend to maintain their initial opinions more often. As depicted in Fig. 6, all three agents adhere to their original stances throughout the entire discussion and hence never converge to the correct answer.
A.3 Demonstration of \method with Convincing Samples
On the contrary, when convincing samples are present, we show how the explanations of all agents change during the course of a discussion (see Fig. 7). Initially, Claude and Bard provide incorrect answers, but as the discussion unfolds, both agents revise their initial predictions, ultimately arriving at the correct answer.
A.4 Demonstration of Single-Model Multi-Agent Debate Struggling with Echo Chamber
In Figure 8, we provide an illustration of multi-agent debate, implemented with multiple instances of the same underlying ChatGPT model. In this case, an incorrect answer is initially provided, but because external feedback from diverse models is lacking, all agents persist with the same incorrect response throughout the interaction.