Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, Zhaopeng Tu
Introduction
Modern large language models (LLMs) like ChatGPT, GPT-4 OpenAI (2023) and Bardhttps://bard.google.com/, have shown remarkable performance on general language tasks Jiao et al. (2023); Wu et al. (2023); Bang et al. (2023) but still struggle on complex reasoning tasks Zhu et al. (2022); Gou et al. (2023), which drives the research on cognitive behaviors of LLMs to explore human-like problem-solving strategies. In particular, self-reflection Madaan et al. (2023); Shinn et al. (2023), a concept that usually refers to the process of introspection and examination of a person’s own thoughts, has been explored to solve intricate tasks that could be challenging for a zero-shot generation or even chain-of-thought (CoT) prompting Wei et al. (2022). Specifically, self-reflection involves an iterative refinement process such that the LLM generates a new answer based on the answers and feedback in previous iterations and then provides feedback for the new answer. While self-reflection can be effective in creating better solutions, it is highly dependent on the self-evaluation capabilities of LLMs, which are not formally guaranteed Shinn et al. (2023).
In this work, we focus on the Degeneration-of-Thought (DoT) problem in self-reflection, which is proposed and defined by us for the first time. Formally, DoT describes the following scenario:
Once the LLM has established confidence in its answers, it is unable to generate novel thoughts later through self-reflection even if the initial stance is incorrect.
To demonstrate this problem, we define the average disagreement as the percentage of opposition between two debaters in debate (or self-confliction in self-reflection) for each question. As Figure 1 seen, we calculate the disagreement of stances between every two iterations in self-reflection and show the trends. The low disagreement of self-reflection suggests that the LLM sticks to the incorrect answers predicted by CoT and is unable to engage in meaningful self-reflection. There are various factors that could result in DoT, and we outline three here: (1) Bias and Distorted Perception. Self-perception can be influenced by biases, preconceived notions, and distorted thinking patterns, which can be learned from the massive amount of data during pretraining. If an LLM’s self-reflection is clouded by such biases or distorted thinking, it can lead to inaccurate conclusions instinctively. (2) Rigidity and Resistance to Change. Self-reflection often involves challenging one’s beliefs, assumptions, and behaviors. If an LLM is resistant to change or holds rigid beliefs, it may struggle to engage in meaningful self-reflection that leads to better solutions. (3) Limited External Feedback. Self-reflection is primarily an internal process, but external feedback can provide valuable perspectives and insights. Without seeking or considering external feedback, an LLM may miss important blind spots or alternative viewpoints that can enrich its self-reflection.
To address the DoT issue, we leverage another fundamental characteristic of human problem-solving, i.e., debate, to encourage divergent thinking in LLMs. Specifically, we propose the MAD framework, short for Multi-Agent Debate, where two agents express their own arguments in the state of “tit for tat” and a judge monitors and manages the debate process to obtain a final solution. The nature of MAD determines that (1) The distorted thinking of one LLM can be corrected by the others; (2) The resistance to change of one LLM will be complemented by the others; and (3) each agent can obtain external feedback from the others. Therefore, MAD is less susceptible to the factors of DoT, and can explore divergent chain-of-thoughts to achieve accurate solutions.
We conducted experiments on both natural language generation and understanding through two challenging tasks, namely, Commonsense Machine Translation (Common MT) and Counter-Intuitive Arithmetic Reasoning (Counter-Intuitive AR). The common characteristic of the two tasks is that our instincts are mostly incorrect based on only the superficial expressions of the questions, and deeper levels of contemplation are required for better solutions. Experimental results demonstrate that our MAD framework performs much better than the baseline methods, especially, MAD with GPT-3.5-Turbo can surpass the performance of GPT-4 on Common MT.
The contributions of this work are summarized as follows:
We propose and define the Degeneration-of-Thought (DoT) problem in self-reflection, and address it by proposing the Multi-Agent Debate (MAD) framework to explore divergent chain-of-thoughts.
We demonstrate the effectiveness of MAD on two challenging tasks, and find that GPT-3.5-Turbo with MAD can even surpass GPT-4 on the Common MT dataset.
Extensive analyses suggest that the adaptive break strategy of debate and the modest level of “tit for tat” state are required for MAD to obtain good performance. More interestingly, we find that LLMs might not be a fair judge if different LLMs are used for agents.
Multi-Agent Debate Framework
Algorithm 1 illustrates the detailed process of MAD. Generally, our MAD framework is composed of three components which are elaborated as follows:
We use meta prompts to introduce the topic to be solved, the number of debaters, the iteration limit, and other requirements. For example, we require the agents to “tit for tat” so as to create an atmosphere of debate.
Debaters.
There are debaters involved in the framework. In each debate iteration, the debaters speak one by one in a fixed order and express their arguments based on the previous debate history , i.e., . An example of a debater prompt appears below:
You are a debater. Hello and welcome to the translation competition, which will be conducted in a debate format. It’s not necessary to fully agree with each other ’ s perspectives, as our objective is to find the correct translation.
Judge.
We also design a judge to manage and monitor the whole debate process. The judge contains two different modes: (a) Discrinative Mode, in which the judge decides whether the correct solution can be obtained after all the debaters finish their arguments in the current iteration:
If it is True, the debate is over. Otherwise, the debate continues. (b) Extractive Mode, in which the judge needs to extract the final solution based on the whole debate history: , since no correct solution is identified within the iteration limit of debate. An example of a judge prompt appears below:
You are a moderator. There will be two debaters involved in a translation debate competition. They will present their translations and discuss their perspectives on the correct English translation of the given Chinese text: "吃掉敌人一个师。". At the end of each round, you will evaluate the candidates’ translation submissions.
Challenging Testbeds
We conduct experiments on two challenging tasks, namely, commonsense machine translation (i.e., Common MT), and counter-intuitive arithmetic reasoning (i.e., Counter-Intuitive AR), which require deep levels of contemplation for LLMs.
The Common MT dataset is composed of ChineseEnglish translation examples He et al. (2020), which are used to examine the ambiguity resolution ability of translation models. Within the challenging part of Common MT, each source sentence contains an ambiguous word. While these ambiguous words might appear to have a straightforward translation, such a literal interpretation is erroneous. Failure to identify and address such ambiguities may result in inaccurate translations. In this work, we adopt the lexical ambiguity test set in the following experiment. Table 1 lists an example, where the source word “吃掉” should be translated to “destroy” rather than the straightforward translation “eat up” by considering the common sense in the real world.
2 Counter-Intuitive Arithmetic Reasoning
Previous studies on thinking hierarchy Kong et al. (2022); Wei et al. (2022) suggest that we humans have a fast and intuitive system and a slow and logical system, and tend to run the lower level system before the higher level one. Inspired by this, we created a more challenging dataset named Counter-Intuitive Arithmetic Reasoning (Counter-Intuitive AR) to evaluate the reasoning abilities of LLMs at deep levels.
Our Counter-Intuitive AR dataset contains 50 questions collected from elicitation questions Kong et al. (2022)https://elicitation.info/questionnaire/1/, web datahttps://www.geeksforgeeks.org/puzzles/ and manual collection. Compared to the commonly-used datasets, e.g., MultiArith Roy and Roth (2015), GSM8K Cobbe et al. (2021), our dataset presents two distinct challenges:
Resistance to Intuition. The questions in our dataset are embedded in hidden traps designed to elicit intuitive and appealing answers that are often incorrect. This feature evaluates the abilities of LLMs to resist the traps of superficial expressions.
Multi-Step Reasoning. Each correct answer within the dataset requires a rigorous multi-step reasoning process, thereby evaluating the capacity of LLMs to engage in complex decision-making and problem-solving.
Dataset Format.
In our Counter-Intuitive AR dataset, each example contains three key components (see Table 2 for an example). We elaborate on the details below:
Questions. The questions in our dataset are designed to stimulate counter-intuitive thinking, which aims to challenge conventional decision-making by presenting situations where the immediate, intuitive response is often incorrect.
Answers. Each question is provided with a correct answer, which requires deep comprehension of the question and commonsense knowledge. Additionally, we also provide a plausible yet incorrect answer for comparison.
Explanations. We provide a detailed explanation for each correct answer. The explanation outlines the step-by-step reasoning process that leads to the correct answer. Each incorrect answer is also complemented by an explanation demonstrating a seemingly logical reasoning process but ultimately leading to the incorrect answer. This reasoning process highlights the potential pitfalls and misconceptions during decision-making, especially when intuition is prioritized over rigorous logical reasoning.
Experiment
In this work, we mainly use three agents in our MAD framework, including two debaters (i.e., affirmative and negative) and a judge. Unless other stated, we use GPT-3.5-Turbo as the backbone model for all agents by default.
Compared Methods.
Generally, we compare our MAD framework with GPT-3.5-Turbo, GPT-4, and Self-Reflect on both tasks. We also include other baseline methods individually, namely, Rerank and MAPS for Common MT, CoT and Self-Consistency for Counter-Intuitive AR. Below elaborates the details of them:
Self-Reflect Shinn et al. (2023): This approach requires the LLM to scrutinize and refine its translation until it deems the current output satisfactory.
Rerank He et al. (2023): We sample the translations from the LLM for four times, from which we select the best candidate based on a quality estimation (QE) scorerWe use wmt21-comet-qe-da as the QE scorer.. This approach can be seen as analogous to self-consistency Wang et al. (2022), where the majority voting is replaced by an external QE scorer.
MAPS He et al. (2023): This method enables the LLM to mimic the human translation process: analyze and then translate, which can be viewed as a chain-of-thought method applied to translation task.
CoT Kojima et al. (2022): This approach concatenates a trigger sentence “Let’s think step by ste” to the test question.
Self-Consistency Wang et al. (2022): This method samples multiple responses from LLMs and determines the final answer through a majority vote.
We implement the methods on top of GPT-3.5-Turbo. The implementation details are described in Appendix A.1.
Evaluation Metrics.
For Counter-Intuitive AR, we report the accuracy (ACC) of predictions. For Common MT, we adopt automatic metrics like COMEThttps://github.com/Unbabel/COMET/, Unbabel/wmt22-cometkiwi-da and BLEURThttps://github.com/google-research/bleurt, BLEURT-20, which are widely adopted evaluation metrics for LLM-based translation literature He et al. (2023); Hendy et al. (2023); Garcia et al. (2023); Pilault et al. (2023). Moreover, we also employ human evaluation for the translation results in terms of two aspects: ambiguity resolution accuracy and direct assessment of translation quality in range $$.
2 Common MT
Table 3 presents the experimental results. MAPS and Self-Reflec achieve improvements over baseline GPT-3.5-Turbo. Remarkably, our proposed MAD, by utilizing GPT-3.5 as the backbone model, has demonstrated significant advancements over GPT-4 across both automatic and human evaluation metrics.
Case Study.
Table 4 shows example translations generated by baseline GPT-3.5-Turbo and the proposed MAD. We can find that the baseline GPT-3.5-Turbo (even the more powerful GPT-4) incorrectly translates the source words literally. Because of the DoT issue, Self-Reflect cannot rectify the literal translation. The proposed MAD framework, which explores divergent chain-of-thoughts, can generate the free translation of the underlined words within the source sentences. The detailed debate process of translation examples can be found in Appendix A.2.
3 Counter-Intuitive AR
Table 5 lists the experimental results in terms of reasoning accuracy. We can observe that Self-Reflect does not improve over the baseline GPT-3.5-Turbo, while CoT and Self-Consistency bring some improvements. Our MAD framework, though not as good as GPT-4, outperforms all the other compared methods based on GPT-3.5-Turbo, which further demonstrates its effectiveness.
Case Study
Table 6 presents two example outputs on Counter-Intuitive AR. We find both CoT and Self-Reflect fail to reach the right answer. With divergent thinking, our MAD framework emerges “we need to consider both the rotation around circle B and the rotation of circle A itself” and find the correct answer. The detailed debate process can be found in Appendix A.2.
Analysis
We conduct extensive analyses to gain a deeper understanding on our MAD framework. By default, we use the Common MT dataset.
We first investigate the stopping strategy of debate. For each iteration, we force the judge to extract the final answer () instead of adaptively breaking the debate as in Algorithm 1. Figure 3 shows the results. We can observe that MAD performs better than self-reflection as the iteration increases. However, the highest COMET score appears at the first iteration and is also lower than the result of the adaptive break. It indicates that, for most examples, MAD can generate good translations at the first iteration such that the debate should be stopped. Forcing the debate to continue will harm the translation results, which demonstrates the reasonableness of our adaptive break strategy.
Essense of “Tit for Tat” State.
We then study how the intensity of “tit for tat” affects the performance of MAD. To achieve so, we design different prompts (see Appendix B) to initialize the debate process. As shown in Figure 4, asking the debaters to “tit for tat” (i.e., higher disagreement) is necessary for MAD to achieve good performance. However, we find that “must disagree with each other on every point ” (with a disagreement of 0.988) does not lead to the best performance. We speculate that continuous disagreement without finding common ground can contribute to polarization, where the debate becomes more about winning the argument than seeking truth or understanding. This can reinforce pre-existing biases and make it difficult to reach a consensus or meaningful decision.
Behavior of Agents.
We study the behavior of agents by calculating how many times the judge chooses the answers of each debater as the final solution. The results are listed in Table 7 and we have the following observations: (1) Comparing row \small{1}⃝ and \small{2}⃝, we find that the judge consistently favors the negative side, which is believed to contribute significantly to the performance improvement in MAD. When encountering complex tasks, the affirmative side tends to make mistakes that should be corrected to achieve improvements. (2) Comparing row \small{3}⃝ and \small{4}⃝ (or row \small{4}⃝ and \small{5}⃝), we find the judge shows a preference to the side with the same LLM as the backbone. This bias indicates that LLMs might not be a fair judge Wang et al. (2023) when different LLMs are used for the agents.
Related Work
Recently, Wei et al. (2022) has proposed chain-of-thought (CoT) prompting to improve the reasoning ability of LLMs. Specifically, CoT prompts LLMs to generate a series of intermediate steps that lead to the final answer of a multi-step problem. Most earlier work primarily concentrates on two main aspects: prompt design and decoding strategies. Zero-shot CoT Kojima et al. (2022) employs the trigger sentence “Let’s think step by step” to provide guidance for the decoding of LLMs. Advanced sampling strategies have been explored to improve CoT by generating diverse reasoning paths, e.g., Self-Consistency Wang et al. (2022), Auto-CoT Zhang et al. (2022), Active-Prompting Diao et al. (2023), Complexity-based Consistency Fu et al. (2022), Multi-Chain Reasoning Yoran et al. (2023), and Progressive-Hint Prompting Zheng et al. (2023).
With the emergence of powerful LLMs, approaches based on self-evaluation have attracted increasing attention. These approaches involve the generation of initial output, followed by evaluating the output to acquire feedback, which is then utilized to refine the output. Evaluation feedback can come from the model itself, e.g., Self-refine Madaan et al. (2023) and Tree of Thoughts Yao et al. (2023)) or external environments, e.g., QAaP Zhu et al. (2023a) and Reflection Shinn et al. (2023). The intuition behind these approaches involves the utilization of robust LLMs to mimic the human cognition process.
Generative Agents
Recently, LLMs-based multi-agent intelligent, e.g., Generative Agents Park et al. (2023), Ghost in the Minecraft Zhu et al. (2023b), GPT-Bargaining Fu et al. (2023), has drawn significant attention for enabling simulations of human behavior. Our work follows this research line to address the DoT problem of LLMs. Concurrent with our work, a few studies Xiong et al. (2023); Du et al. (2023) also explore the multi-agent debate framework to enhance the reasoning ability of LLMs. The main differences between the proposed MAD framework and these approaches are: (1) our work aims to address the DoT problem, which is an inherent deficiency of LLMs; and (2) we empirically find that our MAD framework can yield enhanced performance by employing agents with the identical backbone LLM.
Conclusion
We propose and define the Degeneration-of-Thought (DoT) problem in self-reflection, and address it by proposing the Multi-Agent Debate (MAD) framework to explore divergent chain-of-thoughts. We demonstrate the effectiveness of MAD on two challenging tasks and find that GPT-3.5-Turbo with MAD can even surpass GPT-4 on the Common MT dataset. Extensive analyses suggest that the adaptive break strategy of debate and the modest level of “tit for tat” state are required for MAD to obtain good performance. More interestingly, we find that LLMs might not be a fair judge if different LLMs are used for agents. Future works may include scheduling more agents in the debate, multi-agents for board games, and AI feedback for model alignment.
References
Appendix A Example Appendix
Figure 5 displays a typical template of debate history, formatted according to the Turbo API.
A.2 Debate Case
Table 9 and Table 10 present the debate process of example translations in Section 4.2. Table 10 and Table 11 show the debate process of example answers in Section 4.3.
We observe that the affirmative side ( ) often relies on direct intuition, which can lead to incorrect or inappropriate responses. Conversely, the negative side ( ) demonstrates an ability to identify and rectify his mistakes.
Appendix B Level Control of “tit for tat” State
We modulate the level of “tit for tat” state outlined in Section 5 through appending natural language instructions to the debaters’ meta prompt. All the corresponding prompts are itemized in Table 12.