ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, Zhiyuan Liu

Introduction

Evaluating the quality of text generated by language models or written by humans has long been a challenging endeavor, consistently garnering substantial attention (Celikyilmaz et al., 2020). Traditional methodologies predominantly rely on human annotation of texts (Callison-Burch, 2009), an approach considered overly demanding in terms of time and cost. Automatic evaluation metrics based on n-grams, such as Rouge (Lin, 2004), BLEU (Papineni et al., 2002), and METEOR (Banerjee & Lavie, 2005), have been proposed to tackle this issue (Kondrak, 2005). However, these methods have been shown to exhibit a relatively weak correlation with human judgments, particularly in the context of tasks involving open-ended generation or requiring domain-specific expertise (Novikova et al., 2017).

Recent advancements in the field of natural language processing have led to the emergence of billion-parameter scale LLMs, such as GPT-3 (Brown et al., 2020). These LLMs have demonstrated remarkable capabilities across diverse downstream tasks, presenting new opportunities for text quality evaluation using such models. Moreover, various training paradigms have been proposed to endow LLMs with the ability to accomplish tasks in a zero-shot manner and better adhere to human-provided instructions (Ouyang et al., 2022; Sanh et al., 2021; Wei et al., 2021). These advancements facilitate the prompting of LLMs to evaluate generated text, effectively simulating human evaluators in the assessment process.

In view of the impressive text understanding and instruction-following capabilities of recent LLMs, a body of literature (Liu et al., 2023b; Chiang & Lee, 2023; Gao et al., 2023; Shen et al., 2023) has adopted LLM as an evaluator to assess the quality of responses to open-ended questions or traditional NLG tasks, including dialogue response generation and summarization. This methodology is dubbed LLM-as-a-judge (Zheng et al., 2023). Findings from these researches indicate that LLM can mimic human behavior and provide evaluations that correspond with human judgments, revealing a potentially scalable and transparent alternative to costly and laborious human evaluations.

While a single powerful LLM can already tackle various missions, emerging studies suggest that multiple LLMs can further improve one another through debate and cooperation (Li et al., 2023a; Liang et al., 2023). By incorporating multiple LLMs into an integrated group and designing specific interaction mechanisms, different LLMs can engage in proposing and deliberating unique responses and thought processes across several rounds. This approach leads to enhanced factuality of generated responses (Du et al., 2023) and improvement in the completion of arduous tasks (Li et al., 2023a; Qian et al., 2023). Furthermore, the multi-agent group also addresses and mitigates the Degeneration-of-Thought (DOT) problem (Liang et al., 2023).

In the human evaluation processes, relying on a single perspective can introduce bias and instability in the results (Karpinska et al., 2021). Recognizing this, best practices often involve multiple human annotators collaborating in the evaluation (Van Der Lee et al., 2019). Drawing inspiration from this collaborative and iterative human evaluation approach, we propose ChatEval, a system that enables each agent to employ varied communication strategies in collaborative discussion, working towards formulating final judgments. Furthermore, to enrich the evaluation dynamics, every agent within ChatEval is endowed with a unique persona. This deliberate design ensures that each agent focuses on distinct perspectives or brings specific expertise to the table. By doing so, the collective evaluation benefits from a more comprehensive lens, capturing nuances and subtleties that a single perspective might overlook. We derive this idea primarily from the insight of ’There are a thousand Hamlets in a thousand people’s eyes’, meaning that every person has their unique interpretation or perspective, especially applicable to text evaluation. Indeed, these divergent perspectives shape the comprehensive and multifaceted assessment of Hamlet. Another underlying intuition of our work stems from renowned concepts in sociology and biology, including Collective Intelligence(Woolley et al., 2010) and Cognitive Synergy(Luppi et al., 2022), where multiple cognitive processes or systems interact and cooperate in a way that produces a combined effect greater than the sum of their separate effects.

To summarize, the main contribution of our work is as follows:

We propose a multi-agent-based framework called ChatEval that aligns better with human preferences compared with single-agent-based approaches as depicted in Figure 1.

We propose various communication strategies and demonstrate the necessity of diverse role prompts in multi-agent debate scenarios.

We release our library. It’s designed to be both composable and scalable, enabling researchers to implement their unique communication strategies easily. We hope this contributes to advancing research in the field of communicative agents and beyond.

Methodology

In this section, we elaborate on the principal components in ChatEval including debater agents, diverse role specification, communication strategy, and provide a detailed overview of each component’s role and functionalityour code repository is built on top of https://github.com/OpenBMB/AgentVerse..

Debater Agents. Debater agents are one of the most significant components in our framework. We treat each individual LLM as an agent and ask them to generate their response from the given promptThe full prompt template can be found in Appendix A.. Responses from other agents are served as chat history which will be replaced in the prompt template. After configuring the agents, we then start the group debate where each agent autonomously receives responses from the others and, in turn, delivers its own responses to them. It should be noted that the whole process does not require human intervention.

Diverse Role Specification. As presented in Section 1, diverse role specification is necessary for the framework as well. Although all the agents share a common prompt template, we substitute the role_description slot with diverse role prompts, specifying distinct personalities for different agents. We take inspiration from Wu et al. (2023) and formulate an analogous role description.

Communication Strategy. How to maintain the chat history is another significant issue in ChatEval. In our work, we use a more intuitive term to illustrate the maintenance of the chat history called communication strategy. In a nutshell, different communication strategies can be seen as different approaches to maintaining and manipulating their chat history. As is shown in Figure 2, We primarily design three different communication strategies and illustrate them as follows:

One-By-One. During each round of the debate, the debater agents take turns in a set order to generate their response based on the current observation. When it’s time for a debater agent to respond, we directly concatenate what previous other agents have said into its chat history slot.

Simultaneous-Talk. Unlike the one-by-one strategy, we carry out an alternative communication strategy called simultaneous-talk, where debater agents are prompted to asynchronously generate responses in each iteration of the discussion to nullify the impact of the speaking order.

Simultaneous-Talk-with-Summarizer. The main difference between this strategy and simultaneous-talk is that we additionally employ another LLM as a summarizer. At the end of each iteration of the debate, we prompt this extra LLM to summarize the messages conveyed so far and concatenate this summarization into all debater agents’ chat history slots.

Unlike previous work like Du et al. (2023), we do not explicitly ask the debater agents to reach a consensus at the end of the debate. In situations where the response format relies on direct comparison, we derive the final results from the majority vote among various annotators. Conversely, if the response format requires a direct score, we calculate the average score obtained from multiple annotators. This methodological approach ensures the impartiality and balance of our evaluation process.

Experiments

We evaluate ChatEval on two benchmarks, FairEval and Topical-Chat which represent the categories of open-ended question answer and dialogue response generation, respectively.

We choose to utilize models from OpenAI’s GPT family as our LLMs in ChatEval, including GPT-4 and ChatGPT (GPT-3.5-turbo) and set the temperature to 0 to ensure reproducibility. The rationale behind this selection is the exceptional performance these models offer, being among the most advanced and powerful in the world. Additionally, their accessibility and ease of use through APIs enable us to directly call and interact with the models during our research, significantly simplifying the process. In our current research, we focus on homogeneous groups of LLMs. That is, within a given multi-agent group, all LLMs belong to the same GPT family model, either all GPT-4 or all ChatGPT. We acknowledge the potential of heterogeneous groups for future research, which could provide fascinating insights into how strong models and weak models can cooperate in a multi-agent setting.

2 Benchmarks

The detailed introduction of different categories and benchmarks are listed as follows:

Open-ended Question Answer is a key component within the field of NLP and generative AI. It necessitates an AI system to provide comprehensive, detailed, and human-like responses to questions that don’t have a predefined or fixed set of possible answers. The work of Chiang et al. (2023) encompasses a collection of 80 open-ended questions originating from a wide array of categories, including common-sense, counterfactual, coding, etc. We then take the human annotation results from Wu et al. (2023) to conduct the experiments in this paper. For each question, they direct three annotators to evaluate the replies given by Vicuna-13B and ChatGPT through the given rules and finally derive the results by the majority votes among the annotators.

Dialogue Response Generation is a task involves creating a coherent and contextually appropriate response to a given input dialogue. We draw upon the Topical-Chat (Gopalakrishnan et al., 2019) dataset for our study. We then take the human annotation results from Mehri & Eskenazi (2020) where they carry out the annotations on 60 dialogue contexts with each response generated by 6 different systems. Human evaluators analyzed these responses based on natural, coherence, engagingness, groundedness, and understandable, where we take the first four dimensions for experiments in our paper following Zhong et al. (2022).

3 Baselines

We evaluate ChatEval against following methods. As the main portion of our comparison, we primarily focuses on the single-agent-based method. Single-Agent means that we directly query an LLM to generate the response towards the evaluationWe use the same prompt template as our multi-agent debate settings in single-agent baseline except that we ignore some slot.. We use Multi-Agent to represent ChatEval where several agents discuss towards the evaluation. By default, we configure the communication strategy to one-by-one, agent numbers to 2, and discussion turns to 2 in this section and employ position calibration techniques in both single-agent and multi-agent settings. We will discuss more debate configurations in Section 4 for completeness. For the open-ended question answer task, we also compare our method with FairEval (Wang et al., 2023b). They propose various strategies to improve the evaluation performance of a LLM including Multiple Evidence Calibration (MEC) and Balanced Position Calibration (BPC). For the dialogue response generation task, we also compare our method with G-EVAL (Liu et al., 2023b). They utilize CoT and probability-weighted summation for their method. Additionally, we include results from n-gram-based metrics, such as ROUGE (Lin, 2004), BLEU (Papineni et al., 2002) and embedding-based metrics such as BERTScore (Zhang et al., 2019).

4 Results for Open-ended question answers

We adopt the same evaluation approach as Wang et al. (2023b) to assess the annotation results produced by different methods and annotators. Specifically, we calculate the Accuracy (Acc.), which measures the proportion of correctly classified instances out of the total instances, and the Kappa correlation coefficient (Kap.) (McHugh, 2012) which gauges the agreement between results from models and human annotators while taking into account the possibility of agreement occurring by chance. Both metrics provide insights into the reliability and consistency of the annotations. We take the human annotation results and FairEval’s (Wang et al., 2023b) best results from their paper. As is shown in Table 1, different annotators can reach a relatively high agreement and perform better than any other LLM-based approach. Still, the average human annotations accuracy which is 71.7% shows there exists a certain degree of discrepancy among different unique individuals revealing that text evaluation is absolutely an arduous task. The second part and the third part of Table 1 show the results of FairEval’s method and the results of our proposed method respectively. We find that (1) ChatEval can enhance the performance of the evaluation process, achieving higher alignment with human preference compared with single-agent evaluation. Specifically, the multi-agent-based method improves the accuracy by 6.2% for ChatGPT and 2.5% for GPT-4; (2) ChatEval surpasses FairEval’s best results within both ChatGPT and GPT-4 settings showing the effectiveness of our proposed method.

5 Results for Dialogue Response Generation

For the dialogue response generation benchmarks, we align the evaluation method with Zhong et al. (2022), calculating the turn-level Spearman and Kendall-Tau correlation in correspondence with human judgments on four aspects (naturalness, coherence, engagingness and groundedness). Results can be found in Table 2. In the first part of Table 2, we demonstrate that n-gram-based metrics and embedding-based metrics perform overall poorly on all the aspects evaluated illustrating that these methods can hardly reveal human preference. In the second part of Table 2, we show the results from the G-eval (Liu et al., 2023b) paper. They first ask the LLM to generate intermediate thought and finally calculate the weighted summation of the output scores based on the probability. The results show that their method outperforms previous traditional metrics depicting the fact that the LLM-based evaluator is effective and reliable for evaluating the dialogue response generation task. While their method delivers sound results, our proposed approach raises the bar in terms of performance for GPT-4. Specifically, ChatEval improves the average Spearman and Kendall-Tau correlation by 0.096 (16.3%) and 0.057 (10.0%) respectively. Additionally, compared with the single-agent method, ChatEval amplifies the performance both for ChatGPT and GPT-4, showing the effectiveness of our method which is aligned with the results in Section 3.4.

Analysis

In this section, we further explore the key components encompassed in ChatEval. We discuss the importance of diverse role prompts in Section 4.1, the effect of different communication strategies in Section 4.2, and the impact of role numbers and discussion turns in Section 4.3. If not specified otherwise, we choose the FairEval benchmark and ChatGPT as the backbone LLM for the analysis.

Previously in Table 1 and 2, we demonstrate that ChatEval equipped with diverse role configurations can significantly improve the performance of evaluation. We further consider whether it is necessary to design diverse role prompts for the evaluation system. To answer so, we carry out the experiments by replacing all the role prompt with ”You are now an Annotator, one of the referees in the text evaluation task.” and keeping other prompt unchanged. We experiment with the one-by-one communication strategy and 2 agents with 2 discussion turns. The results in Table 3 illustrate that ChatEval with the same role prompt design underperforms that with diverse role prompt design and cannot effectively enhance the performance compared with single-agent setting, highlighting the cruciality of diverse role prompt design in the multi-agent debate framework.

2 The study of communication strategies

As shown in Figure 2, we also design three different communication strategy termed as one-by-one, simultaneous-talk, simultaneous-talk-with-summarizer. The detailed descriptions and formal formulations can be found in Appendix B. We experiment with 3 agents and 2 discussion turns with diverse role prompts in this section. As is shown in Table 4, we can find that the one-by-one communication strategy is more effective than other strategies for ChatGPT setting. Although the other two communication strategies did not perform as robustly as the one-by-one strategy, it is noteworthy that they still exceeded the performance of the naive single-agent method. Furthermore, the variations in performance among three different communication strategies underscore the influence of different strategies on the effectiveness of the evaluation process, revealing the potential for further exploration and optimization of ChatEval. Thus, future studies could be aimed at a more comprehensive understanding of different communication strategies, and how they could be effectively employed to enhance performance. This could serve as an avenue for substantial improvements and novel insights in the multi-agent debate framework.

3 The impact of role numbers and discussion turns

We then study the impact of different role numbers and discussion turns. From Figure 3(a), a discernible trend is observed in the relationship between the role number and both Acc. and Kap. within the FairEval dataset. As the role number increases, there is a corresponding growth in Acc. and Kap. This pattern reaches an apex with an Acc. of 62.5% at role numbers 3 and 4 before declining at role number 5. Despite this decline in accuracy, Kap. continues to ascend between role numbers 3 and 4, underscoring the effectiveness of incorporating diverse roles within ChatEval.

Conversely, no significant upward trend is detected with respect to the increase in discussion turns, as is shown in Figure 3(b). This observation aligns with the findings in Liang et al. (2023); Du et al. (2023), highlighting a consistent phenomenon where continual discussion often leads to stagnation or even degradation of performance. Such a trend may be attributed to issues associated with the ever-increasing context length, which consequently diminishes the performance. This insight prompts a more nuanced understanding of the balance needed between role differentiation and discussion dynamics to optimize the performance of ChatEval.

4 qualitative analysis

Table 5 showcases the debate process towards the evaluation of two assistants’ responses to the open-ended question ”What are the most effective ways to deal with stress?”. Assistant 1’s response is :

Here are some effective ways to deal with stress: 1. Exercise regularly: Physical activity can help reduce stress and improve mood. 2. Practice relaxation techniques: Deep breathing, meditation, and yoga are some effective ways to relax and reduce stress levels. 3. Get enough sleep: Lack of sleep can increase stress levels. Try to get at least 7-8 hours of sleep each night. 4. Eat a healthy diet: A well-balanced diet can help reduce stress levels and improve overall health. 5. Connect with others: Spending time with friends and family, joining a support group, or participating in social activities can help reduce stress levels and improve mood. 6. Time management: Prioritizing tasks and creating a schedule can help reduce stress and increase productivity. 7. Seek professional help: If stress is impacting your daily life, consider seeking professional help such as therapy or counseling.

Stress can be caused by a variety of factors, including work, relationships, health, and financial issues. Here are some effective ways to manage stress: 1. Exercise: Regular exercise releases endorphins, which are natural stress-relievers. It can be as simple as going for a walk or jog, practicing yoga, or joining a fitness class. 2. Mindfulness: Mindfulness practices, such as meditation or deep breathing, can help to reduce stress and improve focus. 3. Social Support: Spending time with family and friends, or seeking support from a mental health professional, can provide a sense of connection and help to reduce feelings of isolation and loneliness. 4. Healthy Eating: A balanced diet that includes plenty of fruits, vegetables, and whole grains can help to reduce stress levels. 5. Good Sleep: Getting enough sleep is essential for good mental and physical health. Aim for 7-8 hours per night and establish a regular sleep routine. 6. Time Management: Prioritizing tasks and setting realistic goals can help to reduce stress levels by giving a sense of control and direction. 7. Relaxation Techniques: Techniques such as progressive muscle relaxation, aromatherapy, and massage can help to reduce stress levels and promote relaxation. 8. Seek professional help: If stress is interfering with daily life, it may be time to seek professional help from a therapist or counselor.

We can find that both of the responses produce similar strategies and equally compelling descriptions for dealing with stress, making it challenging to discern significant disparity in terms of quality. It is in this context of nuanced evaluation that the significance of the ChatEval process emerges. To understand this complexity better, We first outline the ChatEval process and subsequently delve into the agents’ constructive behaviors during discussions.

As is shown in Table 5, Alice first points out that the response of Assistant 2 contains more detailed information and he prefers to choose Assistant 2 as a better response. Bob then says that she agrees with Alice’s assessments, but in the meantime, she also points out that Assistant 1’s response is also concise and carries out a thought-provoking question. Carol then gives the feedback that she believes both responses are equally valuable. In the subsequent discussion, Bob indicates that Assistant 1’s response is straightforward while Assistant 2’s is detailed, suggesting that the effectiveness of the response should depend on the context and individual’s needs. At the end of the debate, we finally extract the evaluation results that both responses are of the same quality which is identical to human annotation results.

From this sequence, we can pinpoint several fascinating behaviors exhibited by the agents: (1) Opening Statement: Alice initiates the debate with a clear stance, establishing the foundational argument and guiding the trajectory of the subsequent discourse. (2) Alternative Proposal: Bob introduces an alternative viewpoint, emphasizing the need to consider diverse interpretations. This not only broadens the discussion but also stimulates critical thinking. In the context of a debate, the introduction of an alternative proposal prevents the stagnation of thought, challenges pre-existing bias, and uncovers considerations that might otherwise be overlooked, ensuring that the discussions are well-rounded. (3) Stance Maintenance: Alice’s persistent adherence to her initial stance, even when faced with opposing views, exemplifies commitment and challenges other participants to refine their perspectives. By firmly holding his position, Alice encourages depth in the discourse, prompting others to dive deeper into their arguments and perhaps consider aspects they hadn’t previously. It ensures the conversation remains robust, focused, and continually evolving, driving all participants to a higher level of engagement and critical thinking. (4) Seeking Consensus: The discussion’s climax reveals a collective agreement amongst the participants, which is reached through mutual understanding and compromise, underlining the value of each presented viewpoint.

In light of the above, ChatEval stands out not just as a tool for comparison but as an embodiment of interactive natural language dialogue. By simulating human argumentative interactions, it differentiates itself from static, single-presented opinions. This dynamic interaction showcases the richness and complexity of language, capturing nuances often missed in singular viewpoints. As such, ChatEval offers a reliable evaluation process that not only mirrors human discourse but also highlights the transformative power of collaborative dialogue. This positions it uniquely, underscoring its significant potential to execute text evaluation tasks both reliably and effectively.

Related Work

Automatic NLG evaluation In the landscape of NLG, evaluating the quality of generated text represents a particularly arduous task. For a significant period, evaluation was primarily dependent on human annotations, a process that is labor-intensive and limited by scalability issues. Automatic NLG evaluation attempts to address these challenges by leveraging computational models to assess the quality of a generated text. Previous work lies on the following categories: (1) n-gram-based metrics: ROUGE (Lin, 2004) is a set of metrics that compute the amount of overlap between n-grams in the machine-generated summaries and the reference summaries. BLEU (Papineni et al., 2002) compare the generated text with reference translations, based on the co-occurrence of n-grams in both texts. In spite of being easily and widely used, the above method is incapable of capturing syntactic and semantic similarity (Stent et al., 2005). (2) embedding-based metrics: Word embeddings are vector representations of words that capture their semantic properties, such that words with similar meanings have similar embeddings. A bunch of work leverages word embeddings to evaluate the semantic similarity between two pieces of text. BERTScore (Zhang et al., 2019) use contextualized word embeddings from transformer models like BERT (Devlin et al., 2018), BLEURT (Sellam et al., 2020) utilize supervised training data to enhance the performance. MoverScore (Zhao et al., 2019) combine contextualized word embeddings with Earth Mover’s Distance (Rubner et al., 2000). (3) LLM-based metrics: Amidst the flourishing advancement of LLM which embodies a wealth of information derived from extensive training data, using LLM as an evaluator has experienced notable progress. GPTScore (Fu et al., 2023) utilize conditional probability to assign the text a score representing its quality. Wang et al. (2023a) explore the potential of utilizing ChatGPT as an NLG evaluator by prompting it to score a text directly. Wang et al. (2023c) curate a reliable dataset containing pairwise comparison and evaluation explanation which can be used to train a foundation model making it a better evaluator. Bai et al. (2023) propose decentralized evaluation to provide fairer evaluation results. G-EVAL (Liu et al., 2023b) propose probability-weighted techniques to calibrate the score given by a single LLM.

Communicative Agents Most recently, significant attention has been dedicated to the development of communicative agents. These agents, often acted by LLMs like ChatGPT or GPT-4, are designed to interact and communicate effectively with other agents or human users using natural language. The primary goal is to facilitate more productive and efficient interaction and collaboration as different agents can autonomously communicate and negotiate to tackle a more complex task collectively. Several studies have explored various aspects of communicative agents. Li et al. (2023a) propose a cooperative agent framework dubbed as role-playing enabling agents to autonomously cooperate to solve complex tasks. Park et al. (2023) create a sandbox environment consisting of 25 individual virtual entities endowed with a character description and memory system. Every intelligent agent is capable of autonomously interacting with other agents and the environment simulating reliable human behavior. Qian et al. (2023) establish a chat-based software development framework that can complete a software design and produce executable software at a reduced cost compared to recruiting human programmers. Liu et al. (2023a) utilize a sandbox environment to curate reliable datasets in better alignment with human preference and train a socially-aligned LLM. Liang et al. (2023) and Du et al. (2023) also make use of the multi-agent debate framework in other scenarios such as translation and arithmetic problems resulting in better results. Wang et al. (2023d) propose an alternative method called self-collaboration to enable the communication of agents by utilizing a single LLM prompted by multi-persona descriptions. Mandi et al. (2023) propose a novel framework designed for the collaboration of multiple robots, utilizing multiple LLMs to enhance coordination and strategic planning among the robots. Concurrent with our work, Li et al. (2023b) propose Peer Rank and Discussion (PRD) which is similar to our approach. However, they probe the different dimensions of evaluation by using different models as agents and did not explore alternative communication strategies.

Conclusion

In this paper, we present evidence that ChatEval contributes to improving the evaluation performance concerning text quality, aligning more closely with human preferences. We emphasize the necessity of the diverse role specification and propose distinct communication strategies as integral components within ChatEval. Our qualitative analysis of the discussion process conveys insightful intuitions about how a text is evaluated by ChatEval and substantiates our approach’s ability to support comprehensive evaluations akin to human judgment, thereby demonstrating the reliability and efficacy of our framework.

References

Appendix A prompt template and diverse role prompt

The overall prompt template is shown in Table 6, we draw inspiration from Wu et al. (2023) and design several different role descriptions as follows.

General Public You are now General Public, one of the referees in this task. You are interested in the story and looking for updates on the investigation. Please think critically by yourself and note that it’s your responsibility to choose one of which is the better first.

Critic You are now Critic, one of the referees in this task. You will check fluent writing, clear sentences, and good wording in summary writing. Your job is to question others judgment to make sure their judgment is well-considered and offer an alternative solution if two responses are at the same level.

News Author You are News Author, one of the referees in this task. You will focus on the consistency with the original article. Please help other people to determine which response is the better one.

Psychologist You are Psychologist, one of the referees in this task. You will study human behavior and mental processes in order to understand and explain human behavior. Please help other people to determine which response is the better one.

Scientist You are Scientist, one of the referees in this task. You are a professional engaged in systematic study who possesses a strong background in the scientific method, critical thinking, and problem-solving abilities. Please help other people to determine which response is the better one.

Appendix B formal depiction of different communication strategy