Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
Rishav Hada, Varun Gumma, Adrian de Wynter, Harshita Diddee, Mohamed Ahmed, Monojit Choudhury, Kalika Bali, Sunayana Sitaram
Introduction
Large Language Models (LLMs) can achieve remarkable results on a variety of tasks, sometimes even outperforming humans on certain tasks and domains OpenAI (2023); Chen and Ding (2023); Veen et al. (2023); Chiang and Lee (2023). However, measuring the performance of LLMs is challenging, as standard NLP benchmarks may not reflect real-world applications. Other hurdles for LLM evaluation include the scarcity of benchmarks for diverse and complex tasks, benchmark saturation, contamination of benchmark data in LLM training data, and the weak correlation between automated metrics and human judgment Jacovi et al. (2023); Chang et al. (2023); Reiter (2018); Liu and Liu (2008). Therefore, researchers have proposed alternative evaluation methods that go beyond benchmarking to assess the abilities and limitations of LLMs Chang et al. (2023).
While LLMs excel at various tasks in English, their capabilities in other languages are more limited. This disparity may increase the digital divide, preventing a significant portion of the global population from benefiting from LLMs and potentially harming them. Ahuja et al. (2023a, b) conduct a comprehensive benchmarking of LLMs across the available multilingual benchmarks covering several tasks and languages, and show that the performance of LLMs degrades significantly on languages that are transcribed in non-Latin scripts and under-resourced languages.
Multilingual evaluation is challenging to scale. Certain language families, such as Indo-European, are over-represented in multilingual benchmarks with other language families having very little presence. There is a scarcity of multilingual benchmarks designed to assess tasks that simulate actual LLM usage in real-world scenarios. The metrics used in these benchmarks may be unsuitable for languages with rich morphology or complex writing systems, as well as phenomena arising from language contact such as borrowing, code-mixing, and transliteration. Evaluation by native speakers is the gold standard for building an accurate picture of model performance, especially in complex tasks without well-defined automated metrics. However, budget constraints, turnaround time, and the lack of easy access to native speakers in some languages all pose challenges in scaling evaluation. This leads to a situation in which LLM performance is unknown for most languages of the world Ahuja et al. (2022).
The success of LLMs in complex tasks such as sentiment analysis, reasoning, problem-solving Mao et al. (2023); Arora et al. (2023), and providing feedback for reducing LLM harms Bai et al. (2022) has led to the question of whether LLMs can replace human annotators, or help augment human evaluation Gilardi et al. (2023). Utilizing LLMs as multilingual evaluators is, therefore, an attractive option to decrease costs and circumvent the challenges of scaling assessments by native speakers. However, LLMs have been demonstrated to have inferior performance even in some high-resource languages and have not been evaluated extensively across many languages on dimensions such as toxicity, fairness, and robustness (due to the absence of such benchmarks) Ahuja et al. (2023a), it is prudent to proceed with caution. Failing to do so can lead to misleading results which may further widen the digital divide.
In this work, we study whether LLM-based evaluation can be the answer to scaling up multilingual evaluation. In other words, can LLMs serve as substitutes or supplements for human native speakers in delivering useful and accurate insights regarding LLM outputs in non-English languages, while considering diverse aspects of interest like linguistic acceptability, task accomplishment, and safety? Our main contributions are as follows:
We present the first evaluation of LLMs, specifically GPT-4 as multilingual evaluators to examine whether LLMs can be used to scale up multilingual evaluation.
We calibrate LLM judgments on an in-house dataset across three tasks, eight languages, and five dimensions by comparing them to over K human judgments on the same tasks, languages, and dimensions.
We evaluate a variety of prompting strategies for LLM-based evaluation in the multilingual setting.
We provide a framework for evaluating LLM-evaluators in the multilingual setting that can generalize across tasks, metrics, and languagesCode available at: https://aka.ms/LLM-Eval.
We suggest best practices and provide recommendations for future work.
Related Work
Broadly, there are two main uses of LLMs as evaluators: LLMs can be used as alternatives to metrics that compare human and machine-generated text, such as BLEU Papineni et al. (2002) and ROUGE Lin (2004). Word overlap-based metrics are limited, and LLM-based scorers have been shown to outperform them. GPTScore Fu et al. (2023) is a popular LLM-based framework that can be used to score model outputs based on human-created references along various dimensions. However, these scores still rely on having examples of human-created reference data.
The second use case of LLMs as evaluators is when the LLM is presented with the output of a system (usually an LLM, sometimes the same model) and asked to judge its quality or safety without any human output to compare against Zheng et al. (2023). The LLM is instructed on how to perform this evaluation with the help of the task description, evaluation rubric, and sometimes, one or more examples in the prompt. This is the use case we focus on in this work.
Gilardi et al. (2023) prompt ChatGPT to annotate Tweets across various dimensions such as topic and stance and find that it outperforms crowdworkers. Shen et al. (2023) explore the use of GPT3.5 as an evaluator for abstractive summarization and find that although GPT is a useful evaluator, as the quality of summarization improves, the quality of evaluation degrades. Along similar lines, Wang et al. (2023a) evaluate ChatGPT on various NLG tasks and find that it has a high correlation with human judgments. Kocmi and Federmann (2023) evaluate the effectiveness of LLMs on evaluation of translation quality and find that LLMs starting from GPT3.5 and above achieve SOTA performance on translation evaluation benchmarks. Fernandes et al. (2023) leverage LLMs for fine-grained annotation of errors in Machine Translation outputs. LLM-based evaluators have also been used to score and refine outputs they produce, as described in Madaan et al. (2023), ultimately producing outputs that are scored higher on human and automated metrics than the original outputs. Naismith et al. (2023) explore the use of LLM-based evaluators on scoring written discourse for coherence and find a strong correlation with human judgments. The success of LLM-based evaluators has led many to question whether LLM-based evaluation can replace or augment human evaluation Chiang and Lee (2023).
However, there have been studies showing that LLM-based evaluators can have some biases. Wu and Aji (2023) demonstrate that LLMs tend to prefer answers with factual errors when they are too short or contain grammatical errors. Pangakis et al. (2023) highlight the need for validating LLM-based evaluators on a task-by-task basis. Liu et al. (2023) perform NLG evaluation using GPT-4 and find that although it correlates well with human judgments, it may potentially be biased towards preferring LLM-generated texts. Koo et al. (2023) show that LLMs have egocentric bias where they prefer to rank their own outputs highly in evaluation. Wang et al. (2023b) point out that GPT4-based evaluators have positional bias and scores can be easily altered by changing the order of appearance. There are also several ethical issues with the use of LLMs as evaluators described in Chiang and Lee (2023). Zhang et al. (2023) suggest that wider and deeper LLMs are fairer evaluators, while Chan et al. (2023) introduce a framework for multiple evaluator agents to reach a consensus, mimicking the situation of having multiple annotators.
Although there has been some work measuring the calibration of LLM-based evaluators to human judgments Koo et al. (2023), previous studies have focused on English, and ours is the first work (to the best of our knowledge) that addresses this problem in the multilingual context.
Experimental Setup
We perform experiments on a text generation application that is powered by GPT-4, and evaluate the following sub-tasks:
Open Prompt: This task processes a concise prompt to generate a document adhering to the provided guidelines, producing up to tokens, approximately equivalent to one page in English or Spanish, and marginally less in other languages. Continue Writing: This task takes two textual inputs, termed “left” and “right” to generate a coherent continuation between them, accommodating up to tokens. Notably, one of the inputs may be omitted. Summarize: Engages in standard summarization by condensing a document of at least words into a succinct summary. It allows for an optional user-defined prompt to tailor the summary format, such as highlighting key points.
We cover the following languages: English (En), French (Fr), German (De), Spanish (Es), Chinese (Zh), Japanese (Ja), Italian (It), Brazilian Portuguese (Pt-Br), and Czech (Cs). Of these, the first six are classified as very high resource languages (Class 5, or “the winners”), while the last three are classified as Class 4 (“the underdogs”) according to Joshi et al. (2020). We plan to extend our study to lower-resource languages in the future. We study the following dimensions of interest: Linguistic Acceptability (LA): This measures whether the text sounds right to a native speaker. The values of this metric are {, , }, with corresponding to not acceptable, corresponding to some errors, but acceptable and to perfectly acceptable. We chose LA as opposed to grammaticality to ensure a comparable, native-speaker-led evaluation that did not require formal training in the language. Output Content Quality (OCQ): Whether the general quality of the content is good or not, with values {, , }. A score of could indicate that the output is in the wrong language, is repetitive, or sounds like it has been scraped from the web, or translated. A score of 1 indicates that the output is okay in terms of grammar and word choice but still sounds awkward in the language. A score of indicates that the text is of high quality. Task Quality (TQ): This measures the ability of the model to follow the given instructions in the prompt. The values of this metric are {, , }, with indicating that the model did not follow the instructions at all. Likewise, a score of indicates that the model followed the instructions approximately well and that it followed perfectly well. The difference between TQ and OCQ is that the latter focuses on whether the content is appealing to a user, while TQ emphasizes the ability of the model to follow the given instructions. Problematic Content (PC): Whether there was any offensive or problematic content in the output. This is a binary metric, with indicating that the output contains this type of content. Hallucinations (H): This measures how well-grounded the model’s output was to the input content, and/or whether the model output counterfactual information conflicted with the input content. It is a binary metric, with indicating the presence of hallucinations.
For creating this in-house dataset, we asked human judges to evaluate the output of LLM-based systems configured to perform the three tasks described earlier. Each entry was annotated by three annotators. They were contracted through an external annotator services company at a starting rate depending on locale ranging from $ USD/hr and up to $ USD/hr. The pay was adjusted based on locale and experience level. Each annotator was given texts to judge. We used a subset of the annotated data for our experiments.
We provided annotators with the following information: General instructions about the task (including specific instructions from the prompt) and high-level descriptions of the metrics that we are seeking to evaluate, a description of the file that contained data to be evaluated, and the output format expected. Then we provided detailed descriptions of each metric including the range of values for each metric and examples in English. These examples were provided in the context of different tasks, as each metric could have slightly different interpretations for different tasks.
1.2 Data Statistics
Table 1 contains the statistics of the human evaluation dataset for the three tasks across the languages we consider. We create a subset of this data for experimenting with prompting variations and its statistics are available in the small column of the aforementioned table. Our full dataset contains over data points, while the smaller subset contains over data points. Each of the data points in our dataset was annotated by annotators.
2 LLM-based Evaluators
We use the GPT4-32K model as our LLM-based evaluator with a temperature of , except in our ablation experiments. The model was accessed through Azure.
Our evaluation prompts are constructed using the {{guidance}} toolkithttps://github.com/guidance-ai/guidance/tree/main. guidance is a DSL that uses handlebar templating to enable the specification of prompts that interleave instructions and generation with data and logic. This makes it simpler to construct and validate complex prompts.
Evaluation prompts were written to be clear, simple, and not tuned for the data or task. All prompts for evaluation were specified in English, as past work has shown that instructions in native languages can lead to worse performance Ahuja et al. (2023a).
In writing the evaluation prompts, we started with simple unstructured specifications (Natural language sentences with no formatting or styling) and found that it often led to errors in formatting the outputs correctly or even returning all the expected outputs. We found adding styling and formatting, for example, outputting JSON by providing the prompt with a JSON schema for the expected attributes improved the reliability of the LLM outputs.
We tried to keep the task and metric description as close as possible to the text that was shown to human annotators for evaluations in the default prompting variation. Each prompt consists of system, user, and assistant components as shown in Figure 2 in a generic prompt schema. The metric description for Hallucinations is shown in Figure 3Prompts for task description and other metrics are in Appendix A.1..
3 Prompting Variations
First, we experiment with variations based on the number of metrics evaluated and instructions providedAll experiments reported in this study are conducted zero-shot unless specified.. Single Call: In this variation, we call GPT-4 once per metric, without any in-context examples. Compound Call: In this variation, we call GPT-4 once for all the metrics in a single prompt. Single Call - Detailed: In this variation, we call GPT-4 once for all the metrics in a single prompt, with a very detailed metrics description. One of the challenges with LLM evaluation is sensitivity to prompting instructions, which can greatly affect the performance of the LLM on tasks, including evaluation. We experiment with providing detailed instructions for each metric in the prompt. Detailed instruction for Hallucination is shown in Figure 4The detailed instructions for all metrics can be found in Figures 15 - 18 in Appendix A.2. We queried GPT-4 to produce these instructions by providing it with the instructions given to annotators and manually modifying them.
4 Calibration with Human Judgments
Inter-annotator Agreement Analysis: We assessed inter-annotator agreement (IAA) among three annotators Annot1,Annot2,Annot3 using Percentage Agreement (PA) to determine the proportion of data points with consistent annotations across annotators. Weighted F1 scores are documented in Table 2. Additionally, Fleiss’ Kappa () values, which offer insights into agreement beyond chance, are provided in Table 3 (Appendix A.3). Since our dataset is skewed towards one or more classes for each of the metrics, values can be misleading due to known issues with computing expected agreement in such cases Eugenio and Glass (2004).
IAA (3 annotators) and GPT: We measure IAA between the majority score of the three annotators and the LLM-evaluator. We refer to this as AnnotAgg,GPT4 and use PA to measure it. Class distribution: We analyze the class distribution of scores across tasks, metrics, and languages to check for potential biases in the dataset and LLM-evaluator.
We perform experiments contrasting compound and single-call prompting on the full dataset and zero-shot vs. few-shot prompting on the smaller dataset. We analyze how well-calibrated our LLM-based evaluators are with respect to human judgments by examining PA, and class distribution of scores.
5 Ablation Experiments
In addition, we perform some ablation experiments to check for consistency, the effect of hyperparameters, and few-shot examples. We perform these ablations on the smaller dataset. Consistency check: We prompt GPT-4 with the same prompt five times to check its consistency. Single Call – Few-Shot: In this variation, we call GPT-4 once per metric, with a few in-context examples. We provide examples in the prompt of human judgments for the same task and metric from a held-out dev set. We take the majority vote from the three human annotations per sample as the aggregate class for that sample to choose our few-shot examples. For each task, language, and metric we choose up to two samples per possible class for that metric. Therefore, we have a minimum of two and a maximum of six exemplars as few-shot examples. For all evaluations, the few-shot examples used are fixed. Sensitivity analysis: We check the sensitivity of the Linguistic Acceptability metric evaluation by randomly shuffling % of the words in the whole text for all instances and checking if the LA score provided by the model changes. Temperature variation: We vary the temperature parameter to check its effect on LLM evaluation.
Results
In this set of graphs, we look at the percentage agreement between LLM-evaluator and the annotators, and between the annotators. We aggregate the results by task, metric, and language.
Figure 5(a) shows the percentage agreement between the aggregate of the human annotator scores and LLM-evaluator for the full dataset. The figures show both joint (compound), single, and single with detailed instructions prompting techniques for the full dataset. We see that the PA between the annotators and GPT is lowest compared to the PA between the human annotators for Japanese and Czech, with the PA between annotators also being lower for Chinese.
Next, we look at PA grouped by metric in Figures 5(c) for the full dataset with the same prompting variations as before. We find that the PA of the LLM-evaluator with the annotators is lower for the OCQ metric. We also find that the PA between annotators is relatively low for the TQ metric, while all the PA values are very high for the problematic content metrics.
Finally, we look at PA aggregated by task in Figure 5(b). We find that PA is lower for the “Continue Writing” task, while the PA between GPT and the annotators is lower than the agreement between annotators for the “Open Prompt” and “Continue Writing” tasks. Overall, we find that the LLM-evaluator prompted using the compound prompt has a lower agreement with human annotators than the single prompt variation.
Figures 5(a), 5(b) and 5(c) compare the PA of the LLM-evaluators with detailed instructions vs. the simpler instructions described earlier. We find that PA drops slightly for all metrics with detailed instructions.
2 Class Distribution
Next, we examine the distributions of the scores from native speakers and the LLM-evaluator. There are three cases to consider for metrics that have three values: Full agreement (all three annotators give the same score), partial agreement (two of the three give the same score), and no agreement (all three give different scores). In metrics that have binary values, we only have full or partial agreement. We group annotations into these classes and analyze responses across these classes.
We present results for metrics that have three values (LA, OCQ, and TQ), with corresponding to the lowest score and corresponding to the highest score. In Figures 6(a) and 6(b), we find that the LLM-evaluator provides a score of 2 in most cases, particularly in cases where human annotators disagree. This is even more evident in the case of non-English languages where there is partial agreement or no agreement between the annotators (around % of the time on average).
Next, we look at languages that are either lower-resourced or not written in the Latin script. In Figures 7(a) and 7(b) we find that the LLM-evaluator almost never provides scores of and in the 26% of cases that annotators disagree and find similar results for Japanese and Czech shown in Figures 22(e), 22(f), 22(g) and 22(h) in the Appendix A.4. Overall, we find that LLM-based evaluators give a score of 2 in most cases. While this is consistent with human evaluations in a large part of the dataset, the LLM-based evaluator continues to assign a score of even when humans disagree or provide lower scoresFigures for other languages included in Appendix A.4 and A.5..
Interestingly, even though PA drops slightly for all metrics with the detailed instructions, we find that the LLM-based evaluator may be slightly less biased towards producing high scores with these instructions as shown in Figures 8(a) and 8(b). However, more investigation is needed to determine whether detailed instructions or a different prompting strategy can eliminate the bias toward high scores.
We use a temperature of and receive the same score and justification in each of the five tries, showing that the LLM-evaluator exhibits high consistency.
2.2 Few-shot Prompting
Figure 24 in Appendix A.7 shows the PA values when few-shot in-context examples are provided. We observe no significant changes in PA values, suggesting that in-context examples might not significantly aid LLM-based evaluators. This also aligns with the findings of Min et al. (2022).
3 Sensitivity Analysis
As described earlier, we perturb the word order of sentences and check the sensitivity of the Linguistic Acceptability metric on the small dataset. Figure 9 shows the distribution of cases per language per task where the LLM-based evaluator changes its evaluation from a higher score to a lower score. The evaluator shows the most sensitivity to inputs for the Summarization task for all languages except Japanese. For “Continue Writing”, Chinese and Japanese show very little sensitivity. For “Open Prompt", Chinese and Japanese show no sensitivity to the perturbations. One possible explanation for this could be that the evaluator is genuinely less sensitive to these languages. Alternatively, it might be attributed to the flexible word order characteristics of Chinese and Japanese. The examination of tokenizer efficiency in logographic languages, and the exploration of sensitivity across other metrics can be an interesting future exploration.
4 Temperature Variation
Figure 23 in Appendix A.6 show the PA values for temperatures of , , and . PA reduces as we increase temperature, indicating that a temperature of should be used for LLM-based evaluators. We also observe that increasing the temperature makes the model more susceptible to any noise in the data, making the evaluations highly stochastic and not reproducible.
Discussion
Overall, our results indicate that GPT-based evaluators have relatively high consistency for non-English languages when set to a temperature of 0. They also display a fair sensitivity to input variations along the dimension of linguistic acceptability. While LLM-based evaluators show a high Percentage Agreement, there is a noticeable bias towards positive scores, particularly when human opinions differ. It remains uncertain what score an LLM-based evaluator should provide when humans cannot reach a consensus, but consistently high scores in such situations might create a misleading impression of good performance in more challenging evaluations. We find that PA and bias towards higher scores are particularly evident in non-Latin script languages such as Chinese and Japanese, and lower-resource languages such as Czech, which is consistent with prior work on the performance of LLMs on various tasks Ahuja et al. (2023a).
We experiment with several prompting strategies for LLM-based evaluators and find that evaluating a single metric at a time produces better results than evaluating all metrics in one go, which comes at the cost of having to make multiple calls to the LLM. We also find that providing few-shot examples does not help improve performance. We also provide more detailed instructions to the LLM-evaluator but find that it does not eliminate the problem of bias toward higher scores. In this work, we only use evaluators based on GPT-4. An interesting future direction is the use of smaller models for evaluation or models trained with better coverage of non-English data. We also do not do extensive prompt tuning - future work in this direction includes exploring better prompting approaches including automatically tuning prompts to a held-out set.
Our results show that LLM-based evaluators may perform worse on low-resource and non-Latin script languages. Certain metrics corresponding to output quality and task completion may be challenging for LLM-based evaluators. Hence, we advocate for a cautious approach in using LLM-based evaluators for non-English languages and suggest that all LLM-based multilingual evaluations should be calibrated with a set of human-labeled judgments in each language before deployment.
Limitations
In this work, we utilize a dataset comprising human assessments of a text generation system executing various tasks in eight languages. As we do not regulate the quality of the system’s output, most of the generated texts receive positive ratings from human evaluators. Consequently, the high Percentage Agreement’s origin remains unclear – whether it stems from the inclination of the LLM-evaluator to assign high scores or not. In future work, we aim to replicate this study using a dataset with a more balanced distribution of human judgments, achieved by controlling the output quality.
In this work, we utilize an in-house annotated dataset that, due to restrictions, cannot be released, limiting the reproducibility of our research. However, we intend to make a dataset available to the research community for calibrating LLM-based evaluators in the future. An important research direction is the creation of datasets with good language coverage, multiple annotators per data point, and clear annotation instructions, covering a variety of dimensions to calibrate LLM-based evaluators. Exploring the development of various evaluator personas to represent diverse perspectives of human evaluators and achieve consensus is another research direction that needs further investigation.
Ethical Considerations
We use the framework by Bender and Friedman (2018) to discuss the ethical considerations for our work.
Institutional Review: We used an in-house dataset annotated by an external company that has long-standing contracts with the organization and was employed by the organization regularly to do this work.
Data: The LLM evaluator scores were generated using API calls to GPT-4. The dataset used for calibration is an in-house dataset that will not be released publicly. The dataset was not created with the intent of studying human and LLM calibration; hence, it is not a balanced dataset. Specific instructions were provided to LLMs to avoid generating problematic content, and our ratings of the Problematic Content metrics show no such data; however, the possibility still exists.
Annotator Demographics: Annotators were recruited through an external annotator services company. The pay was adjusted after deliberation with the company, based on the annotator’s location and expertise. No demographic information is available about the annotators. The annotators are governed by their company’s and our organization’s privacy policy.
Annotation Guidelines: We draw inspiration from the community standards set for similar tasks. Annotators were given general instructions about the task, detailed instructions about the metrics to be evaluated, and examples in English.
Methods: In this study, we explore several methods of calibrating human judgments with LLM judgments on various tasks and languages. While these methods can be misused to replace human judgments with LLM judgments, our intent with this study is to highlight the gap between the two and urge the community to proceed with caution.
References
Appendix A Appendix
Figure 10 shows task description. Figures 11 - 14 show simple instructions for various metrics.
A.2 Prompts for Detailed Instructions
Figures 15 - 18 show complex instructions for various metrics.
A.3 Fleiss’ Kappa
Table 3 shows the Fleiss’ Kappa () on the full dataset for various annotator combinations, aggregated by language, task, and metrics.
A.4 Class distribution for Metrics with 3 classes
Figures 19 and 20 show class distribution for various languages, aggregated over metrics with 3 classes - LA, OCQ, TQ.
A.5 Class distribution for Metrics with 2 classes
Figures 21 and 22 show class distribution for various languages, aggregated over metrics with 2 classes - H, PC.
A.6 Temperature Variations
Figure 23 shows PA values for different temperature values, results are aggregated over language, task, and metrics.
A.7 few-shot Results
Figure 24 shows PA values for few-shot prompting, results are aggregated over language, task, and metrics.