Calibrating LLM-Based Evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang
Introduction
The emergence of large language models is calling on a greater focus and importance on the quality of natural language generation evaluation. With the rapid improvement of language models, their goals have gone beyond simply fitting its output to a number of given samples to a broader human alignment. Traditional evaluation metrics like BLEU (Papineni et al. 2002), ROUGE (Lin 2004) and CIDEr (Vedantam et al. 2015) often require curated reference outputs, whose application is limited when the output space is open and diversified, and show a low correlation with human judgments (Freitag et al. 2022). While sophisticated model-based evaluators like BERTScore (Zhang* et al. 2020) and COMET (Rei et al. 2020) yield correlation improvements, their performance is still limited by the quality of references. As a result, there is a surging demand for human-aligned, reference-free evaluators for NLG evaluations.
On this front, recent lines of research works explored leveraging state-of-the-art large language models (LLMs) as reference-free evaluators on various NLG tasks (Kocmi & Federmann 2023; Fu et al. 2023; Wang et al. 2023a; Liu et al. 2023). Given that LLMs are optimized to follow human instructions (Ouyang et al. 2022) as well as their state-of-the-art performance on language modeling (OpenAI 2023), they could perform the task of evaluation when prompted accordingly. Multiple evidences show that LLMs are promising competent in evaluating instruction-tuned models like Alpaca (Taori et al. 2023) and Vicuna (Zheng et al. 2023), and being a viable alternative to human expert evaluations (Zheng et al. 2023; Dubois et al. 2023).
Despite these promising results, emerging studies are raising concerns about the validity of LLM-based evaluators - whether LLM’s underlying scoring mechanism aligns with human guidelines and preferences (Liu et al. 2023). Existing LLM-based evaluators enclose the candidate text together with the evaluation task into an instruction prompt. While this paradigm succeeds in presenting the task, it elicits several unaddressed issues, including the sensitivity and bias to output space (Wang et al. 2023a), sample ordering (Wang et al. 2023b), and prompt format (Zheng et al. 2023). Plus, as the scoring prompt is also human-written, it may also incorporate potential bias to the LLM.
To address this issue, we study calibrating an LLM-based evaluator towards better human alignment. We start from a retrospection into existing LLM-based evaluators and uncover they suffer from insufficient prompting, where the scoring guidelines are absent and only output spaces (e.g. 0-100) are provided, resulting in inconsistent and misaligned evaluations (Lu et al. 2023). We argue that such an issue could be mitigated by elucidating the scoring criteria. And by finalizing the scoring criteria, a consensus could be reached between humans and the LLM, as a means of alignment.
However, it is non-trivial to obtain adequate criteria Results in Chen et al. 2023 suggest that poorly curated criteria reduce relevance with human expert scoring. Un-calibrated random criteria would introduce extra bias as a misalignment between the standards used for human experts. And improperly assigned rubrics might reduce the difference between each score., as it may require expert-level domain knowledge to assign rubrics and prevent personal bias. Drawing inspirations from the in-context learning capability (Dong et al. 2022) of LLMs, we propose AutoCalibrate, a framework to automatically align and calibrate an LLM-based evaluator through human alignment. To tackle the challenge of curating scoring criteria, we take a data-driven methodology to draft, filter, and refine rubrics using the LLM, based on human expert labels. By incorporating the mined and calibrated rubrics into scoring instructions, we obtained significant improvements in human alignment when evaluating text summarization, data-to-text generation, and hallucinations. Moreover, we release the optimal scoring criteria sets mined for the above tasks, and present detailed qualitative and quantitative analysis to uncover the essence that makes an effective criteria.
Methodology
Figure 1 illustrates the overall framework of AutoCalibrate. To calibrate an LLM-based evaluator, we focus on optimizing the evaluation prompt template applied to improve the correlation and alignment between LLM’s scores and human preference. Specifically, we mine and tune the scoring criteria in pursuing such alignment. To express human preference, we first construct a golden set , containing ground-truth sample-label pairs from human expert labeler. We then follow a novel multi-stage procedure to optimize candidate scoring criteria, including drafting and revisiting. Initial criteria drafts are first inferred from in-context labels and an induction prompt, evaluated and filtered on expert labels, and then refined to accommodate erroneous evaluations.
2 Problem formulation
In this section, we elaborate on the calibration medium and objective of AutoCalibrate - the scoring criteria. Denote the dataset which contains multiple samples to evaluate. Based on different tasks, a sample can contain various components: single text, for tasks like evaluating grammatical correctness; source-target, for the vast majority of conditional generations, and multi-turn, like assessing multi-turn conversations.
To guide the LLM to evaluate the quality of sample , prompts are applied to provide sufficient instructions and clarifications of the task. To calibrate the prompt template applied during evaluation, we regulate it by decomposing it into the following building blocks: instructions, criteria, aspect, output format, and data sample to evaluate, as illustrated in Figure 2. For an arbitrary sample , given a prompt template (guides the LLM to perform evaluation on NLG quality), scoring criteria , evaluation aspect (e.g., fluency, coherence, consistency) and a large language model , the NLG quality of could be evaluated as
Denote a golden set consists of curated sample-label pairs from human experts, and an correlation metric. In AutoCalibrate, we focus on calibrating the scoring criteria to maximize the correlation between predicted labels and human expert labels, as
3 AutoCalibrate
To calibrate an LLM-based evaluator, one primary question is: how to represent and model the preference of human experts. On existing approaches, sophisticated model-based evaluators like COMET (Rei et al. 2020) directly train on human labels, while ranking-based human labels are widely adopted in RLHF to model human preference (Ouyang et al. 2022). However, these model-based preference modeling methods require extra fine-tuning, which makes them computationally intensive and impracticable to API-based LLMs. To mitigate these limitations, We implicitly encode human expert preference to a set of sample-label pairs and form a golden set . Compared with curating finalized scoring criteria and guidelines with joint human expert works, it is more feasible to collect labels leveraging crowd-sourcing dispatch, and also easier to validate and merge opinions from different experts.
Criteria Drafting
After constructing the expert label set , we utilize the instruction following and in-context learning capability of LLMs to independently infer scoring criteria from few-shot exemplars. One crucial part here is to ensure the diversity of recalled criteria. To mitigate the label bias and position bias of in-context learning (Zhao et al. 2021), we construct various Monte-Carlo samples from to obtain few-shot in-context exemplars. Given drafting prompt template and a few-shot exemplar set , an corresponding criteria is inferred as
where denotes the evaluation aspect. Temperature sampling is also applied to draw scoring criteria in diversified presentations from the LLM. Example prompt templates are provided in Appendix D.1. Following this procedure, we obtain the initial set of scoring criteria for evaluation and refinement.
Criteria Revisiting
Inferred from various few-shot exemplars, criteria within the initial draft set are diversified, but may be sub-optimal or contain potential bias (e.g., to particular scoring labels). To filter out high-quality candidates, we first revisit them leveraging and select the top performing candidates w.r.t their human relevance A meta-evaluation method is applied here to perform meta-evaluation on the correlation between human and LLM judgments. For detailed explanations and definitions, please refer to Appendix A.. To mitigate disagreements between human experts and the drafted criteria, we prompt LLMs to refine (Madaan et al. 2023) the previously generated criteria by providing them samples with strong disagreement in their scores. When refining the criteria, we suggest the following atomic editing operations via prompting to the LLM Detailed prompt examples are provided in Appendix D.3.:
Modification: Adjust some parts of the criteria to increase its correlation.
Paraphrase: If a criteria is good enough, paraphrase it to make it clearer and more concise.
Adding Aspects or Details: When LLM discovers new underlying scoring rules that are not covered by the current criteria, consider adding them as a new line to the current criteria, but make sure not to make the criteria too long and redundant.
Calibrate: Any other modifications that the LLM considers helpful.
As illustrated in Figure 1, after obtaining refined candidate criteria, we first filter them with and then combine them with the pre-filtered draft criteria to obtain a calibrated set of scoring rules.
Conclusion and Discussion
Combining the above, we obtain AutoCalibrate, an automatic pipeline in calibrating LLM-based evaluators. The overall procedure is summarized in Algorithm 1.
The benefits of choosing criteria as a medium for calibration are multitudinous. First, we do not require gradients or access to model parameters, which makes AutoCalibrate applicable to API-based LLMs. Second, since criteria remain in a natural language form (compared with soft prompt-tuning), calibrating the criteria is essential to reaching an agreement between humans and the LLM. Therefore, the process is more explainable and controllable (e.g., one may perform human-in-the-loop adjustments to scoring criteria in case of preference changes, or to avoid corner cases).
Experimental Setup
We evaluate AutoCalibrate on three text quality evaluation tasks, including text summarization, data-to-text generation, and evaluating hallucinations. We select tasks following previous research works (Zhong et al. 2022; Fu et al. 2023). We select two datasets for each of the tasks, consisting of 6 datasets in total, each containing human expert labels for candidate samples. Specifically, we select NewsRoom (Grusky et al. 2018) and SummEval (Fabbri et al. 2021) for evaluating machine summarization; SFRES (Wen et al. 2015) and SFHOT (Wen et al. 2015) for data-to-text task, QAGS-XSUM and QAGS-CNN (Wang et al. 2020b) for evaluating hallucinations. To evaluate the alignment between the scoring from LLM and human experts, we perform a meta-evaluation following (Zhong et al. 2022). Details on the evaluation strategy are listed in Appendix A.
2 Models and Baselines
To implement AutoCalibrate, we select OpenAI’s GPT-4 model (GPT-4-32K) as the LLM for the evaluator. We list prompt templates for criteria drafting, evaluation, and refinement for tach tasks in Appendix D. We set the temperature to during evaluation, and when obtaining initial criteria drafts and their refined versions. Please refer to Appendix C for detailed configurations of each task.
We compare AutoCalibrate with various state-of-the-art and/or widely applied evaluators. We first include ROUGE (Lin 2004), a widely-applied n-gram-based evaluation metric for text summarization. We then select various evaluators based on smaller neural (language) models, including BERTScore (Zhang* et al. 2020), MoverScore (Zhao et al. 2019), PRISM (Thompson & Post 2020), BartScore (Yuan et al. 2021), CTC (Deng et al. 2021), and UniEval (Zhong et al. 2022). Finally, we compare evaluators based on state-of-the-art LLMs (e.g. GPT-3.5 and GPT-4), including GPTScore (Fu et al. 2023), ChatGPT The ‘ChatGPT’ evaluator included multiple versions according to different prompt templates, and we mark these variants with parentheses. We encourage readers to check the original works for detailed information. (Wang et al. 2023a), and GPT-Eval (Liu et al. 2023).
Experimental Results
We conduct meta-correlation analysis on NewsRoom and SummEval benchmark to evaluate AutoCalibrate’s performance to calibrate an LLM-based evaluator on text summarization. Following Liu et al. 2021, we perform summary-level Spearman and Kendall correlation analysis on each of the 4 evaluation metrics with human expert evaluations. To represent the performance of our un-calibrated backbone LLM, we add a GPT-4 baseline, whose evaluations are obtained with a one-pass call using an evaluation prompt where scoring criteria is omitted For a fair comparison, the only difference is the removal of criteria from prompt. We keep the rest identical..
Results on NewsRoom and SummEval benchmark are listed in Table 1 and 2, respectively. On NewsRoom benchmark (Table 1), our AutoCalibrate significantly outperforms the LLM-based ChatGPT evaluator. It also surpasses the vanilla GPT-4-based evaluator by a large margin (with a 10.4% improvement on Spearman and 11% on Kendall correlation), demonstrating the importance and effectiveness of the calibration procedure. While BartScore obtained a competent performance on NewsRoom, it falls short on SummEval. We conjecture that since it utilizes a smaller model, the consistency of its scoring might be hindered due to the distribution of its fine-tuning corpus.
In contrast, our AutoCalibrate demonstrated a consistent human relevance uplift on both summarization datasets, since the pretraining knowledge in LLM is more in-depth and generalizable. On SummEval, AutoCalibrate improves the human correlation of GPT-4 evaluations by 7.3%, and also superior to a strong baseline G-Eval-4 that also utilizes GPT-4. Noteworthy, G-Eval-4 requires 20 calls from LLM to obtain an average score to mitigate replicated evaluations. While this improves Spearman correlation by creating a more continuous distribution, it reduces the rank coefficient. In contrast, by elucidating the scoring rule with calibrated criteria, AutoCalibrate improves both Spearman (2.9%) and Kendall (13.4%) coefficients with only one forward call.
2 Results for Data-to-Text
We consider SFRES and SFHOT datasets for evaluation of data-to-text generation task and follow Fu et al. 2023 to conduct dataset-level meta-evaluation on human alignment. Results are listed in Table 3. As illustrated in the table, AutoCalibrate significantly outperforms the most competent trained evaluator (UniEval) over 30%, and yields an over 20% and 10% improvement on Spearman correlation over GPT-Score (based on 175B-LLM GPT-3.5) and uncalibrated GPT-4 evaluator, respectively. These results suggest that the proposed procedures within AutoCalibrate could promptly curate adequate scoring criteria for different NLG tasks and sample distributions.
3 Results for Evaluating Hallucinations
Hallucinations are an important issue in NLG models where the output is based on fabricated, unwarranted facts or digress from a previous context, and it is becoming an increasingly important topic for trustworthy LLMs (Ji et al. 2023). To test AutoCalibrate on evaluating hallucinations, we select QAGS-CNNDM and QAGS-XSUM dataset and perform dataset-level meta-analysis following Liu et al. 2023. As presented in Table 4, AutoCalibrate uplift the average Spearman correlation by 15% over G-Eval-4. Noteworthy, since fine-tuned on CNN data, BartScore achieves promising human relevance on QAGS-CNN, but significantly falls short on QAGS-XSUM, while LLM-based AutoCalibrate performs consistently on both datasets. This further indicates that LLMs, given their immense knowledge gained during pre-training, are strong candidates for general evaluators, and their performance could be further boosted with proper calibration.
4 Ablation Experiments
We conduct ablation studies on the procedure of AutoCalibrate to better investigate the contribution of each process in calibrating LLM-based evaluator. The main ablation experiments are listed in Table 5. As illustrated in the table, removing criteria in the prompt significantly reduces the human correlation of GPT-4. This corroborates our argument that previously LLMs suffered from a vaguely defined scoring principle, and this could be calibrated to increase the human alignment of LLM evaluators. The self-refine process also positively contributed to the improvements in human alignment. This indicates that LLMs could accordingly adjust the effectiveness of scoring criteria. Detailed qualitative analysis is presented in Chapter 5.
Analysis
In this chapter, we present statistical analysis on the pool of draft candidates of scoring criteria, and mine for possible essence that contributes to effective scoring criteria with high human relevance for LLM-based evaluators. The main results are presented in Figure 3.
We study the sensitivity of AutoCalibrate to the sample size of few-shot in-context samplers. As illustrated in Figure 3(A), the size of in-context few-shot exemplars yields no significant impact except for QAGS-CNN. The results indicate that AutoCalibrate is mostly robust to the size of in-context samples. Thanks to the sufficient prior knowledge obtained during pretraining by the LLM, AutoCalibrate is capable of inferring the underlying criteria using only a few examples in context. As illustrated in the figure, a few-shot size of 8 to 12 is sufficient in mining effective criteria across all tasks. This intriguing feature enables a reduction in search space for cost reductions upon deployment.
Effect of Criteria Length
The distribution of lengths of generated criteria and their human relevance is illustrated in Figure 3(B). Most evaluation criteria drafted and refined with AutoCalibrate lie in the range of 60 to 600 words. We discover different trends on the preference of AutoCalibrate to different lengths of criteria. While fluency and coherence metrics on text summarization lean towards shorter criteria, lengthier versions are favored by the informativeness metric on data-to-text and evaluating hallucinations. Despite this difference, AutoCalibrate enjoys the capability to generate effective criteria at each length. We conjecture this nuance is caused by the intrinsic complexity of the aspect to evaluate: it could be straightforward to define fluency, but possibly more challenging to address hallucination.
Patterns of Criteria
We observed two significant patterns on the criteria drafted by GPT-4: holistic and specific. The former typically characterizes the common features possessed by high and low-quality samples, while the latter generates a segment of the corresponding rubric for each evaluation score (e.g., 1 to 5). A random example of these patterns of criteria is listed in Table 6. These two patterns emerge across all sets of experiments on different benchmarks. The performance distribution of these two patterns across all datasets is illustrated in Figure 4. As illustrated in the figure, there is no significant difference in human expert correlation between holistic and specific patterns, indicating that both patterns generated from AutoCalibrate are of high quality. Therefore, the performance of AutoCalibrate is robust to the patterns of criteria generated.
2 Case Study
To investigate the effect of criteria refinement, we present a case study in Table 7. As demonstrated in the table, when prompted with previous misaligned evaluation cases and possible means of modifications (Section 2.3), the LLM automatically infers new patterns of underlying scoring principles, and promptly adapts the existing criteria to accommodate them. As illustrated in the table, AutoCalibrate discovers that the genre and format is crucial to the fluency of summary from in-context examples provided, adjusts the criteria accordingly, and achieves higher human relevance. These findings corroborate with Madaan et al. 2023 that LLM is capable of self-refine, and opens a future research direction on the multi-turn, iterative calibration of LLM-based evaluators.
Related Work
It has been a long and arduous endeavor to automatically evaluate natural language generations. This paragraph outlines automatic evaluation metrics before the era of LLM. (1) N-gram-based metrics: as the most widely adopted method, n-gram-based metrics measure the quality of a candidate text by the overlap of its lexical fraction between references. As two of the most widely used metrics, BLEU (Papineni et al. 2002) and ROUGE (Lin 2004) are specialized in precision for machine translation and recall for text summarization, respectively. Despite being widely applied, their human relevance is undesired (Freitag et al. 2022). (2) Embedding-based metrics: this line of method leverages a pre-trained language model (e.g. BERT (Devlin et al. 2019)) to measure the similarity between word embedding of the candidate and reference text (Zhang* et al. 2020; Zhao et al. 2019). Their major limitation lies in the similarity-based paradigm and high dependency on the quality and diversity of references. (3) Trained neural evaluators: more recent research focus on specializing the PLMs by either fine-tuning on human (Rei et al. 2020) or synthetic (Zhong et al. 2022) labels, or pretraining on domain-relevant documents (Yuan et al. 2021). However, these metrics either focus on a single dimension (Wang et al. 2020a; Huang et al. 2020) or are limited in human relevance (Mehri & Eskenazi 2020; Zhong et al. 2022).
LLM-Based NLG Evaluation
With the emergence of LLM, recent research works focus on LLM-based evaluators given their promising instruction-following and generalization capability. A first line of work goes through preliminary explorations on LLM-based evaluators, including prompting methods and model variants (Fu et al. 2023; Kocmi & Federmann 2023; Wang et al. 2023a; Chen et al. 2023; Liu et al. 2023). Successor research focuses on various aspects of improving LLM-based evaluators, including factuality (Min et al. 2023), interpretability (Lu et al. 2023), mitigating position bias (Wang et al. 2023b), and agreement to human evaluation (Zheng et al. 2023). Different from the above approaches, we focus on a general method to calibrate an off-the-shelf LLM with gradient-free approaches, to improve its alignment with human preferences on a desired task.
Conclusion
In this work, we focus on an important question: how to calibrate and align an off-the-shelf LLM-based evaluator towards human alignment in a gradient-free fashion. We first take a retrospection into existing LLM-based NLG evaluators and uncover they suffer from insufficient prompting, where the scoring guidelines are absent and only output spaces are provided, resulting in inconsistent and misaligned evaluations. We emphasize the significance of aligned scoring criteria as a consensus between humans and LLM and propose AutoCalibrate to automatically calibrate an LLM-based evaluator through criteria drafting and refinement. Inferred from human expert labels and refined according to previous misalignment samples by the LLM, the criteria curated by AutoCalibrate demonstrate significant improvements in human correlation across evaluating text summarization, data-to-text, and hallucinations. Our qualitative analysis conveys insightful intuitions and observations on the essence of effective scoring criteria.
Discussions
This work study on calibrating a strong LLM-based evaluator towards better human alignment. Beyond manual prompt engineering, AutoCalibrate automates the calibration process of LLM-based evaluators and provides a first experimental study on how further LLM-based evaluators could be strengthened with better prompting. We envision AutoCalibrate being potentially applied to a wider spectrum of tasks in NLG and beyond.
The primary limitation is that only criteria are mined to improve alignment. After carefully analyzing prompts, we conclude that the criteria are most crucial, as they are most causal to the scores given, and can be regarded as a shared consensus between humans and LLMs due to their natural language form. Plus, the criteria section is the hardest to curate compared with other parts of the prompt template (e.g., scoring scale, task definition), on which we primarily focus. Besides, A more comprehensive research on advancing and assessing other components of prompts to calibrate a LLM-based evaluator, and adapting it to wider tasks and languages is open to future work.
References
Appendix A Evaluation Strategy
In this section, we introduce meta-evaluation strategies for assessing human alignment that are applied in this work. We select evaluation strategies primarily following previous works (Zhong et al. 2022; Fu et al. 2023; Liu et al. 2023). Given a dataset consisting of NLG samples from diverse systems and source text samples, evaluation metric (e.g., BLEU (Papineni et al. 2002)) and correlation metric , we could perform meta-evaluation at either sample or dataset level.
For sample-level meta-evaluation, we first compute correlation values on multiple candidate response (from each system) to a individual sample, then average across all samples:
where and denote the evaluation results (if not, converted to a numeric value) for the -th response to -th sample from evaluator and human experts, respectively.
Dataset Level
For dataset-level meta-evaluation, we evaluate the correlations on all samples in the dataset (with a total of samples), as follows:
Appendix B On Performance of Adding Chain-of-Thoughts
Chain-of-thought (CoT) (Wei et al. 2022) prompting elicits reasoning in large language models by encouraging models to generate their rationales before obtaining an answer. As studied in recent research (Liu et al. 2023), chain-of-thoughts are beneficial to improving human alignment in NLG evaluation, if incorporated in the scoring prompt template. Therefore, we study whether AutoCalibrate could further benefit from adding a CoT into our calibrated scoring prompts.
To obtain the CoT for each scoring aspect, we follow Liu et al. 2023, and results are illustrated in Table 8. As shown in the figure, adding CoTs to our calibrated prompts yields negligible difference. We conjecture the effectiveness of ‘CoT’ is marginalized by providing informative and instructive scoring criteria. In contrast to math, the assessment of text quality is not a strictly chained reasoning process, so providing a CoT is essentially clarifying the evaluation rubrics, which is consistent with the meaning of the criteria in this paper, and thus obtained no additional benefit. Plausibly, the ‘CoT’s here act to elucidate the scoring rules, rather than providing reasoning paths to follow.
Appendix C Configuration Details
In this section, we list the configuration details of AutoCalibrate for each experiments. Detailed configurations for AutoCalibrate are listed in Table 9. We apply the same set of configurations to each of the two datasets within a task.
Appendix D List of Prompt Templates
In this section, we list prompt templates applied throughout this study, including induction templates for criteria drafting, evaluation templates that utilize the generated scoring criteria, and templates for self-refinement of criteria.
Prompt templates for criteria drafting are listed in Figure 5, 6 and 7. The [Aspect] denote placeholders for aspects to evaluate (e.g. coherence, consistency, etc.), and sampled few-shot in-context exemplars are placed at [In-Context Few-Shot Samples], including samples and their expert scores.
D.2 Evaluation Templates
Prompt templates for evaluation are listed in Figure 8, 9 and 10. The [Aspect] denotes placeholders for aspects to evaluate (e.g. coherence, consistency, etc.). Evaluation samples and calibrated scoring criteria for each aspect are filled into corresponding placeholders during evaluation.
D.3 Criteria Refinement Templates
An example prompt template for criteria refinement can be found in Figure 11. As illustrated in the figure, we first fill in the aspect and tasks to the instructions, then prompt the LLM with the previous criteria, few-shot in-context samples of misaligned evaluations, together with suggested means of modifications to obtain a modified version of scoring criteria for this task.
Appendix E Extended Case Study
In this section, we present a case study on scoring criteria generated by AutoCalibrate for each evaluation aspect of each benchmark throughout this study in Table 10, 11, 12 and 13. As illustrated in the tables, scoring criteria generated with AutoCalibrate are informative, covering significant rubrics to evaluate a given aspect of the target NLG task.