Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
Introduction
Evaluating the quality of outputs produced by language models (LMs) is progressively becoming difficult, as the outputs cover an extremely diverse distribution of text and complex tasks. To address this issue, language model-based evaluation has emerged as a scalable and cheap paradigm for assessing LM-generated text (Li et al., 2024; Gao et al., 2024). In this paradigm, LMs are either prompted to output a scalar indicator of quality (denoted as direct assessment) (Zheng et al., 2023; Liu et al., 2023b; Ye et al., 2023; Kim et al., 2023) or to determine which of two outputs are preferred (denoted as pairwise ranking) (Wang et al., 2023b; Li et al., 2023b; Lambert et al., 2024). Prior works employing proprietary LMs as evaluators have demonstrated not only high correlations with human evaluations but also increased speed and cost-effectiveness (Zheng et al., 2023; Liu et al., 2023b; Dubois et al., 2023; Ye et al., 2023).
However, relying on proprietary LMs for evaluation poses significant challenges. The lack of transparency about their training data compromises both fairness and compliance, making it problematic to use them in evaluation pipelines. Additionally, concerns regarding controllability and affordability also persist (Kim et al., 2023). To address these issues, recent works have focused on developing evaluator LMs that are open-access, transparent, and controllable (Kim et al., 2023; Wang et al., 2023a, b; Li et al., 2023a; Zhu et al., 2023; Jiang et al., 2023b, c; Lee et al., 2024). Yet, these models often yield scoring decisions that do not correlate well enough with human judgments or those made by proprietary LMs, failing to effectively simulate them. Moreover, open evaluator LMs are not flexible since they are typically trained only to perform either direct assessment or pairwise ranking and assess based on general public preferences like helpfulness and harmlessness, limiting their ability to handle diverse real-life scenarios.
To close the gap with proprietary LMs, we investigate unifying the two model-based evaluation paradigms - direct assessment and pairwise ranking - to train a robust unified evaluator LM. We propose a recipe based on merging the weights of two evaluator LMs trained separately on direct assessment and pairwise ranking formats. Our key empirical observation is that weight merging can yield an evaluator LM that not only works in both formats, but also outperforms evaluator LMs that are jointly trained or only trained on a single format.
To demonstrate our approach, we develop the Preference Collection, a new fine-grained pairwise ranking feedback dataset that builds on the Feedback Collection Kim et al. (2023), which is a direct assessment feedback dataset. We choose Mistral-7B (Jiang et al., 2023a) and Mixtral-8x7B (Jiang et al., 2024) as our base models, and merge the weights of evaluator LMs separately trained on the Feedback Collection and the Preference Collection to obtain our resulting models, Prometheus 2 (7B & 8x7B).
On four direct assessment benchmarks (Vicuna Bench, MT Bench, FLASK, Feedback Bench), the Prometheus 2 models demonstrate the highest correlation with both human evaluators and proprietary LM-based judges compared to existing open evaluator LMs, with the Pearson correlation surpassing other baselines by 0.2 units across all datasets. Similarly, on four pairwise ranking benchmarks (HHH Alignment, MT Bench Human Judgment, Auto-J Eval, Preference Bench), the Prometheus 2 models show the highest agreement with human evaluators among all the open evaluator LMs we tested, reducing the performance gap with GPT-4 in half.
Our contributions are summarized as follows:
We introduce Prometheus 2 (7B & 8x7B), state-of-the-art open evaluator LMs that score high correlations with both human evaluators and proprietary LM-based judges on both direct assessment and pairwise ranking.
We introduce a pairwise ranking feedback dataset called the Preference Collection, which includes 1K custom evaluation criteria beyond helpfulness and harmlessness.
We show that merging the weights of evaluator LMs trained on direct assessment and pairwise ranking feedback datasets results in a unified evaluator LM that excels in both schemes.
Related Work
To assess the generation capabilities of LMs, prior works such as the GEM benchmark (Gehrmann et al., 2021, 2022) employed Rouge (Lin, 2004), BLEU (Papineni et al., 2002), and BERTScore (Zhang et al., 2019) as their metric, which measures the lexical or semantic similarity between a reference answer and a response. However, these conventional metrics are prone to false negatives because they are not expressive enough to recognize responses that are of good quality but differ from the reference answer (Schluter, 2017; Freitag et al., 2020; Hanna and Bojar, 2021).
Recently, employing language models as a judge has gained attention as a promising paradigm to mimic the depth and granularity that human evaluation offers (Zheng et al., 2023; Liu et al., 2023b; Li et al., 2023b; Chan et al., 2023; Ye et al., 2023). To reduce the over-reliance on proprietary LMs, follow-up works suggest training language models specialized in evaluations (Cui et al., 2023; Kim et al., 2023; Jiang et al., 2023b, c; Li et al., 2023a; Lee et al., 2024). Yet, open evaluator LMs do not possess the flexibility to function in different evaluation schemes and show weak evaluation performances compared to proprietary LMs. We aim to bridge this gap by introducing Prometheus 2.
2 Weight Merging
Prior works have demonstrated that weight merging can enhance performances across various domains, including language modeling (Li et al., 2022; Matena and Raffel, 2022; Ilharco et al., 2022; Don-Yehiya et al., 2022; Gururangan et al., 2023; Yadav et al., 2024; Sukhbaatar et al., 2024), instruction-tuning (Jang et al., 2023b; Yu et al., 2023), and aligning to user perferences (Jang et al., 2023a; Rame et al., 2024; Wang et al., 2024). In our work, we specifically focus on enhancing the evaluation capabilities of open evaluator LMs. By merging models trained on different assessment formats—specifically, direct assessment and pairwise ranking—we aim to obtain an evaluator LM that not only functions in both formats but also shows as good evaluation performances as proprietary LMs.
Methodology
We propose a new recipe for training a unified evaluator LM based on merging the weights of models trained for direct assessment and pairwise ranking. We begin with background on direct assessment and pairwise ranking for evaluator LMs (Section 3.1, 3.2), followed by the construction process of our training data (Section 3.3). Finally, we present our methods to train the state-of-the-art evaluator LM, Prometheus 2 models (Section 3.4).
Prior works have identified several recipes to align the scores provided by evaluator LMs () and the scores assigned by humans (). For instance, Liu et al. (2023a) and Zheng et al. (2023) have shown that it is crucial to add a reference answer as input to the evaluator LM to maximize the correlation between and . Also, Zheng et al. (2023) and Ye et al. (2023) showed that prompting the language model to write verbal feedback before also improves the correlation between and . Lastly, Ye et al. (2023) and Kim et al. (2023) showed that by explicitly integrating evaluation criteria , users can define the standards for model assessment, ensuring evaluations are flexible to specific needs rather than generic qualities. Specifically, is represented as a score rubric including a description for the criteria itself and a set of descriptions for each score between the scoring range. This is expressed as:
2 Pairwise Ranking
Pairwise ranking is mapping an instruction and two pair of responses , into either or , such as .
Similar to direct assessment, prior works have identified that integrating a reference answer and verbal feedback into the evaluation pipeline is crucial (Zheng et al., 2023; Li et al., 2023b, a). In addition, to support granular assessment under custom criterion, we add the evaluation criteria as input to the evaluator LM (Ye et al., 2023; Kim et al., 2023). To the best of our knowledge, we are the first to study such fine-grained evaluation in pairwise ranking settings. This is expressed as:
In pairwise ranking, the evaluation criteria do not include a set of descriptions for each score; instead, only the description of the evaluation criterion itself. Also, it is noteworthy that the verbal feedback compares the commonalities and differences between and concerning .
3 The Preference Collection
Popular pairwise ranking datasets such as HH-RLHF (Bai et al., 2022) or Ultra Feedback (Cui et al., 2023) do not include an evaluation criteria and a verbal feedback . To obtain an evaluator LM that could assess based on what users care about, we construct the Preference Collection that includes 1K evaluation criteria.
To construct the Preference Collection, we apply two modifications to the Feedback Collection. First, since the Feedback Collection includes five responses for each instruction, each corresponding to a scoring decision between 1 and 5, we pair two out of the five responses, resulting in a total of ten combinations per instruction. Using the existing scoring decisions for each response, we determine which response is better and assign a new scoring decision for that pair (i.e., “Response A is better” or “Response B is better”). Second, to generate new verbal feedback for each pair of responses, we prompt GPT-4-1106 to identify the commonalities and differences of the two responses.
The statistics of the resulting dataset are listed in Table 1 along with the Feedback Collection. We explain about our quality verification process of the Preference Collection in Appendix A. Also, we include the prompts we use for the augmentation process in Appendix F.
4 Employing Evaluator Language Models
Prompting involves querying an LM to make judgments in a specified evaluation format without training on any feedback dataset.
Single-Format Training
Single-Format training involves training a base model on either on a direct assessment feedback dataset or a pairwise ranking feedback dataset .
Joint Training
Joint training involves training a base model on both a direct assessment feedback dataset and a pairwise ranking feedback dataset . This enables the resulting evaluator LM to function across both evaluation formats.
Weight Merging
Weight Merging involves training two models, and , separately on a direct assessment feedback dataset and a pairwise ranking feedback dataset . Then, we obtain the final evaluator LM with linear merging :
We conduct experiments by using . In Section 6.3, we observe how altering the coefficient affects downstream performance on each evaluation scheme. We empirically find that this simple recipe work best when we choose Mistral-7B as our base model. In addition to linear merging, we also test different merging techniques including:
Task Arithmetic merging (Ilharco et al., 2022) which can be expressed as follows:
where is the weight of the base model. However, we empirically find that the resulting evaluator LM often does not generate valid scoring decisions (e.g., generating an integer during pairwise ranking).
TIES merging (Yadav et al., 2024), while similar to Task Arithmetic merging, adds (1) a Trim operation to remove redundant weights in the task vector and and (2) Elect and Disjoint operations to resolve disagreement (i.e., opposite directed weights) between and .
DARE merging (Yu et al., 2023), while also similar to Task Arithmetic and TIES merging, performs a Random Drop and Re-scale operations in the task vector and to remove redundant weights. We find that DARE merging work best when we choose Mixtral-8x7B as our base model.
Experimental Setup
In this section, we explain our experimental setup to assess evaluator LMs. We first explain the benchmarks and metrics we employ (Section 4.1) and the baselines we use as evaluator LMs (Section 4.2).
The statistics of all the benchmarks are in Table 2.
The four direct assessment benchmarks are:
Vicuna Bench (Chiang et al., 2023): A single-turn chat benchmark that includes 80 test prompts, 80 hand-crafted score rubrics from Kim et al. (2023), and 320 responses obtained by WizardLM-13B, Vicuna-13B, Llama-2-Chat-13B, GPT-3.5-Turbo-0613.
MT Bench (Zheng et al., 2023): A multi-turn chat benchmark that consists of 80 test prompts, 80 hand-crafted score rubrics from Kim et al. (2023), and 320 responses obtained by WizardLM-13B, Vicuna-13B, Llama-2-Chat-13B, GPT-3.5-Turbo-0613.
FLASK (Ye et al., 2023): A fine-grained evaluation benchmark comprised of 200 test prompts, 12 score rubrics, and 2000 responses acquired from Alpaca-7B, Vicuna-13B, Bard, GPT-3.5-Turbo-0613. In addition to scores from proprietary LMs, this benchmark also includes scores marked by human evaluators.
Feedback Bench (Kim et al., 2023): The test set of the Feedback Collection with 1K score rubrics, 200 instructions, and 1K responses that do not overlap with the train data.
The four pairwise ranking benchmarks are:
HHH Alignment (Askell et al., 2021): A benchmark consisting of 221 prompts; 4 score rubrics (helpfulness, harmlessness, honesty, and other) and 221 response pairs (graded as ‘win’ or ‘lose’) judged by human evaluators.
MT Bench Human Judgment (Zheng et al., 2023): A benchmark that shares the same 80 prompts as MT-Bench. In addition, it provides 3,360 response pairs (graded as ‘win’, ‘tie’, or ‘lose’) judged by human evaluators.
Auto-J Eval (Li et al., 2023a): A benchmark consisted of 58 prompts and 1,392 response pairs (graded as ‘win’, ‘tie’, or ‘lose’) judged by human evaluators. This benchmark is used as the in-domain test set of Auto-J.
Preference Bench: Our in-domain test set for the Prometheus models. Similar to how the Preference Collection was made with the Feedback Collection, we adjust the Feedback Bench and pair two out of the five responses, resulting in a test set with 200 prompts, 2,000 response pairs (graded as ‘win’ or ‘lose’), and 200 evaluation criteria.
In direct assessment, we conduct reference-based evaluations by appending the reference answer as the input. We use Pearson, Spearman, and Kendall-Tau as performance metrics to measure scoring correlations against reference evaluators.
In pairwise ranking, we conduct reference-free evaluations. Based on judgments assigned by humans, we use accuracy as our metric to measure agreement between evaluator LMs and humans.
Also, the MT Bench Human Judgment and Auto-J test set includes a ‘tie’ option assessed by human evaluators. We evaluate in two ways: by excluding all ‘tie’ options for pairwise ranking (denoted as ‘w/o tie’), or by using direct assessment where responses scored as ‘ties’ are grouped, and pairwise rankings are applied to the remaining responses with differing scores (denoted as ‘w/ tie’).
2 Baselines
We employ Llama-2-Chat-7,13,70B (Touvron et al., 2023); Mistral-7B-Instruct-v0.2 (Jiang et al., 2023a); and Mixtral-8x7B-Instruct-v0.1 (Jiang et al., 2024) as our baselines. It’s worth noting that models not explicitly trained on feedback data often fail to generate responses in the required format, making it extremely difficult to parse scoring decisions. Although it is impractical for regular use, we make a fair comparison by infinitely looping until scores can be parsed. Also, we include proprietary LMs such as GPT-3.5-Turbo-0613; GPT-4-1106; and Claude-3-Opus.
Single-Format Trained Evaluator LMs
For single-format trained evaluator LMs, we test Prometheus-7,13B (Kim et al., 2023) (direct assessment); UltraRM-13B (Cui et al., 2023) (pairwise ranking); and PairRM-0.4B (Jiang et al., 2023c) (pairwise ranking). In addition, we also report the performances of single-format training Mistral-7B-Instruct-v0.2 and Mixtral-8x7B-Instruct-v0.1 on either direct assessment or pairwise ranking.
Jointly Trained Evaluator LMs
For jointly trained evaluator LMs, we test Auto-J (Li et al., 2023a). In addition, we report the performances of jointly training Mistral-7B and Mixtral-8x7B on both direct assessment and pairwise ranking.
Weight Merging
Prometheus 2 (7B & 8x7B) models are our weight merging baselines.
Details on the hyper-parameters for training and inference along with the prompt templates are all listed in Appendix B, G, H.
Experimental Results
The direct assessment results are shown in Table 3. The scoring decisions of Prometheus-2 models (7B & 8x7B), GPT-4-1106, Claude-3-Opus, and human evaluators all strongly correlate with each other, yielding Pearson correlations higher than 0.5 regardless of the reference evaluator and benchmark. On the other hand, base LMs, single-format trained LMs, and jointly trained LMs show lower correlations with GPT-4-1106, Claude-3-Opus, and humans, mostly falling below 0.5.
Notably, Prometheus 2 models outperform Prometheus and Auto-J by at least 0.2 units across benchmarks in their correlation with proprietary LMs. Moreover, on the FLASK benchmark, while the correlation between humans and GPT-4 is 0.679, the highest correlation previously achieved by Prometheus-13B with humans was 0.449, but Prometheus-2-8x7B achieves a correlation of 0.555 with humans, effectively halving the gap.
2 Pairwise Ranking Results
The pairwise ranking results are shown in Table 4. We exclude the results of Pair RM, Ultra RM on ‘w/ Tie’ settings since they could not give tie options.
On all of the 4 benchmarks, the Prometheus 2 models achieve the highest scores, showing that they could effectively simulate human judgments. Notably, while HHH Alignment is an in-domain test set for Pair RM, and Auto-J Eval is for Auto-J, Prometheus-2-8x7B achieves higher scores. This shows that training a large LM (i.e., Mixtral-8x7B) with feedback data could be an effective strategy to obtain a robust evaluator LM that could generalize beyond its training data. Moreover, the Prometheus 2 models at least halve the performance gap with proprietary LMs compared to existing evaluator LMs on out-of-domain test sets.
3 Consistency Across Evaluation Formats
In addition to obtaining high correlation and accuracy, achieving high consistency is another important aspect for evaluator LMs. Specifically, we conduct an experiment testing if evaluator LMs could achieve consistent scores across different evaluation formats. To do this, we use pairwise ranking benchmarks and measure the performance differences when prompted with direct assessment formats and pairwise ranking formats. Specifically, following Kim et al. (2023), to process pairwise ranking datasets in a direct assessment scheme, we evaluate each response separately and compare the scoring decisions. We mark it as correct if the evaluator LM provides a higher score for the human-chosen response over the rejected one. As shown in Table 5, the results show that Prometheus 2 models show lower performance differences across evaluation formats, indicating their robustness.
Discussions
To understand the effectiveness of our proposed weight merging method in the context of evaluations, we address the following research questions:
RQ1: Is Weight Merging more effective compared to Joint Training? (Section 6.1)
RQ2: Is the effectiveness of Weight Merging due to model ensembling? (Section 6.2)
RQ3: To what extent does learning with direct assessment help pairwise ranking performance, and vice versa? (Section 6.3)
Table 6 compares the performance of evaluator LMs trained via weight merging and joint training. Alongside this, we also add and compare the results of prompting and single-format training.
Surprisingly, we observe that evaluator LMs trained via joint training often show lower performance compared to single-format trained evaluator LMs, which indicates negative task transfer. Specifically, evaluator LMs trained only on direct assessment formats obtain higher correlations compared to jointly trained evaluator LMs across different model scales. Similarly, evaluator LMs trained only on pairwise ranking formats obtain higher average accuracy compared to multi-task trained evaluator LMs when using Mixtral-8x7B as a base model.
On the other hand, evaluator LMs trained via weight merging show superior performance not only compared to jointly trained evaluator LMs but also single-format trained evaluator LMs, indicating positive task transfer. Also, while both benefit each other, merging the pairwise ranking evaluator LM weights improves direct assessment performance more significantly than the reverse.
2 Is the Effectiveness of Weight Merging due to Model Ensembling?
While we empirically find that weight merging works effectively, it is unclear what might be the reason. One natural assumption might be that weight merging works effectively due to the effect of ensembling multiple models. To check the validity of this hypothesis, we conduct an ablation experiment by training multiple evaluator LMs on different random seeds and merging them. Specifically, we merge two evaluator LMs trained on direct assessment formats (denoted as ‘Direct Assessment & Direct Assessment’) and two evaluator LMs trained on pairwise ranking formats (denoted as ‘Pairwise Ranking & Pairwise Ranking’). We use Mistral-7B-Instruct as our base model.
Results are shown in Table 7. Against our expectations, we observe that in the majority of cases, merging evaluator LMs trained on the same evaluation format does not improve evaluation performances. Specifically, on direct assessment benchmarks, merging two evaluator LMs trained on direct assessment harms performance on average. Similarly, on pairwise ranking benchmarks, merging two evaluator LMs trained on pairwise ranking also harms performance on average. In contrast, by merging two evaluator LMs each trained on direct assessment and pairwise ranking formats, the resulting evaluator LM shows superior performance compared to different settings. This suggests that the positive task transfer that occurs from weight merging comes from unifying different evaluation formats, not by ensembling multiple models.
3 Quantifying Positive Transfer across Evaluation Formats
To explore how training on direct assessment feedback data influences pairwise ranking accuracy and vice versa, we experiment by adjusting the value during linear merging. We evaluate the average performance using all eight benchmarks in our experiments. To illustrate the average performance (colored in black), we adjust the scale by multiplying direct assessment Pearson correlations, originally from 0 to 1, by 100 before averaging with pairwise ranking accuracy.
The results are shown in Figure 3. For direct assessment benchmarks, evaluator LMs obtain the optimal performance when is set to 0.5. This indirectly indicates that both pairwise ranking and direct assessment feedback data contribute equally. On the other hand, for pairwise ranking benchmarks, the performance is optimal when is set to 0.3. This also indirectly implies that while both benefit each other, training on pairwise ranking improves direct assessment performance more significantly than the reverse.
Conclusion
We introduce Prometheus 2, an open-source language model specialized in evaluating other responses. Unlike existing open evaluator language models that cannot effectively process both direct assessment and pairwise ranking—the two most prevalent evaluation schemes— the Prometheus 2 models demonstrate superior performance and consistent results on both schemes, significantly narrowing the gap with proprietary LM-based evaluations. To train the Prometheus 2 models, we develop the Preference Collection, the first pairwise ranking dataset that includes over 1,000 instance-wise evaluation criteria beyond basic qualities such as helpfulness and harmlessness. Notably, we find that merging evaluator LMs trained on either direct assessment or pairwise ranking formats can lead to a unified evaluator LM with strong performance. We hope that our work encourages more research on using open-source language models as evaluators, moving away from reliance on proprietary models for fair and accessible evaluations.
Acknowledgements
We thank Sungdong Kim, Seonghyeon Ye, Sohee Yang, Dongkeun Yoon, and Hyeonbin Hwang for their helpful comments and discussions.
References
Appendix A Quality Verification of the Preference Collection
To ensure the quality of the Preference Collection, particularly the generated verbal feedback , we employ five annotators with backgrounds in natural language processing. We randomly sample 200 instances with different instructions and conduct a three-part verification process. First, we assess the coherence of with the scoring decision (i.e., ’A is better’ or ’B is better’). Second, we evaluate the suitability of against the evaluation criteria . Lastly, to determine the criticality of the feedback, we compare the newly generated with a concatenation of and . This aims to determine if effectively leverages the mutual information between and . Annotators then vote on whether or the concatenation of and is more critical. The results are shown in Table 8.
Appendix B Training and Inference Hyperparameters
The configurations we used for prompting and training evaluator LMs are shown in Table 9, 10, 11. For Auto-J, PairRM and UltraRM, we utilize their prompt template, inference hyperparameter mentioned in the model cards or github repositories in order to ensure the configuration is optimal for a fair performance comparison. For proprietary LMs, Prometheus 1, and Prometheus 2 models, we use the same prompt template and evaluation configurations.
Appendix C Direct Assessment Results: Extended
Table 12 and 13 (on the next page) shows the extended results Table 3. Even when changing the metrics to either Kendall-Tau and Spearman, the overall trends are maintained. Prometheus 2 shows superior evaluation performances among the open evaluator LMs, achieving high correlations with humans and proprietary LMs.
Appendix D Consistency Experiment Results: Extended
We test if evaluator LMs could give consistent scoring decisions in direct assessment formats. We inferencing multiple times with non-deterministic decoding (e.g., using temperature 1.0). Following the experimental design from Ye et al. (2023), we choose to inference 3 times and report the Krippendorff’s alpha value. As shown in Table 14, the results indicate that training on feedback data only slightly improves consistency. On the other hand, we find that the LMs with a large number of parameters achieve high consistency. This indicates the importance of selecting a large LM as the base model when training an evaluator LM. Notably, Prometheus-2-8x7B obtains the highest correlation among open evaluator LMs.
Moreover, to evaluate consistency in pairwise ranking settings (Table 15), we measure transitivity (i.e., a higher score for response B over A, and for C over B, results in a higher score for C over A). As shown in Table 15, the Prometheus 2 models achieve performances on par with GPT-4, showing that they could provide robust judgments in pairwise ranking schemes.
Appendix E Merging Method Ablation
In Table 16, we try different merging methods introduced in our previous section. We empirically find that merging evaluator LMs with Task Arithmetic (Ilharco et al., 2022) and TIES merging (Yadav et al., 2024) constantly results in a model that degenerates. On the other hand, for the Mistral-7B based evaluator LMs, we find that linear merging and DARE merging (Yu et al., 2023) results in a model that does not degenerate and could process both evaluation formats. Also, for Mixtral-8x7B based evaluator LMs, we find that only DARE merging works effectively for both base models.