The Devil is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André F. T. Martins, Graham Neubig, Ankush Garg, Jonathan H. Clark, Markus Freitag, Orhan Firat

Introduction

Evaluating natural language generation systems has always been challenging, and as the output quality of these systems has improved, evaluation has become even more challenging and critical. For example, in Machine Translation (MT), a field where evaluation has garnered considerable attention, previous standard automatic surface-level metrics such as BLEU Papineni et al. (2002) are becoming less reliable as the quality of generation systems improves, with little remaining correlation with human judgments Freitag et al. (2022).

To keep pace with the constantly improving quality of MT output, the next generation of automatic metrics is rapidly evolving. Learned automatic metrics that leverage human-judgments to finetune language models Sellam et al. (2020); Rei et al. (2022a) currently represent the state-of-the-art in automatic evaluation benchmarks like the WMT Metrics task Freitag et al. (2022), and show high correlation with human judgments. However, these metrics typically output a single, uninterpretable quality score, making it difficult to understand the type and extent of errors identified by them. The lack of insights makes it difficult for model developers to leverage these metrics to improve their systems.

Unlike automatic metrics that only provide a single scalar value as quality score, state-of-the-art human evaluation methodologies like Multidimensional Quality Metrics (MQM; Lommel et al., 2014; Freitag et al., 2021a) ask professional annotators to identify and label error spans with a category and severity. This much richer feedback can be used to gain a better understanding of the current limitations of the model under evaluation and improve it.

In this paper, we ask whether large language models (LLMs) in combination with a few human annotations can be used to design an automatic metric that generates rich feedback similar to that generated by human experts in MQM. This work is motivated by recent papers that demonstrated that LLMs can be used as automatic metrics Liu et al. (2023b) to generate a single quality score. In particular, Kocmi and Federmann (2023) showed that LLMs can be prompted to assess the quality of machine-generated translations, even achieving state-of-the-art performance on assessing system-level quality. However, previous work only provides a limited view of the capabilities of LLMs for machine translation evaluation: the focus has predominantly been on score prediction (i.e. predicting a numerical value for quality), without considering the use of any annotated data (either through in-context learning or finetuning), and only in high-resource language pairs.

We provide a large-scale study of the capabilities of LLMs (from the PaLM and PaLM-2 families; Chowdhery et al., 2022; Anil et al., 2023) for machine translation evaluation (both with and without a reference translation), provide a novel comparison between prompting and finetuning, and investigate the performance in the low-resource scenario. Inspired by findings that the performance of LLMs can be improved by prompting them for rationales of their predictions Wei et al. (2022); Lu et al. (2023), we also propose AutoMQM, a prompting technique for MT evaluation that asks LLMs to identify error spans in a translation and to classify these errors according to the MQM framework, with a quality score derived automatically from the identified errors. A key advantage of AutoMQM is its interpretability, as users can inspect the errors responsible for a score (Figure 1).

Our contributions can be summarized as follows:

We confirm the finding of Kocmi and Federmann (2023) that LLMs are zero-shot state-of-the-art system-level evaluators, but show low correlation with human judgment compared to learned metrics at the segment-level.

We show that finetuning an LLM with human judgment mitigates its low segment-level performance (particularly for smaller LLMs), showing similar correlations with human judgment at both the system-level and segment-level to state-of-the-art learned metrics.

We are the first to evaluate LLM-based evaluation methods on low-resource language pairs. We find that their performance is promising, but lags behind state-of-the-art learned metrics.

We find that, with AutoMQM, PaLM-2 models can be prompted to generate rich MQM-like annotations, outperforming their score prediction counterparts at the segment-level.

Furthermore, annotations predicted by PaLM-2 models correctly identify over 50% of words that are part of major errors, and are comparable to the ones produced by state-of-the-art supervised word-level evaluators.

Our findings might have significant implications for not only MT evaluation, but evaluation of machine-generated text in general, and further highlight the potential of using LLMs to provide AI Feedback Fernandes et al. (2023).

Background: MT Evaluation

Machine translation evaluation is one of the most well-studied evaluation problems in NLP Callison-Burch et al. (2008); Freitag et al. (2022). In this task, given

a candidate translation in a (target) language

an evaluation metric assesses the quality of the candidate translation by how well it conveys the meaning of the source sentence while considering other factors like fluency. Like many other natural language generation evaluation problems, this task is difficult because the set of correct translations for a given source sentence is often very large and not entirely known in advance. To simplify the problem of machine translation evaluation, often (3) a reference translation (typically created by a professional human translator) is included as additional information when assessing the candidate translation. This sub-problem is known as reference-based evaluation (as opposed reference-less evaluation or quality estimation).

Up until recently, human evaluation of machine translation was carried out predominantly with the aim of assigning a single quality score to a candidate translation. Consequently, learned metrics, which leverage collected human judgment data, are trained for and evaluated on the same task of score prediction (i.e., assigning a single quality score to a candidate translation), and can achieve high correlation with human-provided scores Freitag et al. (2022).

However, framing machine translation evaluation as a score prediction task is problematic: any scoring or ranking of translations is implicitly based on an identification of errors in the candidate translations, and asking raters to solely provide a single score can lead to rushed and noisy judgments Freitag et al. (2021a).

This insight has led to the adoption of the Multidimensional Quality Metrics (MQM) framework Lommel et al. (2014); Freitag et al. (2021a) as the gold standard for evaluating machine translation. The MQM framework asks human evaluators to identify error spans in candidate translations and classify those errors according to various dimensions, e.g., fluency, accuracy, … (see Appendix A for a more detailed description of MQM). Importantly, the MQM framework does not ask annotators to provide a quality score for each translation, and instead derives one automatically from the identified error spans and their classifications. However, despite its richness, most automatic metrics that leverage MQM data only use the final quality score produced by the framework and discard the error span information and classification.

Related Work

The success of learned machine translation metrics Sellam et al. (2020); Rei et al. (2022a); Freitag et al. (2022); Qin et al. (2022), which finetune neural network models pretrained on large amounts of (unsupervised) data, highlighted the importance of leveraging transfer learning to achieve metrics with better correlation with human judgments. More recently, generative LLMs OpenAI (2023); Anil et al. (2023) have consistently demonstrated impressive results in natural language understanding and zero- and few-shot transfer and, naturally, interest in employing these models for (translation) evaluation has increased. Kocmi and Federmann (2023) first explored the use of GPT models for evaluating machine translation tasks, showing their potential as zero-shot evaluators, and others have since extended GPT-based evaluation to other generation problems Jain et al. (2023); Liu et al. (2023b).

Perrella et al. (2022) first highlighted that MQM annotations could be leveraged to allow pretrained models to predict major and minor errors and, similarly to AutoMQM, used the identified errors to automatically score translations. However, their approach relied on weaker encoder-only or encoder-decoder language models, required supervised data to work, and overall underperformed other top metrics. We compare against their MaTASe metric in our experiments. Lu et al. (2023) showed that doing error analysis, a prompting technique similar to AutoMQM, could lead to better ChatGPT-based evaluators. However, they still relied on the LLM to provide a score once it identified errors (rather than do it automatically using something like the MQM framework). Furthermore, they provided a very limited meta-evaluation using only 40 examples per language pair. Concurrently with our work, Xu et al. (2023) proposed InstructScore, a LLaMA-based evaluator that asks models to identify and categorize errors in translation (as well as providing a natural language explanation for each error). However, the authors only explore a 7B parameter model and don’t leverage zero- and few-shot capabilities of models as in this work. Instead, they rely on a more complex approach of distilling the knowledge of a more capable GPT-4 LLM.

Additionally, WMT Word-Level Quality Estimation shared tasks Fonseca et al. (2019); Zerva et al. (2022) leverage MQM data by converting span-level annotations of errors (normally of major severity) to word-level tags and Task 2 in the WMT19 Quality Estimation shared task evaluation explicitly evaluated submissions of span-level annotations (although most submissions still consisted of models that predicted word-level tags which were converted to spans). We also compare against state-of-the-art word-level quality estimation models.

Using LLMs to Predict Quality Scores

Recent works have shown that large language models are versatile, general-purpose models that can be used to tackle many problems in NLP, including evaluation Kocmi and Federmann (2023); Jain et al. (2023); Liu et al. (2023b). We begin by exploring how LLMs can be used for machine translation evaluation through score prediction.

We start by measuring how far we can push the performance of LLMs with just prompting Liu et al. (2023a): by defining the task of MT evaluation and quality estimation as textual templates (with a general description of the problem and “slots” for the inputs and outputs), we can use general-purpose LLMs to perform these tasks at inference-time, without any parameter updates.

Throughout the paper, we choose to use Kocmi and Federmann (2023)’s GEMBA-SQM prompt (Figure 2), which asks models to generate (a string representation of) a score from 0-100. We choose this prompt for two reasons: firstly, early explorations with theirs and other prompts showed that this generally performed well. Secondly, using a single prompt ensures a fairer comparison between the capabilities of different models.While this prompt wasn’t the best for system-level, it led to the best segment-level performance in GEMBA.

A surprising emergent capability of LLMs is their ability to improve on prompting-based tasks by including a very small amount of labeled data as part of the prompt/context Brown et al. (2020) and without parameter updates, a technique called in-context learning (ICL). We thus investigate the impact that ICL has on LLMs’ ability to assess translation quality. Recent works have shown that the impact of ICL is tightly tied with the exact examples included in the prompt, with a poor selection procedure leading to no improvements or even worse performance than the zero-shot case Jain et al. (2023). We therefore explore two sampling approaches to select in-context examples from a pre-defined “pool” of translation quality assessments: uniform sampling and stratified sampling, where the example pool is bucketed by score ranges and examples are sampled from each bucket.

2 Finetuning

It has previously been shown that LLMs are capable of zero-shot evaluation Kocmi and Federmann (2023), but the extent to which finetuning on human judgment data can further boost the performance of LLMs has not been studied. In the WMT’22 Metrics Shared Task (Freitag et al., 2022), all top submissions were learned metrics; that is, pretrained models finetuned on human judgment dataWhile these metrics all leverage powerful pretrained (language) models, these generally aren’t considered LLMs.

Thus, we investigate whether LLMs are amenable to finetuning on human judgment data. LLMs used in top-performing metrics are generally much larger than the pretrained language models leveraged by previous learned metrics (which generally have fewer than 1 billion parameters). Moreover, most learned metrics leverage pretrained encoder-only rather than (decoder-only) prefix language models. We experiment with finetuning LLMs using two objectives:

Regression (R): Commonly used for training learned metrics Rei et al. (2022a), the objective here is a regression loss (e.g., mean squared error) between continuous scores obtained from the model (for example, with a regression head) and the human scores.

Generative Classification (GC): We bucket scores into discrete classes (see §6.1) and treat the MT evaluation task as a text-to-text classification problem Raffel et al. (2020).

Using LLMs to Predict Error Spans

While producing quality scores that correlate with human judgments is an important part of translation quality assessment, metrics that solely do score prediction suffer from problems of interpretability: if a metric assigns a low score, the downstream users are left in the dark about which parts of the translation were responsible for the score and thus need to be corrected. This is especially problematic in cases where the metric assigns a wrong score to a translation, as it is much harder to diagnose why the evaluation model made a mistake, and identify and prevent similar mistakes in the future. In fact, reducing translation quality to a single score has proven problematic even for human annotators: asking raters to solely provide a single score can lead to rushed and noisy judgments Freitag et al. (2021a) and the current gold standard for translation quality evaluation involving human annotators is instead based on methodologies like the MQM framework (see §2) , which provide richer feedback by identifying error spans, categorizing them, and evaluating their severity.

Interestingly, another emergent phenomenon in LLMs is the success of chain-of-thought prompting Wei et al. (2022): when defining a prompt for a particular task, if we instruct the model to produce a series of intermediate reasoning steps (“let’s think step-by-step”), it tends to generate a free-text rationale before generating an output, and this often improves the performance on the task at hand Liu et al. (2023b). Furthermore, this chain-of-thought prompting can be used to obtain structured rationales from LLMs, and this can lead to better performance than with free-text rationales Lu et al. (2023).

Motivated by these findings, we propose AutoMQM, a prompting technique for translation quality assessment that instructs LLMs to identify errors in a translation, and categorize the type of error according to the MQM framework Lommel et al. (2014). Furthermore, we don’t ask the model to produce a score, as the MQM framework provides an algorithmic procedure to obtain one from identified errors: the total score is the sum of penalties for all errors identified, where (roughly) major errors get penalized with −5-5 and minors with −1-1 (see Appendix A for a more detailed description of the scoring algorithm).This is similar to methods that leverage external executors to improve the performance of LLMs Gao et al. (2022) Figure 3 shows the main AutoMQM prompt used in this paper.

Importantly, obtaining meaningful AutoMQM results in a zero-shot setting is a substantially more challenging task compared to score prediction: we found that, without any in-context examples, LLMs tend to produce outputs that are either uninformative or difficult to parse. Thus we only consider the AutoMQM task in the few-shot scenario. Based on the findings from §6.2, we explore the impact of in-context learning by sampling from the example pool using stratified sampling extended with a set of rejection criteria (Appendix B), which ensures that the example set has a balance between major and minor errors as well as diversity in the categories of errors.

Experiments

The metrics in this work are evaluated on both high-resource and low-resource language pairs. The three high-resource language pairs come from the WMT’22 Metrics Shared Task (Freitag et al., 2022): en→\rightarrowde, zh→\rightarrowen, and en→\rightarrowru. The ground-truth translation quality scores are derived from MQM ratings in which expert annotators marked error spans in the translations with different severity levels which are automatically converted to a numeric score (see §2). The four low-resource language pairs come from the WMT’19 Metrics Shared Task (Ma et al., 2019): en↔\leftrightarrowgu and en↔\leftrightarrowkk. Since MQM ratings are not available for the low-resource pairs, the ground truth quality scores are direct assessment (DA) scores. DA scores are quality assessments assigned by non-expert raters on a scale from 0-100, then normalized per rater. See Table 1 for statistics about the number of MT systems and segments for every language pair.

Additionally, in our experiments, AutoMQM required in-context examples with MQM annotations to work, so we restrict our evaluation of AutoMQM to en→\rightarrowde and zh→\rightarrowen because there are available MQM ratings from the WMT’21 Metrics Shared Task Freitag et al. (2021b) that we can use as in-context learning example pools.

We base most of our experiments on the following LLMs:

PaLM: A 540 billion parameter autoregressive Transformer model trained on 780 billion tokens of high-quality text Chowdhery et al. (2022). It showed remarkable performance on a wide-range of NLP tasks, including Machine Translation Vilar et al. (2022).

PaLM-2: The successor to PaLM, the PaLM-2 family of LLMs Anil et al. (2023) builds upon recent research insights, such as compute-optimal scaling, a more multilingual and diverse pre-training mixture, and architectural/optimization improvements. We mainly use two model sizes in the family: PaLM-2 Bison and (the larger) PaLM-2-Unicorn.Information about exact number of parameters of PaLM-2 models is not publicly available. In addition we explore the impact of instruction-tuning by using a Unicorn model finetuned on the FLAN dataset Wei et al. (2021).

For score prediction, we compare PaLM and PaLM-2 against the GPT family of LLMs Brown et al. (2020); OpenAI (2023) by leveraging the results and outputs from the GEMBA evaluator Kocmi and Federmann (2023). We then evaluate the performance of AutoMQM with only PaLM-2 models (which performed best in score prediction).

Additionally, for the high-resource languages, we compare to a set of strong baseline evaluation metrics, MetricX-XXL and COMET-22, which were the two top-performing metrics in the WMT’22 Metrics Shared Task. MetricX-XXL and COMET-22 are both finetuned regression models trained on DA data from WMT that are initialized with mT5 (Xue et al., 2021) and XLM-R (Conneau et al., 2020), respectively.

For the AutoMQM experiments, we also compare against MaTESe, a comparable submission to the WMT’22 Metrics Shared task that finetuned a XLM-R model to identify major and minor errors, and computed a score automatically. Since we were unable to obtain the span-level predictions for the MaTESe submission, we also compare against the top submission to the WMT’22 Word-Level Quality Estimation Shared Task Zerva et al. (2021): word-level CometKiwi (COMET-WL) Rei et al. (2022b), also based on an XLM-R model trained on a combination of sentence- and word-level data. To do so, we re-run this model on the WMT’22 Metrics Shared Task data, and convert the predicted word-level OK/BAD tags into spans.We consider a span as any maximal consecutive sequence of words marked as BAD, assigning every span the major severity.

For regression finetuning, we use a real-valued logit, extracted from a fixed index in the first target token’s logit vector, as the quality signal. (In particular, we leverage a special, unused, vocabulary token.) This was the technique used to train MetricX-XXL in the WMT 2022 Shared Task submission (Freitag et al., 2022). The regression-based model was trained on WMT direct assessment (DA) data from the years 2015 through 2020.

For generative classification, we bucket the scores in the training data into five classes, where class boundaries are assigned so that each class contains an equal number of training examples. We then map labels to verbal ratings from the following set, based on their bucket: ["very bad", "bad", "ok", "good", "very good"]. To evaluate the model, predictions are mapped back to integer labels from 1 to 5. Any predictions not containing a substring in the label set are considered invalid and are mapped to 0. We experimented with finetuning on both DA and MQM 2020 (Freitag et al., 2021a) data, and found that the latter performed slightly better.

To assess the impact of model size, we also finetune two additional (smaller) PaLM-2 models, which we call SS and MM, comparing their finetuned and zero-shot performance.We use a small variation of the zero-shot prompt, asking models for scores from the same 5 buckets used in finetuning.

The quality of an automatic evaluation metric is estimated by comparing the agreement between the metric scores and ground-truth quality scores on a large number of translations from different MT systems, a process known as metric meta-evaluation. This work reports three different agreement scores, as follows.

The first is system-level accuracy, which calculates the percent of system pairs that are ranked the same by the metric and ground-truth scores, micro-averaged over a set of language pairs (Kocmi et al., 2021). System-level scores are defined as the average score across all segments.

At the segment-level, the standard correlation that is reported by WMT is Kendall’s τ\tau. However, recent work pointed out problems with Kendall’s τ\tau with respect to ties (Deutsch et al., 2023). In short, different variants of τ\tau are inconsistent with respect to ties and even biased against metrics that predict ties, as our metrics do in this work. Deutsch et al. (2023) recommend reporting a pairwise accuracy score, which rewards metrics for correctly ranking translations as well as correctly predicting ties, in combination with a tie calibration procedure that automatically introduces ties into metric scores so that the meta-evaluation is fairer. This accuracy score, denoted acc∗, ranges between 0 and 1, and a random metric would achieve 33% accuracy. We report the “group-by-item” variant of the pairwise accuracy score from Deutsch et al. (2023) in addition to Pearson’s ρ\rho, a complementary signal to rank-based correlations that measure the strength of the linear relationship between two variables (and one of the standard correlations reported in WMT).

Since AutoMQM provides not only scores but also the identified error spans, we can compare the predicted spans with the errors marked by annotators in the MQM annotations. We evaluate quality of predicted spans using: (1) Span Precision (SP), which measures the overlap of predicted spans and gold (annotated) spans; and (2) Major recall (MR), which captures the percentage of gold major errors that were predicted as errors (either minor or major).

Intuitively, we care for overall precision (regardless of severity) since we want to make sure predicted errors tend to be marked by annotators as well, but for recall we care mostly for major errors, as these have a larger impact on translation quality and are more critical to identify. Additionally, we also report the (3) Matthews Correlation Coefficient (MCC), one of the official metrics in the word-level quality estimation tasks Zerva et al. (2022).

2 Results

Table 2 summarizes the meta-evaluation results, at the system and segment level, for both the zero-shot prompting and finetuning settings.

A first observation is almost all zero-shot LLM evaluators have higher system-level performance than learned metrics (with and without references), with PaLM 540B and PaLM-2 Unicorn achieving the best performance. At the segment level, the story is more complicated: similarly to Kocmi et al. (2022), we find that none of the LLMs we explored was able to consistently outperform the baseline learned metrics. We see that PaLM-540B is a particularly poor reference-based evaluator, which is surprising given its system-level performance. Unexpectedly, instruction-tuning with FLAN seems to degrade performance, with FLAN-PaLM-2 Unicorn achieving poor performance at both the system and segment levels.Note that this might be a problem with the FLAN dataset and not instruction-tuning in general, as the GPT models are also instruction-tuned and perform well.

Figure 4 shows the distribution of scores produced by PaLM- and PaLM-2-based evaluators. We find that, despite being prompted to give a score in the 0-100 range, these models almost always output one of a very limited set of scores (e.g. 0, 50, 90, 95). Given Kocmi and Federmann (2023)’s similar findings with GPT models, it seems that this is a consequence of the pretraining objective.

Despite their already-great performance in the zero-shot setting, we find that finetuning LLMs can further improve LLM evaluators’ segment-level scores. This is particularly obvious for the reference-less evaluators, where a finetuned PaLM-2 Bison achieves state-of-the-art performance in segment-level correlations and comparable system-level accuracy across all language pairs. Moreover, when we look at how performance scales with parameter count (Figure 5), we observe an interesting trend: while smaller models are not capable of being effective zero-shot evaluators, finetuning them leads to competitive performance, and only a slight decrease when compared to their larger finetuned counterparts.

Figure 6 shows the mean and interquartile range (IQR) of the performance as we increase the number of in-context examples kk (with 100 example sets per kk) sampled with stratified sampling (see Appendix C for uniform). Surprisingly, despite evidence of the benefits of in-context learning for many tasks, we found that including in-context examples during evaluation (almost) never led to better performance, either with uniform or stratified sampling.

To investigate the cause of this disappointing performance, we looked at how particular in-context example sets affect the distribution of scores produced by LLM-based evaluators. Figure 7 shows the distribution of scores over the whole test set for the 1-shot and 2-shot settings, with different in-context examples sets. We can see that output distribution is heavily biased by the scores in the in-context examples: despite never predicting 79 in the zero-shot setting, when a single example with that score is included, it starts to dominate the model predictions. This seems to hint that LLMs “overfit” to the specific scores provided as examples, rather than generalizing to the broader evaluation task, which could explain the lackluster performance of in-context learning.

3 Low Resource Languages

Table 4 shows the performance of PaLM-2 models at score prediction for low-resource translation. Overall, we find that similar to high-resource LPs, these models are good zero-shot evaluators, with system-level accuracies around 90%. However, zero-shot LLMs underperform learned metrics, even when these metrics also weren’t exposed to data in these low-resource languages.

Figure 8 shows the mean and interquartile range (IQR) of the performance of PaLM-2 Bison with AutoMQM, as we increase the number of in-context examples (again, with 100 example sets per kk). Contrary to the performance with score prediction, we find that performance with AutoMQM seems to (mostly) scale with the number of in-context examples: performance increases monotonically with up to 4 in-context examples and plateaus thereafter. Additionally, the variance across the in-context learning sets seems to be lower, with most example sets exhibiting less than 0.05 Pearson difference from the best-performing sets. All this suggests that LLM evaluators are much more robust to the choice of in-context examples when prompted for AutoMQM rather than for score prediction. We also find that the behavior of in-context learning is quite similar for both reference-based and reference-less evaluation tasks. Finally, we observe that the example sets that perform well for one task generally work well for the other, with performance on both settings given a fixed in-context set being highly correlated, as shown in Figure 9.

Table 5 shows the meta-evaluation results for PaLM-2 Bison and Unicorn prompted with AutoMQM (using the best-performing in-context learning sets in Figure 8). For ease of comparison, we also report their performance when prompted for score prediction, as well as the performance of the baselines. Overall, prompting LLMs with AutoMQM seems to lead to significant improvements in evaluating machine translation quality, particularly for larger models: Unicorn achieves better performance (across all meta evaluations) with it than when prompted for score prediction, and its reference-less version is competitive with the best learned metric even at the segment level. However, for the smaller Bison, the benefits of AutoMQM are less clear, with both techniques performing comparably. This hints that scale is necessary for zero- and few- shot fine-grained evaluation (like with AutoMQM). We also find that the distribution of scores produced by LLMs prompted with AutoMQM is much closer to the gold MQM distribution, with models outputting a much larger set of scores, and in the same ranges as annotators do (see Figure 10).

Finally, when evaluating the error spans produced by LLMs prompted with AutoMQM (Table 6), we find that PaLM-2 models are able to identify most of the major errors. However, it does seem to over-predict errors (with errors predicted by Unicorn having on average ∼\sim5 words per span vs ∼\sim2 words in the ground truth) and have overall low span precision. Similarly to overall score correlations, scale also seems to be important for the quality of spans produced by AutoMQM, with Unicorn outperforming Bison at most metrics. Additionally, Unicorn prompted with AutoMQM predicts spans of comparable quality to the ones produced by current state-of-the-art learned word-level evaluators (trained on a considerable number of fine-grained annotations derived from MQM): while word-level models are more precise, their overall span correlation (MCC) is comparable, and they miss considerably more major errors than LLMs (despite only leveraging a handful of annotations).

Conclusion

In this study, we have systematically investigated the capabilities of large language models for machine translation evaluation through score prediction, and proposed AutoMQM, a novel prompting technique that leverages the Multidimensional Quality Metrics (MQM) framework for interpretable MT evaluation using LLMs.

We demonstrated that just prompting LLMs for score prediction leads to state-of-the-art system-level evaluators, but still falls short of the best learned metrics at the segment-level (with finetuning being necessary to close this gap). Then we showed that AutoMQM can further improve the performance of LLMs without finetuning while providing interpretability through error spans that align with human annotations.

Our findings surrounding finetuning LLMs for score prediction hint that LLMs’ performance in machine translation evaluation could be further improved by finetuning these models on fine-grained human judgment data (like MQM) and is a direction we are actively pursuing. Additionally, the general-purpose nature of LLMs may enable the application of similar prompting techniques (leveraging some fine-grained evaluation schemes) to other evaluation problems Wu et al. (2023).

Acknowledgements

We would like to thank Ricardo Rei, Marcos Treviso and Chryssa Zerva for helping run the word-level QE baselines, and George Foster who provided feedback on an earlier version of this work. This work was partially supported by EU’s Horizon Europe Research and Innovation Actions (UTTER, contract 101070631), the P2020 program MAIA (LISBOA-01-0247-FEDER-045909), the Portuguese Recovery and Resilience Plan, and the Fundação para a Ciência e Tecnologia through contracts SFRH/BD/150706/2020 and UIDB/50008/2020.

References

Appendix A Multidimensional Quality Metric (MQM)

The Multidimensional Quality Metrics (MQM) framework is a flexible human-evaluation framework developed to evaluate and categorize errors in translations. Annotators are instructed to identify all errors within each segment in a document, paying particular attention to document context. See Table 7 for the annotator guidelines provided.

Annotators are asked to assign both an error severity and category. Error severity (either major or minor) is assigned independently of category. Spans with no marked errors have neutral severity and no category. Possible error categories are displayed in Table 8.

Since MQM doesn’t ask annotators for quality scores, those scores are derived automatically from the identified error spans and their classifications, based on a weighting of each error severity and category. Table 9 summarizes this weighting scheme, in which segment-level scores can range from 0 (perfect) to 25 (worst). The final segment-level score is an average over scores from all annotators. In some settings (e.g. calculating correlation for learned metrics), the scores are negated.

We use the same weighting to obtain scores from errors identified by AutoMQM.

Appendix B Sampling in-context learning examples for AutoMQM

Figure 11 shows the rejection criteria used when sampling example sets as discussed in §4.

Appendix C Additional Results

Figures 12, 13, and 14 present additional experimental results.