Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration
Daniel Deutsch, George Foster, Markus Freitag
Introduction
Kendall’s is a widely used correlation statistic (Kendall, 1938). It is easy to grasp intuitively, being based on pairwise rank ordering. This makes it complementary to other well-known statistics such as Pearson or Spearman.
In the context of machine translation (MT), Kendall plays a key role assessing the performance of evaluation metrics, a process known as meta-evaluation: it has been the main statistic for measuring a metric’s ability to score segment-level translations in the Workshop on Machine Translation (WMT) metrics shared tasks over the years (Freitag et al., 2022b, inter alia).
Several recent developments in MT—common to other areas of generative AI—have highlighted an important weakness in Kendall’s , namely how it deals with ties (§4). First, as MT systems get better, they (1) produce more “perfect” outputs, which get assigned the same score by human raters; and (2) necessitate error-based analyses such as MQM Lommel et al. (2014); Freitag et al. (2021a), which often produce tied scores due to integer error counts. Second, on the metric side, the use of recently-proposed LLM-based metrics (Kocmi and Federmann, 2023) and metrics that model MQM annotations (Perrella et al., 2022) can also lead to small and discrete score ranges that assign many ties.
In this paper, we examine the problems caused by ties in Kendall’s , using data from the WMT metrics tasks. We first show that there are simple phenomena that are not handled properly by any of the existing Kendall variants, which mostly differ in how they treat ties (§5.1). We also demonstrate the possibility of gaming the meta-evaluation by exploiting how ties are handled by existing ’s, resulting in large improvements in certain evaluation settings (§5.2).
We propose instead to meta-evaluate metrics with a version of pairwise accuracy that is robust to these problems, assigning proper credit for correctly predicting ties (§6). Although there is a modification to that is closely related to pairwise accuracy, we argue that the accuracy formulation is easier to interpret, being just the proportion of correctly ranked pairs (including tied pairs).
However, pairwise accuracy comes with its own problem, namely that it can discriminate against metrics that rarely assign ties. To counter this, we also propose an algorithm called tie calibration that automatically introduces ties into metric scores in order to optimize its correlation (§7). We argue, and show empirically, that these two modifications result in a fairer assessment of MT metric performance (§8.1).
Finally, we analyze different aspects of pairwise accuracy and tie calibration, including assessing the generalization of tie calibration across datasets (§8.2), the score ranges where ties are introduced (§8.3), and how more fine-grained statistics can be used to better understand metric behavior (§8.4).
While our experimental setting is limited to MT metrics, our work should be applicable to meta-evaluation for other generative AI metrics with similar characteristics.
Background & Related Work
We begin by justifying our exclusive focus on ranking-based statistics, like Kendall’s , then provide some background on MT metric meta-evaluation, and finally contextualize our work by discussing Kendall variants.
Pearson’s and Spearman’s are two other widely-used correlation coefficients. The Pearson coefficient captures linear correspondence between two input vectors, defined as their covariance divided by the product of their variances. Spearman is equivalent to Pearson applied to the ranks of the inputs. As shown in Figure 1, Pearson is complementary to Kendall; it assigns a much higher score to the noisy but globally linear metric1, but a much lower score to the perfectly-ordered but non-linear metric2. Spearman is a compromise, siding with Pearson for metric1 and for Kendall for metric2.
For applications where linear correspondence with a gold standard and correct ranking decisions are both important, it is advisable to measure both Pearson and Kendall, as is typically done in the MT evaluations described below. Although we focus on problems with Kendall here, it is worth noting that Pearson has problems of its own, notably sensitivity to outliers Mathur et al. (2020). For instance, adding the point to metric1 produces an almost perfect correlation of compared to for Kendall.
2 Metric Meta-Evaluation
For over 10 years, the Workshop on Machine Translation (WMT) has run a metrics shared task that meta-evaluates automatic metrics. Meta-evaluation quantifies a metric’s performance by calculating the agreement or correlation between the metric’s scores and human-annotated scores on a large number of translations. In WMT, metrics are meta-evaluated at either the system- or segment-level, as follows.
First, metric and human scores are collected for translations produced by systems for source segments. System-level correlations are calculated between the metric and human scores per system, typically calculated by averaging over the segment scores. In WMT, the system-level correlation is often Pearson, or more recently, a ranking-based pairwise agreement that is similar to our proposed statistic (§6), except that it does not need to account for ties since ties are very rare at the system-level (Kocmi et al., 2021).
Segment-level correlations evaluate metric scores on individual translations rather than aggregated system scores. They can be calculated in several different ways (see Appendix A for equation definitions):
No-Grouping: Calculate the correlation between the translation scores
Group-by-Item: Calculate the average correlation between the translation scores grouped by source segment“Item” is used to keep the terminology generic so it can be applied to other generation tasks. Here, “item” refers to the source segment.
Group-by-System: Calculate the average correlation between the translation scores grouped by system
Segment-level correlations are better than system-level correlations at discriminating between metrics (Freitag et al., 2022b), and they are more closely related to applications where metrics can be used to improve generation, such as Minimum Bayes Risk decoding (Freitag et al., 2022a; Fernandes et al., 2022).
Historically, WMT has evaluated metrics at the segment-level using the group-by-item method, however no-grouping was used in WMT’21 and all three were used in WMT’22. The standard correlation function that is used is some variant of Kendall’s , described next.
3 The Landscape of Kendall’s τ𝜏\tau
Kendall’s is a ranking-based correlation coefficient. Although there are many different variants of , intuitively, it counts how frequently the metric and human scores agree (concordant) or disagree (discordant) on the ranking of all possible pairs of translations. Importantly, there cannot be a tie in either the metric or human score for a pair to be considered concordant or discordant. Each ranges from -1 to 1, with the extremes resulting from the metric and human scores being perfectly discordant/concordant and 0 meaning random chance.
Some variants of are generic and included in libraries like SciPy, whereas others were proposed by WMT metrics shared task organizers and tailored to the application of MT metric meta-evaluation. Table 1 shows the definitions of the different variants of using the notation in Table 2.
The main differences between the variants are how they handle ties. The standard variants, and , are modifications of designed to ensure the values can reach -1 and 1 in the presence of ties. In contrast to our proposal, the versions proposed by WMT do not include ties in the human scores, and penalize ties in the metric scores. This is due to the fact that the metrics shared task organizers either did not want to penalize small differences in metric scores when the human score is tied (Callison-Burch et al., 2010) or only evaluated on pairs that had a large difference in DA score in order to ensure the pair’s ranking was reliable (Bojar et al., 2017).
Overall, none of the ’s directly rewards the metric for correctly predicting ties in the human score. We view our work as a next step in updating the meta-evaluation to account for properties of today’s metrics and human scores.
Analysis Setup
Our analysis is performed on the Multidimensional Quality Metrics (MQM; Lommel et al., 2014; Freitag et al., 2021a) ratings collected by the WMT’22 metrics shared task (Freitag et al., 2022b) for three language pairs: ende, zhen, and enru. We use the MQM scores as the ground-truth human scores that the automatic metrics’ scores are evaluated against. The language pairs have 13-15 systems and around 1300-1900 segments per system with MQM ratings.
Automatic Metrics
We explore how the choice of meta-evaluation statistic affects the rankings of the primary metric submissions to the WMT’22 shared task, in addition to the recently proposed GEMBA metrics (Kocmi and Federmann, 2023). We also discuss and examine various different metrics in more detail, including the top 2 performing metrics in the WMT’22 shared task, Metric-X and COMET-22 Rei et al. (2022), in addition to BLEURT-20 (Sellam et al., 2020), MaTESe (Perrella et al., 2022), and GEMBA. The former three metrics are regression-based metrics that predict floating point translation quality scores. MaTESe predicts span-level errors that are combined into an overall score based on an error severity weighting. The GEMBA metrics predict quality scores using 0-shot prompting with GPT-3.5 and GPT-4 (Brown et al., 2020). Importantly, the predicted scores from MaTESe and GEMBA tend to come from a small set of values rather than a large range of possible floating point scores, which has significant implications for the number of ties they predict (see §4) and how they are treated by different variants of .
Why Ties are Important
There are several motivations for incorporating ties into a ranking-based meta-evaluation statistic like Kendall’s .
First, ties in human scores from recent WMT shared tasks are much more trustworthy than they were previously. Since WMT’20, the human scores are MQM scores instead of direct assessment (DA) scores. The MQM scores come from expert translators and are more reliable than the crowdsourced DA scores. As such, ties (or minor differences) between scores are more likely representative of actual ties (or minor differences) in translation quality.
Second, ties in MQM scores are very common. For instance, up to 53% of possible pairs in en-de have tied MQM scores (see Table 3), the majority of which have MQM scores of 0, meaning there are no errors in the translations. As the quality of MT systems improves, the number of tied translations is likely to increase since there will be fewer differences between systems. If ties in the MQM scores are removed from the meta-evaluation (as is done by some Kendall variants), we throw away a valuable metric quality signal and lose the ability to discriminate between metrics that reliably detect ties and those that do not (see next section).
Finally, recently proposed metrics, such as MaTESe or those based on large language models (GEMBA) predict a large number of ties (see Table 4). These metrics should be directly rewarded for correctly predicting ties in the human scores, which is not the case with existing Kendall variants.
Shortcomings of Kendall’s Variants
The way ties are handled by existing variants of Kendall’s introduces blind spots in the meta-evaluation and opens the door for metrics to exploit -specific properties to improve correlations. We demonstrate the shortcomings of existing ’s through a motivational example and experimental analysis.
Due to how existing ’s handle ties, they are unable to discriminate between metrics that accurately predict ties and those that do not. Figure 2 contains an example of such an instance.
When considering ties, metric only incorrectly ranks 1 out of the 15 possible pairs, whereas incorrectly ranks 6 pairs. However, because existing ’s do not give credit to metrics for correctly predicting ties, the correlation coefficients either consider to be either approximately equal or worse than . This blind spot of existing ’s means they are inadequate for meta-evaluating metrics in the presence of ties.
2 The NaN Problem
Another consequence of how the correlations handle ties is what we refer to as the “NaN problem.” In the event that either the metric or human scores are a constant vector (therefore, all pairs are tied), many of the values are not defined, or NaN. When the segment-level correlation is calculated by grouping by either item or system and one of the groups’ correlations is NaN, the correlation is removed from the average in practice. This happens most often when grouping by item because the size of the input vectors is the number of systems, , which is generally rather small (15).
A metric could take advantage of this property of the segment-level correlation by introducing ties for difficult-to-score groups, resulting in NaN scores. This has the effect of removing the challenging groups from the meta-evaluation, resulting in higher correlations.Another possibility would be to assign a neutral correlation value of 0. However, this has the disadvantage of penalizing metrics that assign ties when all human scores are also tied or are close to being tied. Note that both Pearson and Spearman are also NaN for constant vectors and therefore are also susceptible to gaming. Indeed, we find that this is possible.
To introduce ties, we mapped Metric-X’s scores to integers by assigning each score to an equal-width bucket. This bucketing results in ties in challenging pairs because similar quality translations likely have close metric scores, so when the scores are converted to integer buckets, their scores become the same value. Figure 3 plots the group-by-item (the coefficient used in WMT’22) and the number of non-NaN groups as a function of the number of buckets.
When the number of buckets is small, the number of non-NaN segments is reduced, and the resulting correlations improve over the original values by very large margins. Because the correlations with different numbers of buckets are computed over different non-NaN subsets of the full dataset, their values are not fairly comparable. Indeed, in §8, we demonstrate that WMT’22 metrics submissions were evaluated on different non-NaN groups, and directly comparing their correlations leads to erroneous conclusions.
A metric could have taken advantage of the NaN problem in order to game the WMT’22 metrics shared task since the number of non-NaN segments is not taken into account in the metric meta-evaluation. A method for handling ties that made correlations for constant vectors well defined would close this loophole.
Evaluating with Pairwise Accuracy
Instead of using Kendall’s as the ranking-based meta-evaluation statistic, we propose to use a version of pairwise accuracy that includes ties. We define the pairwise accuracy to be the proportion of all pairs that the metric either ranks correctly or correctly predicts are tied. The equation for our proposal, denoted (“eq” for equality, as in ties), is included in Table 1. This statistic now directly incorporates ties in the human and metric scores.
Although there is a modification of Kendall’s that corresponds to (denoted in Table 1), we advocate for reporting accuracy instead. Accuracy is more intuitive than since its value is between 0 and 1 and it can be read as the proportion of pairs that the metric correctly ranks/predicts as ties. This stands in contrast to that is between -1 and 1, which does not have an easy-to-communicate interpretation. Pairwise accuracy has the additional benefit of aligning how metrics are meta-evaluated at the system- and segment-levels (Kocmi et al., 2021). The results related to metric rankings in this work apply equally to and .
Pairwise accuracy (and ) does not suffer from the same issues as the ’s that were presented in §5: strongly prefers , the metric with fewer incorrectly ranked pairs (Figure 2; §5.1). Because its value is never NaN, it does not suffer from the NaN problem (§5.2); all examples are always used for evaluation.
Pairwise accuracy effectively evaluates the automatic metrics as 3-way classifiers that decide between predicting a tie or one of the two possible rankings for each pair. This formulation nicely allows for further decomposition into class-specific precision, recall, and F1, which can be used to further understand metric performance. Class-specific evaluations help to address a potential class imbalance problem between tied and non-tied pairs that may be hidden by accuracy.
Table 5 contains the definitions of precision and recall with respect to “ties” and “correct ranking.” The “ties” statistics calculate the precision of the metric when predicting a tie and its recall of human ties. The “correct ranking” statistics calculate the proportion of correctly ranked pairs out of all pairs it predicts are not tied and the proportion of all human non-tied pairs correctly ranked by the metric. These additional statistics help provide a more holistic view of metric performance.
Tie Calibration
Although we argue that properly addresses ties in human and metric scores, some metrics do not frequently predict exact ties between translations. Regression metrics, such as BLEURT (Sellam et al., 2020) and COMET (Rei et al., 2020), practically never predict tied scores for two different translations (see Table 4), so they will not able to correctly predict a tie in the human score, putting them at a disadvantage. This is undesirable because it prevents a fair comparison between metrics that do and do not predict ties.
To address this shortcoming, we propose an algorithm called tie calibration for automatically introducing ties into metric scores so that metrics that do and do not predict ties can be fairly compared. The algorithm is based on the intuition that, although regression metrics do not frequently predict ties, the difference between two translations’ scores is sometimes small enough to be considered a tie.
Tie calibration searches for an value that maximizes a rank-based correlation statistic (e.g., or ) such that any two translations with a difference in score less than is considered to be a tie. We experimented with relative differences between scores and found little difference compared to absolute differences. Our implementation considers all possible differences between the pairs of translations as candidates for and selects the one that maximizes the desired ranking-based statistic. The algorithm runs in , where is the number of translations. In practice, when is large, we downsample the number of pairs to consider when searching for , which significantly improves runtime. Experimentally, this appears to be a rather good approximation (see Appendix E). Detailed psuedocode for tie calibration is included in Appendix D.
Because tie calibration introduces an optimal number of tie predictions, metrics are not penalized for under-predicting ties, and therefore metrics that do and do not predict ties can be fairly compared. An added benefit of tie calibration is that the resulting optimal improves the interpretability of metric scores. Its value can be understood as the threshold for which a difference in metric scores should be considered significant (at least with respect to a specific dataset; see §8.2).
Henceforth we use to denote a statistic that has been calculated with tie calibration (e.g., ) and the optimal tie threshold found by the algorithm.
In principle, tie calibration can be used to find an optimal value of any correlation statistic in which the presence of ties changes the value, being one of them. However, care needs to be taken to ensure that the statistic handles ties in a desirable way. For example, omits all ties from its formula, so tie calibration could convert a discordant pair into a tie to improve the value of , which, if the human scores are not tied, is undesirable ( would not reward this change). The combination of tie calibration and a statistic that does not properly handle ties may lead to unexpected results.
Analysis
In this section, we analyze several different aspects related to our proposal of pairwise accuracy and tie calibration. We address the following questions:
§8.1: How does the choice of meta-evaluation statistic affect metric ranking?
§8.2: How does the selected value of generalize across datasets?
§8.3: Does the selected value introduce ties uniformly across score values for a metric?
§8.4: What insights can be drawn from evaluating metrics on predicting tied versus non-tied pairs?
Table 6 shows group-by-item correlations calculated with various ’s and pairwise accuracy. We also report the performance of a “constant metric” that predicts a tie for every pair as a baseline comparison. From the existing ’s, we report and since they are the most recent and used most frequently by WMT. See Appendix C for the results for each language pair, type of segment-level correlation, and correlation statistic. Clearly, the choice of meta-evaluation statistic significantly affects the metric rankings, with the largest changes happening to MaTESe and GEMBA, the two metrics that output the most ties.
Under , GEMBA-GPT-4 and MaTESe are the top ranked metrics. However, this result can be partially explained by the NaN problem (§5.2). MaTESe’s correlation is calculated on 773 non-NaN segments, compared to 1133 for Metric-X. When both metrics are evaluated on the same 773 segments, Metric-X’s correlation is higher (0.296 versus 0.281). This result highlights how correlations calculated on different source segments cannot be fairly compared.
If is used to rank the metrics, MaTESe and GEMBA fall to the bottom at of the ranking. This result can be explained by the fact that is systematically biased against metrics that output a large number of ties because ties are penalized as if they are discordant pairs. Predicting a tie can only decrease a correlation. In fact, the values of MaTESe and GEMBA can be improved by large margins simply by randomly breaking all ties since around half of the pairs will now become concordant, while the other half remain penalized as they were before. For example, randomly breaking ties improves MaTESe and GEMBA-GPT-3.5’s correlations by around 0.5-0.6 points (MaTESe: to 0.15, GEMBA: to 0.15). In contrast, COMET-22’s correlation only improves by 0.005 due to the fact that it predicts few ties (see Table 4).
In contrast, when the metrics are ranked by with tie calibration, denoted , MaTESe and the GEMBA metrics are ranked 4th, 6th, and 15th. Because and are never NaN, all values are fairly comparable. Further, there is no systematic bias for or against ties; Randomly breaking or introducing ties runs the risk of changing a correct prediction of a tie or concordant pair into a discordant pair or an incorrect tie prediction. Clearly, the choice of correlation statistic matters, and we argue that is the most fair and reliable method compared to the variants.
2 Generalization of Epsilon
The previous analysis selected on the same dataset that is used to rank the metrics. Here, we examine what happens if the value is selected on a held-out dataset. For this analysis, the MQM ratings from the WMT’21 metrics shared task (Freitag et al., 2021b) are used as a held-out set.
Figure 4 shows the different and values for BLEURT-20 when is selected on one dataset and applied to the other for en-de and zh-en. For en-de, the epsilon value changes by 0.03, and the calculated on the held-out changes by relative 2%, suggesting the results are rather stable.
However, zh-en behaves quite differently. From the plots, it is clear that for WMT’21, there is almost never an incentive to predict a tie, as evidenced by the very low , and the corresponding does not generalize well to WMT’22 (or vice versa). Our hypothesis is that this result is due to the fact that the WMT’21 zh-en data has far fewer ties than the WMT’22 data (23% versus 41%).
These results indicate that values are not likely to generalize across dissimilar datasets under current metrics. Such a property would be desirable—and an interesting challenge for metric developers—since it would make score differences more interpretable. However, we argue that treating as a latent variable calibrated on the current test set allows for fair comparisons of metrics even in the absence of this property. Other evaluation protocols have also involved optimizations on the test set, for example using an oracle sentence segmenter to evaluate MT for speech (Matusov et al., 2005).
3 Where are Ties Introduced?
Since most of the human score ties are for error free translations (see Table 3), it is worth understanding if the tie threshold introduces ties for high scoring translations to predict error-free translation ties or if the ties are introduced more uniformly across the score distribution.
Figure 5 plots the distribution of the average score per pair where ties are introduced by for Metric-X on the WMT’22 zh-en dataset. In comparison to the distribution of all pairs’ average scores, the tied distribution is skewed toward higher predicted scores. Since Metric-X has a relatively strong correlation to MQM scores, this suggests that the newly introduced ties mostly predict perfect translations, which are assigned high scores according to the metric. An extension of our tie calibration procedure could first identify a threshold to predict a perfect translation, then run tie calibration on the remaining pairs.
4 Class-Specific Statistics
Figure 6 plots the ties-F1, correct-rank-F1 (see §6.1), and pairwise accuracy for COMET-22 on en-de. The ties-F1 is much higher than the correct-rank-F1 for almost every , demonstrating that the metric more reliably predicts tied pairs than the correct rank for non-tied pairs. This is likely due to the fact that the number of perfect translations is large, and the values are biased toward introducing ties to predict perfect translations (§8.3).
If a statistic other than pairwise accuracy is better aligned to how a metric is being used in practice, the tie calibration procedure can be used to select an that strikes the desired balance of performance with respect to the class-specific statistics.
Conclusion
In this work, we demonstrated the importance of taking ties into account when calculating rank-based correlation statistics. We argued existing variants of Kendall’s are inadequate for the current state of meta-evaluation. We advocated to instead use pairwise accuracy, which rewards metrics for both predicting correct pair rankings and correctly predicting ties, in combination with a tie calibration procedure that allows for comparing metrics that do and do not predict ties. Although our experiments were specific to MT, the methods proposed are generally applicable to any metric meta-evaluation in NLP.
Limitations
The tie calibration algorithm introduced in §7 makes an assumption that absolute differences in metric scores reflect the same amount of change in quality for any value of the metric. That is, the difference in predicted quality between translations with scores 0.2 and 0.1 is the same as with scores 100.2 and 100.1. An alternative version of the tie calibration algorithm could introduce ties based on relative differences between scores instead of absolute differences. We experimented with relative differences and did not see a significant different in results. However, it may be that a metric that we did not experiment with performs better with relative instead of an absolute difference.
Since the tie decision operates at the pair level, the value does not induce a global ordering of translations. For example, if there are scores 1, 2, and 3 with , pairs (1, 2) and (2, 3) are tied but (1, 3) is not. A similar limitation can also be observed in pairwise statistical significance testing.
Finally, although we argue that our meta-evaluation proposal is more fair, we are unaware of any way to prove that this is true. Instead, we rely on experimental results and the fact that our proposals are not susceptible to known issues with existing methodologies.
Acknowledgments
The authors would like to thank Juraj Juraska, Mara Finkelstein, Ricardo Rei, Chi-kiu Lo, Tom Kocmi, and Alon Lavie for their helpful discussions and feedback related to this work.
References
Appendix A Correlation Definitions
This section more explicitly defines the three different types of segment-level correlations.
Let and denote the human and metric scores for the translation produced by system on source segment . Define to be a correlation coefficient, such as Pearson’s , Spearman’s , Kendall’s , or any such function that calculates an agreement score over a set of paired observations, like the pairwise accuracy statistic proposed in this work. There are three different segment-level correlations that can be computed.
Appendix B WMT Tabular Notation
WMT’14 (Macháček and Bojar, 2014) developed a tabular notation to describe how Kendall’s was calculated. For completeness, we include a mapping of the notation from this work in Table 2 to the tabular notation in Table 7. The tabular versions of , , and are reproduced in Table 8. The tabular versions of and are included in Table 9.
Re-using the notation from WMT’14, a value can be computed using the tabular notation via the following equation:
is defined as the coefficient in the tabular notation and is the number of pairs that fall into the corresponding bucket.
Appendix C Additional Results
Table 10 contains more statistics related to the number of tied pairs in the WMT’22 MQM scores, including the number of pairs that are tied with a score of 0 (i.e., an error free translation).
The full correlation results and metric ranks according to the different s across different language pairs and segment-level correlations is included in this section. See Table 11 for the listing of the individual tables.
Appendix D Tie Calibration Psuedocode
Algorithm 1 contains the pseudocode for the tie calibration procedure (§7) when applied to two vectors of human and metric scores. The algorithm runs in where is the number of scored translations. The bottleneck is sorting all of the possible pairs. When is too large, we approximate the search for the optimal by downsampling the number of pairs. See Appendix E for an analysis of how lossy this approximation is.
In practice, the tie calibration is applied to matrices of human and metric scores, where each row corresponds to a group (see §2). The algorithm is very similar to Algorithm 1 except there is additional bookkeeping required to match each pair to the group that it came from. The extra bookkeeping only adds an overhead.
Appendix E Epsilon Search Approximation
Finding the exact value that maximizes pairwise accuracy requires considering all possible choices of . For specific segment-level correlations, such as the no-grouping variant, can be prohibitively large, on the order of hundreds of millions of pairs (see Table 3).
When the number of pairs is too large, we instead find an approximate best by sampling from all possible pairs. Figure 7 plots the values calculated on a subset of the data and the value of for those values. Even with as little as 10% of the possible pairs, the approximations are quite precise. The largest observed differences over 30 iterations for were 2.5e-3 and for were 4.3e-5. Overall, downsampling appears to be a safe approximation to improve the run time of the tie calibration algorithm.
Appendix F Unbabel Normalization
The experiments in this paper calculate MQM scores for translations using the normalization technique advocated for by Google: a translation’s MQM score is the sum of the weights of each of the errors. The alternative method used by Unbabel normalizes the sum of error weights by the length of the translation. The Unbabel normalization will thus result in fewer human score ties than the Google normalization.
We repeated the analysis from §8.1 using the Unbabel normalization method and calculated the rankings of the different metrics under and variants for the en-ru language pair. The results for the group-by-item segment-level correlation are shown in Table 21.
Overall, the fewer ties did not make a significant impact on whether or not it was possible to demonstrate that the meta-evaluation statistics are biased toward or against ties. For instance, favors metrics with ties, such as GEMBA-GPT-4, and is still biased against metrics that predict ties. We suspect this is due to the fact that the majority of ties occur for perfect translations, which will remain tied in either normalization method. Further, the number of non-perfect ties (MQM score of 0) only decreased by 7% (from 44% to 37%). Therefore, we argue that the results presented in this work apply to either normalization technique, but larger changes will likely be observed under using the Google method due to the increase in number of ties.