Quality-Aware Decoding for Neural Machine Translation
Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, André F. T. Martins
Introduction
The most common procedure in neural machine translation (NMT) is to train models using maximum likelihood estimation (MLE) at training time, and to decode with beam search at test time, as a way to approximate maximum-a-posteriori (MAP) decoding. However, several works have questioned the utility of model likelihood as a good proxy for translation quality (Koehn and Knowles, 2017; Ott et al., 2018; Stahlberg and Byrne, 2019; Eikema and Aziz, 2020). In parallel, significant progress has been made in methods for quality estimation and evaluation of generated translations Specia et al. (2020); Mathur et al. (2020b), but this progress is, by and large, not yet reflected in either training or decoding methods. Exceptions such as minimum risk training (Shen et al., 2016; Edunov et al., 2018) come at a cost of more expensive and unstable training, often with modest quality improvements.
An appealing alternative is to modify the decoding procedure only, separating it into two stages: candidate generation (§2.1; where candidates are generated with beam search or sampled from the whole distribution) and ranking (§2.2; where they are scored using a quality metric of interest, and the translation with the highest score is picked). This strategy has been explored in approaches using -best reranking (Ng et al., 2019; Bhattacharyya et al., 2021) and minimum Bayes risk (MBR) decoding (Shu and Nakayama, 2017; Eikema and Aziz, 2021; Müller and Sennrich, 2021). While this previous work has exhibited promising results, it has mostly focused on optimizing lexical metrics such as BLEU or METEOR (Papineni et al., 2002; Lavie and Denkowski, 2009), which have limited correlation with human judgments Mathur et al. (2020a); Freitag et al. (2021a). Moreover, a rigorous apples-to-apples comparison among this suite of techniques and their variants is still missing, even though they share similar building blocks.
Our work fills these gaps by asking the question:
“Can we leverage recent advances in MT quality evaluation to generate better translations? If so, how can we most effectively do so?”
To answer this question, we systematically explore NMT decoding using a suite of ranking procedures. We take advantage of recent state-of-the-art learnable metrics, both reference-based, such as COMET and BLEURT Rei et al. (2020a); Sellam et al. (2020), and reference-free (also known as quality estimation; QE), such as TransQuest and OpenKiwi Ranasinghe et al. (2020); Kepler et al. (2019). We compare different ranking strategies under a unified framework, which we name quality-aware decoding (§3). First, we analyze the performance of decoding using -best reranking, both fixed according to a single metric and learned using multiple metrics, where the coefficients for each metric are optimized according to a reference-based metric. Second, we explore ranking using reference-based metrics directly through MBR decoding. Finally, to circumvent the expensive computational cost of the latter when the number of candidates is large, we develop a two-stage ranking procedure, where we use -best reranking to pick a subset of the candidates to be ranked through MBR decoding. We explore the interaction of these different ranking methods with various candidate generation procedures including beam search, vanilla sampling, and nucleus sampling.
Experiments with two model sizes and four datasets (§4) reveal that while MAP-based decoding appears competitive when evaluating with lexical-based metrics (BLEU and ChrF), the story is very different with state-of-the-art evaluation metrics, where quality-aware decoding shows significant gains, both with -best reranking and MBR decoding. We perform a human-study to more faithfully evaluate our systems and find that, while performance on learnable metrics is not always predictive of the best system, quality-aware decoding usually results in translations with higher quality than MAP-based decoding.
Candidate Generation and Ranking
We start by reviewing some of the most commonly used methods for both candidate generation and ranking under a common lens.
An NMT model defines a probability distribution over a set of hypotheses , conditioned on a source sentence , where are learned parameters. A translation is typically predicted using MAP decoding, formalized as
A stochastic alternative to beam search is to draw samples directly from with ancestral sampling, optionally with variants that truncate this distribution, such as top- sampling (Fan et al., 2018) or -nucleus sampling (Holtzman et al., 2020) – the latter samples from the smallest set of words whose cumulative probability is larger than a predefined value . Deterministic methods combining beam and nucleus search have also been proposed (Shaham and Levy, 2021).
Unlike beam search, sampling is not a search algorithm nor a decision rule – it is not expected for a single sample to outperform MAP decoding (Eikema and Aziz, 2020). However, samples from the model can still be useful for alternative decoding methods, as we shall see. While beam search focus on high probability candidates, typically similar to each other, sampling allows for more exploration, leading to higher candidate diversity.
2 Ranking
We assume access to a set containing candidate translations for a source sentence, obtained with one of the generation procedures described in §2.1. As long as is relatively small, it is possible to (re-)rank these candidates in a post-hoc manner, such that the best translation maximizes a given metric of interest. We highlight two different lines of work for ranking in MT decoding: first, -best reranking, using reference-free metrics as features; second, MBR decoding, using reference-based metrics.
In its simplest form (which we call fixed reranking), a single feature is used (e.g., an estimated quality score), and the candidate that maximizes this score is picked as the final translation,
When multiple features are available, one can tune weights for these features to maximize a given reference-based evaluation metric on a validation set (Och, 2003; Duh and Kirchhoff, 2008) – we call this tuned reranking. In this case, the final translation is
2.2 Minimum Bayes Risk (MBR) Decoding
While the techniques above rely on reference-free metrics for the computation of features, MBR decoding uses reference-based metrics to rank candidates. Unlike MAP decoding, which searches for the most probable translation, MBR decoding aims to find the translation that maximizes the expected utility (equivalently, that minimizes risk, Kumar and Byrne 2002, 2004; Eikema and Aziz 2020). Let again be a set containing hypotheses and a utility function measuring the similarity between a hypothesis and a reference (e.g, an automatic evaluation metric such as BLEU or COMET). MBR decoding seeks for
where in Eq. 4 the expectation is approximated as a Monte Carlo (MC) sum using model samples .We also consider the case where are obtained from nucleus sampling or beam search. Although the original MC estimate is unbiased, these ones are biased. In practice, the translation with the highest expected utility can be computed by comparing each hypothesis to all the other hypotheses in the set.
Quality-Aware Decoding
While recent works have explored various combinations of candidate generation and ranking procedures for NMT Lee et al. (2021); Bhattacharyya et al. (2021); Eikema and Aziz (2021); Müller and Sennrich (2021), they suffer from two limitations:
The ranking procedure is usually based on simple lexical-based metrics (BLEU, chrF, METEOR). Although these metrics are well established and inexpensive to compute, they correlate poorly with human judgments at segment level Mathur et al. (2020b); Freitag et al. (2021c).
Each work independently explores -best reranking or MBR decoding, making unclear which method produces better translations.
In this work, we hypothesize that using more powerful metrics in the ranking procedure may lead to better quality translations. We propose a unified framework for ranking with both reference-based (§3.1) and reference-free metrics (§3.2), independently of the candidate generation procedure. We explore four methods with different computational costs for a given number of candidates, .
1 Reference-based Metrics
Reference-based metrics are the standard way to evaluate MT systems; the most used ones rely on the lexical overlap between hypotheses and reference translations Papineni et al. (2002); Lavie and Denkowski (2009); Popović (2015). However, lexical-based approaches have important limitations: they have difficulties recognizing correct translations that are paraphrases of the reference(s); they ignore the source sentence, an important indicator of meaning for the translation; and they do not always correlate well with human judgments, particularly at segment-level Freitag et al. (2021c).
In this work, apart from BLEU (computed using SacreBLEU nrefs:1|case:mixed|eff:no|tok:13a |smooth:exp|version:2.0.0 Post (2018)) and chrF, we use the following state-of-the-art trainable reference-based metrics for both ranking and performance evaluation of MT systems:
BLEURT (Sellam et al., 2020; Pu et al., 2021), trained to regress on human direct assessments (DA; Graham et al. 2013). We use the largest multilingual version, BLEURT-20, based on the RemBERT model Chung et al. (2021).
COMET (Rei et al., 2020a), based on XLM-R Conneau et al. (2020), trained to regress on quality assessments such as DA using both the reference and the source to assess the quality of a given translation. We use the publicly available model developed for the WMT20 metrics shared task (wmt20-comet-da).
These metrics have shown much better correlation at segment-level than previous lexical metrics in WMT metrics shared tasks Mathur et al. (2020b); Freitag et al. (2021c). Hence, as discussed in §2.2, they are good candidates to be used either indirectly as an optimization objective for learning the tuned reranker’s feature weights, or directly as a utility function in MBR decoding. In the former, the higher the metric correlation with human judgment, the better the translation picked by the tuned reranker. In the latter, we approximate the expected utility in Eq. 4 by letting a candidate generated by the model be a reference translation – a suitable premise if the model is good in expectation.
2 Reference-free Metrics
MT evaluation metrics have also been developed for the case where references are not available – they are called reference-free or quality estimation (QE) metrics. In the last years, considerable improvements have been made to such metrics, with state-of-the-art models having increasing correlations with human annotators (Freitag et al., 2021c; Specia et al., 2021). These improvements enable the use of such models for ranking translation hypotheses in a more reliable way than before.
In this work, we explore four recently proposed reference-free metrics as features for -best reranking, all at the sentence-level:
COMET-QE (Rei et al., 2020b), a reference-free version of COMET (§3.1). It was the winning submission for the QE-as-a-metric subtask of the WMT20 shared task (Mathur et al., 2020b).
TransQuest (Ranasinghe et al., 2020), the winning submission for the sentence-level DA prediction subtask of the WMT20 QE shared task (Specia et al., 2020). Similarly to COMET-QE this metric predicts a DA score.
MBART-QE (Zerva et al., 2021), based on the mBART Liu et al. (2020) model, trained to predict both the mean and the variance of DA scores. It was a top performer in the WMT21 QE shared task (Specia et al., 2021).
OpenKiwi-MQM (Kepler et al., 2019; Rei et al., 2021), based on XLM-R, trained to predict the multidimensional quality metric (MQM; Lommel et al. 2014).MQM annotations are expert-level type of annotations more fine-grained then DA, with individual errors annotated. This reference-free metric was ranked second on the QE-as-a-metric subtask from the WMT 2021 metrics shared task.
Experiments
We study the benefits of quality-aware decoding over MAP-based decoding in two regimes:
A high-resource, unconstrained, setting with large transformer models (6 layers, 16 attention heads, 1024 embedding dimensions, and 8192 hidden dimensions) trained by Ng et al. (2019) for the WMT19 news translation task (Barrault et al., 2019), using English to German () and English to Russian () language pairs. These models were trained on over 20 million parallel and 100 million back-translated sentences, being the winning submissions of that year’s shared task. We consider the non-ensembled version of the model and use newstest19 for validation and newstest20 for testing.
A more constrained scenario with a small transformer model (6 layers, 4 attention heads, 512 embedding dimensions, and 1024 hidden dimensions) trained from scratch in Fairseq (Ott et al., 2019) on the smaller IWSLT17 datasets Cettolo et al. (2012) for English to German () and English to French (), each with a little over 200k training examples. We chose these datasets because they have been extensively used in previous work Bhattacharyya et al. (2021) and smaller model allows us to answer questions about how the training methodology affects ranking performance (see § 4.2.2). Further training details can be found in Appendix A.
We use beam search with a beam size of 5 as our decoding baseline because we found that it resulted in better or similar translations than larger beam sizes. For tuned N-best reranking, we use Travatar’s Neubig (2013) implementation of MERT (Och, 2003) to optimize the weight of each feature, as described in §3.2. Finally, we evaluate each system using the metrics discussed in §3.1, along with BLEU and chrF (Popović, 2015).
2 Results
Overall, given all the metrics, candidate generation, and ranking procedures, we evaluate over 150 systems per dataset. We report subsets of this data separately to answer specific research questions, and defer to Appendix B for additional results.
First, we explore the impact of the candidate generation procedure and the number of candidates.
We generate candidates with beam search, vanilla sampling, and nucleus sampling. For the latter, we use based on early results showing improved performance for all metrics.We picked nucleus sampling over top- sampling because it allows varying support size and has outperformed top- in text generation tasks (Holtzman et al., 2020). For -best reranking, we use up to 200 samples; for MBR decoding, due to the quadratic computational cost, we use up to 100.
Figure 2 shows BLEU and COMET for different candidate generation and ranking methods for the WMT20 and IWSLT17 datasets, with increasing number of candidates. The baseline is represented by the dashed line. To assess the performance ceiling of the rankers, we also report results with an oracle ranker for the reported metrics, picking the candidate that maximizes it. For the fixed -best reranker, we use COMET-QE as a metric, albeit the results for other reference-free metrics are similar. Performance seems to scale well with the number of candidates, particularly for vanilla sampling and for the tuned -best reranker and MBR decoder. (Lee et al., 2021; Müller and Sennrich, 2021). However, all the rankers using vanilla sampling severely under-perform the baseline in most cases (see also §4.2.2). In contrast, the rankers using beam search or nucleus sampling are competitive or outperform the baseline in terms of BLEU, and greatly outperform it in terms of COMET. For the larger models, we see that the performance according to the lexical metrics degrades with more candidates. In this scenario, rankers using nucleus sampling seem to have an edge over the ones that use beam search for COMET.
Based on the findings above, and due to generally better performance of COMET over BLEU for MT evaluation (Kocmi et al., 2021), in following experiments we use nucleus sampling with the large model and beam search with the small model.
2.2 Impact of Label Smoothing
Label smoothing (Szegedy et al., 2016) is a regularization technique that redistributes probability mass from the gold label to the other target labels, typically preventing the model from becoming overconfident (Müller et al., 2019). However, it has been found that label smoothing negatively impacts model fit, compromising the performance of MBR decoding (Eikema and Aziz, 2020, 2021). Thus, we train a small transformer model without label smoothing to verify its impact in the performance of -best reranking and MBR decoding. Figure 3 shows that disabling label smoothing really helps when generating candidates using vanilla sampling. However, the performance degrades for candidates generated using nucleus sampling when we disable label smoothing, hinting that the pruning mechanism of nucleus sampling may help mitigate the negative impact of label smoothing in sampling based approaches. Even without label smoothing, vanilla sampling is not competitive with nucleus sampling or beam search with label smoothing, thus, we do not experiment further with it.
2.3 Impact of Ranking and Metrics
We now investigate the usefulness of the metrics presented in §3 as features and objectives for ranking. For -best reranking, we use all the available candidates (200) while, for MBR, due to the computational cost of using 100 candidates, we report results with 50 candidates only (we found that ranking with tuned -best reranking with and MBR with takes about the same time). We report results in Table 1, and use them to answer some specific research questions.
We consider a fixed -best reranker with a single reference-free metric as a feature (see Table 1, second group). While none of the metrics allows for improving the baseline results in terms of the lexical metrics (BLEU and chrF), rerankers using COMET-QE or MBART-QE outperform the baseline according to BLEURT and COMET, for both the large and small models. Due to the aforementioned better performance of these metrics for translation quality evaluation, we hypothesize that these rankers produce better translations than the baseline. However, since the sharp drop in the lexical metrics is concerning, we will verify this hypothesis in a human study, in §4.2.4.
We consider a tuned -best reranker using as features all the reference-free metrics in §3.2, and optimized using MERT. Table 1 (3rd group) shows results for . For the small model, all the rankers show improved results over the baseline for all the metrics. In particular, optimizing for BLEU leads to the best results in the lexical metrics, while optimizing for BLEURT leads to the best performance in the others. Finally, optimizing for COMET leads to similar performance than optimizing for BLEURT. For the large model, although none of the rerankers is able to outperform the baseline in the lexical metrics, we see similar trends as before for BLEURT and COMET.
Table 1 (4th group) shows the impact of the utility function (BLEU, BLEURT, or COMET). For the small model, using COMET leads to the best performance according to all the metrics except BLEURT (for which the best result is attained when optimizing itself). For the large model, the best result according to a given metric is obtained when using that metric as the utility function.
Looking at Table 1 we see that, for the small model, -best reranking seems to perform better than MBR decoding in all the evaluation metrics, including the one used as the utility function in MBR decoding. The picture is less clear for the large model, with MBR decoding achieving best values for a given fine-tuned metric when using it as the utility; this comes at the cost of worse performance according to the other metrics, hinting at a potential “overfitting” effect. Overall, -best reranking seems to have an edge over MBR decoding. We will further clarify this question with human evaluation in § 4.2.4.
Table 1 shows that, for both the large and the small model, the two-stage ranking approach described in §3 leads to the best performance according to the fine-tuned metrics. In particular, the best result is obtained when the utility function is the same as the evaluation metric. These results suggest that a promising research direction is to seek more sophisticated pruning strategies for MBR decoding.
2.4 Human Evaluation
Our experiments suggest that, overall, quality-aware decoding produces translations with better performance across most metrics than MAP-based decoding. However, for some cases (such as fixed -best reranking and most results with the large model), there is a concerning “metric gap” between lexical-based and fine-tuned metrics. While the latter have shown to correlate better with human judgments, previous work has not attempted to explicitly optimize these metrics, and doing so could lead to ranking systems that learn to exploit “pathologies” in these metrics rather than improving translation quality. To investigate this hypothesis, we perform a human study across all four datasets. We ask annotators to rate, from 1 (no overlap in meaning) to 5 (perfect translation), the translations produced by the 4 ranking systems in §3, as well as the baseline translation and the reference. Further details are in App. C. We choose COMET-QE as the feature for the fixed -best ranker and COMET as the optimization metric and utility function for the tuned -best reranker and MBR decoding, respectively. The reasons for this are two-fold: (1) they are currently the reference-free and reference-based metrics with highest reported correlation with human judgments (Kocmi et al., 2021), (2) we saw the largest “metric gap” for systems based on these metrics, hinting of a potential “overfitting” problem (specially since COMET-QE and COMET are similar models).
Table 2 shows the results for the human evaluation, as well as the automatic metrics. We see that, with the exception of T-RR w/ COMET, when fine-tuned metrics are explicitly optimized for, their correlation with human judgments decreases and they are no longer reliable indicators of system-level ranking. This is notable for the fixed -best reranker with COMET-QE, which outperforms the baseline in COMET for every single scenario, but leads to markedly lower quality translations. However, despite the potential for overfitting these metrics, we find that tuned -best reranking, MBR, and their combination consistently achieve better translation quality than the baseline, specially with the small model. In particular, -best reranking results in better translations than MBR, and their combination is the best system in 2 of 4 LPs.
2.5 Improved Human Evaluation
To further investigate how quality-aware decoding performs when compared to MAP-based decoding, we perform another human study, this time based on expert-level multidimensional quality metrics (MQM) annotations (Lommel et al., 2014). We asked the annotators to identify all errors and independently label them with an error category (accuracy, fluency, and style, each with a specific set of subcategories) and a severity level (minor, major, and critical). In order to obtain the final sentence-level scores, we require a weighting scheme on error severities. We use weights of , , and to minor, major, and critical errors, respectively, independently of the error category. Further details are in App. D. Given the cost of performing a human study like this, we restrict our analysis to the translations generated by the large models trained on WMT20 (EN DE and EN RU).
Table 3 shows the results for the human evaluation using MQM annotations, including both error severity counts and final MQM scores. As hinted in §4.2.4, despite the remarkable performance of the F-RR with COMET-QE in terms of COMET (see Table 2), the quality of the translations decreases when compared to the baseline, suggesting the possibility of metric overfitting when evaluating systems using a single automatic metric that was directly optimized for (or a similar one). However, for both language pairs, the T-RR with COMET and the two stage approach (T-RR + MBR with COMET) achieve the highest MQM score. In addition, these systems present the smallest number of errors when combining both major and critical errors.
Although the performance of all systems is comparable for ENDE, both the T-RR and the T-RR+MBR decoding markedly reduce the number of grammatical register errors related to using pronouns and verb forms that are not compliant with the register required for that translation, at the cost of increasing the number of lexical selection errors (see Figure 4). For ENRU, however, the number of lexical selection errors produced when using the T-RR or the T-RR+MBR decoding is approximately a half of the ones produced by the baseline (see Figure 5). In this case, this comes at apparently almost no cost in other error types, leading to significantly better results, as shown in Table 3.
Related Work
Inspired by the work of Shen et al. (2004) on discriminative reranking for SMT, Lee et al. (2021) trained a large transformer model using a reranking objective to optimize BLEU. Our work differs in which our rerankers are much simpler and therefore can be tuned on a validation set; and we use more powerful quality metrics instead of BLEU. Similarly, Bhattacharyya et al. (2021) learned an energy-based reranker to assign lower energy to the samples with higher BLEU scores. While the energy model plays a similar role to a QE system, our work differs in two ways: we use an existing, pretrained QE model instead of training a dedicated reranker, making our approach applicable to any MT system without further training; and the QE model is trained to predict human assessments, rather than BLEU scores. Leblond et al. (2021) compare a reinforcement learning approach to reranking approaches (but not MBR decoding, as we do). They investigate the use of reference-based metrics and, for the reward function, a reference-free metric based on a modified BERTScore Zhang et al. (2020). This new multilingual BERTScore is not fine-tuned on human judgments as COMET and BLEURT and it is unclear what its level of agreement with human judgments is. Another line of work is generative reranking, where the reranker is not trained to optimize a metric, but rather as a generative noisy-channel model (Yu et al., 2017; Yee et al., 2019; Ng et al., 2019).
MBR decoding (Kumar and Byrne, 2002, 2004) has recently been revived for NMT using candidates generated with beam search (Stahlberg et al., 2017; Shu and Nakayama, 2017) and sampling (Eikema and Aziz, 2020; Müller and Sennrich, 2021). Eikema and Aziz (2021) also explore a two-stage approach for MBR decoding. Additionally, there is concurrent work by Freitag et al. (2021b) on using neural metrics as utility functions during MBR decoding: however they limit their scope to MBR with reference-based metrics, while we perform a more extensive evaluation over ranking methods and metrics. Amrhein and Sennrich (2022) also concurrently explored using MBR decoding with neural metrics, but with the purposes of identifying weaknesses in the metric (in their case COMET), similarly to the metric overfitting problem we discussed in §4.2.4. A comparison with -best re-ranking was missing in these works, a gap our paper fills. A related line of work is minimum risk training (MRT; Smith and Eisner 2006; Shen et al. 2016), which trains models to minimize risk, allowing arbitrary non-differentiable loss functions (Edunov et al., 2018; Wieting et al., 2019) and avoiding exposure bias (Wang and Sennrich, 2020; Kiegeland and Kreutzer, 2021). However, MRT is considerably more expensive and difficult to train and the gains are often small. Incorporating our quality metrics in MRT is an exciting research direction.
Conclusions and Future Work
We leverage recent advances in MT quality estimation and evaluation and propose quality-aware decoding for NMT. We explore different candidate generation and ranking methods, with a comprehensive empirical analysis across four datasets and two model classes. We show that, compared to MAP-based decoding, quality-aware decoding leads to better translations, according to powerful automatic evaluation metrics and human judgments.
There are several directions for future work. Our ranking strategies increase accuracy but are substantially more expensive, particularly when used with costly metrics such as BLEURT and COMET. While reranking-based pruning before MBR decoding was found helpful, additional strategies such as caching encoder representations Amrhein and Sennrich (2022) and distillation (Pu et al., 2021) are promising directions.
Acknowledgments
We would like to thank Ben Peters, Wilker Aziz, and the anonymous reviewers for useful feedback. This work was supported by the P2020 program MAIA (LISBOA-01-0247- FEDER-045909), the European Research Council (ERC StG DeepSPIN 758969), the European Union’s Horizon 2020 research and innovation program (QUARTZ grant agreement 951847), and by the Fundação para a Ciência e Tecnologia through UIDB/50008/2020.
References
Appendix A Training Details
For the experiments using IWSLT17, we train a small transformer model (6 layers, 4 attention heads, 512 embedding dimensions, and 1024 hidden dimensions) from scratch, using Fairseq (Ott et al., 2019). We tokenize the data using SentencePiece (Kudo and Richardson, 2018), with a joint vocabulary with 20000 units. We train using the Adam optimizer (Kingma and Ba, 2015) with and and use an inverse square root learning rate scheduler, with an initial learning rate of and with a linear warm-up in the first steps. For models trained with label smoothing, we use the default value of .
Appendix B Additional Results
For completeness, we include in Table 4 results to evaluate the impact of the metrics presented in §3 as features and objectives for ranking using the other language pairs: (large model) and (small model).
Appendix C Human Study
In order to perform human evaluation, we recruited professional translators who were native speakers of the target language on the freelancing site Upwork.https://upwork.com. Freelancers were paid a market rate of 18-20 US dollars per hour, and finished approximately 50 sentences in one hour. 300 sentences were evaluated for each language pair, sampled randomly from the test sets after a restriction that sentences were no longer than 30 words. All translation hypotheses for a single source sentence were first deduplicated, and then shown to the translator side-by-side in randomized order to avoid any ordering biases.
Sentences were evaluated according to a 1-5 rubric slightly adapted from that of Wieting et al. (2019):
There is no overlap in the meaning of the source sentence whatsoever.
Some content is similar but the most important information in the sentence is different.
The key information in the sentence is the same but the details differ.
Meaning is essentially equal but some expressions are unnatural.
Meaning is essentially equal and the sentence is natural.
Appendix D MQM Framework
Human evaluations were performed by Unbabel’s PRO Community, made of professional translators and linguists with relevant experience in linguistic annotations and translation errors annotations. In order to properly assess translations quality, annotators must be native speakers of the target language and with a proven high proficiency of the source language, so that they can properly capture errors and their nuances. The systems’ outputs were evaluated by using the annotation framework adopted internally at Unbabel, which is an adaptation of the MQM Framework (Lommel et al., 2014).
We asked the annotators to identify all errors and independently label them with an error category and a severity level. We consider three categories (each of them containing a set of different subcategories) that may affect the quality of the translations:
Accuracy, if the target text does not accurately reflect the source text (e.g., changes in the meaning, addition/omission of information, untranslated text, MT hallucinations);
Fluency, if there are issues that affect the reading and the comprehension of the text (e.g., grammar and spelling errors);
Style, if the text has stylistic problems (e.g., gramatical and lexical register).
Additionally, each error is labeled according to three severity levels (minor, major, and critical), depending on the way they affect the accuracy, the fluency, and the style of the translation. The final sentence-level score is obtained using a weighting scheme where minor, major, and critical errors are weighted as , , and , respectively.
Figures 4 and 5 show the counts of errors breakdown by typology and severity level for ENDE and ENRU, respectively.