Machine Translation into Low-resource Language Varieties
Sachin Kumar, Antonios Anastasopoulos, Shuly Wintner, Yulia Tsvetkov
Introduction
Despite tremendous progress in machine translation (Bahdanau et al., 2015; Vaswani et al., 2017) and language generation in general, current state-of-the-art systems often work under the assumption that a language is homogeneously spoken and understood by its speakers: they generate a “standard” form of the target language, typically based on the availability of parallel data. But language use varies with regions, socio-economic backgrounds, ethnicity, and fluency, and many widely spoken languages consist of dozens of varieties or dialects, with differing lexical, morphological, and syntactic patterns for which no translation data are typically available. As a result, models trained to translate from a source language (src) to a standard language variety (std) lead to a sub-par experience for speakers of other varieties.
Motivated by these issues, we focus on the task of adapting a trained srcstd translation model to generate text in a different target variety (tgt), having access only to limited monolingual corpora in tgt and no src–tgt parallel data. tgt may be a dialect of, a language variety of, or a typologically-related language to std.
We present an effective transfer-learning framework for translation into low resource language varieties. Our method reuses srcstd MT models and finetunes them on synthesized (pseudo-parallel) src–tgt texts. This allows for rapid adaptation of MT models to new varieties without having to train everything from scratch. Using word-embedding adaptation techniques, we show that MT models which predict continuous word vectors (Kumar and Tsvetkov, 2019) rather than softmax probabilities lead to superior performance since they allow additional knowledge to be injected into the models through transfer between word embeddings of high-resource (std) and low-resource (tgt) monolingual corpora.
We evaluate our framework on three translation tasks: English to Ukrainian and Belarusian, assuming parallel data are only available for EnglishRussian; English to Nynorsk, with only English to Norwegian Bokmål parallel data; and English to four Arabic dialects, with only EnglishModern Standard Arabic (MSA) parallel data. Our approach outperforms competitive baselines based on unsupervised MT, and methods based on finetuning softmax-based models.
A Transfer-learning Architecture
We first formalize the task setup. We are given a parallel srcstd corpus, which allows us to train a translation model that takes an input sentence in src and generates its translation in the standard veriety std, . Here, are the learnable parameters of the model. We are also given monolingual corpora in both the standard std and target variety tgt. Our goal now is to modify to generate translations in the target variety tgt. At training time, we assume no src–tgt or std–tgt parallel data are available.
Our solution (Figure 1) is based on a transformer-based encoder-decoder architecture (Vaswani et al., 2017) which we modify to predict word vectors. Following Kumar and Tsvetkov (2019), instead of treating each token in the vocabulary as a discrete unit, we represent it using a unit-normalized -dimensional pre-trained vector. These vectors are learned from a std monolingual corpus using fasttext (Bojanowski et al., 2017). A word’s representation is computed as the average of the vectors of its character -grams, allowing surface-level linguistic information to be shared among words. At each step in the decoder, we feed this pretrained vector at the input and instead of predicting a probability distribution over the vocabulary using a softmax layer, we predict a -dimensional continuous-valued vector. We train this model by minimizing the von Mises-Fisher (vMF) loss—a probabilistic variant of cosine distance—between the predicted vector and the pre-trained vector. The pre-trained vectors (at both input and output of the decoder) are not trained with the model. To decode from this model, at each step, the output word is generated by finding the closest neighbor (in terms of cosine similarity) of the predicted output vector in the pre-trained embedding table.
We train in this fashion using src–std parallel data. As shown below, training a softmax-based srcstd model to later finetune with tgt suffers from vocabulary mismatch between std and tgt and thus is detrimental to downstream performance. By replacing the decoder input and output with pretrained vectors, we separate the vocabulary from the MT model, making adaptation easier.
Now, to finetune this model to generate tgt, we need tgt embeddings. Since the tgt monolingual corpus is small, training fasttext vectors on this corpus from scratch will lead (as we show) to low-quality embeddings. Leveraging the relatedness of std and tgt and their vocabulary overlap, we use std embeddings to transfer knowledge to tgt embeddings: for each character -gram in the tgt corpus, we initialize its embedding with the corresponding std embedding, if available. We then continue training fasttext on the tgt monolingual corpus (Chaudhary et al., 2018). Last, we use a supervised embedding alignment method (Lample et al., 2018a) to project the learned tgt embeddings in the same space as std. std and tgt are expected to have a large lexical overlap, so we use identical tokens in both varieties as supervision for this alignment. The obtained embeddings, due to transfer learning from std, inject additional knowledge in the model.
Finally, to obtain a srctgt model, we finetune on psuedo-parallel src–tgt data. Using a stdsrc MT model (a back-translation model trained using large std–src parallel data with standard settings) we (back)-translate tgt data to src. Naturally, these synthetic parallel data will be noisy despite the similarity between std and tgt, but we show that they improve the overall performance. We discuss the implications of this noise in §4.
Experimental Setup
We experiment with two setups. In the first (synthetic) setup, we use English (en) as src, Russian (ru) as std, and Ukrainian (uk) and Belarusian (be) as tgts. We sample 10M en-ru sentences from the WMT’19 shared task (Ma et al., 2019), and 80M ru sentences from the CoNLL’17 shared task to train embeddings. To simulate low-resource scenarios, we sample 10K, 100K and 1M uk sentences from the CoNLL’17 shared task and be sentences from the OSCAR corpus (Ortiz Suárez et al., 2020). We use TED dev/test sets for both languages pairs (Cettolo et al., 2012).
The second (real world) setup has two language sets: the first one defines English as src, with Modern Standard Arabic (msa) as std and four Arabic varieties spoken in Doha, Beirut, Rabat and Tunis as tgts. We sample 10M en-msa sentences from the UNPC corpus (Ziemski et al., 2016), and 80M msa sentences from the CoNLL’17 shared task. For Arabic varieties, we use the MADAR corpus (Bouamor et al., 2018) which consists of 12K 6-way parallel sentences between English, MSA and the 4 considered varieties. We ignore the English sentences, sample dev/test sets of 1K sentences each, and consider 10K monolingual sentences for each tgt variety. The second set also has English as src with Norwegian Bokmål (no) as std and its written variety Nynorsk (nn) as tgt. We use 630K en-no sentences from WikiMatrix (Schwenk et al., 2021), and 26M no sentences from ParaCrawl (Esplà et al., 2019) combined with the WikiMatrix no sentences to train embeddings. We use 310K nn sentences from WikiMatrix, and TED dev/test sets for both varieties (Reimers and Gurevych, 2020).
Preprocessing
We preprocess raw text using Byte Pair Encoding (BPE, Sennrich et al., 2016) with 24K merge operations on each src–std corpus trained separately on src and std. We use the same BPE model to tokenize the monolingual std data and learn fasttext embeddings (we consider character -grams of length 3 to 6).We slightly modify fasttext to not consider BPE token markers “@@” in the character -grams. Splitting the tgt words with the same std BPE model will result in heavy segmentation, especially when tgt contains characters not present in std.For example, both ru and uk alphabets consist of 33 letters; ru has the letters Ёё, ъ, ы and Ээ, which are not used in uk. Instead, uk has Ґґ, Єє, Ii and Її. To counter this, we train a joint BPE model with 24K operations on the concatenation of std and tgt corpora to tokenize tgt corpus following Chronopoulou et al. (2020). This technique increases the number of shared tokens between std and tgt, thus enabling better cross-variety transfer while learning embeddings and while finetuning. We follow Chaudhary et al. (2018) to train embeddings on the generated tgt vocabulary where we initialize the character -gram representations for tgt words with std’s fasttext model wherever available and finetune them on the tgt corpus.
Implementation and Evaluation
We modify the standard OpenNMT-py seq2seq models of PyTorch (Klein et al., 2017) to train our model with vMF loss (Kumar and Tsvetkov, 2019). Additional hyperparameter details are outlined in Appendix B. We evaluate our methods using BLEU score (Papineni et al., 2002) based on the SacreBLEU implementation (Post, 2018).While we recognize the limitations of BLEU (Mathur et al., 2020), more sophisticated embedding-based metrics for MT evaluation (Zhang et al., 2020; Sellam et al., 2020) are unfortunately not available for low-resource language varieties. For the Arabic varieties, we also report a macro-average. In addition, to measure the expected impact on actual systems’ users, we follow Faisal et al. (2021) in computing a population-weighted macro-average () based on language community populations provided by Ethnologue Eberhard et al. (2019).
1 Experiments
Our proposed framework, LangVarMT, consists of three main components: (1) A supervised srcstd model is trained to predict continuous std word embeddings rather than discrete softmax probabilities. (2) Output std embeddings are replaced with tgt embeddings. The tgt embeddings are trained by finetuning std embeddings on monolingual tgt data and aligning the two embedding spaces. (3) The resulting model is finetuned with pseudo-parallel srctgt data.
We compare LangVarMT with the following competitive baselines. Sup(srcstd): train a standard (softmax-based) supervised srcstd model, and consider the output of this model as tgt under the assumption that std and tgt may be very similar. Unsup(srctgt): train an unsupervised MT model (Lample et al., 2018a) in which the encoder and decoder are initialized with cross-lingual masked language models (MLM, Conneau and Lample, 2019). These MLMs are pre-trained on src monolingual data, and then finetuned on tgt monolingual data with an expanded vocabulary as described above. This baseline is taken from Chronopoulou et al. (2020), where it showed state-of-the-art performance for low-monolingual-resource scenarios. Pivot: train a Unsup(stdtgt) model as described above using std and tgt monolingual corpora. During inference, translate the src sentence to std with the Sup(srcstd) model and then to tgt with the Unsup(stdtgt) model. We also perform several ablation experiments, showing that every component of LangVarMT is necessary for good downstream performance. Specifically, we report results with LangVarMT but using a standard softmax layer (softmax) to predict tokens instead of continuous vectors.Additional ablation results are listed in Appendix C.
Results and Analysis
Table 1 compares the performance of LangVarMT with the baselines for Ukrainian, Belarusian, Nynorsk, and the four Arabic varieties. For reference, note that the enru, enmsa, and enno models are relatively strong, yielding BLEU scores of , , and , respectively.
Considering std and tgt as the same language is sub-optimal, as is evident from the poor performance of the non-adapted Sup(srcstd) model. Clearly, special attention ought to be paid to language varieties. Direct unsupervised translation from src to tgt performs poorly as well, confirming previously reported results of the ineffectiveness of such methods on unrelated languages Guzmán et al. (2019).
Translating src to tgt by pivoting through std achieves much better performance owing to strong Unsup(stdtgt) models that leverage the similarities between std and tgt. However, when resources are scarse (e.g., with 10K monolingual sentences as opposed to 1M), this performance gain considerably diminishes. We attribute this drop to overfitting during the pre-training phase on the small tgt monolingual data. Ablation results (Appendix C) also show that in such low-resource settings the learned embeddings are of low quality.
Finally, LangVarMT consistently outperforms all baselines. Using 1M uk sentences, it achieves similar performance (for enuk) to the softmax ablation of our method, Softmax, and small gains over unsupervised methods. However, in lower resource settings our approach is clearly better than the strongest baselines by over 4 BLEU points for uk (10K) and 3.9 points for be (100K).
To identify potential sources of error in our proposed method, we lemmatize the generated translations and test sets and evaluate BLEU (Qi et al., 2020). Across all data sizes, both uk and be achieve a substantial increase in BLEU (up to +6 BLEU; see Appendix D for details) compared to that obtained on raw text, indicating morphological errors in the translations. In future work, we will investigate whether we can alleviate this issue by considering tgt embeddings based on morphological features of tokens (Chaudhary et al., 2018).
Real-world Setup
The effectiveness of LangVarMT is pronounced in this setup with a dramatic improvement of more than 18 BLEU points over unsupervised baselines when translating into Doha Arabic. We hypothesize that during the pretraining phase of unsupervised methods, the extreme difference between the size of the msa monolingual corpus (10M) and the varieties’ corpora (10K) leads to overfitting. Additionally, compared to the synthetic setup, the Arabic varieties we consider are quite close to msa, allowing for easy and effective adaptation of both word embeddings and enmsa models. LangVarMT also improves in all other Arabic varieties, although naturally some varieties remain challenging. For example, the Rabat and particularly the Tunis varieties are more likely to include French loanwords Bouamor et al. (2018) which are not adequately handled as they are not part of our vocabulary. In future work, we will investigate whether we can alleviate this issue by potentially including French corpora (transliterated into Arabic) to our tgt language corpora. On average, our approach improves by 2.3 BLEU points over the softmax-based baseline (cf. 7.7 and 10.0 in Table 2 under ) across the four Arabic dialects. For a population-weighted average (), we associate the Doha variety with Gulf Arabic (ISO code: afb), the Beirut one with North Levantine Arabic (apc), Rabat with Moroccan (ary), and the Tunis variety with Tunisian Arabic (aeb). As before, LangVarMT outperforms the baselines. The absolute BLEU scores in this highly challenging setup are admittedly low, but as we discuss in Appendix D, the translations generated by LangVarMT are often fluent and input preserving, especially compared to the baselines.
Finally, due to high similarity between no and nn, the Sup(enno) model also performs well on nn with 11.3 BLEU, but our method yields further gains of over 4 points over the baselines.
Discussion
Fairness The goal of this work is to develop more equitable technologies, usable by speakers of diverse language varieties. Here, we evaluate the systems along the principles of fairness. We evaluate the fairness of our Arabic multi-dialect system’s utility proportionally to the populations speaking those dialects. In particular, we seek to measure how much average benefit will the people of different dialects receive if their respective translation performance is improved. A simple proxy for fairness is the standard deviation (or, even simpler, a performance) of the BLEU scores across dialects (A higher value implies more unfairness across the dialects) Beyond that, we measure a system’s unfairness with respect to the different dialect subgroups, using the adaptation of generalized entropy index Speicher et al. (2018), which considers equities within and between subgroups in evaluating the overall unfairness of an algorithm on a population Faisal et al. (2021) (See Appendix F for details and additional discussion). Table 2 shows that our proposed method is fairer across all dialects, compared to baselines where only msa translation produces comprehensible outputs.
Negative Results Our proposed method relies on two components: (1) quality of tgt word embeddings which is dependent on std and tgt shared (subword) vocabulary, and (2) the psuedo-parallel src–tgt obtained by back-translating tgt data through a stdsrc model. If std and tgt are not sufficiently closely related, the quality of both of these components can degrade, leading to a drop in the performance of our proposed method. We present results of two additional experiments to elucidate this phenomenon in Appendix E.
Related Work We provide an extensive discussion of related work in Appendix A.
Conclusion
We presented a transfer-learning framework for rapid and effective adaptation of MT models to different varieties of the target language without access to any source-to-variety parallel data. We demonstrated significant gains in BLEU scores across several language pairs, especially in highly resource-scarce scenarios. The improvements are mainly due to the benefits of continuous-output models over softmax-based generation. Our analysis highlights the importance of addressing morphological differences between language varieties, which will be in the focus of our future work.
Acknowledgements
This research was supported by Grants No. 2017699 and 2019785 from the United States-Israel Binational Science Foundation (BSF), by the National Science Foundation (NSF) under Grants No. 2040926 and 2007960, and by a Google faculty research award. We thank Safaa Shehadi for evaluating our model outputs, Xinyi Wang and Aditi Choudhary for helpful discussions, and the anonymous reviewers for much appreciated feedback.
References
Appendix A Related Work
Early work addressing translation involving language varieties includes rule-based transformations (Altintas and Cicekli, 2002; Marujo et al., 2011; Tan et al., 2012) which rely on language specific information and expert knowledge which can be expensive and difficult to scale. Recent work to address this issue only focuses on cases where parallel data do exist. They include a combination of word-level and character-level MT (Vilar et al., 2007; Tiedemann, 2009; Nakov and Tiedemann, 2012) between related languages or training multilingual models to translate to/from English to different varieties of a language (e.g., Lakew et al. (2018) work on Brazilian–European Portuguese and European–Canadian French). Such parallel data, however, are typically unavailable for most language varieties.
Unsupervised translation models, which require only monolingual data, can address this limitation (Artetxe et al., 2018; Lample et al., 2018a; Garcia et al., 2020, 2021). However, when even monolingual corpora are limited, unsupervised models are challenging to train and are quite ineffective for translating between unrelated languages (Marchisio et al., 2020). Considering varieties of a language as writing styles, unsupervised style transfer (Yang et al., 2018; He et al., 2020) or deciphering methods (Pourdamghani and Knight, 2017) to translate between different varieties have also been been explored but have not been shown to perform well, often only reporting BLEU-1 scores since they obtain BLEU-4 scores which are closer to 0. Additionally, all of these approaches require simultaneous access to data in all varieties during training and must be trained from scratch when a new variety is added. In contrast, our presented method allows for easy adaptation of srcstd models to any new variety as it arrives.
Considering a new target variety as a new domain of std, unsupervised domain adaptation methods can be employed, such as finetuning srcstd models using pseudo-parallel corpora generated from monolingual corpora in target varieties (Hu et al., 2019; Currey et al., 2017). Our proposed method is most related to this approach; but while these methods have the potential to adapt the decoder language model, for effective transfer, std and tgt must have a shared vocabulary which is not true for most language varieties due to lexical, morphological, and at times orthographic differences. In contrast, our proposed method makes use of cross-variety word embeddings. While our examples only involve same-script varieties, augmenting our approach to work across scripts through a transliteration component is straightforward.
Appendix B Implementation Details
We modify the standard OpenNMT-py seq2seq models of PyTorch (Klein et al., 2017) to train our model with vMF loss (Kumar and Tsvetkov, 2019). We use the transformer-base model (Vaswani et al., 2017), with 6 layers in both encoder and decoder and with 8 attention heads, as our underlying architecture. We modify this model to predict pretrained fasttext vectors. We also initialize the decoder input embedding table with the pretrained vectors and do not update them during model training. All models are optimized using Rectified Adam (Liu et al., 2020) with a batch size of 4K tokens and dropout of . We train srcstd models for 350K steps with an initial learning rate of with linear decay. For finetuning, we reduce the learning rate to and train for up to K steps. We use early stopping in all models based on validation loss computed every 2K steps. We decode all the softmax-based models with a beam size of 5 and all the vMF-based models greedily.
We evaluate our methods using BLEU score (Papineni et al., 2002) based on the SacreBLEU implementation (Post, 2018). While we recognize the limitations of BLEU (Mathur et al., 2020), more sophisticated embedding-based metrics for MT evaluation (Zhang et al., 2020; Sellam et al., 2020) are simply not available for language varieties.
Appendix C Additional English-Ukrainian Experiments
On our resource-richest setup of enuk translation using 1M uk sentences and ru as std, we compare our method with the following additional baselines. Table 3 presents these results.
Lample-Unsup(srctgt): This is another unsupervised model, based on Lample et al. (2018a) which initializes the input and output embedding tables of both encoder and decoder using cross-lingual word embeddings trained on src and tgt monolingual corpora. The model is trained in a similar manner to Chronopoulou et al. (2020) (Unsup(srctgt)) with iterative backtranslation and autoencoding.
Pivot:Lample(stdtgt): This baseline is similar to the Pivot baseline, where we replace the unsupervised model with that of Lample et al. (2018a).
Pivot:DictReplace(stdtgt): Here we first translate src to std using Sup(srcstd), and then modify the std output to get a tgt sentence as follows: We create a std–tgt dictionary using the embedding map suggested by Lample et al. (2018b). This dictionary is created on words tokenized with Moses tokenizer (Hoang and Koehn, 2008) rather than BPE tokens. We replace each token in the generated std sentence which is not in the tgt vocabulary using the dictionary (if available). We consider this baseline to measure lexical vs. syntactic/phrase level differences between Russian and Ukrainian.
In addition to baseline comparison, we report the following ablation experiments.
(1) To measure transfer from std to tgt embeddings, we finetune the Sup(srcstd) model using tgt embeddings trained from scratch (as opposed to initialized with std embeddings).
(2) To measure the impact of initialization during model finetuning, we compare with a randomly initialized model trained in a supervised fashion on the psuedo-parallel src–tgt data.
On the unsupervised models based on Lample et al. (2018a), we observe a similar trend as that of Chronopoulou et al. (2020), where the Lample-Unsup(srctgt) model performing poorly (0.4) with substantial gains when pivoting through Russian (9.0 BLEU).
Pivot:DictReplace(stdtgt) gains some improvement over considering the output of Sup(srcstd) as tgt, probably due to syntactic similarities between Russian and Ukrainian. This result can potentially be further improved with a human-curated ru–uk dictionary, but such resources are typically not available for the low-resource settings we consider in this paper.
Ablations
As shown in Table 3, training the srctgt model on a randomly initialized model (LangVar-random) results in a performance drop, confirming that transfer learning from a srcstd model is beneficial. Similarly, using tgt embeddings trained from scratch (LangVarMT w/ poor embeddings) results in a drastic performance drop, providing evidence for essential transfer from std embeddings.
Appendix D Analysis
To better understand the performance of our models, we perform additional analyses.
For uk and be, we lemmatize each word in the test sets and the translations and evaluate BLEU scores. The results, depicted in Table 4, very likely indicate that our framework often generates correct lemmas, but may fail on the correct inflectional form of the target words. This highlights the importance of considering morphological differences between language varieties. The high BLEU scores also demonstrate that the resulting translations are quite likely understandable, albeit not always grammatical.
Translation of Rare Words
On the outputs of the enuk model, trained with 100K uk sentences, we compute the translation accuracy of words based on their frequency in the tgt monolingual corpus for LangVarMT, our best baseline Sup(srcstd)+Unsup(srctgt) and the best performing ablation Softmax. These results, shown in Table 5, reveal that LangVarMT is more accurate at translating rare words (with frequency less than 10) compared to the baselines.
Examples
We provide some examples of en-uk and en-Beirut Arabic translations generated by the three models in Tables 6 and 7. As evaluated by native speakers of the Beirut Arabic, we find that despite a BLEU score of only 8, in a majority of cases our baseline model is able to generate fluent translations of the input, preserving most of the content, whereas the baseline model ignores many of the content words. We also observe that in some cases, despite predicting in the right semantic space of the pretrained embeddings, it fails to predict the right token, resulting in surface form errors (e.g., predicting adjectival forms of verbs). This phenomenon is known and studied in more detail in Kumar and Tsvetkov (2019).
Appendix E Negative Results
We present results for the following experiments: (a) adapting an English to Thai (enth) model to Lao (lo). We use a parallel corpus of around 10M sentences for training the supervised enth model from the CCAligned corpus (El-Kishky et al., 2020), around 140K lo monolingual sentences from the OSCAR corpus (Ortiz Suárez et al., 2020) and TED2020 dev/tests for both th and loAlthough Thai and Lao scripts look very similar, they use different Unicode symbols which are one-to-one mappable to each other: https://en.wikipedia.org/wiki/Lao_(Unicode_block) (Reimers and Gurevych, 2020). (b) adapting an English to Amharic Model (enam) to Tigrinya (ti). We use training, development and test sets from the JW300 corpus (Agić and Vulić, 2019) containing 500K en–am parallel corpus and 100K Tigrinya monolingual sentences.
As summarized in Table 8, our method fails to perform well on these sets of languages. Although Thai and Lao are very closely related languages, we attribute this result to little subword overlap in their respective vocabularies which degrade the quality of the embeddings. This is because Lao’s writing system is developed phonetically whereas Thai writing contains many silent characters. Considering shared phonetic information while learning the embeddings can alleviate this issue and is an avenue for future work. On the other hand, Amharic and Tigrinya, while sharing a decent amount of vocabulary, use different constructs and function words (Kidane et al., 2021) leading to a very noisy psuedo-parallel corpus.
Appendix F Measuring Unfairness
When evaluating multilingual and multi-dialect systems, it is crucial that the evaluation takes into account principles of fairness, as outlined in economics and social choice theory Choudhury and Deshpande (2021). We follow the least difference principle proposed by Rawls (1999), whose egalitarian approach proposes to narrow the gap between unequal accuracies.
A simple proxy for unfairness is the standard deviation (or, even simpler, a performance) of the scores across languages. Beyond that, we measure a system’s unfairness with respect to the different subgroups using the adaptation of generalized entropy index described by Speicher et al. (2018), which considers equities within and between subgroups in evaluating the overall unfairness of an algorithm on a population. The generalized entropy index for a population of individuals receiving benefits with mean benefit is
Using following Speicher et al. (2018), the generalized entropy index corresponds to half the squared coefficient of variation.The coefficient of variation is simply the ratio of the standard deviation to the mean of a distribution.
If the underlying population can be split into disjoint subgroups across some attribute (e.g. gender, age, or language variety) we can decompose the total unfairness into individual and group-level unfairness. Each subgroup will correspond to individuals with corresponding benefit vector and mean benefit . Then, total generalized entropy can be re-written as:
The first term corresponds to the weighted unfairness score that is observed within each subgroup, while the second term corresponds to the unfairness score across different subgroups.
In this measure of unfairness, we define the benefit as being directly proportional to the system’s accuracy. For a Machine Translation system, each user receives an average benefit equal to the BLEU score the MT system achieves on the user’s dialect. Conceptually, if the system produces a perfect translation (BLEU=1) then the user will receive the highest benefit of 1. If the system fails to produce a meaningful translation (BLEU) then the user receives no benefit () from the interaction with the system.