Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task

Thibault Sellam, Amy Pu, Hyung Won Chung, Sebastian Gehrmann, Qijun Tan, Markus Freitag, Dipanjan Das, Ankur P. Parikh

Introduction

The recent progress in machine translation models has led researchers to question the use of n-gram overlap metrics such as BLEU, which focus solely on surface-level aspects of the generated text, and thus may correlate poorly with human evaluation Papineni et al. 2002; Lin 2004; Ma et al. 2019; Mathur et al. 2020; Belz and Reiter 2006; Callison-Burch et al. 2006. This has led to a surge of interest for more flexible metrics that use machine learning to capture semantic-level information Celikyilmaz et al. 2020. Popular examples of such metrics include YiSi-1 Lo 2019, ESIM Mathur et al. 2019, BERTscore Zhang et al. 2020, the Sentence Mover’s Similarity Zhao et al. 2019; Clark et al. 2019, and Bleurt Sellam et al. 2020. These metrics utilize contextual embeddings from large models such as BERT Devlin et al. 2019 which have been shown to capture linguistic information beyond surface-level aspects Tenney et al. 2019.

The WMT Metrics 2020 Shared Task is the reference benchmark for evaluating these metrics in the context of machine translation. It tests the evaluation of systems that are to-English (X→En\texttt{X}\rightarrow\texttt{En}) and to other languages (X→Y\texttt{X}\rightarrow\texttt{Y}), which requires a multilingual approach. An additional challenge for learned metrics is that human ratings are not available for all language pairs, and therefore, the models must use unlabeled data and perform zero-shot generalization.

We describe several learned metrics based on Bleurt (Sellam et al. 2020), originally developed for English data. We first extend Bleurt to the multilingual setup, and show that our approach achieves competitive results on the WMT Metrics 2019 Shared Task. We use the following languages for fine-tuning and/or testing: Chinese, Czech, German, English, Estonian, Finnish, French, Gujarati, Kazakh, Lithuanian, Russian, and Turkish. In addition, we also pre-train on Inuktitut, Japanese, Khmer, Pastho, Polish, Romanian, and Tamil. We also present several simple BERT-based baselines, which we submit for analysis. Finally, we focus on English to German and enhance Bleurt’s performance by combining its predictions with those of YiSi Lo 2019 as well as by using alternative references.

Background and Notations

Reference-based NLG evaluation seeks to assign a score to a triplet of sentences (input, reference, candidate), where input is a sentence in the source language, reference is a reference translation kept secret at inference time, and candidate is a translation produced by an MT system.

Bleurt

where W{\bm{W}} and b{\bm{b}} are the weight matrix and bias vector respectively.

Extending BLEURT Beyond English

An approach to extend BLEURT would be to use mBERT, the public version of BERT pre-trained on 104 languages, and “mid-train” with non-English signals as described above. Yet, the evidence we gathered from early experiments were inconclusive. On the other hand, we did observe that models trained on several languages were often more accurate than monolingual models, possibly due to the larger amount of fine-tuning data. Thus, we opted for a simpler approach where we start with a multilingual BERT model and fine-tune it on all the human ratings data available for all languages (X→Y\texttt{X}\rightarrow\texttt{Y} and X→En\texttt{X}\rightarrow\texttt{En}). In most cases, we found that such models could perform zero-shot evaluation: if a language Y does not have human ratings data, the metric can still perform evaluation in this target language as long as the base multilingual BERT model contains unlabeled data for Y, as observed in the past literature Karthikeyan et al. 2019; Pires et al. 2019.

We experiment with two pre-trained multilingual models: mBERT and mBERT-WMT, a custom multilingual variant of BERT. The mBERT-WMT model is larger that mBERT (24 Transformer layers instead of 12), and it was pre-trained on 19 languages of the WMT Metrics shared task 2015 to 2020.

We trained mBERT-WMT model with an MLM loss Devlin et al. 2019, using a combination of public datasets: Wikipedia, the WMT 2019 News Crawl (Barrault et al.), the C4 variant of Common Crawl (Raffel et al. 2020), OPUS (Tiedemann 2012), Nunavut Hansard (Joanis et al. 2020), WikiTitles https://linguatools.org/tools/corpora/wikipedia-parallel-titles-corpora/, and ParaCrawl (Esplà-Gomis et al. 2019). We trained a new WordPiece vocabulary (Schuster and Nakajima 2012; Wu et al. 2016), since the original vocabulary of mBERT does not support the alphabets of Pashto, Khmer and Inuktitut. The model was trained for 1 million steps with the LAMB optimizer (You et al. 2020), using the learning rate 0.0018 and batch size 4096 on 64 TPU v3 chips.

2 Experimental Setup

At the time of writing, no human ratings data is available for WMT Metrics 2020. Therefore, we use the human ratings from WMT Metrics years 2015 to 2019 for both training and evaluation. We do so in two stages. In the first stage, we use 2015 to 2018 for training (216,541 sentence pairs in 8 languages), setting 10% aside for early stopping. We use 2019 as a development set, to choose hyper-parameters and to support high-level modeling decisions. In the second stage, we use 2015 to 2019, that is, all the data available, for training and uniformly sample 10% of the data for early stopping and hyper-parameter tuning. This adds 289,895 sentence pairs and 4 additional languages to our training set, approximately doubling the size of the training data. We report our results on the first setup, but submit our predictions to the shared task using the second setup.

Hyper-parameters

We run grid search on the learning rate and export the best model, using values {5e-6, 8e-6, 9e-6, 1e-5, 2e-5, 3e-5}. We use batch size 32 and evaluate the model every 1,000 steps on a 10% held-out data set to prevent over-fitting. During preliminary experiments, we additionally experimented with the batch size, dropout rate, frequency of continuous evaluation, balance of languages, pre-training schemes, WordPiece vocabularies, and model architecture.

3 Additional Models and Baselines

We fine-tune a new Bleurt checkpoint, following the methodology described above. The main difference with Sellam et al. 2020 is that we incorporate the to-English ratings of year 2019, which were not previously available.

Monolingual baselines based on BERT

We experiment with three baselines and submit the results to the WMT Metrics Shared Task for analysis. BERT-L2-base and BERT-L2-large are two regression models based on BERT and trained on to-English ratings. We use the same setup as English Bleurt, but we omit the mid-training phase. A similar approach was described in Shimanaka et al. 2019. BERT-Chinese-L2 is similar to BERT-L2-base, but it uses BERT-Chinese and it is fine-tuned on to-Chinese ratings.

Other Systems

We compare our setups to other state-of-the-art learned metrics: BERTscore (Zhang et al. 2020), and Yisi (Lo 2019) all apply rules on top of BERT embeddings while ESIM (Mathur et al. 2019) is a neural sentence similarity model. PRISM (Thompson and Post 2020) trains a multilingual translation model that is used as a zero-shot paraphrasing system. All the aforementioned systems take sentences pairs as input. Concurrent work has investigated incorporating the source with great success Rei et al. 2020. We leave this line of research for future work.

Results

Tables 1 and 2 show the results in the X→En\texttt{X}\rightarrow\texttt{En} direction, at the segment- and system-level respectively. In the majority of cases, one of the Bleurt configurations yields the strongest results. The original Bleurt metric seems to perform better at the segment-level. At the system-level it may be dominated by PRISM (3 out of 7 language pairs) or by one of the simpler BERT-based models (4 out of 7 language pairs).

Tables 3 and 4 present the results for the other languages. mBERT-WMT yields solid results at the segment-level (it achieves the highest correlations for 7 out of 11 language pairs), in particular for the “zero-shot” setups, En→Gu\texttt{En}\rightarrow\texttt{Gu}, En→Kk\texttt{En}\rightarrow\texttt{Kk}, and En→Lt\texttt{En}\rightarrow\texttt{Lt}. It outperforms mBERT consistently, except for En→Ru\texttt{En}\rightarrow\texttt{Ru} and En→Zh\texttt{En}\rightarrow\texttt{Zh} where it lags behind the other metrics. The results are consistent at the system-level.

Based on these results, we make two “competitive” submissions. We present Bleurt as described above, which we ran on all the X→En\texttt{X}\rightarrow\texttt{En} sentence pairs. Additionally, we submitted a multilingual system that combines mBERT-WMT (for all languages except Chinese) and BERT-Chinese-L2 (for Chinese). We ran the multilingual system for all language pairs including to-English, as the large amount of non-English fine-tuning data made available in 2019 may benefit this setup too. We also release the predictions of BERT-Base-L2, BERT-Large-L2, and mBERT for analysis.

Additional Improvements on English→\toGerman

For English→\toGerman, the organizers of WMT20 provide three different reference translations: two standard references and one additional paraphrased reference. Given this novel setup, we investigate how to combine our predictions. Moreover, we use a similar framework to ensemble the predictions of different metrics. In particular, we average the predictions of BLEURT, YiSi-1 and YiSi-2. All three metrics are different in their approaches. While BLEURT and YiSi-1 are reference-based metrics, Yisi-2 is reference-free and calculates its score by comparing translations only to the source sentence. BLEURT is fine-tuned on previous human ratings, while YiSi-1 is based on the cosine similarity between BERT embeddings of the reference and the candidate.

In the remainder of this section, we report Bleurt results using the mBERT-WMT setup unless specified otherwise. We use a different checkpoint from the one described in Section 4. The model was trained for 880K steps instead of 1 million, and it uses a sequence length of 256 tokens instead of 128.

Before combining BLEURT and YiSi, we perform a series of modifications to YiSi-1 and evaluate their impact on English→\toGerman.

All experimental results are summarized in Table 5. We report both segment-level (DARR) and system-level (Kendall τ\tau) correlations. To replicate the multi-reference setup of 2020, we compute correlations with the standard WMT references as well as the paraphrased reference from Freitag et al. 2020.

Improving YiSi’s Predictions

Our baseline is similar to the YiSi-1 submission from WMT 2019 Lo 2019: we run YiSi-1 with the public multilingual mBERT checkpoint. We then experiment with the underlying checkpoint. We continued pre-training mBERT on the in-domain German NewsCrawl dataset. The resulting model +pre-train NewsCrawl layer 9 increases the correlation for both reference translations. We improve the correlation further on the paraphrased reference by using the 8th instead of the 9th layer.

Other experiments

We tried pre-training BERT on forward translated sentences from German NewsCrawl, to adapt the word embeddings to MT outputs. We also trained a BERT model from scratch on the German NewsCrawl data. These experiments did not result in higher correlations with human ratings.

2 Combining Bleurt, YiSi-1 and YiSi-2 on Multiple References

We describe our two submissions to WMT 2020, YiSi-comb and All-comb, which result from our efforts to use multiple references for automatic evaluation. YiSi-comb is a multi-reference version of the YiSi score Lo 2019 aimed at achieving better system-level correlations. All-comb leverages metrics from Bleurt, YiSi-1, and YiSi-2 on multiple references to achieve better segment-level correlation.

YiSi scores are F1F_{1} scores of YiSi precision and YiSi recall. For the YiSi-comb submission, we take the minimum of the YiSi recalls for the three different references as the multi-reference recall, and the maximum of the YiSi precision as the multi-reference precision. Using the same notations as in (Lo 2019), the final score is the F1F_{1} of the recall and precision computed with α = 0.7\alpha~\textrm{=}~0.7 (see Figure 1). This submission aims to maximize the system-level correlation.

As shown in Table 5, YiSi-1 has the highest system-level correlation on paraphrased references. Given that we used α = 0.7\alpha~\textrm{=}~0.7, YiSi scores are quite similar to YiSi recalls (when α = 1.0\alpha~\textrm{=}~1.0, YiSi scores are equal to YiSi recalls). YiSi-1 scores for paraphrased references are usually much lower than those of standard references, therefore taking the minimum recall is oftentimes equivalent to taking the YiSi recall from the paraphrased references. Furthermore, we found that using the maximum precision, in combination with aggregating recalls, usually performs the best.

All-comb

We combined the predictions of YiSi-1 with those of Bleurt and YiSi-2. YiSi-2 usually performs worse than the reference-based metrics, but we found that incorporating its predictions can help. Having three different metrics (Bleurt, YiSi-1, YiSi-2) and three different reference translations, we take all seven predictions and average the scores for each segment. The combined prediction All-comb outperforms every single metric at the segment level, though the system-level correlation drops in comparison to the best YiSi-1 score on paraphrased references. This submission aims to maximize the segment-level correlation.

Summary

We submit the following systems to the WMT Metrics shared task:

Bleurt as previously published, fine-tuned on the human ratings of the WMT Metrics shared task 2015 to 2019, to-English.

A multi-lingual extensions of Bleurt based on a 20 languages variant of mBERT and BERT-Chinese.

Three baseline systems based on BERT-base, BERT-large, and mBERT.

Two combination methods for English to German that use YiSi and alternative references, YiSi-comb and All-comb.

Acknowledgements

Thanks to Xavier Garcia and Ran Tian for advice and proof-reading.

References