UniTE: Unified Translation Evaluation
Yu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek F. Wong, Lidia S. Chao
Introduction
Automatically evaluating the translation quality with the given reference segment(s), is of vital importance to identify the performance of Machine Translation (MT) models Freitag et al. 2020; Mathur et al. 2020a; Zhao et al. 2020; Kocmi et al. 2021. Based on the input contexts, translation evaluation can be mainly categorized into three classes: 1) reference-only evaluation (Ref) approaches like BLEU Papineni et al. 2002 and BLEURT Sellam et al. 2020a, which evaluate the hypothesis by referring the golden reference at target side; 2) source-only evaluation (Src) methods like YiSi-2 Lo 2019 and TransQuest Ranasinghe et al. 2020b, which are also referred as quality estimation (QE). These methods estimate the quality of the hypothesis based on the source sentence without using references; 3) source-reference-combined evaluation (Src+Ref) works like COMET Rei et al. 2020, where the evaluation exploits information from both source and reference. With the help of powerful pretrained language models (Devlin et al. 2019; Conneau et al. 2020, PLMs,), model-based approaches (e.g., BLEURT, TransQuest, and COMET) have shown promising results in recent WMT competitions Ma et al. 2019; Mathur et al. 2020b; Freitag et al. 2021; Fonseca et al. 2019; Specia et al. 2020; Specia et al. 2021.
Nevertheless, each existing MT evaluation work is usually designed for one specific task, e.g., BLEURT is only used for Ref task and can not support Src and Src+Ref tasks. Moreover, those approaches preserve the same core – evaluating the quality of translation by referring to the given segments. We believe that it is valuable, as well as feasible, to unify the capabilities of all MT evaluation tasks (Ref, Src and Src+Ref) into one model. Among the promising advantages are ease of use and improved robustness through knowledge transfer across evaluation tasks. To achieve this, two important challenges need to be addressed: 1) How to design a model framework that can unify all translation evaluation tasks? 2) How to make the powerful PLMs better adapt to the unified evaluation model?
In this paper, we propose UniTE - Unified Translation Evaluation, a novel approach which unifies the functionalities of Ref, Src and Src+Ref tasks into one model. To solve the first challenge as mentioned above, based on the multilingual PLM, we utilize layerwise coordination which concatenates all input segments into one sequence as the unified input form. To further unify the modeling of three evaluation tasks, we propose a novel Monotonic Regional Attention (MRA) strategy, which allows partial semantic flows for a specific evaluation task. For the second challenge, a multi-task learning-based unified pretraining is proposed. To be concrete, we collect the high-quality translations and degrade low-quality translations of NMT models as synthetic data. Then we propose a novel ranking-based data labeling strategy to provide the training signal. Finally, the multilingual PLM is continuously pretrained on synthetic dataset with multi-task learning manner. Besides, our proposed models, named UniTE-MRA and UniTE-UP respectively, can benefit from finetuning with human-annotated data over three tasks at once, not requiring extra task-specific training.
Experimental results demonstrate the superiority of UniTE. Compared to various strong baseline systems on each task, UniTE, which unifies Ref, Src and Src+Ref tasks into one single model, achieves consistently absolute improvements of Kendall’s correlations at 1.1, 2.3 and 1.1 scores on English-targeted translation directions of WMT 2019 Metric Shared task Fonseca et al. 2019, respectively. Meanwhile, after introducing multilingual-targeted support for our unified pretraining strategy, a single model named UniTE-MUP also gives dominant results against existing methods on non-English-targeted translation evaluation tasks. Furthermore, our method can also achieve competitive results over WMT 2020 QE task compared with the winner submission Ranasinghe et al. 2020b. Ablation studies reveal that, the proposed MRA and unified pretraining strategies are both important for model performance, making the model preserve the outstanding performance and multi-task transferability concurrently.
Related Work
In this section, we briefly introduce the three directions of translation evaluation.
Ref assesses the translation quality via comparing the translation candidate and the given reference. In this setting, the two inputs are written in the same language, thus being easily applied in most of the metric tasks. In the early stages, statistical methods are dominant solutions due to their strengths in wide language support and intuitive design. These methods measure the surface text similarity for a range of linguistic features, including n-gram (Papineni et al. 2002, BLEU,), token (Snover et al. 2006, TER,), and character (Popovic 2015; Popovic 2017, ChrF & ChrF++,). However, recent studies pointed out that these metrics have low consistency with human judgments and insufficiently evaluate high-qualified MT systems Freitag et al. 2020; Rei et al. 2020; Mathur et al. 2020a.
Consequently, with the rapid development of PLMs, researchers have been paying their attention to model-based approaches. The basic idea of these studies is to collect sentence representations for similarity calculation (Zhang et al. 2020, BERTScore,) or evaluating probabilistic confidence Thompson and Post 2020, PRISM-ref,; Yuan et al. 2021, BARTScore,. To further improve the model, Sellam et al. 2020a pretrained a specific PLM for the translation evaluation (BLEURT), while Lo 2019 combined statistical and representative features (YiSi-1). Both these methods achieve higher correlations with human judgments than statistical counterparts.
2 Source-Only Evaluation
Src, which also refers to quality estimation Refer to “quality estimation” or “reference-free metric” in WMT (http://www.statmt.org/wmt19/qe-task.html, http://www.statmt.org/wmt21/metrics-task.html)., is an important translation evaluation task especially for the scenario where the ground-truth reference is unavailable. It takes the source-side sentence and the translation candidate as inputs for the quality estimation. To achieve this, the methods are required to model cross-lingual semantic alignments. Similar to reference-only evaluation, statistical-based (Ranasinghe et al. 2020b), model-based Ranasinghe et al. 2020b, TransQuest,; Thompson and Post 2020, PRISM-src,, and feature combination Lo 2019, YiSi-2, are typical and advanced methods in this tasks.
3 Source-Reference-Combined Evaluation
Aside from the above tasks that only consider either source or target side at one time, Src+Ref takes both source and reference sentences into account. In this way, methods in this context can evaluate the translation candidate via utilizing the features from both sides. As a rising paradigm among translation evaluation tasks, Src+Ref also profits from the development of cross-lingual PLMs. For example, finetuning PLMs over human-annotated datasets (Rei et al. 2020, COMET,) achieves new state-of-the-art results among all evaluation approaches in WMT 2020 Mathur et al. 2020b.
Methodology
As mentioned above, massive methods are proposed for different automatic evaluation tasks. On the one hand, it is inconvenient and expensive to develop and employ different metrics for different evaluation scenarios. On the other hand, separate models absolutely overlook the commonalities among these evaluation tasks, of which knowledge potentially benefits all three tasks. In order to fulfill the aim of unifying the functionalities on Ref, Src, and Src+Ref into one model, in this section, we introduce UniTE (Figure 1).
By receiving a data example composing of hypothesis, source, and reference segment, UniTE first modifies it into concatenated sequence following the given setting as Ref, Src, or Src+Ref:
where is the model size of PLM. According to Ranasinghe et al. 2020b, we use the first output representation as the input of feedforward layer.
Compared to existing methods Zhang et al. 2020; Rei et al. 2020 which take sentence-level representations for evaluation, the advantages of our architecture design are as follows. First, our UniTE model can benefit from layer-coordinated semantical interactions inside every one of PLM layers, which is proven effective on capturing diverse linguistic features He et al. 2018; Lin et al. 2019; Jawahar et al. 2019; Tenney et al. 2019; Rogers et al. 2020. Second, for the unified approach of our model, the concatenation provides the unifying format for all task inputs, turning our model into a more general architecture. When conducting different evaluation tasks, our model requires no further modification inside. Note here, to keep the consistency across all evaluation tasks, as well as ease the unified learning, is always located at the beginning of the input sequence.
For training, we encourage the model to reduce the mean squared error with respect to given score :
However, for the pretraining of most PLMs (Conneau et al. 2020, e,g., XLM-R,), the input patterns are designed to receive two segments at most. Thus there exists a gap between the pretraining of PLM and the joint training of UniTE where the concatenation of three fragments is used as input. Moreover, previous study Takahashi et al. 2020 shows that directly training over Src+Ref by following such design leads to worse performance than Ref scenario. To alleviate this issue, we propose two strategies: Monotonic Regional Attention as described in §3.2 and Unified Pretraining in §3.3.
2 Monotonic Regional Attention
To fill the modeling gap between the pretraining of PLM and the joint training of three downstream tasks, a natural idea is to unify the number of involved segments when modeling semantics for Src, Ref and Src+Ref tasks. Following this, we propose to modify the attention mask of Src+Ref to simulate the modeling of two segments in Src and Ref. Specifically, when calculating the attention logits, semantics from a specific segment are only allowed to derive information from two segments at most. Considering the conventional attention module:
where stores the index pairs of all masked areas.
Following this idea, the key of MRA is how to design the matrix . For the cases where interactions inside each segment, we believe that these self-interactions are beneficial to the modeling. For other cases where interactions are arranged across segments, three patterns are included: hypothesis-reference, source-reference, and hypothesis-source. Intuitively, the former two parts are beneficial for model training, since they might contribute the monolingual signals and cross-lingual disambiguation to evaluation, respectively. This leaves the only case, where our experimental analysis also verifies (see §5.1), that interaction between hypothesis and source leads to the performance decrease for Src+Ref task, thus troubling the unifying.
To give more fine-grained designs, we propose two approaches for UniTE-MRA, which apply the MRA mechanism into UniTE model (Figure 2):
Hard MRA. Only monotonic attention flows are allowed. Interactions between any two segments are strictly unidirectional through the entire PLM, where stores the index pairs of unidirectional interactions of , and , where “” denotes the direction of attention flows.
Soft MRA. Specific attention flows are forbidden inside each attention module. The involved two segments may interact inside a higher layer. In practice, index pairs which denoting or between source and hypothesis are stored in .
Note that, although the processing in source and reference may be affected because their positions are not indexed from the start, related studies on positional embeddings reveal that, PLM can well capture relative positional information Wang and Chen 2020, which dispels this concern.
3 Unified Pretraining
To further bridge the modeling gap between PLM and the joint training of UniTE mentioned in §3.1, we propose a unified pretraining strategy including the following main stages: 1) collecting and downgrading synthetic data; 2) labeling examples with a novel ranking-based strategy; 3) multi-task learning for unified pretraining and finetuning.
As our approach aims at evaluating the quality of translations, generated hypotheses with NMT models are ideal synthetic data. To further improve the diversity of synthetic data quality, we follow existing experiences Sellam et al. 2020a; Wan et al. 2021 to apply the word and span dropping strategy to downgrade a portion of hypotheses. The collected data totally contains triplets composing of hypothesis, source and reference segments, which is formed as .
Data Labeling
After obtaining the synthetic data, the next step is to augment each data pair with a label which serves as the signal of unified pretraining. To stabilize the model training, as well as normalize the distributions across all score systems and languages, we propose a novel ranking-based approach. This method is based on the idea of Borda count Ho et al. 1994; Emerson 2013, which provides more precise and well-distributed synthetic data labels than Z-score normalization.
where is the list storing all the sorted descendingly. Then, we use the conventional Z-score strategy to normalize the scores:
Compared to related approaches which apply Z-score normalization Bojar et al. 2018, or leave the conventional labeled scores as signals for learning (Kim and Rush 2016; Phuong and Lampert 2019, i.e., knowledge distillation,), our approach can alleviate the bias of chosen model for labeling and prior distributional disagreement of scores. For example, different methods may give scores with different distributions. Especially for translation directions of low-resource, scores may follow skewed distribution Sellam et al. 2020a, which has a disagreement with rich-resource scenarios. Our method can unify the distribution of all labeling data into the same scale, which can also be easily applied by the ensembling strategy.
Multi-task Pretrainig and Finetuning
To unify all evaluation scenarios into one model, we apply multi-task learning for both pretraining and finetuning. For each step, we arrange three substeps for all input formats, yielding , , and , respectively. The final learning objective is to reduce the summation of all losses:
Experiments
Following Rei et al. 2020; Yuan et al. 2021, we examine the effectiveness of the propose method on WMT 2019 Metrics Ma et al. 2019. For the former, we follow the common practice in COMET https://github.com/Unbabel/COMET Rei et al. 2020 to collect and preprocess the dataset. The official variant of Kendall’s Tau correlation Ma et al. 2019 is used for evaluation. We evaluate our methods on all of Ref, Src and Src+Ref scenarios. For Src scenario, we further conduct results on WMT 2020 QE task Specia et al. 2020 referring to Ranasinghe et al. 2020a for data collection and preprocessing. Following the official report, the Pearson’s correlation is used for evaluation.
Model Pretraining
As mentioned in §3.3, we continuously pretrain PLMs using synthetic data. The data is constructed from WMT 2021 News Translation task, where we collect the training sets from five translation tasks. Among those tasks, the target sentences are all in English (En), and the source languages are Czech (Cs), German (De), Japanese (Ja), Russian (Ru), and Chinese (Zh). Specifically, we follow Sellam et al. 2020a to use Transformer-base Vaswani et al. 2017 MT models to generate translation candidates, and use the checkpoints trained via UniTE-MRA approach for synthetic data labeling. We pretrain two kinds of models, one is pretrained on English-targeted language directions, and the other is a multilingual version trained using bidirectional data. Note that, for a fair comparison, we filter out all pretraining examples that are involved in benchmarks.
Model Setting
We implement our approach upon COMET Rei et al. 2020 repository and follow their work to choose XLM-R Conneau et al. 2020 as the PLM. The feedforward network consists of 3 linear transitions, where the dimensionalities of corresponding outputs are 3,072, 1,024, and 1, respectively. Between any two adjacent linear modules inside, hyperbolic tangent function is arranged as activation. During both pretraining and finetuning phrases, we divided training examples into three sets, where each set only serves one scenario among Ref, Src and Src+Ref to avoid learning degeneration. During finetuning, we randomly extracting 2,000 training examples from benchmarks as development set. Besides UniTE-MRA and UniTE-UP which are derived with MRA (§ 3.2) and Unified Pretraining (§ 3.3), we also extend the latter with multilingual-targeted unified pretraining, thus obtaining UniTE-MUP model.
Baselines
As to Ref approaches, we select BLEU Papineni et al. 2002, ChrF Popovic 2015, YiSi-1 Lo 2019, BERTScore Zhang et al. 2020, BLEURT Sellam et al. 2020a, PRISM-ref Thompson and Post 2020, BARTScore Yuan et al. 2021, XLM-R+Concat Takahashi et al. 2020, and RoBERTa+Concat Takahashi et al. 2020 for comparison. For Src methods, we post results of both metric and QE methods, including YiSi-2 Lo 2019, XLM-R+Concat Takahashi et al. 2020, PRISM-src Thompson and Post 2020 and multilingual-to-multilingual MTransQuest Ranasinghe et al. 2020b. For Src+Ref, we use XLM-R+Concat Takahashi et al. 2020 and COMET Rei et al. 2020 as strong baselines.
2 Main Results
Results on English-targeted metric task are conducted in Table 1. Among all involved baselines, for Ref methods, BARTScore Yuan et al. 2021 performs better than other statistical and model-based metrics. As to Src scenario, MTransQuest Ranasinghe et al. 2020b gives dominant performance. Further, COMET Rei et al. 2020 performs better than XLM-R+Concat Takahashi et al. 2020 on Src+Ref scenario.
As for our methods, we can see that, UniTE-MRA achieves better results on all tasks, demonstrating the effectiveness of monotonic attention flows for cross-lingual interactions. Moreover, the proposed model UniTE-UP, which unifies Ref, Src, and Src+Ref learning on both pretraining and finetuning, yields better results on all evaluation settings. Most importantly, UniTE-UP is a single model which surpasses all the different state-of-the-art models on three tasks, showing its dominance on both convenience and effectiveness.
Multilingual-Targeted
As seen in Table 2, the multilingual-targeted UniTE-MUP gives dominant performance than all strong baselines on Ref, Src and Src+Ref, demonstrating the transferability and effectiveness of our approach. Besides, the UniTE-UP also gives dominant results, revealing an improvement of 0.6, 0.3 and 0.9 averaged Kendall’s correlation scores, respectively. However, we find that UniTE-MUP outperforms strong baselines but slightly worse than UniTE-UP on English-targeted translation directions (see Table 3). We think the reason lies in the curse of multilingualism and vocabulary dilution Conneau et al. 2020.
Quality Estimation
The results for UniTE approach on WMT 2020 QE task are concluded in Table 4. As seen, it achieves competitive results on QE task compared with the winner submission Ranasinghe et al. 2020b.
Ablation Studies
In this section, we conduct ablation studies to investigate the effectiveness of regional attention patterns (§5.1), unified training (§5.2), and ranking-based data labeling (§5.3). All experiments are conducted by following English-targeted setting.
To investigate the effectiveness of MRA, we further collect experiments in Table 5. As seen, MRA can give performance improvements than full attention, and preventing the interactions between hypothesis and source segment can improve the performance most. We think the reasons behind are twofold. First, the source side is formed with a different language, whose semantic information is rather weak than the reference side. Second, by preventing direct interactions between source and hypothesis, semantics inside the former must be passed through reference, which is helpful for disambiguation. Besides, not allowing the source to derive information from the hypothesis is better than the opposite direction. Wang and Chen 2020 found that the positional embeddings in PLM are engaged with strong adjacent information. We think the reason why SH performs worse than HS lies in the skipping of indexes, which corrupts positional similarities in alignment calculation.
Additionally, when we combined two methods together, i.e., unified pretraining and finetuning with Src+Ref UniTE-MRA setting, model performance drops to 34.9 over English-targeted tasks on average. We think that both methods all intend to solve the problem of unseen Src+Ref input format, and MRA may not be necessary if massive data examples can be obtained for pretraining. Nevertheless, UniTE-MRA has its advantage on wide application without requiring pseduo labeled data.
2 Unified Training
Experiments for comparing unified and task-specific training are concluded in Table 6. As seen, when using the unified pretraining checkpoint to finetune over the specific task, performance over three models reveals performance drop consistently, indicating that the unified finetuning is helpful for model learning. This also verifies our hypothesis, that the cores of Ref, Src, and Src+Ref tasks are identical to each other. Moreover, unified pretraining and finetuning are complementary to each other. Also, utilizing task-specific pretraining instead of unified one reveals worse performance. To sum up, unifying both pretraining and finetuning only reveals one model, showing its advantage on the generalization on all tasks, where one united model can cover all functionalities of Ref, Src and Src+Ref tasks concurrently.
3 Ranking-based Data Labeling
To verify the effectiveness of ranking-based labeling, we collect the results of models applying different pseudo labeling strategies. After deriving the original scores from the well-trained UniTE-MRA checkpoint, we use Z-score and proposed ranking-based normalization methods to label synthetic data. For both methods, we also apply an ensembling strategy to assign training examples with averaged scores deriving from 3 UniTE-MRA checkpoints. Results show that, Z-score normalization reveals a performance drop when applying score ensembling with multiple models. Our proposed ranking-based normalization can boost the UniTE-UP model training, and its ensembling approach can further improve the performance.
Conclusion
In the past decades, automatic translation evaluation is mainly divided into Ref, Src and Src+Ref tasks, each of which develops independently and is tackled by various task-specific methods. We suggest that the three tasks are possibly handled by a unified framework, thus being ease of use and facilitating the knowledge transferring. Contributions of our work are mainly in three folds: (a) We propose a flexible and unified translation evaluation model UniTE, which can be adopted into the three tasks at once; (b) Through in-depth analyses, we point out that the main challenge of unifying three tasks stems from the discrepancy between vanilla pretraining and multi-tasks finetuning, and fill this gap via monotonic regional attention (MRA) and unified pretraining (UP); (c) Our single model consistently outperforms a variety of state-of-the-art or winner systems across high-resource and zero-shot evaluation in WMT 2019 Metrics and WMT 2020 QE benchmarks, showing its advantage of flexibility and convincingness. We hope our new insights can contribute to subsequent studies in the translation evaluation community.
Acknowledgements
The authors would like to send great thanks to all reviewers and meta-reviewer for their insightful comments. This work was supported in part by the Science and Technology Development Fund, Macau SAR (Grant No. 0101/2019/A2), the Multi-year Research Grant from the University of Macau (Grant No. MYRG2020-00054-FST), National Key Research and Development Program of China (No. 2018YFB1403202), and Alibaba Group through Alibaba Research Intern Program.
References
Appendix A Collection of Pretraining Data
Considering the English-targeted model, we select Czech (Cz), German (De), Japanese (Ja), Russian (Ru), and Chinese (Zh) as source languages, and English (En) as target. For each translation direction, we collect 1 million samples, finally yielding 5 million examples in total for unified pretraining. As to the multilingual-targeted model, we further collect 1 million synthetic data for each language direction of En-Cz, En-De, En-Ja, En-Ru, and En-Zh. Finally, we construct 10 million examples for the pretraining of the multilingual version by adding the data of the English-targeted model. Note that, for a fair comparison, we filter out all pretraining examples that are involved in benchmarks.
Appendix B Reproducibility
All the models reported in this paper were finetuned on a single Nvidia V100 (32GB) GPU. Specifically for UniTE-UP and UniTE-MUP, the pretraining is arranged on 4 Nvidia V100 (32GB) GPUs. Our framework is built upon COMET repository Rei et al. 2020. For the contribution to the research community, we release both the source code of UniTE framework and the well-trained evaluation models as described in this paper at https://github.com/NLP2CT/UniTE.