CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C. Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte M. Alves, Alon Lavie, Luisa Coheur, André F. T. Martins
Introduction
Quality Estimation (QE) is the task of automatically assigning a quality score to a machine translation output without depending on reference translations Specia et al. (2018). In this paper, we describe the joint contribution of Instituto Superior Técnico (IST) and Unbabel to the WMT22 Quality Estimation shared task, where systems were submitted to three tasks: (i) Sentence and Word-level Quality Prediction; (ii) Explainable QE; and (iii) Critical Error Detection.
This year, we leverage the similarity between the tasks of MT evaluation and QE and bring together the strengths of two frameworks, Comet (Rei et al., 2020), which has been originally developed for reference-based MT evaluation, and OpenKiwi (Kepler et al., 2019), which has been developed for word-level and sentence-level QE. Namely, we implement some of the features of the latter, as well as other new features, into the Comet framework. The result is CometKiwi, which links the predictor-estimator architecture with Comet training-style, and incorporates word-level sequence tagging.
Given that some language pairs (LPs) in the test set were not present in the training data, we aimed at developing QE systems that achieve good multilingual generalization and that are flexible enough to account for unseen languages through few-shot training. To do so, we start by pretraining our QE models on Direct Assessments (DAs) annotations from the previous year’s Metrics shared task as it was shown to be beneficial in our previous submission (Zerva et al., 2021). Then we fine-tune our models with the data made available by the shared task.For zero-shot LPs we use only the 500 training examples. We experimented with different pretrained multilingual transformers as the backbones of our models, and we developed new explainability methods to interpret them. We describe our systems and their training strategies in Section 3. Overall, our main contributions are:
We combine the strengths of Comet and OpenKiwi, leading to CometKiwi, a model that adopts Comet training features useful for multilingual generalization along with the predictor-estimator architecture of OpenKiwi.
Following our previous work (Zerva et al., 2021), we show the importance of pretraining QE models on annotations from the Metrics shared task.
We show that we can improve results for new LPs with only 500 examples without harming correlations for other LPs.
We propose a new interpretability method that uses attention and gradient information along with a head-level scalar mix module that further refines the relevance of attention heads.
Our submitted systems achieve the best multilingual results on all tracks by a considerable margin: for sentence-level DA our system achieved a 0.572 Spearman correlation (+7% than the second best system); for word-level our system achieved a 0.341 MCC score (+2.4% than the second best system); and for Explainable QE our system achieved 0.486 R@K score (+10% than the second best system). The official results for all LPs are presented in Table 7 in the appendix.
Background
The transformation maps rows to distributions, with softmax being the most common choice, . Multi-head attention is computed by evoking Eq. 1 in parallel for each head :
Implemented Systems
Every year, the WMT News Translation shared task organizers collect human judgments in the form of DAs. The collective corpora of 2017, 2018, and 2019 contain 24 LPs and a total of 657k samples with source, target, reference, and DA score. We follow our experiments from last year Zerva et al. (2021) and start by pretraining our QE models on this data using the learning objective proposed by UniTE Wan et al. (2022), which incorporates reference translations into training and thus acts as data augmentation.
We follow the recent trend (Kepler et al., 2019; Ranasinghe et al., 2020) and experiment with three different pretrained multilingual transformers as the encoder layer of our models: XLM-R Large (Conneau et al., 2020),https://huggingface.co/xlm-roberta-large InfoXLM Large (Chi et al., 2021),https://huggingface.co/microsoft/infoxlm-large and RemBERT (Chung et al., 2021).https://huggingface.co/google/rembert XLM-R and InfoXLM consist of 24 encoder blocks with 16 attention heads each, whereas RemBERT has 32 encoder blocks with 18 attention heads each.
1 Task 1: Quality prediction
After the pretraining phase, we adapt our models to the released QE data using source and translation (i.e., in this phase we do not include references) to the different type of quality assessments provided, namely, DA and HTERHTERs are available only for word-level subtasks. from the MLQE-PE corpus Fomicheva et al. (2022) and MQM annotations from WMT 2020 and 2021 Freitag et al. (2021a, b).
For the sentence-level QE task we consider a multi-task setting (using sentence scores alongside supervision from ok/bad tags) and the sentence-level only setting, with supervision only from the sentence-level quality assessment . We found that adding the word-level supervision was beneficial for models built on top of InfoXLM. For the sentence-level supervision we used both DA and MQM scores. In this multi-task setting we use a combined loss as described in Eq. 5:
Since in this shared task submissions are tested on 5 LPs for which there is no official training data (km-en, ps-en, en-ja, en-cs, en-yo), we experimented with few-shot adaptation using half of the data released in the official development set. The official development set has 1K examples for each language pair (except en-yo for which there is no available data). To perform few-shot language adaptation we split the data into two halves: one for fine-tuning and another for validation.
For our final submission for Direct Assessments we combine six multilingual systems using different hyperparameters by computing an weighted average of their outputs, where the weights for each language pair were tuned with Optuna (Akiba et al., 2019). The major difference between the ensembled models comes from the underlying encoder and whether or not they used word-level supervision. Three models of our final ensemble use word-level supervision while the other three use only sentence-level supervision. Regarding the encoder, three models use InfoXLM, two models use RemBERT and a single model uses XLM-R.
Our final submission for MQM predictions was an ensemble of eleven multilingual systems, which combined the six systems used in the DA ensemble as well as five additional systems. For these additional systems, we made two major adjustments to the fine-tuning process. First, we filtered the DA data to the languages that were included in the MQM LPs, namely ru-en, en-zh, and en-de. Second, we incorporated the MQM data into the fine-tuning process, either as an additional fine-tuning step after fine-tuning on the language-filtered DA data, or by concatenating the DA and MQM data together. All additional systems used word-level supervision in addition to sentence-level and used InfoXLM as encoder.
1.2 Word-level quality prediction
Similarly, for the word-level QE tasks we experimented with both the multi-task setting and word-labels only ( and ). Overall, we found that adding the sentence-level supervision was beneficial, especially for the languages pairs included in the test-set. Nonetheless, for some LPs, ignoring sentence-level supervision showed superior performance. Due to the mix of high-, mid- and low-resource languages in the data, the distribution of ok and bad tags differs substantially between LPs leading to inconsistent performance in terms of MCC (see Table 5 in the appendix). To mitigate this, for the word-level subtask, we prepend a language prefix token to the beginning of the source and target segments during training and testing.
Extending the pretraining on Metrics data, we pretrain the word-level models on two corpora that include both word-level labels and sentence (HTER) scores, namely QT21 Specia et al. (2017) and APEQuest Ive et al. (2020). We compute the sentence-level score, using translation edit rate (TER) Snover et al. (2006) between the target and the corresponding post-edited sentence.
For word-level we followed a similar ensembling technique used for sentence-level, namely we combine multiple systems trained with different hyperparameters, encoders and pre-training setups. In the case of word-level predictions however, we need to resolve how to aggregate multiple predictions into OK/BAD tags. We use Optuna (Akiba et al., 2019) to choose how to weight and combine the models based on performance for each language pair on our internal test-set and we compare three different approaches:
A naive “best-only” approach: we identify the best model for each LP and use its predictions.
We ensemble the logits of each model: for each input segment we compute an ensembles of logits as , where is the set of models, is the weight of each model and the model logit vector. We use Optuna to find the optimal weight for each model in each LP.
We ensemble the predicted tags of each model: for each input segment we compute an ensembles of tags as , where is the predicted class and is the weight given for the bad class. We use Optuna to find the optimal weights for each model and the optimal bad weight for each LP.
In the final submission we combine five models for the post-edit originated LPs: a RemBERT based model, an InfoXLM based model pretrained on APEQuest and QT21, and three checkpoints that are based on InfoXLM but use different parameters for the bad/ok weights and learning rate that were found via Optuna. For MQM we also combine five models, but this time instead of choosing three checkpoints based on optimising weights and learning rate, we use three different checkpoints with different training data mix on the relevant DA LPs, as this seemed to impact the performance on MQM word-level more than the weight ratios. Refer to §4 and Table 3 for more details.
2 Task 2: Explainable QE
The goal of the Explainable QE task is to identify machine translation errors without relying on word-level label information. In other words, it can be cast as an unsupervised word-level quality estimation problem, where explanations can be seen as highlights, representing the relevance of input words w.r.t. the model’s prediction via continuous scores, aiming at identifying tokens that were not properly translated.
Head Mix: We reformulate the scalar mix module (Eq. 2) to consider different weights for representations coming from different attention heads as follows:
Furthermore, since all of our sentence-level models use subword tokenization, to get explanations for an entire word we follow Treviso et al. (2021) and sum the scores of its word pieces.
In our final submissions we average the explanation scores of different attention heads and layers to create a final explainer. We decided which heads and layers to aggregate together by looking at their performance on the dev set, selecting the top-5 with the highest explainability score.
3 Task 3: Critical Error Detection
Critical translations are defined as translations with strongly semantic deviations from the original source sentence, with the potential to lead to negative impacts in critical applications. The goal of this task is to predict sentence-level scores indicating whether a translation contains a critical error. Since the evaluation metrics automatically account for different binarization thresholds to separate good translations from bad ones, for this task we employed a single sentence-level InfoXLM model from Task 1 that was trained on DA data. Moreover, we participated only in the constrained setting, meaning that we did not trained our systems specifically for this task. Therefore, our goal for this task was to validate whether our QE system from Task 1 was able to detect and differentiate translations with critical errors.
Experimental Results
As we have seen in Section 3, for our experiments we split the provided development sets into two equal size halves creating a new internal devset and an internal testset. The resulting sets contain 500 segments per language pair for both DA and MQM, word and sentence-level. As for baselines we used our submitted systems from previous shared tasks: for Task 1 we used the M1M-adapt Zerva et al. (2021), and for Task 2 we used the explainer Treviso et al. (2021). The official results for Task 1 and Task 2 are shown in Table 7.
Sentence-level submissions were evaluated using the Spearman’s rank correlation. Pearson’s correlation, MAE, and RMSE were also used as secondary metrics, but here we report only Spearman correlation since it was the primary metric used to rank systems. Word-level submission were evaluated using MCC, -OK, and -BAD, but we report only MCC as it was considered the main metric. The submitted systems were independently evaluated on in-domain and zero-shot LPs for direct assessments and MQM.
Results for sentence-level DAs can be seen in Table 1. The results show that the training strategies employed in CometKiwi, namely (i) pretraining models using Metrics data and (ii) incorporating references into training, lead to a correlation close to our best system from last year while disregarding the data from the MLQE-PE corpus. When fine-tuning on MLQE-PE data, we get overall improvements of , and further fine-tuning on new LPs gives overall improvement. Still, for the unseen LPs (km-en, ps-en, en-ja, en-cs), we got improvements between 2-3% with just 500 samples. Among the three backbone transformers, we noticed that InfoXLM is the one that leads to a higher Spearman correlation (+1.7% than XLM-R and RemBERT). Furthermore, including word-level supervision always maintains or improves the results, especially for InfoXLM. In contrast, RemBERT does not seem to benefit from this signal. We suspect that, for this task, the benefit of word-level supervision is not higher because the word-level information is coming from post-editions, which are conceptually different from DA annotations.
Results for sentence-level MQM systems are shown in Table 2. The results show that the two main techniques used for adapting to MQM data, filtering DA data to the three MQM LPs and using MQM data for fine-tuning, improved Spearman correlations for all LPs over the pure DA baseline, for both sentence-level and multi-task systems. However, these techniques improved certain LPs more than others, so combining them together improved multilingual scores even further. Overall, we noticed that our results for MQM data have a high variance. To mitigate this, we concatenated the DA and MQM datasets together for a single fine-tuning, resulting in our best individual system on our internal test set. Due to these peculiarities in the MQM LPs, we decided to ensemble systems tuned on both DA and MQM data. Our final ensemble did not have as strong results as the individual systems on our internal test set, yet, it showed superior performance upon submission to codalab leader-board.
For the word-level task we tuned models separately for the LPs that consisted of post-edit-derived word tags and the ones consisting of MQM-derived word tags; we report the Matthew’s correlation coefficient (MCC) in Table 3. We experimented with multi-tasking by adding sentence-level supervision to the word-level task and found that it boosts performance especially for the out-of-English translations. For the non-MQM LPs we used the HTER scores as sentence level targets as we found they lead to significantly higher correlations. We can also see that using the sentence-mix and the language prefix boosted the performance for all LPs, both in the MQM and post-edit originated LPs. Overall, the results show further improvements when we use the HTER scores of APEQuest and QT21 as additional pretraining data, but only for specific LPs. These findings merit further investigation, since the directionality of the LPs seems to have impacted our experiments. Finally, ensembling led to better results across all languages. Ensembling the logits led to better results for the post-edit originated LPs, while word-level ensembling helped more the MQM-originated LPs. Yet, in the submitted versions we found that the difference in performance between the three ensembling methods yielded similar results, with only 1-2% difference, while in the averaged multilingual versions these differences were even smaller, varying less than 0.1%.
2 Explainable QE
Since the explanations are given as continuous scores, they are evaluated against the ground-truth word-level labels in terms of the Area Under the Curve (AUC), Average Precision (AP), and Recall at Top-K (R@K) metrics only on the subset of translations that contain errors. Although R@K was considered the main metric for this task, we optimized internally for the average of all three metrics. The results are shown in Table 4.
Official Results
We present the official results of our submissions alongside the results from other competitors in Section B for all three tasks. For sentence-level, our submissions achieved the best results for 6/9 LPs. For word-level, we obtained the best results for 5/9 LPs. For the explainable QE track, we obtained the best results for all but two LPs (km-en and ps-en). Although the critical error detection task had no other competitor for the constrained setting, our submission vastly surpassed the organizers’ baseline. We also obtained the best results for the multilingual settings (including and excluding en-yo) for all tasks. Finally, when averaging the results for all LPs, our submissions place on top for all tasks.
Conclusions and Future Work
We presented the joint contribution of IST and Unbabel to the WMT 2022 QE shared task. We found that incorporating references during pretraining improves performance across several LPs on downstream tasks, and that jointly training with sentence and word-level objectives yields a further boost. For Task 1, our final submissions were ensembles of models finetuned with different pretrained language models as encoders, boosting the results when compared to the previous year submission. For Task 2, we take inspiration on the literature of explainability and propose to use gradient information in tandem with attention weights, and to further refine the impact of attention heads towards the prediction via the Head Mix component. Besides leading to better explainability performance for some LPs, this strategy is potentially useful to identify good attention heads at inference time for zero-shot LPs, and deserves more investigation. Overall, our submissions achieved the best results for all tasks (including Task 3) for almost all LPs by a considerable margin.
One of the challenges of leveraging big ensembles is the burdensome weight of parameters and inference time. For future work we will extend our recent work, Cometinho Rei et al. (2022) and explore how to effectively distill large ensembles into small and more practical QE systems.
Acknowledgements
This work was supported by the P2020 program MAIA (contract 045909), by the European Research Council (ERC StG DeepSPIN 758969), and by the Fundação para a Ciência e Tecnologia through contract UIDB/50008/2020.
References
Appendix A Data Information
The data used for finetuning our QE systems is shown in Table 5. For DA data, we split the original development set to generate a new dev/test split, therefore the reported numbers in the table correspond to this “internal” dev split.
Appendix B Official Results
Submissions for this task were evaluated in terms of ranking using R@K and MCC as metrics. In Table 6, we report only MCC scores as it was the main metric for this task.
Table 7 shows the official results for sentence-level QE (top) in terms of Spearman’s correlation, word-level QE (middle) in terms of MCC, and explainable QE (bottom) in terms of R@K.