Scaling up COMETKIWI: Unbabel-IST 2023 Submission for the Quality Estimation Shared Task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, André F. T. Martins
Introduction
Quality Estimation (QE) is the task of automatically assigning a quality score to a machine translation output without depending on reference translations Specia et al. (2018). This paper details the collaborative effort of Unbabel and Instituto Superior Técnico (IST) in the WMT23 Quality Estimation shared task, which encompassed two primary tasks: (i) sentence- and word-level quality prediction and (ii) fine-grained error span detection.
As of last year, some language pairs in the test set were absent from the training data. To address this, following a similar approach to the previous year, our systems were developed to achieve good multilingual generalization and to accommodate previously unseen languages. To achieve this, we start by leveraging the direct assessments (DA) labeled data obtained from the WMT Metrics shared task from 2017 to 2020, the MLQE-PE dataset Fomicheva et al. (2022), and the training data (DA) specifically annotated for Indian languages in the 2023 shared task edition. In total, these datasets encompass close to 1M annotations covering 38 language pairs. We start by constructing generic models using this corpus. These generic QE models were subsequently fine-tuned for this year’s subtasks.
For Task 1 – sentence-level, we fine-tuned our generic models exclusively with this year’s DA data. The architecture of these models remains consistent with our submission from the previous year, but we employ XLM-R XL and XXL as pretrained encoders Conneau et al. (2020). For the word-level quality prediction task, we follow the successful approach of combining the sentence- and word-level signals into one loss during the finetuning step, which has yielded positive results in previous iterations Rei et al. (2022b). For fine-grained error span detection, we conducted experiments exploring various approaches that build upon our word-level and sentence-level strategies. In terms of contrasting systems, we explored UnbabelQihttps://qi.unbabel.com/ and GPT-4 OpenAI (2023). For GPT-4, we used a prompt designed to predict both the location and severity of errors in each translation, akin to the approach used in AutoMQM Fernandes et al. (2023).
Overall, our main contributions are: (i) we introduce approaches for multilingual machine translation quality estimation that are consistently first-ranked at word-, span-, and sentence-level granularity; (ii) we explore different approaches to predict the span of problematic translations along with their error severities (ok, Minor, Major); (iii) we publicly release two of our best models for research purposes (CometKiwi -XLhttps://huggingface.co/Unbabel/wmt23-cometkiwi-da-xl and -XXLhttps://huggingface.co/Unbabel/wmt23-cometkiwi-da-xxl). To the best of our knowledge, these are the largest open-source QE models publicly released.
Our submitted systems attain the top multilingual results in all tasks: For Task 1 sentence-level prediction, our multilingual system achieves 59.4 Spearman correlation points, surpassing the second-best system by nearly 4 absolute points. For word-level, our system achieves a 31.7 MCC score, outperforming the second-best system by almost 2 absolute MCC points. For error span prediction, our multilingual system achieves a 22 F1.0 score, beating the second-best system by more than 5 points.
Overview of the shared-task
submissions were evaluated on correlation with DA annotations for 5 language pairs (en-mr, en-hi, en-ta, en-te, en-gu) and MQM annotations for 3 language pairs (en-de, zh-en and he-en). Training data was released or is available for all directions but he-en.
submissions were evaluated on tags inferred from post-editions for 2 language pairs (en-fa, en-mr), and MQM annotations for 3 language pairs (en-de, zh-en and he-en). No additional training or development data with word-level tags were made available. To the best of our knowledge, no word-level data is available for en-fa and he-en.
submissions were evaluated on error spans obtained via MQM annotations for 3 language pairs (en-de, zh-en and he-en). No training nor development data is available for he-en.
Implemented Systems
Similarly to (Rei et al., 2022b), we employ InfoXLM L (Chi et al., 2021).https://huggingface.co/microsoft/infoxlm-large Additionally, we experiment with scaled-up multilingual encoders, including XLM-R XL,https://huggingface.co/facebook/xlm-roberta-xl and XLM-R XXL.https://huggingface.co/facebook/xlm-roberta-xxl InfoXLM L comprises 24 encoder blocks with 16 attention heads each, totaling 550M parameters. XLM-R XL and XLM-R XXL have 32 attention heads for each encoder block, 36 and 48 encoder blocks and a total of 3.5B and 10.7B parameters, respectively.
We create, for each model size, a generic model that will then be further adapted to each separate task. To train these models, we use the collective corpora from 2017 to 2019 DA annotations of the WMT Translation shared task, and the MLQE-PE corpus Fomicheva et al. (2022). We include the human annotations respective to the language pairs of this year’s shared task for 7 different language pairs: DA annotations for en-mr, en-hi, en-ta, en-te, en-gu, and MQM annotations for en-de and zh-en. Overall, the generic models are trained on sentence-level quality prediction with over 940k samples with source, translation and quality score on 38 different language pairs. When presented with multiple DA scores for the same sentence pair, we used the z-score of the DAs for training but we first normalize the DAs between 0 and 1, where 1 represents a perfect translation and 0 a random one.
After having obtained the generic models, we will train models for each separate stream of the shared-task, i.e., sentence-level, word-level or error span prediction. To do so, we consider the multi-task optimization from Rei et al. (2022b) wherein sentence scores can be used alongside supervision from word-level tags. Formally,
For error span detection, we evaluate UnbabelQi, an Unbabel demo QE system, alongside GPT4 (OpenAI, 2023). We prompt GPT4 to produce an MQM annotation for each source-target pair, based on five-shot examples which vary across language pairs but are consistent within segments of the same language pair. We also apply this system in Task 1, deriving a sentence-level score from error spans, in alignment with the MQM framework. This approach bears similarity to AutoMQM (Fernandes et al., 2023).
1 Task 1: Quality prediction
After the pretraining phase, we further separately adapt the generic models to the released DA and MQM data for this year’s shared task.
To further adapt the models to this year’s language pairs, we fine-tuned the generic models using, exclusively, the newly released DA annotations from this year. This approach yields additional improvements for those languages. In the case of the MQM language pairs, our preliminary experiments revealed that attempting significant performance improvements on the MQM data led to noteworthy drops in correlations for the other language pairs using DAs. Consequently, for the MQM language pairs, we opted to employ the generic models as they are.
Similarly to Rei et al. (2022b), we use Optuna Akiba et al. (2019) to assemble four models – two XL and two XXL – into a single system. We do so by finding the optimal weights for each language pair among these four multilingual models, and combining their predictions according to those weights. Notably, the XXL models are generic models, whereas the two XL checkpoints were further optimized with this year’s shared task data. As expected, the XL models carry more weight for Indian languages, while the XXL generic models were deemed more crucial for MQM languages.
1.2 Word-level quality prediction
For the word-level QE tasks, we experimented with both the multi-task setting and word-labels only.
This year, no training or development data with word-level tags were made available. As such, the training data for our models consists of the training data used in Rei et al. (2022b), combined with the development sets from the 2022 WMT Shared Task. As the word-level task was going to be tested in a zero-shot scenario for two out of five language pairs (en-fa, he-en), contrary to Rei et al. (2022b), we do not prepend a language prefix to the beginning of the source and target segments during training. Moreover, for the post-edit (PE) models, we removed samples from two language pairs (ps-en and en-cs) from the training data. We did so to assess, during validation, the models’ capability to generalise in a zero-shot scenario. For the MQM models, we used all available annotations, including those in en-ru.
For word-level we followed a similar ensembling technique used for sentence-level. Specifically, we combined multiple systems trained with different hyperparameters, encoder size and pre-training setups. In the case of word-level predictions, we aggregate multiple predictions into OK/BAD tags by following the ensemble-tags procedure from Rei et al. (2022b). In this approach, we combine the predicted tags of each model: for every input segment, we get a combined tag, , where is the tag predicted by the model and is the weight for the bad tag. We use Optuna to determine the optimal weights for each model and the optimal bad weight for each LP. In the final submission, we combine six models (five PE models and one MQM model). Five of these models use InfoXLM as the encoder model, and one PE model uses XLM-R XL.We found it hard to obtain performance boosts by scaling up to XLM-R XL on the word-level task. As such, we did not experiment with XLM-R XXL. Refer to Table 3.2 for the test set results.
2 Task 2: Fine-grained error span detection
In this task, we investigated three distinct approaches. The first approach extends word-level models by modifying their output predictions. More precisely, it involves transforming consecutively predicted bad tags into character-level error spans, rather than categorizing individual words based on the first subword. To determine the error severities of these spans, we considered two options: labeling all the subwords within the span as either minor or major. Our best results were achieved with the latter approach.
The second approach leverages xComet (TBA)Further details about xComet will be provided soon. in conjunction with a pseudo-reference obtained from DeepL or Google Translate.We choose the best translation using the generic XXL model from task 1. Similar to our models from Task 1 word-level, xComet is trained with a multitask objective. Additionally, xComet is simultaneously optimized for both reference-free and reference-based evaluation, following UniTE Wan et al. (2022). During inference, xComet can leverage a reference translation to enhance error identification. Since we employ a pseudo-reference that may contain translation errors, we initially assess the quality of the pseudo-reference using a generic QE system from Task 1 (reference_score). For all pseudo-references with a score below , we run xComet with QE-only input. For pseudo-references scoring above , the input weights for xComet are determined as follows:
Here, src_weight represents the weight assigned to the source-only input, ref_weight denotes the typical metric input (reference-only input), and uni_weight represents a unified input where the model receives all three sentences (translation, source, and reference). Notably, for pseudo-references with a QE score of 1, we rely solely on a reference-only input and the unified input. We refer to this approach as xcomet-ps-ref.
We also contrast the aforementioned approaches with two unconstrained QE systems: UnbabelQi and GPT-4, as mentioned in Section 3. We refer to these approaches as UnbabelQi and GPT4-QE, respectively.
We present the results on the official test set for each of the tasks for multiple model/data configurations. Sentence-level submissions were evaluated using the Spearman rank correlation. Pearson and Kendall correlation were also used as secondary metrics, but here we report only Spearman since it was the primary metric used to rank systems. word-level submission were evaluated using MCC, -OK, and -BAD, but we report only MCC as it was considered the main metric. Error span detection was evaluated using score in which the positive labels are all the characters belonging to erroneous spans. Furthermore, each true positive is downweighted to half if the system failed to classify the error span’s severity (e.g., minor instead of major). The submitted systems were independently evaluated on in-domain and zero-shot LPs for direct assessments and MQM.
Results for sentence-level are presented in Table 1. Results indicate that retraining the system from the previous year, specifically CometKiwi with InfoXLM, using data that encompasses this year’s DA, leads to significant improvements. Remarkably, this improvement in correlations is achieved while maintaining the same level of correlations for en-de (a high-resource language pair for which both models share the same data) and he-en, a language pair that both models had not seen during training. Surprisingly, there was a drop in correlations for zh-en even though both models saw the same zh-en data. Nevertheless, the overall performance of the newly retrained version improved by 4.1 Spearman points.
As anticipated, among the three backbone transformers, the XXL model is the top performer, with significant improvements across all language pairs when compared to InfoXLM. Moreover, additional finetuning on this year’s training data results in further improvements for the Indian languages. Notably, concerning the MQM data, this supplementary finetuning step not only preserves performance but sometimes even increases it. Similar to last year, the ensemble of high-performing models once again makes up our best submission.
Finally, despite performing well in Task 2, GPT4-QE shows poor correlations at sentence-level prediction with the exception of the en-de for which GPT4-QE, although lagging behind the ensemble approach, surpasses our individual models.
We report the best individual systems Table 3.2. Our best individual systems were trained on top of the InfoXLM L generic model. For PE models, we used multi-task objective in Eq. 4, as we found that combining the sentence-level and word-level loss was beneficial. However, for MQM models, we trained word-level only models, by setting and .
Interestingly, we found that PE models are very competitive on MQM language pairs. For example, the best overall performance for he-en was actually obtained with a PE word-level model. This is also reflected on the Optuna weights obtained for our final ensemble, wherein the weights of the PE models are significantly higher than those of the MQM models for all language pairs but en-de. In fact, our final ensemble for en-zh and en-he consists solely of PE models trained with different learning rates, , and . Further investigation on two different vectors may lead to improved word-level models: (i) balancing DA and MQM word-level annotations, and (ii) appropriately leveraging the larger capacity of scaled up encoder models.
Final Remarks
Results for fine-grained error span detection are shown in Table 4.1. Using a word-level model to obtain error span predictions leads to reasonable performance, comparable to our unconstrained submission, UnbabelQi, a model directly tasked with error span detection. That said, xcomet-ps-ref, an error span detection model, surpassed both of the previous approaches. We attribute the improved performance to this system being an ensemble of two significantly larger models, and to the usage of a pseudo-reference. We found the latter to be particularly beneficial on he-en, a language pair for which we had no training data.
The best approach in terms of average was GPT4-QE, mostly due to the improved performance on en-de. While this is a promising finding for LLM-based quality estimation systems, there are limitations. First, obtaining a sentence-level score from the error spans (as per the MQM framework) leads to poor correlations with human judgements derived from DA (see Table 1) and with low-resource language-pairs like he-en. Second, despite being useful in practice and leading to gains in , it is hard to control GPT’s precision and recall. We found that the number of examples included in the prompt, their ordering, and the number of errors within each example led to noticeable changes in the system’s propensity to flag errors. Thirdly, running QE with a system such as GPT-4 is expensive and slow even for a shared task exercise.
We describe Unbabel and IST joint submission to WMT23 QE shared task. Our approaches correlate well with human judgements for all the three granularities of translation quality prediction, ranking first in all multilingual tasks and surpassing the previous state-of-the-art model, CometKiwi-22, by up to 10 Spearman correlation points. Overall, our models follow the same architecture of last year’s participation, CometKiwi. However, this year we leverage more data and larger encoder models. Our best final systems are ensembles of different models trained on DA, post-edits or MQM scores that complement each other. Interestingly, our best systems surpass GPT-4 by a large margin for sentence-level translation quality prediction, and they are comparable to GPT-4 at error span detection.