Findings of the WMT 2023 Shared Task on Discourse-Level Literary Translation: A Fresh Orb in the Cosmos of LLMs
Longyue Wang, Zhaopeng Tu, Yan Gu, Siyou Liu, Dian Yu, Qingsong Ma, Chenyang Lyu, Liting Zhou, Chao-Hong Liu, Yufeng Ma, Weiyu Chen, Yvette Graham, Bonnie Webber, Philipp Koehn, Andy Way, Yulin Yuan, Shuming Shi
Introduction
In past decades, the evolution of machine translation (MT) has undergone significant improvements in accuracy and efficiency, leading to many practical applications in various fields Bojar et al. (2014); Barrault et al. (2019); Farhad et al. (2021); Kocmi et al. (2022). Despite its success, MT still struggles in certain intricate scenarios to deliver translations that meet high standards Läubli et al. (2018); Koehn and Knowles (2017). Translating literary texts is considered to be the greatest challenge for MT due to its complex nature Toral and Way (2018); Toral et al. (2018); Ghazvininejad et al. (2018):
Rich Linguistic and Cultural Phenomena: literary texts contain more complex linguistic and cultural knowledge than non-literary ones Voigt and Jurafsky (2012); Ghazvininejad et al. (2018). To generate a cohesive and coherent output, MT models require an understanding of the intended meaning and structure of the text at discourse level Wang et al. (2016, 2018a, 2018b, 2019, 2023b). Furthermore, it demands skillful adaptation of cultural references, idioms, and subtle expressions to capture the essence of the original work in target languages.
Limited Data: existing document-level datasets are news articles and technical documents Liu and Zhang (2020); Thai et al. (2022); there is limited availability of copyrighted, discourse-level, parallel data in the literature domain. This makes it difficult to develop models that are able to handle the complexities of literary translation.
Long-Range Context: literature such as novels have much longer contexts than texts in other domains (e.g. news articles). Translation models need to acquire the capacity of modeling long-range context for learning translation consistency and lexical choice Wang et al. (2017); Wang (2019); Matusov (2019); Du et al. (2023).
Unreliable Evaluation Methods: literary evaluation needs to measure the meaning and structure of the text, and the nuances and complexities of the source language. A single automatic evaluation using a single reference is unreliable. Thus, professional translators with well-defined error typologies and targeted automatic evaluation are considered a complement Matusov (2019).
With the swift progression of MT and the notable advancements in Large Language Models (LLM) Ouyang et al. (2022b); OpenAI (2023), our curiosity is piqued regarding the efficacy of MT and LLM in the realm of literary translation. We aim to explore the extent to which these technologies can aid in addressing the intricate challenges of translating literary works. Therefore, we hold the first edition of the Discourse-Level Literary Translation in WMT 2023. Literary texts encompass a wide range of forms, including novels, short stories, poetry, plays, essays, and more. Among these, web novels, also known as online or internet novels, represent a unique and rapidly growing subset of literature. Their popularity, accessibility, and diverse genres set them apart. As they provide not only an extensive volume of text but also exhibit distinctive linguistic features, cultural phenomena, and simulations of societies, web novels can serve as valuable resources and challenging for MT research.
This year, the shared task mainly focuses on document-level web novels, and we introduce a document-level benchmark dataset and establish human evaluation criteria specifically tailored to address the challenges of literary translation:
Benchmark Dataset: We build and release a copyrighted and high-quality Chinese-English training corpus, comprising 2 million sentences sourced from 179 web fictions. This dataset preserves both book-level and chapter-level contexts, and features manually-aligned sentence pairs. We also provide three types of testsets, varying in distribution and document length (in Section 2).
Evaluation Methods: In order to evaluate the translation quality of the participating systems we used both automatic and human evaluation methods. About automatic evaluation, we employ document-level sacreBLEU (d-BLEU) as our metric, which is computed by matching n-grams in the whole document Liu et al. (2020); Post (2018). In terms of human evaluation, we propose a well-defined criteria by adapting multidimensional quality metrics (MQM) Lommel et al. (2014) to fit the context of literary translation. Note that all evaluations are case-sensitive (in Section 3).
We introduce the task overview and submission form in Section 4. This year, 14 submissions were received from 7 different teams, which are detailed in Section 5. We report the evaluation results in Section 6 followed by the conclusion in Section 7.
The GuoFeng Webnovel Corpus
We release a copyrighted and high-quality Chinese-English corpus on web novels. Additionally, we provide in-domain pretrained models as supplementary resources. As shown in Figure 1, a total of 45 institutes and companies from various regions have downloaded our dataset, showing that the prposed tasks and data have garnered widespread interest.
Copyright is a crucial consideration when it comes to releasing literary texts, and it is also one of the primary reasons for limiting the scale of data in this domain. We, Tencent AI Lab and China Literature Ltd., are the copyright owners of the web fictions included in this dataset. In order to promote the advancement of research in this field, we make this data available to the research community, subject to certain terms and conditions.
After registration, WMT participants can use the corpus for non-commercial research purposes and follow the principle of fair use (CC-BY).
Modifying or redistributing the dataset is strictly prohibited.
You should cite the this paper and claim the original download link.
The web novels are originally written in Chinese by web novel writers and then translated into English by professional translators. Our data processing involves a combination of automated and manual techniques: 1) we match Chinese books with its English counterparts based on bilingual titles; 2) within each book, Chinese-English chapters are aligned using Chapter ID numbers; 3) within each chapter, we build a MT-based sentence aligner to align sentences in parallel, preserving the sentence order in the chapter; 4) human annotators are engaged to review and correct any discrepancies in sentence-level alignment. To ensure the retention of discourse information, we permit null alignments. We totally spent 6 months addressing copyright issues and around 40,000 euros for human annotation. Figure 2 shows the final format of our corpus.
Training/Validation/Testing Data
Table 2.1 lists data statistics of our dataset. As seen, the training set contains 23K continuous chapters from 179 web novels, covering 14 genres such as fantasy science and romance. To enable participants to evaluate model performance by themselves, we provide two unofficial validation/testing sets with one reference. For dataset1, books overlap with the training data, whereas dataset2 contains unseen books. The participants can regard each chapter as a document to train and test their discourse-aware models. Apart from this, parallel training data in the General MT Task can also be used for data augmentation. In the final testing stage, participants use their systems to translate the official testing set (Testfinal). We select around 20 consecutive chapters from each book. Thus, we participants could treat all chapters within a book as a long documentThe participants can still regard one chapter as a document, which depends on the models’ length capability.. As seen, the document length of Testfinal is quite longer than other sets. The final testset contains two references: Reference 1 is translated by human translators and Reference 2 is bult by manually aligning bilingual text in web page. The genres in the valid and test sets are sampled evenly.
2 Pretrained Models
Apart from training dataset from web novels, we also provide in-domain pretrained models as supplementary resources. These models can be used to finetune or initialize MT models.
RoBERTa (base): The original model features a 12-layer encoder and is trained on the Chinese Wikipedia Liu et al. (2019). It has a hidden size of 768 and a vocabulary size of 21,128 using whole word masking. We continuously train it with Chinese literary texts (84B tokens) Wang et al. (2023a).
mBART (CC25): This original model is equipped with a 12-layer encoder and a 12-layer decoder, having been trained on a web corpus spanning 25 languages Liu et al. (2020). It boasts a hidden size of 1024 and a vocabulary size of 250,000. We continuously train it with English and Chinese literary texts (114B tokens) Wang et al. (2023a).
Besides, general-domain pretrained models listed in General MT Track are also allowed in this task: mBART, BERT, RoBERTa, sBERT, LaBSE.
Evaluation Methods
It is still an open question whether human and automatic evaluation metrics are complementary or mutually exclusive in measuring the document-level and literary translation quality. Thus, we report both automatic and human evaluation methods, and officially rank the systems based on the overall human judgments.
We use widely-used sentence- and document-level evaluation metrics: 1) sentence-level: we employ sacreBLEU Post (2018), chrF Popović (2015), TER Snover et al. (2006) and pretraining-based COMET Rei et al. (2020); 2) document-level: we mainly use document-level sacreBLEU (d-BLEU) Liu et al. (2020), which is computed by matching n-grams in the whole document. For d-BLEU, We combine all sentences in each document as one line and then conduct sacreBLEU metric. Note that all evaluations are case-sensitive. We employ sacrebleuhttps://github.com/mjpost/sacrebleu with signature: nrefs:2|case:mixed|eff:no|tok:13a|smooth:exp |version:2.3.1. to calculate sacreBLEU, chrF, TER and d-BLEU with sacrebleu using two references. The command is: cat output | python -m sacrebleu reference*. We employ unbabel-comethttps://github.com/mjpost/sacrebleu. to calculate COMET score using Reference 1. The command is: comet-score -s input -t output -r reference1 (default model).
2 Human Evaluation
The human evaluation was performed by professional translators using an adaptation of the multidimensional quality metrics (MQM) framework Lommel et al. (2014). For example, we consider the preservation of literary style and the overall coherence and cohesiveness of the translated texts. As shown in Table 7, we put forth an industry-endorsed criteria to guide human evaluation process. The main error types are:
Accuracy (Acc.): The target text does not accurately reflect the source text, allowing for any differences authorized by specifications.
Fluency (Flu.): Issues related to the form or content of a text, irrespective as to whether it is a translation or not.
Style (Sty.): The text has stylistic problems.
Terminology (Ter.): A term (domain-specific word) is translated with a term other than the one expected for the domain or otherwise specified.
Locale Convention (Loc.): The text does not adhere to locale-specific mechanical conventions and violates requirements for the presentation of content in the target locale.
Others (Oth.): Other issues such as the signs of MT, gender bias and source errors.
MQM utilizes a scorecard format to quantify the quality assessment results. Evaluators assign numerical values to identified translation errors based on error types, severity, etc., making the assessment results more intuitive. The overall quality score is calculated based on per-word translation accuracy:
where where we set four error severity levels: Neutral (Neu.), Minor (Min.), Major (Maj.), Critical (Cri.) with 0/5/10/25 severity penalty. denotes the number of errors. The “Total Word Count” is calculated based on source input (Chinese word). Considering our task is centered on Zh-to-En translation, we engaged four evaluators who are native English speakers and also fluent in Chinese.
Task Description
The shared task will be the translation of literary texts between ChineseEnglish. Participants will be provided with two types of training datasets: (1) discourse-level GuoFeng Webnovel Corpus; (2) General MT Track Parallel Training Data. Additionally, they are provided two types pretrained models: (1) in-domain pretrained models, including In-domain RoBERTa (base) and In-domain mBART (CC25). (2) other general-domain pretrained models listed in General MT Track. Note that basic linguistic tools are allowed in the constrained condition as well as pretrained language models released before February 2023.
In the final testing stage, participants use their systems to translate an official testing set. The translation quality is measured by a manual evaluation and automatic evaluation metrics. All systems will be ranked by human judgement according to our professional guidelines and translators. Participants can submit either constrained (i.e. only use the training data specified above) or unconstrained (i.e. it allows the participation with a system trained without any limitations) systems with flags, and we will distinguish their submissions.
Goals
Encourage research in machine translation for literary texts.
Provide a platform for researchers to evaluate and compare the performance of different machine translation systems on a common dataset.
Advance the state of the art in machine translation for literary texts.
Submission and Format
Submissions will be done by sending us an email to our official email. Each team can submit at most 3 MT outputs per language pair direction, one primary and up to two contrastive. The requirements of submission format are (1) Keep 12 output files that are identical to the testing input files. (2) In the output files, ensure that each line is aligned with the corresponding input line.
Participants’ and Baseline Systems
Here we briefly introduce each participant’s systems and refer the reader to the participant’s reports for further details. Table 1 shows the summary of systems and participant teams.
The team from University of Southern California, Information Sciences Institute introduce three translation systems. The Primary System is built on a paragraph-level transformer, trained on a paragraph-aligned corpus (with a source side cap of 256 characters), executing translations at the paragraph level. The Contrastive System 1 deploys a sentence-level transformer, capitalizing on the sentence alignment data available in the datasets. The Contrastive System 2 adopts a paragraph-level Mega model Ma et al. (2022). The Mega model proposed a single-head gated attention mechanism equipped with an exponential moving average, which achieves comparable performance compared to Transformers having with fewer parameters. In pre-processing, the team opted for Byte-Pair Encoding (BPE) for tokenization. And they employed Jaccard similarity for sentence alignment during the post-processing phase.
2 MAKE-NMT-VIZ (constrained)
The team from Université Grenoble Alpes introduced three translation systems. The Primary System finetune the mBART (CC50) model using Train, Valid1, Test1 of the GuoFeng Corpus, adopting settings similar to those described by Lee et al. (2022). Specifically, they finetune models for 3 epochs, utilizing the GELU activation function, a learning rate of 0.05, a dropout rate of 0.1, and a batch size of 16. For decoding, a beam search of size 5 was employed. The Contrastive System 1 is implemented upon a finetuned concatenation transformer Lupo et al. (2023) with two training steps: (1) a sentence-level transformer is trained for 10 epochs using General, Valid1, Test1 datasets; (2) a document-level transformer is finetuned using pseudo-document data (3-sentence concatenation) from Train, Valid2, Test2 data for 4 epochs. They use ReLU as an activation function, along with an inverse square root learning rate, a dropout rate of 0.1, and a batch size of 64. For decoding, a beam search of size 4 was employed. The Contrastive System 2 is a sentence-level transformer model trained for 10 epochs using General, Valid1, Test1 datasets. The training adopted an inverse square root scheduled learning rate, a dropout rate of 0.1, and a batch size of 64. Decoding was done using a beam search of size 4.
3 TJUNLP (constrained)
The team from Tianjin University introduced a Primary System based on a sentence-level Transformer model. The training consists of two phases: initially, it undergoes 100k steps on a dense model, followed by a 50k step fine-tuning on mixture of experts (MOE). They adopt the Polynomial Decay as their learning rate scheduling strategy, with a learning rate set at 2e-4, a dropout rate of 0.1, and a batch size encompassing 4096 tokens. For decoding, a beam search of size 5 was employed. For pre-processing, the team opted for SentencePiece Model (SPM) for tokenization.
4 NTU (unconstrained)
The Nantong University team introduce a Primary System. It is based on a pretrained MT model, Opus-MT,https://huggingface.co/Helsinki-NLP/opus-mt-zh-en., which is trained on OPUS dataset Tiedemann and Thottingal (2020). The model is finetuned on one NVIDIA Tesla A100 80 GB where the learning rate is 5e-5, batch size is 64, max length is 512 and the epoch number is 10.
5 DLUT (unconstrained)
The team form Dalian University of Technology introduce a Primary System based on GPT-3.5-turbo Brown et al. (2020). They mainly propose prompt engineering, data filtering, and document segmentation to activate the capabilities of LLMs for discourse-level translation Zhao et al. (2023).
6 HITer-WMT (unconstrained)
The team form Harbin Institute of Technology (Harbin) introduce two translation systems. The Primary System centers on instruction fine-tuning, executed through the Llama-7b model within the Parrot framework Jiao et al. (2023).https://github.com/wxjiao/ParroT. Specifically, they build an instruction dataset from two comprehensive chapters of our existing training corpus according to methodologies in Peng et al. (2023). This dataset was fine-tuned using Llama-7b over 3 epochs with a learning rate of 2e-5. The Contrastive System utilizes the GuoFeng mBART Model provided by the shared task. This model was trained over 10 epochs at a learning rate of 1e-4, with gradient clipping applied to stabilize training.
We report the automatic evaluation scores of all submissions in Table 5.6. The evaluation metrics includes 1) sentence-level BLEU, chrF, COMET, TER; and 2) document-level d-BLEU. To calculate d-BLEU, we first concatenate all continuous sentences in one book as on line, and then employ sacreBLEU to obtain scorers. To compute d-BLEU, we merge all the consecutive sentences from a single book into one continuous line, and then utilize the sacreBLEU to generate the scores.
Among constrained Primary systems, the MAKE-NMT-VIZ system shows impressive performance and achieves the best in terms of all metrics. Similarly, the HW-TSC⋆ Primary system achieves the best in constrained settings. As introduced in Section 5, MAKE-NMT-VIZ mainly finetune the mBART pretrained model while HW-TSC⋆ train a doc2doc Transformer model using a number of data augmentation methods.
In the majority of teams, the primary system exhibits superior performance compared to the corresponding contrastive system. The exceptions to this trend are noted in the cases of HITer-WMT⋆ and HW-TSC⋆, where this pattern does not hold. Among the baseline systems, Google Translate, a commercial translation service, outperforms both commercial and open-source LLMs (GPT-4 API and Llama-MT) in terms of d-BLEU scores. Interestingly, both the top-1 ranked Primary constrained and the top-2 ranked unconstrained systems surpass the performance of the commercial MT system.
2 Human Evaluation
Table 3 presents the results of the human evaluation and system rank for the Primary submissions. We enlisted four human annotators to evaluate 5 documents, comprising a total of 2,194 words sourced from distinct books within the final testset for each translation system.
As seen, the MAKE-NMT-VIZ system outperforms the other three constrained systems, while DLUT⋆ ranks first among the four unconstrained systems. This is not fully consistent with the automatic evaluation results in Table 5.6. Moreover, the top-2 unconstrained systems outperform the best constrained system, highlighting the benefits of external knowledge. This observation is consistent with that of automatic evaluation.
Among the baseline systems, the LLM system performs the best, whereas the MT system shows the poorest performance, diverging from the observations of automatic evaluation. Interestingly, the literary MT-enhanced models perform comparable with some systems such as MaxLab and Google Translate.
3 Analysis
We engaged four annotators to independently review an identical document (i.e. 601 words) selected from the testset. Table 4 outlines the individual scores given by each annotator and the corresponding average scores. The findings illustrate that (1) while there is variance in the exact scores assigned by different annotators, their scoring trends align; (2) the results on this sample may diverge from those obtained from a larger dataset, highlighting the necessity of human evaluation on a larger scale.
In our effort to understand the consistency among the human evaluators, we conducted a Pearson correlation analysis on their scoring patterns. Table 5 illustrates the pairwise Pearson correlation coefficients for the scores given by each annotator. The results indicate a high degree of agreement among the annotators. For example, Annotator 2 demonstrated a very high correlation with Annotator 3 () and Annotator 4 (). Besides, the Average Scores also reveal strong evaluator consensus on translation quality. This consistency underscores the reliability of the evaluators’ judgments across the assessed translations.
Error Type
We further analyze the error distribution in human-annotated results. Figure 4 classifies and counts the errors identified in the evaluated documents by their severity. This visualization allows for a direct comparison of the error profiles of each system, highlighting their strengths and weaknesses in different aspects of translation quality.
In the baseline systems analysis, GPT-4⋆ registers a higher frequency of Minor errors, particularly in Fluency and Style, indicating areas where refinement could enhance the translation’s naturalness and adherence to stylistic norms. Llama-MT⋆, by contrast, has a pronounced incidence of Major and Critical errors in Accuracy and Terminology, raising concerns about the fidelity and technical precision of its translations. Google⋆ stands out with its Fluency errors, suggesting potential issues in maintaining a coherent and natural flow compared to the language models.
Regarding the constrained systems, MAKE-NMT-VIZ displays an even spread of errors, with relatively fewer instances in each category, which points to a well-rounded performance in capturing nuances across various aspects of translation. Both MaxLab and TJUNLP exhibit an increased number of Accuracy and Fluency errors, suggesting challenges in delivering translations that are not only faithful to the source material but also exhibit a seamless and natural flow in the target language.
The unconstrained systems, particularly HW-TSC⋆ and DLUT⋆, show a notable reduction in errors related to Accuracy and Fluency when compared to their constrained counterparts. This trend suggests that the lack of constraints may afford these systems more flexibility, resulting in translations that are more accurate and fluid. However, the overall error distribution across different systems highlights the complex trade-offs and challenges inherent in machine translation, underscoring the need for continued innovation and optimization in the field. In the future, we will also consider hallucination errors Zhang et al. (2023).
We believe that the WMT2023 Shared Task on discourse-level literary translation will be a valuable contribution to the field of machine translation and will encourage further research in this area. We discuss the potential limitations of this edition of the shared task as follows:
Language Pair. This year, we only focus on ChineseEnglish direction. However, we have a long-term plan to continuously organize this task, and will extend the copyrighted dataset into Chinese-Russian and Chinese-German language pairs next year.
Literary Genre. This year, we mainly used the Web Fiction Corpus which is only one type of literary text. We use Web Fiction for two reasons: (1) its literariness is less complicated than others (e.g. poetry, masterpiece); (2) such bilingual data are numerous and continuously increased. We will consider to extend more literary genres such as poetric translation in the next year.
Discourse Benchmark. We have accumulated some discourse- and context-aware benchmarks Xu et al. (2022, 2023); Wang et al. (2023a). These benchmarks are pivotal for assessing the proficiency of LLMs in handling complex language structures and contextual nuances. As participation of LLM-based systems in our shared tasks increases, we anticipate integrating these benchmarks more comprehensively into our future evaluations to better measure and understand the evolution of LLM capabilities in linguistic context and discourse comprehension.
Machine translation of web novels not only holds research value but also offers practical application prospects Huang et al. (2021); Lyu et al. (2023). This shared task serves to spur competitive innovation and fosters the advancement of sophisticated machine translation systems capable of navigating the intricate nuances of literary works. Anticipating the future, our objective is to broaden the engagement in the forthcoming shared task, inviting an extensive range of collaborators from industry and academia alike to contribute their unique insights and expertise.
We would like to thank the WMT2023 organizers for providing us the opportunity to explore this new task. We also express our gratitude to the experts on the Shared Task Committee for their efforts in organization, evaluation, and advisory roles:
Longyue Wang, Zhaopeng Tu, Dian Yu, Chenyang Lyu, Shuming Shi (Tencent AI Lab)
Yan Gu, Yufeng Ma, Weiyu Chen (China Literature Ltd.)
Siyou Liu, Yulin Yuan (University of Macau)
Liting Zhou, Andy Way (Dublin City University)