Document-Level Machine Translation with Large Language Models

Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, Zhaopeng Tu

Introduction

In the past several years, machine translation (MT) has seen significant advancements with the introduction of pre-trained models such as BERT Devlin et al. (2018), GPT-2 Radford et al. (2019), and T5 Raffel et al. (2020). These models have demonstrated impressive performance on MT Zhu et al. (2020); Guo et al. (2020); Xue et al. (2021). However, most of the existing work has focused on sentence-level translation, which can result in translations that lack coherence and context. Recent years have seen a growing interest in document-level translation, which is a crucial task that involves translating entire documents Wang et al. (2017); Bawden et al. (2018); Wang (2019); Zhang et al. (2022) while modelling specific discourse phenomena Wang et al. (2016); Voita et al. (2018); Wang et al. (2018a, b, 2019); Voita et al. (2019b); Wang et al. (2023b). The most popular large language model (LLM) – ChatGPThttps://chat.openai.com. All corresponding results were obtained from GPT-3.5 and GPT-4 in March 2023. The reproducibility is discussed in Section Limitation. shows the ability of maintaining long-term coherence and consistency in a conversation by conditioning on previous conversational turns. Additionally, the model is trained on a large dialogue dataset, which allows it to learn the patterns and conventions of human communication, further improving its ability to document-level understanding and generation (as shown in Figure 1).

In this paper, we are particularly interested in how LLMs such as ChatGPT perform for modeling document-level text, encompassing discourse phenomena such as entity consistency, referential expressions, and coherence. Taking document-level MT as a testbed, we conduct an empirical study from three in-depth perspectives:

Effects of Context-Aware Prompts: ChatGPT needs a prompt as guidance to trigger its translation ability. Thus, we enable prompts to guide ChatGPT to consider document-level contexts as long as possible. Jiao et al. (2023) has found that the candidate prompts generally work well and show minor performance differences on sentence-level translation. In this work, we further investigate the effects of prompts on the translation quality and specific discourse phenomena.

Comparison of Advanced Translation Models: While ChatGPT has demonstrated remarkable abilities in long-text NLP tasks, we are specifically interested in how it performs on document-level translation. Consequently, we conduct a systematic comparison of commercial MT products and advanced document-level approaches, utilizing both automatic and human evaluations to assess their discourse awareness.

Analysis of Discourse Modelling Abilities: A more challenging question is the extent to which ChatGPT capture and utilize discourse knowledge. To answer this question, we introduce a probing method through contrastive testing and explanation. In addition, the impact of various training techniques on the ability of LLMs to model discourse has not been thoroughly investigated. We compare variant models of ChatGPT that incorporate techniques such as code pretraining, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF). However, this is not a strict comparison because there are other confounding variables employed during the evolution of ChatGPT. In general, we hope to pose this open question that stimulates reflection and sparks further investigation.

We conduct experiments on a variety of document-level translation benchmarks, covering three language pairs (i.e. Chinese⇒\RightarrowEnglish, English⇒\RightarrowGerman and English⇒\RightarrowRussian) and seven domains (i.e. news, social, fiction, Q&A, TED, Europarl, and subtitle). We adopt a comprehensive set of evaluation methods to measure the performance of the models on document-level translation, including general-purpose metrics, discourse-specific metrics, and human evaluation. The main contributions are:

Our empirical study shows the superior capabilities of LLMs over advanced MT systems and methods on document-level translation, indicating their potential to form a new paradigm.

We establish a benchmark with a probing method to thoroughly assess the document-level translation quality and the ability of learning discourse knowledge, which will be made available for future research.

To facilitate future research on document MT, we publicly release the instruction-based benchmark, system outputs as well as human annotations.

Table 1 shows statistics of document-level datasets used in our experiments. About Group #1, we utilized the latest datasets, mZPRT Wang et al. (2022) and WMT2022 Kocmi et al. (2022), for evaluation to ensure that the testing data had not been used in commercial systems (e.g. Google Translate and ChatGPT) before. As seen, this covers four domains (i.e. news, social media, web fiction, and Q&A forum) in Chinese⇒\RightarrowEnglish. Regarding Group #2, we utilized four widely-used benchmarks to compare established document-level methods with GPT-like applications. This covers three domains (i.e. TED talks, news commentary, European Parliament) in Chinese⇒\RightarrowEnglish and English⇒\RightarrowGerman. In Group #3, we employed an English⇒\RightarrowRussian contrastive testset Voita et al. (2019b) that specifically targets discourse phenomena, such as deixis, ellipsis, and lexical cohesion. We use this dataset to further exploit models’ capacity for modeling discourse knowledge.

As seen, we also report the average length of a document (|W|/|D|), which can be considered a measure of the complexity of discourse modeling. As the length of the document increases, it becomes more challenging to model accurate cohesive devices and discourse structure. From this perspective, the mZPRT Fiction and IWSLT TED datasets pose a greater challenge compared to others.

We conducted ChatGPT in March 28∼\sim31 2023 with the official notice “the training data is up to September 2021”.https://platform.openai.com/docs/models/gpt-4. Different from previous work that leaned on dated or single-type datasets to assess ChatGPT’s capabilities Jiao et al. (2023); Lu et al. (2023), we carefully chosen both lastst, public and diverse testsets to mitigate the risks associated with data contamination. Taking Probing Discourse Knowledge (Section 5.1) for example, while the contrastive testset used for the prediction task originated in 2019, the evaluation on the explanation task remains devoid of public references. Our conclusions are comprehensively made by considering both prediction and explanation results, balancing out any potential data contamination concerns. Despite precautions, there remains a risk of data contamination, given that publicly available datasets are easily incorporated into LLM training (e.g. pre-training, SFT, or RLHF). A better way is to consistently integrate and employ the latest datasets when evaluating LLMs.

2 Evaluation Method

We evaluate different approaches and systems using classical sentence- and document-level evaluation metrics. About sentence-level metrics, we employ the commonly-used sacreBLEU Post (2018) and TER Snover et al. (2006). Additionally, we utilize COMET Rei et al. (2020), which leverages pretrained language models to achieve high correlations with human quality judgments. About document-level metrics, we report document-level sacreBLEU (d-BLEU) Liu et al. (2020), which is computed by matching n-grams in the whole document. Note that all evaluations are case-sensitive. To facilitate sentence-level evaluation of document outputs, we implemented automatic sentence splitting/alignmenthttps://github.com/rsennrich/Bleualign. on the output and then manually corrected errors.

Discourse Awareness

To target specific discourse phenomena, we utilized two targeted metrics, namely, cTT and aZPT, which respectively evaluate the consistency of terminology translation and accuracy of zero pronoun (ZP) translation. Regarding cTT, one repeated terminology should keep the same translation throughout the whole document Xiao et al. (2011). We adopt a lexical translation consistency metric Lyu et al. (2021):

Regarding ZP, it is a discourse phenomenon that appears frequently in pronoun-dropping (pro-drop) languages such as Chinese and Japanese. Recovering ZPs in a target language (non-pro-drop) needs an understanding of the discourse structure. We used the aZPT score to measure the accuracy of ZP translation Wang et al. (2022):

where ZP\bf ZP is the list of zero pronouns in the source sentences, tzt_{z} is the generated translation for the zero pronoun zz, and A(tz∣z)A({t_{z}|z}) is a binary scorer to judge whether tzt_{z} is the correct translation of zz.

Human Evaluation

To thoroughly validate our conclusions, we also conduct a human evaluation (see Section Limitation.). We establish two sets of evaluation criteria: 1) general quality, covering aspects such as fluency and adequacy; 2) discourse-aware quality, including factors such as consistency, word choice, and anaphora. The detailed scoring criteria are listed in Appendix§ A.2. Accordingly, each output will be assigned two distinct scores (0∼\sim5). For each domain subset, we assessed 100 instances, with each instance containing outputs from 5 different systems. This amounted to an evaluation of roughly 70K words in total. The scores were assigned to each window of neighboring sentences, taking into account the context provided by the entire document. Our intent was for evaluators to consider discourse properties beyond single sentences, while also avoiding the difficult task of evaluating an entire document. We employed two professional evaluators for our study. The payment and background is detailed in Section Ethical Considerations and Appendix§ A.2, respectively. Besides, our annotators were given practice items, and the annotations reaches 0.86 Cohen’s kappa scores McHugh (2012), demonstrating that the annotators work efficiently and consistently under this guideline.

Existing document NMT methods can be mainly classified into two categories: multi-sentence Wang et al. (2017); Voita et al. (2018); Tu et al. (2018) and whole-document Macé and Servan (2019); Bao et al. (2021) approaches. ChatGPT is capable of not only handling long text in a single conversational turn but also recalling the entire context in the chat box. Accordingly, we design prompts to trigger the document-level translation ability of ChatGPT.

The prompt engineering is necessary to ensure ChatGPT’s robust ability to interpret instructions and to model long-term dependencies. Our research confirms the neutrality and representativeness of various prompts, allowing other researchers to utilize them with confidence, unburdened by concerns of unintended biases.

2 Comparison of Different Prompts

We query ChatGPT itself for advice and obtain a number of candidate prompts, and then refine them into three document-level prompts as shown in Table 3. We utilize P1 to translate a document sentence by sentence, with each sentence placed in a single conversational turn and the entire document contained within one chat box. This mainly takes advantage of ChatGPT’s long-term modeling ability in the chat box. P2 and P3 combine multiple continuous sentences and translate them in one conversational turn until the entire document is finished. This aims to maximize the length of document as input. The only difference is whether or not the sentential boundary tag “[]” is inserted into each sentence.

We compare the three candidate prompts on the Zh⇒\RightarrowEn translation task using two testsets, WMT2022 News and mZPRT Fiction. Table 2 shows the translation quality in terms of a variety of automatic evaluation metrics. In general, ChatGPT reliably performs well with three candidate prompts, showing only minor variations in performance. This aligns with prior findings in sentence-level translation with ChatGPT Jiao et al. (2023). Out of the three prompts, the prompt involved multi-turn contexts without sentence boundaries (P3) achieves the best scores in most evaluation metrics, except for COMET. Regarding discourse phenomena, P3 outperforms other candidates with better consistency of terminology translation and higher accuracy of ZP translation. Upon examining the output samples, we noticed that ChatGPT may sometimes forget the sentential boundary tag of P2 and combine all sentences together. Takeaway: (1) Despite translating a document sentence by sentence, ChatGPT’s ability to model long-term dependencies already exists within the chat box. (2) Increasing document length as a input can further enhance translation quality and discourse awareness. (3) ChatGPT tends to translate a document without adhering to strict sentential boundaries, mirroring a natural approach adopted by humans during document translation, which doesn’t necessitate sentence-to-sentence translation.

In this section, we compare various systems and methods for the document-level translation task. In the following experiments, we use the P3 prompt for ChatGPT and the same document-level window size for MT models as the default setting.

Commercial systems are known for their high accuracy and efficiency in translation, making them a strong contender for any machine translation evaluation. By comparing with commercial systems, we can gauge ChatGPT’s performance relative to the best available MT technologies. We compare GPT-3.5/GPT-4 with three commercial translation products, including Google Translate,https://translate.google.com. DeepL Translate,https://www.deepl.com. and Tencent TranSmart Huang et al. (2021).https://transmart.qq.com. We employ both automatic (d-BLEU) and human evaluation (general/discourse-aware quality) as detailed in Section 2.2.

Table 4 shows the results. When evaluated using d-BLEU, commercial MT systems generally outperform LLM-based systems, except for the Q&A domain, which involves informal spoken language. While the difference in performance is not significant in the news domain (e.g. the gap between DeepL and GPT-4 is only 0.6 points), it is considerable in the social media and web fiction domains (i.e. the gaps are 3.3 and 1.9 points). A surprising finding is that GPT-4 and GPT-3.5 perform significantly better than MT systems in terms of human evaluation. The potential reasons may be: (1) d-BLEU only measures the similarity of the n-grams between the MT output and the reference translations. However, human takes into account additional factors such as coherence, fluency, and naturalness of the translation, which may not necessarily correlate with d-BLEU scores. (2) ChatGPT and MT systems may have different strengths and weaknesses. For example, ChatGPT may be better at modeling long-term dependencies and capturing discourse-level information, which could result in higher human evaluation. On the other hand, MT systems may perform better in terms of word-level accuracy, which is reflected in d-BLEU. Note that, our findings is distinguished from Neubig and He (2023). Focusing on long-text translation, we compare ChatGPT with MT systems, and underscore ChatGPT’s enhanced capacity to model long-term dependencies in comparison to MT systems. On the other hand, Neubig and He (2023) investigate the varying performances of GPT models based on sentence length. They found that GPT models perform better on shorter sentences while worse on longer ones. Karpinska and Iyyer (2023) recently highlighted that GPT-3.5 has the capability to utilize document-level context effectively for literary translation, yet it is not free from critical errors. While the Fiction testset in our work is categorized under literary, we did not find obvious omission errors in the output. A more detailed comparison is earmarked for future exploration. Karpinska and Iyyer (2023) recently pointed that GPT-3.5 can effectively leverage document-level context for literary translation, but critical errors persist. Although the Fiction subset belongs to literary, we did not find omission errors in the output and we leave this fine-grained comparison for future work. Takeaway: (1) There is a certain degree of discrepancy discrepancy between human and automatic evaluation, which potentially provide complementary reference points when measuring the quality of document-level translations; (2) This discrepancy underlines the complexity inherent in accurately evaluating the capabilities of such systems. We further explore evaluation methods in Section 5.1.

2 ChatGPT vs. Document NMT Methods

Document NMT methods are specifically designed to handle part or entire documents, making them a relevant point of comparison for evaluating ChatGPT’s ability to model long-term dependencies and discourse phenomena. We compare with five advanced document-level NMT models:

MCN Zheng et al. (2020): A multi-channel network that integrates a hierarchical encoder and a parallel decoder, which leverages the document structure and semantics for translation.

G-Trans Bao et al. (2021): A graph-based transformer that incorporates document-level discourse structure as a directed acyclic graph, enhancing the representation of the context.

Sent2Sent: A superior sentence-level baseline that employs a transformer architecture to translate each sentence independently and then merges the translations into a document-level output.

MR-Doc2Doc and MR-Doc2Sent: Sun et al. (2022) explore to resolve document translation with the end-to-end, namely document-to-document (Doc2Doc) pattern, and utilize Multi-resolutional Training, which combines documents with shorter segments like sentences or paragraphs to improve translation quality (denoted as MR-Doc2Doc). Additionally, they reproduce the document-to-sentence baseline (MR-Doc2Sent) that introduces extra model modules to capture contextual information.

To enable a fair comparison with previous work, we use four widely used document-level translation benchmarks: TED (ZH-EN and EN-DE), News (EN-DE), and Europarl (EN-DE). We adopt tokenized case-insensitive BLEU and d-BLEU as the evaluation metrics. As MR-Doc2Doc and ChatGPT generate document-level translations that are difficult to separate into individual sentences, we only report d-BLEU scores for these models.

Table 3.2 lists the results. The MR-Doc2Doc with extra model pre-training achieves the best document-level performance among previous models. Thanks to the document-level LM pre-training, ChatGPT easily outperforms MR-Doc2Doc⋆ on TED (EN-DE) and News (EN-DE) datasets, obtaining similar performance on TED (ZH-EN) dataset. Surprisingly, ChatGPT performs poorly on the Europarl (EN-DE) dataset, even worse than Sent2Sent. We suspect this phenomenon may be caused by the domain distribution bias of the training data. Moreover, we find that ChatGPT is unstable, and its translation results sometimes exhibit omissions and obvious copying behaviors. Note that, the commonly-used datasets were created between 2012 and 2017, a time frame that raises the possibility of these datasets being incorporated into the training data of newer language models. Takeaway: (1) ChatGPT has exhibited superior performance and may become a new promising paradigm for document-level NMT; (2) It is still debatable whether these benchmarks can be considered as appropriate measures for evaluating document-level translation methods. We advocate for greater transparency from model developers regarding their training datasets. Additionally, this highlights the importance of designing innovative evaluation techniques that can reliably assess model capabilities while sidestepping concerns related to data contamination.

We analyze the ability of LLMs to capture discourse knowledge from two perspectives: (1) probing the discourse knowledge encoded in LLMs, and (2) examining the impact of different training techniques on discourse modeling.

In order to verify whether LLMs truly learn to utilize the context to resolve discourse inconsistencies, we adopt the contrastive test sets proposed by Voita et al. (2019b). This dataset includes deixis, lexicon consistency, ellipsis (inflection), and ellipsis (verb phrase) for evaluating discourse phenomena in English-Russian translations. Each instance has a positive translation and a few negative ones that differ by only one specific word. The goal is to determine if a model is more likely to generate a correct translation than incorrect variations. In this experiment, we compare GPT-3.5/GPT-4 with advanced methods, such as Sent2Sent, MR-Doc2Doc, CADec Voita et al. (2019b) and DocRepair Voita et al. (2019a), where CADec and DocRepair introduce context-aware post-editing modules to refine the sentence-level translations. For these baselines, we adopt force decoding to generate scores for all translation candidates in each instance. If the score of the positive translation is the highest, then this instance is counted as correct. For ChatGPT, we query them with the prompt P4 in Table 6 to obtain responses and correspondent explanations for each instance. Then some heuristic rules and manual verification are used to calculate final performance.

As shown in Table 7, GPT-3.5 performs worse than DocRepair (discouse-enhanced method) across all discourse phenomena, with particularly significant gaps present in deixis and lexical consistency tests. These results show that it is difficult to handle deixis and lexical consistency phenomena with large-scale document-level pre-training. GPT-4 exhibits significant improvements in these areas, but it still lags behind DocRepair in deixis, lexical consistency, and ellipsis (inflection) phenomena. Takeaway: (1) GPT-3.5 demonstrates lower accuracy in contrastive prediction compared to conventional translation models, whereas GPT-4 exhibits significant improvement. (2) As there is no detailed technical report available for GPT-4, we argue that its significant improvements are likely due to the use of supervised data and RLHF. We further explore this in Section 5.2.

Evaluation on Explanation

We conduct human evaluations to assess the quality of LLM-generated explanations. This provides an additional way to explore the discourse knowledge contained within LLMs. As illustrated in Table 8, we randomly select 100 examples for each contrastive test set and request native speakers to evaluate whether the models’ responses contain the correct prediction and explanation, respectively. Then the Phi coefficient (rϕr_{\phi}) is further calculated to better measure the correlation between two binary variables (i.e., prediction and explanation). We can observe that the accuracy of explanation is often not reflective of the accuracy of prediction, indicating a mismatch in utilizing discourse knowledge for prediction and explanation. In addition, GPT-3.5 is not good at explaining the reason for selecting the correct translation, while GPT-4 exhibits high performance in this aspect and brings better accuracy of prediction. Takeaway: (1) GPT-4 demonstrates a strong ability to explain discourse knowledge. (2) Despite GPT-4’s superior performance in prediction and explanation, the correlation between prediction and explanation does not appear to be significantly improved compared to GPT-3.5.

2 Potential Impacts of Training Techniques

LLMs have become the foundation of natural language processing research Brown et al. (2020), with recent advancements such as learning from source code Chen et al. (2021) and RLHF showing promise in improving LLM performance Ouyang et al. (2022). To investigate the potential impacts of these approaches on discourse modelling, we conduct experiments on Chinese⇒\RightarrowEnglish Fiction and English⇒\RightarrowRussian datasets using different variants of LLMs trained with distinct techniques (detailed in §A.3). Accordingly, we use P3 and P4 prompts.

Table 9 shows the results. When SFT with high-quality demonstration examples, the translation performance can achieve 14.1 d-BLEU, which reaches to an acceptable level (InstructGPT +FeedME-1 vs. +SFT). Moreover, code pretraining can improve the document translation quality by 2 d-BLEU points and the discourse knowldge probing by 4/1.5 (CodexGPT +FeedME-2 vs. InstructGPT +FeedME-1). When further adding PPO method, it outperforms all other combination strategies on translation quality, discourse awareness and discourse knowledge probing (CodexGPT +FeedME-2 +PPO vs. others). This shows that RLHF strongly enhances LLM’s capability of translation. Lastly, GPT-3.5 and GPT-4 excel in d-BLEU and human evaluation scores as well as probing performance, demonstrating the importance of contextual information for complex, lengthy document translation. Takeaway: (1) Methods like code pretraining, SFT and RLHF appear to enhance the performance of document translation and discourse modeling; (2) However, it is quite challenging to explore the non-open source systems due to various confounding factors introduced during their development. Therefore, we advocate for organizations like OpenAI to provide greater transparency to aid researchers in such explorations.

We provide a comprehensive evaluation of LLMs (such as GPT-3.5 and GPT-4) for document-level machine translation. Our evaluation covers three main aspects: (1) the effects of discourse-aware prompts, (2) comparison of advanced translation models, and (3) analysis of discourse modelling abilities. With the release of the GPT-4 model, the discourse-aware performance has been significantly improved, making it a promising paradigm for document-level translation. Despite its prowess in generative tasks, it struggles with discerning subtle distinctions for ranking.

In our future work, we plan to explore more document-level evaluation method Castilho (2021); Jiang et al. (2023); Kocmi and Federmann (2023), latest long-text benchmarks Wang et al. (2023a); Thai et al. (2022); Wang et al. (2023c) and other MT scenarios Ghazvininejad et al. (2023); Guerreiro et al. (2023); Lyu et al. (2023). Furthermore, we intend to delve into a more detailed analysis and comparison in future work. For instance, we will employ appropriate significant test methods to account for multiple comparisons (e.g. non-parametric Kruskal-Wallis test, Bonferroni correction) and conduct a power analysis Card et al. (2020); Graham et al. (2020); Vilar et al. (2022); Hendy et al. (2023). About annotation consistency, we further apply the Krippendorff’s alpha coefficient and check the confidence interval Krippendorff (2011).

We list the main limitations of this work as follows:

Potential Inaccuracy of Conclusions. Our conclusions are derived from experiments conducted on a limited set of datasets, which may not guarantee accuracy or applicability across all contexts. These limitations might inadvertently introduce bias or overlook certain phenomena, potentially impacting the comprehensiveness of our findings. In response to this concern, we strive to use an extensive array of the most recent datasets, spanning various language pairs and domains. This broad coverage aims to encompass a wide range of linguistic and contextual variations, thereby enhancing the generalizability of our findings.

Model Updates in ChatGPT and Reproducibility. When the underlying model or its parameters of ChatGPT are updated or changed, the conclusions derived from prior evaluations may no longer be valid or entirely accurate. To mitigate this issue, this paper has tried utmost to ensure the reproducibility of our findings: (1) We release all system outputs accompanied by exact timestamps and change logs. This ensures that researchers can reliably reproduce and validate our results. (2) We evaluated all systems at two distinct points: in March and August 2023. While there were minor variations in the exact performance figures between these two evaluations, our overarching conclusions and core findings remained unchanged and consistent.

Criteria of Human Evaluation and Refinement. The design of the criteria still has room for improvement. For example, in the "Discourse Awareness" category, there is only a slight difference between Scores 5 and 4, but a more significant gap between Scores 3 and 2. Given the scarcity of benchmark standards on discourse evaluation from past research, we published the detailed scores for further analysis and highlight this area as an opportunity for refinement.

Annotation Process and Annotator Profile. A one-week trial annotation phase was conducted for bidding (five companies participated) on our enterprise-level annotation platform. The authors answered questions posted by the annotators of these companies and updated annotation guidelines accordingly. The Q&A history is recorded and updated in the formal annotation phase. After evaluating the trial annotations based on accuracy and consistency, we selected a professional language service company (large enterpriseIt is based on the information provided by https://www.kanzhun.com/firm.) headquartered in Beijing, China. Their annotators were experts in both source and target languages, with a background in translation and linguistics (detailed in TableA.2). To understand any potential biases, annotators were inquired about their familiarity with the specific translation models under evaluation and any affiliations with AI or translation companies. We ensured that none of the annotators had conflicts of interest, and their annotations were routinely cross-checked for consistency. In terms of compensation, annotators received an hourly wage of $37.4. This rate aligns closely with the mean hourly wages observed for U.S. interpreters/translators and foreign language teachers.It is based on the information provided by https://www.bls.gov/oes/current/oes273091.htm and https://www.bls.gov/oes/current/oes251124.htm.

Annotator Consent and IRB Review. Prior to the commencement of the study, all participating annotators gave their informed consent, confirming their understanding of the study’s objectives and the intended research use of their annotations. An exhaustive IRB review was undertaken and finalized before the project’s onset (IRB Protocol Number: IRB-2023-00067 and Approval Date: 01/12/2023). The protocol employed in this work was approved by the Tencent Institutional Review Board. We remain steadfast in our commitment to uphold and adhere to the institution’s highest ethical and professional benchmarks throughout our research endeavors.

Reproducibility Challenges and Mitigation Strategies. The evolving nature of closed commercial platforms indeed presents challenges to the reproducibility of research. As these platforms continue to upgrade and improve, previous versions of models may be retired or modified, which can make replication efforts problematic. To tackle this issue, we have made several additions to our paper: (1) Documentation of Specifics: We have included exact versions of models we evaluated, along with the precise date of evaluation and other pertinent details. This allows for a clear record of the conditions under which our study was conducted. (2) Release of System Outputs: We release all system outputs, which ensures that researchers can reliably reproduce and validate our results. (3) Advocacy for Archiving Historical Versions: We emphasize the importance for both the AI community and commercial organizations to maintain archives of previous model iterations. By doing this, researchers can readily access and evaluate past versions, ensuring continuity in analysis even as new model versions emerge.

We are grateful to the anonymous reviewers, area chairs and ethics committee for their insightful comments and suggestions which will serve to improve the paper considerably. Their insights are invaluable, not only in helping us present our findings more cautiously, but also in educating us beyond the scope of this single paper.

Appendix A Appendix

For automatic evaluations, we used the non-parametric one-tailed Wilcoxon signed-rank test Woolson (2007). About results in Table 4, the significance test contrasting GPT-3.5/GPT-4 with others yields a p-value of less than 0.05, indicating they do significantly boosts translation quality. For human evaluation, we employ Welch’s t-test Bl (1947). Table 10 shows the overall significance test by combining all datasets, and the results for each domain are also consistent.

Probing Task

We performed a Welch’s t-test Bl (1947) with unequal variances to verify the significance between GPT-3.5 and GPT-4 in Table 7 and 8. We find that the corresponding two-tailed p-value is smaller than 0.001, which indicates the significance between them.

A.2 Human Evaluation Guidelines

Table 12 presents the human evaluation criteria for document-level translation, with scores ranging from 0 to 5. A score of 5 indicates excellent overall translation quality, with no grammatical errors, accurate word choice, consistent key terms, and consistent context and tone throughout the passage. A score of 0 indicates poor overall translation quality, with more than half of the translation being mistranslated or missing, inconsistent key terms, and poor fluency and clarity. In between scores reflect varying degrees of translation quality, with factors such as fluency, accuracy, consistency of key terms, and context and tone consistency affecting the score.

Human Evaluators

We employed two professional evaluators for our study through Tencent’s designated supplier. Table A.2 detailed their background related to this task.