A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models
Haoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan Awadalla
Introduction
Generative (decoder-only) large language models (LLMs) such as GPT models (Brown et al., 2020; OpenAI, 2023), PaLM (Chowdhery et al., 2022), OPT (Zhang et al., 2022), BLOOM (Scao et al., 2022), LLaMA (Touvron et al., 2023a; b), and others have exhibited remarkable capabilities across various NLP tasks. However, for the translation task, only very large models such as GPT-3.5 and GPT-4 can rival the supervised encoder-decoder state-of-the-art (SoTA) models like NLLB (NLLB TEAM et al., 2022), while they still fall short in translation for low-resource languages (Hendy et al., 2023; Jiao et al., 2023). The discrepancy becomes more evident when comparing other LLMs with traditional translation models (Zhu et al., 2023a). For instance, the OPT-175B model trails behind the NLLB-1.3B model by an average of more than 15 BLEU (Papineni et al., 2002) points for languages within the Indo-European-Romance family. The gap is even larger in smaller LLMs; for example, XGLM (Lin et al., 2021), with a parameter size of 7B, lags behind the NLLB-1.3B by a substantial 30 BLEU points (Zhu et al., 2023a). Therefore, there is an urgent need to narrow this performance gap between LLMs and conventional SoTA models.
As exemplified by NLLB-1.3B, traditional machine translation models demonstrate proficiency in producing high-quality translations with a small number of parameters. By extension, smaller LLMs should similarly possess the capability to adeptly manage the translation task. Recent research has sought to enhance translation performance by commencing with smaller LLMs (Yang et al., 2023; Zeng et al., 2023; Chen et al., 2023; Zhu et al., 2023b; Li et al., 2023; Zhang et al., 2023b), especially 7B or 13B parameters. Nevertheless, the achieved improvements remain modest and limited. As depicted in Figure 1, contemporary studies such as Balyling (Zhang et al., 2023b) and BigTranslate (Yang et al., 2023), which use LLaMA as their backbone, exhibit a maximum increment of 3 to 4 BLEU or COMET in relation to the zero-shot performance of LLaMA on the WMT’22 test set (8 directions).All COMET scores in the paper is COMET-22 (Unbabel/wmt22-comet-da) (Rei et al., 2022). While these gains represent promising research direction for smaller LLMs in the translation task, a significant performance chasm persists when benchmarked against very large LLMs such as GPT-3.5-text-davinci-003 and SoTA translation models such as NLLB-54B. We posit that the modest translation gains observed in prior studies can be ascribed to an unsuitable training recipe.
We hypothesize that an efficacious training recipe ought to follow two stages: learning general multilingual linguistic knowledge and inducing (instructing) models toward translation generation. Consequently, we propose a two-stage fine-tuning approach and introduce the LLM developed through this strategy as Advanced Language Model-based trAnslator (ALMA). Specifically, given most LLMs are trained on English-dominant data, the first stage is fine-tuning non-English monolingual data to enhance the model’s proficiency in other languages involved in the translation task. Secondly, drawing inspiration from the recognized significance of data quality in other applications (Zhou et al., 2023; Maillard et al., 2023; Gunasekar et al., 2023), we fine-tune the model with a small amount of high-quality parallel data.
Our main contributions are summarized as follows:
Diminished Necessity of Parallel Data Traditional translation frameworks rely on large amounts of parallel data, which may lead to a false impression that such data is essential for the translation task with LLMs. Prior studies have fine-tuned LLMs with datasets containing over 300M parallel instances (Yang et al., 2023). However, our empirical evaluations suggest that this strategy may not be optimal, and even harm the translation capabilities of LLMs.
LLM Via A New Training Recipe: ALMA We introduce a novel two-stage fine-tuning method for translation with decoder-only LLMs. Leveraging LLaMA-2 as the base model, we attain an average improvement of more than 12 BLEU and COMET scores over its zero-shot performance across 10 translation directions from WMT’21 and WMT’22 test datasets. Notably, the performance surpasses all previous work and is even better than the NLLB-54B model and GPT-3.5-text-davinci-003.
Efficient Computational Cost Our ablation study reveals both stages are crucial factors for achieving large improvements. The most computationally intensive part is monolingual data fine-tuning, however, we show that only fine-tuning 1B monolingual tokens is sufficient to have comparable performance to NLLB-54B in 10 translation directions, which only requires around 18 hours to complete with 16 MI200 GPUs.
Preliminary
We consider a decoder-only transformer model parameterized by for machine translation. Let represent the source sentence and its corresponding target sentence. We utilize a fixed prompt template, denoted as , to guide the model in generating translation. The log-likelihood loss of the parallel sentence (, ) with regard to the model parameters can be formulated as follows:
where is length of the target sentence, and is the -th target token. The loss is a standard causal language modeling (CLM) loss, which predicts the next token based on prior tokens. We use the same sentence-level translation prompt template suggested by Hendy et al. (2023), and illustrate the prompt and the model input/target in Figure 2. Note that we do not compute the loss for the prompt template and source sentence during training (Zhang et al., 2023a). In Appendix A, we show that CLM is more suitable for the translation task compared with other modeling methods, such as prefix language modeling (Wang et al., 2022) and mixture-of-denoisers (Tay et al., 2022a).
2 A Backbone LLM for Translation
We seek a robust LLM to serve as our foundational model. With the recent emergence of numerous LLMs, we prioritize evaluating the zero-shot translation performance of these models before delving into optimal training recipes. As most of these models provide a 7B version, our comparative analysis centers on this magnitude: OPT-7B (Zhang et al., 2022), Falcon-7B (Almazrouei et al., 2023), BLOOM-7B (Scao et al., 2022), MPT-7B (MosaicML, 2023), LLaMA-1-7B (Touvron et al., 2023a), and LLaMA-2-7B (Touvron et al., 2023b). We additionally present results from GPT-3.5-text-davinci-003 (hereinafter referred to as GPT-3.5-D) and GPT-3.5-turbo-0301 (hereinafter referred to as GPT-3.5-T) to show the performance gap. https://beta.openai.com/docs/model-index-for-researchers
Zero-Shot Evaluation We conduct zero-shot evaluations on 5 English-centric language pairs, considering both from English and to English directions: German (de), Czech (cs), Icelandic (is), Chinese (zh) and Russian (ru), where Icelandic test data is from WMT’21 and the others are from WMT’22. We choose these test dataset because they are the recent and less likely to overlap the training data used by LLMs, and importantly, they have high-quality data to avoid problems of “translationese” (Zhang & Toral, 2019). The beam size is 5. We report sacreBLEU (zh tokenizer for Chinese and 13a for the others) (Post, 2018). We also report COMET (Unbabel/wmt22-comet-da) (Rei et al., 2022) because BLEU only reflects the degree of lexical match. In this paper, We rely more on COMET than BLEU due to its better alignment with human evaluations (Freitag et al., 2022).According to Freitag et al. (2022), COMET holds the 2-nd position in alignment with human ratings, whereas BLEU is situated at the 19-th spot among 20 metrics
LLM Translation Performance The overall results for the LLMs are presented in Figure 3, with scores averaged across five languages for translations to and from English. Among the 7B LLMs, LLaMA-2-7B exhibits superior performance translating into English, while MPT-7B leads in translations out of English, as measured by BLEU. However, when evaluated with COMET, LLaMA-2-7B wins in both directions. We show the numeric results in Appendix B. Consequently, we select LLaMA-2-7B and MPT-7B for further investigation into the necessity of parallel data for LLMs.
Do LLMs Have an Appetite for Parallel Data?
Conventional machine translation training predominantly relies on utilizing large volumes of parallel datasets within encoder-decoder frameworks. This trend is not confined to training models from scratch but also pertains to strategies that fine-tune pre-trained LLMs, often involving millions of parallel sentences (Rothe et al., 2020; Liu et al., 2020; Xu et al., 2021; 2023; Yang et al., 2023). In this section, we examine whether the recently proposed decoder-only LLMs retain a dependence on substantial parallel data and adhere to the traditional training paradigm.
Following Section 2.2, our focus narrows to fine-tuning LLaMA-2-7B and MPT-7B. To allow for a deep analysis, we concentrate on one language pair, EnglishRussian (enru). We opted for a language pair that is translating out of English and to a non-Latin language, since those categories show larger gaps with SoTA models in our initial investigation in Section 2.2. We use the clean data filtered from 75M parallel sentences from Hendy et al. (2023) and split the data size in 5 levels: 10K, 100K, 1M, 5M, and 20M. We use the same prompt template and training scheme as described in Section 2.1, and train the model by updating all parameters. Detailed training settings can be found in Appendix C.
2 Observations
The fine-tuning results for LLaMA-2-7B and MPT-7B at each data size step are presented in Figure 4. Additionally, we benchmark these against the performance of the NLLB-54B model to show the disparity with one of the SoTA multilingual translation models.
Small Training Data Is Enough According to COMET, there is a notable difference in the curve of LLaMA-2-7B and MPT-7B: LLaMA-2-7B peaks with 10K and 100K training data before experiencing a decline, while MPT-7B exhibits continuous improvement with more training data. LLaMA-2-7B requires only limited training examples (10K and 100K) to achieve competent translation. However, a surplus of examples (5M or 20M) seems to dilute its existing knowledge in Russian. Conversely, MPT-7B, potentially due to its inherently weaker translation capability, exhibits improved performance with an increase in training data. This may suggest that LLaMA-2 or other well-trained LLMs may not necessitate substantial parallel data.
Large Parallel Data Wash Out the Knowledge Both LLMs eventually achieve similar BLEU and COMET with 20M training data, regardless of their performance on smaller data. We hypothesize that this phenomenon is caused by catastrophic forgetting (French, 1999; Kirkpatrick et al., 2017), suggesting that too many parallel data wash out the pre-existing knowledge. To validate this hypothesis, we consider an extreme case: training the model from scratch using 20M data, thereby erasing all prior knowledge.We initialize parameters randomly based on the LLaMA-2-7B model, but use the same vocabulary. As expected, it tends up with a similar performance in both BLEU and COMET evaluations (triangle in Figure 4), strengthening our speculation regarding the dilution of LLM’s intrinsic knowledge with extensive data training.
Beyond BLEU COMET reveals a decline in translation performance for LLaMA-2-7B as the amount of parallel data increases, a trend not captured by BLEU which shows an increase. This discrepancy arises since BLEU primarily evaluates lexical overlap, and the extensive WMT training data, being similar in domain to the test set, likely enhances this measure. This highlights the necessity of utilizing additional metrics (like COMET) for a comprehensive evaluation of translation.
From our observations, LLaMA-2 (potentially other well-trained LLMs) should not adopt the same training approach as earlier models—-whether randomly initialized or pre-trained—that rely heavily on vast amounts of training data.
A New Training Recipe
We demonstrate that LLMs like LLaMA-2-7B do not voraciously consume parallel data. We introduce a novel training strategy that markedly enhances translation performance without relying heavily on parallel data. The recipe comprises two stages: continuous monolingual data fine-tuning and high-quality parallel data fine-tuning. After applying our training recipe to LLMs, we name the resulting model as ALMA (Advanced Language Model-based trAnslator).
Monolingual Data Fine-tuning LLMs like LLaMA are pre-trained on English-dominated corpora. This potentially explains their inadequate translation performance which necessitates cross-lingual capabilities. To ameliorate this, our first stage is fine-tuning LLMs with monolingual data of non-English languages involved in translation tasks, enhancing their proficiency in these languages. Note that we also add English monolingual data during fine-tuning to prevent English knowledge forgetting. Previous studies also offer some clues that monolingual data help in translation. For instance, Tan et al. (2023) utilizes a monolingual target corpus to bridge the gap in translation mismatches caused by domain discrepancies. BigTranslate (Yang et al., 2023) and PolyLM (Wei et al., 2023) use a huge amount of Chinese monolingual data and improve translation to or from Chinese. Furthermore, Li et al. (2023) utilizes monolingual generation instructions to improve translation. In Section 6.1, we show that utilizing small monolingual data and modest computational cost (e.g., 1B monolingual tokens mixed by 6 languages and fine-tuning under 18 hours), can facilitate significant improvements in 10 translation directions. Note that we employ full-weight fine-tuning at this stage.
High-Quality Data Fine-tuning Drawing on insights from Section 3.2 that LLMs may require only small parallel data, coupled with previous research emphasizing training data quality (Zhou et al., 2023; Maillard et al., 2023; Gunasekar et al., 2023), we fine-tune the model using a small, yet high-quality parallel dataset in this stage. To ensure the data quality, we collect human-written datasets from WMT test data and Flores-200 (NLLB TEAM et al., 2022) development and test sets. Here, we explore both full-weight and lightweight Low-Rank Adaptation (LoRA) (Hu et al., 2022; Mangrulkar et al., 2022) fine-tuning, where we apply LoRA to the down-projection layer in each feed-forward network.
Experiments
For our parallel training data, we collect human-written test datasets from WMT’17 to WMT’20, plus the development and test sets from Flores-200 (NLLB TEAM et al., 2022), resulting in a total of 58K training examples across all languages. For the test data, we still use the same 10 translation directions to be consistent with our study in Section 2: csen, deen, isen, zhen, ruen, where isen is from WMT’21 and the others are from WMT’22. Test data in WMT’21 (except for is) is used for the development dataset (a total of 8K parallel sentences).There is no development dataset for Icelandic. The monolingual dataset is sourced from OSCAR (Ortiz Su’arez et al., 2019; Kreutzer et al., 2022). We mix the monolingual data and fine-tune the model with a sampling ratio of 20%, 14%, 8%, 19%, 22%, and 17% respectively for de, cs, is, zh, ru and en. We explain the reasoning behind the sampling ratios and show the detailed parallel data information in Appendix D.
2 Training Setup
We train the model in a many-to-many multilingual translation manner, and use LLaMA-2-7B (or 13B) as our backbone model given its best zero-shot performance. Our two-stage fine-tuning process yields two model types, differentiated based on the utilization of LoRA:
ALMA-7B/ALMA-13B Full-Weight fine-tuning on monolingual data followed by Full-Weight fine-tuning on high-quality parallel data for LLaMA-2-7B or -13B models.
ALMA-7B-LoRA/ALMA-13B-LoRA Full-Weight fine-tuning on monolingual data followed by LoRA fine-tuning on high-quality parallel data for LLaMA-2-7B or -13B models.
If using LoRA, the LoRA rank is 16 and only updates 0.1% parameters (7.7M for 7B and 12M for 13B model). Both monolingual data fine-tuning and human-written data fine-tuning basically share the same hyperparameter settings. Specifically, we fine-tune LLaMA-2 with a batch size of 256, a warm-up ratio of 0.01, and a sequence containing a maximum of 512 tokens. For monolingual data fine-tuning, we train the LLaMA-2-7B up to 20B tokens and LLaMA-2-13B up to 12B tokens. However, it is very likely that the model would be better in translation with more monolingual data fine-tuning. For human-written data fine-tuning, we train the model for 2 epochs (enough to see a clear convergence) and pick the best model with the lowest validation loss. For both stages, we adopt deepspeed (Rasley et al., 2020) to accelerate our training.
3 Baselines
We evaluate our method against two baseline categories. First, we consider prior studies with the goal aligning with ours: leveraging LLMs for translation. Secondly, we benchmark against the current SoTA translation models. It’s worth noting that this comparison isn’t entirely fair due to discrepancies in training data and model architectures (e.g., 175B GPT-3.5 vs. our 7B models). Nevertheless, utilizing the same test set provides insights into our model’s current standing.
Prior Similar Work We compare our model with BigTranslate (Yang et al., 2023), which extends LLaMA-1-13B to over 100 translation directions; TIM (Zeng et al., 2023), which uses correct and incorrect examples to help LLM to learn translation; SWIE (Chen et al., 2023), which improves LLM in translation via instruction augmentation; and BayLing (Zhang et al., 2023b), which uses interactive translation instructions. Given that the same test data and evaluation metrics are utilized, we directly report BLEU and COMET from their papers (except for BigTranslate, we assess their released model using the prompt they provided).
SoTA Models We consider the NLLB-54B model, which is the largest and best translation model released in the NLLB family (NLLB TEAM et al., 2022); and the zero-shot performance of GPT-3.5-text-davinci-003 (GPT-3.5-D) and GPT-3.5-turbo-0301 (GPT-3.5-T). Additionally, we present the zero-shot results for GPT-4.GPT-4 results are sourced from Zhang et al. (2023b).
4 Results
We show our main results of enxx and xxen respectively in Table 1 and 2. In summary, our best system (ALMA-13B-LoRA) outperforms all previous studies, NLLB-54B, and GPT-3.5-D, while it marginally underperforms compared to GPT-3.5-T and GPT-4.
Comparing With LLaMA-2 Zero-Shot For all 10 translation directions and both 7B and 13B models, LLaMA-2 trained by our recipe significantly outperforms its original zero-shot performance. For instance, ALMA-7B achieves +16.12 BLEU and +17.61 COMET for enxx on average. It is worth noting that LLaMA-2-13B suffers from the off-target issue in enxx zero-shot translation. However, it can be substantially alleviated by few-shot in-context learning (Brown et al., 2020), but still largely lag behind our methods (e.g., over 10 BLEU and COMET when translating from English). We discuss this further in Appendix E.
Compared with Prior Similar Studies ALMA significantly outperforms all prior studies. BigTranslate, which is fine-tuned on Chinese corpus and 300M parallel corpus, struggles to surpass LLaMA-2’s zero-shot performance, except for enzh. This observation also aligns with our findings that an excessive amount of parallel data may damage the model, whereas target monolingual data is helpful to translation. Both TIM and SWIE specifically target two high-resource languages, de and zh. Their performance, however, is predominantly determined by their backbone models: effective translation is observed for zh but is lackluster for de when using BLOOMZ, and vice versa with LLaMA-1. In contrast, ALMA is versatile, showcasing strong results across all directions.
Compared with SoTA models Our best model (ALMA-13B-LoRA) substantially outperforms NLLB-54B and GPT-3.5-D on average. In enxx direction, it even outperforms GPT-3.5-T on average COMET (87.00 vs. 86.56) and has close performance when it comes to xxen. Notably, SoTA models typically excel with high-resource languages but falter with low-resource languages such as is. With our recipe, the performance of is remains strong and performs the best.
Analyses
In our main results, we present ALMA with our best settings, fine-tuned on either 20B or 12B tokens. Yet, we snapshot all ALMA models after every 1B monolingual tokens (and human-written parallel data) they have been fine-tuned with, and evaluate all their translation performance. As illustrated in Figure 5, we report the ALMA-7B’s average performance across all directions after fine-tuning every 1B tokens. The test dataset remains the same, i.e., the 10 aforementioned directions. We provide detailed numeric results and similar analysis for ALMA-13B to Appendix F. Importantly, merely fine-tuning on 1B monolingual tokens, followed by fine-tuning on human-written data, yields performance comparable to NLLB-54B and GPT-3.5-D. In practice, we employ 16 MI200 GPUs with a batch size of 256 and sequence length of 512, which requires only 18 hours to complete the fine-tuning of 1B tokens and an additional hour allocated for human-written data fine-tuning. It takes around 19 hours of training to have a strong MMT model.
2 The Effect of Monolingual Data and Parallel Data Quality
To scrutinize the impact of monolingual data, we juxtapose LLaMA-2-7B models fine-tuned with and without monolingual data (20B tokens), while keeping the same parallel data. Furthermore, to evaluate the impact of parallel data quality, we introduce three distinct parallel datasets for stage 2 fine-tuning. The first dataset is the human-written data (HW) utilized in prior experiments. The second is the filtered data (Filtered) referenced in Section 3.1. Lastly, we employ a randomly selected dataset (Random) sourced from the comprehensive WMT data. We anticipate the quality hierarchy as HW, followed by Filtered, and lastly, Random. For both Filtered and Random, each translation direction has 10K parallel data, aligning the total training dataset size with HW. We show the ablation results in Table 3. Using the LLaMA-2-7B as our foundational model, it’s evident that with the same parallel data, incorporation of monolingual data largely enhances translation results, e.g., an increase from 74.35 to 83.98 in enxx COMET scores when training on the same Filtered data. Moreover, regardless of the monolingual data’s presence, models fine-tuned with higher-quality data exhibit better performance. Both monolingual and human-written data emerge as critical factors in improving translation. Detailed results for each language pair are deferred to the Appendix G.
3 Other Analyses
We also explore additional in-depth analyses and elaborate on them in the appendix: 1) The impact of the volume and domain of human-written data on translation performance is explored in Appendix H; 2) A comparison between stage 2 fine-tuning (parallel data fine-tuning) and in-context few-shot learning can be found in Appendix I; 3) An evaluation of the zero-shot cross-lingual capabilities of LLaMA-2 after stage 1 fine-tuning on other tasks is presented in Appendix J.
Conclusion
In this paper, we show that LLMs do not require as extensive a collection of parallel data as traditional translation models do. Subsequently, we introduce a novel training recipe for decoder-only LLMs in translation, resulting in strong translation models, ALMA. When using our LLaMA-2 as our foundational model, ALMA exceeds the zero-shot translation performance of LLaMA-2 by more than 12 BLEU and COMET scores across 10 directions on average. Moreover, ALMA models surpass all preceding studies and even outperform NLLB-54B and GPT-3.5-D.
We extend our gratitude to Hieu Hoang, Marcin Junczys-Dowmunt, Yunmo Chen, Steven Tan, Huda Khayrallah, Thamme Gowda, Vikas Raunak, Matt Post, Anoop Kunchukuttan, Roman Grundkiewicz, Tom Kocmi, Kenton Murray and Arul Menezes for their insightful and valuable suggestions.
References
Appendix A Comparing LLM Training Objectives For Machine Translation
We evaluate three potential training objectives for decoder-only LLM in machine translation.
We first consider a standard language modeling loss that predicts the next token based on all prior tokens.
For decoder-only models, a prefix can be defined with a non-causal attention mask. Analogous to standard language modeling, the model predicts each token outside the prefix based on previous tokens. In the context of machine translation, the provided prompt serves as the prefix, as depicted in Figure 2.
The UL2 model (Tay et al., 2022a) introduces a unified approach to masking methods, utilizing a mixture-of-denoisers (MoD) strategy, which has also been implemented in the fine-tuning of PaLM (Tay et al., 2022b). This strategy is grounded in three objectives:
Regular Denoising: In this approach, noise is sampled in spans and replaced with sentinel tokens, aligning with the standard span corruption technique delineated in Raffel et al. (2020). The parameters set for this objective include a mean of 3 and a corruption rate of 15
Extreme Denoising: This method amplifies the noise to a comparatively ’extreme’ level, characterized by a mean length of 32 and a corruption rate reaching up to 25
Sequential Denoising: This is known as the Prefix LM objective previously mentioned.
In our training process, we allocate a 25% probability each for both regular and extreme denoising, and a 50% probability for sequential denoising.
We employ the MPT-7B as our backbone model. Our investigation considers four distinct training data sizes: 0 (zero-shot), 100K, 1M, and 5M, with translation directed from Russian to English. We use the parallel dataset previously described in Section 3.1. For each data size, the MPT-7B is fine-tuned using the corresponding training objective, noting that all trainings utilize full-weight fine-tuning.
The results of the comparison between training objectives can be viewed in Figure 6. Although three objectives end up with similar performance under 5M training data, both prefix LM and MoD markedly lag behind CLM under limited parallel data (100K or 1M). Surprisingly, with 100K, models fine-tuned using prefix LM and MoD even underperform their zero-shot performance. Conversely, CLM demonstrates a healthy improvement as the amount of parallel data increases. Consequently, we adopt CLM as our primary training objective for machine translation.
Appendix B Full Results of Zero-Shot Evaluation
In Section 2.2, we present the average zero-shot translation performance of recently released LLMs. Detailed results for each translation direction can be found in Table 4.
Appendix C Training Details
We fine-tune the backbone model using a warm-up ratio of 0.01, a maximum sequence length of 512 tokens, and a weight decay of 0.01. The test data from WMT’21 serves as our development set. The training spans 3 epochs (for MPT-7B as detailed in Section 3, and 2 epochs for LLaMA-2 human-written data fine-tuning). The best model is selected based on the lowest validation loss, with validation performed every 10% of the total training progress. We utilize 16 MI200 GPUs for training; each GPU manages 4 batches and has a gradient accumulation step of 4, yielding an effective batch size of 256. The peak learning rate is set at 2e-5 , with an inverse square learning rate decay to 0. The training operates under fp16 precision, facilitated by deepspeed Rasley et al. (2020), employing ZeRO stage 2.
Appendix D Data Information
In Table 5, we observe a substantial imbalance in the volume of monolingual data available for different languages, denoted by their respective word countshttps://huggingface.co/datasets/oscar-corpus/OSCAR-2301. Specifically, the English language dataset contains 523.9B words, vastly outnumbering other languages, such as Icelandic, which contains 0.3B words. Utilizing an unmodified concatenation and shuffling approach for this data would disproportionately prioritize English, undermining our objective of enhancing the model’s proficiency in non-English languages. To address this, we straightforwardly set the sampling ratio for English as , thereby ensuring a balanced learning emphasis. The remaining of the probability allocation employs temperature sampling, as suggested by Aharoni et al. (2019), a technique prevalently adopted in the processing of unbalanced multilingual machine translation. Consequently, the process of selecting a monolingual example from language adheres to the following distribution:
where is the amount of the data in language , is the temperature, and is the set of all languages except for English. The temperature we use is 6.
D.2 Data Statistics
We show data statistics in Table 5. The training parallel data is sourced from the WMT’17 to WMT’20. The development data was acquired from WMT’21, and the test data was derived from WMT’22, with the exception of the Icelandic dataset, which was procured from WMT’21. This means, Icelandic does not have development dataset. Additionally, the monolingual data was extracted from the Oscar dataset.
Appendix E Off-Target Issue for LLaMA-2-13B
In the zero-shot scenario, the performance of LLaMA-2-13 is reasonable for translations into English. However, we identify a significant off-target issue with LLaMA-2-13B when translating from English to other languages. This issue is highlighted in Table 6 using a red highlighted box. An illustrative example of the off-target issue is provided below:
English: Plug the wall charger (not included) to a power outlet, and then connect your eReader to the wall charger.
Russian: Comment: I’m voting to close this question as off-topic because it is not about programming.
Expectedly, the model should produce translations in Russian. Yet, LLaMA-2-13B outputs “I’m voting to …”, indicating a misinterpretation of the task, potentially linked to its pre-training phase. We address this off-target behavior through two methods.
One approach is to utilize prompts in the target language (Raunak et al., 2023). For instance, when translating from English to Chinese, the preferred prompt is: ”将其从英文翻译成中文:\n英文:source sentence\n中文:” as opposed to ”Translate this from English to Chinese:\nEnglish:source sentence\nChinese:”. Employing this technique markedly enhances the zero-shot performance of LLaMA-2-13B. Specifically, the BLEU score escalates from 0.87 to 20.80 for encs, and from 0.59 to 22.66 for enru.
Employing in-context few-shot learning by including several examples within the prompt has proven effective. We investigate both 1-shot and 5-shot learning scenarios. As delineated in Section I, we utilize two sets of examples: Filtered, extracted from the WMT training data, and another set randomly chosen from human-written data, termed HW. Table 6 demonstrates that both 1-shot and 5-shot configurations effectively counteract the off-target challenges. Few-shot learning exhibits performance comparable to the strategy of using prompts in the target language. Moreover, echoing observations from Section I, examples of human-written quality outperform those from the Filtered set.
Nevertheless, both strategies trail behind our proposed solution by a margin of approximately 5 BLEU and COMET points during translations into English, and by over 10 BLEU and COMET points in translations originating from English.
Appendix F Numeric Results for Models Fine-Tuned With Every 1B Tokens
In Table 7 and 8, results for LLaMA-2-13B and LLaMA-2-7B are presented. Both models were fine-tuned at every 1B-token interval (comprising six languages) before subsequent fine-tuning with human-written parallel data. Full-weight fine-tuning was employed to ensure a consistent comparison. During inference, the 7B models utilized a beam search of size 5, while the 13B models adopted a greedy search strategy. For 13B models, we only utilize a beam size 5 for the final models we reported in the main manuscript (Table 1 and 2).
The data from these tables highlight that fine-tuning only 1B tokens, followed by human-written data fine-tuning, is adequate to compete with or even outperform the state-of-the-art (SoTA) models.
Appendix G Detailed Results in Ablation Study
We show the detailed results of the ablation study on the effect of monolingual data and the quality of the data in Table 9.
Appendix H Is More Human-Written Parallel Data Better?
The composition of our human-written data consists of the prior-year WMT test sets (approximately 10K parallel sentences per pair) and Flores data (around 2K per pair). In this analysis, we assess the impact of additional human-written parallel data. Specifically, we compare models (LLaMa-2-7B after stage 1) fine-tuned exclusively on Flores against those fine-tuned on both Flores and WMT data. Results can be found in Table 10. Notably, upon integrating WMT data into the training set, we discern a modest improvement in COMET scores. However, there’s an uptick in BLEU scores, particularly for translations into English. We attribue the increase in lexical match (BLEU) to the domain alignment of WMT data. Consequently, our hypothesis is that while an augmented volume of human-written data might marginally enhance segment-level human judgment correlation (COMET), in-domain data can significantly enhance lexical matching.
Appendix I Parallel Data Fine-tuning vs. In-Context Learning
An alternative way to instruct the model to have better translation is in-context learning (ICL) (Brown et al., 2020), as opposed to additional fine-tuning on parallel data. However, ICL is limited to only a few shots given the length of translation examples, while fine-tuning can leverage entirely available data. For ICL, we consider 5-shot evaluations. 5 examples are randomly selected from Filtered data (the Quality-Random examples used by Hendy et al. (2023)). We also consider another 5 examples randomly from the human-written data to examine the impact of example quality. We here compare the performance of our fine-tuning method and 5-shot ICL.Due to ICL’s extended prompt length, employing a large beam size is impractical; hence, we opt for a beam size of 1 for all to ensure a fair comparison. We assess the LLaMA-2-13B after stage 1 (12B token fine-tuning) and present results in Table 11.
Interestingly, ICL also holds the same property that higher quality data leads to better performance (Filtered 5-shot vs. HW 5-shot). Moreover, as expected, ICL substantially underperforms our stage 2 fine-tuning possibly due to the small examples provided, which aligns with the findings in the previous work (Liu et al., 2022; Mosbach et al., 2023). This could also clarify why implementing ICL subsequent to stage 2 yields no additional benefits, as all high-quality data has already been incorporated during stage 2 fine-tuning (the last row in the Table).
Appendix J Cross-Lingual Proficiency of Fine-Tuned Models
We explore the cross-lingual competencies of our models derived from LLaMA-2 after fine-tuning them on monolingual data. Our aim is to discern whether augmenting monolingual data enhances performance in cross-lingual tasks. Experiments were conducted on zero-shot cross-lingual tasks encompassing three benchmarks: Cross-lingual language understanding (XNLI) Conneau et al. (2018), XStoryCloze—a translation of the English StoryCloze dataset into ten languages (Mostafazadeh et al., 2017), and XWinograd—a multilingual compilation of Winograd Schemas Tikhonov & Ryabinin (2021). Evaluations were restricted to languages overlapping with our fine-tuned languages, namely, German (with only XNLI being inclusive), Chinese, Russian, and English. Unfortunately, none of these datasets covers Icelandic. We first consider baselines for some widely used models: XLM-R large (Conneau et al., 2020), XGLM-7.5B (Lin et al., 2021), BLOOM-7B (Scao et al., 2022), and MPT-7B (MosaicML, 2023). In these comparisons, LLaMA-2 demonstrates the top performance for the tested languages. Subsequent fine-tuning with either 1B or 20B monolingual tokens on both LLaMA-2-7B and 13B models yields substantial enhancements for non-English languages across all tasks. A consistent trend observed was that increased monolingual data corresponds to greater performance boosts. Only English is observed for a negligible difference after fine-tuning monolingual data, which is an anticipated outcome given LLaMA-2’s proficient grasp of English. The tool we utilize for LLM evaluation is lm-evaluation-harness (Gao et al., 2021).https://github.com/EleutherAI/lm-evaluation-harness/tree/big-refactor