Extrapolating Large Language Models to Non-English by Aligning Languages
Wenhao Zhu, Yunzhe Lv, Qingxiu Dong, Fei Yuan, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, Lei Li
Introduction
The language ability of LLMs is often imbalanced across languages (Zhu et al., 2023; Yang et al., 2023; Zhang et al., 2023), because both the pre-training corpus (Blevins & Zettlemoyer, 2022) and the instruction-tuning data (Wang et al., 2023b) are English-dominated. As a result, LLMs usually perform poorly on non-English languages, especially on languages that are dissimilar to English (Bang et al., 2023; Huang et al., 2023).
There have been some attempts to enhance LLMs’ non-English abilities by continued pre-training with large scale monolingual corpus (Cui et al., 2023; Yang et al., 2023). However, learning a language from monolingual data may need large scale data and computing.
In this paper, we elicit the non-English ability of pre-trained LLMs by building semantic alignment between English and non-English. To extrapolate the English ability to a particular non-English language, we propose a multi-task setting which combines translation tasks and cross-lingual general tasks during instruction-tuning. The translation tasks are used to stimulate the semantic alignment between languages, while the cross-lingual general tasks enhance the instruction-following capabilities of models (Fig 1). These cross-lingual instruction tuning (CoIT) brings out a cross-lingual model tailored to a specific non-English language.
Next, we explore to extrapolate LLM to multiple languages simultaneously through multilingual instruction-tuning (MuIT) with mixed multilingual resources (Fig 1). We consider two specific settings in our study. In the first setting, we simply combine all available resources for instruction-tuning to obtain multilingual LLM. In the second setting, we consider a practical scenario where instruction-tuning is performed under a specific data budget. To achieve optimal data allocation, we formulate the task as non-linear programming based on our previously discovered scaling laws. The objective of this optimization is to maximize the averaged multilingual performance.
In the experiments, we use LLaMA-7B as the pre-trained LLM and consider six challenging target languagesThe six languages are Arabic (Ar), Greek (El), Hindi (Hi), Turkish (Tr), Vietnamese (Vi), Chinese (Zh). that share little alphabet with English. For each language, a separate x-LLaMA is obtained with language-specific data. And a m-LLaMA is obtained with mixed multilingual data.
Experiment results on two cross-lingual benchmarks XQUAD and MLQA show that x-LLaMAs outperform the model tuned with English instructions (Alpaca-7B) by an average of 27.83%. Notably, the accuracy of x-LLaMAs on non-English tasks is comparable to the performance of Alpaca-7B on English tasks. We also observe that x-LLaMAs exhibit strong translation ability without the need of massively continued pre-training. Evaluation results on multilingual translation dataset Flores-101 show that x-LLaMAs outperforms previous LLaMA-based models by an average of 18.89% and even outperforms the supervised multilingual translation system M2M-12B (Fan et al., 2021) in half of the evaluated translation directions.
In our first setting of multilingual instruction-tuning, we discover that m-LLaMA can achieve comparable performance to strong x-LLaMAs on individual languages. Moreover, m-LLaMA is now capable of following multilingual instructions. Further analysis on response content and representation space reveals that m-LLaMA has a tendency to generate non-English response based on its English memory and multilingual semantic space becomes aligned in the middle layers of m-LLaMA, demonstrating the effectiveness of our methods. In the resource-constrained setting, experimental results show that our optimized data allocation yields higher multilingual performance than a uniform allocation, which showcases a practical usage of our formulated scaling laws.
The main contribution of this paper can be summarized as:
We explore cross-lingual instruction-tuning (CoIT) and multilingual instruction-tuning (MuIT) to elicit the non-English ability of LLMs.
Experiment results demonstrate that our instruction-tuning methods can simultaneously boost LLM’s non-English language ability, e.g., following multilingual instructions, generating multilingual response, and translation ability.
We formulate the scaling law in cross-lingual instruction-tuning and devise a data allocation strategy based on formulated laws for resource-constrained multilingual instruction-tuning.
We compare the scaling law of cross-lingual instruction-tuning and continued pre-training, and show that aligning language is a more efficient choice.
Background
To unlock the potential of pre-trained LLMs, Wei et al. (2022) propose instruction-tuning. In this stage, LLM will be fed with instruction data , where is a task instruction that describes the task requirement. is an optional input and is the desired output for the given task. The objective of this optimization stage is to minimize following the negative log-likelihood.
where denotes learnable parameters of the LLM and represents the instruction tuning dataset. The instruction-tuning dataset often covers diverse tasks, which is found beneficial for generalization to unseen instructions and tasks (Wei et al., 2022). However, we notice that commonly-used instruction-tuning datasets, e.g., Alpaca (Taori et al., 2023), FLAN (Longpre et al., 2023) are English-dominant, which limits LLM’s potential on following non-English instructions and solving non-English tasks.
Eliciting LLM’s non-English Ability
Empowering LLM on more languages beyond English is non-trivial. Training a LLM from scratch for each non-English language is almost prohibitive due to the huge cost of data collection and computation. In this paper, we explore to elicit pre-trained LLM’s non-English ability by strengthening semantic alignment between English and target languages. We begin by targeting a single language by performing cross-lingual instruction-tuning (CoIT, §3.1). To better understand the potential of aligning languages, we design a formulation to describe the scaling law in CoIT (§3.2). In the end, we introduce multilingual instruction-tuning (MuIT), aiming at eliciting LLM’s language ability on multiple non-English languages simultaneously, and present a potential usage of the scaling law in multilingual data allocation (§3.3).
To elicit LLM’s non-English ability, we perform cross-lingual instruction-tuning (illustrated in Fig 1) with multi-task data, including cross-lingual general task instruction data and translation task instruction data .
Considering that commonly-used instruction-tuning dataset is almost in English, we translate it to a foreign version with a translation engine. We then utilize both English and non-English version as cross-lingual general task instruction data. This approach aims to encourage LLM to better comprehend and follow cross-lingual instructions.
Intuitively, translation data is a valuable resource for learning semantic alignment. Previous researches have also shown that LLM’s translation performance can be enhanced by using expert-annotated translation data Jiao et al. (2023); Zhang et al. (2023) for instruction-tuning. Unlike them, we use publicly available parallel corpora, e.g., WikiMatrix (Schwenk et al., 2021), NewsCommentary (Tiedemann, 2012), to construct translation task instruction data, making our method more reproducible, scalable and extendable to more languages. While both En-X (translating English to non-English) and X-En (translating non-English to English) translation data are beneficial for learning semantic alignment, we find that placing non-English text on the target side of translation data yields better performance improvements for LLMs on non-English tasks compared to placing it on the source side. This finding will be further demonstrated in the upcoming experiments.
2 Scaling Law of Cross-lingual Instruction-tuning
We use bilingual translation performance as an indicator of semantic alignment and find that the scale of translation task instruction data has a huge impact on it. To quantify the relationship between translation performance and translation data scale , we formulate the underlying scaling law based on following intuitions: (1) The upper bound of is 100, which is the maximum score of frequently used translation quality metrics such as COMET and BLEU. (2) Translation performance tends to improve as the translation data scale increases. (3) Languages that are less similar to English require a larger amount of translation data to establish semantic alignment compared to languages more similar to English. Consequently, we present our final formulation as follows:
where and are parameters to estimate, is the language similarityWe calculate language similarity following the approach of Pan et al. (2021) using multi-way translation data. between the target language and English. When estimating the scaling law for a specific language, we first compute and then estimate and with observed data points. In the following subsection, we will further demonstrate how these scaling laws can assist us in optimizing data allocation when constructing multilingual LLM in a resource-constrained scenario.
3 Multilingual Instruction-tuning
While cross-lingual instruction-tuning is effective, serving customized LLMs for each language can be costly, particularly as the number of languages increases. Therefore, we take a step further and investigate the possibility of extrapolating a singe pre-trained LLM to multiple non-English languages simultaneously.
To achieve this goal, we perform multilingual instruction-tuning with a combination of multilingual resources. This includes general task instruction data in multiple languages and translation task instruction data from multiple directions. By leveraging these resources, the instruction-tuned LLM can establish alignment between English and multiple languages, enabling it to comprehend and follow multilingual instructions.
As for data mixture, we consider two settings. In the first setting, we straightforwardly combine all available resources for instruction-tuning. But a potential drawback of this approach is that instruction-tuning LLM with large-scale multilingual data may take huge computational cost.
Therefore we also consider a practical scenario, where the applied instruction data is constrained by a specific data budget. For instance, the total amount of utilized parallel data is a fixed number. To achieve the optimal data combination in this scenario, we propose to formulate data allocation as a non-linear programming problem. The objective of this programming problem is to maximize the averaged multilingual performance:
There are two constraints in this formulation: (i) data budget, the total amount of translation task instruction data is limited by a fixed budget . (ii) data availability, the maximum number of available translation data for language is .
Experiment Setting
We take LLaMA-7B as the pre-trained LLM, which is trained on trillions of tokens (mainly in English) and found to be competitive with state-of-the-art LLMs (Touvron et al., 2023). We construct x-LLaMAs for six challenging target languages: Arabic (Ar), Greek (El), Hindi (Hi), Turkish (Tr), Vietnamese (Vi) and Chinese (Zh), which share little alphabet with English.
Baseline LLMs
For comparison, we include several models that are built by instruction tuning on LLaMA: Alpaca-7B (Taori et al., 2023), which is tuned with English instructions; Parrot-7B (Jiao et al., 2023), which is tuned with human annotated translation data; Bayling-7B (Zhang et al., 2023), which is tuned with human interactive translations and English instruction data. We also present results from Chinese-Alpaca-7B (Cui et al., 2023) and Bigtrans-13B (Yang et al., 2023) for reference. Both these two models extend the vocabulary of LLaMA and use a large scale monolingual data for continued pre-training.
Instruction tuning details
For translation task instruction data, we use publicly available parallel corpora, WikiMatrixhttps://opus.nlpl.eu/News-Commentary.php (Schwenk et al., 2021) and NewsCommentaryhttps://github.com/facebookresearch/LASER/tree/main/tasks/WikiMatrix (Tiedemann, 2012). These corpora are more accessible and scalable compared to high-cost expert-annotated data (Jiao et al., 2023; Zhang et al., 2023). The statistics of two datasets are presented in Table 1. For multilingual general task instruction data, we incorporate Alpaca dataset (Taori et al., 2023), which consists of 52k English questions and corresponding response, and we obtain its foreign version with in-house translation engine. We use stanford_alpacahttps://github.com/tatsu-lab/stanford_alpaca as the code base. More training details are provided in Appendix A.
Evaluation Dataset
To evaluate LLM’s performance on non-English languages, we use two benchmark cross-lingual datasets, XQUAD (Artetxe et al., 2020) and MLQA (Lewis et al., 2020), which requires the model to reason over the given context and answer the given question. In addition, we create a new multilingual evaluation set MI-Eval (introduced in Appendix B) to assess the capability of LLM in following multilingual instructions. These multilingual multi-way test sets also allow us to compare language ability across languages. To evaluate LLM’s translation ability, we follow Zhu et al. (2023) and use multilingual translation dataset Flores-101 (Goyal et al., 2022). Details of the prompts used for all these tasks are provided in Appendix C.
Evaluation Metrics
On XQUAD, MLQA and MI-Eval, we follow Liu et al. (2023) and Wang et al. (2023a) to use ChatGPT for generation quality evaluation. On XQUAD, MLQA, we also report exact-matching results in Appendix D. For translation tasks, we use COMET (Rei et al., 2020), BLEURT (Sellam et al., 2020) and sentence-piece BLEU (Papineni et al., 2002) as metricsSpecifically, we report COMET score computed by wmt22-comet-da model and report BLEURT score computed by BLEURT-20 model.. More evaluation details can be referred to Appendix D.
Main Results
Table 2 presents experimental results for non-English question answering tasks. We can see that Alpaca-7B performs poorly on non-English, although it achieves 95% answer accuracy on corresponding English questions. Notably, x-LLaMA outperforms its counterpart (Alpaca-7B) by an average of 27.83% across six non-English languages. More importantly, x-LLaMA’s answer accuracy on non-English tasks is approaching Alpaca-7B’s answer accuracy on English tasks. This indicates that cross-lingual instruction-tuning is an effective way to elicit LLM’s non-English ability.
Table 2 also reports multilingual performance of other two representative LLaMA-based models. Both Chinese-Alpaca-7B (trained on a large-scale Chinese corpus) and Bayling-7B (trained on chinese-annotated interactive translation data) shows impressive performance on Chinese task. But they do not perform well on five other languages. Since their used training data can not be easily obtained, it is also hard to extend their training frameworks to cover more languages.
x-LLaMA shows impressive translation performance
Figure 2 and Figure 9 (Appendix E) present the performance of x-LLaMA and baseline systems in translating between English and non-English, which serves as crucial evidence for language alignment. Compared with other LLaMA-based LLMs, x-LLaMA exhibits higher translation performance in all evaluated directions. Notably, x-LLaMA even outperforms strong supervised baseline M2M-12B (Fan et al., 2021) on four En-X directions (translating English to non-English) and two X-En directions (translating non-English to English), and is approaching strong commercial translation engines, ChatGPT and Google Translate.
Scaling laws of cross-lingual instruction-tuning
Now, we investigate x-LLaMA’s translation performance under varying translation data scales and present the advantages of using scalable translation task instruction data for language alignment. As illustrated in Figure 3, adding translation data is always beneficial for strengthening semantic alignment. Encouragingly, our designed formulation (represented by the dotted line) effectively captures the trend and provide quantified relationship between translation performance and translation data scale. In subsequent experiments, we will demonstrate the practical applications of these formulated scaling laws, e.g., optimizing data allocation, analyzing learning efficiency.
2 Results on Multilingual Instruction-tuning
Now we combine all available resources to construct m-LLaMA. We compare the performance of m-LLaMA with that of x-LLaMAs customized for individual languages. Figure 4 shows that m-LLaMA can achieve comparable performance to x-LLaMAs on non-English QA tasks and multilingual translation tasks. This indicates the feasibility of extrapolating pre-trained English LLaMA to multiple non-English languages simultaneously.
m-LLaMA is able to handle multilingual instructions according to its English memory.
More importantly, incorporating multiple languages into a single m-LLaMA enables it to follow multilingual instructions. Evaluation results on MI-Eval (Fig. 4) show that m-LLaMA can achieve comparable response quality to x-LLaMAs when provided with instructions in different languages.
Additionally, we find that our instruction-tuning approach has minor impact on LLM’s English proficiencyWe draw this conclusion by comparing the answer quality of m-LLaMA and Alpaca-7B on English test set of MI-Eval. Generally, their answer quality is close. On 24% of test cases, m-LLaMA wins. On 32% of test cases, m-LLaMA loses. On the rest of test cases, the two models tie. but makes m-LLaMA have a tendency to generate response to non-English instructions with its English memory. Table 3 shows two representative cases where m-LLaMA produces similar response when given instructions in different languages. This phenomenon suggests that English and non-English becomes aligned within LLM after our instruction-tuning.
Visualization results show that multilingual semantic space becomes aligned in the middle layers of m-LLaMA.
For comprehensive analysis, we investigate the representation space of m-LLaMA and Alpaca-7B. Specifically, we use them to encode multilingual multi-way data from Flores-101 dataset and compare encoded representations across different layers. Figure 5 displays visualization results. For Alpaca-7B, representations of different languages always stay apart from bottom layers to top layers. In contrast, we observe representation overlap in m-LLaMA, especially in the middle layers, which offers another evidence that multilingual instruction-tuning encourages language alignment.
In our second setting, we study data-constrained multilingual instruction-tuning and explore the usage of formulated scaling laws. We compare our devised allocation approach with uniform data allocation in Table 4. The experiment results is mixed. When the data budget is low, e.g., 300k, the gap between different data allocation strategies is minor. When the data budget reaches 1.2M, our optimized allocation achieves significantly higher averaged multilingual translation performance than uniform allocation in all three metrics, which meets our initial optimization objective.
Analysis and Discussion
We conduct experiments with different combinations of instruction data for ablation study (Table 7). Instruction tuning LLaMA-7B with Chinese Alpaca data is better than English Alpaca data on the Chinese task. Jointly using two versions of Alpaca data brings further improvement. Interestingly, using translation task instruction data alone can reach moderate answer accuracy. And we find that putting Chinese on the target side of translation data is more useful for boosting LLaMA’s Chinese ability. Jointly using cross-lingual general task instruction data and En-Zh translation task instruction data reaches the highest accuracy.
Using translation data is far more efficient than monolingual data for building semantic alignment.
Using monolingual corpus of target language for continued pre-training is another way to help LLM to understand non-English and improve translation performance (Yang et al., 2023). For comparison, we use Chinese monolingual corpus mC4 (Xue et al., 2021) for continued pre-training and use cross-lingual general task data for instruction-tuning. Figure 7 compares scaling laws of two approach. We observe that using parallel data is far more efficient than using monolingual data for accomplishing semantic alignment.
Discussion on extending vocabulary for non-English.
Unlike previous work (Cui et al., 2023; Yang et al., 2023), we do not extend vocabulary for target non-English languages. The effect is dual. Our approach does not require a large-scale non-English corpus to learn embedding of extended tokens. On the other hand, since LLaMA usually tokenizes non-English tokens to bytes, our model is slower in encoding and decoding non-English sequence than those models equipped with extended vocabulary. We leave the exploration on vocabulary manipulation as our future work.
Conclusion
In this paper, we focus on extrapolating pre-trained large language models to non-English by building semantic alignment across languages. Specifically, we explore two approach: cross-lingual instruction-tuning (CoIT) and multilingual instruction-tuning (MuIT). Experiment results show that our cross-lingual models, x-LLaMAs, achieve great improvements on non-English, e.g., outperforming its English counterpart (Alpaca-7B) by 27.83% on question answering tasks and by 18.89% on translation tasks. After training on mixed multilingual resources, our m-LLaMA model can achieve comparable performance to strong x-LLaMAs on individual languages and is capable of following multilingual instructions. Further analysis of response consistency and representation space reveals that multilingual semantic space becomes aligned in the middle layers of m-LLaMA. In the setting of resource-constrained multilingual instruction-tuning, we show the usage of formulated scaling laws to achieve optimal data allocation. Overall, our approach and findings illuminate the potential for developing more potent LLMs for non-English languages.
Acknowledgement
We would like to thank Yinquan Lu for his support to this project. Shujian Huang is the corresponding author.
References
Appendix A Details of Our Instruction-tuning
For each experiment, we instruction-tune LLaMA’s full parameters for 3 epoch on 8A100. The learning rate is set as 2e-5 and batch size is set as 128. For training acceleration, we adopt FSDP training strategy (Zhao et al., 2023).
Appendix B Details of Our Constructed MI-Eval Dataset
We follow the fashion of “self-instruct” (Wang et al., 2022) and generate new English instructions with Alpaca dataset as seed. Then we translate these instructions to six non-English languages with strong multilingual machine translation system NLLB Costa-jussà et al. (2022) and obtain the multilingual multi-way evaluation set MI-Eval.
Appendix C Our Used Prompts for Downstream Tasks
We report all our used prompts in Figure 8. For question answering tasks, i.e., XQUAD and MLQA, we apply language-specific prompt when evaluate LLM’s performance on the target language. Table 5 lists two cases for better illustration. For machine translation tasks, i.e. Flores-101, we use English instruction for multilingual translation in our experiments.
Appendix D Details of Conducting Evaluation with ChatGPT
In this paper, we use ChatGPT to automatically evaluate the quality of question answering and instruction following. The evaluation prompts are reported below. Considering the API cost of evaluating with ChatGPT, we use the first one hundred questions in XQUAD and MLQA as representatives for experiments. Table 6 also reports exact-matching results on full test set. However, we notice two limits of exact-matching evaluation: (1) it does not penalty answer that heavily copies the given context. (2) it does not favor answer that is correct but different from the reference answer.
Appendix E Translation Performance on Reverse Translate Directions
Due to the page limit of main text, we report translation performance of different systems on translating non-English to English here (Fig. 9). The findings are similar to those in §5.1.