Multilingual Instruction Tuning With Just a Pinch of Multilinguality
Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, Matan Eyal
Introduction
Instruction tuning is a fundamental aspect of building modern general-purpose large language models (LLMs), involving fine-tuning a pre-trained model on pairs of instructions and corresponding responses (Mishra et al., 2022; Wei et al., 2022; Sanh et al., 2022; Ouyang et al., 2022). For these models to be globally applicable, they must operate on a wide range of languages, yet, most instruction tuning datasets are typically limited to English. While curating naturally occurring instructions and responses for every language is challenging, cross-lingual transfer has emerged as a promising approach, in which a model is fine-tuned using one language, and acquiring similar abilities in another (Pires et al., 2019; Wu and Dredze, 2019; Artetxe and Schwenk, 2019; K et al., 2020; Conneau et al., 2020a, b). The ability to follow instructions for languages seen only at pre-training can significantly expand the applicability of LLMs, allowing them to be used by more people worldwide. In this work, we show that instruction-tuning of multilingual LLMs transfers across languages better than previously known, and that even minimal language diversity in the tuning set can further unlock instruction-following generalization to languages that are unseen during instruction tuning.
We investigate the effect of multilingual data on instruction-following across languages using an LLM pre-trained on hundreds of languages Anil et al. (2023), and high-quality, open-ended instructions and responses Zhou et al. (2023); Köpf et al. (2023) translated into 11 languages, across different families and writing systems. Initially, we examine the transferability of monolingual instruction tuning across different languages. Naturally, tuning using each language individually enhances performance within that language. Notably, we find that this also translates into instruction-following capabilities across other languages, and that tuning with English, Italian, or Spanish yields the best average multilingual performance.
Inspired by this result, we turn to ask how much multilingual data is required to improve multilingual instruction-following, while preserving English performance. We find that replacing even just 40 English training examples with multilingual examples, significantly improves instruction-following in those languages. Surprisingly, this small amount of language-diverse examples also improves performance for languages that are only seen during pre-training and are not represented in the instruction tuning set at all.
The next question we tackle is whether increasing the number of languages in the tuning set can enhance generalization to new languages from the pre-training corpus. We find that tuning using a few languages enables better performance for languages unseen during tuning, compared to monolingual tuning with the same number of examples.
Finally, we test two potential factors that might influence the degree of cross-lingual transfer: language similarity and the amount of language-specific pre-training data, but find no significant correlations. Overall, our results provide recipes for multilingual instruction tuning that improves cross-lingual generalization, while preserving performance on English, under a fixed budget. In particular, we find that capable multilingual instruction-following models can be tuned even with a minimal amount of multilingual data.
Measuring Multilingual Instruction-Following
Our objective is to discover how multilinguality during instruction tuning affects general-purpose instruction-following across languages. We break this down to multiple questions, including how well can monolingual instruction tuning transfer to other languages, how many multilingual examples can enhance multilingual instruction-following while preserving English performance, and whether increasing the number of languages can result in improved cross-lingual generalization. In this section we elaborate on the data, evaluation protocol, models we use, and the human annotation process to ensure the models quality.
We use datasets of high-quality open-ended instructions and responses, rather than classic task-specific datasets. Our training data contains 1,000 English instructions and responses from LIMA (Zhou et al., 2023) and 3,640 from OpenAssistantWe focus on single-instruction/single-response interactions so we keep only the first prompt and response from conversations in OpenAssistant similarly to Li et al. (2023). (Köpf et al., 2023). These examples resemble real world scenarios of users interacting with chatbots, with queries like "Can you explain Fermat’s Last Theorem?" and "How to keep a dog hydrated?", that enable efficient tuning even with a small training set (Zhou et al., 2023). For evaluation, we use 617 instructions from AlpacaFarm Dubois et al. (2023), originated from Self-Instruct (Wang et al., 2023), Vicuna (Chiang et al., 2023), Koala (Geng et al., 2023), and hh-rlhf (Bai et al., 2022).We exclude AlpacaFarm’s evaluation instructions from OpenAssistant, as we tune using its training set.
We use the Google Translate APIhttps://cloud.google.com/translate/docs/reference/api-overview to translate the instruction-response pairs of the training set and the instructions of the evaluation set to 11 languages, creating parallel training and evaluation sets in Arabic, Chinese, Czech, English, Estonian, Finnish, Hebrew, Hindi, Italian, Russian, Spanish, and Swahili.Languages are selected from Table 21 in Anil et al. (2023), describing the top-50 languages the model (§2) was pre-trained on. While translated data is different from naturally sourced data per language, it allows for more control as the data size and semantics are similar for all languages. A overview of the languages, their language codes, families and scripts is described in Table 2 in Appendix A.
Evaluation
We conduct a side-by-side automatic evaluation protocol (Bubeck et al., 2023; Dubois et al., 2023; Dettmers et al., 2023; Gudibande et al., 2023; Zheng et al., 2023), in which an LLM assesses two responses for the same instruction, with the goal of identifying the superior one. We follow the common practice of presenting both responses to the model twice, alternating the order of the two responses (Zheng et al., 2023; Zhang et al., 2023). The exact prompt we use is shown in Figure 9 in Appendix B. We define a “win" for a certain response if the judge selects it twice irrespective of the order, and a “tie" if the model selects a different response for each order. We use a discounted-tie (Zhou et al., 2023) scoring method, in which a model receives a score of 1 for a win, 0.5 for a tie, and 0 for a loss. We average the scores of individual instructions to get the score over the evaluation set and present it in percentages. To validate that the LLM judge decisions align with human preferences across languages, we conduct a human annotation study and find good aggregated agreement scores of 79.5% for English, 77% for Spanish, and 76.5%, and 75% for Russian and Hebrew, receptively. Further details on validating the LLM judge are provided in Appendix D.
Instruction-Following Score Per Language
Throughout this work we measure instruction-following per language by comparing the performance of a model that was tuned on some training set , to a model that was monolingually tuned on the target language , by using the full training set in this language, . Formally, we define our instruction-following () metric for language :
Where is the side-by-side protocol applied on and , which are the models instruction-tuned on and , respectively. A score of 0% means that loses on all instructions, and 50% means the performance of and in are indistinguishable when aggregated over the evaluation set.
Model
We use the PaLM 2 model family of Transformer-based (Vaswani et al., 2017) LLMs that were pre-trained on hundreds of languages (Anil et al., 2023). We use PaLM 2-S as our pre-trained model for all the instruction tuning experiments, and an instruction-tuned PaLM 2-L as the judge for the side-by-side evaluation. The training and inference hyperparameters we use are described in Appendix C.
Human Validation
Our evaluation protocol relies on the quality of our monolingually tuned models. To validate their usage as high bar baselines in their respective languages, we conduct a human annotation study in 4 languages: English, Spanish, Russian and Hebrew. Namely, we sample 50 random instructions per language, and ask 2 native speakers to assign a score of excellent, pass, or fail (Zhou et al., 2023) to the responses generated by the model that was monolingually tuned using that language. Results in Figure 2 show that our tuned models indeed demonstrate strong instruction-following abilities. Notably, the scores across languages are similar or better than the reported numbers by Zhou et al. (2023) in English.The differences can be attributed both to the pre-trained model and to the size of the instruction tuning dataset.
How Much Multilinguality Is Needed For Multilingual Instruction Tuning?
We now describe our controlled experiments, designed to quantify the effect of multilingual data during instruction tuning of multilingual LLMs, following the research questions defined in §2.
To explore zero-shot cross-lingual transfer of instruction tuning in multilingual LLMs, we tune models on a single language and evaluate them on all of the rest. We find that all of those models are able to transfer non-negligible instruction-following abilities to other languages.
We instruction-tune 12 models, each one using the full train set in a different language. We generate responses using every such model to the evaluation instructions in all other languages. Finally, we calculate their per language scores as described in §2.
Results
Figure 1 shows the results, where rows represent training languages and every column is an independent heatmap of the results over a single evaluation language. Most importantly, tuning using each single language yields a model with some multilingual instruction-following capabilities across languages. For context, even the model with the lowest average score, the one tuned on Hindi, achieves a score of over 30% in 9 out of 11 cases.For example, a score of 30% can be obtained by wining 30% of the instructions and losing 70%, or by achieving a tie on 60% of the instructions and losing 40%. The model with the best average score is the one tuned on English, when Italian and Spanish also enable consistently high scores.
Notably, we manually inspect the generations and find that our tuned models consistently respond in the same language as their instruction, regardless of the language they were instruction-tuned on, in contrast with findings in previous work Touvron et al. (2023a); Chen et al. (2023). We hypothesize that this comes from the multilingual nature of PaLM 2s’ pre-training, compared to the more English-centric LLaMA Touvron et al. (2023a), further details are in Appendix E. In addition to our main setup, we also compare the generations of these models to the ones of the pre-trained model that was not instruction-tuned. Results shown in Figure 10 in Appendix F further demonstrate that instruction tuning in every language separately, greatly improves instruction-following abilities across different languages.
2 A Few Dozen Examples Improve Multilingual Instruction-following
Naturally, multilingual tuning, as opposed to English-exclusive tuning under a fixed training examples budget, should result in better downstream performance for non-English languages, and might hurt performance on English. Therefore, we ask how many multilingual examples can improve the instruction-following abilities across languages, while preserving English performance. To that end, we tune models on subsets of the English examples combined with subsets of multilingual examples in different ratios. We find a significant boost in multilingual instruction-following abilities even when using just a few dozen multilingual examples.
We create data mixtures with examples that are evenly split among all 12 languages, and the rest English examples.Every example appears exactly once in every mixture, in a single language. We create such a train set for every from 10 to 100, incremented by tens, and also for , for which only 40 multilingual examples are included from across all 11 non-English languages, and the rest are English examples. Finally, we evaluate every tuned model on every one of the 12 languages as defined in §2.
Results
Figure 3 visualizes the results. As expected, multilingual examples in the train set improve the score on their languages (Red), and diluting the number of English examples hurts the performance in English (Green). Notably, the significant multilingual improvement comes from replacing only of the English examples by multilingual ones, which translates to 40 examples evenly distributed across the training languages. These results on the effect of such a small amount of language-diversity extend findings regarding task-diversity by Zhou et al. (2023), which demonstrated that a capable monolingual instruction-following model can be tuned using only 1,000 high-quality examples. A second trend is that these models often outperform their monolingually-tuned counterparts on the very language the latter were exclusively tuned on (blue markers above the 50 line). For example, the model tuned using the uniform set () preforms similarly or better than the individual monolingually-tuned models in 8 of 12 languages, despite being trained on 12 times less instruction-response pairs for each language. This suggests that for some languages, multilingual tuning can enable better instruction-following abilities compared to a traditional monolingual tuning with the same number of examples.
3 A Few Dozen Examples Improve Cross-lingual Generalization
Combining the lessons on cross-lingual generalization from monolingual tuning and the effect of a small amount of multilingual examples from previous sections, we turn to examine how multilingual examples in the tuning set affect language generalization. Specifically, we conduct a similar experiment to the one in §3.2, this time using only half of the languages for tuning while the rest of languages are unseen. In line with the results from §3.2, we find that a very small amount of multilingual examples also improve performance on languages that were not in the tuning set.
We repeat the setup from §3.2, this time with only English and 5 more languages: Arabic, Finnish, Italian, Russian, and Swahili, and evaluate models again on all 12 languages.
Results
Results in Figure 4 show similar trends to the ones in Figure 3. Specifically, the average score over non-English training languages (red) again improves very quickly, even with . Strikingly, this is also true for languages that the model has only seen during pre-training, and are not represented at all in the instruction tuning dataset (orange). This suggests that very few multilingual examples can not only improve performance for the languages of those examples, but also enable better cross-lingual instruction-following generalization.
4 Even a Small Number of Languages Improves Cross-Lingual Generalization
Given the results on the impact of a small number of multilingual examples from a fixed set of languages, we ask whether a small number of languages can also enhance cross-lingual generalization. We experiment with different numbers of languages in the tuning set and indeed observe that the transfer to languages only seen during pre-training improves from the very first additional languages.
We instruction-tune models on a single language and up to 6 languages. At each step, we add a language to the tuning set, and split the same examples budget uniformly among the current set of languages. We use the 6 training languages from §3.3, and follow 3 different permutations that determine the order in which we add languages to the mix. These permutations are shown in Table 4 in Appendix G. We evaluate every model on each of the remaining 6 languages, and average scores per evaluation language across models that are tuned using the same number of languages.
Results
Results on Figure 5 show that adding languages to the tuning set improves cross-lingual generalization. The average score (red) increases from tuning on monolingual data to tuning on bilingual data, and even more when using 3 and 4 languages, where the average score gets to almost 50. At that point, there is an indication for saturation, as more languages does not seem to improve transfer further. These findings demonstrate that diversifying the instruction tuning data with only a few different languages can improve cross-lingual transfer to new languages, only seen during pre-training.
Bilingual Tuning Sets
To show this holds for even more combinations of languages, we randomly split all languages to pairs, and tune models using of the examples in the one language and in the other. We evaluate each of these models on the remaining 10 languages, and compare their score to the ones of the two models tuned using the full monolingual sets. Results on Figure 6 reveal that bilingual tuning helps generalize to new languages better than monolingual tuning.
Potential Factors of Transferability
Following the results from the previous sections, a natural question arises: what factors can predict the degree of cross-lingual transfer? We explore two immediate candidates. Initially, we examine the relation of various aspects of language similarity to transferability within language pairs. Next, we look into whether the proportion of language-specific data in the pre-training corpus correlates with the amount of cross-lingual transfer of instruction tuning using the given language.
A intuitive hypothesis is that aspects of language similarity like the script or mutual intelligibility might affect the levels of instruction tuning cross-lingual transfer between languages. We test this using a case study of 4 Slavic languages, looking into possible effects of such aspects. However, we do not find a signal indicating these factors strongly correlate with cross-lingual transfer for this setting.
We train models on monolingual versions of the data in Russian, Serbian, Croatian, and Slovenian, and evaluate their transfer to each other. These languages can be divided along several linguistic lines that are summarized in Table 1. First, Russian is East Slavic, and the rest are South Slavic. Second, Russian and Serbian both use the Cyrillic script, while Croatian and Slovenian use Latin. Moreover, Serbian and Croatian share a significant degree of mutual intelligibility.
Results
Results are displayed on Figure 7. As shown, there is no a strong signal indicating that any of the aspects above is correlated with better mutual cross-lingual transfer. Russian tend to transfer instruction-following abilities best, and even though Russian and Serbian both use Cyrillic, it is Croatian that transfers capabilities to Russian better in our study. Moreover, Despite being largely mutually intelligible, Croatian and Serbian do not seem to share cross-lingual abilities more than the others. Our results align with recent findings that language similarity does not impact transferability or interference in machine translation given sufficient data and model capacity Fernandes et al. (2023); Shaham et al. (2023).
2 Fraction of Data in Pre-training
A second possible predictor of the degree of cross-lingual transfer from a particular language is the extent to which the model was exposed to it during pre-training. Generally, a model’s downstream performance on a specific language correlates with the fraction of data in that language in the pre-training corpus (Muennighoff et al., 2023). In contrast, Figure 8 suggests this is not necessarily the case for the cross-lingual transfer from a specific language. We find a weak Pearson correlation of 0.22 between the average cross-lingual score of each language and the number of documents in that language in pre-training corpus (Table 21 in Anil et al. (2023)).
Related work
The success of the pre-training–fine-tuning paradigm Devlin et al. (2019) ignited a new line of work on cross-lingual transfer. Pires et al. (2019) and Wu and Dredze (2019) showed that the multilingual variant of BERT can be fine-tuned on a specific task in one language and preform this task on another language, and Artetxe and Schwenk (2019) reported similar findings with a Recurrent Neural Network. Conneau et al. (2020a) introduced XLM-R, a multilingual pre-trained encoder with strong cross-lingual abilities. Phang et al. (2020) showed that intermediate training on an English task improves XLM-R’s transfer across languages further, and Pfeiffer et al. (2020) suggested an adapter-based framework to improve cross-lingual and task generalization. Hu et al. (2020) proposed a benchmark for cross-lingual generalization consists of 40 languages across 9 NLP tasks.
K et al. (2020) found that the depth of the network matters for cross-lingual transfer, and Conneau et al. (2020b) showed that parameter sharing is more important than shared vocabulary. Choenni et al. (2023) delved into the influence of specific examples from the training data on the performance in other languages, and Malkin et al. (2022) investigated how pre-training BERT-based models using different language pairs affects cross-lingual downstream performance. Going beyond encoder-only models, Xue et al. (2021) proposed mT5, a multilingual variant of T5 Raffel et al. (2020), and showed the significance of model scaling for cross-lingual transfer in generation tasks. Ye et al. (2023) explored trasferability in English-centric models Touvron et al. (2023a) using four tasks.
In contrast to most cross-lingual transfer literature that is focused on task-specific fine-tuning, we explore trends of cross-lingual generalization for general-purpose instruction-following LLMs.
Multilingual Instruction Tuning
Initially, works on instruction tuning Mishra et al. (2022); Wei et al. (2022); Sanh et al. (2022) focused on cross-task generalization in English. Subsequently, a large body of work was dedicated to multilingual instruction tuning. Muennighoff et al. (2023) found that tuning models with English datasets enables zero-shot cross-lingual abilities to new languages. The authors also found that this holds for languages that the model has never intentionally seen during pre-training, and that multilingual training improves generalization to new tasks. Chen et al. (2023) investigated the effects of full parameter training vs low-rank adaptation Hu et al. (2022) and monolingual vs multilingual instruction tuning using the Stanford Alpaca Taori et al. (2023) data, machine translated into 5 languages. Lai et al. (2023) trained multilingual instruction-following models for 26 languages with reinforcement learning from human feedback Ouyang et al. (2022), and Zhang et al. (2023) suggested instruction tuning LLMs by prepending the instruction and response translated into a pivot language (e.g English) to the response in the target language.
In this work, we consider transfer from monolingual instruction tuning from 12 languages, rather than exclusively on English. Furthermore, we examine multilingual instruction-following using an LLM pre-trained on hundreds of languages, which might be a key to unlocking more transfer to languages not represented during tuning. Importantly, we unveil the potential of just a small amount of language diversity in the instruction tuning set for this cross-lingual generalization.
Conclusion
We demonstrate that cross-lingual transfer offers a promising avenue for building multilingual instruction-following LLMs. Our findings across different languages suggest that even monolingual instruction tuning using only one language can result in improved instruction-following capabilities in other languages. Moreover, incorporating even a small set of a few dozen multilingual examples can significantly enhance instruction-following performance for both the languages the model is tuned on, and ones that were only seen during pre-training. Additionally, training on such multilingual datasets achieves comparable or even superior performance compared to monolingual tuning for some languages. We observe a similar trend when exploring the effect of total number of languages in the tuning set, as even splitting the train set to only two languages improves generalization to new languages, compared to monolingual tuning. These findings pave the way for efficient and scalable development of multilingual LLMs capable of understanding and following instructions across languages with minimal multilingual supervision.
Limitations
Limitations of our work include the use of translation for expanding datasets to multilingual settings, the number of languages we evaluated on, and number of models we experimented with. We now discuss each of them.
One limitation of our work is that our data is translated using the Google Translate API, and not originally sourced by native speakers. Automatic translation is inherently imperfect and may introduce noise to the tuning sets. However, translation also allows to for a controlled setup with parallel data, in which the content of all training and evaluation examples is the same for all languages.
Number of languages
A second limitation is that we use 12 languages in our main experiments (§3), with 3 additional languages in the language similarity experiment (§4.1). Clearly, multilingual instruction-following models need to successfully operate in many more languages, and we leave work on scaling this number to future work.
Number of models
Lastly, we experiment with PaLM 2, and results may vary with different LLMs. Nevertheless, our focus on PaLM 2 highlights the potential of multilingual pre-training for future advancements in LLMs.
Acknowledgments
We thank Omer Levy, Or Honovich, Alon Jacovi, Avi Caciularu, and Omer Goldman for their valuable feedback.
References
Appendix A Languages
The languages we use, their language families, scripts ,and language codes are shown in Table 2.
Appendix B Side-By-Side Evaluation
Figure 9 shows the prompt given the the LLM judge for the side-by-side evaluation.
Appendix C Training and Inference Details
We now describe the hyperparameters we use in our experiments. We tune every model for 2,000 steps, using a fixed learning rate of 1e-5, a batch size of 128, and a dropout rate of 0.05. We limit inputs to 1,024 tokens and targets to 512 tokens. We sample a development set of 250 examples from every training set and select the checkpoint based on the development RougeL Lin (2004) score. During inference, we generate responses of up to 512 tokens using nucleus sampling (Holtzman et al., 2020) with and temperature of 0.7. For the judge, we use greedy decoding to generate the ID of the better response (1 or 2).
Appendix D Judge-Human Agreement
To measure PaLM 2-L agreement with human judgments across language, we conduct a human annotation process on four languages, English, Spanish, Russian, and Hebrew. For every language we sample 50 instructions and let two native speakers select the better response out of two options, similarly to the task we assign the LLM judge (Figure 9). We always present the response by the model that was monolingually tuned using the evaluation language, alongside a response by model selected at random from the of the monolingually tuned ones described in §3.1. The agreement score on a single instruction is 1 if the LLM judge and human agree, 0.5 if exactly one of them selects a tie, and 0 if each selects a different response (Zhou et al., 2023). Table 3 shows the results. Overall, the LLM judge agreement with humans is strong for all four languages, yet there is some room of 2.5-7 points from inter human agreement in all languages. As expected, the models’ highest agreement with humans is in English with 79.5%,. In the rest of the languages the agreement is a few points lower.
Appendix E Response Language
When a user prompts a model in a specific language, they usually expect to receive a response in that same language. However, pre-trained LLMs often respond in a different language than the language of their prompt Touvron et al. (2023a); Chen et al. (2023). This poses a challenge also for evaluation of open-ended queries, since those are commonly evaluated with an LLM-as-a-judge Zheng et al. (2023) protocol, and the judges often ignore whether the response language match the prompt language, even when instructed not to Chen et al. (2023). Usually, this is handled by forcing the lowest score to such response Chen et al. (2023), which does not account for all cases.For example, a response in English to a prompt in French can still be very helpful, or when the prompt is a request for translation or code. To verify our trained models respond in the same language as their prompt, we manually annotate the language of responses to evaluation instructions in all languages. For every language, we randomly sample 20 responses from the pool of models tuned monolingually in other languages, to end up with a total of 240 generations from various models. We find that 239 responses are in the same language as the prompt, as desired. This is a major difference in the behavior of our PaLM 2-based instruction-tuned models and the commonly used Chen et al. (2023) LLaMA-based ones Touvron et al. (2023a, b). We hypothesize this stems from the multilingual emphasis in the pre-training of PaLM 2, compared to the more English-centric LLaMA.
Appendix F Comparison to The Base Model
The scores of models of model instruction tuned monolingually compared to the pre-trained model that was not instruction tuned, as opposed to our main evaluation setup, are shown in Figure 10. As evident, instruction tuning the model on each of the languages separately unlocks instruction-following abilities across all languages.
Appendix G Languages Permutations
We use 3 different permutations of 6 languages to determine the order in which we add languages to the tuning set in the experiment described Section 3.4. The permutations are displayed in Table 4.