Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?

Tannon Kew, Florian Schottmann, Rico Sennrich

Introduction

Conversational instruction tuning is a popular method for aligning large pretrained language models (LLMs) with user expectations such that they can effectively respond to a user’s input query and follow natural language instructions Ouyang et al. (2022); Wang et al. (2023); Chiang et al. (2023). An implicit expectation of conversational chatbots is that the language of a model’s response should match that of the user’s input query. For instance, unless otherwise specified, a German-language input should result in a German-language output. However, since the vast majority of pretraining and tuning data used to develop these models is in English, many instruction-tuned LLMs struggle to respond consistently in other languages Touvron et al. (2023b); Chen et al. (2023); Ye et al. (2023); Zhang et al. (2023b).

Despite this limited exposure, English-centric LLMs such as Llama 2 can achieve near-perfect input/output (IO) language agreement when tuned with relatively few multilingual conversational instructions. Figure 1 depicts IO language agreement, as measured by OpenLIDhttps://github.com/laurieburchell/open-lid-dataset Burchell et al. (2023) and compares multilingual tuning on the language-diverse Guanaco dataset Dettmers et al. (2023); Köpf et al. (2023) with monolingual tuning on an English-only subset of instructions. As can be seen, multilingual tuning elicits strong IO language agreement, well within the bounds of language identification error rates, for non-English languages seen to differing degrees during pretraining and finetuning without degrading performance on English.

This observation raises two major questions, which we aim to address in this paper: Q1: How much multilinguality is required during finetuning to elicit few- and even zero-shot cross-lingual generalisation in English-centric LLMs? Q2: Which languages and tasks benefit most from multilingual instruction tuning of English-centric LLMs?

To investigate these questions, we instruction-tune English-centric LLMs, varying the number of languages seen during finetuning and evaluate performance in multiple target languages and on different tasks. We consider both high- and low-resource language settings with different scripts and investigate performance on generative tasks, such as open-ended chat and extractive question answering, as well as structured tasks aimed at assessing commonsense reasoning and language understanding.

Our results indicate that multilingual instruction tuning is crucial for eliciting cross-lingual transfer on generative tasks that assume IO language agreement while being less important in structured tasks that are commonly used to benchmark LLM performance. Furthermore, we empirically show that only a small number of finetuning languages is required to promote cross-lingual transfer. This highlights that tuning data for all potential target languages is not necessary to derive a capable polyglot chat model from an English-centric LLM.

Related Work

Instruction tuning describes a supervised finetuning (SFT) strategy that aims to provide a model with a diverse set of demonstrations of user input queries paired with desirable model outputs. Unlike task-specific SFT, instruction tuning aims to promote cross-task generalisation, allowing for a ‘generalist’ model that is capable of completing any text-based task on the basis of natural language instructions provided at inference time Mishra et al. (2022); Wei et al. (2022); Wang et al. (2022); Sanh et al. (2022); Longpre et al. (2023). Meanwhile, framing instructions in a conversational manner and over multiple dialogue turns has been shown to be effective at deriving performant chat models Taori et al. (2023); Conover et al. (2023); Chiang et al. (2023); Dettmers et al. (2023); Ding et al. (2023). Furthermore, LLM instruction tuning remains effective given relatively limited labelled data (Ouyang et al., 2022; Touvron et al., 2023b; Zhou et al., 2023), parameter efficient methods Hu et al. (2021); Zhang et al. (2023a) and model quantisation Dettmers et al. (2023); Li et al. (2023).

2 Cross-lingual Transfer in English-centric LLMs

Despite numerous pushes towards improving NLP support for non-English and low-resource languages Costa-jussà et al. (2022); Le Scao et al. (2023), the vast majority of today’s publicly available LLMs are English-centric. For instance, GPT-3’s training data consisted of approximately 93% English documents with the remaining 7% pertaining to other languages Brown et al. (2020).https://github.com/openai/gpt-3/blob/master/dataset_statistics This trend is further reflected in popular open-weight LLMs (see Table 1). One potential reason for this could be the “the curse of multilinguality” Conneau et al. (2020) which describes how as multilinguality increases, evaluation performance decreases. For example, multilingual LLMs have been shown to sacrifice performance on English tasks due to having to share a finite amount of pre-training tokens and model parameters across more languages Lin et al. (2022); Le Scao et al. (2022). Consequently, the cross-lingual transfer abilities of more performant, English-centric models is important for many NLP practitioners looking to deploy these models in real-world settings.

A number of works have reported on the multilingual abilities of proprietary models such as GPT-3 and its derivatives across a range of NLU and NLG benchmarks Lai et al. (2023a); Holmström et al. (2023); Armengol-Estapé et al. (2022). On tasks such as machine translation (MT), GPT-3.5-Turbo has been shown to outperform dedicated systems for some high-resource languages Hendy et al. (2023); Lu et al. (2023); Jiao et al. (2023); Bang et al. (2023), particularly when translating between two non-English languages Laskar et al. (2023). However, performance on low-resource languages generally lags behind.

Meanwhile, other studies have focused on open-weight LLMs. Ye et al. (2023) compare multilingual reasoning capabilities of pretrained BLOOM Le Scao et al. (2023) and LLaMA Touvron et al. (2023a) and find that despite minimal amounts of non-English pretraining data, LLaMA has strong cross-lingual transfer abilities. Muennighoff et al. (2023) find that multilingual multitask finetuning improves BLOOM’s zero-shot abilities across all target languages compared with English-only finetuning. Finally, Chen et al. (2023) show that large-scale multilingual instruction tuning can improve performance on open-ended chat in multiple target languages.

This last work is most closely related to our current work but differs in a few key factors. Firstly, we investigate multilingual chat capabilities of more recent English-centric models, pretrained with significantly more data. Secondly, inspired by Zhou et al. (2023) we adopt a less-is-more finetuning approach, controlling the amount and distribution of multilingual data. In doing so, we focus our analysis on the minimal amount of multilinguality needed to elicit cross-lingual transfer. Finally, our evaluations consider more target languages and distinct downstream tasks.

Experimental Setup

To explore the multilingual capabilities of English-centric LLMs, we instruction-tune a series of models on a fixed-size set of examples, varying the number of languages available. Following this, we evaluate the resulting models in multiple target languages on four distinct tasks that are representative of how LLMs may be used in downstream applications.

A prevailing trend in the development of recent LLMs is a clear focus on scaling up the size of the pretraining corpus Hoffmann et al. (2022). For instance, open-weight LLMs such as Falcon Almazrouei et al. (2023) and Llama 2 Touvron et al. (2023b) were trained on 1.5 and 2 trillion tokens respectively. Yet, although these numbers far surpass the 300 billion tokens used to train GPT-3 Brown et al. (2020), the distribution of language coverage remains similar across models with more than 90% pertaining to English (see Table 1). For our main experiments, we focus on Llama 2 7b, but also consider Llama 2 70b (§5.2) and Falcon 7b (Appendix E) to study the effect of model scaling and different training approaches respectively.

2 Instruction-tuning Data

For instruction tuning, we take inspiration from Dettmers et al. (2023) and finetune on high-quality conversations from the OpenAssistant dataset Köpf et al. (2023). These conversations comprise multiple dialogue turns between crowdworkers who were asked to either interact with or assume the role of a helpful AI assistant.

In contrast to Dettmers et al. (2023), who use all 9,846 top-rated conversations to train their ‘Guanaco’ models, we subsample training instances from the Guanaco dataset in order to control the amount of multilinguality. Specifically, we sample 3,200 unique English instances as an initial monolingual dataset, which we refer to as ‘Mono’. To construct datasets for multilingual finetuning we sample 200 unique training examples from each of five most frequent non-English languages in Guanaco (Spanish, Russian, German, Chinese, French). Given these subsets, we incrementally substitute English examples in Mono for non-English ones, one language at a time, following the order of how frequently each language appears in Guanaco. The resulting multilingual datasets are denoted as Multi-ii, where ii equals the number of distinct languages included. For comparison, we also train models on the full Guanaco dataset, which includes more than 30 distinct languages.

3 Instruction-tuning and Inference Settings

For efficient instruction tuning, we use LoRA adapters (R=64R=64, α=16\alpha=16), leveraging Hugging Face’s TRL libraryhttps://github.com/huggingface/trl. We train all models for 2,000 update steps, using an effective batch size of 64 and a learning rate of 1e−51^{e-5}. Sequences longer than 1024 tokens are truncated. For 7-billion parameter models, we use two NVIDIA GeForce RTX 3090 with 24GB of memory. The time required for each training run is approximately 8 hours. For the larger 70-billion parameter model, discussed in §5.2, we use the same hyperparameters and train on four NVIDIA A100 GPUs with 80GB of memory each. Here, a single training run takes approximately 20 hours.

At inference time, we load the instruction-tuned models with vLLM Kwon et al. (2023). For open-ended generation tasks, we use top-p sampling (p=0.9) with a temperature of 0.8 and frequency penalty of 0.2. For the more constrained generation tasks we reduce the temperature to 0.001. To account for randomness under these sampling parameters, we report aggregated results for all tasks by repeating inference experiments with three random seeds.

4 Evaluation Tasks and Languages

As evaluation tasks, we consider open-ended chat, extractive question answering, commonsense reasoning and natural language inference. As all of these tasks differ in terms of the availability and representation of ground-truth labels, we describe the specific evaluation strategies in the following section.

Finally, we investigate performance on different target languages, which we select based on the makeup of Llama 2’s pretraining data (see Table 1). In doing so, we aim to investigate i) the supervised setting, where a target language has been seen during pretraining and finetuning; ii) the zero-shot cross-lingual setting, where a target language may have been seen during pretraining but not during finetuning Wu and Dredze (2019); and iii) the extreme low-resource setting, where the target language has not intentionally been seen at any stage of the training procedure. According to Touvron et al. (2023b), German (de), French (fr), Swedish (sv), Chinese (zh), Spanish (es) and Russian (ru) represent the most frequently seen non-English pretraining languages with an estimated 3.4 to 2.6 billion tokens. Catalan (ca), Norwegian (no) and Bulgarian (bg) represent languages less frequently seen during pretraining with an estimated 800 to 400 million tokens. And finally, we select Icelandic (is), Hindi (hi) and Greek (el) to represent extremely low-resource languages, whose frequency in Llama 2’s pretraining data is not available but may still appear in small amounts due to intentional or unintentional contamination Blevins and Zettlemoyer (2022). Unfortunately, existing benchmark evaluation datasets do not cover all of these languages and thus we limit the target languages in those tasks to the available subset of our target languages.

Experiments

General-purpose chatbots are a popular application of instruction-tuned LLMs. To evaluate chat performance, we make use of the AlpacaEval prompt dataset Dubois et al. (2023), which includes a diverse set of prompts for open-ended questions, creative writing, brainstorming, and other tasks. We randomly sample 300 prompts and translate these into each target language. As a translation engine, we follow Lai et al. (2023b) and use GPT-3.5-TurboTo translate chat prompts from AlpacaEval, we use gpt-3.5-turbo-0613.. In contrast to using dedicated translation systems, employing GPT-3.5-Turbo for this purpose has the advantage of being able to explicitly specify instructions that allow for preserving code blocks, tables and terminology, which we include as part of our translation prompt (see Figure B). Furthermore, since GPT-3.5-Turbo is trained on instruction- and conversational-style data, we expect it to perform well at translating in this domain.

To automatically evaluate open-ended chat responses, we leverage LLM-as-a-judge (Zheng et al., 2023). Following Zhou et al. (2023), given an input prompt and model’s response, we ask GPT-3.5-TurboFor chat evaluation, we use gpt-3.5-turbo-1106 due to its longer context window and cheaper inference costs. to grade the helpfulness of the response on a 6-point Likert scale (see Figure 16 for the prompt used). For each evaluation instance, we provide the prompt and model-generated response directly in the target language, which we found to be on par with evaluating via first translating responses into English (cf. Hada et al., 2023) (see Appendix C for more details). As noted by Chen et al. (2023), GPT-3.5-Turbo sometimes ignores the fact that the output language differs from the input language. Therefore, we force a score of 1 (indicating least helpful) if the language of the response does not match the intended target language.

Figure 2 shows the helpfulness scores assigned by GPT-3.5-Turbo for all models for each of the target languages.

For English, we observe a reasonably strong and stable performance across all incremental multilingual instruction tuning settings, indicating that chat performance does not diminish with the increase in non-English training examples. In contrast, for most non-English target languages, performance increases significantly when moving from monolingual to bilingual instruction tuning. Most notably though, performance tends to plateau when instruction tuning with as few as three different languages. This observation holds for languages in both the supervised setting and the zero-shot setting, indicating that three finetuning languages is sufficient to elicit zero-shot cross-lingual transfer on this task. For languages in the extremely low-resource category, performance remains low despite multilingual finetuning. Manual inspection of these outputs reveal that, while these responses generally look convincing at first glance and are sufficient for language identification (see Figure 1), they are mostly nonsensical. This suggests that useful zero-shot cross-lingual transfer abilities of Llama 2 7b are limited to languages seen with higher frequency during pretraining.

2 Extractive Question Answering

In contrast to open-ended questions commonly used to query LLMs in chat settings, extractive question answering requires the model to identify relevant answer spans within longer context passages provided as part of the prompt. This kind of task closely resembles a retrieval augmented generation (RAG) setting, which is a popular method for extending an LLM’s knowledge with additional data not available during training Lewis et al. (2020); Izacard and Grave (2021). To evaluate our models on this task in multiple target languages, we use XQuAD Artetxe et al. (2020).We report results measured on the validation split of XQuAD since labels for the test split are not publicly available. This provides 1,190 QA pairs that were professionally translated into different languages.

For this task, we consider both a multilingual prompt setting and a monolingual prompt setting. The former presents the task instruction in English and provides the context passage and question in the relevant target language (en:xx), while the latter is entirely in the target language (xx:xx). As a starting point, we borrow the English prompt from Lai et al. (2023a) and manually translate it into each of the target languages considered. Additionally, we include a standardised response prefix as part of the prompt, effectively force-decoding the response “Based on the passage, the answer to the question is”. This allows us to better isolate the relevant answer string in the generative model’s output. A response is considered correct if the model’s answer string matches the ground truth answer after minimal postprocessing.We find that some postprocessing of model outputs is necessary for certain languages. Specifically, when queried with German and Russian prompts, all models consistently repeated the question before providing the extracted answer. To handle such edge cases, we strip away the question and truncate the system output to a maximum of 50 characters or the first line break, whichever comes first. An example of our prompting strategy for this task and model outputs is shown in Table 2.

Figure 3 shows the performance on each of our target languages in XQuAD given incremental multilingual instruction tuning.

Similar to the results on the chat task (§4.1), we observe no performance degradation on English as the ratio of multilingual instructions increases. For most non-English target languages in the supervised setting, we see that when using the multilingual prompting strategy (en:xx), performance remains relatively uniform, with the exception of Chinese. In contrast, given the monolingual prompting strategy (xx:xx), multilingual instruction tuning can lead to small but noticeable gains in performance over monolingual tuning. Specifically, for German and Chinese, performance improves when moving from monolingual to bilingual instruction tuning and again plateaus rather quickly with as few as three languages. For languages in the extremely low-resource setting (Hindi and Greek), multilingual instruction tuning fails to improve the model’s ability to handle these languages. While performance gains on this task are generally less pronounced than on the open-ended chat task, we note that zero-shot extractive QA is inherently challenging for LLMs tuned on conversational instructions as they tend to generate verbose responses rather than the single word or entity that correctly matches the ground truth.

3 Commonsense Reasoning

Effectively understanding natural language requires some representation of the natural world and how concepts can relate to one another. Commonsense knowledge describes the set of general facts that reflects this. To evaluate how well English-centric models can reason across languages, we use the X-CSQA dataset.We report results for X-CSQA measured on the validation set of 1,000 questions. This dataset consists of questions paired with multiple choice answers which aim at assessing general, language-agnostic world knowledge involving different types of commonsense reasoning Talmor et al. (2019); Lin et al. (2021). Given a question and a set of five possible answers from A-E, we prompt the model to output the letter corresponding to the answer that best answers the question. Again, we borrow the English prompt template for this task from Lai et al. (2023a) as a starting point and translate it into each target language. An example of the prompt used is given in Table 3. Finally we assess performance with the monolingual prompting strategy (xx:xx).

Figure 4 shows the accuracy on X-CSQA given incremental multilingual instruction tuning. Again, for English, we observe that multilingual instruction tuning does not significantly degrade performance. However, for non-English target languages, incremental multilingual instruction tuning fails to deliver improvements over monolingual tuning. This contrasts with the results on the previous tasks considered.

4 XNLI

Natural language inference (NLI) is an important skill for LLMs, especially as input and output sequences grow to consist of multiple sentences. Given two sentences, this task aims to recognise a relationship between them either as entailment, contradiction or neutral. To assess an LLM’s ability to solve this task given multilingual instruction tuning, we use XNLI (Conneau et al., 2018) and evaluate performance on the official test split.

For this task, we use the implementation in the lm-evaluation harness Gao et al. (2023)https://github.com/EleutherAI/lm-evaluation-harness, which uses language-specific prompts for each target language (i.e. monolingual prompting (xx:xx)). Instead of querying the model to generate the desired label, multiple queries are constructed for each test instance (one for each possible label) and scored by the model. The sequence with the highest likelihood is chosen as the model’s answer.

Figure 5 shows the accuracy measured on our target languages from XNLI given incremental multilingual instruction tuning. Strikingly, performance for all target languages remains uniform regardless of the amount of multilinguality used in finetuning. Furthermore, the equal performance between the Guanaco baseline and our Mono and Multi-ii finetuned models suggests that scaling up instruction diversity and the number of languages seen during finetuning beyond just six languages is still insufficient for delivering noticeable gains on this task.

Further Analysis

The results on four distinct evaluation tasks show that multilingual instruction tuning benefits open-ended chat in non-English languages most strongly (§4.1). In this section, we focus on this task to investigate the impact of instruction diversity and model scaling.

Diversity is a key factor for sample efficient instruction tuning Zhou et al. (2023). Since our experiments in §4 make use of native non-English training instances from the OpenAssistant dataset, a potential confounding factor could be that adding more languages during finetuning also introduces more diverse training instructions. To investigate this, we retrain Multi-ii models using translated instruction-tuning data from Mono in place of native non-English examples, following the same incremental recipe as described in §3.2. This ensures that the data distribution remains constant as multilinguality increases. As a translation engine, we again use GPT-3.5-TurboSince conversational training instances can be quite long, sometimes exceeding the default 4k token context window, we use gpt-3.5-tubo-16k for this translation task. and the prompt template provided in Figure 10.

Figure 6 compares the resulting performance on the open-ended chat task on a subset of our target languages (results for the remaining target languages are provided in Figure 15). Notably, we observe no significant differences between tuning with distinctly native non-English examples compared to those derived via automatic translation. This indicates that the gains attributed to increased multilinguality are not conflated with an increase in the diversity of instructions.

2 Scaling up Model Size

In order to investigate the effect of model scaling, we also repeat our chat experiments using Llama 2 70b as the underlying base model. Figure 7 shows the resulting helpfulness scores assigned by the LLM judge. Most notably, performance on non-English languages seen relatively frequently during pretraining is dramatically improved, often matching that of English. Secondly, we observe that the larger model’s performance on most non-English target languages tends to plateau with just two finetuning languages, unlike the smaller model that required three. Finally, while the performance on languages in the extremely low-resource scenario remains low, it exhibits a substantial relative improvement compared to the 7-billion parameter model. These results indicate that model scaling is extremely beneficial for exploiting the multilingual capabilities in English-centric models. This aligns with the assertion of Shaham et al. (2023) that larger models are more adept at handling multilinguality.

Discussion

Our findings show that multilingual instruction tuning can elicit zero-shot cross-lingual transfer in English-centric models, though its effectiveness on downstream performance varies across tasks. Notably, we observe significant performance gains for open-ended chat and noticeable gains on extractive QA with monolingual prompting. In comparison, for highly structured tasks such as X-CSQA and XNLI that impose a strict constraint on the output space regardless of the input language (e.g., choosing from options ‘A’, ‘B’, ‘C’, etc.), multilingual instruction tuning does not seem to provide substantial benefits. This distinction highlights that multilingual instruction tuning is most beneficial for open-ended generative tasks where there is an implicit agreement between the input query language and the language output by the model.

We posit that our findings align with the superficial alignment hypothesis Zhou et al. (2023). While instruction tuning guides the model towards a desirable ‘subdistribution of formats’ to use, a small amount of multilingual instruction data allows the model to learn a simple mapping between input and output languages. In essence, it helps steer the model towards language-specific ‘subdistributions of formats’ that agree with the language of the input. Therefore, we conclude that multilingual instruction tuning is of most benefit in scenarios where IO language agreement is required.

Most surprisingly, we find that multilingual instruction tuning with just three languages is sufficient to encourage IO language agreement in smaller LLMs, resulting in improved cross-lingual transfer on generative tasks. As model size increases, however, the number languages required decreases from three to two, indicating that larger models can learn this mapping easier. In both model sizes, adding more instruction tuning languages – including the target language itself – provides no significant gains.

Despite benefits from cross-lingual transfer, our results highlight a clear gap in performance between English and non-English target languages on all tasks. This indicates that small-scale multilingual instruction tuning is insufficient to improve the intrinsic multilingual capabilities of English-centric LLMs and that more advanced interventions may be required (e.g. Pfeiffer et al., 2020; ImaniGooghari et al., 2023).

Conclusion

In this paper, we investigated the use of multilingual instruction tuning to elicit multilingual capabilities of English-centric LLMs. Our results show that finetuning with as few as three languages can promote cross-lingual transfer and allow models to better exploit the relatively small amounts of non-English data seen during pretraining. Experiments conducted on four distinct tasks revealed that this can lead to significant performance improvements on open-ended generative tasks that assume input and output language agreement. Furthermore, we observed that these effects are amplified in larger models, requiring only two finetuning languages for successful multilingual instruction tuning. The effectiveness of cross-lingual transfer, which reduces the need for extensive multilingual instruction tuning, is good news for LLM developers. However, future work could explore other methods to encourage smaller models to better leverage the limited non-English data seen during pretraining and whether there are tasks for which language-specific instruction tuning is of higher importance.

Limitations

This work focused on detecting and evaluating non-English responses from English-centric LLMs. Our experiments demonstrated that cross-lingual transfer, which can be elicited given small amount of multilingual instruction tuning, can be effective for improvinging the non-English performance of English-centric models. That said, our analysis focused primarily on Indo-European languages, except for Chinese. As a consequence, it remains unclear how well our results generalise to more distant or topologically diverse languages, however, performance gains observed on Chinese suggest that generalisation is more dependent on exposure during pretraining than language similarity.

For our experiments and analysis on model training data, we relied on automatic language identification. To this end, we employed the OpenLID model from Burchell et al. (2023). Despite low error rates achieved by this model, language identification is not perfect and can lead to some texts being misidentified. To mitigate the risk of language contamination Blevins and Zettlemoyer (2022) in our controlled experiments, we only included training examples whose language is identified with a confidence threshold ≥\geq 0.8.

In §5.1 we investigated the impact of multilingual diversity versus training example diversity. While our findings revealed that there is no significant difference between these two settings, we note that even when finetuning with the original native non-English examples, task diversity may be inherently limited by design of the data collection. For instance, regardless of the language used, crowdworkers were asked to follow the same set of guidelineshttps://projects.laion.ai/Open-Assistant/docs/guides/guidelines when creating the data.

Acknowledgements

We sincerely thank our friends and colleagues, Fabian Aiolfi, Thea Bächler, Anastassia Shaitarova, Finnur Ágúst Ingimundarson, Cui Ding, Andrianos Michail, Janine Aeberhard, Arnisa Fazla, Farhad Nooralahzadeh, Omnia Ibrahim, Xuan Lam and Manasi Muglikar, for helping us with language-specific questions and validating translations used in our experiments, as well as Lena Bolliger and Patrick Haller for providing valuable feedback. Rico Sennrich acknowledges funding by the Swiss National Science Foundation (project MUTAMUR; no. 213976).

References

Appendix A Pretraining Data for English-centric LLMs

LLMs are pretrained on enormous amounts of unlabelled text gathered from existing corpora and the web. With the notable exception of BLOOM (Le Scao et al., 2023), very limited information is offered about the distribution of languages represented in datasets used to train LLMs and the subsequent performance on non-English languages. Table 1 provides an overview of the document-level language distributions of the LLMs used in this paper. For Llama 2, the information is taken from the original paper Touvron et al. (2023b), in which the authors analyse the training data using a fastText language classifier on corpus documents with a threshold of 0.5. For GPT-3, we use the official dataset statistics made available on Githubhttps://github.com/openai/gpt-3/blob/master/dataset_statistics/languages_by_document_count.csv which provide document-level language identification information. To gather statistics for Falcon, we inspected a sample of the RefinedWeb dataset Penedo et al. (2023a) that was constructed to train this family of models.https://huggingface.co/datasets/tiiuae/falcon-refinedweb Using the OpenLID fastText model from Burchell et al. (2023), we identify the most frequent document-level languages based on approximately 320 million examples from this corpus. Languages are identified with a confidence threshold of ≥\geq 0.6 and predictions below this threshold are aggregated under ‘unknown’.

Appendix B Translating AlpacaEval Prompts

To translate AlpacaEval prompts from English into each target language, we query GPT-3.5-Turbo using the template in Figure 8. To validate the translated prompts, we manually inspected a sample of the outputs in various languages. This inspection revealed that the translations were typically decent, although often included literal translations for metaphorical expressions rather than how a native speaker might express themselves. For instance, English ‘bullet points’ was translated literally into Hindi, rather than an arguably more appropriate phrasing such as ‘important points’. Additionally, in languages that distinguish between formal and informal or gendered pronouns (e.g., German, French, Hindi), the formal and male forms are dominant, which may be less representative of how native speakers actually interact with LLM chatbots.

Appendix C Direct vs. Translated Evaluation for Non-English Chat Responses

While using a powerful LLM to evaluate the outputs of other models has been shown to achieve reasonable agreement with human judgements in English (Zheng et al., 2023; Chiang and Lee, 2023), it is unclear whether this agreement transfers to all languages under investigation. Recent work by Hada et al. (2023) has shown that agreement between human and LLM judges tends to be lower for non-English languages, especially in the case of low-resource and non-Latin scripted languages, where the LLM judge tends to be overly optimistic in its assessment. However, for certain assessment criteria, such as linguistic acceptability and general content quality, they also confirm that inter-annotator agreement between LLM-based evaluators and humans is inline with that of multiple human annotators.

To investigate this potential bias, we compared scores assigned by the LLM judge on model-generated responses directly in each non-English target language and on their English translations. For each non-English prompt-response pair, we use GPT-3.5-Turbo (gpt-3.5-turbo-1106) and the prompt shown in Figure 8, specifying English as the target language. The resulting translated response is then paired with its corresponding AlpacaEval prompt in its original English form and presented to the LLM judge using the prompt in Figure 16.

Figure 9 shows that the distribution of assigned scores in the direct and translated evaluation settings is very similar for most languages. For languages that use non-Latin scripts (e.g., Chinese, Russian, Bulgarian), we observe that GPT-3.5-Turbo tends to assign slightly higher scores more frequently when evaluating directly on the non-English prompt/response pairs. This finding agrees with those from Hada et al. (2023) and indicates that LLM-based evaluations in non-English languages tend to be slightly overly optimistic and should be considered with caution. Nevertheless, we observe that the discrepancy between direct and translated evaluation is relatively minor for the languages considered and leads to negligible differences in the overall average helpfulness score. Because of this, and to keep evaluation costs to a minimum, we opt to use the direct evaluation strategy for our experiments.

Appendix D Translating Non-English Guanaco Training Examples into English

In order to investigate the effect of language diversity compared to instruction diversity, we translate a subset of Guanaco’s English training examples into non-English target languages and use these to create MT-based Multi-ii instruction-tuning datasets. By default, speaker roles in Guanaco are denoted with ‘### Human:’ and ‘### Assistant’. To ensure that these are never translated and the dialogue structure is maintained, we substitute them with special tokens ‘’ and ‘’ and explicitly tell the model to leave these tokens intact (see Figure 10). Before training, the special tokens back to their original form.

Appendix E Results with Falcon 7b

In order to assess whether our findings generalise to other LLMs, we repeat our experiments using Falcon 7b Almazrouei et al. (2023).

Figure 11 shows the helpfulness scores assigned by our LLM judge for Falcon 7b given incremental multilingual instruction tuning across all target languages. Similarly to our results with Llama 2 7b (see Figure 2), cross-lingual transfer is elicited after finetuning with relatively few languages, and no additional gains observed when including more than three languages. That said, Falcon 7b appears to show strong performance on French, even without multilingual finetuning, indicating that, despite being an English-centric model, it has strong capabilities in French out of the box. For Spanish, German and Chinese, performance is comparable to that of Llama 2 7b. However, for all other languages, responses are often ranked least helpful, indicating that Falcon 7b’s multilingual capabilities are limited strictly to major European languages and Chinese.

E.2 Extractive Question Answering

Figure 12 shows the results of Falcon 7b on XQuAD. While performance is generally lower than that of Llama 2 7b on this task (see Figure 3), we observe a similar, albeit weaker, effect of multilingual finetuning within the supported languages (Spanish, German, Chinese).

E.3 Commonsense Reasoning

Figure 13 shows the results of Falcon 7b on X-CSQA. Strikingly, in contrast to the results achieved with Llama 2 7b (Figure 4), Falcon 7b fails to score above random performance across all target languages. Regarding the effect of multilingual instruction tuning, we again see that it fails to deliver any performance improvements on this highly structured task.

E.4 XNLI

Figure 14 shows the results of Falcon 7b on XNLI. Similar to the results attained with Llama 2 7b (see Figure 5), we observe no significant differences in performance given different degrees of multilingual instruction tuning.

E.5 Discussion

Overall, these experiments indicate that Falcon’s multilingual capabilities are considerably narrower than Llama 2’s. We hypothesise that this may be a result of more stringent filtering of web-scraped pretraining data Penedo et al. (2023b), possibly resulting in less accidental contamination. Interestingly though, according to our estimated statistics based on the sample of the RefinedWeb corpus (Table 1), which makes up a large portion of Falcon’s pretraining data, Russian appears as frequently as Chinese and Swedish. However, we find that Russian texts are, on average, three times longer when tokenized with Falcon’s tokenizer compared to Llama 2’s. This indicates that, despite having a much larger vocabulary (65k for Falcon vs. 32k for Llama 2), Cyrillic-scripted languages such as Russian and Bulgarian are outside of the model’s intended domain and performance on these languages is inherently limited.