PolyLM: An Open Source Polyglot Large Language Model
Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, Tianxiang Hu, Shangjie Li, Binyuan Hui, Bowen Yu, Dayiheng Liu, Baosong Yang, Fei Huang, Jun Xie
Introduction
Large language models (LLMs) are trained on vast amounts of data in a self-supervised fashion, which has shown promising performance in a variety of zero-shot and few-shot tasks (Brown et al., 2020; Chowdhery et al., 2022). Fine-tuning these models on a diverse set of tasks allows them to handle unseen tasks following natural language instructions (Ouyang et al., 2022; Longpre et al., 2023; Taori et al., 2023; Anand et al., 2023). These properties have attracted significant attention from the Artificial Intelligence community and offering a potential path towards artificial general intelligence. Unfortunately, most LLMs are developed for English, such as LLaMA (Touvron et al., 2023), BLOOM (Scao et al., 2022), Chinchilla (Hoffmann et al., 2022), OPT (Zhang et al., 2022). A main reason stems from recent findings that model performance is closely related to the scale of the training dataset (Kaplan et al., 2020; Rae et al., 2021; Biderman et al., 2023; Touvron et al., 2023), leading to predominant focus on resource-rich languages, particularly English.
The relatively high concentration of studies on English limits the research and usage of LLMs in other languages. For instance, Thai and Indonesian have over 300 million (M) speakers, yet the size of these two languages in common crawl-based dataset such as mC4 (Xue et al., 2020) is only 80 billion (B) tokens, comprising a mere 3% of the English data. Due to the insufficient high-quality internet data, LLM capabilities on low-resource languages fail to be easily improved through expanding their data size like English (Kaplan et al., 2020; Rae et al., 2021; Biderman et al., 2023). As a result, existing open-source LLMs such as XGLM (Lin et al., 2022), BLOOM (Scao et al., 2022), and LLaMA (Touvron et al., 2023) perform relatively poor on these languages, some of which are entirely overlooked. It is crucial to explore multilingual LLMs to bridge this gap and achieve academic and social significance.
Our goal is to enhance the exploration and utilization of LLMs for non-native English speakers. In this work, we fill three significant gaps in this field: 1) the absence of an open-source multilingual LLM; 2) the inadequate availability of multilingual instruction data; and 3) the lack of a unified evaluation benchmark for multilingual settings.
Concretely, we first develop an open-source multilingual LLM from scratch, called Polyglot Large Language Model (PolyLM, Section 3). Contrary to existing open-source multilingual LLMs that lack 13B model, we release PolyLM-13B and PolyLM-1.7B to facilitate its usage. To construct PolyLM, we leverage a massive dataset of 640B tokens, culled from publicly available sources such as Wikipedia, mC4 (Xue et al., 2020), CC-100 (Conneau et al., 2019). This dataset contains over 30% of non-English languages, specifically covering 18 of the most commonly spoken languages.According to https://www.ethnologue.com/insights/most-spoken-language/. Some languages with interchangeable and more widely used official languages are not given priority, such as Hindi, Wu Chinese, and Cantonese. To alleviate the problem of insufficient data for low-resource languages, we propose a curriculum learning strategy. The training schedule increases the amount of data available for training in English during the initial phases, then ramping up the ratio of high-quality, low-resource languages as training progresses. We expect the method to enable the transfer of general knowledge from English to other languages, leading to significant improvements in overall performance.
In light of the supervised fine-tuning (SFT) stage, we construct a multilingual instruction dataset termed MultiAlpaca with 132,701 samples (Section 4). At present, there is a dearth of high-quality open-source multilingual SFT datasets. On the one hand, extant multilingual SFT datasets, e.g. xP3-MT (Muennighoff et al., 2022), are acquired via machine translation, which potentially yields a style of translationese, a lack of cultural nuances, as well as translation errors. On the other hands, manually annotating instructions is a laborious and costly process that does not lend itself well to the incorporation of creative flourishes. Drawing inspiration from recent advances in self-instruct (Wang et al., 2022; Taori et al., 2023), we devise a multilingual self-instruct method to automatically generate instruction data. Utilizing 175 English seeds as a starting point, our method leverage multilingual seed translation, instruction generation, and filtering mechanisms to deliver high quality multilingual instruction data.
In order to assess the multilingual capabilities of LLM, we curate a benchmark derived from existing multilingual tasks (Section 5.1), including QA (Clark et al., 2020), understanding (Conneau et al., 2019; Yang et al., 2019; Tikhonov & Ryabinin, 2021; Ponti et al., 2020), generation (Chen et al., 2021), and cross-lingual machine translation (Barrault et al., 2020). The benchmark is constructed with meticulously prompting and finally covers 10 tasks across 15 languages. Extensive experiments (Section 6) demonstrate that our pretrained model outperforms open-source models of comparable model size (e.g. BLOOM, LLaMA, etc.) in non-English languages. Through in-depth analyses, we identify finding that the proposed curriculum training strategy boosts the multilingual performance while maintain the English proficiency. In addition, the use of multilingual instruction data markedly enhances the ability of PolyLM to tackle multilingual zero-shot tasks.
Preliminary
In this section, we begin with a review of the background on language modeling. We then examine previous research on knowledge transferring, and instruction learning of pre-trained LLMs, with a focus on their relevance to PolyLM. Finally, we outline our rationale for training PolyLM.
Language Modeling refers to the process of estimating the probability of a sequence of tokens, i.e. . This is also commonly referred to as autoregressive sequence modeling, as it involves predicting the future token at each time-step based on the preceding context. The initial language models were predominantly -gram models that evaluate the likelihood of a sequence of tokens based on the frequency of its occurrence in a training corpus. Over the last two decades, neural networks have proven to be effective in the task of language modeling, including feed-forward models (Mikolov et al., 2010) and recurrent neural networks (Bengio et al., 2000). More recently, Transformer (Vaswani et al., 2017), a self-attention based neural network, has shown unparalleled language model performance (Devlin et al., 2019; Radford et al., 2018), and become the de facto backbone of LLMs emerged in the past three years, such as GPT3 (Brown et al., 2020), Gopher (Rae et al., 2021), PaLM (Anil et al., 2023), BLOOM (Scao et al., 2022), Chinchilla (Hoffmann et al., 2022), GLM (Zeng et al., 2022) and LLaMA (Touvron et al., 2023).
Transfer Learning is a rapidly evolving field of research that has garnered significant interest in recent years. In this scenario, models are initially trained on extensive unlabeled data, and then their acquired knowledge is applied to various downstream tasks through fine-tuning. Some of the most prominent works in this area include the ELMo (Peters et al., 2018), BERT (Devlin et al., 2019) and GPT (Radford et al., 2018) have demonstrated remarkable success. These developments subsequently prompt work (Raffel et al., 2020; Radford et al., 2019; Xue et al., 2020) on better results by adopting larger scale data and parameters to further improve model performance. Although pretraing-then-finetuning is still effective in achieving high performance with limited labeled data, recent advancements has shown that language models with extremely large scale parameters can perform tasks without further optimization. The most exemplary model is GPT3 (Brown et al., 2020), which utilizes a contextualized approach by incorporating multiple input-output demonstrations and presenting them alongside the query. This effectively stimulates the model to generate accurate predictions, showcasing encouraging outcomes in zero/few-shot situations.
Instruction Learning aims to bring together various natural language processing tasks by framing them as question-answering exercises that operate over a given context. This approach enhances the value of LLMs by leveraging their existing knowledge. With the success of language models, there has been a growing interest in exploring their potential to comprehend and execute instructions. Several advanced researches (Ouyang et al., 2022; Wei et al., 2022; Peng et al., 2023; Ye et al., 2023; Zhou et al., 2023) have demonstrated a remarkable ability to generalize to new zero-shot tasks. However, they rely heavily on human-generated instruction data, which is frequently constrained in terms of quantity, diversity, and creativity, which is very time-consuming and labor-intensive. Wang et al. (2022) make an effort to construct a self-Instruct framework for improving the instruction-following capabilities of LLMs. Similarly, Xu et al. (2023) propose an evol-instruct framework to automatically rewrite simple human-written instructions step by step into more complex ones, to further improve instruction-followed LLMs.
In this paper, we propose PolyLM to address the following blanks and limitations in current LLM research, offering a comprehensive and innovative solution to advance this field.
We provide a 13B scale model that is proficient in the major non-English languages spoken worldwide, such as Spanish, Russian, Arabic, Japanese, Korean, Thai, Indonesian, and Chinese etc. It is a perfect complement to the existing open-source models, including: (1) LLaMA, English is predominant among the whole dataset. (2) BLOOM, lack of 13B version and fail to address languages spoken by significant populations, such as Japanese, Korean and Thai. (3) XGLM (Lin et al., 2022), the maximum version is 7B. (4) mGPT (Shliazhko et al., 2022), only 1.3B version is available.
We suggest an advanced curriculum learning approach that facilitates the transfer of commonsense knowledge, acquired mainly in English, to diverse non-English languages and specific NLP downstream tasks such as machine translation.
We propose MultiAlpaca to complement Alpaca (Taori et al., 2023) and Chinese-Alpaca (Cui et al., 2023), making LLMs better follow multilingual instructions, particularly those coming from non-native English speakers.
PolyLM: a polyglot large language model
In this section, we present the design of PolyLM, which includes a detailed description of its training dataset (Section 3.1), architecture (Section 3.2), and training process (Section 3.3).
The composition of the pre-training dataset used for PolyLM is shown in Table 1. Our pre-training dataset contains 640B tokens in total, of which English data accounts for 68%. To develop PolyLM with multilingual capabilities, the pre-training dataset has about 32% non-English multilingual data, which is a higher percentage of non-English data than most previous open-sourced large language models (Biderman et al., 2023; Zhang et al., 2022; Touvron et al., 2023; Penedo et al., 2023). To be concrete, the English data contains documents with 425B tokens from multiple sources, such as The Pile (Gao et al., 2020), mC4 (Xue et al., 2020), and Wikipedia. While the 204B multilingual data tokens come from CC-100 (Conneau et al., 2019), mC4 (Xue et al., 2020), Wikipedia. The multilingual data mainly covers the following languages: zh, ar, es, fr, de, it, nl, ru, id, pl, pt, ja, th, tr, he, ko, vi, with the distribution given in Table 2. To enable the model ability of code understanding and generation, we also incorporate code data of 7.5B tokens from GitHub with permissioned licenses into our pre-training dataset. In order to further improve the cross-lingual and multilingual ability of the PolyLM, similar to PaLM2 (Anil et al., 2023), we employ parallel multilingual data of 1B tokens into our pre-training dataset.
To build the pre-training dataset, we also develop a comprehensive data pre-processing pipeline that implements multiple techniques for data cleaning and filtering. The pipeline consists of the following stages:
1) Language identification. We classify documents according to their primary languages and remove those with low confidence in classification, leveraging inexpensive n-gram models (e.g., fastText (Joulin et al., 2016)).
2) Rule-based filtering. Following Rae et al. (2021); Scao et al. (2022), we eliminate irrelevant or low-quality content using various rules and heuristics, including repetition removal (the document with the excessive line, paragraph, or n-gram repetitions is removed), document-wise filtering (removing outlier documents by overall length, symbol-to-word ratio, the ratio of ellipsis, invisible characters, numbers, and dates, etc.), and line-wise corrections (such as URL filtering, long words removal, and whitespace standardization).
3) ML-based quality filtering. We further filter low-quality multilingual documents using several small n-gram-based language models (e.g., KenLM (Heafield, 2011)) for different languages trained on their gold-standard corpora. In addition, similar to Raffel et al. (2020); Smith et al. (2022), we also train a 2-gram fastText (Joulin et al., 2016) classifier to filter the low-quality English documents. This classifier uses Wikipedia, and Books from The Pile (Gao et al., 2020) as the positive samples and CommonCrawl web documents as the negative samples. To sum up, about 28.3% data are filtered with Rule-based filtering and ML-based quality filtering.
4) Deduplication. In line with Raffel et al. (2020), we remove similar documents to reduce data redundancy with MinHashLSH-based fuzzy deduplication technology, where 23.1% English documents and 18.6% non-English documents are removed.
Based on the PolyLM multilingual pre-training dataset, we derived a vocabulary with 256K token entries using Byte-Pair Encoding (BPE) (Sennrich et al., 2015) with the implementation from SentencePiece (Kudo & Richardson, 2018). To enhance the mathematical capabilities of our model, we follow Touvron et al. (2023) to split all numbers into individual digits. The unknown characters are fallback to byte encoding of UTF-8 to guarantee the coverage of rare words (e.g., emoji, and special symbols). For tokenizer training, we sample multilingual documents with a similar distribution as Conneau et al. (2019) used to increase the number of vocabulary tokens associated with low-resource languages and alleviate the bias towards high-resource languages. We compare the compression rate on different language corpora of different tokenizers. We use XLM-R (Conneau et al., 2019) tokenizer, which supports 100 languages, as the baseline (the compression rate of XLM-R tokenizer is set to 1). As shown in Figure 1, PolyLM has achieved significantly better compression rates in most covered languages, while maintaining the compression rate in English as BLOOM (Scao et al., 2022), LLaMA (Touvron et al., 2023), GPT-2 (Radford et al., 2019), and GPT-4 (OpenAI, 2023). Note that some open source models that are not friendly to language extensions, for example, LLaMA (Touvron et al., 2023) only contain a 32K size vocabulary mostly composed of English tokens, which is not friendly to non-Latin languages. When improving a certain non-Latin language ability, the vocabulary needs to be expanded like Chinese-LLaMA (Cui et al., 2023). On the contrary, PolyLM allows researchers to improve the model’s ability in a covered language by simply continuing monolingual pre-training without expanding the vocabulary.
2 Architecture
It has become apparent that the computational cost of exploring different architectural designs for LLMs is prohibitive. Therefore, we present the distinctive design options of PolyLMRecent research indicates that Rotary Position Encoding (RoPE) (Su et al., 2021) yields superior performance. Accordingly, we will switch to the latest Megatron-LM branch and promptly release 13B and 1.7B versions featuring RoPE. in this section.
Following some endeavours on large language models, we develop a decoder-only autoregressive Transformer architecture detailed in Radford et al. (2019). To stabilize the training, we adopt Pre-LN (Xiong et al., 2020), i.e. (where indicates the layer function) for layer normalization, and apply the Xavier normal initialization (Glorot & Bengio, 2010) with bias terms are initialized to zero. To improve FFNs in Transformer, we replace ReLU with GeLU activation (Hendrycks & Gimpel, 2016).
In this paper we present two Transformer language models with 1.7 billion and 13 billion parameters, respectively. The architectural details are displayed in Table 3.
3 Training
We train all models with a 2048 token context window, using the Adam (, ) optimizer. We warm-up the learning rate from to the maximum learning rate over the first 2000 steps, and then decay it to 10% of the maximal learning rate using a cosine schedule. We use a weight decay of 0.1 and gradient clipping of 1.0.
PolyLM was trained using Megatron-LM https://github.com/NVIDIA/Megatron-LM on a cluster of 32 A100 GPU (880G) servers. We apply tensor model parallelism within a single node, setting tensor-model-parallel-size as 8. When training a 13B-parameter model, our code processes around 1170 tokens/sec/GPU, thus training over our dataset containing 640B tokens takes approximately 29 days. However, we faced numerous unforeseen spikes and deviations in losses, which prolonged the entire training process to a duration of two months. There are several possible conditions that result in training collapses, and our unique choices to enhance training stability.
Lower Maximal Learning Rate. Learning rate is an important hyperparameter in neural network models that controls the magnitude of parameter updates. In our first few attempts, we drew inspiration from previous research which indicated that smaller models tend to benefit from higher learning rates. As such, we opted to set the learning rate to . Without exception, all attempts to train PolyLM-13B have resulted in loss spikes with this choice in early stage, which tend to occur more frequently as the training progresses, as illustrated in Figure 2(a). We have noticed that the gradient norm shows significant fluctuations during the warm-up phase, when the learning rate is increasing linearly (see Figure 2(b)).
The fundamental issue with instability during training is that a large learning rate can cause the gradient to grow too large, surpassing the model’s capacity and resulting in a gradient explosion that prevents parameter updates. The problem is handled via reducing learning rate to , i.e. a proper learning rate located before the step where the initial spike in loss occurs (Cf. Figure 2(c)).
Mixed-Precision. Despite the potential instabilities associated with training models using half precision (float16) activations and model parameters that arise from the limited numerical range, it has been proposed that the numbers represented by bfloat16 allow for training of models and can avoid performance degradation compared to full float32 training. Thus, we incorporate the bfloat16 numerical format to reduce memory and increase training efficiency. However, similar to OPT-175B (Zhang et al., 2022), BLOOM-176B (Scao et al., 2022) and GLM-130B (Zeng et al., 2022), the training of PolyLM-13B still faces frequent loss spikes while lowering learning rate. We attempted to address such challenge via manually skipping data and restart the straining, it unfortunately tends to become increasingly severe as the training does on (Cf. Figure 3(a)).
After conducting two weeks of investigation, we have come to the realization that the instabilities we are encountering may not be due to the training data under the mutlilingual scenario (with the vocabulary up to 256,000), but rather due to the model itself. Specifically, we suspect that there could be a risk of overflow in the attention or residual connectivity layers. Taking this into account, we have configured the residual connection and attention layers to have a numerical precision of float32 to ensure optimal performance, resulting in a highly stable training process (Cf. Figure 3(b)).
Curriculum Learning. Optimizing LLMs to learn knowledge encoded in multiple languages simultaneously is a significant challenge. We concretely formulate this problem as transferring general knowledge to low-resource languages while maintaining the advantage of high-resource language in the model. To address this issue, we adopt a curriculum learning strategy (Bengio et al., 2009; Kumar et al., 2010; Jaegle et al., 2021) that ramps up the ratio of high-quality and low-resource languages during training. Specifically, the training process is divided into two stages. In the first stage, we use the whole pre-training dataset to train a base model yields commonsense generalization ability. In the second stage, we transition to a subset of the pre-training dataset that boasts superior quality and a greater proportion of multilingual content, to further strengthen the model’s multilingual capabilities. Figure 4 compares the language distribution of training data in two stages, indicating that the proportion of most low-resource languages has been increased in the sub-dataset.
To build the sub-dataset for curriculum learning, we first manually evaluate the quality of publicly available data source in the pre-training dataset, and sample about 97B tokens from the high-quality sources while increasing the proportion of languages other than Chinese and English. We also enhance the proportion of parallel data (OPUS) to facilitate the modeling of cross-lingual representation. The detail of the sub-dataset are illustrated in Figure 5. According to our established setup, the curriculum training process is highly stable (Cf. Figure 3(c)).
MultiAlpaca: A Multilingual Self-Instruction Dataset
Fine-tuning LLMs with instruction-based tasks has been proven effective in practice (Ouyang et al., 2022; Wei et al., 2022; Peng et al., 2023; Ye et al., 2023). By providing accurate task instructions during the SFT phase, LLMs can not only learn to understand the requirements of each task via the instruction part, but also show extensive abilities to cope with other types of tasks which are even unseen during training (Wei et al., 2022). Nevertheless, tuning multilingual LLMs is still troubled by the scarcity of current SFT datasets. On the one hand, most instruction-based datasets are mainly in resource-rich languages (e.g., English or Chinese). To the best of our knowledge, there is currently no high-quality multilingual instruction-based SFT dataset for LLM training. On the other hand, most instructions are manufactured by experienced language speakers (e.g., Wei et al., 2022). Although the quality of instructions is well preserved, the amount of tasks is rather scarce for fine-tuning LLMs.
To overcome these two drawbacks, we determine to extend the generality of our proposed PolyLM via creating a multilingual SFT dataset – MultiAlpaca (Figure 6). Following the self-instruct paradigm proposed by recent studies (Wang et al., 2022; Taori et al., 2023), we query the available LLM for responses, iteratively collecting and filtering self-instruct examples to build our dataset. MultiAlpaca delivers comprehensive support on multilingualism, covering up to 11 languages including Arabic (Ar), German (De), Spanish (Es), French (Fr), Indonesian (Id), Japanese (Ja), Korean (Ko), Portuguese (Pt), Russian (Ru), Thai (Th), and Vietnamese (Vi). For each language, the number of tasks in MultiAlpaca varies from 9,515 to 14,671, yielding 132,701 tasks in total.
We first form the format of our tasks by referring to Taori et al. (2023), where each task contains three parts: 1) “instruction” describes the requirements of the corresponding task; 2) “input” can complement the “instruction” to a complete question; and 3) “output” is a correct answer of the question. We notice that, Taori et al. (2023) constructed their dataset where each “instruction” can be equipped with multiple “input-output” instances. For simplicity, we only assign each “instruction” with one “input-output” instance.
2 MultiAlpaca Construction
As shown in Figure 7, we construct the MultiAlpaca dataset based on the following steps:See Appendix A for more details.
We first obtain 175 seed tasks from Taori et al. (2023) to construct the multilingual ones for MultiAlpaca. After manually checking them, we remove the cases where answering the questions requires cultural backgrounds (e.g., idiom explanation, character-level riddle, and lyrics generation). Then, we marked the cases whose original “input” or “output” should be reserved (e.g., single-choice question, translation, bias identification, and code generation), where those tasks will directly use the original “input” or “output” across different languages for MultiAlpaca. Finally, we filter out 13 inappropriate seed tasks, and modified 23 ones marked due to the reuse of “input” or “output” parts. We translate the remaining 162 tasks into the other 11 languages, yielding multilingual seed tasks for each language.
Iterative Progress
We manage the MultiAlpaca dataset construction progress as an iterative one with multiple rounds. For each round, we manage the following five substeps in order:
Prompt Construction We follow Taori et al. (2023) to construct the prompts for MultiAlpaca when querying LLM for completion. When handling each involved language, for each prompt, we sample two seed tasks and one MultiAlpaca task as the demonstrations, and guide the LLM to complete the other 17 tasks in the response. For each round, we construct 100 prompts for querying the completion by LLM.Except for the first round where the task pool is empty, we arrange 10 prompts for completion due to the small number of available tasks for demonstrations.
Response Collection We collect the responses from ChatGPT via the OpenAI API service. The model we use is “gpt-3.5-turbo-0301”, which supports the processing of tokens up to 4,096.
Format Checking When checking the format, we first remove the last task if the response is stopped due to the exceeding of max sequence length. Then, we use the pre-defined task format to help split the response string, so as to make sure each of the tasks contains “instruction”, “input”, and “output” parts.
Similarity Checking After that, to preserve the diversity of MultiAlpaca dataset, we further check the similarity between the tasks that are newly collected and those from the task pool. Following Taori et al. (2023), we compute the Rouge-L F-scores between the instruction of each newly collected task and those of all collected ones. For each newly collected task, it would be added to the task pool only if all the scores are lower than 0.7.
Task Pool Updating In the end, we update the task pool by adding the newly collected tasks, and arrange the next round for collecting MultiAlpaca self-instruct tasks.
MultiAlpaca Dataset Export
Totally, we arrange 10 rounds in the iterative progress when constructing the MultiAlpaca dataset. We export all tasks from the task pool as the MultiAlpaca dataset for SFT learning.
Multilingual Benchmark
We aim to assess the capabilities of PolyLM from various perspectives: 1) the ability of large language models (LLMs) to understand and generate natural languages, as well as the ability to grasp world knowledge; 2) the performance of LLMs across different languages; and 3) their capacity to handle cross-lingual tasks. Following the experiment design of previous work (Scao et al., 2022; Ahuja et al., 2023), we gather a subset of datasets from previous NLP tasks to construct a multilingual benchmark. The brief statistics of all datasets in the benchmark can be found in Table 4. The details of how we frame all the tasks with prompting are listed in Appendix B.
All the datasets in the above multilingual benchmark can be divided into four groups: Natural Language Understanding, Knowledge, Natural Language Generation and Machine Translation. The details of each dataset that we use for benchmarking are given below.
To assess the comprehension capability of large models across various languages, we collect the multilingual versions of datasets from seberal wide-used NLP benchmarks (Wang et al., 2018; 2019).
XNLI (Conneau et al., 2019) serves as a benchmark to evaluate a model’s proficiency in predicting textual entailment. The task entails the evaluation of whether two given sentences, A and B, convey the same meaning, are contradictory, or are unrelated. The dataset has been professionally translated into 14 languages from the original English XNLI dataset.
PAWS-X (Yang et al., 2019) is a benchmark to evaluate the model’s ability to judge whether one sentence is the paraphrase of another. It is professionally translated from the PAWS (Zhang et al., 2019) dataset into 6 diverse languages.
XWinograd (Tikhonov & Ryabinin, 2021) serves as a benchmark to measure a model’s common sense reasoning ability. Specifically, the task entails presenting the model with a brief contextual passage and requiring it to select the accurate term from a set of two options for a pronoun in the passage.
XCOPA (Ponti et al., 2020) is another benchmark intended to assess the proficiency of models in commonsense reasoning across languages. The dataset comprises translations and re-annotations of the English COPA Gordon et al. (2011), spanning 11 languages around the globe. Based on the given premise and prompt, the task is to choose the more plausible response between two answer choices that can be inferred from the premise.
TyDi QA (Clark et al., 2020) is a question-answering dataset covering 11 typologically diverse languages with 200K question-answer pairs. We use this dataset to evaluate the ability to grasp knowledge from natural text. Unlike previous datasets such as MLQA (Lewis et al., 2020) and MKQA (Longpre et al., 2020), this dataset is collected directly in each language without the use of translation. We select 5 languages out of 11 that are included in the pretraining corpora of PolyLM. Following the PaLM (Chowdhery et al., 2022), we evaluate models on the Gold passage task, which requires answering questions based on a passage that is guaranteed to contain the answer.
MTG (Chen et al., 2021) is used to assess the efficacy of large language models in generating longer responses across diverse usage scenarios and multiple languages. MTG covers four different generation tasks: Story Ending Generation (SG), Title Generation (TG), Question Generation (QG), and Summarization (Summ). The datasets are originally written in English, subsequently extended into four other languages (German, French, Spanish, and Chinese) through the use of machine translation and human annotation. The effectiveness of LLM-generated responses is evaluated using the average of Rouge1, Rouge2, and RougeL.
WMT20 (Barrault et al., 2020) is used to study the cross-lingual proficiency of large language models in accomplishing translation tasks, as the process of translation entails both comprehending the semantic of the input in one language and expressing it in another. We select translation tasks between English and each of the following languages as benchmark languages: German, Japanese, Russian, and Chinese. The results are evaluated using the SacreBLEU (Post, 2018) and the scores for BLEU (Papineni et al., 2002) on the test set are reported.
2 Evaluation Design
For metric evaluation, the tasks included in our benchmark can be divided into two categories: classification-style tasks and generation-style tasks.
Classification-style tasks require selecting the correct option from several options, such as the XNLI dataset. To evaluate these tasks, following the way in Gao et al. (2021), we design the problem in the form of a cloze test, where each option is filled in to construct a complete sentence. We then choose the correct answer by separately calculating the log-likelihood of each completed sentence and selecting the one with the highest value.
Generation-style tasks, such as machine translation, require generating answers with several natural sentences. For these tasks, we adopt greedy decoding for deterministic results. Considering the efficiency of decoding, we restrict the maximum number of generated tokens to 256. For foundation models, we choose the result before the first ‘\n’ as the answer, while for models that have undergone instruction tuning, we decode until the EOS token appears.
In evaluating foundation models, considering that models have not been able to understand instructions, we adopt in-context learning (Brown et al., 2020) to evaluate the model for generation-style tasks. We generally choose no more than five examples due to the model’s context window limitation. For tasks that have well-divided training/development sets, we randomly draw examples from them for each test sample. Otherwise, we draw examples randomly from the test sets except for the current sample.
Experiments
In this section, we provide separate comparison results for the pre-training and SFT models. Then, we analyze the effectiveness of our model in three aspects: curriculum learning, multilingual instruction finetuning, and the scaling for model size.
For the pre-trained models, we selected two mainstream open-source models as our baselines.
LLaMA (Touvron et al., 2023) is a pre-trained language model released by MetaAI, which includes 7B, 13B, 30B, and 65B versions. The pre-training dataset is sourced from publicly available corpora. The 33B and 65B models are trained on 1.4 T tokens, while the 7B and 13B models are trained on 1 T tokens. To ensure an equal parameter count comparison with PolyLM, we mainly take the 13B version into consideration.
BLOOM (Scao et al., 2022) is a multilingual model that covers 46 natural languages and 13 programming languages with a maximum of 176B parameters. Since BLOOM has not released a 13B version, we opt for the BLOOM-7.1B model as our baseline.
We evaluate PolyLM across various multilingual tasks, covering natural language understanding (NLU), knowledge, natural language generation (NLG) and machine translation (MT). To make a clearer comparison of the multilingual capabilities of different models, we present the results using radar charts, with detailed results available in the C.
Natural Language Understanding. Figure 8 shows the results on four NLU tasks under the zero-shot setting. PolyLM-13B shows comparable performance to the English-centric LLaMA-13B model in the English scenario. Moreover, it yields substantial improvements of 7.2% and 19.1% on PAWS-X and XNLI respectively. For languages other than English (the multilingual column), PolyLM-13B outperforms LLaMA-13B with average improvement up to 7.6%, 5.6%, 3%, and 11% on XCOPA, PAWS-X, XWinagrad, and XNLI, respectively. When compared to the multilingual language model BLOOM-7.1B, PolyLM-13B outperforms with an average improvement of 4.2%, 4.1%, 3.4%, and 4% points on the respective tasks. This improvement can be attributed to the higher percent of multilingual text during pre-training and curriculum learning strategy.
Knowledge. We evaluate our model on grasping multilingual knowledge by using the TyDiQA benchmark in the one-shot setting. Upon careful analysis of Figure 9(a), it is evident that BLOOM-7.1B experiences significant performance drops in the Korean (ko) and Russian (ru) language directions, whereas LLaMA-13B and PolyLM-13B exhibit better balance across all five languages. Furthermore, PolyLM-13B has an additional advantage of an average 1.2-point lead over LLaMA-13B.
Natural Language Generation. Figure 9(b) displays the Rouge scores of four diverse NLG tasks in multilingual settings. From a multilingual perspective, PolyLM-13B outperforms all other models across four languages, namely Chinese (zh), Spanish (es), French (fr), and German (de). Moreover, in terms of task types, PolyLM-13B performs the best in question generation (QG) and summarization (Sum) tasks, while also showing comparable performance to the best model LLaMA-13B in the text generation (TG) task. Across all MTG tasks and languages, PolyLM-13B has an average score advantage of 1.6 and 2.3 compared to LLaMA-13B and BLOOM-7.1B, respectively.
Machine Translation We focus on evaluating the translation performance on four typologically diverse languages from WMT20 datasets, including translation directions both from and to English. Results of Figure 9(c) show that PolyLM-13B achieves similar performance to LLaMA-13B in the multilingual to English directions and surpasses LLaMA-13B and BLOOM-7.1B with average BLEU scores of 5.4 and 15.8 in the English to multilingual directions.
2 Comparisons between Instruction-followed Models
This section focuses on evaluating the effectiveness of instruction-followed models founded on the pre-trained language models discussed in Section 6.1. We conduct a comparative analysis of PolyLM-MultiAlpaca-13B that is fine-tuned on PolyLM-13B using MultiAlpaca, against two other publicly available models:
BLOOMZ-MT-7B is initially pre-trained on BLOOM-7B, and later fine-tuned on the multilingual task mixture xP3-MT (Muennighoff et al., 2022).
LLaMA-Alpaca-13B is built based on the pre-trained model LLaMA-13B and fine-tuned on the English self-instruction dataset Alpaca (Taori et al., 2023).
Figure 10 and 11 present the performance comparisons of instruction-followed models with the zero-shot setting, considering various tasks and language directions. The results indicate that PolyLM-MultiAlpaca-13B is comparable or superior to LLaMA-Alpaca-13B on all English tasks, although the latter is primarily trained on English-only instructions. On other non-English tasks, PolyLM-MultiAlpaca-13B significantly outperforms LLaMA-Alpaca-13B. This superiority can be attributed to the inclusion of more well-balanced multilingual datasets during the pre-training and instruction fine-tuning. In comparison to BLOOMZ-MT-7B, PolyLM-MultiAlpaca-13B has demonstrated consistent improvements across all tasks and languages. We have observed an outlier MTG, and we speculate that this may be due to the fact that MTG testsets are part of the xP3 dataset. We plan to refine our instruction tuning process for PolyLM by utilizing the xP3 dataset in order to delve deeper into this inconsistency.
Note that it is not feasible to fully assess the effectiveness of the model’s performance through downstream NLP tasks after instruction fine-tuning. Therefore, we have presented selected examples for qualitative analysis, which are fully outlined in Appendix D.
3 Analysis
We validate the effectiveness of the curriculum learning strategy in NLU and MT tasks of multilingual benchmark (Section 5.1) by comparing the following variants:
(1) w/o CL PolyLM-13B trained without curriculum learning, which is only optimized in pretrained dataset.
(2) w/ CL PolyLM-13B trained with curriculum learning, using about 100B high-quality multilingual data selected from the pretrained dataset.
Please note that we only focus on the languages included during curriculum learning. Referring to Figure 12, the model with curriculum learning has achieved stable progress in mainly all languages in both NLU and MT tasks. First of all, the model performance is enhanced in most low-resource languages, indicating that the general knowledge can be effectively transferred to these languages through raising data proportion. Additionally, the model retains its superior performance in English, which illustrates that improving data quality for high-resource languages can achieve competitive results to training with larger amounts of data. Finally, it is worth noting that introducing more multilingual parallel data during the curriculum learning significantly boost the model performance on translation task.
Multilingual Self-instruction.
Here we highlight the advantages of MultiAlpaca over English-only Alpaca (Taori et al., 2023), particularly in cross-lingual tasks (i.e., machine translation). As illustrated in Table 5, compared to the model fine-tuned only using Alpaca, PolyLM-MultiAlpaca-13B exhibits substantial improvements in TyDiQA and multiple WMT20 translation tasks, with enhancements of +10 BLEU and +1.4% F1. These results suggest that MultiAlpaca is capable of simulating the cross-lingual alignment ability of the foundational, as well as facilitating the comprehension of multilingual instructions.
Scaling for Model Size.
In addition to the 13B model, we also release a smaller 1.7B model. Recent studies highlight the critical role of model size in the performance of large language models (LLMs), with much of this work focusing on English (Kaplan et al., 2020; Rae et al., 2021; Biderman et al., 2023; Touvron et al., 2023). In this section, we present results for PolyLM-13B and PolyLM-1.7B to investigate the impact of model size on multilingual abilities. Consistent with the aforementioned experimental setup for the validation of base model, we compare the two models using a one-shot setting. As illustrated in Figure 13, the 13B model significantly outperforms the 1.7B model across all compared multilingual tasks. We posit that multilingual problems are more complex than their monolingual counterparts and may depend more heavily on the model’s throughput. Moving forward, we plan to release additional models of varying sizes, with the ultimate goal of refining the scaling law for multilingualism.
Conclusion
Multilingualism poses an inevitable challenge for LLM due to the scarcity of resources. In this work, we release PolyLM – a new multilingual LLM, alone with MultiAlpaca – a multilingual instruction dataset, and a multilingual benchmark. Quantitative and qualitative analyses demonstrate the superiority of PolyLM over open-source models in non-English languages. We find that incorporating curriculum learning strategy can boost the performance of LLM on non-English languages, without impeding its English proficiency. In addition, fine-tuning LLM with multilingual instruction data can considerably improve zero-shot performance on these languages.
There is still ample opportunity for refinement in our work endeavors. For instance, while we briefly assess the model’s capacity to comprehend multilingual instructions, there is potential for further optimization through the amalgamation of data sources (Wang et al., 2023; Longpre et al., 2023), evolutionary methods (Xu et al., 2023) and diversification strategies (Zhou et al., 2023). Moreover, in our current version, we adopt absolute position encoding, which adheres to the early default configuration in Megatron toolkit (Shoeybi et al., 2020). Future iterations should incorporate techniques that facilitate the expansion of window size, such as rotary position encoding (Su et al., 2021; Chen et al., 2023) or ALiBi (Press et al., 2022).
Language serves as a conduit for culture, and the unique contributions of various languages enrich and diversify our global community. Nevertheless, the advancement of LLM may inadvertently amplify the influence of prominent languages and present a formidable obstacle for low-resource languages. In light of these concerns, we aspire that our research will motivate further inquiry and innovation in the field of multilingual LLM.
Ethics Statement
In this paper, we propose PolyLM, an LLM which offers a wider support on non-English languages. Our contributions are fully methodological: adding the support of multilingualism to LLM during training and SFT phases. However, when building our PolyLM model, it is unavoidable that our PolyLM might exhibit several common deficiencies of language models, e.g., hallucination and toxicity. Specifically, as the collected MultiAlpaca dataset are generated by ChatGPT, the pseudo tasks might give inappropriate pseudo tasks which are hardly filtered out, e.g., hallucinated reasoning and anti-fact statements (Brown et al., 2020; OpenAI, 2023). Besides, PolyLM may deliver toxic texts, which might be gender- or race-biased like other existing LLMs (Taori et al., 2023; Cui et al., 2023).
Despite the ethical concerns above, we think that those problems are of vital importance to the AI community to study the deficiencies of LLMs. We recommend that the users of PolyLM and MultiAlpaca deploy our released materials only for research proposals. Besides, we suggest the users better identify the deficiencies of those contents, and welcome the following researchers to facilitate further research on the alignment between the LLM outputs and human values with PolyLM and MultiAlpaca materials.
References
Appendix A Detailed Setting for MultiAlpacaDataset Construction
We show the used prompt when constructing MultiAlpaca dataset in Table 6. We mainly refer to Taori et al. (2023), and adopt our prompt to multilingual scenarios after minor revisions. Briefly, in the prompt, we list several requirements of the self-instruct tasks in the prompt, i.e., the used language, the format, the diversity, and the lengths of tasks within each single response. We also add three demonstrations to help the model generate the tasks which follow the pre-defined format.
A.2 Format and Similarity Checking
After collecting the pseudo tasks for each language, we first remove the cases which contain website links. Then, we tokenize the “instruction”, “input”, and “output” with available tokenization toolkits.nltk.tokenize.wordpunct_tokenize for Ar, De, Es, Fr, Id, Pt, Ru, and Vi; kytea for Ja; Okt for Ko; thai_tokenizer for Th.
Besides, we found that some of the tasks may give redundant information within both “instruction” and “input” part. We revise those examples to make them available as much as possible. In detail, if the “input” part can be a sub-string of the “instruction”, we mark the “input” part as an empty one (using the placeholder “
We show the percentage of remaining examples after similarity checking in Figure 14. For each language, we show the number of self-instruct tasks in MultiAlpaca in Table 7.
Appendix B Details of Task Formatting
We list the detailed format of all the tasks in our multilingual benchmark in the following tables.