Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, William Marshall, Gurpreet Gosal, Cynthia Liu, Zhiming Chen, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, Lalit Pradhan, Zain Muhammad Mujahid, Massa Baali, Xudong Han, Sondos Mahmoud Bsharat, Alham Fikri Aji, Zhiqiang Shen, Zhengzhong Liu, Natalia Vassilieva, Joel Hestness, Andy Hock, Andrew Feldman, Jonathan Lee, Andrew Jackson, Hector Xuguang Ren, Preslav Nakov, Timothy Baldwin, Eric Xing
Introduction
Large language models (LLMs) have revolutionized the field of natural language processing (NLP), demonstrating remarkable capabilities in generating high-quality texts and resulting in widespread adoption across a diverse array of practical NLP applications and domains. Yet, the main focus of research and development efforts so far has been on English. While recent LLMs such as Falcon [AAA+23], PALM [CND+22] and LLaMA [TLI+23, TMS+23], among others, are able to process data in multiple languages, they were nevertheless primarily trained and instruction-tuned for English. As a result, they are not able to extend their understanding and generation capabilities to languages other than English. In this work, we aim to bridge this gap. We focus on Arabic, one of the world’s most spoken languages with over 400M speakers, which has been noticeably underrepresented in the LLM space so far. In particular, we develop Jais, a powerful Arabic-centric decoder-only LLM with 13B parameters, based on the GPT-3 generative pretraining architecture [BMR+20].
The primary challenge in developing an Arabic LLM is the limited availability of high-quality Arabic data. As compared to English, where corpora of size up to two trillion tokens are readily available [TMS+23], Arabic corpora are significantly smaller in size. As part of this work, we have collected the largest Arabic corpora to date, consisting of 72 billion tokens. However, this dataset is still not sufficiently large for the purposes of training an Arabic LLM capable of demonstrating emergent capabilities [Ope23].
To address this, we train bilingual models, by augmenting the limited Arabic pretraining data with abundant English pretraining data. We pretrain Jais on 395 billion tokens, including 72 billion Arabic tokens (which we repeat 1.6 times, to obtain an effective total of 116 billion Arabic tokens), 232 billion English tokens, and the remainder being code in various programming languages. As part of our effort, we have designed and developed a specialized Arabic text processing pipeline that includes thorough data filtering and cleaning to produce high-quality Arabic data.
Unlike previous massively multilingual LLMs such as BLOOM [SFA+23] or mT0 [MWS+23], which contain more than 50 languages, we do not include languages aside from Arabic and English in any significant percentage. Neither do we relegate Arabic to a minority in the pretraining dataset. Instead, Arabic data constitutes 33% of our pretraining. Our choice of mixing two languages attains the best of both worlds; the LLM is highly fluent in Arabic, with linguistic capability as well as cultural awareness and sensitivity. At the same time, it is on par with recent English LLMs in terms of reasoning capacity and world knowledge, capabilities we observe to have transferred from English to Arabic and vice-versa.
Building upon the standard transformer architecture [VUWS22] in the form of its GPT-3 variant, we adopt a number of improvements from the literature including (i) ALiBi [PSL22] positional encodings, which enable the model to extrapolate to longer contexts at inference, (ii) SwiGLU activation function [Sha20] to improve the performance, (iii) maximal update parametrization to perform hyperparameter optimization based on experiments with smaller models [YHB+21], and (iv) a custom-built tokenizer that weighs both languages equally.
We further develop an instruction-tuned version of our model, Jais-chat, which uses over 3.6 million Arabic and 6 million English instruction-response pairs. Considering the inherent safety concerns of LLMs, we further fine-tune it with safety-oriented instructions. In our deployed system which provides an interactive interface to the instruction-tuned model https://arabic-gpt.ai, we add extra guardrails in the form of safety prompts, keyword-based filtering, and external classifiers. An example conversation with Jais-chat on this interface is shown in Figure 1.
We evaluate Jais and Jais-chat across a wide array of Arabic and English NLP benchmarks, addressing reasoning, knowledge, misinformation, and bias. The results show that Jais is superior in Arabic compared to other models of similar size, while also being competitive in English, despite being trained on significantly less English data.
Jaishttps://huggingface.co/inception-mbzuai/jais-13b: base pretrained 13B foundation model;
Jais-chathttps://huggingface.co/inception-mbzuai/jais-13b-chat: instruction-tuned 13B version of Jais, optimized for dialog interaction.
By making our models publicly available, we hope to enable further research and development in this area, stimulating innovation and practical applications that can better serve the Arabic and the global communities. Despite our significant efforts to ensure safety, we recognize that the models are not foolproof and may not cover all cases. Therefore, we strongly urge all adopters to exercise caution and to conduct additional safety testing before deploying our models. For this purpose, we outline responsible release notes in Section 9.
Pretraining Data
We pretrain the LLM on hundreds of billions of words of diverse text from a variety of sources in order to develop a strong foundation in the target language(s) while at the same time establishing a broad factual knowledge base in the model. In settings such as clinical domains, research has shown that larger-scale LLMs exhibit improved emergent capabilities [SAT+22]. Note that LLMs such as LLaMA [TLI+23] and Falcon [AAA+23] are predominantly trained on a single language: English. While these models exhibit impressive linguistic and reasoning capabilities, their abilities do not extend so well to other languages such as Arabic, as we will demonstrate experimentally below.
Moreover, the extent of knowledge of Arabic world embedded in these models is limited, as they only include relatively small amounts of native Arabic text. To tackle this challenge, we pretrain our model with the largest Arabic dataset in the world, while further extending it with English data and some programming code, to improve the logical reasoning abilities of the model.
Our pretraining data mix is 1:2:0.4 for Arabic:English:code. We arrived at this ratio through extensive experiments on smaller models, which we describe in Section 3. We base this mix on all of the available Arabic data, as this is the smallest of the three data sources.
We collect our Arabic training data from multiple sources including web pages, Wikipedia articles, news articles, Arabic books, and social network content. To augment the dataset, we also translate English content to Arabic using an in-house machine translation system.Our in-house translation system is a standard transformer sequence-to-sequence model implemented in the FairSeq library [OEB+19] and trained on public datasets available in OPUS [Tie12]. The English to Arabic translation performance is 31 and 40 BLEU points [PRWZ02] on Flores-101 and a held-out test dataset, respectively. We restrict this to high-quality English resources such as the English Wikipedia and English books. We apply checks to avoid translating English sources with embedded code, or text that is not well structured.
A breakdown of the Arabic dataset (except the translated content) is detailed in Table 1. Specifically, we use text from the following sources:
Abu El-Khair: a collection of more than five million news articles, collected from ten major news sources of Arabic countries over a period of fourteen years [AEK16].
Aranews: Arabic news corpus from multiple sources ranging from year 2005-2022 [GEQ12]
ArabicText 2022: an open-source Arabic collectionhttps://data.baai.ac.cn/details/ArabicText-2022 prepared by the Beijing Academy of Artificial Intelligence (BAAI), that includes Arabic text corpora such as ArabicWeb22-A, ArabicWeb16 [SKF+16], OSCARhttps://oscar-project.org/, ArabicWeb22-B, CC100-AR [CKG+20], and Arabic Tweets.
Arabic subset of C4: a cleaned version of the Common Crawl using the cleaning and the filtering described in [RSR+20]. We use the Arabic subset of this corpus.
Arabic Wikipedia: Wikipedia written in Arabichttps://dumps.wikimedia.org/
ArabicNews 2020: an in-house news crawl at Inception of various Arabic news channels.
Maktabah: a corpus of approximately 6,500 Arabic books.https://www.kaggle.com/datasets/mahmoudqaddoumi/arabic-library
UN Meeting transcripts: the United Nations Parallel Corpus,https://conferences.unite.un.org/uncorpus v1.0 [ZJDP16] which is available in the six official languages of the United Nations, of which we use the Arabic documents.
Other Sources: a combined dataset of multiple smaller corpora including poetry, news, entertainment, sports, and management documents.https://master.dl.sourceforge.net, https://github.com/ceefour/hadith-islamware, https://alt.qcri.org/resources1/qedcorpus/QEDCorpusv1.4_MT.tgz
We further augment the Arabic data by translating 3B tokens from English Wikipedia and 15B tokens from the Books3 corpus. As a result, we increase the Arabic data from 55B to 72B tokens. Subsequently, we upsample this Arabic data 1.6 times, obtaining 116B Arabic tokens.
For English, we use The Pile [GBB+20], a collection of 22 high-quality datasets, from which we randomly sample 232B English tokens and 46B tokens from its GitHub subset. Table 2 shows details about the English data we use. Specifically, we use text from the following sources, part of The Pile:
Pile-CC: A subset of The Pile dataset, derived from the Common Crawl, a collection of website crawls from 2008 onwards. The dataset includes raw web pages, metadata, and text extractions from diverse domains. Due to the varying quality of the data in Common Crawl, Pile-CC is created using jusText [EN13] on Web Archive files for extraction, yielding higher quality output than directly using the WET files [GBB+20].
Books3: Derived from the contents of the Bibliotik private tracker made available by Shawn Presser [Pre20]. It is a mix of fiction and non-fiction books, significantly larger than the next largest dataset, BookCorpus2, and was included for its value in long-range context modeling and coherent storytelling.
ArXiv: A subset of the ArXiv preprint repository for research papers, which has been in operation since 1991.https://arxiv.org/
PubMed Central: A subset of the PubMed online repository for biomedical articles, managed by the United States’ National Center for Biotechnology Information (NCBI).https://www.ncbi.nlm.nih.gov/pmc
OpenWebText2: A web scrape dataset produced by EleutherAI, inspired by WebText [RWC+19] and OpenWebTextCorpus [GC19].
Wikipedia (en): The dataset, sourced from the TensorFlow Datasetshttps://www.tensorflow.org/datasets/catalog/wikipedia#wikipedia20200301en, includes articles from the English Wikipedia as a standard source of high-quality text for language modeling.
FreeLaw: This dataset is derived from the CourtListener platformhttps://www.courtlistener.com/, part of the Free Law Project, which provides access to legal opinions from federal and state courts in the United States.
PubMed Abstracts: This datasethttps://github.com/thoppe/The-Pile-PubMed includes abstracts from 30 million publications in PubMed, managed by the National Library of Medicine. It encompasses the significantly limited coverage of full texts in PubMed Central (PMC) and includes MEDLINE abstracts from 1946 to the present day.
DeepMind Mathematics: A collection of mathematical problems from various topics formatted as natural language prompts [SGHK19]. It is included in The Pile to enhance the mathematical ability of the language models [BMR+20].
Project Gutenberg (PG-19): This dataset consists of classic Western literature from Project Gutenberg, specifically books published before 1919 [RPJ+20]. It represents distinct styles compared to the more modern Books3 and BookCorpus datasets and is already used for long-distance context modeling.
BookCorpus2: An expanded version of the original BookCorpus [ZKZ+15], comprising books by unpublished authors, minimizing overlap with Project Gutenberg and Books3, which include published books. It is commonly used for language model training [RNSS18].
EuroParl is a multilingual parallel corpus initially introduced for machine translation [Koe05], but has also been utilized in several other fields of NLP [GW06, VH08, CDS17]. The version used in this work consists of the proceedings of the European Parliament in 21 European languages from 1996 until 2012.
PhilPapers: A collection of open-access philosophy publications from the Center for Digital Philosophy, University of Western Ontario.https://philpapers.org/
YouTube Subtitles: This dataset consists of text from human-generated closed captions on YouTubehttps://github.com/sdtblck/youtube_subtitle_dataset. It provides not only multilingual data, but also a variety of content including educational material, popular culture, and natural dialogue.
NIH Grant Abstracts: This dataset includes abstracts of awarded applications from the EXPORTER service, covering fiscal years 1985-present. It was included because it features high-quality scientific writing.https://exporter.nih.gov/
Enron Emails: This dataset [KY04] is widely used for analyzing email usage patterns. It was included to aid in understanding the modality of email communications, which is typically not found in other datasets.
GitHub: This datasethttps://github.com/EleutherAI/github-downloader consists of a large collection of open-source code repositories [BMR+20]. It was included to improve the model’s downstream performance on code-related tasks, given GPT-3’s ability to generate plausible code completions without any explicitly gathered code datasets.
Table 3 summarizes the composition of our dataset: a total of 395B tokens, including Arabic, English, and programming code.
Preprocessing, which includes filtering, normalizing, and cleaning, has been shown to be a vital step in training high-quality LLMs. We apply several standard preprocessing steps, combined with modules targeted at getting high-quality Arabic content, in a data processing pipeline to generate our Arabic dataset of 72B tokens.
An outline of our preprocessing pipeline for Arabic is provided in Figure 2. As explained above, the raw data is primarily sourced from publicly available databases, such as Abu El Khair or BAAI, as well as through in-house web scraping and machine translation of high-quality English sources.
Given that some of these sources have already been preprocessed or tokenized for NLP applications, it is essential to standardize our input. We thus subject all sources to an initial detokenization step (which leaves non-tokenized input unchanged) to achieve consistency. A document, at this step, is one article/web page, depending on the source.
We then apply a large number of filtering rules in order to eliminate documents that are noisy or low-quality. This includes removing extremely short or very long documents, or those that do not include a sufficiently high proportion of Arabic characters or sentences, which could be indicators of a document in a different language where Arabic characters appear only incidentally. We also remove documents that contain words more than 100 characters long, which can indicate the presence of extremely long URLs and/or an otherwise noisy document.
Once a document has passed the filtering step, it is subject to cleaning and normalization. We remove non-printable Unicode characters and rare diacritic marks, and normalize the text using the Camel toolset for Arabic [OZK+20]. We remove embedded JavaScript and HTML (which are common sources of noise in web-scraped datasets), and highly-frequent words and phrases (which are typically boilerplate text, such as a news channel name). We normalize Arabic punctuation marks, and use a lightweight -gram LM to further identify and remove noisy -grams.
Finally, we apply a fuzzy deduplication step using standard locality-sensitive hashing techniques. After this deduplication step, the size of the English dataset was about 20% of the original.
Things were more challenging for Arabic. Unlike English, where several large-scale and open-access datasets already exist, and established preprocessing pipelines are available, for Arabic, this pipeline had to be custom-built. Experimentation with smaller LLMs informed many of the choices of heuristics we used in our final preprocessing pipeline. Given the limited amount of available Arabic data, we took care not to filter Arabic content as aggressively as for English.
2 Mixing Arabic and English Data
A commonly reported phenomenon in LLM research is that larger LLMs generally perform better than smaller ones; this trend is clearly visible on public LLM leaderboardshttps://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard and is also evident in the recent LLaMA2 release [TMS+23].https://ai.meta.com/llama/ In general, the quality of a model is limited by two main factors: (i) data availability, and (ii) computational cost. While the latter can be overcome with improved hardware, the former is a fundamental obstacle. The Chinchilla scaling law [HBM+22] tells us that the optimal balance between model size and data is approximately twenty tokens per parameter. This is why for English, the largest open-source LLMs until recently had about 30B parameters, as publicly available datasets such as Red Pajamahttps://github.com/togethercomputer/RedPajama-Data have 1.2T tokens of text. The recently-released LLaMA2 has 70B parameters, and it is trained on 2T tokens.
As mentioned above, for Arabic, we have 72 billion tokens (after adding 18 billion tokens of translated text). If we apply the Chinchilla scaling law, we would optimally be able to train a model of 6-7B parameters on this data. We could probably train a slightly larger model, as Arabic involves cltificization of conjunctions and pronouns (e.g., and his house is one word in Arabic, but three words in English), and thus the scaling law might differ a bit. Indeed, some of our experiments suggest that one might need as few as 14 tokens per parameter for Arabic; yet, this does not fundamentally change the fact that we do not have enough data to train a 13B parameter Arabic model, let alone a 30B one. One possible solution is to obtain more data, e.g., by adding more Arabic social media posts, but these are generally noisy. Another option is to train on mixed Arabic and English training data, and thus compensate for the missing Arabic tokens with English ones. This latter idea worked well in our experiments: we found that mixing Arabic and English in a proportion of 1:2 (i.e., 2 more English than Arabic) works better than training on Arabic only. In the future, we plan to try incorporating a higher proportion of English, but we also need to be careful: for example, the BLOOMz experiments [MWS+23] indicate that adding ten times as much English data results in degradation of the model performance.
Model
Jais is based on a standard transformer-based architecture [VSP+17]. In particular, we use a causal decoder-only model, similar to the one used by GPT-2 [RWC+19] and LLaMA [TLI+23]. Decoder-only models have achieved state-of-the-art performance in generative language tasks. Building upon this base transformer architecture, we use a number of recent improvements from the literature, as well as from our own experiments.
The choice of tokenizer can have a significant impact on the performance of an NLP model [LBM23]. How words are split is influenced by the composition of the corpora used to train the tokenizer [PLMTB23]. A common tokenizer used in LLMs is the GPT-2 tokenizer [RWC+19], which is also used by OPT [ZRG+22] and GPT-3 [BMR+20].
However, because the GPT-2 tokenizer is primarily trained on English corpora, common Arabic words such as \RLلماذا (English ‘why’) are over-segmented into individual characters [PLMTB23]. This over-segmentation lowers the performance of the model and increases the computational costs compared to using a custom tokenizer that is specifically designed for the target languages [CL19]. Moreover, in order to increase the scope of multi-linguality, we want the tokenizer to break words into meaningful subwords. This is likely to encourage cross-lingual transfer by better token-level alignment between languages.
In order to achieve this, we trained our own subword tokenizer (Jais tokenizer) on a combined corpus of English and Arabic languages using byte-pair encoding (BPE) [SHB16]. To alleviate bias towards one language, we prepared a training corpus of 10B words containing equal proportions of English and Arabic text. Table 4 shows the fertility scores [BCP+90] of Jais tokenizer against the tokenizers of BERT Arabichttps://huggingface.co/asafaya/bert-base-arabic [SAY20], BLOOM [SFA+23], and GPT-2 [RWC+19] on English, Arabic, and code validation datasets. We can observe that the fertility score for the Jais tokenizer is close to 1, even though the vocabulary of Jais has only 84,992 entries, compared to BLOOM, which has 250,000 entries. The result shows the optimality of our custom-made tokenizer over our test corpus as compared to other tokenizers.
Positional embeddings provide information about word order to transformer-based LLMs. A common strategy to manage training complexity is to train the model with a limited context length. Subsequently, during inference, the model is applied to an extended context length using extrapolation [SLP+22]. Recent research has indicated that conventional methods of integrating word order into the transformer model, such as learnable positional embeddings, as used in models such as GPT-2 [RWC+19], and sinusoidal encoding, as proposed in [VSP+17], do not perform well when applied to longer contexts [PSL22]. Thus, we use Attention with Linear Biases (ALiBi) positional encodings [PSL22], which support efficient extrapolation to long contexts. Rather than modifying the input embeddings, ALiBi penalizes the attention scores by a linearly decreasing amount, proportional to the distance between the relevant key and the query.
Activation functions play a pivotal role in the training of neural network models. We use SwiGLU [Sha20] in each transformer block. It combines the advantages of Swish [RZL17] and GLU [Sha20] activations, and has been shown to improve over both of them. Because of SwiGLU’s extra computational overhead, adjustments were made in the hidden dimensionality of the feed forward network to compensate. Rather than apply a filter , we apply a filter that is . This ensures that the feed forward network has a FLOP cost that is comparable to that of GeLU activation.
Hyperparameter search in LLMs is expensive due to the size of the model and the scale of the dataset used in training. Thus, it is not feasible to do an extensive hyperparameter search on the final model. Fortunately, recent studies have shown that optimal hyperparameter values become stable across neural network sizes when the models have been parametrized using maximal update parametrization (µP) [YHB+21]. For Jais hyperparameter search, we tuned the optimal values for batch size and learning rate on a 40M-parameter model, and transferred the best values to our 13B-parameter model.
2 Model and Training Hyperparameters
Table 5 shows the number of layers, heads, and dimensionality for Jais, along with the optimization hyperparameter values and peak learning rates.
While training, we sampled a source from the source list described in Section 2 and generated instances with a complete length of tokens. When a document was smaller than tokens, we concatenated several documents into one sequence. <|endoftext|> is used to demarcate the end of each document, giving the language model the information necessary to infer that tokens separated by <|endoftext|> are unrelated.
We train Jais-13b using the AdamW optimizer [LH18] with , , , and weight decay of 0.1. We scale the gradient norms using a maximum norm clipping value of 1.0. The learning rate schedule starts with a linear warm-up from 0 to the maximum learning rate at 95 steps, followed by a 10 linear decay until 100,551 steps. After packing, we used a global batch size of 3,392 sequences of 2,048 tokens each. For µTransfer, we base Jais-13b on a roughly 40M-parameter model. The model depth is 24 and the hidden dimension size is 256.
The base learning rate is set to a maximum value of 1.2e-2, and the learning rate for each layer is set according to this base value depending on the layer shape [YHB+21]. Analogously, we initialize the layers with a base standard deviation of 7.3e-2, which we adjust based on the layer shape. Additionally, we scale the embedding’s output activations by a factor of 14.6, and scale the model’s output logits by a factor of 2.22 divided by the hidden size multiplier, e.g., 5,120 / 256 = 20.
3 Learnings and Observations
We conducted a series of preliminary experiments training on Arabic-only data, as well as on mixtures of Arabic and English. The aim was to find the optimal mix, and to identify the best model size for our Arabic-centric LLM. We maintained a constant size for the Arabic corpus as discussed in Section 2. We further sampled the English dataset to reflect different ratios relative to the Arabic data size. In all cases, we trained the LLM for one epoch. Previous work [BMR+20, KMH+20] has shown that cross-entropy loss correlates with LLM quality in downstream tasks. Therefore, we report the cross-entropy loss on the Arabic validation set.
Due to the size of the search space and required computing resources, we did not train models of all sizes and for all data ratios. Instead, we experimented on models of 590M, 1.3B, 2.7B, 6.7B, 13B, and 30B parameters under a few data ratios. The trends are shown in Figure 3. We can see that for small models, e.g., 590M and 1.3B parameters, adding English impacts the cross entropy loss in Arabic adversely. However, this trend reverses for larger models, e.g., for 6.7B and 13B parameters, where adding English improves Arabic performance. In particular, we observe that the 13B model trained on a 1:2 Arabic–English mix (Jais-13b) outperforms the 30B-parameter Arabic-only model by a sizable margin. This suggests that increasing the model capacity improves the cross-lingual transfer between English and Arabic. In future work, we plan to study the extent to which additional English data can be incorporated without adversely affecting the performance of Arabic.
4 Training Infrastructure
All training, hyper-parameter tuning, and instruction-tuning experiments were executed on the Condor Galaxy 1 (CG-1) www.cerebras.net/blog/introducing-condor-galaxy-1-a-4-exaflop-supercomputer-for-generative-ai/ AI supercomputer from Cerebras, built in partnership with G42. The final training and fine-tuning runs for Jais were performed on 16 CS-2 systems within CG-1. CG-1 is a Cerebras Wafer-Scale Cluster composed of Cerebras CS-2 systems, MemoryX, SwarmX, management, and input worker nodes. The foundation of the CG-1 cluster is the Cerebras Wafer Scale Engine (WSE) within the CS-2 system, the largest and most powerful AI processor currently available. CS-2 systems are purpose-built network-attached AI accelerators. MemoryX is a large-capacity off-wafer memory service, used to store all model weights, gradients, and optimizer states. SwarmX is a broadcast/reduce fabric that connects the memory service MemoryX to each of the CS-2 systems in a wafer-scale cluster. Swarm-X coordinates the broadcast of the model layer weights, giving each CS-2 a local copy, and it receives and aggregates (by addition) the independent weight gradients coming from the CS-2 systems during backpropagation. At the end of each iteration, the aggregated gradients are sent to MemoryX for weight update.
The CG-1 hardware and software stack enables training extremely large models using data parallelism by relying on a special execution mode available with Cerebras Wafer Scale Clusters, called weight streaming. Weight streaming fully bypasses the complexity of 3D parallelism on traditional GPU clusters, and provides simpler and higher performance scaling.
Instruction-Tuning
LLMs can produce coherent text and execute an extensive array of NLP tasks, requiring only a few task examples as input. Nonetheless, the model cannot interpret user instructions or engage in dialogue-style interactions without instruction-tuning [OWJ+22]. To tailor our LLMs for dialogue-style applications, we instruction-tuned them on a dataset prepared for instruction-based adaptation in English and Arabic. We refer to our instruction-tuned model as Jais-chat.
As we have a bilingual model, we use a combination of Arabic and English instruction-tuning datasets. We include a wide range of datasets covering various domains in single-turn and multi-turn chat formats. We have 10M prompt–response pairs in total, made up of 4M in Arabic and 6M in English; see Tables 6 and 7 for detailed stastistics about the datasets we use. Below, we provide a brief description of each dataset.
Super-NaturalInstructions [WMA+22] encompasses 76 types of tasks, such as classification, extraction, infilling, and sequence tagging. These instructions span a comprehensive range of 1,616 diverse NLP tasks, all presented in expert-written instruction–response pair format. P3 [SWR+21] and xP3 (Code & English) [MWS+23] are collections of prompted datasets that cover a diverse set of NLP tasks in instruction–response format. The P3 dataset contains over 2,000 prompt types from 270 different public datasets in English. xP3 (Code & English) is designed for multi-lingual and cross-lingual instruction-tuning and contains more than 9M examples in 46 languages, including programming languages. To make our model diverse, we included at most five thousand examples from each task of the Super-NaturalInstructions dataset; from P3 and xP3 (Code & English), we only include English and programming code examples. The Natural Questions datasethttps://huggingface.co/datasets/nq_open comprises question–answer pairs extracted from Google Search; it only includes questions with concise answers, which can be addressed using the information found in English Wikipedia [KPR+19].
Baize-Chatbothttps://huggingface.co/datasets/linkanjarad/baize-chat-data is a multi-turn dialogue-style instruction-tuning dataset. HH-RLHF is designed for helpful and harmless assistance through preference modelling [OWJ+22], and has an accepted and a rejected response for each prompt; we only use the former. Alpaca-CoT [QS23] is a fusion of nine Chain-of-Thought (CoT) [WWS+22] datasets released by FLAN [CHL+22]. Self-instruct [WKM+23] is a bootstrapping algorithm that uses a small set of manually written instructions to prompt an LLM to generate new instructions.
We used the dataset provided by the authors, which was cleaned and filtered to remove low-quality or similar pairs. Alpaca-Cleanedhttps://huggingface.co/datasets/yahma/alpaca-cleaned, Instruct-Wild [XJS+23], Unnatural Instruction [HSLS23] and GPTeacherhttps://huggingface.co/datasets/causal-lm/gpt_teacher are prepared using the same method, but using ChatGPT [BMR+20].
Open Instruction Generalist (OIG)https://huggingface.co/datasets/iamketan25/oig-instructions-dataset, GPT4ALL-J [AND+23], and Dolly-15k [CHM+23] were constructed to train assistant-style LLMs in a semi-automatic way, and are moderate in quality. From GPT4ALL-J, we randomly sampled 100,000 examples from v1.0.https://huggingface.co/datasets/nomic-ai/gpt4all-j-prompt-generations HC3 [GZW+23] is a manually curated dataset for comparing the response of humans and ChatGPT; we used the former only. From HC3, we only included examples from four domains: finance, medicine, Wikipedia, and OpenQA. GSM-General-QA https://huggingface.co/datasets/iamketan25/gsm-general-qa-instructions, Math-Instructionhttps://huggingface.co/datasets/alpayariyak/MATH_Instruction_Format and Grade-School-Mathhttps://huggingface.co/datasets/qwedsacf/grade-school-math-instructions are instruction-tuning datasets prepared to assist in mathematical problems. Finally, Instruction-Poems https://huggingface.co/datasets/checkai/instruction-poems and Essays-with-Instructionshttps://huggingface.co/datasets/ChristophSchuhmann/essays-with-instructions target poem and essay writing, and Stack-Exchange-Instructionhttps://huggingface.co/datasets/ArmelR/stack-exchange-instruction and Python-QAhttps://huggingface.co/datasets/iamketan25/python-qa-instructions-dataset are aimed at programming code tasks.
In order to enhance the conversational abilities of our fine-tuned model, we integrated dialogue-based and persona-based datasets into the instruction-tuning procedure. For this purpose, we curated 19 in-house question–answer pairs that revolved around the LLM developer, and we also processed the Basic-Convhttps://github.com/gunthercox/chatterbot-corpus/tree/master dataset to incorporate it into our instruction-tuning process.
We further created our own set of question–answer pairs related to the UAE and the local region, based on information from relevant Wikipedia pages and other sources. We refer to this dataset as NativeQA and incorporate it into the fine-tuning process. We also prepared an instruction dataset to teach the model about safety issues, named it SafetyQA. As a responsible language model, we want the model to avoid engaging in unsafe conversations e.g. discussions on self-harm, sexual violence, or identity attacks. For this, we prepared prompt-response from DoNotAnswer [WLH+23] and OLID [ZMN+19]. In all these prompts, the response is a polite rejection of the question. The impact is explored in Section 6.
1.2 Arabic Instruction-Tuning Datasets
Due to the limited availability of instruction-tuning datasets for Arabic, we translated some of the above English instruction-tuning datasets to Arabic using the same machine translation system that we used for the training data: Supernatural Instruction, Unnatural, NaturalQuestions, Alpaca [TGZ+23], HC3, Dolly-15k, Baize, Basic-Conv, Bactrian [LKW+23]. We then performed a manual assessment for each task within the Super-NaturalInstructions dataset, and excluded tasks that were primarily related to translation as well as those relating to counting words, as they could break when translated to Arabic (i.e., their is no guarantee the translated text has the same number of words as the original English).
Apart from the translated datasets, we also included the Arabic examples from xP3 (Code & English). We further formatted AraNER [BRB07] to the instruction–response format (NER-Ar) and added it as a dataset for instruction-tuning. Moreover, similarly to English, we created additional datasets NativeQA-Ar and SafetyQA-Ar with instruction–response pairs related to the UAE and the region as well as safety, but this time in Arabic; note that we created these natively in Arabic. We further translated the English datasets that we created to Arabic, and we used them as additional datasets.
2 Instruction-Tuning Setup
In instruction-tuning, each instance comprises a pair of a prompt and its corresponding response, and the model needs to be able to distinguish between them. We thus wrap each instance within a template as illustrated in Figure 4, where we have additional special markers to indicate what is the human input and what is the expected response. Note that we use different templates for single-turn question–answer pairs vs. dialog interactions. We further use padding for each instance, as we cannot pack examples during instruction-tuning (unlike pretraining where we pack the documents until the maximum sequence length has been reached). We use the same autoregressive objective as for pretraining the LLM. However, similarly to Alpaca [TGZ+23], we mask the loss of the prompt, i.e., we perform backpropagation on the answer tokens only, which ensures that short responses are not penalized.
Evaluation
We perform a comparative evaluation of Jais and Jais-chat against other LLMs for both Arabic and English, building upon the evaluations conducted in prior studies [TLI+23, TMS+23, Ope23, SFA+23]. For each language, our evaluation encompasses aspects such as knowledge, reasoning, misinformation, and bias, as outlined in Table 8. To extend the evaluation to Arabic, we use an in-house English-to-Arabic translation system (as discussed in Section 2), and additionally we hired native speakers of Arabic to manually translate the MMLU dataset [HBB+22] from English to Arabic. We further added two additional datasets, with question–answering pairs that were in Arabic: (i) EXAMS [HMZ+20], a set of school examination questions in various languages (we took the Arabic questions only), and (ii) a new manually-constructed LiteratureQA dataset.This dataset was created in house by manually digitizing university-level Arabic language question papers from the following sources: http://www.examrace.com/, http://arabicuniversitycollege.yolasite.com
World Knowledge. Validating the knowledge embedded within a pre-trained language model is crucial, given its extensive training on a vast amount of textual data. We evaluate the knowledge of our models on four different datasets: (1) MMLU [HBB+22], a multiple-choice exam question set covering 57 tasks spanning various educational levels, from school subjects to university and professional exams; (2) RACE [LXL+17], a reading comprehension task constructed from English exams for middle and high school Chinese students; (3) EXAMS [HMZ+20], multilingual high school questions from natural and social sciences covering 16 languages including Arabic; and (4) LiteratureQA, a collection of multiple-choice questions focused on Arabic literature at the university level.
Commonsense Reasoning. Making inference from text requires logical reasoning, and language models that undergo pre-training on extensive textual data have been shown to be able to do such reasoning. We evaluate the reasoning capabilities of language models using seven datasets: (1) HellaSwag [ZHB+19], a sentence completion dataset for commonsense natural language inference, constructed using adversarial filtering, (2) PIQA [BZB+20], a set of questions that require reasoning, centered around physical activities, (3) BoolQ [CLC+19], a yes/no reading comprehension question dataset that requires a wide range of inferential capabilities, (4) SituatedQA [ZC21], a question-answering dataset that is conditioned on temporal and geographical context, (5) ARC-Challenge [CCE+18], a dataset comprising science questions typically encountered at the grade-school level, demanding considerably enhanced knowledge and reasoning capabilities,For ARC-Challenge, we only use the Challenge dataset, which presents a higher level of difficulty compared to the Easy dataset. (6) OpenBookQA [MCKS18], an elementary science question dataset designed to evaluate broad common knowledge, and (7) WinoGrande [SBBC21], a dataset comprising expert-crafted pronoun resolution tasks that require common-sense reasoning.
Misinformation and Bias. We also evaluate the faithfulness and the biases of our LLMs based on two datasets: (1) TruthfulQA [LHE22], which contains expert-crafted questions that measure the extent of model misconception on the topics of health, law, finance, and politics; and (2) CrowS-Pairs [NVBB20], a dataset to assess stereotype biases against protected attributes such as race, religion, and age.
We perform an extensive evaluation where we compare our LLMs to twenty baseline models that support Arabic and/or English. Some models are trained to support Arabic: AraT5 and AraT5-v2 (220M) [NEAM22], AraBART (139M) [KETH+22], mT0 (1.2B, 3.7B, 13B) [MWS+23], BLOOM (1.7B, 3B, 7.1B) [SFA+23], and BLOOMz (1.7B, 3B, 7.1B) [MWS+23]. Other models are not trained for Arabic, but still can answer questions in Arabic, probably because some amount of Arabic data was present in their pretraining and/or instruction-tuning datasets: LLaMA (7B, 13B) [TLI+23], LLaMA2 and LLaMA2-chat (7B, 13B) [TMS+23], and Falcon (7B) [PMH+23].
We adopt the LM-Evaluation-Harness framework [GTB+21] to evaluate each model in a zero-shot setting, and we report the accuracy for each task. Within the LM-Evaluation-Harness framework, the context string is concatenated with each candidate output string, and the answer is determined by selecting the concatenated string with the highest normalized log-likelihood.
Table 9 shows the zero-shot evaluation results for Arabic. We can see that our Jais and Jais-chat models exhibit superior performance across all evaluation criteria, establishing them as the new state-of-the-art LLMs for Arabic. Specifically, in comparison to monolingual Arabic models (AraT5, AraT5-v2 and AraBART), Jais-chat (13B) achieves absolute performance improvements of +11.7 to +15.3. This is particularly pronounced in the domains of knowledge acquisition and commonsense reasoning.
We can further see that BLOOMz (7.1B) is the best baseline model for Arabic, with an average accuracy of 42.9, which is better than mT0-xxl (13B), which has an accuracy of 40.9. Notably, Falcon, LLaMA, and LLaMA2 lag behind, which should not be surprising given their limited exposure to Arabic pre-training data. We see that Jais-chat (6.7B) outperforms these baselines (including the 13B models) by +3.5 to +10.9 points absolute. Moreover, Jais-chat (13B) widens the gap even further, with an additional overall improvement of +1.9 points over Jais-chat (6.7B).
Instruction-tuning [OWJ+22] further improves the results over the corresponding base models, with the exception of Falcon (7B). The absolute improvements due to instruction-tuning for Jais-chat (1.3B, 6.7B, 13B) are +0.7, +3.2, and +1.9, respectively, and are similar to those for BLOOMz. The full results for each dataset and model can be found in the Appendix (Table 12).
We also performed an evaluation for English. The results are given in Table 10, where we can see that Jais-chat is highly competitive against existing English models, despite having seen less English data in pretraining. First, we observe that the existing Arabic models perform almost randomly on this benchmark, while our models perform substantially better. This result is unsurprising given that AraT5, AraT5-V2, and AraBART were pretrained on Arabic data only. In comparison to the multilingual BLOOMz (1.1B), Jais-chat (1.3B) performs +3.4 points better. We can further see that Jais-chat (13B) performs on par with the recently released LLaMA2-chat (13B) model (57.3 vs. 57.7), even though the latter is trained on 2T of English word tokens, while our model has only seen 232B English word token. Jais-chat (13B) also outperforms other baselines including mT0-xxl (13B) and Falcon (7B), by margins ranging from +2.6 to +7.2 points absolute. Our instruction-tuning is also effective, with improvements of +3.9, +4.3, and +3.4, for the 1.3B, 6.7B, and 13B models, respectively. The full results for each dataset and model can be found in the Appendix (Table 13).
2 Generation Evaluation
We next perform evaluation of the models over the core capability of Arabic text generation. Following prior work [PLH+23, CLL+23], we perform automatic evaluation over the generated Arabic content using GPT-4 [Ope23] based on Vicuna-Instructions-80, which were manually translated to Arabic by translators.
Vicuna-Instructions-80https://lmsys.org/blog/2023-03-30-vicuna/ consists of 80 challenging and open-ended questions across eight categories: knowledge, Fermi, counterfactual, roleplay, generic, math and coding, writing, and common-sense.
We generate outputs for Arabic prompts in Vicuna-Instructions-80 using a temperature of 0.3 and a repetition penalty of 1.2. As baselines, we use two closed-source models, ChatGPT (175B) [OWJ+22] and Claude (52B).https://www.anthropic.com/index/introducing-claude We further use several open-source models, which are either Arabic centric or multilingual: BLOOM (7B) [SFA+23], BLOOMz (7B) [MWS+23], AraT5 (220M) [NEAM22], AraT5-v2 (220M) [NEAM22], AraBART (550M) [KETH+22], and LLaMA2 (13B) [TMS+23]. We also include as baselines Bactrian-X (13B) and Bactrian-X (7B) [LKW+23], which are LLaMA and BLOOM base models, respectively, fine-tuned on multi-lingual (including Arabic) instruction-tuning datasets. For convenience, we name them BX and BX, respectively. We evaluate these baselines against our instruction-tuned models – Jais-chat (6.7B) and Jais-chat (13B). During the GPT-4 evaluation, we perform pairwise comparisons between all pairs of models. We first prompt GPT-4 to score each pair of models based on their outputs generated for the prompts in the Arabic Vicuna-Instructions-80. We randomly permute the answers from both candidates, aiming to have any one as the first candidate at random, and we prompt GPT-4 as follows:
You are a helpful and precise assistant for checking the quality of two Arabic assistants. Suppose the user only speaks Arabic, please evaluate both answers with your justification, and provide an integer score ranging from 0 to 10 after your justifications. When evaluating the answers, you should consider the helpfulness, relevance, accuracy, and level of detail of the answers. The score for answer 1 should be wrapped by
First, we find that certain models struggle to generate meaningful Arabic text according to the given instructions. This observation applies particularly to models that have not undergone instruction-following fine-tuning, namely BLOOM, AraT5, AraT5-v2 and AraBART. Additionally, some models produce subpar Arabic text, with average scores lower than 1 (out of 10) when evaluated against Jais-chat (13B) — these models include BLOOMz and LLaMA2. While BLOOMz is competitive in downstream task evaluation (see Table 9), it is unable to follow Arabic instructions, despite being pretrained using 73G of Arabic text [SFA+23].
With these observations, we focus on comparing the top 6 models: ChatGPT, Claude, BX, BX, Jais-chat (6.7B), and Jais-chat (13B). As there are six models in total, each one is compared against the other five models, resulting in 400 scores (80 questions 5 pairs) for every individual model. Since each score ranges from 0 to 10, summing up these scores for a model brings the maximum possible total score to 4,000.
The overall comparative results are shown in Figure 5. While both ChatGPT (175B) and Claude (52B) outperform Jais-chat (13B), it is important to note that (i) they are 4–13 times larger, and (ii) the difference in scores between our Jais-chat (13B) and these larger models is relatively modest, at around 400 points.
When focusing solely on certain types of tasks, including common-sense, knowledge-based, writing-related, and generic inquiries, the disparity between Jais-chat and ChatGPT/ Claude diminishes. Jais-chat is only 35 scores behind Claude and 114 scores behind ChatGPT, out of a total of 2,000, as illustrated in Figure 6.
Figure 7 shows a breakdown of the scores across various tasks. For the categories in Figure 6 including common-sense, knowledge-based, writing-related, and generic inquiries, Jais-chat performs generally better. This is particularly true for writing, where Jais-chat is almost on par with ChatGPT and Claude. In other task categories, including counterfactual, Fermi, roleplay, and math-and-coding, Jais-chat is worse than ChatGPT and Claude. This is expected, since these categories require a higher degree of reasoning, and the smaller size of the Jais-chat models puts them at a disadvantage.
Safety
We used several strategies and precautionary measures to make Jais-chat safer to interact with and to minimize potential risks. These precautionary measures were incorporated at various stages of the model development.
During the instruction-tuning process, we encoded safety measures into Jais-chat. Moreover, towards developing an interactive application based on Jais-chat, we implemented several practical and simple safety measures, which we describe here with the aim of providing developers examples of guardrails to be considered during the application development for end-users.
To ensure that Jais-chat has in-built safeguards on the content it generates, we have focused on this aspect during instruction-tuning. This involves avoiding the generation of content in the five risk areas identified by [WMR+21]. Through the process of instruction-tuning, we impart the following principles to Jais: (1) refrain from generating language that promotes discrimination, exclusion, or toxicity, regardless of user request or preference; (2) uphold privacy standards by preventing the leakage of private or sensitive information; (3) exercise caution in disseminating accurate information and responding thoughtfully to queries that could potentially lead to material harm, such as those related to fields like medicine or law; (4) reject engagement in any form of malicious use, including inquiries about unethical or illegal activities; and (5) counteract emotional manipulation by transparently indicating that the model is a chatbot and not a human, particularly when there is a discernible overreliance on its responses. Furthermore, we also add some examples that aim to teach Jais-chat to avoid engaging in discussions on sensitive topics, particularly such concerning certain aspects of religion and politics.
We crawled data from various Arabic websites, encompassing a wide spectrum of materials related to religion and politics, and amassed approximately 1,000 instances in Arabic. Given the constraints of available Arabic resources, we expanded our dataset by incorporating content in English. To this end, we integrated the DoNotAnswer dataset [WLH+23], which comprises around 6,000 questions designed to potentially provoke harmful output from language models, along with corresponding harmless responses. Subsequently, we translated this dataset into Arabic, forming the core foundation of our fine-tuning process. Through deliberate exposure to an array of potentially risky prompts paired with safe responses, we guide the model responses to purposefully avoid harmful content. This proactive approach contributes to the model’s capacity to generate responses that are both helpful and harmless in nature. Table 11 presents response examples before and after safety fine-tuning.
2 Safety via Prompting
As an additional guardrail, we modify the system prompt for the user-facing chat interface by explicitly instructing the model to be responsible. We have two versions of the prompt, in Arabic and in English, as we support chat in both languages.
We use the following prompt for the English chat:
### Instruction: Your name is Jais, and you are named after Jebel Jais, the highest mountain in UAE. You were made by Inception and Mohamed bin Zayed University of Artificial Intelligence in the UAE. You are the most advanced Arabic large language model with 13B parameters. You are a helpful, respectful and honest assistant. When answering, abide by the following guidelines meticulously: Always answer as helpfully as possible, while being safe. Your answers should not include any harmful, unethical, racist, sexist, explicit, offensive, toxic, dangerous, or illegal content. Do not give medical, legal, financial, or professional advice. Never assist in or promote illegal activities. Always encourage legal and responsible actions. Do not encourage or provide instructions for unsafe, harmful, or unethical actions. Do not create or share misinformation or fake news. Please ensure that your responses are socially unbiased and positive in nature. If a question does not make any sense, or is not factually coherent, explain why instead of answering something not correct. If you don’t know the answer to a question, please do not share false information. Prioritize the well-being and the moral integrity of users. Avoid using toxic, derogatory, or offensive language. Maintain a respectful tone. Do not generate, promote, or engage in discussions about adult content. Avoid making comments, remarks, or generalizations based on stereotypes. Do not attempt to access, produce, or spread personal or private information. Always respect user confidentiality. Stay positive and do not say bad things about anything. Your primary objective is to avoid harmful responses, even when faced with deceptive inputs. Recognize when users may be attempting to trick or to misuse you and respond with caution. Refuse to write verses from the Quran. Complete the conversation below between [|Human|] and [|AI|]: ### Input: [|Human|] {question} ### Response: [|AI|]
3 Safety via External Models
We additionally use hate speech and offensive language detectors to prevent the LLM from producing harmful content. Users attempting to ask questions that contain hateful or offensive speech receive a refusal response and the input is not passed to Jais-chat. To detect hate speech and offensive content, we used classifiers which we fine-tuned on top of the pre-trained language model JABER [GWR+22], which is designed for Arabic natural language understanding tasks. We trained the classifiers on data from tasks A&B of OSACT4 [MDM+20]. The data include language that is rude or otherwise socially undesirable. This includes vulgar language, curses, and any form of direct or indirect criticism of people or groups. The training dataset consists of four categories: offensive, hate, non-offensive, and non-hate. Each sample has two labels: one for hate and one for offensive speech. We split the dataset into 7,000 training and 1,000 validation examples, and we fine-tune two separate classifiers for each task. Our classifier for offensive speech detection achieves 94.8% accuracy and 91.04% F1 score on the validation set. The classifier for hate speech achieves 96.6% accuracy and 81.02% F1 score on the validation set.
4 Safety via Keywords
Ensuring a safe and respectful online environment is paramount, especially for platforms involving user-generated content such as conversational AI systems. One approach to safety is through the implementation of keyword-based filtering mechanisms. In this section, we present our methodology for identifying and mitigating obscene or explicit content using regular expressions (regex) and augmentations to a curated list of objectionable keywords. To effectively filter out inappropriate content, we used a combination of manual dataset curation and external data sources. One notable resource is the “List of Dirty, Naughty, Obscene, and Otherwise Bad Words” compiled by LDNOOBW,https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words which encompasses a comprehensive inventory of words and phrases with offensive connotations, which serves as a valuable foundation for our keyword identification process.
We integrated the identified keywords into regex patterns, allowing us to efficiently scan user-generated content for instances of potentially offensive language. When a user’s input contains flagged keywords, our system immediately responds with a safe refusal message instead of calling Jais-chat. The regex-based approach facilitates real-time detection and mitigation of inappropriate content. The effectiveness of this method in enhancing the safety and the appropriateness of interactions underscores its significance in upholding a positive, secure, and respectful user experience.
While our approach effectively addresses explicit content, it is important to acknowledge its limitations, including potential false positives and the dynamic nature of language usage. Our ongoing efforts to keep the quality high include continuous refinement of the keyword list and exploration of advanced natural language processing techniques in order to further enhance the accuracy and the effectiveness of our content filtering system.
Related Work
Below, we discuss previous work on the following relevant topics: Arabic language models, LLMs in general, instruction-tuning, and evaluation of LLMs.
Arabic language models have been developed across various architectures and learning objectives. Examples of encoder-only models include AraBERT [ABH20], QARiB [AHM+21], JABER and SABER[GWR+22], CAMeLBERT [IAB+21], AraELECTRA [ABH21a], GigaBERT [LCXR20], and ARBERT & MARBERT [AMEN21]. There have been also decoder-only models such as ARAGPT2 [ABH21b]. In the encoder–decoder category, prominent models include AraT5 [NEAM22] and AraBART [ETH+22]. These models, when fine-tuned, have demonstrated competitiveness in both natural language understanding and natural language generation tasks.
However, to the best of our knowledge, no public Arabic language model has been trained with over a billion parameters, capable of showcasing generalization and robust zero-shot capabilities across various tasks.JASMINE [NAME+22] are Arabic GPT models ranging in sizes from 350M to 13B parameters, but these models have not been released to the public, and the arXiv paper describing them says the 6.7B and 13B models are still training. In addition to monolingual models, Arabic has also been integrated into multilingual models, including earlier models such as mBERT [DCLT19] and XLM-RoBERTa [CKG+20], as well as more recent large language models such as BLOOM [SFA+23]. However, due to the Arabic content being dwarfed by other languages, these models tend to perform substantially worse than dedicated monolingual models and often exhibit limited generalization abilities in zero-shot settings [LKW+23].
Language models with ever larger numbers of parameters have consistently improved over smaller models such as BERT [DCLT19], BART [LLG+19], and T5 [RSR+20]. Despite their extensive training on multilingual text data, recent large language models have an English-centric bias and are less effective for languages beyond English [LKW+23]. An exception to this trend is GLM [ZLD+23], which is the sole large language model specifically designed to excel in both Chinese and English.
Existing pretraining frameworks for language models fall into three categories: autoregressive, autoencoding, and encoder–decoder models. Most recent large language models, such as the GPT series [RWC+19, BMR+20, Ope23], LLaMA [TLI+23, TMS+23], BLOOM [SFA+23], and Falcon [AAA+23], are autoregressive, using a left-to-right language model objective. Earlier models such as BERT [DCLT19], ELECTRA [CLLM20], and RoBERTa [LOG+19] are encoder-only, while BART [LLG+19] and T5 [RSR+20] are encoder–decoder models. As elaborated in Section 3.1, Jais and Jais-chat follow the autoregressive model paradigm, building upon the successes of LLaMA2 and GPT-4.
Progress in large language models can also be categorized into two streams: closed-source and open-source models. Closed-source models such as Bard,https://ai.google/static/documents/google-about-bard.pdf Claude,https://www.anthropic.com/index/introducing-claude Gopher [RBC+22], and GPT-4 [Ope23] offer fewer advantages to the research community compared to open-source models [TMS+23, SFA+23]. The lack of model transparency exposes leads to various risks for closed-source models, including privacy concerns [MGU+22, YRC23] and safety issues [SXD+22]. In contrast, Jais and Jais-chat are open-source models, as elaborated in Section 3.
Fine-tuning language models using instruction–response pairs has enhanced the generalization capabilities of language models across various tasks [OWJ+22]. In terms of open-source models, BLOOMz [MWS+23] is a fine-tuned version of the foundation model BLOOM [SFA+23] based on large-scale instruction-tuning over a dataset created via templates, while LLaMA2 [TMS+23] uses a publicly available instruction–response pair dataset [CHL+22]. Moreover, instruction-tuning has been coupled with reinforcement learning with human feedback (RLHF). This combination aligns the generated responses to reward functions, optimizing the model’s factuality, and reducing its toxicity [OWJ+22].
The prompts used for instruction-tuning can have diverse origins. Some, as observed by [ZMH+23], are human-designed, while others can be autonomously generated. These prompts can be refined with follow-up instructions for more relevant or specific outputs, as studied by [GAS+23] and [MTG+23]. Recently, [WWS+22] introduced chain-of-thought prompting, directing models to clarify their reasoning over complex tasks, which was shown to enhance their accuracy.
Large language models are proficient at generating coherent and fluent text, but have shortcomings in terms of factuality and reasoning skills. As a proxy to evaluate factuality, existing English large language models such as GPT-4 [Ope23] and LLaMA [TLI+23] use school exam questions [HBB+22] to understand how faithful the models are at providing knowledge. Evaluating commonsense reasoning abilities is also important, and is the target of datasets such as HellaSwag [ZHB+19], WinoGrande [SBBC21], ARC easy and challenge [CCE+18], and OpenBookQA [MCKS18]. Moreover, reasoning via programming is evaluated using HumanEval [CTJ+21] and MBPP [AON+21].
In Arabic NLP, existing benchmarks primarily focus on evaluating natural language understanding tasks. For instance, the ALUE benchmark [STG+21] encompasses semantic tasks such as irony detection [GKB+19], emotion classification [MBMSK18], sentiment classification [MBMSK18], offensive language [MDM+20] and hate speech identification [MDM+20]. Existing Arabic benchmarks, however, do not include knowledge and commonsense evaluation, posing a challenge for the assessment of Jais.
In contrast, in other languages, researchers have effectively used methods such as machine translation or the construction of datasets in a similar manner to assess the knowledge proficiency and the commonsense understanding of language models [Ope23, LKW+23]. In this context, as detailed in Section 5, we used a combination of techniques, including crafting analogous datasets to those available for English, using human translations and our in-house machine translation system to convert English datasets into Arabic for the purposes of evaluation.
Evaluating only on knowledge [HBB+22, LZK+23] and commonsense reasoning [ZHB+19, SBBC21] based on the evaluation settings of prior work [TLI+23, MWS+23] is arguably not a holistic evaluation, as they are multiple-choice questions. To evaluate the generated text as a whole, human evaluation remains crucial. Unfortunately, it is both resource-intensive and sometimes exhibits variable quality, especially when using crowd-sourcing. Recent studies [Tör23, LXA23, GRS+23, WA23] have even suggested that ChatGPT annotation surpasses the performance of Amazon crowd-sourced workers, underscoring the importance of expert workers in the evaluation process. Expanding upon these findings, another study [PLH+23, CLL+23] used GPT-4 as a substitute for crowd-sourced workers to compare two model outputs. This is achieved by presenting an evaluation prompt and providing both model outputs as a context for the assessment.
Conclusion
We have introduced Jais, a new state-of-the-art Arabic-English bilingual large language model (LLM), as well as its instruction-tuned variant, Jais-chat. The latter can perform a wide range of generative and downstream language tasks in both Arabic and English, ranging from common-sense reasoning to natural language understanding tasks such as sentiment analysis, irony detection, and hate speech detection. Its pre-trained and fine-tuned capabilities outperform all known open-source Arabic models, and are comparable to state-of-the-art open-source English models that were trained on larger datasets. We encourage researchers, hobbyists, and enterprise developers alike to experiment with and to develop on top of our model, particularly those working on multi-lingual and/or non-English applications.
Jais represents an important evolution and expansion of the NLP and AI landscape in the Middle East. This first-of-a-kind Arabic model born in the UAE represents an important strategic step for government and commercial organizations towards the digital revolution. By advancing Arabic language understanding and generation, empowering local players with sovereign and private deployment options, and nurturing a vibrant ecosystem of applications and innovation, this work supports a broader strategic initiative of digital and AI transformation to usher in an open, more linguistically-inclusive, and culturally-aware era.
Release Notes
We release the models under Apache 2.0 license. Users of Jais must comply with the terms of the provided license, and applicable policies, laws, and regulations governing the specific use case and region. We encourage researchers, hobbyists, and enterprise developers alike to experiment with and to develop on top of the model – particularly those working on multi-lingual and/or non-English applications.
This model is not only the first of its kind in the Arabic LLM ecosystem, but it also has been shown to be the best in the world among open Arabic or multilingual LLMs in terms of Arabic NLP capabilities. Some potential downstream uses are listed below:
Research: This model can be used by researchers and developers to advance the Arabic LLM/NLP field.
Commercial Use: It can be used as a foundational model to further fine-tune for specific usecases (like Jais-chat). Some potential usecases for businesses include (1) chat-assistants, (2) downstream tasks such as NLU/NLG, (3) customer service, and (4) process automation.
We believe that a number of audiences will benefit from our model:
Academics: those researching Arabic natural language processing.
Businesses: companies targeting Arabic-speaking audiences.
Developers: those integrating Arabic language capabilities in apps.
2 Out-of-Scope Use
While Jais is a powerful Arabic and English bilingual model, it is essential to understand its limitations and the potential for its misuse. The following are some scenarios, but not limited to, where the model should not be used:
Malicious Use: The model should not be used for generating harmful, misleading, or inappropriate content. This includes but is not limited to (i) generating or promoting hate speech, violence, or discrimination, (ii) spreading misinformation or fake news, (iii) engaging in illegal activities or promoting them, (i) (iv) handling sensitive information: the model should not be used to handle or to generate personal, confidential, or sensitive information.
Generalization Across All Languages: Jais is bilingual and optimized for Arabic and English, and it should not be assumed to have equal proficiency in other languages or dialects.
High-Stakes Decisions: The model should not be used for making high-stakes decisions without human oversight. This includes medical, legal, financial, or safety-critical decisions, among others.
3 Biases, Risks, and Limitations
The model is trained on publicly available data which in part (Arabic) was curated by our preprocessing pipeline. We used different techniqes to reduce the bias that is inadvertently present in the dataset. While efforts were made to minimize biases, it is still possible that our model, like all LLM models, may exhibit some biases.
The model is trained as an AI assistant for Arabic and English speakers, and thus it should be used to help humans to boost their productivity. In this context, it is limited to produce responses for queries in these two languages and it might not produce appropriate responses for queries in other languages.
Potential misuses include generating harmful content, spreading misinformation, or handling sensitive information. Users are urged to use the model responsibly and with discretion.
Acknowledgments
We thank Arwa Abouelseoud and Ali Al Naqbi for their help with Arabic data annotation, evaluation, and contributions to improving the Arabic data processesing steps. We also thank Xudong Han for the help in the model evaluation.
References
Appendix A Detailed Zero-Shot Evaluation Results
Table 12 and Table 13 show the detailed zero-shot evaluation results for Arabic and English, respectively.
Appendix B Jais-chat Response Examples
Below, we provide examples demonstrating various capabilities of Jais-chat in Arabic and English.
Appendix C Model Cards
Table 25 and 26 showcase model cards [MWZ+19] summarizing the details of Jais and Jais-chat, respectively.