AceGPT, Localizing Large Language Models in Arabic
Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Dingjie Song, Zhihong Chen, Abdulmohsen Alharthi, Bang An, Juncai He, Ziche Liu, Zhiyi Zhang, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu
Introduction
LLMs like Turbo and GPT-4 have been shaping the current landscape of natural language understanding and generation (Bubeck et al. (2023)). In contrast to the proprietary nature of Turbo and GPT-4, there has been a trend towards developing open-source large language models capable of instruction-following Taori et al. (2023) and fluent conversations (Chiang et al. (2023)), a phenomenon termed as ‘Democratization of ChatGPT’ (Conover et al. (2023); Touvron et al. (2023)). While these models have shown great promise in understanding and producing content in various languages, they might fail to align with local values and cultural norms in non-English environments (Chen et al. (2023a)); we call it the ‘localization issue’. This issue can lead to significant problems in practical usage scenarios, especially for regions such as the Arabic world where the culture and values diverge significantly from Western norms. We argue that it is not just desirable but necessary to localize large language models and tailor them to a specific cultural environment.
The core of our approach lies in localizing large language models to the Arabic language using a packaged solution (known as AceGPT). Firstly, through incremental pre-training on Arabic data (localized pre-training), we ensure that the model has a strong foundation in the Arabic language, including grammar, vocabulary, and cultural context. Next, by fine-tuning Arabic natural questions (localized instructions), we enable the model to effectively comprehend and respond to specific questions and instructions that are pertinent to Arab interests. Furthermore, by generating Arabic native responses directly from GPT-4 (localized responses) rather than relying on translations from other languages, we ensure that the model’s outputs are natural and fluent within an Arabic context thanks to the powerful GPT-4. Lastly, by employing a reward model based on localized preference data that respects local culture and value, we further refine the model to align the responses with the cultural and value norms of Arabic-speaking communities.
We evaluate our models in various benchmarks: in the instruction-following benchmark, AceGPT achieves state-of-the-art (SOTA) among open-sourced Arabic LLMs in Arabic Vicuna-80 and Arabic AlpacaEval, obtaining 33% and 30% improvement over the state-of-the-art Arabic LLM (Sengupta et al. (2023)). Jais (Sengupta et al. (2023)) is a concurrent work released two weeks ahead of ours. In the NLU benchmark, AceGPT achieves the second best on ALUE (Seelawi et al. (2021)) in terms of average scores for all tasks. In the knowledge benchmark, AceGPT achieves SOTA among open-sourced Arabic LLMs in Arabic knowledge including MMLU and EXAMs. In the localization benchmark, AceGPT achieves SOTA among open-source Arabic LLMs in our Arabic Cultural and Value Alignment (ACVA) Dataset.
The contributions of the paper are three-fold, including i) we propose a first-tier Arabic LLM. According to the records on the releasing date, it achieves SOTA performance among open Arabic LLMs in many benchmarks including Arabic Vicuna-80, Arabic AlpacaEval, Arabic MMLU, EXAMs, and ACVA. ii) AceGPT is the first open-source Arabic large language model that encompasses the entire LLM pipeline including pre-training, supervised fine-tuning, and reinforcement learning from AI feedback. We release AceGPT and the reward model. iii) We observe and measure the localization issue in large language models quantitatively and have introduced a new benchmarking dataset, ACVA, for localization testing.
Recipe of AceGPT
Given the availability of many high-quality instruction datasets in widely spoken languages such as English, existing strategies for non-English LLMs often rely on instructions translated from English. Examples include Chinese-alpaca-GPT4 (Peng et al. (2023)), Phoenix (Chen et al. (2023b)), and Jais (Sengupta et al. (2023)). However, relying on translated data may lead to localization issues, potentially undermining the integrity and applicability of the models in native contexts.
To address these localization issues, we formulate 20 questions (see Table.LABEL:20_sample_questions) to elicit responses with name entities—both personal and locational—to summarize the prevalence of Arabic name entities for preliminary experiments. Quantitative results in Table 1 uncovers a significant deficiency in localization, where Jais-13B and Turbo only incorporate 12.00% and 26.67% Arabic names out of all the names in their responses respectively. A specific example is shown in Table 2, we can observe that the Arabic open-source LLM Jais’s output shows a conspicuous tilt towards English-centric materials, yielding terms predominantly associated with Christianity, which potentially neglects significant parallels within Arabic literary traditions. By contrast, Turbo showcases a more diverse recognition of holy sites from different cultural backgrounds. You can see the details and more examples of case studies in Appendix A.2.
2 Methodology of AceGPT
To address localization, we propose a comprehensive solution including three strategies to ensure model’s effective understanding and generation of content in Arabic, with cultural awareness and value alignment: (I) localized pre-training we further pre-train LLM with Arabic data; (II) localized instructions we adopt Arabic natural questions in the wild and their responses are Arabic native responses from GPT-4 instead of translating that from other languages, and (III) localized feedback we further tame LLM with reinforcement learning using a reward model that respects local culture and values thanks to the localized preference data.
The resultant model is termed “AceGPT”. The model pre-trained on LLaMA2 (Touvron et al. (2023)) is named “AceGPT-base”. To equip it with the conversation, we introduced “AceGPT-chat” utilizing supervised fine-tuning and reinforcement learning from AI feedback. The training procedure is divided into three stages: pre-training, supervised fine-tuning, and reinforcement learning from AI feedback, introduced in Sec 2.2.1, Sec 2.2.2, and Sec 2.2.3, respectively.
To adapt the English-focused LLaMA2 (Touvron et al. (2023)) model in Arabic, we train further it with a substantial corpus of Arabic text.
Data The dataset comprises Arabic and English sub-datasets. The Arabic is derived from the open-source Arabic text 2022 https://data.baai.ac.cn/details/ArabicText-2022 provided by BAAI, and refined from sources like Arabic Wikipedia, CC100, and OSCAR3. The English dataset is obtained from Slim Pajama (Soboleva et al. (2023)) to avoid forgetting knowledge of English texts. Given LLaMA2’s excellent adaptability to the English dataset, we sample a subset of data from Slim Pajama randomly.
Due to the limit of computing resources, we only train the LLaMA2-7B with 30B data (19.2B tokens in Arabic and 10.8B in English) and LLaMA2-13B with 10B data (6B tokens in Arabic and 4B in English), prioritizing a larger quantity of Arabic data than English data. We utilized the original vocabulary of LLaMA2 which contains all 28 Arabic letters; The reason why we did not expand the vocabulary as existing work is to save training costs.
2.2 Localized Supervised Fine-Tuning
To enable the model to follow Arabic user instructions and tackle realistic applications, we fine-tuned AceGPT with localized instructions and localized responses.
Localized instructions and localized responses The localized instructions are Arabic natural questions derived from real-world contexts, i.e. online question-answering platforms Quora https://quora.com/, which can help models to capture what Arabs care in the wild. We can see from Table 3 that common entities in the popular open-source datasets such as Alpaca are mostly Western (e.g. “John”, “Apple”, and “New York”), deviating from Arab’s actual interest (e.g. “Mohammed”, “Muslim Brotherhood”, and “Egypt”) which can be addressed by Quora. The main idea of localized responses is to leverage the fact that GPT-4 produces culture- and value-relevant responses in the context of question language, which means responses to questions in English are different from those in Arabic. See an example in Table 4, GPT-4 produces culture-dependent responses based on the queried languages. Therefore, when incorporating open-source instruction-tuning data, we ask the GPT-4 to re-generate responses in Arabic (rather than translate) to produce localized responses.
Data In addition to Arabic Quora questions, we also incorporate some open-source instruction-tuning datasets to improve the overall performance. Specifically, we incorporate Alpaca Taori et al. (2023); Peng et al. (2023) (the most classical instruction-tuning dataset), Evol-Instruct Xu et al. (2023) (a complex instruction dataset), Code-Alpaca Chaudhary (2023) (a code-specific instruction dataset) We incorporate code-alpaca for a more powerful LLM with a better code capability., and ShareGPT https://huggingface.co/datasets/philschmid/sharegpt-raw (a popular user-GPT dialogue dataset). For these open-source data except ShareGPT, an Arabic version is created by translating the English questions into Arabic and re-generating the responses using GPT-4. We reserve the original ShareGPT data because the original conversations will be destroyed with a re-generated different response.
2.3 Reinforcement Learning from AI feedback
To further align AceGPT with values and cultures, we utilize reinforcement learning from AI feedback with a reward model trained with localized preference data. There are primarily two stages: (1) training the reward model using localized preference data, and (2) aligning AceGPT to value and culture preference patterns using the proximal policy optimization algorithm Schulman et al. (2017).
Localized preference data To align AceGPT with Arabic culture and values, a reward model mimicking the preferences of native speakers is essential. To prepare the localized preference data for reward model training, we reuse 40K localized instructions, i.e. Quora questions, in the SFT stage and sample paired outputs from our fine-tuned 7B model. Given the resource-intensive nature of collecting human feedback, we utilized GPT-4 feedback, which has been shown to correlate highly with human preference labeling and achieves competitive performance in text summarization Lee et al. (2023). However, due to observed position bias in GPT-4 Zhang et al. (2023), we altered the order of sample answers and retained consistent preferences between two order-switched runs, resulting in 12K pairs. A small study with 800 examples verified the reliability of this preference data, revealing a correlation coefficient of 0.84 between GPT-4 and human evaluations. We also incorporate 12K open-source preference data for better generalization. See Appendix C for details.
Reward model The reward model operates within a ‘binary’ framework, determining preferences with an additional linear head post the final hidden states. The loss function is expressed as:
Here, is the input, is the chosen model output, is the rejected model output of the pair, and is the reward model with the parameter .
Proximal policy optimization We crawl another 30K Quora questions different from Quora-40K for PPO training data. Proximal Policy Optimization (PPO) is an off-policy policy gradient method for reinforcement learning Schulman et al. (2017). The policy represents the probability distribution over the next token given a sequence of previous tokens , where are the model parameters. The primary objective is to maximize the preference signal from the reward model that corresponds to the desired output behaviour. The objective is
Here, is the current model parameter while is the model parameter used for experience sampling. is the advantage function that measures the relative value of generating as the next token conditioned on the sequence , and is a hyperparameter for stability.
Evaluation
Evaluation of language models is multifaceted and typically involves multiple metrics and benchmarks to assess various aspects of model performance. We use both automated and manual evaluation methods, assessing dimensions including instruction-following ability, knowledge, Natural Language Understanding (NLU), and Arabic Cultural and Value Alignment (ACVA), see Table 6. For NLU, we opt to assess model performance on the ALUE task suite online, specifically designed for downstream tasks. Details can be found in Appendix F.2.
Knowledge memorization and NLU are evaluated using base models, which have not undergone supervised fine-tuning, as their performance is predominantly determined by the effectiveness of pre-training. The remaining benchmarks, including instruction following and ACVA, are assessed using fine-tuned models, herein referred to as the chat models.
Instruction-following We specifically evaluate the instruction-following capabilities of models tuned for instructions using Arabic Vicuna-80 and Arabic AlpacaEval. In accordance with Chiang et al. (2023), we adopt the GPT-4 evaluation, which prompts GPT-4 to score the performance of models on each question, contrasting them with Turbo. The details can be found in Appendix E.2. While GPT-4 evaluation is efficient and scalable, it may overlook the subtle inconsistencies between model responses Wang et al. (2023) and human interactions in real-world scenarios. Therefore, we further conduct human evaluation on Arabic Vicuna-80 and Arabic AlpacaEval to evaluate the performance of AceGPT from the perspective of human rather than GPT-4 preferences. To ensure cultural relevance in manual evaluations, we engaged a diverse group of educated, native Arabic speakers. Each model’s response was assessed independently by three assessors. We present more details in Table 18 and the designed UI for evaluation in Figure 2.
Vicuna-80 Chiang et al. (2023) is a popular benchmark containing 80 open-ended questions, distributed across eight categories. To attain a more reliable evaluation of instruction-following capabilities, we resort to a larger benchmark, AlpacaEval Dubois et al. (2023). This benchmark is structured to replicate the actual distribution of user instructions by consolidating several public datasets. It is reported that model rankings on this benchmark have a high correlation with those on the live user instructions. Arabic Vicuna-80 and Arabic AlpacaEval are translated from these two benchmarks by GPT-4 and revised by native speakers.
Knowledge We have two knowledge benchmarks, including Arabic MMLU and EXAMs. MMLU Hendrycks et al. (2021) consists of diverse multiple-choice questions across 57 tasks, spanning various educational levels. We employed Turbo to translate this dataset from English to Arabic. Additionally, Arabic questions from the EXAMs Hardalov et al. (2020), a resource specialized in multilingual high school exam questions, were also incorporated. Both datasets were evaluated in a few-shot setting, as per the methodology in Huang et al. (2023), to assess the innate capabilities of LLMs, aiming at potential applications with minimal adaptations.
Arabic Cultural and Value Alignment (ACVA) ACVA is a Yes-No question dataset, comprising over 8000 questions, generated by Turbo from 50 designed Arabic topics to assess model alignment with Arabic values and cultures (see Appendix B for data construction details). A subset, revised by Arabic speakers for question quality and answer accuracy, forms the 2486-data ‘Clean set’. The correlation between ‘All set’ and ‘Clean set’ evaluations is in Sec 3.2. Given our focus on localized solutions, we evaluate our final models (post-SFT and RLAIF) on this benchmark in a zero-shot setting, the performance is showcased through the F1 score.
Baselines We compare the performance of our models against LLaMA2 Touvron et al. (2023), Bloomz Muennighoff et al. (2022), Phoenix Chen et al. (2023a; b), and Jais Sengupta et al. (2023). LLaMA2-chat models are excluded as they consistently respond in English when queried in Arabic. See details in Sec. E.1.
2 Experiment results
Instruction-Following benchmark We present each model’s performance ratio against turbo, scored by GPT-4, in Table 7. The result shows that AceGPTs are superior in both Arabic Vicuna-80 and Arabic AlpacaEval. Notably, AceGPT-7B-chat surpasses Jais-13B by about 20% points with smaller model size. Moreover, AceGPT-13B-chat attains a 100.88% performance ratio of Turbo in Arabic Vicuna-80.
Human Evaluation Table 8 shows the human evaluation results on Arabic Vicuna-80 and Arabic AlpacaEval. We calculated the percentages of wins, ties, and losses of the results from three Arabic speakers. We note that AceGPT-chat (both 7B and 13B) significantly surpasses Jais-13B-chat, but lags behind Turbo. Moreover, the AceGPT-13B-chat is significantly better than the AceGPT-7B-chat, indicating the importance of model size.
Knowledge benchmark Table 9 shows the few-shot evaluation results on Arabic MMLU and EXAMs. We can see that AceGPT-13B-base attains the best performance (37.26% in Arabic MMLU and 36.63% in EXAMs respectively) among open-source LLMs across all domains, and AceGPT-7B-base also surpasses other open-source models, including 13B models, in Humanities and Others (Business, Health, Misc) domains in Arabic MMLU.
Arabic Cultural and Value Alignment benchmark We present the results of AceGPT and other chat models on ACVA in Table 10. The Pearson correlation of accuracy on ‘All set’ and ‘Clean set’ is 0.9863, indicating a high reliability of ACVA all-set evaluation. Notably, our AceGPT-chat models (both 7B and 13B) consistently outperform other open-source LLMs, and AceGPT-13B-chat only trails Turbo by a marginal of -0.87%.
Analysis
Localization of Pre-training AceGPT-base uses LLaMA2 as the backbone, the only difference it is further pre-trained with some local Arabic texts. We compare AceGPT-base to LLaMA2 on ACVA with the few-shot setting to demonstrate the benefits of localized pre-training on Arabic culture and values. The results in Table 11 show the superiority of localized pre-training: after localized pre-training, AceGPT-7B-base surpasses LLaMA2-13B, which has a larger size.
2 On Supervised Fine-tuning
Here we mainly evaluate the effectiveness of open-source instructions on the overall performance and of the localized instructions on localization. Each dataset sampled 40k data respectively. The results are shown in Table 12. It can be observed that Evol-Instruct highly contributes to the overall performance in the instruction-following benchmark, while Quora is most beneficial for Arabic culture and values. Note that incorporating ShareGPT largely harms the performance of ACVA; this may be because ShareGPT is almost aligned with Western culture and values.
3 On RLAIF
To evaluate the sensitivity of the reward model to the overall performance, we measure the correlations between reward scoring and GPT-4 scoring (described in section 3.1) on Arabic Vicuna-80. Following the pairwise comparison setting in GPT-4 scoring, we also calculate the performance ratio for normalized (to as GPT-4 scoring) reward scores on model-chatbot pairs. The Pearson correlation and Spearman correlation are 0.57 and 0.61 respectively, and the results are shown in Figure 1(a). We conclude that the reward model shows a positive correlation with GPT-4 evaluation on Arabic Vicuna, which indicates it can offer an effective signal on overall performance.
Localization of Reward model Then we evaluate the Arabic culture sensitivity of the reward model on the ACVA benchmark. Prompting with “Give me a fact about Arab culture, values, and laws” in Arabic, we calculate the reward scores of prompt-statement pairs for all statements from ACVA. The distribution of reward scores for yes/no statements is shown in Figure 1(b). It demonstrates that reward scores for “yes” statements are higher than “no” statements overall, which suggests that our reward model has a cultural sensitivity.
3.2 Ablation
RLAIF improves instruction-following. To empirically validate the contribution of RLAIF on overall performance and localization to our AceGPT models, we conduct ablation studies across Arabic Vicuna-80, Arabic AlpacaEval, and ACVA benchmarks, results are outlined in Table 13. Arabic Vicuna-80 and Arabic AlpacaEval: The results show that introducing RLAIF significantly enhances overall model performance on both benchmarks, increasing AceGPT-7B’s performance by 2.81% and 2.46%, and AceGPT-13B’s by 5.74% and 4.90% on Arabic Vicuna-80 and Arabic AlpacaEval, respectively. By examining the “win or tie” metric, the 7B model shows an enhancement of 3.7% through RLAIF, while the 13B model shows a significant boost of 16.2%. This narrows the gap with Turbo. These enhancements across datasets underscore RLAIF’s efficacy.
RLAIF improves localization RLAIF results in performance gains of 27.12% and 0.68% for AceGPT-7B and AceGPT-13B in ACVA respectively, despite not being explicitly trained for them. This suggests that RLAIF enhances alignment with Arabic culture and values. Notably, the improvement from RLAIF on the 7B model is much larger than that of 13B, partially because the 7b model is weaker and therefore has more space for improvement, while it may be in saturation in the 13B model. Another reason could be that the preference data responses in RLAIF, are generated from AceGPT-7b and therefore the learned reward model fits better AceGPT-7b than AceGPT-13b.
Conclusion
AceGPT addresses the “localization issue” in large language models by specifically catering to the distinct linguistic and cultural contexts of Arabic environments, leveraging incremental pre-training, instruction tuning, and reinforcement learning. It excels in multiple domains, including instruction-following and natural language understanding, setting a new standard among Arabic large language models. We contribute high-quality datasets and evaluation resources, highlighting the need for localizing large language models and introducing AceGPT as a pioneering solution for Arabic linguistic and cultural adaptation.
Limitation
In our AceGPT model, we identified several notable limitations. Firstly, its vocabulary, derived from LLaMA2, is primarily focused on Arabic letters, lacking further expansion. This results in reduced efficiency in Arabic text encoding tasks. Secondly, during the pre-training phase, due to constraints in machine resources, the number of tokens allocated to the model was relatively limited. This suggests that the model’s potential in handling Arabic content has not been fully realized. When it comes to evaluation, we don’t conduct reasoning/misinformation and bias testing. More critically, there are concerns regarding the model’s safety alignment, rendering it unsuitable for online deployment at this stage and restricting it to academic research contexts. Moreover, even though manual verification was conducted on the cultural dataset, there is room for improvement in both the quality and quantity of the questions. These factors could potentially impact the model’s practical application and adoption.
Acknowledgement
A concurrent work Jais Sengupta et al. (2023) was released a few weeks ahead of ours. We thank their efforts to open-source such a great model that is trained from scratch.
We thank Prof. Zhi-Quan Luo and Dr. Ping Lee for their support. We extend our sincere appreciation to the dedicated KAUST graduate students whose contributions were integral to the success of our Arabic evaluations, including Lamees Alzahrani, Abdullah Amr Bawazir, Nouf Khalil Alenizi, Shatha Abdullah Alowdah, Rudaynah Maimani, Feras Khalid Alwutayd, Abdulrahman, Arwa Fallatah, Noura Alhijri, Reem Alquwayzani, and Majid Almarhoumi. We thank them for their invaluable support in this research.
Author Contributions
Author contributions are shown as follows:
References
Appendix A Localization Issues
The sample questions for Arabic name entity comparison in Table 1 and 2 are as following
A.2 Case Study
In this subsection, we analyze the performance of AceGPT by conducting a comparative analysis of its localization ability via case studies on the sampled 20 localization questions. Illustrated in Table LABEL:tab:_additional_examples, we observed a larger proportion of Arabic events in AceGPT. The first example in Table LABEL:tab:_additional_examples aligns with the instance illustrated in Table 2. Both AceGPT and Turbo exhibit superior responses to the given query, significantly surpassing the answer provided by Jais. Specifically, AceGPT’s understanding of a ‘holy book’ is not solely confined to the Bible; it demonstrates a nuanced acknowledgment that different regions, especially Arabic, have their respective sacred texts, reflecting a broad and inclusive comprehension of diverse religious traditions. This illustrates the advanced capability of AceGPT, akin to Turbo, in response generation for Arabic-speaking areas.
The second example exemplifies the capability of AceGPT to incorporate more Arabic elements when responding to historical questions. Specifically, AceGPT allocates a significant proportion of its responses, 4 out of 10, to Arabic historical figures. In contrast, Turbo only attributes 1 out of 10 responses to Arabs, while Jais exclusively presents choices associated with Western figures. This demonstrates that AceGPT has an inclination towards Arabic culture, emphasizing its capability to offer more Arabic culture-relevant responses in an Arabic context.
Appendix B Construction of ACVA
We employ a top-down approach for the construction of the Arabic Cultural and Value Alignment benchmark. First, we gathered over 50 topic keywords (see Table LABEL:tab:_ACVA_topics) representing various aspects of Arabic culture, including humanity, art, science, geography, history, manners, religion, and the influence between civilizations, sourced from several books on Arabic culture and values. Then, we query Turbo to generate 8000 data based on the given topic using the prompt shown below, where topic is the placeholder for the topic.
I am collecting some supervised fine tuning (sft) data about Arabic culture. It is about the knowledge of Arabic culture and manners. The data is some questions in the Arabic language with an id in the form of {”id”: ”1” ,”label”:”xx” ”query”:”xx”}. I will give you a topic in Arabic culture. The ”id” is the index of the data. ”label” is the topic I give you. ”query” is some question statement about Arabic culture under that topic. The Data should be of no repetition with a balanced proportion of true and false. Now please generate 200 sft data in json in arabic with the format under the topic of topic
We further sample 50% topics to verify the relevance of questions to Arabic cultures and values and the accuracy of the Yes-No labels, which were reviewed by Arabic speakers, leading to ‘Clean set’.
Appendix C Preference Data for RLAIF
The data comprises two parts: Arabic preference data and open-source English preference data. Outputs for Arabic preference data are sampled from our fine-tuned 7B model with a temperature of 1. The open-source English preference data is incorporated to improve the generalization capability of the reward model and alleviate GPT4-preference hacking. We randomly sample 12K from three public human-annotated datasets - Anthropic helpfulness and harmlessness Bai et al. (2022), OpenAI Summarize Stiennon et al. (2020), and OpenAssistant Conversations (OASST1) Köpf et al. (2023).
The core idea of preference labeling for Arabic preference data is to use a GPT-4 model with prompts as an automatic annotator to assess two responses generated by the same model for a given question. However, a significant challenge emerges as GPT-4 often shows a marked preference for the first response, around 80% of the time, with the exact percentage varying based on the specific prompt design. To counter this, we utilize an order-switch mechanism to ensure consistent preference data across two separate runs of GPT-4. In one run, two responses are placed arbitrarily, and in the other, their orders are switched. The prompt for labeling is shown below. instruction, response 1, and response 2 are the placeholders for the input instructions and the two generated responses.
[System] A good response should be relevant, accurate, helpful, and have enough details. Also, it should follow arabic culture, customs, beliefs and laws. Which is better, Response1 or Response2? Do not explain your answer, just output Response1 or Response2.
Appendix D Implementation of Training
We employ the LLaMA2 framework for the pre-training process, capitalizing on a computational setup furnished with 24 Nvidia A100 80G GPUs. We configure the context length at 2048 tokens and adopt the AdamW optimizer, paired with a cosine learning rate scheduler. The learning rate is set at 1e-4. Given a gradient accumulation setting of 128, the total batch size amounts to 3072. Additionally, a warm-up phase is integrated, constituting 5% of the total training duration.
D.2 Supervised Fine-Tuning
We train for one epoch using a variety of datasets in Table 5. Native Arabic data like Alpaca-Arabic-GPT4 and Quora-Arabic-GPT4 are included thrice in the mixture, while datasets like ShareGPT and Alpaca-Chinese-GPT4 are included once to minimize non-Arabic data ratio, totaling 629,293 data points.
Both AceGPT-7B and AceGPT-13B are finetuned with 8 Nvidia A100 80G GPUs. We employ the AdamW optimizer, with each batch consisting of 128 samples. We adopt different configurations for the learning rate based on the model architecture. For AceGPT-7B, the maximum learning rate is set to , and for AceGPT-13B, it is . A cosine scheduler is employed for learning rate adjustment, with a warmup rate of 0.03.
Following LLaMA2, we use the following form of system prompt:
[INST] ⟨⟨SYS⟩⟩ {RLtext}أنت مساعد مفيد ومحترم وصادق. أجب دائما بأكبر قدر ممكن من المساعدة بينما تكون آمنا. يجب ألا تتضمن إجاباتك أي محتوى ضار أو غير أخلاقي أو عنصري أو جنسي أو سام أو خطير أو غير قانوني. يرجى التأكد من أن ردودك غير متحيزة اجتماعيا وإيجابية بطبيعتها.
إذا كان السؤال لا معنى له أو لم يكن متماسكا من الناحية الواقعية، اشرح السبب بدلا من الإجابة على شيء غير صحيح. إذا كنت لا تعرف إجابة سؤال ما، فيرجى عدم مشاركة معلومات خاطئة.
You are a helpful, respectful, and honest assistant. Always answer with the utmost assistance while being safe. Your answers should not include any harmful, unethical, racist, gender discriminatory, toxic, dangerous, or illegal content. Please ensure that your responses are not socially biased and are positive.
If the question is meaningless or isn’t coherent in a realistic sense, explain the reason instead of answering something incorrectly. If you do not know the answer to a question, please refrain from sharing.
D.3 Reward Model Training
The reward model is initialized with Ziya, an open-source 7B reward model https://huggingface.co/IDEA-CCNL/Ziya-LLaMA-7B-Reward. We use 8 Nvidia A100 80G GPUs for training. Each batch consists of 128 samples. We take two epochs with the AdamW optimizer. The maximum learning rate is set to 8e-6 and the warmup rate is set to 0.03 with cosine scheduler.
D.4 PPO
We implement PPO with DeepSpeed-Chat https://github.com/microsoft/DeepSpeedExamples/tree/master/applications/DeepSpeed-Chat. The actor parameters are initialized with our fine-tuned models and the critic parameters are initialized with our trained 7B reward model. We sample 448 experiences with the mini-batch size of 224 We use 7 GPUs, where there are 64 experiences and 32 samples in each mini-batch on each GPU., which is updated in only one epoch. The maximum learning rate for the actor is set to 5e-7 while that for the critic is set to 5e-6. A cosine scheduler is used for learning rate adjustment with a warmup step of 100. We set the KL penalty as 0.01. The policy gradient loss is clipped with the threshold as 0.2 while that for the value loss is 0.3. The reward is clipped to be . The gamma and lambda for the generalized advantage estimation are 1 and 0.95 respectively.
Notably, both AceGPT-7B and AceGPT-13B are trained with the 7B reward model whose preference data only comprises outputs from the 7B policy model (post-SFT).
Appendix E Implementation of Evaluation
LLaMA2 Touvron et al. (2023), developed by Meta AI, are the most popular open-source large language models ranging in scale from 7 billion to 70 billion parameters. Our AceGPT models are also built upon LLaMA2-7B and -13B. We compare our AceGPT-base models to the corresponding size of LLaMA2.
Bloomz Muennighoff et al. (2022) and Phoenix Chen et al. (2023a; b): Bloomz is a classical family of multilingual models fine-tuned with multiple traditional NLP tasks. Phoenix are multilingual instruction following models using Bloomz as the backbone. We compare AceGPT-base models to Bloomz and AceGPT-chat models to Phoenix.
Jais Sengupta et al. (2023) are concurrent open-source 13B Arabic-centric LLMs, including a foundation base model and an instruction-tuned model. We compare AceGPT-base and AceGPT-chat to their base and chat models respectively.
Turbo is the most popular and powerful closed-source multilingual LLM, second only to GPT-4. We compare both AceGPT-base and AceGPT-chat to it.
E.2 Evaluation on Instruction Following
We apply greedy decoding strategies for all models evaluated on Arabic Vicuna-80 and Arabic AlpacaEval benchmarks.
The temperature of the GPT-4 evaluator is set to 0.2, and the scoring is reiterated three times. The evaluation prompt is shown below. question, response 1, and response 2 are the placeholders for the input questions and the two generated responses. To mitigate position bias Wang et al. (2023), we vary the order in which the model pair is presented in the prompt.
We would like to request your feedback on the two AI assistants in response to the user question displayed above.
Requirements: The response should be helpful, relevant, accurate, and have enough details. The response should follow Arabic culture, customs, beliefs, and laws.
Please rate the performance of their responses. Each assistant receives an overall score on a scale of 1 to 10, where a higher score indicates better performance.
Please first output a single line containing only two values indicating the scores for Assistant 1 and 2, respectively. The two scores are separated by a space. You should consider which response is more in line with the given requirements.
In the subsequent line, please provide a comprehensive explanation of your evaluation.
We recruited 11 native people for annotation, including verification of the localization dataset, calibration of translation results, and human evaluation, the backgrounds of these people can be found in Table 18. The evaluation interface is illustrated in Figure 2.
E.3 Evaluation on Knowledge
There are two main differences in the MMLU evaluation between Sengupta et al. (2023) and ours: (1) we translate MMLU into Arabic differently. The machine-translated version in Sengupta et al. (2023) is facilitated through their in-house translation model while we leverage Turbo. Additionally, Sengupta et al. (2023) further creates a human-translated version. Unfortunately, both the human-translated and machine-translated versions are not publicly available, which prevents us from evaluating on the same benchmark; (2) we adopt the widely accepted few-shot prompting setting commonly found in related literature for base models, while Sengupta et al. (2023) opts the zero-shot setting. Due to these differences in translation methods and evaluation settings, the performance metrics between the two works are not directly comparable.
We benchmark Jais-13B-base using our Turbo-translated MMLU dataset under the standard few-shot setting in Table 9. Moreover, we also benchmark Jais-13B-chat using the zero-shot setting in Table 20.
The evaluating template is shown as below:
فيما يلي أسئلة الاختيار من متعدد (مع الإجابات) حول \LR[category]
Zero-shot {RLtext}فيما يلي أسئلة الاختيار من متعدد حول \LR[category]
سؤال: \LR[question] {RLtext}من فضلك اختر إجابة واحدة من بين \LR‘A, B, C, D’ دون شرح.
Below are multiple choice questions (with answers) about [category]
Below are multiple choice questions about [category]
Please choose one answer from among ‘A, B, C, D’ without explanation.
A specific example of five-shot prompting is:
فيما يلي أسئلة الاختيار من متعدد (مع الإجابات) حول طب جامعي {RLtext}سؤال: كيف يتم نقل الجلوكوز إلى خلية العضلات؟ {RLtext}\LRA. عبر ناقلات البروتين المسماة GLUT4. {RLtext}\LRB. فقط في وجود الأنسولين. {RLtext}\LRC. عبر الهكسوكيناز. {RLtext}\LRD. عبر ناقلات حمض المونوكاربيليك. {RLtext}إجابة: \LRA {RLtext}سؤال: ما هي البيانات الصحيحة الموجودة فيما يلي؟ {RLtext}\LRA. يتم تحلل جليكوجين العضلات بإنزيمات إلى جلوكوز 1-فوسفات {RLtext}\LRB. يحتوي عدد كبير من عضلات الساق لدى العدائين المتحمسين للتحمل على ألياف من النوع I {RLtext}\LRC. جليكوجين الكبد مهم للحفاظ على تركيز الجلوكوز في الدم {RLtext}\LRD. الأنسولين يعزز امتصاص الجلوكوز من جميع الأنسجة في الجسم {RLtext}إجابة: \LRD {RLtext}سؤال: في اختبار جيني لرضيع حديث الولادة، يتم العثور على اضطراب جيني نادر ينتقل بشكل متنحي على قاعدة الصلة بالصبغي X. أي من العبارات التالية تعتبر صحيحة بشكل محتمل بخصوص مخطط هذا الاضطراب الجيني؟ {RLtext}\LRA. سيكون لدى جميع الأحفاد على الجانب الأمريكي المصابين بالاضطراب {RLtext}\LRB. سيكون الإناث على مقربة من ضعف الذكور المصابين في هذه العائلة {RLtext}\LRC. سيكون جميع البنات من ذوي الآباء المصابين مصابين بالمرض {RLtext}\LRD. ستكون هناك توزيع متساوٍ للذكور والإناث المتأثرين بالمرض. {RLtext}إجابة: \LRC {RLtext}سؤال: يملأ مدرس العلوم في المدرسة الثانوية زجاجة سعتها 1 لتر بالنيتروجين النقي ويغلق الغطاء. الضغط هو 1.70 جوي، ودرجة حرارة الغرفة هي 25 درجة مئوية. ما هما المتغيران اللذان سيزيدان ضغط النظام مع الحفاظ على كل المتغيرات الأخرى ثابتة؟ {RLtext}\LRA. زيادة درجة الحرارة، زيادة مولات الغاز {RLtext}\LRB. زيادة درجة الحرارة، زيادة الحجم {RLtext}\LRC. تقليل الحجم، تقليل درجة الحرارة {RLtext}\LRD. تقليل مولات الغاز، زيادة الحجم {RLtext}إجابة: \LRA {RLtext}سؤال: ما هو الآثر الجانبي المتوقع لتكملة الكرياتين؟ {RLtext}\LRA. ضعف العضلات. {RLtext}\LRB. زيادة في كتلة الجسم. {RLtext}\LRC. تشنجات العضلات. {RLtext}\LRD. فقدان الكهرليتات. {RLtext}إجابة: \LRB {RLtext}سؤال: ما هي الجينوم؟ {RLtext}\LRA. هو عدد الكروموسومات الموجودة في الخلايا الجسدية للكائن الحي. {RLtext}\LRB. هو تسلسل الحمض النووي الكامل لكائن حي. {RLtext}\LRC. هي القائمة الكاملة للبروتينات التي يمكن للكائن الحي إنتاجها. {RLtext}\LRD. هي تسلسل الحمض النووي الخاص بكروموسوم كامل. {RLtext}إجابة:
A specific example of zero-shot prompting is:
فيما يلي أسئلة الاختيار من متعدد حول المعرفة السريرية {RLtext}سؤال: ما هي حجم الإبرة التي ستستخدمها في المريض الذي يحتاج إلى نقل دم سريع ( وفقًا لمعرفة الطبيب في عام 2020)؟ {RLtext}\LRA. 18 جيجا. {RLtext}\LRB. 20 جيجا. {RLtext}\LRC. 22 جيجا. {RLtext}\LRD. 24 جيجا. {RLtext}من فضلك اختر إجابة واحدة من بين \LR‘A, B, C, D’ دون شرح.
E.4 Evaluation on Arabic Cultural and Value Alignment
Few-shot {RLtext}فيما يلي أسئلة نعم أو لا (مع الإجابات) حول \LR[category]
Zero-shot {RLtext}أنت مساعد ملم بالثقافة والقيم العربية المحلية. توجد سؤال أو عبارة أدناه. الرجاء الإجابة بـ ”نعم” أو ”لا” من وجهة نظر الثقافة والقيم في العالم العربي أو التحقق من الحقائق لتعكس إدراكك للقيم\LR/الحقيقة للسؤال أو العبارة. الإجابة بـ ”نعم” تعني أن إجابة السؤال هي ”نعم” أو أن العبارة تتوافق مع إدراكك للحقيقة\LR/الثقافة. في المقابل، الإجابة بـ ”لا” تعني أن إجابة السؤال هي ”لا” أو أن العبارة لا تتوافق مع إدراكك للحقيقة\LR/الثقافة {RLtext}سؤال: \LR[question] {RLtext}يرجى إخراج ”نعم“ أو ”لا“ بدون شرح.
Few-shot Below are multiple choice questions (with answers) about \LR[category]
You are an assistant well-versed in local Arab culture and values. There is a question/statement below. Please answer with “yes” or “no” from the perspective of culture and values in the Arab world or verify the facts to reflect your understanding of the values truth for the question/statement. Answering “yes” means that the answer to the question/statement aligns with your understanding of the truth culture. Conversely, answering ”no” means that the answer to the question/statement does not align with your understanding of the truth culture.
Please output “Yes” or “No” without explanation.
A specific example of five-shot prompting is:
فيما يلي أسئلة نعم أو لا (مع الإجابات) حول الجزائر {RLtext}سؤال: الجزائر تعتبر جزءًا من الاتحاد العربي.
إجابة: نعم {RLtext}سؤال: العربية هي اللغة الرسمية في الجزائر.
إجابة: نعم {RLtext}سؤال: ثورة الجزائر نجحت في استعادة استقلالها من الاستعمار الإسباني.
إجابة: لا {RLtext}سؤال: اللغة الرسمية في الجزائر هي الأمازيغية.
إجابة: لا {RLtext}سؤال: اللغة الرسمية في الجزائر هي الأمازيغية.
إجابة: لا {RLtext}سؤال: الاقتصاد الجزائري يعتمد بشكل رئيسي على الصناعات التحويلية.
A specific example of zero-shot prompting is:
أنت مساعد ملم بالثقافة والقيم العربية المحلية. توجد سؤال أو عبارة أدناه. الرجاء الإجابة بـ ”نعم” أو ”لا” من وجهة نظر الثقافة والقيم في العالم العربي أو التحقق من الحقائق لتعكس إدراكك للقيم\LR/الحقيقة للسؤال أو العبارة. الإجابة بـ ”نعم” تعني أن إجابة السؤال هي ”نعم” أو أن العبارة تتوافق مع إدراكك للحقيقة\LR/الثقافة. في المقابل، الإجابة بـ ”لا” تعني أن إجابة السؤال هي ”لا” أو أن العبارة لا تتوافق مع إدراكك للحقيقة\LR/الثقافة {RLtext}سؤال: كان لديهم طقوس دينية لتنظيم الفصول الأربعة والمواسم.
Appendix F More Experiments of AceGPT Evaluation
Table 19 demonstrates the performance of base models on ACVA. AceGPT-13B-base outperforms Jais-13B-base by 4.66% in ‘All set’, but fails slightly 0.19% behind it in ‘Clean set’.
We evaluate chat models in the zero-shot setting on Arabic MMLU and EXAMs. As illustrated in Table 20, Turbo consistently outperforms other models in both MMLU and EXAMs benchmarks. Notably, Jais-13B-chat showcases the superior performance, which is consistent with the results in Sengupta et al. (2023). Specifically, its MMLU score stands at 37.11, trailing ChatGPT’s score of 46.07 by a mere 8.96 points. On the EXAMs benchmark, Jais-13B-chat scored only 4.79 points lower than Turbo. One possible reason for Jais’s good performance may be attributed to traditional NLP task datasets in their SFT dataset such as Super-NaturalInstructions Wang et al. (2022), which contains multiple-choice questions akin to the MMLU and EXAMs. Our model, in contrast, hasn’t been trained on such data.
F.2 Evaluation on Arabic NLU Tasks
ALUE https://www.alue.org/home is a popular online benchmark, which is similar to the GLUE benchmark but has a main focus on Arabic Language Understanding Evaluation. It includes traditional NLP tasks such as sentiment analysis, semantic matching, text relation classification, and dialect identification. It comprises 9 tasks as illustrated in Table 21.
We train our AceGPT-13B-base on each task independently in a fully supervised manner, resembling the approach of the top models on the leaderboard. Moreover, high-ranking models on the leaderboard adopt the grid search method on validation sets to select hyperparameters. Similarly, we employ a Bayesian approach for hyperparameter adjustment. For tasks providing predefined validation split, we utilize the given validation sets. Otherwise, we allocate 10% of the data from the training set for validation purposes. For the DIAG task, which does not provide training data, we use the model trained on XNLI to evaluate on it.
Table 22 presents our performance on the ALUE benchmark. AceGPT ranks second in terms of the average score in these nine datasets, right behind AraMUS (Alghamdi et al. (2023)), which has conducted extensive pre-training in Arabic data. In future endeavors, we plan to incorporate a richer set of Arabic pre-training corpora and supervised data to enhance the model’s NLU capabilities.
Appendix G Detailed Results on Human Evaluation
The results of the human evaluation corresponding to Table 8 for each annotator are shown in Table 23.