ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic

Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi

Introduction

Language models (LMs) exploiting self-supervised learning such as BERT Devlin et al. 2019 and RoBERTa Liu et al. 2019a have recently emerged as powerful transfer learning tools that help improve a very wide range of natural language processing (NLP) tasks. Multilingual LMs such as mBERT Devlin et al. 2019 and XLM-RoBERTa (XLM-R) Conneau et al. 2020 have also been introduced, but are usually outperformed by monolingual models pre-trained with larger vocabulary and bigger language-specific datasets Virtanen et al. 2019; Antoun et al. 2020; Dadas et al. 2020; de Vries et al. 2019; Le et al. 2020; Martin et al. 2020; Nguyen and Tuan Nguyen 2020.

Since LMs are costly to pre-train, it is important to keep in mind the end goals they will serve once developed. For example, (i) in addition to their utility on ‘standard’ data, it is useful to endow them with ability to excel on wider real world settings such as in social media. Some existing LMs do not meet this need since they were trained on datasets that do not sufficiently capture the nuances of social media language (e.g., frequent use of abbreviations, emoticons, and hashtags; playful character repetitions; neologisms and informal language). It is also desirable to build models able to (ii) serve diverse communities (e.g., speakers of dialects of a given language), rather than focusing only on mainstream varieties. In addition, once created, models should be (iii) usable in energy efficient scenarios. This means that, for example, medium-to-large models with competitive performance should be preferred to large-to-mega models.

A related issue is (iv) how LMs are evaluated. Progress in NLP hinges on our ability to carry out meaningful comparisons across tasks, on carefully designed benchmarks. Although several benchmarks have been introduced to evaluate LMs, the majority of these are either exclusively in English (e.g., DecaNLP McCann et al. 2018, GLUE Wang et al. 2018, SuperGLUE Wang et al. 2019) or use machine translation in their training splits (e.g., XTREME Hu et al. 2020). Again, useful as these benchmarks are, this circumvents our ability to measure progress in real-world settings (e.g., training and evaluation on native vs. translated data) for both cross-lingual NLP and in monolingual, non-English environments.

Context. Our objective is to showcase a scenario where we build LMs that meet all four needs listed above. That is, we describe novel LMs that (i) excel across domains, including social media, (ii) can serve diverse communities, and (iii) perform well compared to larger (more energy hungry) models (iv) on a novel, standardized benchmark. We choose Arabic as the context for our work since it is a widely spoken language (∼400\sim 400M native speakers), with a large number of diverse dialects differing among themselves and from the standard variety, Modern Standard Arabic (MSA). Arabic is also covered by the popular mBERT Devlin et al. 2019 and XLM-R Conneau et al. 2020, which provides us a setup for meaningful comparisons. That is, not only are we able to empirically measure monolingual vs. multilingual performance under robust conditions using our new benchmark, ARLUE, but we can also demonstrate how our base-sized models outperform (or at least are on par with) larger models (i.e., XLM-RLarge, which is ∼3.4×\sim 3.4\times larger than our models). In the context of our work, we also show how the currently best-performing model dedicated to Arabic, AraBERT Antoun et al. 2020, suffers from a number of issues. These include (a) not making use of easily accessible data across domains and, more seriously, (b) limited ability to handle Arabic dialects and (c) narrow evaluation. We rectify all these limitations.

Our contributions. With our stated goals in mind, we introduce ARBERT and MARBERT, two Arabic-focused LMs exploiting large-to-massive diverse datasets. For evaluation, we also introduce a novel ARabic natural Language Understanding Evaluation benchmark (ARLUE). ARLUE is composed of 4242 different datasets, making it by far the largest and most diverse Arabic NLP benchmark we know of. We arrange ARLUE into six coherent cluster tasks and methodically evaluate on each independent dataset as well as each cluster task, ultimately reporting a single ARLUE score. Our models establish new state-of-the-art (SOTA) on the majority of tasks, across all cluster tasks. Our goal is for ARLUE to serve the critical need for measuring progress on Arabic, and facilitate evaluation of multilingual and Arabic LMs. To summarize, we offer the following contributions:

We develop ARBERT and MARBERT, two novel Arabic-specific Transformer LMs pre-trained on very large and diverse datasets to facilitate transfer learning on MSA as well as Arabic dialects.

We introduce ARLUE, a new benchmark developed by collecting and standardizing splits on 4242 datasets across six different Arabic language understanding cluster tasks, thereby facilitating measurement of progress on Arabic and multilingual NLP.

We fine-tune our new powerful models on ARLUE and provide an extensive set of comparisons to available models. Our models achieve new SOTA on all task clusters in 3737 out of 4848 individual datasets and a SOTA ARLUE score.

The rest of the paper is organized as follows: In Section 2, we provide an overview of Arabic LMs. Section 3 describes our Arabic pre-tained models. We evaluate our models on downstream tasks in Section 4, and present our benchmark ARLUE and evaluation on it in Section 5. Section 6 is an overview of related work. We conclude in Section 7. We now introduce existing Arabic LMs.

Arabic LMs

The term Arabic refers to a collection of languages, language varieties, and dialects. The standard variety of Arabic is MSA, and there exists a large number of dialects that are usually defined at the level of the region or country Abdul-Mageed et al. 2020a; Abdul-Mageed et al. 2021a; Abdul-Mageed et al. 2021b. A number of Arabic LMs has been developed. The most notable among these is AraBERT Antoun et al. 2020, which is trained with the same architecture as BERT Devlin et al. 2019 and uses the BERTBase configuration. AraBERT is trained on 2323GB of Arabic text, making ∼70\sim 70M sentences and 33B words, from Arabic Wikipedia, the Open Source International dataset (OSIAN) Zeroual et al. 2019 (3.53.5M news articles from 2424 Arab countries), and 1.51.5B words Corpus from El-Khair 2016 (55M articles extracted from 1010 news sources). Antoun et al. 2020 evaluate AraBERT on three Arabic downstream tasks. These are (1) sentiment analysis from six different datasets: HARD Elnagar et al. 2018, ASTD Nabil et al. 2015, ArsenTD-Lev Baly et al. 2019, LABR Aly and Atiya 2013, and ArSaS Elmadany et al. 2018. (2) NER, with the ANERcorp Benajiba and Rosso 2007, and (3) Arabic QA, on Arabic-SQuAD and ARCD Mozannar et al. 2019 datasets. Another Arabic LM that was also introduced is ArabicBERT Safaya et al. 2020, which is similarly based on BERT architecture. ArabicBERT was pre-trained on two datasets only, Arabic Wikipedia and Arabic OSACAR Suárez et al. 2019. Since both of these datasets are already included in AraBERT, and Arabic OSACAR https://oscar-corpus.com. has significant duplicates, we compare to AraBERT only. GigaBERT Lan et al. 2020, an Arabic and English LM designed with code-switching data in mind, was also introduced. Since GigaBERT is very recent, we could not compare to it. However, we note that our pre-training datasets are much larger (i.e., 15.615.6B tokens for MARBERT vs. 4.34.3B Arabic tokens for GigaBERT) and more diverse.

Our Models

We train ARBERT on 6161GB of MSA text (6.56.5B tokens) from the following sources:

Books (Hindawi). We collect and pre-process 1,8001,800 Arabic books from the public Arabic bookstore Hindawi. https://www.hindawi.org/books/.

El-Khair. This is a 55M news articles dataset from 1010 major news sources covering eight Arab countries from El-Khair 2016.

Gigaword. We use Arabic Gigaword 5th Edition from the Linguistic Data Consortium (LDC). https://catalog.ldc.upenn.edu/LDC2011T11. The dataset is a comprehensive archive of newswire text from multiple Arabic news sources.

OSCAR. This is the MSA and Egyptian Arabic portion of the Open Super-large Crawled Almanach coRpus Suárez et al. 2019, https://oscar-corpus.com/. a huge multilingual subset from Common Crawl https://commoncrawl.org. obtained using language identification and filtering.

OSIAN. The Open Source International Arabic News Corpus (OSIAN) Zeroual et al. 2019 consists of 3.53.5 million articles from 3131 news sources in 2424 Arab countries.

Wikipedia Arabic. We download and use the December 20192019 dump of Arabic Wikipedia. We use WikiExtractor https://github.com/attardi/wikiextractor. to extract articles and remove markup from the dump.

We provide relevant size and token count statistics about the datasets in Table 1.

1.2 Training Procedure

Pre-processing. To prepare the raw data for pre-training, we perform light pre-processing. This helps retain a faithful representation of the naturally occurring text. We only remove diacritics and replace URLs, user mentions, and hashtags that may exist in any of the collections with the generic string tokens URL, USER, and HASHTAG, respectively. We do not perform any further pre-processing of the data before splitting the text off to wordPieces Schuster and Nakajima 2012. Multilingual models such as mBERT and XLM-R have 55K (out of 110110K) and 1414K (out of 250250K) Arabic WordPieces, respectively, in their vocabularies. AraBERT employs a vocabulary of 6060K (out of 6464K). The empty 44K vocabulary bin is reserved for additional wordPieces, if needed. For ARBERT, we use a larger vocabulary of 100100K WordPieces. For tokenization, we use the WordPiece tokenizer Wu et al. 2016 provided by Devlin et al. 2019.

Pre-training. For ARBERT, we follow Devlin et al. 2019’s pre-training setup. To generate each training input sequence, we use the whole word masking, where 15%15\% of the NN input tokens are selected for replacement. These tokens are replaced 80%80\% of the time with the [MASK] token, 10%10\% with a random token, and 10%10\% with the original token. We use the original implementation of BERT in the TensorFlow framework. https://github.com/google-research/bert. As mentioned, we use the same network architecture as BERTBase: 1212 layers, 768768 hidden units, 1212 heads, for a total of ∼163\sim 163M parameters. We use a batch size of 256256 sequences and a maximum sequence length of 128128 tokens (256256 sequences ×\times 128128 tokens = 32,76832,768 tokens/batch) for 88M steps, which is approximately 4242 epochs over the 6.56.5B tokens. For all our models, we use a learning rate of 11e−4-4. We pre-train the model on one Google Cloud TPU with eight cores (v2.82.8) from TensorFlow Research Cloud (TFRC). https://www.tensorflow.org/tfrc. Training took ∼16\sim 16 days, for 4242 epochs over all the tokens. Table 2 shows a comparison of ARBERT with mBERT, XLM-R, AraBERT, and MARBERT (see Section 3.2) in terms of data sources and size, vocabulary size, and model parameters.

2 MARBERT

As we pointed out in Sections 1 and 2, Arabic has a large number of diverse dialects. Most of these dialects are under-studied due to rarity of resources. Multilingual models such as mBERT and XLM-R are trained on mostly MSA data, which is also the case for AraBERT and ARBERT. As such, these models are not best suited for downstream tasks involving dialectal Arabic. To treat this issue, we use a large Twitter dataset to pre-train a new model, MARBERT, from scratch as we describe next.

To pre-train MARBERT, we randomly sample 11B Arabic tweets from a large in-house dataset of about 66B tweets. We only include tweets with at least three Arabic words, based on character string matching, regardless whether the tweet has non-Arabic string or not. That is, we do not remove non-Arabic so long as the tweet meets the three Arabic word criterion. The dataset makes up 128128GB of text (15.615.6B tokens).

2.2 Training Procedure

Pre-processing. We employ the same pre-processing as ARBERT.

Pre-training. We use the same network architecture as BERTBase, but without the next sentence prediction (NSP) objective since tweets are short. It was also shown that NSP is not crucial for model performance Liu et al. 2019a. We use the same vocabulary size (100100K wordPieces) as ARBERT, and MARBERT also has ∼160\sim 160M parameters. We train MARBERT for 1717M steps (∼36\sim 36 epochs) with a batch size of 256256 and a maximum sequence length of 128128. Training took ∼40\sim 40 days on one Google Cloud TPU (eight cores). We now present a comparison between our models and popular multilingual models as well as AraBERT.

3 Model Comparison

Our models compare to mBERT Devlin et al. 2019, XLM-R Conneau et al. 2020 (base and large), and AraBERT Antoun et al. 2020 in terms of training data size, vocabulary size, and overall model capacity as we summarize in Table 2. In terms of the actual Arabic variety involved, Devlin et al. 2019 train mBERT with Wikipedia Arabic data, which is MSA. XLM-R Conneau et al. 2020 is trained on Common Crawl data, which likely involves a small amount of Arabic dialects. AraBERT is trained on MSA data only. ARBERT is trained on a large collection of MSA datasets. Unlike all other models, our MARBERT model is trained on Twitter data, which involves both MSA and diverse dialects. We now describe our fine-tuning setup.

4 Model Fine-Tuning

We evaluate our models by fine-tuning them on a wide range of tasks, which we thematically organize into six clusters: (1) sentiment analysis (SA), (2) social meaning (SM) (i.e., age and gender, dangerous and hateful speech, emotion, irony, and sarcasm), (3) topic classification (TC), (4) dialect identification (DI), (5) named entity recognition (NER), and (6) question answering (QA). For all classification tasks reported in this paper, we compare our models to four other models: mBERT, XLM-RBase, XLM-RLarge, and AraBERT. We note that XLM-RLarge is ∼3.4×\sim 3.4\times larger than any of our own models (∼550\sim 550M parameters vs. ∼160\sim 160M). We offer two main types of evaluation: on (i) individual tasks, which allows us to compare to other works on each individual dataset (4848 classification tasks on 4242 datasets), and (ii) ARLUE clusters (six task clusters).

For all reported experiments, we follow the same light pre-processing we use for pre-training. For all individual tasks and ARLUE task clusters, we fine-tune on the respective training splits for 2525 epochs, identifying the best epoch on development data, and reporting on both development and test data. A minority of datasets came with no development split from source, and so we identify and report the best epoch only on test data for these. This allows us to compare all the models under the same conditions (2525 epochs) and report a fair comparison to the respective original works. For all ARLUE cluster tasks, we identify the best epoch exclusively on our development sets (shown in Table 10). We typically use the exact data splits provided by original authors of each dataset. Whenever no clear splits are available, or in cases where expensive cross-validation was used in source, we divide the data following a standard 80%80\% training, 10%10\% development, and 10%10\% test split. For all experiments, whether on individual tasks or ARLUE task clusters, we use the Adam optimizer Kingma and Ba 2015 with input sequence length of 256256, a batch size of 3232, and a learning rate of 22e−6-6. These values were identified in initial experiments based on development data of a few tasks. NER and QA are expetions, where we use sequence lengths of 128128 and 384384, respectively; a batch sizes of 1616 for both; and a learning rate of 22e−6-6 and 33e−5-5, respectively. We now introduce individual tasks.

Individual Downstream Tasks

Datasets. We fine-tune the language models on all publicly available SA datasets we could find in addition to those we acquired directly from authors. In total, we have the following 1717 MSA and DA datasets: AJGT Alomari et al. 2017, AraNETSent Abdul-Mageed et al. 2020b, AraSenTi-Tweet Al-Twairesh et al. 2017, ArSarcasmSent Farha and Magdy 2020, ArSAS Elmadany et al. 2018, ArSenD-Lev Baly et al. 2019, ASTD Nabil et al. 2015, AWATIF Abdul-Mageed and Diab 2012, BBNS & SYTS Salameh et al. 2015, CAMelSent Obeid et al. 2020, HARD Elnagar et al. 2018, LABR Aly and Atiya 2013, TwitterAbdullah Abdulla et al. 2013, TwitterSaad, www.kaggle.com/mksaad/arabic-sentiment-twitter. and SemEval-2017 Rosenthal et al. 2017. Details about the datasets and their splits are in Section A.1.

Baselines. We compare to the STOA listed in Table 3 and Table 4 captions. For all datasets with no baseline in Table 3, we consider AraBERT our baseline. Details about SA baselines are in Section A.2.

Results. To facilitate comparison to previous works with the appropriate evaluation metrics, we split our results into two tables: We show results in F1PN in Table 3 and F1 in Table 4. We typically bold the best result on each dataset. Our models achieve best results in 1313 out of the 1717 classification tasks reported in the two tables combined, while XLM-R (which is a much larger model) outperforms our models in the 44 remaining tasks. We also note that XLM-R acquires better results than AraBERT in the majority of tasks, a trend that continues for the rest of tasks. Results also clearly show that MARBERT is more powerful than than ARBERT. This is due to MARBERT’s larger and more diverse pre-training data, especially that many of the SA datasets involve dialects and come from social media.

2 Social Meaning Tasks

We collectively refer to a host of tasks as social meaning. These are age and gender detection; dangerous, hateful, and offensive speech detection; emotion detection; irony detection; and sarcasm detection. We now describe datasets we use for each of these tasks.

Datasets. For both age and gender, we use Arap-Tweet Zaghouani and Charfi 2018. We use AraDan Alshehri et al. 2020 for dangerous speech. For offensive language and hate speech, we use the dataset released in the shared task (sub-tasks A and B) of offensive speech by Mubarak et al. 2020. We also use AraNETEmo Abdul-Mageed et al. 2020b, IDAT@FIRE2019 Ghanem et al. 2019, and ArSarcasm Farha and Magdy 2020 for emotion, irony and sarcasm, respectively. More information about these datasets and their splits is in Appendix B.1.

Baselines. Baselines for social meaning tasks are the SOTA listed in Table 5 caption. Details about each baseline is in Appendix B.2.

Results. As Table 5 shows, our models acquire best results on all eight tasks. Of these, MARBERT achieves best performance on seven tasks, while ARBERT is marginally better than MARBERT on one task (irony@FIRE2019). The sizeable gains MARBERT achieves reflects its superiority on social media tasks. On average, our models are 9.83\bf 9.83 F1 better than all previous SOTA.

3 Topic Classification

Classifying documents by topic is a classical task that still has practical utility. We use four TC datasets, as follows:

Datasets. We fine-tune on Arabic News Text (ANT) Chouigui et al. 2017 under three pre-taining settings (title only, text only, and title+text.), Khaleej Abbas et al. 2011, and OSAC Saad and Ashour 2010. Details about these datasets and the classes therein are in Appendix C.1.

Baselines. Since, to the best of our knowledge, there are no published results exploiting deep learning on TC, we consider AraBERT a strong baseline.

Results. As Table 6 shows, ARBERT acquires best results on both OSAC and Khaleej, and the title-only setting of ANT. AraBERT slightly outperforms our models on the text-only and title+text settings of ANT.

4 Dialect Identification

Arabic dialect identification can be performed at different levels of granularity, including binary (i.e., MSA-DA), regional (e.g., Gulf, Levantine), country level (e.g., Algeria, Morocco), and recently province level (e.g., the Egyptian province of Cairo, the Saudi province of Al-Madinah) Abdul-Mageed et al. 2020a; Abdul-Mageed et al. 2021b.

Datasets. We fine-tune our models on the following datasets: Arabic Online Commentary (AOC) Zaidan and Callison-Burch 2014, ArSarcasmDia Farha and Magdy 2020, ArSarcasmDia carries regional dialect labels. MADAR (sub-task 2) Bouamor et al. 2019, NADI-2020 Abdul-Mageed et al. 2020a, and QADI Abdelali et al. 2020. Details about these datasets are in Table D.1.

Baselines. Our baselines are marked in Table 7 caption. Details about the baselines are in Table D.2.

Results. As Table 7 shows, our models outperform all SOTA as well as the baseline AraBERT across all classification levels with sizeable margins. These results reflect the powerful and diverse dialectal representation of MARBERT, enabling it to serve wider communities. Although ARBERT is developed mainly for MSA, it also outperforms all other models.

5 Named Entity Recognition

We fine-tune the models on five NER datasets.

Datasets. We use ACE03NW and ACE03BN Mitchell et al. 2004, ACE04NW Mitchell et al. 2004, ANERcorp Benajiba and Rosso 2007, and TW-NER Darwish 2013. Table E.1 shows the distribution of named entity classes across the five datasets.

Baseline. We compare our results with SOTA presented by Khalifa and Shaalan 2019 and follow them in focusing on person (PER), location (LOC) and organization (ORG) named entity labels while setting other labels to the unnamed entity (O). Details about Khalifa and Shaalan 2019 SOTA models are in Appendix E.2.

Results. As Table 8 shows, our models outperform SOTA on two out of the five NER datasets. We note that even though SOTA Khalifa and Shaalan 2019 employ a complex combination of CNNs and character-level LSTMs, which may explain their better results on two datasets, MARBERT still achieves highest performance on the social media dataset (TW-NER).

6 Question Answering

Datasets. We use ARCD Mozannar et al. 2019 and the three human translated Arabic test sections of the XTREME benchmark Hu et al. 2020: MLQA Lewis et al. 2020, XQuAD Artetxe et al. 2020, and TyDi QA Artetxe et al. 2020. Details about these datasets are in Table F.1.

Baselines. We compare to Antoun et al. 2020 and consider their system a baseline on ARCD. We follow the same splits they used where we fine-tune on Arabic SQuAD Mozannar et al. 2019 and 50%50\% of ARCD and test on the remaining 50%50\% of ARCD (ARCD-test). For all other experiments, we fine-tune on the Arabic machine translated SQuAD (AR-XTREME) from the XTREME multilingual benchmark Hu et al. 2020 and test on the human translated test sets listed above. Our baselines in these is Hu et al. 2020’s mBERTBase model on gold (human) data.

Results. As is standard, we report QA results in terms of both Exact Match (EM) and F1. We find that results with ARBERT and MARBERT on QA are not competitive, a clear discrepancy from what we have observed thus far on other tasks. We hypothesize this is because the two models are pre-trained with a sequence length of only 128128, which does not allow them to sufficiently capture both a question and its likely answer within the same sequence window during the pre-training. In addition, MARBERT is not trained on Wikipedia data from where some questions come. To rectify this, we further pre-train the stronger model, MARBERT, on the same MSA data as ARBERT in addition to AraNews dataset Nagoudi et al. 2020 (8.68.6GB), but with a bigger sequence length of 512512 tokens for 4040 epochs. We call this further pre-trained model MARBERT-v2, noting it has 2929B tokens. As Table 9 shows, MARBERT-v2 acquires best performance on all but one test set, where XLM-RLarge marginally outperforms us (only in F1).

ARLUE

We concatenate the corresponding splits of the individual datasets to form ARLUE, which is a conglomerate of task clusters. That is, we concatenate all training data from each group of tasks into a single TRAIN, all development into a single DEV, and all test into a single TEST. One exception is the social meaning tasks whose data we keep independent (see ARLUESM below). Table 10 shows a summary of the ARLUE datasets. Again, ARLUESM datasets are kept independent, but to provide a summary of all ARLUE datasets we collate the numbers in Table 10. We now briefly describe how we merge individual datasets into ARLUE.

ARLUESenti. To construct ARLUESenti, we collapse the labels very negative into negative, very positive into positive, and objective into neutral, and remove the mixed class. This gives us the 33 classes negative, positive, and neutral for ARLUESenti. Details are in Table A.1.

ARLUESM. We refer to the different social meaning datasets collectively as ARLUESM. We do not merge these datasets to preserve the conceptual coherence specific to each of the tasks. Details about individual datasets in ARLUESM are in B.1.

ARLUETopic. We straightforwardly merge the TC datasets to form ARLUETopic, without modifying any class labels. Details of ARLUETopic data are in Table C.1.

ARLUEDia. We construct three ARLUEDia categories. Namely, we concatenate the AOC and AraSarcasmDia MSA-DA classes to form ARLUEDia-B (binary) and the region level classes from the same two datasets to acquire ARLUEDia-R (4-classes, region). We then merge the country classes from the rest of datasets to get ARLUEDia-C (21-classes, country). Details are in Table D.1. ARLUENER & ARLUEQA. We straightforwardly concatenate all corresponding splits from the different NER and QA datasets to form ARLUENER and ARLUEQA, respectively. Details of each of these task clusters data are in Tables E.1 and F.1, respectively.

2 Evaluation on ARLUE

We present results on each task cluster independently using the relevant metric for both the development split (Table 11) and test split (Table 12). Inspired by McCann et al. 2018 and Wang et al. 2018 who score NLP systems based on their performance on multiple datasets, we introduce an ARLUE score. The ARLUE score is simply the macro-average of the different scores across all task clusters, weighting each task equally. Following Wang et al. 2018, for tasks with multiple metrics (e.g., accuracy and F1), we use an unweighted average of the metrics as the score for the task when computing the overall macro-average. As Table 12 shows, our MARBERT-v2 model achieves the highest ARLUE score (77.4077.40), followed by XLM-RL (76.5576.55) and ARBERT (76.0776.07). We also note that in spite of its superiority on social data, MARBERT ranks top 44. This is due to MARBERT suffering on the QA tasks (due to its short input sequence length), and to a lesser extent on NER and TC.

Related Work

English and Multilingual LMs. Pre-trained LMs exploiting a self-supervised objective with masking such as BERT Devlin et al. 2019 and RoBERTa Liu et al. 2019b have revolutionized NLP. Multilingual versions of these models such as mBERT and XLM-RoBERTa Conneau et al. 2020 were also pre-trained. Other models with different objectives and/or architectures such as ALBERT Lan et al. 2019, T5 Raffel et al. 2020 and its multilingual version, mT5 Xue et al. 2021, and GPT3 Brown et al. 2020 were also introduced. More information about BERT-inspired LMs can be found in Rogers et al. 2020.

Non-English LMs. Several models dedicated to individual languages other than English have been developed. These include AraBERT Antoun et al. 2020 and ArabicBERT Safaya et al. 2020 for Arabic, Bertje for Dutch de Vries et al. 2019, CamemBERT Martin et al. 2020 and FlauBERT Le et al. 2020 for French, PhoBERT for Vietnamese Nguyen and Tuan Nguyen 2020, and the models presented by Virtanen et al. 2019 for Finnish, Dadas et al. 2020 for Polish, and Malmsten et al. 2020 for Swedish. Pyysalo et al. 2020 also create monolingual LMs for 42 languages exploiting Wikipedia data. Our models contributed to this growing work of dedicated LMs, and has the advantage of covering a wide range of dialects. Our MARBERT and MARBERT-v2 models are also trained with a massive scale social media dataset, endowing them with a remarkable ability for real-world downstream tasks.

NLP Benchmarks. In recent years, several NLP benchmarks were designed for comparative evaluation of pre-trained LMs. For English, McCann et al. 2018 introduced NLP Decathlon (DecaNLP) which combines 1010 common NLP datasets/tasks. Wang et al. 2018 proposed GLUE, a popular benchmark for evaluating nine NLP tasks. Wang et al. 2019 also presented SuperGLUE, a more challenging benchmark than GLUE covering seven tasks. In the cross-lingual setting, Hu et al. 2020 provide a Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark for the evaluation of cross-lingual transfer learning covering nine tasks for 4040 languages (1212 language families). ARLUE complements these benchmarking efforts, and is focused on Arabic and its dialects. ARLUE is also diverse (involves 4242 datasets) and challenging (our best ARLUE score is at 77.4077.40).

Conclusion

We presented our efforts to develop two powerful Transformer-based language models for Arabic. Our models are trained on large-to-massive datasets that cover different domains and text genres, including social media. By pre-training MARBERT and MARBERT-v2 on dialectal Arabic, we aim at enabling downstream NLP technologies that serve wider and more diverse communities. Our best models perform better than (or on par with) XLM-RLarge (∼3.4×\sim 3.4\times larger than our models), and hence are more energy efficient at inference time. Our models are also significantly better than AraBERT, the currently best-performing Arabic pre-trained LM. We also introduced AraLU, a large and diverse benchmark for Arabic NLU composed of 4242 datasets thematically organized into six main task clusters. ARLUE fills a critical gap in Arabic and multilingual NLP, and promises to help propel innovation and facilitate meaningful comparisons in the field. Our models are publicly available. We also plan to publicly release our ARLUE benchmark. In the future, we plan to explore self-training our language models as a way to improve performance following Khalifa et al. 2021. We also plan to investigate developing more energy efficient models.

Acknowledgements

We gratefully acknowledge support from the Natural Sciences and Engineering Research Council of Canada, the Social Sciences and Humanities Research Council of Canada, Canadian Foundation for Innovation, Compute Canada and UBC ARC-Sockeye (https://doi.org/10.14288/SOCKEYE). We also thank the Google TFRC program for providing us with free TPU access.

Ethical Considerations

Although our language models are pre-trained using datasets that were public at the time of collection, parts of these datasets might become private or get removed (e.g., tweets that are deleted by users). For this reason, we will not release or re-distribute any of the pre-training datasets. Data coverage is another important consideration: Our datasets have wide coverage, and one of our contributions is offering models that can serve more diverse communities in better ways than existing models. However, our models may still carry biases that we have not tested for and hence we recommend they be used with caution. Finally, our models deliver better performance than larger-sized models and as such are more energy conserving. However, smaller models that can achieve simply ‘good enough’ results should also be desirable. This is part of our own future research, and the community at large is invited to develop novel methods that are more environment friendly.

References

Appendix A Sentiment Analysis

AJGT. The Arabic Jordanian General Tweets (AJGT) dataset Alomari et al. 2017 covers MSA and Jordanian Arabic, with 900900 positive and 900900 negative posts.

AraNETSent. Abdul-Mageed et al. 2020b collect 1515 datasets in both MSA and dialects from Abdul-Mageed and Diab 2012 (AWATIF), Abdul-Mageed et al. 2014 (SAMAR), Abdulla et al. 2013; Nabil et al. 2015; Kiritchenko et al. 2016; Aly and Atiya 2013; Salameh et al. 2015; Rosenthal et al. 2017; Alomari et al. 2017; Mohammad et al. 2018, and Baly et al. 2019. These datasets carry both binary (negative and positive) and three-way (negative, neutral, and positive) labels, but Abdul-Mageed et al. 2020b map them into binary sentiment only.

AraSenTi-Tweet. This comprises 17,57317,573 gold labeled MSA and Saudi Arabic tweets by Al-Twairesh et al. 2017.

ArSarcasmSent This sarcasm dataset is labeled with sentiment tags by Farha and Magdy 2020 who extract it from ASTD Nabil et al. 2015 (10,54710,547 tweets) and SemEval-2017 Task 4 Rosenthal et al. 2017 (8,0758,075 tweets).

ArSAS. This Arabic Speech Act and Sentiment (ArSAS) corpus Elmadany et al. 2018 consists of tweets annotated with sentiment tags.

ArSenD-Lev. The Arabic Sentiment Twitter Dataset for LEVantine dialect (ArSenD-Lev) Baly et al. 2019 has 4,0004,000 tweets retrieved from the Levant region.

ASTD. This is a collection of 10,00610,006 Egyptian tweets by Nabil et al. 2015.

AWATIF. This is an MSA dataset from newswire, Wikipedia, and web fora introduced by Abdul-Mageed and Diab 2012.

BBNS & SYTS. The BBN blog posts Sentiment (BBNS) and Syria Tweets Sentiment (SYTS) are introduced by Salameh et al. 2015.

CAMelSent. Obeid et al. 2020 merge training and development data from ArSAS Elmadany et al. 2018, ASTD Nabil et al. 2015, SemEval Rosenthal et al. 2017, and ArSenTD Al-Twairesh et al. 2017 to create a new training dataset (∼24\sim 24K) and evaluate on the independent test sets from each of these sources.

HARD. The Hotel Arabic Reviews Dataset (HARD) Elnagar et al. 2018 consists of 93,70093,700 MSA and dialect hotel reviews.

LABR. The Large Arabic Book Review Corpus Aly and Atiya 2013 has 63,25763,257 book reviews from Goodreads, www.goodreads.com. each rated with a 11-55 stars scale.

TwitterAbdullah. For ease of reference, we assign a name to this and other unnamed datasets. This is a dataset of 2,0002,000 MSA and Jordanian Arabic tweets manually labeled by Abdulla et al. 2013.

TwitterSaad. This dataset is collected using an emoji lexicon by Moatez Saad in 20192019 and is available on Kaggle. www.kaggle.com/mksaad/arabic-sentiment-twitter-corpus.

SemEval-2017. This is the SemEval-2017 sentiment analysis in Arabic Twitter task datasetby Rosenthal et al. 2017.

A.2 SA Baselines

For SA, we compare to the following STOA:

Antoun et al. 2020. We compare to best results reported by the authors on five SA datasets: HARD, balanced ASTD (which we refer to as ASTD-B), ArSenTD-Lev, AJGT, and the unbalanced positive and negative classes for LABR. They split each dataset into 80/20 for Train/Test, respectively, and report in accuracy using the best epoch identified on test data. For a valid comparison, we follow their data splits and evaluation set up.

Obeid et al. 2020. They fine-tune mBERT and AraBERT on the merged CAMelsent datasets and report in F1F_{1}PN, which is the macro F1F_{1} score over the positive and negative classes only (while neglecting the neutral class).

Abdul-Mageed et al. 2020b. They fine-tune mBERT on the AraNETSent data and report results in F1F_{1} score on test data.

A.3 SA Evaluation on DEV

Table A.2 shows results of SA on DEV for datasets where there is a development split.

Appendix B Social Meaning

Age and Gender. For both age and gender, we use the Arap-Tweet dataset Zaghouani and Charfi 2018, which covers 1717 different countries from 1111 Arab regions. We follow the 80-10-10 data split of AraNet Abdul-Mageed et al. 2020b.

Dangerous Speech. We use the dangerous speech AraDang dataset from Alshehri et al. 2020, which is composed of tweets manually labeled with dangerous and safe tags.

Offensive Language and Hate Speech. We use manually labeled data from the shared task of offensive speech Mubarak et al. 2020. http://edinburghnlp.inf.ed.ac.uk/workshops/OSACT4. The shared task is divided into two sub-tasks: sub-task A: detecting if a tweet is offensive or not-offensive, and sub-task B: detecting if a tweet is hate-speech or not-hate-speech.

Emotion. We use the AraNeTemo dataset from Abdul-Mageed et al. 2020b, which is created by merging two datasets from Alhuzali et al. 2018.

Irony. We use the irony identification dataset for Arabic tweets released by IDAT@FIRE2019 shared task Ghanem et al. 2019, following Abdul-Mageed et al. 2020b data splits.

Sarcasm. We use the ArSarcasm dataset developed by Farha and Magdy 2020.

More details about these datasets are in Table B.1.

B.2 SM Baselines

Age and Gender. We compare to AraNET Abdul-Mageed et al. 2020b age and gender models, trained by fine-tuning mBERT. The authors report 51.4251.42 and 65.3065.30 F1F_{1} on age and gender, respectively.

Dangerous Speech. We compare to Alshehri et al. 2020, who report a best of 59.6059.60 F1 on test with an mBERT model fined-tuned on emotion data.

Emotion. We compare to Abdul-Mageed et al. 2020b, who acquire 60.3260.32 F1F_{1} on test with a fine-tuned mBERT.

Hate Speech. The best results on the offensive and hate speech shared task Mubarak et al. 2020 are at 9595 F1 score and are reported by Husain 2020, who employ heavy feature engineering with SVMs. Since our focus is on methods exploiting language models, we compare to Djandji et al. 2020 who rank second in the shared task with a fine-tuned AraBERT (83.4183.41 F1 on test).

Irony. We compare to Zhang and Abdul-Mageed 2019a who fine-tune mBERT on the irony task, with an auxiliary author profiling task, and report 82.482.4 F1 on test.

Offensive Language. We compare to the best results on the offensive sub-task Mubarak et al. 2020 reported by Hassan et al. 2020. They propose an ensemble of SVMs, CNN-BiLSTM, and mBERT with majority voting and acquire 90.5190.51 F1.

Sarcasm. We compare to Farha and Magdy 2020 who train a BiLSTM model using the AraSarcasm dataset, reporting 46.0046.00 F1 score.

B.3 SM Evaluation on DEV

Table B.2 shows results of the social meaning tasks on development splits.

Appendix C Topic Classification

Arabic News Text. Chouigui et al. 2017 build the Arabic news text (ANT) dataset from transcribed Tunisian radio broadcasts.

Khaleej. Abbas et al. 2011 created the Khaleej from Gulf Arabic websites.

OSAC. Saad and Ashour 2010 collect OSAC from news articles.

C.2 TC Evaluation on DEV

Results of TC tasks on DEV data are in Table C.2.

Appendix D Dialect Identification

We introduce each dataset briefly here and provide a description summary of all datasets in Table D.1.

Arabic Online Commentary (AOC). This is a repository of 33M Arabic comments on online news Zaidan and Callison-Burch 2014. It is labeled with MSA and three regional dialects (Egyptian, Gulf, and Levantine).

ArSarcasmDia. This dataset is developed by Farha and Magdy 2020 for sarcasm detection but also carries regional dialect labels from the set {Egyptian, Gulf, Levantine, Maghrebi}.

MADAR. Sub-task 2 of the MADAR shared task Bouamor et al. 2019 https://camel.abudhabi.nyu.edu/madar-shared-task-2019/. is focused on user-level dialect identification with manually-curated country labels (n=2121).

NADI-2020. The first Nuanced Arabic Dialect Identification shared task (NADI 2020) Abdul-Mageed et al. 2020a https://github.com/UBC-NLP/nadi. targets country level (n=2121) as well as province level (n=100100) dialects.

QADI. The QCRI Arabic Dialect Identification (QADI) dataset Abdelali et al. 2020 is labeled at the country level (n=1818).

Details of the datasets are in Table D.1.

D.2 DIA Baselines

Elaraby and Abdul-Mageed 2018 report three levels of classification on AOC data: (1) MSA vs. DA (87.2387.23 accuracy), (2) regional (i.e., Egyptian, Gulf, and Levantine) (87.8187.81 accuracy), and (3) MSA, Egyptian, Gulf, and Levantine (accuracy of 82.4582.45). Their best results are based on BiLSTM.

Abdelali et al. 2020 fine-tune AraBERT on the QADI dataset. They report 60.660.6 F1.

Zhang and Abdul-Mageed 2019b developed the top ranked system in MADAR sub-task 22, with 48.7648.76 accuracy and 34.8734.87 F1 at tweet level.

Talafha et al. 2020 developed NADI sub-task 1 (country level) winning system, an ensemble of fine-tuned AraBERT (26.7826.78 F1).

El Mekki et al. 2020 developed NADI sub-task 2 (province level) winning system using a combination of word and character n-grams to fine-tune AraBERT (6.086.08 F1).

AraBERT. For ArSarcasmDia, where no dialect id system was previously developed, we consider a fine-tuned AraBERT a baseline.

D.3 DIA Evaluation on DEV

Table D.2 shows results of the dialect identification tasks on development splits.

Appendix E Named Entity Recognition

Table E.1 and Table E.2 show the data splits across our NER datasets, and the results of all our models on the development splits.

E.2 NER Baselines

Khalifa and Shaalan 2019 apply CNNs and BiLSTMs and report F1 scores on test sets, as follows: 88.7788.77 (ANERcorp), 91.4791.47 (ACE03NW), 94.9294.92 (ACE03BN), 91.2091.20 (ACE04NW), and 65.3465.34 (Twitter). We use their exact data splits.

Appendix F Question Answering Datasets

ARCD. Mozannar et al. 2019 use crowdsourcing to develop the Arabic Reading Comprehension Dataset. We use the same ARCD data splits used by Antoun et al. 2020.

MLQA. This MultiLingual Question Answering benchmark is proposed by Lewis et al. 2020. It consists of over 55K extractive question-answer instances in SQuAD format in seven languages, including Arabic.

XQuAD. This Cross-lingual Question Answering Dataset Artetxe et al. 2020 consists of 1,1901,190 question-answer pairs and 240240 paragraphs from SQuAD v1.1 Rajpurkar et al. 2016 translated into ten languages (including Arabic) by professional translators.

TyDi QA. The TyDi QA dataset Artetxe et al. 2020 is manually curated and covers 1111 languages (including Arabic). We focus on the “Gold” passage task only.