SERENGETI: Massively Multilingual Language Models for Africa
Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba Inciarte
Introduction
Pretraining NLP models with a language modeling objective has gained popularity as a precursor to task-specific finetuning Ettinger (2020). Pretrained models like BERT Devlin et al. (2019), ELMo Peters et al. (2018), Roberta Liu et al. (2019), GPT Radford et al. (2018, 2019); Brown et al. (2020a), and BART Lewis et al. (2020) have advanced the state of the art in a wide variety of tasks, demonstrating how these models acquire valuable, generalizable linguistic information during the pretraining process. However, training language-specific models is possible for only a few languages which have large amounts of data. A popular alternative has been pretrained multilingual language models (mPLM) such as mBERT Devlin et al. (2019) and XML-R Conneau et al. (2020). mPLMs are trained on large amounts of unlabelled data from multiple languages so that low resource languages may benefit from shared vocabulary and other linguistic information from high-resource and similar languages in the model. The vast majority of the world’s languages today remain uncovered by mPLMs, however.
African languages are no exception. Although there are few mPLMs that support a small number of African languages Devlin et al. (2019); Ogueji et al. (2021); Nzeyimana and Niyongabo Rubungo (2022); Alabi et al. (2022a); Jude Ogundepo et al. (2022); Conneau et al. (2020), these cover only a total of languages. This is grossly inadequate considering that Africa is believed to be home to languages Eberhard et al. (2021). Each of these languages encapsulates unique features that are essential in preserving linguistic diversity. The same way every species embodies essential value to the natural ecosystem, each language plays a crucial role in the linguistic ecosystem. That is, each language encodes knowledge about people, their traditions, wisdom, and environment, as well as how it is that they interact with the sum of the concepts in their own culture Adebara and Abdul-Mageed (2022). This in turn allows people and communities to preserve and transmit their knowledge, values, unique modes of thinking, meaning and expression, history, culture, traditions, and memory to next generations, while participating in society and constructing their future UNESCO 66260 (2022).
Language technology plays an important role in building inclusive knowledge societies, providing access to education and information, supporting freedom of expression, cultural and linguistic diversity, and further stimulating innovation. This technology thus has great impact on multiple domains, including education, government, health, recreation, among others. This motivates adequate representation of African languages in the ongoing technological revolution. This is also likely to connect Africa to the rest of the world. Building technologies for African languages may also aid languages that may be at risk of falling into a state of disuse at an alarming rate, thus hopefully preventing subsequent language death that may become inevitable Adebara and Abdul-Mageed (2022).
Developing LMs that represent a large number of African languages is therefore very crucial for achieving progress in Afrocentric NLP Adebara and Abdul-Mageed (2022) and indeed in addressing issues related to representation bias in artificial intelligence and linguistic diversity - two research themes of international relevance Bender et al. (2021). Motivated by this call for Afrocentric NLP, we introduce SERENGETI. SERENGETI is a massively multilingual language model exploiting a large manually-curated dataset for African languages and language varieties. These languages belong to language families and are written in different scripts. In addition to these African languages, SERENGETI is also pretrained on the top most spoken languages globally.
We also introduce AfroNLU, an extensive benchmark exploiting different datasets across different languages and language varieties for various NLP tasks. For even richer evaluation, we also apply our models to an African language identification task covering all the languages in our pretraining. To the best of our knowledge, AfroNLU is the most extensive and inclusive evaluation benchmark proposed to date for African NLP.
Our contributions in this work are as follows: (1) we collect a large dataset of African languages and language varieties and exploit it to develop SERENGETI. (2) we propose AfroNLU, a new extensive benchmark for African NLU that has the widest and most inclusive coverage for African NLP today. (3) we benchmark SERENGETI on AfroNLU and show through meaningful comparisons how our model excels and acquire new SOTA. (4) we offer a linguistically motivated analysis of model performance substantiated in language genealogy, allowing us for the first time to derive insights across the widest range of African languages in the African NLP literature to date.
The rest of the paper is organized as follows: In Section 2 we discuss related work. We describe genealogical information in Section 3. Next, we give a detailed description of SERENGETI in Section 4. In Section 5 we describe AfroNLU, the benchmark we create. We present performance of SERENGETI in Section 6 and compare it to other mPLMs. We conclude in Section 7, and outline a number of limitations and use cases for our work in Section 8 and Section 9.
Related Work
Afrocentric NLP. An Afrocentric approach to technology development is crucial for African languages. An afrocentric approach will mean that what technologies to build and how to build, evaluate, and deploy them arises from the needs of local African communities Adebara and Abdul-Mageed (2022). We provide more details in Section B in the Appendix.
African Language Models. Here, we briefly describe language models covering any number of African languages. Since we develop encoder-only models in this work, we will restrict our discussion to this category of models. We provide information about the African languages covered by these models in Table 1.
AfriBERTa Ogueji et al. (2021) is trained using a Transformer with the standard masked language modelling objective and covers African languages. The pretraining corpus for this model is small (only million tokens), when compared to many other models. AfroLM Dossou et al. (2022) supports African languages, the largest number of African languages before SERENGETI. It is trained on a multi-domain dataset from various sources Adelani et al. (2022a); Alabi et al. (2022b); Jude Ogundepo et al. (2022); Niyongabo et al. (2020). It uses a self-active learning framework and achieves SOTA on NER, sentiment analysis, and text classification. Afro-XLM-R Alabi et al. (2022a) uses language adaptation on the most-resourced African languages and three other high-resource foreign languages widely used in Africa (i.e., English, French, and Arabic) simultaneously to provide a single model for cross-lingual transfer learning. Authors show that Afro-XLM-R has competitive results with AfriBERTa and XLM-R on NER, topic classification, and news classification. KINYaBERT Nzeyimana and Niyongabo Rubungo (2022) uses a two-tier BERT architecture that involves a morphological analyzer and explicitly represents morphological information for Kinyawanda–a morphologically rich African language. Authors show that KINYaBERT achieves good convergence and accuracy, and is robust on multiple downstream tasks. mBERT Devlin et al. (2019) is a multilingual variant of BERT trained on languages including four African languages. XLM-R Conneau et al. (2020) uses a Transformer based architecture and obtains SOTA on cross-lingual classification, sequence labeling, and question answering on languages including eight African languages.
Genealogy of African Languages
Genealogical or genetic classification groups languages based on their historical and evolutionary relationships. Genetically related languages are often classified into similar families in a hierarchical tree like structure that shows the level of similarity between the languages. Languages with a higher degree of similarity belong to the same class while languages with a lower degree of similarity are further subdivided into different classes and subclasses. Two closely related languages can therefore be viewed as sisters of the same parent language/ancestor–they are languages that evolved over time and/or space from an older parent language Gerhardt (2020). Typological classification differs from geneological classification in that the former is based on grammatical features or types Vossen (2020). For instance, a typological classification would group tone languages together, or split languages based on their morphological structure into, for instance, isolating or agglutinating languages. Despite this difference, languages that belong to the same family often share similar typological information Gerhardt (2020). For example, most Benue-Congo languages are tone languages Williamson (2006). In the case of African languages, where typological information is scarcely available Adebara and Abdul-Mageed (2022); Güldemann (2018), utilizing genetic classes may be a useful way to determine typological information. If the typological information of one language in a group is known, we may make a sensible assumption that other languages in that group perhaps share similar features with minor variations. We use geneological classification information in evaluating SERENGETI’s behaviour. Specifically, we investigate the relationship between language similarity and model performance in zero-shot scenarios for South African languages in some datasets in our benchmark. We use classification information from Ethnologue Eberhard et al. (2021) in all our analyses. We provide a broad overview of the families in our models under six broad ancestors in Section D in the Appendix.
SERENGETI
SERENGETI is pretrained using GB of data comprising a multi-domain, multi-script collection of texts that we manually curate. The pretraining data covers African languages and the most spoken languages globally (i.e., Arabic, English, French, German, Greek, Italian, Portuguese, Russian, Spanish, and Turkish). The multi-domain dataset comprises texts from religious, news, government documents, health documents, and existing corpora written in five scripts from the set {Arabic, Coptic, Ethiopic, Latin, and Vai}. For the top ten foreign languages, we randomly select M paragraphs from Wikipedia for each language to use in our overall pretraining data. We provide further details of the pretraining data in Section C in the Appendix. We also show all languages in our pretraining data in Tables F.1, F.2, and F.3.
2 Preprocessing
To prepare the raw data for pretraining, we perform light preprocessing to retain a faithful representation of the naturally occurring text. Specifically, we ensure that images and non-text materials are not in our dataset by using regular expression and manual curation techniques. We do not perform any further preprocessing of the data before splitting the text off into tokens. For tokenization, we use a WordPiece tokenizer Song et al. (2021). We experiment with two vocabulary sizes, K and K.
3 SERENGETI Models
We pretrain both Electra style Clark et al. (2020b); Chi et al. (2021) as well as XLM-R style Conneau et al. (2020) models, as follows.
SERENGETI-E110 and SERENGETI-E250. We first pretrain Electra Chi et al. (2021) style models. Electra uses a multilingual replaced token detection (MRTD) objective for training. Unlike other training objectives, the goal of MRTD is to distinguish real input tokens from corrupted tokens. Models built with this objective are pretrained as discriminators rather than generators. We train the models with two vocabulary sizes, K and K, and hence refer to them as SERENGETI-E110 and SERENGETI-E250. Each of these models has 12 layers and attention heads. We pretrain each model for epochs with a sequence length of , a learning rate of and a batch size of and for the SERENGETI-E110 and SERENGETI-E250, respectively. We pre-train the models on Google Cloud TPU with cores (v) from TensorFlow Research Cloud (TFRC). https://www.tensorflow.org/tfrc.
SERENGETI Model. Apart form the Electra models, we also experiment with an XLM-R base architecture. We train the model with a K vocabulary size for epochs. This model has layers and attention heads, a sequence length of and a batch size of . We pre-train this model on M AMD Pod GPUs with G ram. Our XLM-R model has better performance compared to the Electra models as we will show. We provide information about each model we build and compare with in Table 2.
AfroNLU Benchmark
Our goal is to evaluate our models extensively, and so we combine all available datasets we could acquire to create an evaluation benchmark that we refer to as AfroNLU. AfroNLU is composed of seven different tasks, covering both token and sentence level tasks, across different datasets. The benchmark covers a total of different languages and language varieties. In addition we evaluate our best model (SERENGETI) on an African language identification (LID) task covering all the languages in our pretraining collection. For LID, we use two datasets to test SERENGETI. This puts AfroNLU at a total of 20 different datasets and eight different tasks. To the best of our knowledge, our evaluation benchmark is the most extensive compared to previous published research. We provide detailed statistics of the datasets comprising AfroNLU in Table 3. We also provide a detailed comparison of our AfroNLU benchmark with evaluation data from other models in Table 4. We now describe each of the downstream tasks in AfroNLU.
We evaluate our models on NER datasets across multiple languages. We use MasakhaNER data Ifeoluwa Adelani et al. (2021), WikiAnn Pan et al. (2017); Rahimi et al. (2019), Yoruba-Twi NER data Alabi et al. (2020), Distance Supervision NER (DS NER) Data Hedderich et al. (2020) and multiple NER data from SADiLaR. For our experiments, we use the region aggregates on MasakhaNER. Specifically, we use MasakhaNER-east, MasakhaNER-west, and MasakhaNER-eastwest. MasakhaNER-east includes NER data for Amharic, Kinyawanda, Luganda, Luo, and Swahili. MasakhaNER-west includes NER data for Hausa, Igbo, Nigerian-Pidgin, Wolof, and Yoruba. MasakhaNER-eastwest, on the other hand, includes a combination of MasakhaneNER-east and MasakhaneNER-west. Data from SADiLaR cover ten indigenous South African languages and is annotated for person, organisation, location, and miscellaneous named entities. Miscellaneous named entities refer to all rigid designators that do not fall into one of the other categories, including temporal expressions (dates and times), URLs, numerical expressions, publications, names of languages, nationalities, among others. More details about the datasets are in Table 3.
2 Part of Speech Tagging
We test our models on POS tagging datasets for Igbo taken from IgboNLP Onyenwe et al. (2018, 2019). In Table 3, we provide the statistical details for the dataset.
3 Phrase Chunks
We evaluate our models on phrase chunks datasets for ten Indigenous languages of South Africa (see Table 3). The data has annotations for noun, verb, adjective, adverbial, and prepositional phrase chunks. Words not belonging to these phrase types are labelled with the tag O.
4 Sentiment Analysis
We finetune our model on three sentiment analysis datasets, including Bambara Sentiment dataset Diallo et al. (2021), YOSM–a new Yorùbá Sentiment Corpus for Movie Reviews Shode et al. (2022), and the Nigerian Pidgin sentiment dataset Oyewusi et al. (2020), respectively. Some details of these datasets is in Table 3.
5 News classification
We use news classification datasets for Amharic Azime and Mohammed (2021), Kinyarwanda Niyongabo et al. (2020), Kirundi Niyongabo et al. (2020), and Swahili David (2020a, b). The Amharic dataset contains six classes–news, sport, politics, international news, business, and entertainment. The Swahili dataset also has six categories including local news, international, finance, health, sports, and entertainment. The datasets for Kinyarwanda and Kirundi have and categories each, respectively. Again, data statistics are in Table 3.
6 Topic classification
We include topic classification datasets for Yorùbá and Hausa Hedderich et al. (2020). The Yorùbá and Hausa datasets contain news titles collected from VOA Hausa and BBC Yorùbá news sites. The Yorùbá dataset has seven topics–Nigeria, Africa, world, entertainment, health, sports, and politics, while the Hausa dataset is categorized into five topics - Nigeria, Africa, world, health, and politics. In Table 3, we provide details about the data split sizes.
7 Question Answering
We use TYDIA question answering dataset Clark et al. (2020a). The dataset has a primary task and a gold passage task. The primary task has two subtasks, one for passage selection and another that is a minimal answer span. For the passage selection subtask, a list of passages is given and the required response is either the index of the passage where the answer to the question is or null (if no answer exists in the passage). The minimal answer span subtask on the other hand gives a full article and the expected answer is either the start and end byte indices of the minimal span that answers the question, yes or no response, or null (if no minimal answer exists). For the gold passage task, a correct answer is predicted from a passage containing one answer. This is similar to existing reading comprehension. We use the Kiswahili dataset alone, since it is the only African language in the dataset. Details about the data splits can be found in Table 3.
8 Language Identification
We also evaluate SERENGETI on the task of language identification (LID). LID focuses on identifying the human language a piece of text or speech segment belongs to, making automatic LID an important first step in processing human language appropriately Tjandra et al. (2021); Thara and Poornachandran (2021). We use datasets from AfroLID Adebara et al. (2022b) for this task. AfroLID data is a multi-genre, multi-script dataset for African languages. We compare the performance of AfroLID data on our models with performance on AfroLID tool. To ensure a fair comparison, the data used for AfroLID is completely different from the data used for SERENGETI. We also evaluate our LID model on AfriSenti dataset Muhammad et al. (2022); Yimam et al. (2020).
Experimental Setup and Evaluation
We evaluate SERENGETI on eight task clusters in the benchmark, and report results on our Test set in Table 5. We also report performance on our Dev set in Table E.1 (Appendix). For each task cluster, we finetune for a maximum of epochs with a patience value of five. We compare results from SERENGETI, SERENGETI-E110, and SERENGETI-E250 to encoder-only models covering any number of African languages. Specifically, we compare with XLMR, mBERT, Afro-XLMR, and AfriBERTa. We report the results of each experiment as an average of three runs, showing the standard deviation. We also evaluate SERENGETI on language identification and show results on Afrolid in Table 6 and on Afrisenti in Table 7. For multilingual datasets in each task, we show evaluation results per language, comparing the performance of various models in Table E.4 in the Appendix.
We report the results for seven of our eight tasks in Table 5.
Named Entity Recognition (NER). SERENGETI sets a new SOTA on six out of eight datasets on the NER cluster. The lowest across all models are on NCHLT and Yoruba-Twi datasets (on both Dev and Test). SERENGETI achieves best performance on both of these datasets on Test (with on the first and on the second).
Phrase Chunking. SERENGETI outperforms all models on the phrase chunking task on both Dev and Test data, reaching on Test.
Part of Speech (POS) Tagging. In the POS tagging task, SERENGETI outperformed all other models in the Dev. and Test sets.
News Classification. Our SERENGETI outperforms other models on three out of four datasets on Test data (and on two datasets on Dev).Our SERENGETI-E110 outperforms SERENGETI on one dataset in Dev and Test sets. We do not report SOTA results for Amharic, Kirnews, and Kinnews datasets because their authors report performance in accuracy (and so are not comparable to our results). We show performance of SERENGETI on each category in the news classification cluster in Figure E.1 in the Appendix.
Sentiment Analysis. SERENGETI-E250 outperforms other models on one out of three tasks in our sentiment analysis task cluster. Afro-XMLR and AfriBERTa outperform other models on one each. To further investigate performance, we conduct an error analysis on the three sentiment datasets (see Figure E.2 in the Appendix).
Topic Classification. AfriBERTa outperforms other models on both tasks in our topic classification cluster, followed by SERENGETI. We show confusion matrices for Hausa and Yoruba topic classification in Figure E.3 in the Appendix.
Language Identification. SERENGETI outperforms AfroLID on AfroLID and AfriSenti data (see Table 6 and 7 for details). We also compare the performance of SERENGETI to AfroLID, and FrancA publicly available LID tool covering African languages., on the African languages represented in Franc in Table E.3 (Appendix). SERENGETI outperforms AfroLID and Franc with an average score of . SERENGETI outperforms both models on languages and has similar results with AfroLID on languages. Next, we evaluate the performance of SERENGETI on Creole languages. Again, we record improvement in results for Creole languages when compared with AfroLID. SERENGETI outperforms AfroLID in out of languages and acquires similar scores on languages. We assume that the addition of the ten most spoken languages to the pretraining data for SERENGETI may have helped the model learn the Creoles better. This is because Creoles share some features including vocabularies and syntax with some of those top ten languages.
2 Error Analysis
In the sentiment analysis cluster, best performance is recorded for positive categories while negative categories have the worst performance. A fine-grained analysis of the Yoruba sentiment dataset found that SERENGETI failed to correctly categorize sentiment if the polarity item(s) were not seen in training, can be associated with both positive and negative sentiments, the polarity item(s) is a negation, or if ambivalent markers are present in the sentence. We provide a table showing examples of each type of error we found in Table E.2 in the Appendix. For the news classification task, politics and tourism are the best performing classes while education and relationships have the worst performance on kirnews and kinnews respectively. It is important to mention that the worst performing categories do not have the smallest data sizes. For the topic classification, the best performance is on the world class for Hausa topic modelling while entertainment and sport have best performance for Yoruba. The worst performance is on Nigeria and health for Hausa and Yoruba topic datasets respectively.
3 Imbalanced Distribution
We find imbalances in the class distributions for all datasets except YOSM. We find a positive correlation between the size of each category in a dataset and the model accuracy. We also find a positive correlation with the number of examples in a specific class and the accuracy we acquire. We provide confusion matrices that represents the sizes of each category and the performance of SERENGETI in Figures E.4, E.5, and E.6 in the Appendix.
4 Genealogy & Language Contact
Our preliminary analyses show that language similarity may improve model performance in zero-shot settings. This we believe is due to high cross-lingual transfer information Conneau et al. (2020) from similar languages. Similar languages often share many features (e.g., vocabulary, syntax, and script) sometimes up to a point of mutual intelligibility Nassenstein (2019); Arndt (2015); Roy-Campbell (2006). Languages in contact may also have such similarities. By language in contact, we mean all languages that speakers of a specific language interact with and influence. A language can be in contact with another due to trade, geographic proximity, migration, or even colonization. Languages in contact can influence each other in multiple ways, such as borrowing words, grammatical structures, phonology, or orthographic conventions Matras (2009). To illustrate our hypothesis, we select two datasets with South African (SA) languages in AfroNLU - NCHLT-ner and phrase-chunk. We select SA languages because they are contact languages (see Figure D.5 in Appendix for a genealogical classification tree that highlights the SA languages.) Nassenstein (2019); Arndt (2015); Roy-Campbell (2006).
To determine the significance of language similarity and language contact in our own zero-shot settings, we measure the Jaccard similarity between the pretraining data for the SA languages (see Table 8). We find strong similarities between some of these languages (see bolded examples in Table 8). We also finetune a BERT model and compare the performance of BERT with MBERT. We do this because BERT does not include any similar language in its representation.
XLM-R, mBERT, and AfriBERTa are not trained on most SA languages but have high scores in zero-shot settings see Table 9 and Table E.4 in Appendix. We argue that XLM-R in addition to cross-lingual transfers from other languages acquires representation from afr and xho where xho alone shares more than 0.4 similarity with afr, nbl, nso, and zul. mBERT also learns representation from afr while AfriBERTa learns representations from Gahuza which is a code-mixed variety of KIN and RUN. BERT on the other hand significantly performs lower than MBERT in all languages except on ssw, and ven (Phrase chunk). SERENGETI, however, outperforms other models on these languages which demonstrates the impact of pretraining on each of these languages.
These analyses are in no way conclusive, but do provide insights on how linguistic information may impact model performance in zero-shot settings. Future work can further probe the influence of similar languages in a more in-depth fashion. (See Appendix F for detailed analysis).
Conclusion
We reported our efforts to develop SERENGETI, a suite of three massively multilingual language models for African NLP. SERENGETI outperforms mPLMs on datasets across tasks. We provide extensive evaluations of model outputs, including zero-shot performance of the mPLMs. We also offer broad linguistically-motivated analyses of model performance.
Limitations
We identify the following limitations for our work:
Due to limited access to a wide network of native speakers from the majority of languages, we were able to manually inspect only a subset of languages present in our pretraining data. Specifically, we could only manually evaluate Afrikaans, Yorùbá, Igbo, Hausa, Luganda, Kinyarwanda, Chichewa, Shona, Somali, Swahili, Xhosa, Bemba, and Zulu. Future work should focus on increasing the subset of languages evaluated manually in order to ensure quality. We believe automatic analyses are not sufficient before development of models that get deployed in particular applications.
Another limitation is related to our inability to perform extensive analysis of biases and hateful speech present in our pretraining data. Again, this is due to relatively restricted access to native speakers (and even automated tools) to perform this analysis. As a result, we cannot fully ensure that our models is free from biases and socially undesirable effects. Therefore, it is important that these models be used with care and caution, and be analyzed for biases and socially undesirable effects before use.
Additionally, due to unavailability of sufficient computing resources, we were unable to evaluate large language models such as BLOOM, even though it covers African languages.
Finally, even though AfroNLU has diverse tasks at the word and sentence level, these tasks only cover few African languages. We therefore encourage the creation of more datasets for downstream NLU tasks in more (and more diverse) African languages. We believe broader benchmarks will continue to be important for future progress in African NLP.
Ethics Statement and Wider Impacts
SERENGETI aligns with Afrocentric NLP where the needs of African people is put into consideration when developing technology. We believe SERENGETI will not only be useful to speakers of the languages supported, but also researchers of African languages such as anthropologists and linguists. We discuss below some use cases for SERENGETI and offer a number of broad impacts.
SERENGETI aims to address the lack of access to technology in about of the world’s languages, which automatically discriminates against native speakers of those languages. More precisely, it does so by focusing on Africa. To the best of our knowledge, SERENGETI is the first massively multilingual PLM developed for African languages and language varieties. A model with knowledge of African languages, is by far the largest to date for African NLP.
SERENGETI enables improved access of important information to the African community in Indigenous African languages. This is especially beneficial for people who may not be fluent in other languages. This will potentially connect more people globally.
SERENGETI affords opportunities for language preservation for many African languages. To the best of our knowledge, SERENGETI consists of languages that have not been used for any NLP task until now. We believe that it can help encourage continued use of these languages in several domains, as well as trigger future development of language technologies for many of these languages.
To mitigate discrimination and bias, we adopt a manual curation of our datasets. Native speakers of Afrikaans, Yorùbá, Igbo, Hausa, Luganda, Kinyarwanda, Chichewa, Shona, Somali, Swahili, Xhosa, Bemba, and Zulu also manually evaluated a subset of the data to ensure its quality. The data collected for this work is taken from various domains to further ensure a better representation of the language usage of native speakers.
Although LMs are useful for a wide range of applications, they can also be misused. SERENGETI is developed using publicly available datasets that may carry biases. Although we strive to perform analyses and diagnostic case studies to probe performance of our models, our investigations are by no means comprehensive nor guarantee absence of bias in the data. In particular, we do not have access to native speakers of most of the languages covered. This hinders our ability to investigate samples from each (or at least the majority) of the languages.
Acknowledgements
MAM gratefully acknowledges support from Canada Research Chairs (CRC), the Natural Sciences and Engineering Research Council of Canada (NSERC; RGPIN-2018-04267), the Social Sciences and Humanities Research Council of Canada (SSHRC; 435-2018-0576; 895-2020-1004; 895-2021-1008), Canadian Foundation for Innovation (CFI; 37771), Digital Research Alliance of Canada,https://alliancecan.ca UBC ARC-Sockeye,https://arc.ubc.ca/ubc-arc-sockeye Advanced Micro Devices, Inc. (AMD), and Google. Any opinions, conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of CRC, NSERC, SSHRC, CFI, the Alliance, AMD, Google, or UBC ARC-Sockeye.
References
Appendix A Introduction
Appendix B Literature Review
Representation learning is an integral part of modern NLP systems. It has significantly improved the state of the art in natural language understanding (NLU) and natural language generation (NLG). We now discuss Afrocentric NLP, Multilingualism in NLP, Diversity and Inclusion in NLP, MLMs, and LMs for African languages.
More than Indigenous languages are spoken in Africa, which is about a third of all languages spoken in the world Eberhard et al. (2021). Unfortunately, the majority of these languages have not received any NLP attention to date. Rather, most NLP research has focused on higher resource languages. Most of these resourceful languages are typologically very different from Indigenous African languages. Methods used to develop technologies for these languages remain Western-centric, and may not be directly extensible to Indigenous African languages Adebara and Abdul-Mageed (2022). Existing NLP technologies also mostly function within the contexts of values and beliefs that reflect western societies and pose unique challenges if the technologies are applied within African communities.
Afrocentric NLP adopts a holistic approach to NLP throughout the life cycle of NLP policy making to model development and deployment. It discourages the current language data gas flaring policies that have led to the low resource status of many Indigenous African languages. Afrocentric NLP entails an understanding of the need for multi-dimensional policies that influence the language policy in education, media, government, and other domains to create ever-increasing, multi-domain, big data sources for NLP. During the archival and collection of language data, Afrocentric NLP necessitates respect of user consent, data sovereignty, wishes of local communities, and privacy Sutherland (2018); Daigle (2021); Makulilo (2012). For model development, approaches tailored to the unique typological features of African languages are of utmost priority. This also means development of models that understand simple to complex tones–a common feature in about of African languages–serial verb constructions, and many other features Hyman (2003); Creissels et al. (2008). Afrocentric NLP also prioritizes deploying models in formats that people without programming experience can easily use. Furthermore, from an Afrocentric approach, development of certain NLP applications such as language models, language identification tools, spelling checkers, language specific keyboards, and machine translation systems is crucial to advance NLP for African languages.
B.2 Multilingualism in NLP
Multilingualism, the ability to handle multiple languages within a single system or model, is becoming increasingly important as the amount of text and speech data in many languages increase. NLP systems capable of handling multiple languages can provide greater access to information and communication for people who speak languages other than those most commonly used or supported by NLP.
Multilingualism in NLP Ruder (2022) is mainly achieved through building (1) a single model trained on several languages Devlin et al. (2019); Conneau et al. (2020) and (2) transfer learning Raffel et al. (2020); He et al. (2022); Ruder et al. (2019). In the former, large transformer models have achieved state-of-the-art on many tasks while the latter has enabled the use of low-resource languages through finetuned on various NLP tasks. Due to lack of adequate (or good quality) pretraining data Kreutzer et al. (2021), transfer learning is often the most accessible method for a few low resource languages. Unfortunately, about of the world’s languages are either left-behinds, in that it is probably impossible to build NLP resources for them, or scraping-bys with no labelled datasets Joshi et al. (2020). For the left-behinds, labelled and unlabelled data is unavailable and even transfer learning approaches are beyond reach. So far, to the best of our knowledge, the largest multilingual model for African languages is pretrained on only African languages Dossou et al. (2022).
Most multilingual models are often trained with no more than languages because increasing the number of language would mean decreasing its capacity to learn representations of each language Conneau et al. (2020). Nevertheless, increasing model size was shown to ameliorate this problem Goyal et al. (2021). In some cases, these benchmarks are translations from English Artetxe et al. (2020); Nzeyimana and Niyongabo Rubungo (2022); Ponti et al. (2020) and may not necessarily be a good evaluation for the languages. This is because translating from a source language may mask concept gaps and differences in linguistic constituents Segerer (2008) in the target language. That is, translations are at best approximations of the target language Adebara and Abdul-Mageed (2022); Joshi et al. (2020). For example, when translating into English (which marks (in)definiteness morphologically) from Yorùbá (which uses bare nouns but marks these features contextually), ambiguities arise Adebara et al. (2022a).
For evaluation of multilingual models, several benchmarks have been createdArtetxe et al. (2020) with most of these supporting English and other high-resource languages. More recently, a few evaluation sets were introduced for African languages Ifeoluwa Adelani et al. (2021); Shode et al. (2022); Niyongabo et al. (2020).We include these evaluation sets in our benchmark, which we henceforth refer to as AfroNLU.
When evaluating multilingual models, reporting model performance for each language in the benchmark is preferred because reporting the results as a single value on all languages may mask the model’s performance on individual languages Ruder (2022). Large pre-training data, fine-tuning data, and evaluation benchmarks remain open challenging questions for achieving progress in multilingual NLP. For SERENGETI, we report results for each language in each benchmark across the tasks we evaluate on.
B.3 Diversity and Inclusion in NLP
Diversity relates to the level of variety within a system. It is the measure of distinctiveness between the various individuals within a group. Inclusion on the other hand relates to the level of representation or alignment of an individual within a group and the ability for that individual to function to its fullest ability Fosch-Villaronga and Poulsen (2022); Mitchell et al. (2020). Diversity and inclusion in NLP has gained increasing attention in recent years. In general, there is an acknowledgement that over-representation (and under-representation) of certain groups in the data used to train models Mitchell et al. (2020) can be amplified by resulting technologies. This raises concerns about the technology and how it is that it can further existing biases and societal inequalities. But these biases can be exhibited in various ways beyond training data, including the algorithms implemented, the diversity of researchers and engineers developing the models, and the societal and cultural context in which they are used.
Although this is starting to change, often times most of the data exploited in NLP models come from closely related Western languages. Most of these languages are Indo-European Aji et al. (2022); Joshi et al. (2020), and many of them share close geographic proximity and typology. In addition, the people who speak these languages have similar cultures. The implication is that several linguistic phenomena and typologies are under-represented in NLP data while those prevalent in Indo-European languages are over-represented Chakravarthi and Muralidaran (2021). About of the languages whose typology is described in WALS Dryer (2013) have not been used in NLP Joshi et al. (2020). Many ideas and topics, alien to Western cultures have also never been seen Adebara and Abdul-Mageed (2022); Bender (2011) in NLP data. African languages–and indeed many low resource languages–have rich linguistic typology, probably not seen in any other language in the world Bender (2011). An obvious problem with the current lack of diversity in NLP data is that the methods and models developed have overfit to these Indo-European typologies and cannot generalize to other typologies. Similarly, machine translation systems have been found to exhibit gender, racial Bolukbasi et al. (2016); Caliskan et al. (2017); Chakravarthi and Muralidaran (2021) and stylistic biases Hovy et al. (2020) in their outputs perpetuated through the data used for training.
A number of studies have also found that algorithms could exhibit biases Hooker (2021); Buolamwini and Gebru (2018); Dwork et al. (2011). For example, a recent study that investigated performance of Amazon Transcribe and Google Speech-To-Text on British English reported notably higher error rates for second language speakers of different varieties of British English Markl (2022). In another study, an evaluation of automatic speech recognition systems show substantial performance differences between ’standard’ US English and African American English (AAE) varieties Koenecke et al. (2020). In this study, commercial ASR systems developed by Amazon, Apple, Google, IBM, and Microsoft were evaluated and higher rates of errors were recorded for speakers of AAE than speakers of standard US varieties. Similar studies have also recorded higher errors in non-white users of English Wassink et al. (2022); Martin and Tang (2020). Other studies also reported differences in the performance of Youtube’s automatic caption in different settings. One study reported higher accuracy in the transcriptions of US English compared with Indian English Meyer et al. (2020). Another reported lower accuracy scores for women and speakers of Scottish English Tatman (2017) and non-white speakers of English Tatman and Kasten (2017).
Apart from data and algorithmic biases, the diversity crises in AI research is also argued to perpetuate historical biases Freire et al. (2021). A more inclusive and diverse workforce could promote the exploration of questions and solutions beyond currently investigated research questions Fosch-Villaronga and Poulsen (2022). Several initiatives have been adopted to increase diversity in AI, including providing travel grants to marginalized communities to attend conferences, creating mentoring opportunities, special workshops, and community diversity chairs. A number of organizations have also been developed to promote diversity and inclusion in AI and NLP, such as Masakhane, Black in AI, LatinX in AI.
The impact of using biased systems in decision making have been extensively studied. Algorithmic decision-making using biased systems have been shown to have significant discriminatory effects in health Obermeyer et al. (2019); Eubanks (2018), employment Barocas and Selbst (2016), housing Buolamwini and Gebru (2018); Barocas and Selbst (2016), government benefit allocation Eubanks (2018), policing Buolamwini and Gebru (2018); Barocas and Selbst (2016); Angwin et al. (2018), and freedom Angwin et al. (2018). Lack of diversity also has implication on access to technology. Currently, due to the use of a few high resource languages in NLP, there is limited global access to important applications such as machine translation, speech processing, information retrieval, and sentiment analysis. These technologies play an important role in ensuring a language thrives and offer major contributions to ongoing communication, literacy, education, and translation efforts in communities worldwide. These languages which have barely been used for NLP, usually referred to as low-resource languages, represent more than 90% of the world’s languages Joshi et al. (2020). The current focus of NLP on resource-rich languages does also have aggravating effects on the language endangerment problem which has been of serious concern for linguistics and language policy around the world. An alarming of languages have been envisaged to go extinct by the end of the century due to the domination by some of these resource-rich languages Besacier et al. (2014).
Overall, diversity and inclusion in NLP remain active areas of research and comprise pressing issues of international significance. SERENGETI contributes to diversity and inclusion in NLP as follows: (1) We develop SERENGETI, a suite of massively, multilingual language models that support African languages and language varieties. To the best of our knowledge, more than of these languages have never been represented in any language model to date. (2) The languages we support belong to language families. (3) We provide a massive benchmark covering languages across eight different tasks.
B.4 Multilingual Language Models
MLMs have proven effective for cross-lingual NLU and NLG, often outperforming monolingual language models Conneau et al. (2020). Different objectives have been adopted for training Doddapaneni et al. (2021), using Transformer architectures. These LMs use one of the three different variants of Transformer architectures–encoder-decoder, encoder-only and decoder-only Cai et al. (2022).
In the encoder-decoder models, input is encoded by the encoder side and the decoder conducts the operation to predict the sequence one token at a time or just reconstruct it by denoising. MBART Liu et al. (2020), AfriTeva Jude Ogundepo et al. (2022), M2M100 Fan et al. (2020), and MT5 Xue et al. (2021) are representatives for this architecture. Encoder-only models use only the encoder part of the transformer architecture, while decoder-only models use its decoder only. Some examples of encoder-only models are BERT Devlin et al. (2019), XLMR Conneau et al. (2020), and Electra Chi et al. (2021), while BLOOM Scao et al. (2022), GPT Radford et al. (2018, 2019); Brown et al. (2020b), OPT Zhang et al. (2022) are examples of decoder-only models. Most LMs developed for African languages use an encoder-only architecture, except AfriTEVA and AfroT5 which use encoder-decoder architectures.
These models are further finetuned on specific tasks. Finetuning has demonstrated its effectiveness on various NLU and NLG downstream tasks including part of speech tagging Conneau et al. (2020), named entity recognition Ushio and Camacho-Collados (2021); Conneau et al. (2020), and question answering Conneau et al. (2020). Finetuning follows a transfer learning approach which attempts to transfer knowledge from other sources to benefit a current task. This is based on the premise that previous knowledge may improve solutions for a current task Pan and Yang (2010); Raffel et al. (2020); He et al. (2022); Ruder et al. (2019). Transfer learning allows the domains, tasks, and distributions used in training and testing to be different thereby enabling a new task to leverage previously acquired domain knowledge. Potential benefits include faster learning, better generalization, and a more robust system. In the real world, we find many examples of transfer learning where humans transfer previous knowledge while learning or performing a task. For instance, knowing how to play the piano may facilitate learning to play the guitar and knowing how to ride a bicycle may facilitate learning to ride a motorbike. Finetuning is thus done by reusing the LM’s parameters as a starting point, while adding one task-specific layer trained from scratch. Finetuning can be done on an individual or joint basis Kitaev et al. (2019). In the former, a model is finetuned on single language for a specific downstream task. In the later, training data from a combination of multiple languages can be jointly finetuned in a single model.
Appendix C Pretraining Data
We provide details of our pretraining data below: Religious Domain. Our religious data is taken from online Bibles, Qurans, and data crawled from the Jehovah’s witness website. We also include religious texts from the book of Mormon.
News Domain. We collect data from online newspapers Adebara and Abdul-Mageed (2022) and news sites such as Voice of America, Voice of Nigeria, BBC, Global voices, and DW news sites. We collect local newspapers from languages from across Africa.
Government Documents. We collect government documents South African Centre for Digital Language Resources (SADiLaR), and the Universal Declaration of human rights (UDHR) in multiple languages.
Health Documents. We collect multiple health documents from the Department of Health, State Government of Victoria, Australia. We collect documents in Amharic, Dinka, Harari, Oromo, Somali, Swahili, and Tigrinya.
Existing Corpora. We collect corpora available on the web for different African languages, including from Project Gutenberg for Afrikaans, South African News data. for Sepedi and Setswana, OSCAR Abadji et al. (2021) for Afrikaans, Amharic, Somali, Swahili, Oromo, Malagasy, and Yoruba. We also used Tatoeba for Afrikaans, Amharic, Bemba, Igbo, Kanuri, Kongo, Luganda, Malagasy, Sepedi, Ndebele, Kinyarwanda, Somali, Swahili, Tsonga, Xhosa, Yoruba, and Zulu; Swahili Language Modelling Data for Swahili; Ijdutse corpus for Hausa; Data4Good corpora for Luganda, CC-100 for Amharic, Fulah, Igbo, Yoruba, Hausa, Tswana, Lingala, Luganada, Afrikaans, Somali, Swahili, Swati, North Sotho, Oromo, Wolof, Xhosa, and Zulu; Afriberta-Corpus for Afaan / Oromo, Amharic, Gahuza, Hausa, Igbo, Pidgin, Somali, Swahili, Tigrinya and Yoruba; mC4 for Afrikaans, Amharic, Hausa, Igbo, Malagasy, Chichewa, Shona, Somali, Sepedi, Swahili, Xhosa, Yoruba and Zulu.
Appendix D Typology Information for AfroNLU
SERENGETI consists of languages from families including: Afro-Asiatic, Austronesean, Creole-English, Creole-French, Creole-Kongo, Creole-Ngbandi, Creole-Portuguese, khoe-kwadi-Hainum, khoe-kwadi-Nama khoe-kwadi-Southwest, Indo-European, Niger-Congo, and Nilo Saharan. We discuss the classes from AfroNLU which includes Afro-Asiatic, Austronesian, Creole-English, Niger-Congo, and Nilo-Saharan.
Afro-Asiatic (aka Hamito-Semitic) is one of the language families of Africa. It consists of five or six branches: Berber, Chadic, Cushitic, Egyptian, Omotic (or a single Cush-Omotic), and SemiticPorkhomovsky (2020); Comrie (2017). Many Afro-Asiatic languages are spoken in Central, East, North, and West Africa. They are also spoken in the Middle East and in scattered communities in Europe, the United States, and the Caucasus Frajzyngier (2018). In Figure D.1, we show relationship between the Afro-asiatic languages in AfroNLU.
D.2 Austronesian
Austronesian languages are found along Mainland Southeast Asia, through Indonesia, Western New Guniea, and the Madagascar area in Africa Eberhard et al. (2021). Many of them have been shown to exhibit an isolating word structure. This means that the words in these languages are of minimal morphological complexity Gil and Schapper (2020). In Figure D.2, we show the geneology for Malagasy, the only Austronesian language in our benchmark.
D.3 Creole
A creole language is one spoken initially only in situations of contact between speakers of two or more mutually unintelligible languages, and not as a language within an ethnic group Sommer (2020). Historically, creoles have evolved along trade routes or in colonized communities particularly when several groups of people without a common lingua franca are forced to communicate in the presence of a dominant language. Creole languages therefore often include lexical items and grammatical features from multiple contact languages. Usually, one dominant language that is also referred to as the lexifier language contributes a majority of the vocabulary. Creole languages are classified based on their geographical location and are further grouped according to their main lexifier languages, their presumed origins, and the major languages with which they are in contact (i.e., contact languages). Figure D.3 shows the geneology for Nigerian Pidgin, the only Creole in our pretraining collection.
D.4 Indo-European
Afrikaans is the only "Indigenous" Indo-European language spoken in Africa. Although it may also be viewed as not being truly Indigenous to Africa Kirsten (2018). Indo-European languages were originally domiciled in Europe, Iran, Turkey, Western Asia and India Clackson (2007); Eberhard et al. (2021); Comrie (2017); Kirsten (2018). However, due to migration, Indo-European languages are spoken around the world. In 2003, over 2.5 billion people spoke an Indo-European language Clackson (2007). In Figure D.4, we show the geneology for Afrikaans.
D.5 Niger-Congo
Niger-Congo, also referred to as Niger-Kordofanian, is the largest language family in Africa Good (2020); Comrie (2017). It consists of the highest number of languages and speakers in Africa. Niger-Congo languages spread across sub-Saharan Africa, with Benue-Congo, including Bantu languages dominating the southern part of the continent. Figure D.5 shows the Niger-congo languages in our collection. Although we use similar colours for languages which are sisters of the same parent, only some of those languages are mutually intelligible. That is speakers of each individual language understand each other’s language without learning it. Specifically, Kinyawanda (kin) and Kirundi (run) are mutually intelligible Nassenstein (2019). Ndebele, Siswati, Xhosa, and Zulu also share various levels of intelligibility mutually intelligible Arndt (2015); Roy-Campbell (2006). Sepedi, Sotho, and Tswana also share some levels of mutual intelligibility Roy-Campbell (2006).
D.6 Nilo-Saharan
Nilo-Saharan is subdivided into four branches that include North Eastern, Central Sudanic and two disputed branches–Songhay and Koman Dimmendaal et al. (2019); Dimmendaal (2020); Comrie (2017). These branches are further divided into other subgroups, languages, and dialects. Nilo-Saharan languages are spoken predominantly by eastern and central African pastoralists, and includes in its main Chari-Nile branch the Central Sudanic and Eastern Sudanic (also called Nilotic) languages. Figure D.6 shows the Nilo-saharan languages in our pretraining data.
Appendix E Evaluation
In this section, we provide more information about our evaluation procedure and results using visualizations and tables. Figure E.1 shows the confusion matrix for the news classification cluster. Figure E.2 shows the performance of SERENGETI on the sentiment analysis cluster. Each confusion matrix represents each dataset in the sentiment analysis cluster. In Figure E.3, we show SERENGETI performance on each category in the topic classification datasets.
E.2 Error Analysis
In the sentiment analysis cluster, best performance is recorded for positive categories while negative categories have the worst performance. A fine-grained analysis of the Yoruba sentiment dataset found that SERENGETI failed to correctly categorize sentiment if the polarity item(s) were not seen in training, can be associated with both positive and negative sentiments, the polarity item(s) is a negation, or if ambivalent markers are present in the sentence. We provide a table showing examples of each type of error we found in Table E.2 in the Appendix. For the news classification task, politics and tourism are the best performing classes while education and relationships have the worst performance on kirnews and kinnews respectively. It is important to mention that the worst performing categories do not have the smallest data sizes. For the topic classification, the best performance is on the world class for Hausa topic modelling while entertainment and sport have best performance for Yoruba. The worst performance is on Nigeria and health for Hausa and Yoruba topic datasets respectively.
E.3 Imbalanced Distribution
We find imbalances in the class distributions for all datasets except YOSM. We find a positive correlation between the size of each category in a dataset and the model accuracy. The larger the number of examples in a specific class, the better the accuracy, although we find a few exceptions. We provide confusion matrices that represents the sizes of each category and the performance of SERENGETI in Figures E.4, E.5, and E.6.
Appendix F Detailed Geneaology and Language Contact Analysis
In this Section, we use Figures and Tables to provide evidence for the influence of similar languages in zero-shot settings. First, we highlight in purple the similar languages that we perform genealogy analysis on in Figure E.7. In the figure, the languages with mutual intelligibility are presented in similar coloured circles. To determine the significance of language similarity and language contact in our own zero-shot settings, we measure the Jaccard similarity between the pretraining data for the South African languages in AfroNLU (see Table 8). To calculate the Jaccard similarities, we removed digits, emojis, and punctuation marks. We do this to ensure that we reduce interference with the similarity scores. We find strong similarities between some of these languages as in the bolded examples in Table 8.
We find that although XLM-R, mBERT, and AfriBERTa are not trained on most most of these languages, we record high scores in zero-shot settings see Table E.4). We argue that XLM-R in addition to cross-lingual transfers from other languages acquires representation from afr and xho where xho alone shares more than 0.4 similarity with afr, nbl, nso, and zul. mBERT also learns representation from afr while AfriBERTa learns representations from Gahuza which is a code-mixed variety of kin and run. SERENGETI however, outperforms other models on these datasets indicating that learning the representation of each language improves performance.
Next, we finetune a BERT model and compare the performance of BERT with MBERT. We do this because BERT is a monolingual model and does not include any similar language in its representation. In Table 9, BERT significantly performs lower than MBERT in all languages in NCHLT-NER. BERT also has lower performance on the phrase-chunk dataset in all languages except on ssw, and ven.
This analysis is far from being conclusive and future work can further probe the influence of similar languages in more detail. This is necessary to evaluate to what extent similar languages have an influence on performance in zero-shot settings and why in zero shot settings, some monolingual models outperform multilingual ones. For example, in the case of ssw and ven.