Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-Tuning
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, Dietrich Klakow
Introduction
Recent advances in the development of multilingual pre-trained language models (PLMs) like mBERT (Devlin et al., 2019), XLM-R (Conneau et al., 2020), and RemBERT (Chung et al., 2021) have led to significant performance gains on a wide range of cross-lingual transfer tasks. Due to the curse of multilinguality (Conneau et al., 2020) — a trade-off between language coverage and model capacity — and non-availability of pre-training corpora for many low-resource languages, multilingual PLMs are often trained on about 100 languages. Despite the limitations of language coverage, multilingual PLMs have been shown to transfer to several low-resource languages unseen during pre-training. Although, there is still a large performance gap compared to languages seen during pre-training.
One of the most effective approaches to adapt to a new language is language adaptive fine-tuning (LAFT) — fine-tuning a multilingual PLM on monolingual texts in the target language using the same pre-training objective. This has been shown to lead to big gains on many cross-lingual transfer tasks (Pfeiffer et al., 2020), and low-resource languages (Muller et al., 2021; Chau & Smith, 2021), including African languages (Alabi et al., 2020; Adelani et al., 2021). Nevertheless, adapting a model to each target language individually takes large disk space, and limits the cross-lingual transfer abilities of the resulting models because they have been specialized to individual languages (Beukman, 2021).
An orthogonal approach to improve the coverage of low-resource languages is to include them in the pre-training data. An example for this approach is AfriBERTa (Ogueji et al., 2021), which was trained from scratch on 11 African languages. A downside of this approach is that it is resource intensive in terms of data and compute.
Another alternative approach is parameter efficient fine-tuning like Adapters (Pfeiffer et al., 2020) and sparse fine-tuning (Ansell et al., 2021), where the model is adapted to new languages by using a sparse network trained on a small monolingual corpus. Similar to LAFT, it requires adaptation for every new target language. Although it takes little disk space, all target language-specific parameters need to be stored.
In this paper, we propose multilingual adaptive fine-tuning (MAFT), a language adaptation to multiple languages at once. We perform language adaptation on the 17 most-resourced African languages (Afrikaans, Amharic, Hausa, Igbo, Malagasy, Chichewa, Oromo, Naija, Kinyarwanda, Kirundi, Shona, Somali, Sesotho, Swahili, isiXhosa, Yorùbá, isiZulu) and three other high-resource language widely spoken on the continent (English, French, and Arabic) simultaneously to provide a single model for cross-lingual transfer learning for African languages. To further specialize the multilingual PLM, we follow the approach of Abdaoui et al. (2020) and remove vocabulary tokens from the embedding layer that correspond to non-Latin and non-Ge’ez (used by Amharic) scripts before MAFT, thus effectively reducing the model size by 50%.
Our evaluation on two multilingual PLMs (AfriBERTa and XLM-R) and three NLP tasks (NER, news topic classification and sentiment classification) shows that our approach is competitive to performing LAFT on the individual languages, with the benefit of having a single model instead of a separate model for each of the target languages. Also, we show that our adapted PLM improves the zero-shot cross-lingual transfer abilities of parameter efficient fine-tuning methods like Adapters Pfeiffer et al. (2020) and sparse fine-tuning (Ansell et al., 2021).
As an additional contribution, and in order to cover more diverse African languages in our evaluation, we create a new evaluation corpus, ANTC – African News Topic Classification – for Lingala, Somali, Naija, Malagasy, and isiZulu from pre-defined news categories of VOA, BBC, Global Voices, and Isolezwe newspapers. To further the research on NLP for African languages, we make our code and data publicly available.https://github.com/uds-lsv/afro-maft Additionally, our models are available via HuggingFace.https://huggingface.co/Davlan
Related Work
The success of multilingual PLMs such as mBERT (Devlin et al., 2019) and XLM-R (Conneau et al., 2020) for cross-lingual transfer in many natural language understanding tasks has encouraged the continuous development of multilingual models (Luo et al., 2021; Chi et al., 2021; Ouyang et al., 2021; Chung et al., 2021; He et al., 2021). Most of these models cover 50 to 110 languages and only few African languages are represented due to lack of large monolingual corpora on the web. To address this under-representation, regional multilingual PLMs have been trained from scratch such as AfriBERTa (Ogueji et al., 2021) or adapted from existing multilingual PLM through LAFT (Alabi et al., 2020; Pfeiffer et al., 2020; Muller et al., 2021; Adelani et al., 2021). AfriBERTa is a relatively small multilingual PLM (126M parameters) trained using the RoBERTa architecture and pre-training objective on 11 African languages. However, it lacks coverage of languages from the southern region of the African continent, specifically the southern-Bantu languages. In our work, we extend to those languages since only a few of them have large (100MB size) monolingual corpus. We also do not specialize to a single language but apply MAFT which allows multilingual adaptation and preserves downstream performance on both high-resource and low-resource languages.
It is not unusual for a new multilingual PLM to be initialized from an existing model. For example, Chi et al. (2021) trained InfoXLM by initializing the weights from XLM-R before training the model on a joint monolingual and translation corpus. Although they make use of a new training objective during adaptation. Similarly, Tang et al. (2020) extended the languages covered by mBART (Liu et al., 2020b) from 25 to 50 by first modifying the vocabulary and initializing the model weights of the original mBART before fine-tuning it on a combination of monolingual texts from the original 25 languages in addition to 25 new languages. Despite increasing the number of languages covered by their model, they did not observe a significant performance drop on downstream tasks. We take inspiration from these works for applying MAFT on African languages, but we do not modify the training objective during adaptation nor increase the vocabulary.
One of the most effective methods for creating smaller PLMs is distillation where a small student model is trained to reproduce the behaviour of a larger teacher model. This has been applied to many English PLMs (Sanh et al., 2019; Jiao et al., 2020; Sun et al., 2020; Liu et al., 2020a) and a few multilingual PLMs (Wang et al., 2020, 2021). However, it often leads to a drop in performance compared to the teacher PLM. An alternative approach that does not lead to a drop in performance has been proposed by Abdaoui et al. (2020) for multilingual PLM. They removed unused vocabulary tokens from the embedding layer. This simple method significantly reduces the number of embedding parameters thus reducing the overall model size since the embedding layer contributes the most to the total number of model parameters. In our paper, we combine MAFT with the method proposed by Abdaoui et al. (2020) to reduce the overall size of the resulting multilingual PLM for African languages. This is crucial especially because people from under-represented communities in Africa may not have access to powerful GPUs in order to fine-tune large PLMs. Also, Google Colabhttps://colab.research.google.com/ (free-version), which is widely used by individuals from under-represented communities without access to other compute resources, cannot run large models like e.g. XLM-R. Hence, it is important to provide smaller models that still achieve competitive downstream performance to these communities.
One of the challenges of developing (multilingual) PLMs for African languages is the lack of evaluation corpora. There have been many efforts by communities like Masakhane to address this issue ( et al., 2020; Adelani et al., 2021). We only find two major evaluation benchmark datasets that cover a wide range of African languages: one for named entity recognition (NER) (Adelani et al., 2021) and one for sentiment classification (Muhammad et al., 2022). In addition, there are also several news topic classification datasets (Hedderich et al., 2020; Niyongabo et al., 2020; Azime & Mohammed, 2021) but they are only available for a few African languages. Our work contributes novel news topic classification datasets (i.e. ANTC) for additional five African languages: Lingala, Naija, Somali, isiZulu, and Malagasy.
Data
We perform MAFT on 17 African languages Afrikaans, Amharic, Hausa, Igbo, Malagasy, Chichewa, Oromo, Naija, Kinyarwanda, Kirundi, Shona, Somali, Sesotho, Swahili, isiXhosa, Yorùbá, isiZulu) covering the major African language families and 3 high resource languages (Arabic, French, and English) widely spoken in Africa. We selected the African languages based on the availability of a (relatively) large amount of monolingual texts. We obtain the monolingual texts from three major sources: the mT5 pre-training corpus which is based on Common Crawl Corpushttps://commoncrawl.org/ (Xue et al., 2021), the British Broadcasting Corporation (BBC) News, Voice of America Newshttps://www.voanews.com (Palen-Michel et al., 2022), and some other news websites based in Africa. Table 9 in the Appendix provides a summary of the monolingual data, including their sizes and sources. We pre-processed the data by removing lines that consist of numbers or punctuation only, and lines with less than six tokens.
2 Evaluation tasks
We run our experiments on two sentence level classification tasks: news topic classification and sentiment classification, and one token level classification task: NER. We evaluate our models on English as well as diverse African languages with different linguistic characteristics.
For the NER task we evaluate on the MasakhaNER dataset (Adelani et al., 2021), a manually annotated dataset covering 10 African languages (Amharic, Hausa, Igbo, Kinyarwanda, Luganda, Luo, Naija, Kiswahili, Wolof, and Yorùbá) with texts from the news domain. For English, we use data from the CoNLL 2003 NER task (Tjong Kim Sang & De Meulder, 2003) also containing texts from the news domain. For isiXhosa, we use the data from Eiselen (2016). Lastly, to evaluate on Arabic we make use of the ANERCorp dataset (Benajiba et al., 2007; Obeid et al., 2020).
We use existing news topic datasets for Amharic (Azime & Mohammed, 2021), English – AG News corpus – (Zhang et al., 2015), Kinyarwanda – KINNEWS – (Niyongabo et al., 2020), Kiswahili – new classification dataset– (David, 2020), and both Yorùbá and Hausa (Hedderich et al., 2020). For dataset without a development set, we randomly sample 5% of their training instances and use them as a development set.
We use the NaijaSenti multilingual Twitter sentiment analysis corpus (Muhammad et al., 2022). This is a large code-mixed and monolingual sentiment analysis dataset, manually annotated for 4 Nigerian languages: Hausa, Igbo, Yorùbá and Pidgin. Additionally, we evaluate on the Amharic, and English Twitter sentiment datasets by Yimam et al. (2020) and Rosenthal et al. (2017), respectively. For all datasets above, we only make use of tweets with positive, negative and neutral sentiments.
2.2 Newly created dataset: ANTC corpus
We created a novel dataset, ANTC — African News Topic Classification for five African languages. We obtained data from three different news sources: VOA, BBChttps://www.bbc.com/pidgin, Global Voiceshttps://mg.globalvoices.org/, and isolezwehttps://www.isolezwe.co.za. From the VOA data we created datasets for Lingala and Somali. We obtained the topics from data released by Palen-Michel et al. (2022) and used the provided URLs to get the news category from the websites. For Naija, Malagasy and isiZulu, we scrapped news topic from the respective news website (BBC Pidgin, Global Voices, and isolezwe respectively) directly base on their category. We noticed that some news topics are not mutually exclusive to their categories, therefore, we filtered such topics with multiple labels. Also, we ensured that each category has at least 200 samples. The categories include but are not limited to: Africa, Entertainment, Health, and Politics. The pre-processed datasets were divided into training, development, and test sets using stratified sampling with a ratio of 70:10:20. Table 1 provides details about the dataset size and news topic information.
Pre-trained Language Models
For our experiments, we make use of different multilingual PLMs that have been trained using a masked language model objective on large collections of monolingual texts from several languages. Table 2 shows the number of parameters as well as the African languages covered by each of the models we consider.
XLM-R (Conneau et al., 2020) has been pre-trained on 100 languages including eight African languages. We make use of both XLM-R-base and XLM-R-large for MAFT with 270M and 550M parameter sizes respectively. Although, for our main experiments, we make use of XLM-R-base.
AfriBERTa (Ogueji et al., 2021) has been pre-trained only on African languages. Despite its smaller parameter size (126M), it has been shown to reach competitive performance to XLM-R-base on African language datasets (Adelani et al., 2021; Hedderich et al., 2020).
XLM-R-miniLM (Wang et al., 2020) is a distilled version of XLM-R-large with only 117M parameters.
We fine-tune the baseline models for NER, news topic classification and sentiment classification for 50, 25, and 20 epochs respectively. We use a learning rate of 5e-5 for all the task, except for sentiment classification where we use 2e-5 for XLM-R-base and XLM-R-large. The maximum sequence length is 164 for NER, 500 for news topic classification, and 128 for sentiment classification. The adapted models also make use of similar hyper-parameters.
Multilingual Adaptive Fine-tuning
We introduce MAFT as an approach to adapt a multi-lingual PLM to a new set of languages. Adapting PLMs has been shown to be effective when adapting to a new domain (Gururangan et al., 2020) or language (Pfeiffer et al., 2020; Alabi et al., 2020; Muller et al., 2021; Adelani et al., 2021). While previous work on multilingual adaptation has mostly focused on autoregressive sequence-to-sequence models such as mBART (Tang et al., 2020), in this work, we adapt non-autoregressive masked PLMs on monolingual corpora covering 20 languages. Crucially, during adaptation we use the same objective that was also used during pre-training. The models resulting from MAFT can then be fine-tuned on supervised NLP downstream tasks. We name the model resulting after applying MAFT to XLM-R-base and XLM-R-miniLM as AfroXLMR-base and AfroXLMR-mini, respectively. For adaptation, we train on a combination of the monolingual corpora used for AfriMT5 adaptation by Adelani et al. (2022). Details for each of the monolingual corpora and languages are provided in Appendix A.1.
The PLMs were trained for 3 epochs with a learning rate of 5e-5 using huggingface transformers (Wolf et al., 2020). We use of a batch size of 32 for AfriBERTa and a batch size 10 for the other PLMs.
1 Vocabulary reduction
Multilingual PLMs come with various parameter sizes, the larger ones having more than hundred million parameters, which makes fine-tuning and deploying such models a challenge due to resource constraints. One of the major factors that contributes to the parameter size of these models is the embedding matrix whose size is a function of the vocabulary size of the model. While a large vocabulary size is essential for a multilingual PLM trained on hundreds of languages, some of the tokens in the vocabulary can be removed when they are irrelevant to the domain or language considered in the downstream task, thus reducing the vocabulary size of the model. Inspired by Abdaoui et al. (2020), we experiment with reducing the vocabulary size of the XLM-R-base model before adapting via MAFT. There are two possible vocabulary reductions in our setting: (1) removal of tokens before MAFT or (2) removal of tokens after MAFT. From our preliminary experiments, we find approach (1) to work better. We call the resulting model, AfroXLMR-small.
To remove non-African vocabulary sub-tokens from the pretrained XLM-base model, we concatenated the monolingual texts from 19 out of the 20 African languages together. Then, we apply sentencepiece to the Amharic monolingual texts, and concatenated texts separately using the original XLM-R-base tokenizer. The frequency of all the sub-tokens in the two separate monolingual corpora is computed, and we select the top-k most frequent tokens from the separate corpora. We used this separate sampling to ensure that a considerable number of Amharic sub-tokens are captured in the new vocabulary, we justify the choice of this approach in Section 5.3. We assume that the top-k most frequent tokens should be representative of the vocabulary of the whole 20 languages. We chose from the Amharic sub-tokens which covers 99.8% of the Amharic monolingual texts, and which covers 99.6% of the other 19 languages, and merged them. In addition, we include the top 1000 tokens from the original XLM-R-base tokenizer in the new vocabulary to include frequent tokens that were not present in the new top-k tokens.This introduced just a few new tokens which are mostly English tokens to the new vocabulary. We end up with distinct sub-tokens after combining all of them. We note that our assumption above may not hold in the case of some very distant and low-resourced languages as well as when there are domain differences between the corpora used during adaptation and fine-tuning. We leave the investigation of alternative approaches for vocabulary compression for future work.
2 Results and discussion
For the baseline models (top rows in Tables 3, 4, and 5), we directly fine-tune on each of the downstream tasks in the target language: NER, news topic classification and sentiment analysis.
For NER and sentiment analysis we find XLM-R-large to give the best overall performance. We attribute this to the fact that it has a larger model capacity compared to the other PLMs. Similarly, we find AfriBERTa and XLM-R-base to give better results on languages they have been pre-trained on (see Table 2), and in most cases AfriBERTa tends to perform better than XLM-R-base on languages they are both pre-trained on, for example amh, hau, and swa. However, when the languages are unseen by AfriBERTa (e.g. ara, eng, wol, lin, lug, luo, xho, zul), it performs much worse than XLM-R-base and in some cases even worse than the XLM-R-miniLM. This shows that it may be better to adapt to a new African language from a PLM that has seen numerous languages than one trained on a subset of African languages from scratch.
The results of applying LAFT to the XLM-R-base model are shown in the last row of Tables 3, 4, and 5. We find that applying LAFT on each language individually provides a significant improvement in performance across all languages and tasks we evaluated on. Sometimes, the improvement is very large, for example, F1 on Amharic NER and F1 for Zulu news-topic classification. The only exception is for English since XLM-R has already seen large amounts of English text during pre-training. Additionally, LAFT models tend to give slightly worse result when adaptation is performed on a smaller corpus.We performed LAFT on eng using VOA news corpus with about 906.6MB, much smaller than the CC-100 eng corpus (300GB)
2.2 Multilingual adaptive fine-tuning results
While LAFT provides an upper bound on downstream performance for most languages, our new approach is often competitive to LAFT. On average, the difference on NER, news topic and sentiment classification is , , and F1, respectively. Crucially, compared to LAFT, MAFT results in a single adapted model which can be applied to many languages while LAFT results in a new model for each language. Below, we discuss our results in more detail.
We found all the PLMs to improve after we applied MAFT. The improvement is the largest for the XLM-R-miniLM, where the performance improved by F1 for NER, and F1 for news topic classification. Although, the improvement was lower for sentiment classification (). Applying MAFT on XLM-R-base gave the overall best result. On average, there is an improvement of , , and F1 on NER, news topic and sentiment classification, respectively. The main advantage of MAFT is that it allows us to use the same model for many African languagesinstead of many models specialized to individual languages. This significantly reduces the required disk space to store the models, without sacrificing performance. Interestingly, there is no strong benefit of applying MAFT to AfriBERTa. In most cases the improvement is F1. We speculate that this is probably due to AfriBERTa’s tokenizer having a limited coverage. We leave a more detailed investigation of this for future work.
Applying vocabulary reduction helps to reduce the model size by more than before applying MAFT. We find a slight reduction in performance as we remove more vocabulary tokens. Average performance of XLM-R-base-v70k reduces by , and F1 for NER, news topic, and sentiment classification compared to the XLM-R-base+LAFT baseline. Despite the reduction in performance compared to XLM-R-base+LAFT, they are still better than XLM-R-miniLM, which has a similar model size, with or without MAFT. We also find that their performance is better than that of the PLMs that have not undergone any adaptation. We find the largest reduction in performance on languages that make use of non-Latin scripts i.e. amh and ara — they make use of the Ge’ez script and Arabic script respectively. We attribute this to the vocabulary reduction impacting the number of amh and ara subwords covered by our tokenizer.
In summary, we recommend XLM-R-base+MAFT (i.e. AfroXLMR-base) for all languages on which we evaluated, including high-resource languages like English, French and Arabic. If there are GPU resource constraints, we recommend using XLM-R-base-v70k+MAFT (i.e. AfroXLMR-small).
3 Ablation experiments on vocabulary reduction
Our results showed that applying vocabulary reduction reduced the model size, but we also observed a drop in performance for different languages across the downstream tasks, especially for Amharic, because it uses a non-Latin script. Hence, we compared different sampling strategies for selecting the top-k vocabulary sub-tokens. These include: (i) concatenating the monolingual texts, and selecting the top-70k sub-tokens (ii) the exact approach described in Section 5.1. The resulting tokenizers from the two approaches are used to tokenize the sentences in the NER test sets for Amharic, Arabic, English, and Yorùbá. Table 6 shows the number of UNKs in the respective test set after tokenization and the F1 scores obtained on the NER task for the languages. The table shows that the original AfroXLMR tokenizer obtained the least number of UNKs for all languages, with the highest F1 scores. Note that Yorùbá has UNKs, which is explained by the fact that Yorùbá was not seen during pre-training. Furthermore, using approach (i), gave UNKs for Amharic, but with approach (ii) there was a significant drop in the number of UNKs and an improvement in F1 score. We noticed a drop in the vocabulary coverage for the other languages as we increased the Amharic sub-tokens. Therefore, we concluded that there is no sweet spot in terms of the way to pick the vocabulary that covers all languages and we believe that this is an exciting area for future work.
4 Scaling MAFT to larger models
To demonstrate the applicability of MAFT to larger models, we applied MAFT to XLM-R-large using the same training setup as XLM-R-base. We refer to the new PLM as AfroXLMR-large. For comparison, we also trained individual LAFT models using the monolingual dataFor languages not in MasakhaNER, we use the same monolingual data in Table 9. from Adelani et al. (2021). Table 7 shows the evaluation result on NER. Averaging over all 13 languages, AfroXLMR-large improved over XLM-R-large by F1, which is very comparable to the improvement we obtained between AfroXLMR-base ( F1) and XLM-R-base ( F1). Surprisingly, the improvement is quite large ( to F1) for seven out of ten African languages: yor, luo, lug, kin, ibo, and amh. The most interesting observation is that AfroXLMR-large, on average, is either competitive or better than the individual language LAFT models, including languages not seen during the MAFT training stage like lug, luo and wol. This implies that AfroXLMR-large (a single model) provides a better alternative to XLM-R-large+LAFT (for each language) in terms of performance on downstream tasks and disk space. AfroXLMR-large is currently the largest masked language model for African languages, and achieves the state-of-the-art compared to all other multilingual PLM on the NER task. This shows that our MAFT approach is very effective and scales to larger PLMs.
Cross-lingual Transfer Learning
The previous section demonstrates the applicability of MAFT in the fully-supervised transfer learning setting. Here, we demonstrate that our MAFT approach is also very effective in the zero-shot cross-lingual transfer setting using parameter-efficient fine-tuning methods.
Parameter-efficient fine-tuning methods like adapters Houlsby et al. (2019) are appealing because of their modularity, portability, and composability across languages and tasks. Often times, language adapters are trained on a general domain corpus like Wikipedia. However, when there is a mismatch between the target domain of the task and the domain of the language adapter, it could also impact the cross-lingual performance.
Here, we investigate how we can improve the cross-lingual transfer abilities of our adapted PLM – AfroXLMR-base by training language adapters on the same domain as the target task. For our experiments, we use the MasakhaNER dataset, which is based on the news domain. We compare the performance of language adapters trained on Wikipedia and news domains. In addition to adapters, we experiment with another parameter-efficient method based on Lottery-Ticket Hypothesis Frankle & Carbin (2019) i.e. LT-SFT Ansell et al. (2021).
For the adapter approach, we make use of the MAD-X approach Pfeiffer et al. (2020) – an adapter-based framework that enables cross-lingual transfer to arbitrary languages by learning modular language and task representations. However, the evaluation data in the target languages should have the same task and label configuration as the source language. Specifically, we make use of MAD-X 2.0 Pfeiffer et al. (2021) where the last adapter layers are dropped, which has been shown to improve performance. The setup is as follows: (1) We train language adapters via masked language modelling (MLM) individually on source and target languages, the corpora used are described in Appendix A.2; (2) We train a task adapter by fine-tuning on the target task using labelled data in a source language. (3) During inference, task and language adapters are stacked together by substituting the source language adapter with a target language adapter.
We also make use of the Lottery Ticket Sparse Fine-tuning (LT-SFT) approach (Ansell et al., 2021), a parameter-efficient fine-tuning approach that has been shown to give competitive or better performance than the MAD-X 2.0 approach. The LT-SFT approach is based on the Lottery Ticket Hypothesis (LTH) that states that each neural model contains a sub-network (a “winning ticket”) that, if trained again in isolation, can reach or even surpass the performance of the original model. The LTH is originally a compression approach, the authors of LT-SFT re-purposed the approach for cross-lingual adaptation by finding sparse sub-networks for tasks and languages, that will later be composed together for zero-shot adaptation, similar to Adapters. For additonal details we refer to Ansell et al. (2021).
For our experiments, we followed the same setting as Ansell et al. (2021) that adapted mBERT from English CoNLL03 Tjong Kim Sang & De Meulder (2003) to African languages (using MasakhaNER dataset) for the NER task.We excluded the MISC and DATE from CoNLL03 and MasakhaNER respectively to ensure same label configuration. Furthermore, we extend the experiments to XLMR-base and AfroXLMR-base. For the training of MAD-X 2.0 and sparse fine-tunings (SFT) for African languages, we make use of the monolingual texts from the news domain since it matches the domain of the evaluation data. Unlike, Ansell et al. (2021) that trained adapters and SFT on monolingual data from Wikipedia domain except for luo and pcm where the dataset is absent, we show that the domain used for training language SFT is also very important. For a fair comparison, we reproduced the result of Ansell et al. (2021) by training MAD-X 2.0 and LT-SFT on mBERT, XLM-R-base and AfroXLMR-base on target languages with the news domain corpus. But, we still make use of the pre-trained English language adapterhttps://adapterhub.ml/ and SFThttps://huggingface.co/cambridgeltl for mBERT and XLM-R-base trained on the Wikipedia domain. For the AfroXLMR-base, we make use of the same English adapter and SFT as XLM-R-base because the PLM is already good for English language. We make use of the same hyper-parameters reported in the LT-SFT paper.
We train the task adapter using the following hyper-parameters: batch size of 8, 10 epochs, “pfeiffer” adapter config, adapter reduction factor of 8, and learning rate of 5e-5. For the language adapters, we make use of 100 epochs or maximum steps of 100K, minimum number of steps is 30K, batch size of 8, “pfeiffer+inv” adapter config, adapter reduction factor of 2, learning rate of 5e-5, and maximum sequence length of 256. For a fair comparison with adapter models trained on Wikipedia domain, we used the same hyper-parameter settings Ansell et al. (2021) for the news domain.
2 Results and discussion
Table 8 shows the results of MAD-X 2.0 and LT-SFT, we compare their performance to fully supervised setting, where we fine-tune XLM-R-base on the training dataset of each of the languages, and evaluate on the test-set. We find that both MAD-X 2.0 and LT-SFT using news domain for African languages produce better performance ( on MAD-X and on LT-SFT) than the ones trained largely on the wikipedia domain. This shows that the domain of the data matters. Also, we find that training LT-SFT on XLM-R-base gives better performance than mBERT on all languages. For MAD-X, there are a few exceptions like hau, pcm, and yor. Overall, the best performance is obtained by training LT-SFT on AfroXLMR-base, and sometimes it give better performance than the fully-supervised setting (e.g. as observed in kin and lug, wol yor languages). On both MAD-X and LT-SFT, AfroXLMR-base gives the best result since it has been firstly adapted on several African languages and secondly on the target domain of the target task. This shows that the MAFT approach is effective since the technique provides a better PLM that parameter-efficient methods can benefit from.
Conclusion
In this work, we proposed and studied MAFT as an approach to adapt multilingual PLMs to many African languages with a single model. We evaluated our approach on 3 different NLP downstream tasks and additionally contribute novel news topic classification dataset for 4 African languages. Our results show that MAFT is competitive to LAFT while providing a single model compared to many models specialized for individual languages. We went further to show that combining vocabulary reduction and MAFT leads to a 50% reduction in the parameter size of a XLM-R while still being competitive to applying LAFT on individual languages. We hope that future work improves vocabulary reduction to provide even smaller models with strong performance on distant and low-resource languages. To further research on NLP for African languages and reproducibility, we have uploaded our language adapters, language SFTs, AfroXLMR-base, AfroXLMR-small, and AfroXLMR-mini models to the HuggingFace Model Hubhttps://huggingface.co/models?sort=downloads&search=Davlan%2Fafro-xlmr.
Acknowledgments
Jesujoba Alabi was partially funded by the BMBF project SLIK under the Federal Ministry of Education and Research grant 01IS22015C. David Adelani acknowledges the EU-funded Horizon 2020 projects: ROXANNE under grant number 833635 and COMPRISE (http://www.compriseh2020.eu/) under grant agreement No. 3081705. Marius Mosbach acknowledges funding from the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) – Project-ID 232722074 – SFB 1102. We also thank DFKI GmbH for providing the infrastructure to run some of the experiments. We are grateful to CoreWeave and EleutherAI for providing the compute to train AfroXLMR-large. We thank Alan Ansell for providing his MAD-X 2.0 code. Lastly, we thank Benjamin Muller, the anonymous reviewers of AfricaNLP 2022 workshop and COLING 2022 for their helpful feedback.
References
Appendix A Appendix
For training the MAFT models, we make use of the aggregation of monolingual data from Table 9.
For the LAFT models, we make use of existing XLMR-base+LAFT models from the MasakhaNER paper (Adelani et al., 2021). However, for other languages not present in MasakhaNER (ara, mlg,orm, sna, som, xho), we make use of the mC4 corpus except for eng — we use the VOA corpus. For a fair comparison across models, when training the XLM-R-large+LAFT models, we use the same monolingual corpus used to train XLM-R-base+LAFT models.
A.2 News corpora for language adapters and SFTs
Table 10 provides the news corpus we used to train language adapters and SFTs for the cross-lingual settings.