NLNDE at SemEval-2023 Task 12: Adaptive Pretraining and Source Language Selection for Low-Resource Multilingual Sentiment Analysis

Mingyang Wang, Heike Adel, Lukas Lange, Jannik Strötgen, Hinrich Schütze

Introduction

In recent years, natural language processing research has attracted considerable interest. However, most studies remain confined to a small number of languages with large amounts of training data available. Low-resource languages, for example, e.g., African languages, are still underrepresented although they are spoken by over a billion people. In this context, the AfriSenti shared task provides a Twitter dataset for sentiment analysis on 14 African languages, promoting the future development of this field. The shared task consists of three sub-tasks: monolingual (Subtask A), multilingual (Subtask B), and zero-shot cross-lingual sentiment analysis (Subtask C). A detailed description can be found in the shared task description papers Muhammad et al. (2023a, b).

In this paper, we describe our submission as Neither Language Nor Domain Experts (NLNDE)We neither know any African languages nor have prior knowledge of the Twitter domain dataset. to the AfriSenti shared task. Given the key challenge of limited training data, we first adopt the language-adaptive and task-adaptive pretraining approaches Gururangan et al. (2020), i.e., LAPT and TAPT, to adapt a pretrained language model to the language and task of interest. Further pretraining the model with such smaller but more task-relevant corpora leads to performance gains in all subtasks.

Second, Cross-lingual transfer has been shown to be an effective method for enhancing the performance of low-resource languages by leveraging one or more similar languages as source languages Lin et al. (2019); Ruder and Plank (2017); Nasir and Mchechesi (2022). However, the 14 African languages covered in this shared task (see Table 1) come from different language families and, therefore, hold different linguistic characteristics. As dissimilar languages could hurt the transfer performance Lin et al. (2019); Adelani et al. (2022), it is important to choose promising languages as the transfer source. Therefore, our system uses transfer learning with an explicit selection of source languages. We apply this approach to the multilingual and zero-shot cross-lingual sentiment analysis tasks (Subtask B and C) and demonstrate that it benefits the performance of each target language. In addition, we investigate different source language selection strategies and show their impact on the final transfer performance.

Our final submission results are the ensemble of the best models with different random seeds. Our system is ranked first in 6 out of 12 languages in subtask A (monolingual), achieves the first place in subtask B (multilingual), and wins for one of two languages in subtask C for the zero-shot cross-lingual transfer.

System Overview

Our system is based on the AfroXLM-R large model Alabi et al. (2022), which applys multilingual adaptive fine-tuning on XLM-R Conneau et al. (2020) with a special focus on African languages. In all three subtasks, we first apply language- and/or task-adaptive pretraining with language- and/or task-specific data to tailor the vanilla AfroXLM-R to our setting. In subtasks B and C, after the adaptive pretraining, we perform source language selection to improve the multilingual and cross-lingual transfer performance.

Most of the current NLP research is based on large language models that have been pretrained on massive amounts of heterogeneous corpora. Gururangan et al. (2020) demonstrate that it is helpful to further tailor a pretrained model to the domain of the target task. They show that continued pretraining with domain-specific and task-specific data consistently improves performance on domain-specific tasks across different domains and tasks. Specifically, they introduce domain-adaptive pretraining (DAPT), i.e., the continued pretraining of the base model on a large corpus of unlabeled domain-specific text. Analogously, task-adaptive pretraining (TAPT) refers to adapting the pretrained model to the task’s unlabeled training data.

A natural extension is the application of this method to multilingual scenarios. Considering different languages as different domains, language-adaptive pretraining can be viewed as a special case of domain-adaptive pretraining. Therefore, we also explore two types of continued pretraining: First, we pretrain the base model with a language-specific corpus, which we will refer to as language-adaptive pretraining (LAPT).In some works Dossou et al. (2022); Alabi et al. (2022), LAPT is also called language-adaptive fine-tuning (LAFT). Second, we adapt the language model on the unlabeled task dataset, i.e., perform task-adaptive pretraining (TAPT).

For the language-specific pretraining, we collect open-source corpora from the multi-domain Leipzig Corpus Collection Goldhahn et al. (2012), covering Wikipedia, Community, Web, and News corpora. Note that the final set of monolingual corpora depends on their availability. There is not a single corpus covering all these languages. Table 1 provides a summary of the monolingual corpora we used for our language-adaptive pretraining.

2 Transfer Learning and Source Selection

Table 1 provides an overview of the languages covered in the shared task. They come from four language families (Afro-Asiatic, Niger-Congo, English-Creole and Indo-European) and, therefore, have different linguistic characteristics. Even inside the same language family, languages can still exhibit distinct linguistic features. For example, although many languages are from the Niger-Congo family (6 out of 14 languages), affixes are very common in Bantu languages (a subgroup of Niger-Congo), but are not typically used in non-Bantu subgroups like Volta-Niger and Kwa.

Previous work has demonstrated that it can still be beneficial to leverage one or more similar languages for cross-lingual transfer learning to the target language Lin et al. (2019). Nevertheless, languages that are dissimilar to the target language could also hinder performance Adelani et al. (2022); Lange et al. (2021). Therefore, it is crucial to properly select source languages to improve the transfer results on the target language.

In this work, we use transfer learning with selected sources for Subtask B and C, as they both involve transferring from multiple languages to a target language. For source language selection, we perform forward and backward source language selection, similar to the corresponding feature selection approaches Tsamardinos and Aliferis (2003); Borboudakis and Tsamardinos (2019).also known as variable selection. Essentially, feature selection is defined as the problem of selecting a minimal-size subset of features that leads to an optimal, multivariate predictive model for a target of interest Tsamardinos and Aliferis (2003). In our task, we consider each candidate source language as a feature. For each target language, we aim to filter out irrelevant or harmful source languages and only keep the beneficial languages as the transfer source.

Forward feature selection usually starts with an empty set of features and adds variables to it, while backward feature selection starts with a complete set of variables and then excludes variables from it. In particular, for forward language selection, given a target language LtL_{t}, we start with a set Sfwd={Lt}S_{fwd}=\{L_{t}\} containing only the target language. We then add each of the other languages Lsi,i=1…N−1L_{s_{i}},i=1\ldots N-1 at a time and obtain N−1N-1 bilingual sets {(Lt,Lsi)}i=1…N−1\{(L_{t},L_{s_{i}})\}_{i=1\ldots N-1}, each with the target language LtL_{t} and one source language LsiL_{s_{i}}. NN refers to the total number of given languages. We experiment with each bilingual language set to build the training dataset for transfer learning. Additionally, we run NN monolingual experiments (one per target language) and use the monolingual performance as the baseline to determine if a candidate source language leads to positive or negative transfer gains. To be more specific: If a bilingual language set (Lt,Lsi)(L_{t},L_{s_{i}}) yields a score of more than 5% above the monolingual performance with LtL_{t}, we consider LsiL_{s_{i}} as a positive source with respect to the target language LtL_{t}.

For backward selection, we start with the complete language set with all NN languages. For each target language, we exclude each of the other N−1N-1 languages and get N−1N-1 language sets, denoted as {(Lt,Ls1…Lsi−1,Lsi+1…LsN−1)}i=1…N−1\{(L_{t},L_{s_{1}}\ldots L_{s_{i-1}},L_{s_{i+1}}\ldots L_{s_{N-1}})\}_{i=1\ldots N-1}. To get a baseline performance for comparison, we randomly select 500 samples from each language to build a small multilingual set. We choose a constant number of samples per language to avoid side effects of the data size on the performance. Given the set of all languages, we remove each language at a time and compare it with the baseline results from the complete language set to investigate the transfer gain of each candidate language on the final performance. If the performance from the language set (Lt,Ls1…Lsi−1,Lsi+1…LsN−1)(L_{t},L_{s_{1}}\ldots L_{s_{i-1}},L_{s_{i+1}}\ldots L_{s_{N-1}}) is more than 5% below the baseline, it shows that the absence of LsiL_{s_{i}} has a large negative impact on performance.

For each of the NN languages, we need to run N−1N-1 experiments with the bilingual language set and 11 baseline experiment with the monolingual dataset. Therefore, for forward source language selection, we need to run N×NN\times N transfer experiments and then select the source languages with positive transfer gains corresponding to each target language via the performance comparison. Similarly, for backward source language selection, N×NN\times N experiments are required for the source language selection.

We apply both selection strategies for subtasks B and C. In subtask C, the language sets do not contain the target language LtL_{t} (as it is a zero-shot task). In particular, for forward selection, this means that we start with an empty set. For the same reason, We run experiments with the complete datasets of all languages as the baseline for both forward and backward selection, as there is no monolingual dataset for the target language in subtask C. Results for the source language selection of Subtask B and C are given in Table 1.

Experimental Setup

We now provide details on our preprocessing steps, the language models and their training.

We preprocess the raw input tweets by removing extra whitespaces, incorrect repeated characters and punctuation. Similar to Nguyen et al. (2020), we replace all URLs with “HTTPURL” and username mentions with “USER” as they have little to no impact on sentiment analysis. When analyzing the data, we noticed a small portion of samples overlapped in the train and dev sets for some languages. Therefore, to measure the actual generalizability of our models, we remove all the overlapping samples from the dev set. We will use dev set* to denote the processed dev set in the following.

2 Pretrained Language Models

Large multilingual pretrained language models (PLMs) like mBERT Devlin et al. (2019) and XLM-R Conneau et al. (2020) have shown impressive capability on many languages for a variety of downstream NLP tasks. They are also often used as initialization checkpoints for adapting to other languages, such as AfroXLM-R, which is initialized from XLM-R and specialized to African languages. In initial experiments, we compare the performance of several multilingual PLMs, including BERT and XLM-R which are trained on hundreds of languages, and AfroLM Dossou et al. (2022), Afro-XLM-R Alabi et al. (2022) as African language-specific models. AfroXLM-R performs best across all three subtasks. Therefore, we select AfroXLM-R large as our base model and apply adaptive pretraining and source language selection on top of it. In addition, in subtask C, we experiment with translating the tweets into English and apply BERTweet Nguyen et al. (2020), a pretrained language model for English tweets.

3 Training Details

For task- and language-adaptive pretraining, we use the AdamW optimizer Loshchilov and Hutter (2017) with a learning rate of 5e-5 and a batch size of 8. For fine-tuning, we use Adam with a learning rate of 2e-5 and a batch size of 32. In both phases, we use a maximum sequence length of 128. The training was done on Nvidia A100 and V100 GPUs.All experiments ran on a carbon-neutral GPU cluster. The results are evaluated using the weighted F1 score on the test set averaged over 5 random seeds. The final submission comes from the majority vote ensemble of different random seeds of the best models.

Results

In this section, we report our results on the three subtasks and discuss our findings and observed limitations of the current work. Our evaluation is based on the weighted F1 score on the test set averaged over 5 random seeds. We use the majority vote method to ensemble our models from different random seeds for submission, we provide the final submission results, as well as the results from several top-ranked systems in the last lines in Table 2 ∼\sim 4. We refer to Appendix A.1 for the results on the development set.

In subtask A, we mainly study the impact of adaptive pretraining on monolingual sentiment analysis. We use the off-the-shelf AfroXLM-R large model as our baseline and fine-tune it on the training dataset of each language, yielding one fine-tuned model per language. Then, we apply LAPT, TAPT and their combination on top of AfroXLM-R. For combined LAPT and TAPT, we begin with AfroXLM-R and apply LAPT then TAPT for the model adaptation. After pretraining, we fine-tune the adapted model on each monolingual dataset for sentiment analysis.

As shown in Table 2, the performance is remarkably improved with adaptive pretraining for most languages, especially with task-adaptive pretraining (TAPT), which leads to a performance gain of 10.58 F1 score on average. LAPT also increases the performance in general, but does not contribute that much in comparison and even degrades the performance for the languages Hausa (ha) and Kinyarwanda (kr). We speculate that, on the one hand, we use relatively small language-specific corpora for LAPT as the covered African languages are indeed low-resource. In contrast, Gururangan et al. (2020) used much larger adaptation corpora for domain-specific pretraining (DAPT). On the other hand, the mismatch of text domains might be another reason: As described in Section 2.1, we use corpora from domains, such as news and Wikipedia for LAPT, while the actual task dataset consists of multilingual tweets involving many Twitter-specific factors, such as code-mixing, misspellings, emojis, or hashtags.

Combining LAPT and TAPT also shows promising results. However, as analyzed before, we hypothesize that most of the benefits come from the more effective TAPT.

2 Subtask B: Multilingual Sentiment Analysis

In the multilingual subtask, we categorize our experiments into three groups: (1) multilingual training of a single model, (2) monolingual training of language-specific models and (3) transfer learning with selected sources. They differ in the composition of the training datasets. In multilingual training, we use all training data from 12 languages. In monolingual training, we use the same language-specific models as in Subtask A (see Section 4.1) and combine the predictions in the end. In transfer learning with selected sources, we perform forward and backward source selection as described in Section 2.2. With the selected source languages given in Table 3, we build the respective training datasets and fine-tune the model for each language. As a baseline, we further group languages based on their language family. This results in four groups, namely Afro-Asiatic, Niger-Congo, English-Creole and Indo-European, details are given in Table 1.

Our multilingual sentiment analysis results are given in Table 3. First, as in Subtask A, task-adaptive pretraining notably improves classification performance in all task settings. Combining LAPT and TAPT is not better than TAPT only. Therefore, due to time constraints, we apply LAPT and LAPT+TAPT only in monolingual training, but not in multilingual training and the transfer learning with selected sources.

Second, fine-tuning the model individually to each target language (monolingual training) outperforms the joint multilingual training, in both the vanilla training (62.91 vs. 48.41) and adaptive pretraining (73.49 vs. 70.65) cases. Furthermore, selecting source languages with positive gains for each target language can further enhance performance over monolingual training. Grouping languages based on their language families (61.81) shows better results than multilingual training (48.41), but it underperforms monolingual training (62.91) and falls behind the transfer learning with selected sources (66.73 and 66.41) by around 5%. Forward and backward source selection gives different results, but they both contribute to the final results.

Finally, another interesting finding is that, in the presence of TAPT, the advantage of specifying languages as training data, i.e., in the cases of monolingual training and transfer learning with selected sources, becomes less pronounced. Specifically, without TAPT, multilingual training achieves an F1 score of 48.41, while monolingual training achieves 62.91 and transfer learning with selected sources achieves 66.73 and 66.41. However, with TAPT, the multilingual training shows a large improvement, yielding an F1 score of 70.65. Although monolingual (73.49) and transfer with selected languages (73.50 and 74.08) still outperform the multilingual result, they become less advantageous. We suppose this is because, with task-adaptive pretraining, the model already adapts to the target language compared with the vanilla model pretrained on a larger language set. As a result, the effect of additionally specifying the source languages is limited.

3 Subtask C: Zero-shot Cross-Lingual Sentiment Analysis

The zero-shot cross-lingual transfer task is particularly challenging, especially when both the source and target languages are low-resource. In this subtask, we also employ different strategies: (1) multilingual transfer, (2) transfer with selected source and (3) BERTweet with English-translated samples.

First, we perform multilingual training, i.e., use all available training datasets from subtask A to fine-tune the AfroXLM-R model. We also perform task-adaptive pretraining with unlabeled multilingual texts here, as in the previous two subtasks.

Second, we perform forward and backward source language selection for the cross-lingual transfer (as detailed in Section 2.2). Here, we use the top 3 selected languages as the transfer source, as they show better performance than using all selected languages as the source in practice. We apply TAPT by using the unlabelled task-specific data from selected sources and the target language.

Finally, we experiment with translating all tweets from the 14 languages to English using the pygoogletranslate API.https://github.com/Saravananslb/py-googletranslation We investigate how the English BERTweet model performs for sentiment analysis. We also perform TAPT on the BERTweet model to adapt it to the unlabelled translated English dataset and then fine-tune the model with the labelled translated English dataset.

The results of subtask C are given in Table 4. As in the previous subtasks, task-adaptive pretraining largely improves performance in all settings. Second, in 7 out of 8 cases, transfer learning with only selected source languages outperforms the multilingual counterparts trained on all languages. The model with a combination of TAPT and backward source language selection achieves the best overall results, which demonstrates the effectiveness of both strategies in subtask C.

Sentiment analysis based on English translations shows competitive performance, but still underperforms transfer learning with source selection. One possible reason could be that the translation quality is not good enough to accurately translate all relevant words with emotional meanings.

4 Discussion

In summary, our work shows that adaptive pretraining and transfer learning with source language selection are effective approaches to tackle sentiment analysis in low-resource languages. Specifically, we demonstrate that (1) adaptive pretraining, especially task-adaptive pretraining, is generally effective across different subtasks and task settings, and (2) transfer learning with source language selection leads to better results than monolingual training. Using only source languages with positive transfer gains for training increases the available training data size on the one hand, and avoids interference from dissimilar languages on the other hand. Notably, forward and backward source selection outperform groupings based on language families in our multilingual experiments.

Limitation and Future Work

One limitation of our work is that the forward and backward selection strategies require a lot of comparative experiments to determine if a candidate language has a positive or negative effect on the target language. For NN languages, we need to perform N∗NN*N transfer experiments for the comparison (as described in Section 2.2). How to automatically select source languages with little manual work is an interesting research question for future work.

Additionally, we found that forward and backward source selection produce different source language results and thus show different transfer scores. In our experiments, neither method completely outperformed the other. We have no conclusive answer to which method is better. Also, We have not conducted an in-depth study on the relationship between the selected sources for the target language and their linguistic correlation. This is another limitation of the current work that could be addressed in future research – in particular when involving language experts.

Related Work

Large multilingual PLMs, such as mBERT Devlin et al. (2019) and XLM-R Conneau et al. (2020) cover more than 100 languages for natural language processing tasks and exhibit good generalization abilities over a large number of languages. However, most of them include few African languages due to the lack of large open-source monolingual corpora Hedderich et al. (2021). Prior work developed African language-centric PLMs to address this under-representation. Among them, AfriBERTa Ogueji et al. (2021) uses the RoBERTa Zhuang et al. (2021) architecture and trains the model from scratch with corpora from 11 African languages. AfroLM Dossou et al. (2022) proposes to use a novel self-active learning framework and the model is trained from scratch on 23 African languages. Another strategy is to initialize a language model from an existing model and continue to train it with a special focus on African languages. AfroXLM-R Alabi et al. (2022) performs multilingual adaptive fine-tuning based on XLM-R on 17 highest-resourced African languages and 3 other high-resource languages spoken on the African continent. AfroXLM-R performs well on African language tasks, such as named entity recognition and sentiment analysis Alabi et al. (2022); Dossou et al. (2022). We therefore use it as the base model in this shared task.

Adaptive pertaining.

Gururangan et al. (2020) demonstrate that it is helpful to further tailor a pretrained model to a target domain and task. In particular, they introduce domain-adaptive pretraining, which continues the pretraining of the model on domain-specific unlabeled data, and task-adaptive pretraining, which further pretrains the model on the task’s unlabeled data. Experimental results show that these two strategies lead to remarkable performance gains. We adopt this idea to our tasks.

Transfer learning with source selection.

Selecting data for transfer learning has been explored in different prior work, i.e., Ruder and Plank (2017); Lin et al. (2019); Lange et al. (2022). For example, Ruder and Plank (2017) learn to select positive sources using Bayesian optimization. LangRank Lin et al. (2019) considers the source language selection for transfer learning as a ranking problem. They train a ranking model to select languages with a positive transfer gain from a larger set of possible languages. In contrast, we adopt the idea of forward and backward feature selection Tsamardinos and Aliferis (2003) and use a much simpler approach based on transfer score comparison to select source languages.

Conclusion

In this work, we introduce our sentiment analysis system for the AfriSenti shared task, which is ranked first in 8 out of 15 tracks and performs competitively on the others. It consists of language-adaptive and task-adaptive pretraining on top of the AfroXLM-R model, together with transfer learning with source language selection. We demonstrate that tailoring the pretrained model to the target language and task considerably improves the performance across all task settings. Additionally, transfer learning with source language selection further improves the results in the multilingual and zero-shot cross-lingual tasks by avoiding potential negative transfer gains from dissimilar languages. A future research direction is to automatically select source languages with positive transfer gains without the need of manually comparing the source-to-target transfer score.

Acknowledgments

We thank the AfriSenti organizers for their time to prepare the data for a large variety of languages and run the competition in a smooth way. The shared task provides an excellent platform for researchers to collaborate and share their knowledge, and also promotes the future development in the field of NLP for African languages.

References

Appendix A Appendix

Here, we provide the experimental results of all three subtasks on our processed dev set* (for details, please refer to Section 3.1).