UCAS-IIE-NLP at SemEval-2023 Task 12: Enhancing Generalization of Multilingual BERT for Low-resource Sentiment Analysis
Dou Hu, Lingwei Wei, Yaxin Liu, Wei Zhou, Songlin Hu
Introduction
Sentiment analysis is a critical aspect of natural language processing with numerous applications, including public opinion monitoring (Boon-Itt and Skunkan, 2020), healthcare services (Zunic et al., 2020), and recommendation systems (Hu et al., 2021b). However, performing sentiment analysis in low-resource languages poses significant challenges, including the scarcity of labeled data and linguistic resources, as well as the diversity of languages and dialects (Lo et al., 2017; Oueslati et al., 2020). In SemEval-2023 Task 12 (Muhammad et al., 2023b), the focus is on sentiment analysis for African languages in Twitter, which further exacerbates the challenges due to the presence of tone, code-switching, and digraphia phenomena Adebara and Abdul-Mageed (2022).
Although multilingual pre-trained language models (multilingual PTMs) (Conneau and Lample, 2019; Conneau et al., 2020) have shown potential in cross-lingual transfer learning compared to monolingual PTMs (Devlin et al., 2019; Hu et al., 2022a), they have limitations in capturing nuances and cultural differences within a language, especially in the context of dialects and regional variations.
In this paper, we propose a generalized multilingual system named SACL-XLMR to address these limitations and enhance the generalization of multilingual PTMs for under-represented languages, particularly African languages. Our system leverages a lexicon-based multilingual BERT model to facilitate language adaptation and sentiment-aware representation learning. Additionally, we apply a supervised adversarial contrastive learning (SACL) technique (Hu et al., 2023) to learn sentiment-spread structured representations and enhance model generalization.
We present the details of the proposed system and evaluate its performance on SemEval-2023 Task 12. Our system achieves remarkable performance, outperforming baselines by +1.1% weighted-F1 score on multilingual sentiment classification subtask and by +2.8% weighted-F1 score on zero-shot sentiment classification subtask in the AfriSenti-SemEval datasets (Muhammad et al., 2023a). Moreover, following the AfriSenti SemEval Prizeshttps://afrisenti-semeval.github.io/prizes/ and the task description (Muhammad et al., 2023b), our system obtains the 1st rank on the zero-shot classification subtask in the official ranking. We conducted experiments to demonstrate the effectiveness of our approach, highlighting the potential of our system in overcoming the challenges of low-resource sentiment analysis.
Background
The SemEval-2023 Task 12: Sentiment analysis for African languages (AfriSenti-SemEval) (Muhammad et al., 2023b) is the first Afro-centric SemEval shared task for sentiment analysis in Twitter. It consists of three subtasks, i.e., monolingual, multilingual, and zero-shot sentiment classification. Brief descriptions of the last two subtasks that our team focuses on are as follows:
Multilingual Sentiment Classification. Given combined training data of multiple African languages, determine the polarity of a tweet on the combined test data of the same languages (positive, negative, or neutral). This subtask has only one track with 12 languages (Amharic, Algerian Arabic/Darja, Hausa, Igbo, Kinyarwanda, Moroccan Arabic/Darija, Mozambican Portuguese, Nigerian Pidgin, Swahili, Twi, Xitsonga, and Yorùbá), i.e., a multilingual track with 12 African languages.
Zero-Shot Sentiment Classification. Given unlabelled tweets in two African languages (Tigrinya and Oromo), leverage any or all available training datasets of source languages (12 African languages in the multilingual track) to determine the sentiment of a tweet in the two target languages. This task has two tracks, i.e., a zero-shot Tigrinya track and a zero-shot Oromo track.
The AfriSenti datasetshttps://github.com/afrisenti-semeval (Muhammad et al., 2023a) are a collection of multilingual Twitter datasets that consist of 110,000+ tweets in 14 low-resource African languages from four language families for sentiment analysis. The statistics of each monolingual tweet datasets are reported in Table 1. The datasets involve tweets labeled with three sentiment classes (positive, negative, neutral). Each tweet is annotated by three native speakers following the sentiment annotation guidelines Mohammad (2016) and the final label for each tweet is determined by majority voting (Davani et al., 2022). If a tweet conveys both a positive and negative sentiment, the stronger sentiment should be chosen.
Related Work
Sentiment analysis has evolved from lexicon-based approaches to more advanced machine learning and deep learning-based methods (Medhat et al., 2014). Previous works in sentiment analysis have focused on various levels of granularity, such as aspect (Pontiki et al., 2014), sentence (Hu et al., 2021a), and document (Wei et al., 2020), as well as different modalities (Zadeh et al., 2017; Hu et al., 2022b) and languages (Boiy and Moens, 2009; Balahur and Turchi, 2014).
2 Low-resource Sentiment Analysis
Despite the success of polarity classification in high-resource languages, noisy user-generated data in under-represented languages presents a challenge (Yimam et al., 2020). Recently, several studies have proposed approaches for sentiment analysis on low-resource languages (Lo et al., 2017; Yimam et al., 2020). Besides, Moudjari et al. (2020); Adebara and Abdul-Mageed (2022); Muhammad et al. (2023a) have relied on manual annotation by native speakers or expert annotators to build sentiment analysis datasets in low-resource languages.
System Overview
In this section, we describe our system adopted in SemEval-2023 Task 12, where we design a generalized multilingual system named SACL-XLMR for sentiment analysis on low-resource languages.
The network structure of SACL-XLMR consists of a multilingual BERT (i.e., an embedding layer and Transformer encoder) and a sentiment classifier.
We apply a multilingual BERT model (Conneau and Lample, 2019; Alabi et al., 2022) on monolingual corpus to facilitate language adaptation. Besides, sentiment lexicon knowledge for each language is used to enhance sentiment-aware representation learning.
Formally, given an input token sequence where refers to -th token in the -th input sample, and is the maximum sequence length, the model learns to generate the context representation of the input token sequences:
where [CLS] and [SEP] are special tokens, usually at the beginning and end of each sequence, respectively. refers to a token sequence of sentiment lexicon prefix corresponding to the input sequence. indicates the hidden representation of the -th input sample, computed by the representation of [CLS] token in the last layer of the encoder.
Sentiment Classifier
Finally, according to the obtained representations, a sentiment classifier is applied to predict the sentiment label of each sample.
2 Optimization Objective
Supervised contrastive learning (SCL) (Khosla et al., 2020; Gunel et al., 2021) is utilized to learn a generalized feature representation by capturing similarities between examples within a class and contrasting them with examples from other classes. However, directly compressing the feature space of each class can harm fine-grained features, which limits the model’s ability to generalize. Recently, a new technique named supervised adversarial contrastive learning (SACL) (Hu et al., 2023) has been proposed to address this issue by learning class-spread structured representations. The SACL uses both original and adversarial samples to effectively utilize prior information on label consistency and retain fine-grained features.
In this task, we apply the SACL technique to learn sentiment-spread representations and enhance the generalization of multilingual BERT. Formally, let us denote as the set of samples in a batch. Define as the set of indices of all positives in the batch distinct from , and is its cardinality. The loss function of soft SCL is a weighted average of CE loss and SCL loss with a trade-off scalar parameter , i.e.,
and denote the value of one-hot vector and probability vector at class index k, respectively. . . is a pairwise similarity function, i.e., dot product. is a scalar temperature parameter that controls the separation of classes.
At each step of training, under the soft SCL objective, we apply an adversarial training strategy (e.g., FGM (Miyato et al., 2017)) on original samples to generate adversarial samples. These samples can be seen as hard positive examples, which spread out the representation space for each sentiment class and confuse robust-less models. After that, we utilize a new soft SCL on obtained adversarial samples to maximize the consistency of sentiment-spread representations with the same sentiment label. Following the above calculation process of on original samples, the optimization objective on corresponding adversarial samples can be easily obtained in a similar way, i.e., .
The overall loss of SACL is defined as a sum of two soft SCL losses on both original and adversarial samples, i.e.,
Experimental Setup
We compare SACL-XLMR with the following several methods:
Random is based on random guessing, choosing each class/label with an equal probability.
XLM-R (Conneau and Lample, 2019) is a multilingual variant of RoBERTa (Liu et al., 2019). It is pre-trained on filtered CommonCrawl data containing 100 languages. We use xlm-roberta-basehttps://huggingface.co/ to initialize XLM-R.
AfriBERTa (Ogueji et al., 2021) is an Afro-centri multilingual language model pretrained on 11 African languages. It is trained on an aggregation of datasets from the BBC news website and Common Crawl. We use castorini/afriberta_large3 to initialize AfriBERTa.
AfroXLMR (Alabi et al., 2022) is an XLM-R model adapted to African languages. It is obtained by MLM adaptation of XLM-R on 17 African languages covering the major African language families and 3 high resource languages (Arabic, French, and English). We use Davlan/afro-xlmr-large3 to initialize AfroXLMR.
We report the comparison of our SACL-XLMR and the above PTMs in Table 2.
2 Implementation Details
All experiments are conducted on a single NVIDIA Tesla V100 32GB card. Stratified k-fold cross validation (Kohavi, 1995) is performed to split combined training and validation data of 12 African languages into 5 folds. Train/validation sets for Oromo (orm) and Tigrinya (tir) are not used due to the limited size of the data. We only evaluate on them in a zero-shot transfer setting. We choose the optimal hyperparameter values based on the the average result of validation sets for all folds, and evaluate the performance of our system on the test data. Following the scoring program of AfriSenti-SemEval, we report the weighted-F1 (w-F1) score to measure the overall performance.
Our SACL-XLMR is initialized with the Davlan/afro-xlmr-large3 parameters, due to the nontrivial and consistent performance in both subtasks. The network parameters are optimized by using Adam optimizer (Kingma and Ba, 2015). The class weights in CE loss are applied to alleviate the class imbalance problem and are set by their relative ratios in the train and validation sets. The detailed experimental settings on both two subtasks are in Table 3.
To effectively utilize sentiment lexicons of partial languages in the AfriSenti datasets, we concatenate the corresponding lexicon prefix with the original input text. Given the -th input sample, the lexicon prefix can be represented as where is the sentiment label, refers to the corresponding -th lexicon token in the original sequence. For our final system, we only use sentiment lexicons on the zero-shot subtask. We do not use it on the multilingual subtask due to the fact that some languages in the multilingual target corpus do not have available sentiment lexicons, making it difficult for the model to adapt effectively.
Results and Analysis
The overall results for both subtasks are summarized in Table 4 and 5. From the results, it is not surprising that all pre-trained models clearly outperformed the Random baseline. The proposed SACL-XLMR consistently outperformed the comparison methods on both subtasks. Specifically, SACL-XLMR achieved 1.1% and 2.8% absolute improvements on the multilingual and zero-shot sentiment classification subtasks, respectively.
Moreover, we present the official results from several top-ranked systems for the zero-shot sentiment classification subtask in AfriSenti-SemEval Shared Task (i.e., SemEval-2023 Task 12) in Table 1. Our submitted system obtained the 1st overall rank on the zero-shot sentiment classification subtask in the official ranking.
2 Ablation Study
In this part, we conduct ablation studies by removing key components of SACL-XLMRfull to further understand the proposed model:
- w/o Lexicon refers to removing the sentiment lexicon.
- w/o SACL is an ablated model removing the supervised adversarial contrastive learning objective.
- w/o Lexicon - w/o SACL indicates removing both sentiment lexicons and SACL objective, degenerated to AfroXLMR.
Figure 2 shows results of ablation studies on two subtasks for low-resource sentiment analysis. Our SACL-XLMR and SACL-XLMRfull yield the best performance on multilingual and zero-shot sentiment classification subtasks, respectively. When removing SACL objective, the results consistently decline on all subtasks, showing the effectiveness of SACL.
For the multilingual sentiment classification subtask, SACL-XLMRfull obtains sub-optimal results. This is most likely due to the fact that some languages in the target corpus do not have available sentiment lexicons, making it difficult for the model to adapt effectively. Also, another caused factor is the incompleteness and poor quality of lexicon. For the zero-shot sentiment classification subtask, the SACL-XLMRfull yields the best performance on both tir and orm languages. It shows the effectiveness of sentiment lexicons in zero-shot scenarios, even if its quality is not good enough.
3 Error Analysis
Figure 3 shows an error analysis of our system on two subtasks of AfriSenti-SemEval, including a multilingual test set and two zero-shot test sets. The normalized confusion matrices are used to evaluate the quality of the predicted outputs of SACL-XLMR.
From the diagonal elements of the matrices, true positives of non-neutral labels exceed those of the neutral label. The results show that positive and negative features are more likely to adapt to low-resource languages. Besides, the above phenomenon is more obvious for tir and orm languages. It indicates that SACL-XLMR can further facilitate language adaptation for low-resource languages by making full use of existing sentiment lexicons which contain only positive and negative words.
The confusion matrix of SACL-XLMR reveals the most confusing pair of sentiment labels: neutral to negative, especially for tir and orm languages in a zero-shot setting. The performance on orm language is relatively poor. Apart from the complexity of the language and the phenomenon of data scarcity, it is also due to the significant differences between orm and other African languages. Considering the above issues make the task optimization more difficult, there is still a lot of room for improvement.
Conclusion
In this paper, a multilingual system named SACL-XLMR has been proposed for sentiment analysis on low-resource African languages. The system employs a lexicon-based multilingual BERT to facilitate language adaptation and sentiment-aware representation learning. It also uses a supervised adversarial contrastive learning technique to learn sentiment-spread structured representations and enhance model generalization. The system achieved competitive results, largely outperforming the comparison baselines on both multilingual and zero-shot sentiment classification subtasks, and obtained the 1st rank on zero-shot classification subtask in the official ranking.
Acknowledgements
All the work in this paper are conducted during the SemEval-2023 Competition. We thank the SemEval-2023 organizers and AfriSenti-SemEval task organizers for making this research possible. We also appreciate the anonymous reviewers for their insightful and constructive comments that have helped us improve the quality of the paper.