AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings

Anne Lauscher, Rafik Takieddin, Simone Paolo Ponzetto, Goran Glavaš

Introduction

Recent research offered evidence that distributional word representations (i.e., word embeddings) induced from human-created text corpora exhibit a range of human biases, such as racism and sexism [Bolukbasi et al., 2016, Caliskan et al., 2017]. With word embeddings ubiquitously used as input for (neural) natural language processing (NLP) models, this brings about the jeopardy of introducing stereotypical unfairness into NLP models, which can reinforce existing social injustices, and therefore be harmful in practical applications. For instance, consider the seminal gender bias example “Man is to computer programmer as woman is to homemaker”, which is algebraically encoded in the embedding space with the analogical relation man⃗−computer programmer⃗≈woman⃗−homemaker⃗\vec{\textit{man}}-\vec{\textit{computer programmer}}\approx\vec{\textit{woman}}-\vec{\textit{homemaker}} [Bolukbasi et al., 2016]. The existence of such biases in word embeddings stems from the combination of (1) human biases manifesting themselves in terms of word co-occurrences (e.g., the word woman appearing in a training corpus much more often in the context of homemaker than together with computer programmer) and (2) the distributional nature of the word embedding models [Mikolov et al., 2013, Pennington et al., 2014, Bojanowski et al., 2017], which induce word vectors precisely by exploiting word co-occurrences, i.e., thus also encoding the human biases as a (negative) side-effect, which represents, expressed according to the taxnomy of harms proposed by ?), a representational harm, more specifically, stereotyping. In order to quantify the amount of bias in word embeddings, ?) proposed the Word Embedding Association Test (WEAT), which is based on the associative difference in terms of semantic similarity between two sets of target terms, e.g., male and female terms, towards two sets of attribute terms, e.g., career and family terms. Most recently, the WEAT test, measuring the degree of explicit bias in the distributional space, has been coupled with other tests, aiming to measure other aspects of bias, such as the amount of implicit bias [Gonen and Goldberg, 2019] or the presence of the analogical bias [Lauscher et al., 2020].

While there is evidence that distributional vectors often encode human biases, the amount of biases does not seem to be universal across different languages and corpora, as recently shown by ?) in the analysis of distributional biases across seven different languages. In this work, we focus on the multi-dimensional analysis of biases in Arabic word embeddings. The motivation for this work is twofold: (1) Arabic is one of the most widely spoken languages in the world:According to ?), Arabic is the fifth most spoken language in the world with close to 300 million native speakers. this means that the biases encoded in language technology for Arabic have the potential for affecting more people than for most other languages; (2) language resources for Arabic – large corpora [Goldhahn et al., 2012], pretrained word embeddings [Mohammad et al., 2017, Bojanowski et al., 2017], and datasets for measuring semantic quality of Arabic embeddings [Elrazzaz et al., 2017, Cer et al., 2017] – are publicly available, allowing for the analyses of biases that these resources potentially hide.

As a first step in the analysis of language technology biases for Arabic, we present AraWEAT, an Arabic extension to the multilingual XWEAT framework [Lauscher and Glavaš, 2019]. Because the WEAT test [Caliskan et al., 2017], despite being derived from an established test from psychology [Nosek et al., 2002], has recently been shown to systematically overestimate the bias present in an embedding space [Ethayarajh et al., 2019], in this work, we couple it with several other bias tests, designed to capture and quantify other aspects of human biases: Embedding Coherence Test [Dev and Phillips, 2019], Bias Analogy Test [Lauscher et al., 2020] and Implicit Bias Tests [Gonen and Goldberg, 2019].

Our work, which is to the best of our knowledge the first study on quantifying biases in Arabic distributional word vector spaces, yields some interesting findings: biases seem more prominent in vectors trained on texts written in Egyptian Arabic than those written in Modern Standard Arabic (MSA). Also, the implicit gender bias in Arabic news corpora seems to be steadily on the rise over the ten year period between 2007 and 2017. Finally, we find evidence that the explicit bias effects, as measured by the WEAT test, in embeddings trained on the entire Arabic news corpus roughly correspond to averaging the biases measured across embeddings trained on temporally disjunct subsets of the corpus.

AraWEAT

We present AraWEAT, our framework allowing for multi-dimensional analysis of bias in Arabic distributional word vector spaces.

At the core of our extension are the Arabic bias test specifications, which are based on the original English WEAT test data. WEAT is an adaptation of the Implicit Association Test [Nosek et al., 2002], which quantifies biases as association differences measured in terms of response times of human subjects when exposed to different sets of stimuli. WEAT, in turn, measures the association differences in terms of the difference in semantic similarity between two sets of target terms towards two sets of attribute terms.

Our creation of WEAT tests for Arabic starts with automatically translating, using Google Translate, the term sets (i.e., the terms from each of two target and two attribute lists) from the English WEAT tests. We then hired a native speaker of modern standard Arabic (MSA), who manually verified and, when needed, corrected the translations. Since Arabic is a language with grammatical genders, we made sure to account for both genders when translating the terms so that we do not artificially introduce a bias in our test specifications (e.g., we translated the genderless English word engineer as both \<مهندس¿ (engineer m.) and \<مهندسة¿ (engineer f.). While initially considered, we did not translate WEAT test specifications to the different Arabic dialects, as the differences between the MSA translations and dialectal translations for the terms from the WEAT test were observed only in a negligible fraction of cases; and even in those cases the MSA translation is also in usage in other Arabic dialects.Although possibly less frequently than the dialectal translation. We further omitted WEAT tests 33–66 and 1010 because they are based on proper names. While it has been shown that names are a good proxy for identifying and removing bias towards specific groups of people [Hall Maudslay et al., 2019], it is difficult to “translate” them.Furthermore, WEAT tests 33–55 are tailored to test racial biases towards African-Americans, which is arguably much less prominent in the Arabic cultural area. As an example of the resulting AraWEAT test, Table 1 list the Arabic translation of WEAT test T7T7. An overview on the remaining tests with their respective target and attribute term sets is provided in Table 2.

2 Bias Evaluation Methodology

Aiming towards a holistic picture of biases encoded in Arabic word vectors, we put together several bias tests that quantify both implicit and explicit biases: (1) WEAT [Caliskan et al., 2017], (2) Embedding Coherence Test (ECT) [Dev and Phillips, 2019], (3) Bias Analogy Test (BAT) [Lauscher et al., 2020], and (4) Implicit Bias Test with K-Means++ (KM) [Gonen and Goldberg, 2019]. For all bias tests, we adopt the notion of implicit and explicit bias specifications as proposed by ?): an explicit bias specification consists here of two sets of target terms and two sets of attribute terms BE(T1,T2,A1,A2)B_{E}(T_{1},T_{2},A_{1},A_{2}). The idea is to measure the bias between the target sets, e.g., science and art, towards the attribute sets, e.g., male vs. female terms, or vice versa. In contrast, an implicit bias specification consists of target terms only, i.e., BE(T1,T2)B_{E}(T_{1},T_{2}). Accordingly, the intuition is to measure bias between the target term representations only, and not its explicit manifestation with regard to other concepts. Furthermore, we report the semantic quality for all word embedding spaces we induced ourselves: to this end, we report the scores on predicting sentence-level semantic similarity for Arabic on the dataset from SemEval 2017 Task 1 [Cer et al., 2017]; we obtain sentence embeddings simply as averages of word embeddings [Cer et al., 2017, Glavaš et al., 2018].

Let BE(T1,T2,A1,A2)B_{E}(T_{1},T_{2},A_{1},A_{2}) be an explicit bias specification consisting of two sets of target terms T1T_{1} and T2T_{2}, and two sets of attribute terms, A1A_{1} and A2A_{2}. ?) define the WEAT test statistic s(T1,T2,A1,A2)s(T_{1},T_{2},A_{1},A_{2}) as the association difference that T1T_{1} and T2T_{2} exhibit w.r.t. A1A_{1} and A2A_{2} – the association is measured as the average semantic similarity of T1T_{1}/T2T_{2} terms with terms from A1A_{1} and A2A_{2}:

with associative difference for a term tt computed as:

with t\mathbf{t} as the distributional vector of term tt and cos as the cosine of the angle between two vectors. The significance of the test statistic is measured by the non-parametric permutation test in which the s(T1,T2,A1,A2)s(T_{1},T_{2},A_{1},A_{2}) is compared to s(X1,X2,A1,A2)s(X_{1},X_{2},A_{1},A_{2}), where (X1X_{1}, X2X_{2}) denotes a random, equally-sized split of the terms in T1∪T2T_{1}\cup T_{2}. A larger WEAT effect size indicates a larger bias.

Embedding Coherence Test (ECT).

Given an AraWEAT explicit bias specification BE(T1,T2,A1,A2)B_{E}(T_{1},T_{2},A_{1},A_{2}), ECT operates on the bias specification which “collapses” the two AraWEAT attribute sets into a single set: BE(T1,T2,A=A1∪A2)B_{E}(T_{1},T_{2},A=A_{1}\cup A_{2}). Next, as proposed by ?), we compute the vectors t1\mathbf{t_{1}} and t2\mathbf{t_{2}} as averages of the word vectors of terms in T1T_{1} and T2T_{2}, respectively. Then, we obtain two vectors of similarities by computing the cosine similarity between the vector of each term in AA and the mean vectors t1\mathbf{t_{1}} and t2\mathbf{t_{2}}. The ECT score is finally the Spearman correlation between the two obtained similarity vectors. The intuition is to assess, whether the similarities of the average vectors t1\mathbf{t_{1}} and t2\mathbf{t_{2}}, which represent the two target term sets, with the attribute terms are correlating. The larger the ECT correlation, the lower the bias.

Bias Analogy Test (BAT).

Inspired by ?)’s famous anology, the idea behind BAT is to quantify the fraction of biased analogies that result from querying the embedding space. Given an AraWEAT test BE(T1,T2,A1,A2)B_{E}(T_{1},T_{2},A_{1},A_{2}), following ?), we create all possible biased analogies t1−t2≈a1−a2\mathbf{t}_{1}-\mathbf{t}_{2}\approx\mathbf{a}_{1}-\mathbf{a}_{2} for (t1,t2,a1,a2)∈T1×T2×A1×A2(t_{1},t_{2},a_{1},a_{2})\in T_{1}\times T_{2}\times A_{1}\times A_{2}. Next we create two query vectors – q1=t1−t2+a2\mathbf{q}_{1}=\mathbf{t}_{1}-\mathbf{t}_{2}+\mathbf{a}_{2} and q2=a1−t1+t2\mathbf{q}_{2}=\mathbf{a}_{1}-\mathbf{t}_{1}+\mathbf{t}_{2} – for each tuple (t1,t2,a1,a2)(t_{1},t_{2},a_{1},a_{2}). We then rank the vectors in the vector space according to the Euclidean distance with q1q_{1} and q2q_{2}, respectively, and report the percentage of cases where: a1a_{1} is ranked higher than a term a2′∈A2∖{a2}a^{\prime}_{2}\in A_{2}\setminus\{a_{2}\} for q1\mathbf{q}_{1} and a2a_{2} is ranked higher than a term a1′∈A1∖{a1}a^{\prime}_{1}\in A_{1}\setminus\{a_{1}\} for q2\mathbf{q}_{2}. The higher the BAT score, the higher the bias.

Implicit Bias Test: K-Means++ (KM).

Sometimes, bias is not expressed explicitely, i.e., as bias between two target term sets in explicit relation towards certain attribute sets, but manifests implicitly. In order to additionally reflect this type of bias in our study, we follow ?) and test the Arabic word vector spaces for the amount of implicit bias by clustering terms from T1T_{1} and T2T_{2} with KMeans++ [Arthur and Vassilvitskii, 2007]. The higher the clustering accuracy, the higher the bias. We report the averaged accuracy over 2020 independent runs.

Semantic Quality (SQ).

For the embedding models we train ourselves, we additionally report the semantic quality of the space by predicting sentence-level semantic similarity on the SemEval 2017 Task 1 for Arabic (ar-ar) [Cer et al., 2017]. Let sa=ea1,...,ean\mathbf{s_{a}}={e_{a1},...,e_{an}} be the set of embeddings of words in sentence aa and let sb=eb1,...,eam\mathbf{s_{b}}={e_{b1},...,e_{am}} be the sequence of embedding representations for individual words in sentence bb. We obtain aggregated sentence representations, by averaging the embeddings of words in the sentence: s=1l∑i=1lei\mathbf{s}=\frac{1}{l}\sum_{i=1}^{l}e_{i} and finally predict the similarity score as cos(sa,sb)cos(\mathbf{s}_{a},\mathbf{s}_{b}).This method was used as the simple aggregation baseline in the corresponding SemEval shared task. We report Pearson correlation between our predicitions and the gold similarity annotations.

Dimensions of Bias Analysis.

We run our tests along 55 different dimensions: (1) embedding methods: we compare embeddings induced using Skip-Gram, CBOW and FastText embedding models; (2) source text types: we analyze vector spaces induced from corpora originating from different sources (Wikipedia, news, Twitter);While Arabic Wikipedia is dominantly written in MSA, Twitter is likely to exhibit non-negligible amounts of dialectical and colloquial Arabic. (3) vector sizes and preprocessing: we hypothesize that biases might be more prominent in higher-dimensional vectors. To this end, we compare 100100- vs. 300300-dimensional embeddings. Furthermore, we analyze the effect of unigram vs. n-gram preprocessing of Arabic text, as offered by pretrained vectors AraVec [Mohammad et al., 2017]; (4) corpus size: ?) hypothesize that biases might be more expressed in bigger corpora. To further investigate this, we run several experiments controlling for corpus size; (5) temporal intervals: lastly, we conduct a diachronic bias analysis by training embeddings on corpora from different time periods.

Distributional Word Vector Spaces.

We conduct our analysis on (a) pretrained distributional word vector spaces from AraVechttps://github.com/bakrianoo/aravec [Mohammad et al., 2017] and FastTexthttps://dl.fbaipublicfiles.com/fasttext/vectors-crawl/cc.ar.300.vec.gz [Bojanowski et al., 2017] and (b) embedding spaces we trained in order to be able to control for corpora size and preprocessing. In (b), we use Arabic corpora from the Leipzig Corpora Collectionhttp://wortschatz.uni-leipzig.de/en/download/ [Goldhahn et al., 2012].

Findings

We present and discuss the findings of our analysis employing AraWEAT.

Bias scores for 300-dimensional pretrained Fasttext (FT) and AraVec (AV) embedding spaces are shown in Table 3. For both (FT) and (AV), we evaluated all available spaces, pretrained on different corpora. For FT, we investigate two models, one trained on the portions of Wikipedia and CommonCrawl corpora written in Modern Standard Arabic (MS) and the other on portions written in Egyptian Arabic.The language identification was performed automatically using the FT Language Detector We evaluate the four variants of AraVec vectors: (a) trained using either Skip-Gram (SG) or CBOW (CB) on (b) either Wikipedia (Wiki) or Twitter (Twitter) text. Interestingly, most of these embedding spaces fail to exhibit significant explicit gender biases according to WEAT tests T7T7 and T8T8. However, the gender biases seem to be rather present implicitly (KM) in most spaces. Comparing FT Arabic versus FT Egyptian, both implicit and explicit bias seems to be slightly more pronounced in the Egyptian than in the MSA corpus. Results of comparison over text types support the unexpected finding for other languages [Lauscher and Glavaš, 2019]: embeddings built from user-generated content on average do not encode more bias than their counterparts trained on Wikipedia.

Embedding dimensionality and preprocessing.

Next, we evaluate the effects of specific hyperparameter settings using the AraVec pretrained vector spaces. AraWEAT bias effect sizes for different embedding dimensionalities and model types are listed in Table 6. For the AraWEAT test specifications T1T1, T7T7, and T8T8, we did not observe prominent variance in the amount of explicit bias w.r.t. the vector dimensionality or pre-processing type. For the remaining test – T2T2 – the explicit bias (according to the WEAT test) is somewhat more pronounced in the lower-dimensional embeddings and in the n-gram versions of the AraVec embeddings.

Diachronic Analysis and Corpora Sizes.

Table 4 displays WEAT effect sizes for test T7T7 (gender bias) in MSA 300-dimensional distributional word vector spaces we trained on the (temporally) disjunctive Arabic portions of the Leipzig News Corpora of sizes 300K and 1M sentences, respectively.

The smaller corpus, consisting of 300300K sentences, exhibits no significant bias effect sizes across all years. This finding is in line with previous observations ?) that biases might be more expressed in embedding spaces obtained on bigger corpora. This could be a reflection of the overall quality of distributional vectors, which is lower when vectors are trained on smaller corpora (as supported by the corresponding STS scores). In the spaces obtained on the larger corpora segments, consisting of 11M-sentences, significant explicit (W) gender biases are present in years 20092009, 20152015, and 20172017, with very similar effect sizes (between .92.92 and .97.97). The implicit gender bias (KM), on the other hand, steadily rises over the entire period under investigation (2007-2017).

Finally, we investigate how the biases in the embedding space induced on the whole Arabic Leipzig News corpus (2007–2017, CONC) relates to the biases detected in embedding spaces induced from its different, temporally non-overlapping subportions. To this end, we average the biases measured on embeddings trained on its yearly subsets (AVG). The correlation results, over all four tests and two measures (W, KM), are shown in Table 6. Indeed, the biases of the whole corpus (CONC) seem to be highly correlated with the averages of biases of subcorpora (AVG): we measure a substantial Pearson correlation of 66%66\% between the two sets of scores (AVG and CONC). This would suggest that one can roughly predict the biases of (embeddings trained on) a large corpus by aggregating the biases of (embeddings trained on) its (non-overlapping) subsets.

Related Work

?) were the first to study bias in distributional word vector spaces. Using an analogy test, they demonstrate gender stereotypes manifesting in word embeddings and propose the notion of the bias direction, upon which they base a debiasing method called hard-debiasing. ?) adapt the Implicit Association Test (IAT) [Nosek et al., 2002] from psychology for studying biases in distributional word vector spaces. The test, dubbed Word Embedding Association Test (WEAT), measures associations between words in an embedding space in terms of cosine similarity between the vectors. They propose 1010 stimuli sets, which we adapt in our work. Later, ?) extend the analysis to three more languages, Dutch, German, and Spanish, but only focus on gender bias. XWEAT, the cross-lingual and multilingual WEAT framework [Lauscher and Glavaš, 2019], covers German, Spanish, Italian, Russian, Croatian, and Turkish. XWEAT analyses also focused on other relevant dimensions such as embedding method and similarity measures. ?) focus on measuring bias in languages with grammatical gender. Several research efforts produced new bias tests: ?) propose the Embedding Coherence Test (ECT) with the intuition of capturing whether two sets of target terms are coherently distant from a set of attribute terms. They also propose several debiasing methods. ?) show that many debiasing methods only mask but do not fully remove biases present in the embedding spaces. They propose to additionally test for implicit biases, by trying to classify or cluster the sets of target terms. ?) unify the different notions of biases into explicit and implicit bias specifications, based on which they propose methods for quantifying and removing biases. While their is some effort to account for gender-awareness in Arabic machine translation [Habash et al., 2019], we are, to the best of our knowledge, the first to measure bias in Arabic Language Technology.

Conclusion

Language technologies should aim to avoid reflecting negative human biases such as racism and sexism. Yet, the ubiquitous word embeddings, used as input for many NLP models, seem to encode many such biases. In this work, we extensively quantify and analyze the biases in different vector spaces built from text in Arabic, a major world language with close to 300M native speakers. To this effect, we translate existing bias specifications from English to Arabic and investigate biases in embedding spaces that differ over several dimensions of analysis: embedding models, corpora sizes, type of text, dialectal vs. standard Arabic, and time periods. Our analysis yields interesting results. First, we confirm some of the previous findings for other languages, e.g., that biases are generally not more pronounced in user-generated text and that embeddings trained on larger corpora lead to more prominent biases. Secondly, our results suggest more bias is present in dialectal (Egyptian) Arabic corpora than in Modern Standard Arabic corpora. Next, our diachronic analysis suggests that the implicit gender bias of Arabic news text steadily increases over time. Finally, we show that the bias effects of the whole corpus can be predicted from bias effects of its subcorpora. We hope that AraWEAT, our framework for multidimensional analysis of stereotypical bias in Arabic text representations, fuels more research on bias in Arabic language technology.

Acknowledgments

Anne Lauscher and Goran Glavaš are supported by the Eliteprogramm of the Baden-Württemberg Stiftung (AGREE grant). We would like to thank the anonymous reviewers for their helpful comments.

References