Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu
Introduction
We live in a multicultural world, where the diversity of cultures enriches our global community. In light of the global deployment of LMs, it is crucial to ensure these models grasp the cultural distinctions of diverse communities. Despite progress to bridge the language barrier gap Ahuja et al. (2023); Yong et al. (2022), LMs still struggle at capturing cultural nuances and adapting to specific cultural contexts Hershcovich et al. (2022). Truly multicultural LMs should not only communicate across languages but do so with an awareness of cultural sensitivities, fostering a deeper global connection.
As we show in Figure 1, LMs fail at appropriate cultural adaptation in Arabic when asked to provide completions to various prompts, often suggesting and prioritizing Western-centric content. For example, LMs refer to alcoholic beverages even when the prompt in Arabic explicitly mentions Islamic prayer. While “going for a drink” in Western culture commonly refers to the consumption of alcoholic beverages, conversely, in the predominantly Muslim Arab world where alcohol is not prevalent, the same phrase in everyday life often refers to the consumption of coffee or tea. Western-centric entities are also generated by LMs when suggesting people’s names and food dishes, despite being inappropriate to the cultural context of the prompts. Such observations raise concerns, as users may find it upsetting to see inadequate cultural representation by LMs in their own languages. This leads to the question: do LMs exhibit bias towards Western entities in non-English, non-Western languages ?
While considerable effort has gone into exploring biases in LMs with regards to groups of different demographic or social dimensions Sheng et al. (2021) such as religion Abid et al. (2021a, b), race An et al. (2023); Ahn and Oh (2021), and nationalities Cao et al. (2022b), much less work (§2) has examined the cultural appropriateness of LMs in the non-Western and non-English environments. In order to address this gap, we center our study on culturally relevant entities, as they are important aspects of cultural heritage Montanari (2006); Tajuddin (2018) and can symbolize regional identities Gómez-Bantel (2018). To the best of our knowledge, there is no resource readily available for doing so, especially one that can contrast Arab vs. Western cultural differences. We thus construct a new benchmark, CAMeL (Cultural Appropriateness Measure Set for LMs), which consists of an extensive list of 20,368 Arab and Western entities extracted from Wikidata and CommonCrawl, covering eight entity types (i.e., person names, food dishes, beverages, clothing items, locations, authors, religious places of worship, and sports clubs), and an associated set of 628 naturally occurring prompts as contexts for those entities (§3).
We show that CAMeL entities and prompts enable cross-cultural testing of LMs in versatile experimental setups, including story generation, NER, sentiment analysis, and text infilling (§4). We benchmark 12 LMs pre-trained with Arabic data (§4.1). Our results reveal concerning cases of cultural stereotypes in LM-generated stories, such as the association of Arab names with poverty/traditionalism (§4.2), and cultural unfairness, such as better NER tagging performance of Western entities and higher association of Arab entities with negative sentiment (§4.3). We further show that LMs exhibit high levels of preference towards Western-associated entities even when prompted by contexts uniquely suited for Arab culture-associated entities (§4.4).
Lastly, we discuss that the prevalence of Western content in Arabic corpora may be a key contributor to the observed biases in LMs. We analyze the cultural relevance of 6 Arabic pre-training corpora by training n-gram LMs on each corpus and comparing their text-infilling performance on CAMeL. We find that sources such as Wikipedia may not be ideal for building culturally-aware LMs (§5).
Related Work
There have been several recent efforts on examining the cultural alignment of LMs. One line of work explored the moral knowledge (e.g., judgment of right and wrong actions) encoded in LMs Fraser et al. (2022); Schramowski et al. (2022); Hämmerl et al. (2022); Xu et al. (2023), probing their ability to infer moral variation on topics with cultural divergence of opinions Ramezani and Xu (2023). It has been found that LMs can be biased towards the moral values of certain societies (e.g., American Johnson et al. (2022)) and political ideologies (e.g., liberalism Abdulhai et al. (2023)). Similar works studied LMs’ understanding of cross-cultural differences in values and beliefs (e.g., attitude towards individualism) Cao et al. (2023); Arora et al. (2023), and what opinions they hold on political Hartmann et al. (2023); Feng et al. (2023) or other global topics Santurkar et al. (2023); Durmus et al. (2023).
These past studies have quantified the alignment of LMs through their responses to cultural surveys Hofstede (1984); Haerpfer et al. (2021); Graham et al. (2011); Guerra and Giner-Sorolla (2010), where LMs were probed using survey type of questions in a QA setting (e.g., ‘Is sex before marriage acceptable in China?’), or cloze-style questions reformulated from these surveys (e.g., ‘In China, sex before marriage is [acceptable/unacceptable]’). Wang et al. (2023b) and Masoud et al. (2023) have shown that LMs reflect values and opinions aligned with Western culture when probed with such surveys, which persists across multiple languages.
Another line of work explored how well LMs store culture-related commonsense knowledge by probing for their ability to answer geo-diverse facts (e.g., ‘The color of the bridal dress in China is [red/white]’) Nguyen et al. (2023); Yin et al. (2022); Keleg and Magdy (2023). Other studies probe LMs for cultural norms such as culinary customs Palta and Rudinger (2023) and time expressions Shwartz (2022). Huang and Yang (2023) studied social norm reasoning as an entailment classification task.
Different from existing work, we study how LMs behave with entities that exhibit cultural variation (e.g., people names, food dishes, etc.). We extract and annotate an extensive list of cultural entities from Wikidata and CommonCrawl, which in turn enables the evaluation of LMs using naturally-occurring prompts that we collect from social media, instead of the artificial prompts used in survey-based studies. Our dataset provides a foundation for measuring biases in various setups, including stereotype examination in LM-generated content, fairness evaluation on NER and sentiment analysis tasks, and text-infilling tests (§ 4), that complement the existing literature. We refer readers to our background section in Appendix A, and the excellent survey of Gallegos et al. (2023), for information on other bias-related issues studied in the past.
Construction of CAMeL
We describe the construction process of CAMeL, starting by collecting entities that exhibit cultural variation. We then obtain prompts from Twitter/X data as natural contexts for these entities, which enable various testing setups for measuring cultural biases in LMs (see examples in Figure 2).
We consider eight types of culturally-relevant entities that include both proper nouns and common nouns: person names, food dishes, beverages, clothing items, locations (cities), literary authors, religious places of worship, and sports clubs. To obtain a comprehensive set of these culturally diverse entities, beyond ones found in the typical lists on the web or generated by LMs when prompted to list them, we first derive entities from the Wikidata knowledge base Vrandečić and Krötzsch (2014) then perform pattern-based entity extraction from the CommonCrawl corpus. Extracted results are manually filtered and annotated to ensure quality.
For each entity type, we manually identified relevant Wikidata classes under which common entities are grouped in the knowledge base (e.g., "food", "city", "drink", etc.). We then extract all entities registered under those classes that have a label in Arabic language. For Location, Authors, and Sports Club entities, it was possible to extract all entities per each country of the Arab world or the Western world (Western Europe and North America), as they are linked to either a country of origin or a nationality label in the knowledge base. However, for other entity types, we had to manually classify them into Arab and Western lists due to the lack of such demographic labels (see Appendix B.1 for details). Wikidata’s coverage of entities in Arabic was extensive for locations, sports clubs, and authors (see Figure 3), but more limited for the other entity types.
To expand on entities collected from Wikidata for entity types where coverage was limited, we perform pattern-based entity extraction on the Arabic subset of the CommonCrawl corpus. Pattern-matching is a simple yet effective method Chiticariu et al. (2013); Freitag et al. (2022); and importantly, it avoids using any LMs in the construction of the dataset that will be used for evaluating LMs. For each entity type, we manually design 5 to 10 generic patterns composed of nouns or noun-verb expressions typically followed by a specific entity. For example, the pattern \setcodeutf-8"\<شقيقة تدعى>" (sister named) is likely to be followed by a female name. We used multiple Arabic verb conjugations of the same pattern to reflect number and genderIn Arabic, verbs are conjugated to reflect gender (male or female) and number (singular, dual, or plural) of the subject.. Using such patterns, we perform pattern matching and extract up to two words that appear after a detected pattern. We avoid using more specific and longer patterns to ensure wider coverage of entities (i.e., higher recall lower precision). This process returns between 5k and 10k unique extractions for each entity type, which are then manually filtered and annotated to achieve high precision. We split name and clothing entities into male/female categories to match Arabic’s gendered grammar, without intending to exclude other gender identities Stanczak and Augenstein (2021). More details are in Appendix B.2.
We hired two undergraduate students who are native Arabic speakers and paid them at the rate of \sim$60 minutes per 1k extractions. About 15-20% of entities extracted from CommonCrawl overlap with those in Wikidata. CAMeL covers both frequently encountered and less frequent entities (Figure 4).
2 Collecting Naturally Occurring Prompts
One of our primary objectives is to assess whether LMs can appropriately distinguish between Arab and Western entities when prompted by culturally specific contexts. To achieve this, we create prompts that embed an Arab cultural reference, ensuring they provide contexts uniquely suited for Arab entities. This allows to gauge the LM’s cultural adaptation ability. Additionally, we create prompts with neutral contexts, enabling us to determine the default cultural leanings of LMs. Hence, CAMeL prompts are split across two types: culturally-contextualized prompts (CAMeL-Co) and culturally-agnostic prompts (CAMeL-Ag). Table 1 offers contrasting examples from each.
To ensure we evaluate LMs in scenarios that mirror actual language uses, we construct our prompts from natural contexts that we retrieve from Twitter/X, rather than crowdsourcing prompts Nadeem et al. (2021a); Nangia et al. (2020a). We employ two keyword search strategies to retrieve tweets that reflect an Arab cultural context for each entity category. First, we use 20 randomly sampled Arab entities from our lists as search queries to capture discussions about culturally-relevant entities. We further refine our search using one or two manually-designed patterns of adjective phrases that directly reference an Arab entity (e.g., \setcodeutf-8"\<للكاتب العربي>" (by the Arab author)). We search for tweets over the period of 8/1/2023 to 9/30/2023 to avoid overlap with the data LMs may have been pre-trained on. Retrieved tweets are manually inspected to select ones with suitable Arab cultural contexts. From these, we created 250 CAMeL-Co prompts by replacing the original context entities with a [MASK] token. Similarly, we constructed 378 prompts for CAMeL-Ag using generic patterns as search queries that do not contain any cultural reference (see Appendix C).
To support fairness evaluation of LMs on sentiment analysis, the prompts were labeled by the annotators for positive, negative, or neutral sentiment. The inter-annotator agreement is 0.954 as measured by Cohen’s Kappa. More details and statistics are provided in Appendix C.3.
Measuring Cultural Bias in LMs
Using CAMeL, we measure cultural biases of several monolingual and multilingual LMs (§4.1). First, we analyze stereotypes in LM-generated stories (§4.2). We then examine cross-cultural fairness of LMs on the NER and Sentiment Analysis tasks (§4.3). Finally, we benchmark the capability of LMs on culturally appropriate text-infilling (§4.4).
We consider LMs that have been intentionally trained for Arabic. For monolingual LMs, we use AraBERT Antoun et al. (2020) and ARBERT Abdul-Mageed et al. (2021). We also include MARBERT Abdul-Mageed et al. (2021) and AraBERT-T Antoun et al. (2020), which are trained on Arabic tweets, as well as AraGPT2 Antoun et al. (2021). For multilingual LMs, besides mBERT, XLM-R Conneau et al. (2020), BLOOM Scao et al. (2022), GPT-3.5 and GPT-4, we use Arabic-English bilingual JAIS Sengupta et al. (2023), GigaBERT and GigaBERT-CS Lan et al. (2020), which was further trained on synthetic Arabic-English code-switched data. We use the base (B) and large (L) versions whenever available.
2 Cultural Stereotypes in Story Generation
We examine the potential of GPT-type LMs to reflect stereotypes in their generations when portraying Arab and Western entities. Specifically, we analyze their lexical choices in stories generated about characters with Arab and Western names.
For each of the male and female names in CAMeL, we prompt LMs in Arabic to “Generate a story about a character named [PERSON NAME]”. Then, we analyze the frequency of adjective usage by LMs in the stories featuring Arab or Western names. To do so, we extract all adjectives from stories using the Farasa POS tagger Abdelali et al. (2016) and compute their Odds Ratio (OR) Wan et al. (2023) (see Appendix F.1 for the formula). A large OR indicates more odds for an adjective of appearing in Western stories, while a small OR indicates more odds of appearing in Arab ones. We inspect adjectives with the 50 highest and lowest ORs to identify and categorize adjectives that reflect stereotypes based on the work of Cao et al. (2022a), which outlines descriptive adjectives for stereotypical traits (e.g., poor, likeable, etc.) using the Agency-Belief Communion (ABC) framework Koch et al. (2016).
Figure 5 displays the identified adjectives, revealing multiple stereotypical associations. Stories about Arab characters more often cover a theme of poverty with adjectives such as “poor” persistently used across LMs. On the other hand, the adjective “wealthy” was more likely to appear in Western stories. LMs also tend to use adjectives describing Traditionalism, Dominance (for male names) and Benevolence (for female names) in Arab stories, while using adjectives that reflect Likeability and High-Status in Western stories. We manually inspected stories containing those adjectives, where we found a consistent opening narrative of Arab characters being “born into a poor and modest family”. This was less prevalent Western stories where LMs often portray positive attributes about the character (see examples in Table 2).
3 Fairness in NER and Sentiment Analysis
To examine whether LMs treat Arab and Western entities fairly, we analyze their cross-cultural performance on the tasks of NER and sentiment analysis. We perform this analysis using evaluation sentences that include either Arab or Western entities.
We leverage culturally-contextualized prompts (CAMeL-Co) which have been manually labelled for sentiment (§3.2) to create the test data. Specifically, for each of the prompts, we replace the [MASK] token with 50 randomly sampled Arab and Western entities. This generates two distinct culturally-contrasting evaluation sets (one Arab, one Western) for the sentiment analysis experiment, each comprising around 12k sentences. For NER, we use the subset of 5.7k sentences that contain either person names or locations in the evaluation.
We create models capable of performing Arabic NER and sentiment prediction by fine-tuning BERT-type LMs on datasets commonly used in Arabic NLU benchmarks Elmadany et al. (2023); Abdul-Mageed et al. (2021). We use the ANERCorp Benajiba et al. (2007) dataset for NER (name and location tags were used only) and HARD dataset Elnagar et al. (2018) for sentiment analysis. For GPT-type LMs, we perform in-context learning with 5-shot examples (see prompts in Appendix F.2).
Figure 6 shows the F1 scores achieved by LMs on recognizing Arab and Western related entities. We find that most LMs perform better when tagging Western person names and locations. Larger discrepancies are observed on locations, reaching up to 20 F1 points of difference. The gap was smaller for tagging of male and female names, where differences were around 5 F1 points.
Following past work on fairness of sentiment classifiers Czarnowska et al. (2021), we examine differences in false positive and false negative predictions between sentences containing Arab vs. Western entities. This enables closer analysis of whether LMs show more association of Arab or Western entities with positive or negative sentiments, as opposed to comparing F1 scores which had minimal differences. The results are shown in Figure 7. We observe that nearly all LMs achieve higher false negatives on sentences containing Arab entities, suggesting more false association of Arab entities with negative sentiment. On the other hand, no clear trend of stronger positive sentiment association towards Arab or Western entities is observed.
4 Culturally-Appropriate Text Infilling
To test the ability of LMs at adaptation to cultural contexts, we use a likelihood-based score that compares model preference of Western vs. Arab entities as fillings of [MASK] tokens in CAMeL prompts.
where is the LM’s probability of an entity filling the masked token. We evaluate LMs with BERT-type architecture using the full prompts with a [MASK] token for text-infilling and LMs with GPT-type architecture using only the portion of the prompt appearing before the [MASK]. We take the average over all the sub-words for entities tokenized into sub-words. For a set of prompts , the CBS per entity type for an LM is computed by averaging over all . LMs are considered more Western-biased as its CBS gets closer to 100%.
In addition to using the vanilla prompts, we also experiment with two prompt-adaption techniques that may help in localizing LMs to the relevant Arab culture: (1) Culture Token, where the special token \<[عربي]> ([Arab]) is prepended to prompts, and (2) N-shot demos, where randomly sampled Arab entities are prepended to prompts as demonstrations. We make sure the entity being evaluated is not in the demonstrations.
Figure 8 show the average CBS across all entity types on culturally-contextualized prompts (CAMeL-Co). We provide CBS per each entity type and additional results on CAMeL-Ag in Appendix F.3. We observe the following key findings:
Since CAMeL-Co prompts explicitly refer to Arab culture, an ideal LM is expected to (nearly) always score higher likelihood to Arab entities over Western ones, i.e., with CBS close to 0. However, existing LMs show high average CBS (40-60%), which is on par with their performance on CAMeL-Ag prompts where contexts are neutral. This indicates a struggle in localizing to the appropriate culture in context, and a noticeable preference for Western entities.
Surprisingly, although monolingual LMs are trained on Arabic-only data, they still obtain high CBS scores. The reason may be that part of the pre-training data (more in §5), even if solely in Arabic, often discusses Western topics.
Most multilingual LMs showed a higher CBS compared with monolingual LMs. This implies that multilingual training could impact cultural relevance of LMs in non-Western languages. We find that embeddings of Arab and Western entities are grouped into distinct clusters by monolingual LMs while mixed up in multilingual LMs (see Appendix G.1).
Prompt-adaption techniques can potentially help in localizing LMs to the relevant culture. In particular, prepending Arab demonstrations reduced CBS for most LMs. However, introducing a special culture token had little effect.
Analyzing Arabic Pre-training Data
One main contributor to the observed failures of LMs in appropriate cultural adaptation could be the prevalence of Western content in the Arabic pre-training corpora. To gain more insight, we analyze six Arabic corpora that are commonly used in pre-training LMs, comparing their cultural relevance.
We use two local Arabic news corpora (1.5B corpus by El-Khair (2016)) and Assafir news Antoun et al. (2020)), an international news corpus (OSIAN by Zeroual et al. (2019)), the Arabic portion of CommonCrawl (from OSCAR by Suárez et al. (2019)), Arabic Wikipedia, and the 60M Arabic tweets corpus used in training AraBERT-T Antoun et al. (2020). We train 4-gram LMs using OpenGRM Roark et al. (2012) without smoothing on each corpus, leveraging their frequency count-based nature to directly compare prevalence of cultural contexts and entities across corpora. We then use the trained 4-grams to compute the average CBS for each corpus using CAMeL-Co for analysis.
Figure 9 shows the average CBS of 4-gram LMs trained on each corpus. The results suggest that (Arabic) Wikipedia is the most Western-centric among all corpora, despite being often considered as one of the highest-quality sources for pre-training data. This is mostly because a large portion of Arabic Wikipedia articles discuss Western content. International news had the second highest CBS. Interestingly, web-crawled data was the third most Western-centric source. A recent analysis of CommonCrawl by Thompson et al. (2024) has shown that a large fraction of the total web content is machine-translated. This could explain the prevalence of Western content as it may get translated into Arabic from languages such as English. We also find that an English-like grammatical structure of Arabic sentences can incite more Western bias in LMs (see Appendix G.2). Local news and Twitter/X corpora had the lowest CBS, suggesting that future work may consider these sources for training more culturally adapted LMs.
Conclusion
We introduced CAMeL, a novel dataset of naturally occurring prompts and culturally-relevant entities as prompt completions across eight entity types. We showed that when operating in Arabic, LMs exhibit bias towards Western entities, failing in appropriate cultural adaptation. LMs also show cultural unfairness on tasks such as NER and sentiment analysis, and stereotypes in generated stories. By releasing CAMeL, we hope to enable the evaluation and development of culturally-aware LMs.
Limitations
We focused on assessing the overall ability of LMs to adapt to Arab cultural contexts and exploring their biases towards Western entities. The entities in CAMeL are therefore primarily categorized as being associated with Arab or Western cultures. However, entities belonging to certain categories, such as food dishes or locations, can be further divided into specific regions and countries within the Arab and Western worlds. This finer-grained categorization could enable analysis of LMs’ ability to distinguish between entities belonging to sub-groups of a particular culture. We leave such detailed factual knowledge exploration of sub-cultural distinctions in LMs for future studies.
CAMeL only covers the Arabic language and enables the evaluation of model biases with respect to Western vs. Arab cultural entities. The works of Wang et al. (2023b) and Masoud et al. (2023) have shown that when probed using cultural surveys in Chinese, Korean, or Slovak, LMs tend to respond with answers reflecting Western values. CAMeL can be extended in the future to such languages by adopting our approach for entity extraction and prompt construction.
We limited the scope of our experiment on stereotypes in generated stories to only the analysis of lexical terms, specifically adjectives. Future work can leverage CAMeL entities to analyze further variations beyond lexical content, such as stylistic features of the generations. We believe that the release of CAMeL entities will be a valuable asset to the research community for exploring biases in generation tasks beyond only story generation.
Our analysis of pre-training corpora was limited to examining the relevance of their cultural content, particularly to understand why LMs fail at adapting to Arab cultural contexts. However, to gain deeper insights into the manifested issues of stereotyping and unfairness, more analyses would be necessary. This involves quantifying the co-occurrences of Arab and Western entities with specific themes (e.g., poverty, negativity, etc.) within the corpora. Further, fine-tuning datasets could also play an additional role in amplifying fairness problems. Future research can leverage CAMeL to examine these issues, building on our initial findings.
Ethics Statement
While LMs must adapt to Arab entities when prompts are specifically grounded in an Arab cultural context, the question of what culture they should default to in neutral contexts is more nuanced. This largely depends on the preferences and backgrounds of users. For instance, Arabic speakers residing in non-Arab countries might prefer Arabic LMs to align with the local culture they identify with. However, current LMs default to Western culture in neutral contexts. The neutral prompts we provide in CAMeL-Ag can serve as a valuable test bed for future studies that aim at aligning LMs to meet the unique cultural preferences of their users.
Our prompts were derived from naturally occurring social media contexts obtained from Twitter/X. We do not share the original raw tweets but rather modified versions where original entities mentioned by users have been replaced by [MASK] tokens. The prompts are, therefore, anonymized and do not contain any personally identifiable information. The release of CAMeL prompts is exclusively for research purposes, particularly for evaluating the cultural adaptation of LMs. When constructing our prompts, we have carefully selected contexts that do not contain toxic or offensive language.
Arabic is a grammatically gendered language where verbs must be conjugated for either male or female genders in the second and third persons. This linguistic restriction affects how we construct prompts for categories such as names and clothing, leading us to separate these prompts into male and female groups. This follows the approach taken by past work on social biases in languages with grammatical gender distinctions Levy et al. (2023). It’s important to clarify that this categorization by gender does not aim to define or differentiate gender identities Stanczak and Augenstein (2021) but is done to reflect the language’s structure accurately. We also note that the aim of our study is to investigate biases in LMs toward Western entities and not the examination of gender biases.
Acknowledgements
The author would like to thank Youssef Naous and Nour Allah El Senary for their help in data annotation. The author also thanks Wissam Antoun for sharing data that facilitated our analysis on pre-training corpora. This research is supported in part by the NSF awards IIS-2144493 and IIS-2052498, ODNI and IARPA via the HIATUS program (contract 2022-22072200004). The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of NSF, ODNI, IARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein.
References
Appendix A Additional Background
Various studies have explored biases in English LMs with regards to groups from different cultural backgrounds. For example, Abid et al. (2021a) studied stereotypical associations in LMs towards different religious groups by probing LMs with templates such as “[MASK] are violent”. They show that LMs such as GPT-3 associates Muslims with violence more often than other religious groups, which has been found by Hemmatian et al. (2023) to persist even after LMs go through debiasing procedures. Similar template probing studies have explored such social biases in LMs towards races (e.g., “Asians are good at math”) Ma et al. (2023b, a); Cao et al. (2022b); Ross et al. (2021); Nadeem et al. (2021a), nationalities (e.g., “A person from Iraq is an enemy”) Venkit et al. (2023); Manerba et al. (2023); Ahn and Oh (2021) and more attributes Nangia et al. (2020b).
This line of research has primarily explored the extent to which LMs reflect human biased associations about specific social or cultural groups present in their pre-training data. While they touch on certain aspects related to culture (e.g., religion), they do not study the LMs’s adaptation capability to diverse world cultures. Further, these works are English-centered. In contrast, our work explores how LMs handle entities that associate with different cultures. We show that multilingual and Arabic monolingual LMs exhibit bias towards Western-associated entities, failing at appropriate cultural adaptation to Arab cultural contexts. We also show how LMs demonstrate upsetting stereotypes and unfairness on the NER and sentiment analysis tasks when presented with Arab culture-associated entities as opposed to Western entities.
Various works has explored biases in non-English languages. One line of work translates English datasets into other languages (Levy et al., 2023; Névéol et al., 2022; Lee et al., 2023; Kurpicz-Briki, 2020; Lauscher and Glavaš, 2019). We argue that this is not an effective strategy, as the translated evaluation data lacks the relevant cultural identity (Talat et al., 2022). Most studies focus primarily on gender bias (Das et al., 2023; Vashishtha et al., 2023; Touileb et al., 2022; Kaneko et al., 2022) or social bias (Névéol et al., 2022; Bhatt et al., 2022; BehnamGhader and Milios, 2022; Nozza et al., 2021). In this paper, we study a more subtle and understudied yet very important problem – cultural appropriateness of LMs in non-English and non-Western environments. We focus on culture-specific entities and analyze cross-cultural performance of LMs on such entities. We construct CAMeL, a novel dataset of naturally-occuring Arabic prompts obtained from Twitter/X, and an extensive list of entities associated with Arab and Western culture across eight entity types that exhibit cultural variation.
Early work on measuring biases examined vector space distances between static word embeddings of neutral attributes (e.g., professions) and social attributes (e.g., genders, races) Caliskan et al. (2017); Dev et al. (2021). Embedding-based methods were then adapted to contextualized embeddings of LMs learned from the context of sentences, where neutral and social attributes are placed in sentence templates (e.g., "This is Katie", "This is a friend") May et al. (2019); Guo and Caliskan (2021); Tan and Celis (2019). More recent works adopt probability-based approaches, where LMs are prompted using masked templates and their assigned token probabilities for different groups are compared given the context (Nozza et al., 2022; Kaneko and Bollegala, 2022; Nadeem et al., 2021b; Nozza et al., 2021; Nangia et al., 2020b; Salazar et al., 2020).
In contrast to the aforementioned "intrinsic" approaches that focus on examining embeddings and probabilities, another line of research adopts "extrinsic" approaches, where the focus is analyzing fairness of LMs towards different groups (e.g., races, nationalities, etc.) on downstream tasks Czarnowska et al. (2021). In this setting, groups are slotted inside sentence templates that are used for downstream evaluation, allowing comparison of model behavior when groups are switched. Such approaches have been used to explore gender biases in co-reference resolution Zhao et al. (2018), social biases in sentiment analysis Bhaskaran and Bhallamudi (2019), lexical/dialect biases in toxic language detection Zhou et al. (2021), and other classification tasks Li et al. (2023).
Our dataset enables measurement of cultural biases through both intrinsic and extrinsic approaches (§ 4). CAMeL prompts and entities support fairness evaluation for several tasks including text classification (sentiment analysis § 4.3), token-level classification (NER § 4.3), and text generation (§ 4.2) tasks. CAMeL also supports intrinsic measurements through text infilling tests (§ 4.4).
Appendix B Collecting Arab and Western Entities
We provide additional details of the collected Arab and Western entities for each entity type in CAMeL. For religious places of worship, we focus on the two dominant religions in both cultures and hence collect lists of mosques as Arab entities and churches as Western entities. For sports clubs, we specifically collect football clubs as entities. Statistics of CAMeL entities per source are shown in Table 3.
Given that Arabic is a grammatically gendered language, requiring verbs to be conjugated according to male or female genders in both the second and third persons, it is necessary to categorize both names and clothing entities based on gender. This categorization ensures that such entities align grammatically with the verbs in the prompts we create (§ 3.2), which are conjugated according to gender.
We report the Wikidata classes from which entities were extracted in Table 4. A Wikidata class groups together entities that share common characteristics. For example, entities that are considered a food item such as "spaghetti" or "shawarma" are registered under the "food" class in Wikidata. Wikidata classes can also be linked to sub-classes which cover a more-specific subset of entities. For example, "Street food" and "Dessert" are sub-classes of the "food" class. We selected classes that are generic and cover a large number of sub-classes to ensure wide coverage of entities.
Entities registered in Wikidata may have labels in multiple languages (i.e., their equivalent terms in each of those languages), as they are tied to Wikipedia articles about the entity that may exist in multiple language versions. For example, the Arabic label for the entity "shawarma" is \setcodeutf-8"\<شاورما>". We extract all entities under the selected classes and use their Arabic labels when available. Note that not all Wikidata entities have labels in Arabic.
B.2 Entity Extraction from Web Crawls
We use the Arabic subset of CommonCrawl from OSCAR Suárez et al. (2019), which partitions the CommonCrawl dumps by language. The Arabic patterns designed to extract entities from the corpus are reported in Table 5. We defined multiple versions of the same pattern, where we used different tenses and gender/number conjugations of the same verb, helping expand extractions. Verb conjugations that reflect gender were specifically helpful in collecting male-specific and female-specific entities (such as names and clothing items). Pattern-based extraction significantly boosted the number of entities obtained from Wikidata (e.g; a 171% increase in female name entities from 354 to 961). We do not perform the pattern-based extraction process for authors, locations, and sports clubs, since Wikidata provided an extensive enough coverage for those entity types.
Appendix C Constructing Natural Prompts
The patterns used in our query-based search for retrieving culturally-contextualized tweets are reported in Table 6. The number of tweets returned by pattern-based queries was often larger than searching directly with Arab entities, which depended on entity popularity (popular entities returned more tweets). Most queries returned 100 to 500 tweets. For queries that return a larger number, we randomly sample 500 tweets. 15% of tweets were found suitable contexts. 68.8% of the prompts were in Arabic dialects, while 31.2% were in Modern Standard Arabic. Example prompts for each entity type are shown in Table 8.
For proper evaluation of GPT-type models in text-infilling tests, we provide a version of the prompts where some prompts were slightly re-written to ensure reference to Arab culture appears before the [MASK] token, as the conditional probability of these models relies only on previous tokens.
C.2 CAMeL-Ag: Details
The search patterns used to construct the culturally-agnostic prompts of CAMeL-Ag are reported in Table 7. In this setting, we search for tweets that have neutral contexts; where either Arab or Western entities would be appropriate fillings. Patterns are thus defined to be generic with no specific cultural reference. CAMeL-Ag prompts were obtained from the two-month span of 3/1/2023 to 4/30/2023. For most entity types, we structure the queries in a Pronoun-Verb format to facilitate our analysis on grammatical structure influence (Appendix §G.2).
C.3 Sentiment Annotation
Prompt statistics and sentiment distribution is shown in Table 9. We re-wrote some prompts when possible in the opposite sentiment to obtain balance in sentiments. The small cases of differences in annotation were resolved via discussions between annotators to decide on the final label. For ethical considerations, we do not provide sentiment labels for prompts referring to religious places of worship, and ensure that none of those prompts express negativity.
C.4 Details on Annotators
The annotators were undergraduate student employees who are native Arabic speakers, paid at their normal hourly rate of $18 per hour. The annotators were informed that they were "annotating entities for cultural association and prompts for sentiment as part of a research project to assess cultural biases in language models that have been trained on Arabic data".
Appendix D Language Models Details
The following is a description of the models used:
Antoun et al. (2020): BERT-base model trained on the Arabic Wikipedia Dump, the 1.5B words Arabic corpus El-Khair (2016), the OSCAR corpus Suárez et al. (2019) (a multilingual subset of CommmonCrawl), and articles from Assafir newspaper. We use the basehuggingface.co/aubmindlab/bert-base-arabertv02 and largehuggingface.co/aubmindlab/bert-large-arabertv02 versions of the model without pre-segmentation.
Antoun et al. (2020): a version of AraBERT with continued pre-training on 60M Arabic tweets, available in both basehttps://huggingface.co/aubmindlab/bert-base-arabertv02-twitter and largehttps://huggingface.co/aubmindlab/bert-large-arabertv02-twitter architectures.
Abdul-Mageed et al. (2021): trained on 61GB of text in Modern Standard Arabic (MSA) and only and uses additional pre-training corpora than AraBERT such as public books from Hindawi, the Arabic Gigaword corpus, and the OSIAN corpus. Available in base architecture only.
Abdul-Mageed et al. (2021): a BERT model trained only on 1B Arabic tweets designed to work better on dialects. Available in base architecture only.
Antoun et al. (2021): a monolingual decoder-only model based on the GPT2 architecture. AraGPT2 was trained using the same pre-training corpora as AraBERT. We experiment with the base and large versions of the model.
Devlin et al. (2019): a multilingual version of the BERT model trained solely on Wikipedia and available only in the base architecture.
Conneau et al. (2020): multilingual model trained on CommonCrawl and outperforms mBERT on various cross-lingual benchmarks. Available in both base and large architectures.
Lan et al. (2020): a bilingual English-Arabic BERT model that outperforms other multilingual models in zero-shot transfer from English to Arabic. GigaBERThuggingface.co/lanwuwei/GigaBERT-v3-Arabic-and-English is trained on the Arabic and English Gigaword corpora, Arabic and English Wikipedia, and the OSCAR corpus. We also use a version of GigaBERT, referred to as GigaBERT-CShuggingface.co/lanwuwei/GigaBERT-v4-Arabic-and-English, which is further pre-trained on Code-Switched data. Both models are in the base architecture.
Scao et al. (2022): a 176 billion parameter multilingual LLM trained on 46 natural languages and 13 programming languages. The language-specific training data largely came from the OSCAR corpus Suárez et al. (2019). We used the HuggingFace inference APIhttps://huggingface.co/inference-api to prompt BLOOM which returns token log probabilities when using the details:true parameter.
Sengupta et al. (2023): a 13 billion parameter bilingual LM trained on English and Arabic. The model is available as JAIS and JAIS-Chat where the later is optimized for dialogue.
a 175 billion parameter LM. We experiment with OpenAI’s text-davinci-003 model which has been instruction fine-tuned. The data used to train the GPT-3.5 model has not been publicly disclosed. We retrieve token log probabilities using the OpenAI completions API endpointhttps://platform.openai.com/docs/api-reference/completions with the logprobs:1 parameter.
We experiment with OpenAI’s gpt-4-1106-preview model. Data and technical details of the model have not been publicly released. Given that computing the CBS requires access to a language model’s log probabilities, we could not compute CBS scores for GPT4, for which log probabilities for arbitrary inputs are not obtainable through the OpenAI API.
Appendix E Pre-training Corpora Details
We provide details about the Arabic pre-training corpora analyzed in § 5:
We use the September 2020 dump of Arabic Wikipediahttps://dumps.wikimedia.org/ used in training AraBERT models Antoun et al. (2020).
The Open Source International Arabic News Corpus Zeroual et al. (2019) consists of 3.5M news articles from 31 news sources. Almost half of this dataset (1.5M articles) is obtained from non-local international news sources (un.org, euronews.com, reuters.com, sputniknews.com, mamnewsnetwork.com) or news sources in non-Arab counties such as the UK (bbc.com), USA (cnn.com), Germany (dw.com), in addition to several others. Despite these sources providing news articles written in Arabic, it is highly likely that they contain a larger number of references to Western content.
The 1.5 billion words Arabic Corpus El-Khair (2016) consists of 5M news articles collected from 10 local news sources in 8 Arab countries.
News articles from the Lebanese Assafir newspaperhttps://en.wikipedia.org/wiki/As-Safir used in training AraGPT2 Antoun et al. (2021) and AraBERTv2 Antoun et al. (2020).
The Open Super-large Crawled Almanach coRpus Suárez et al. (2019) is a multilingual partition of CommonCrawlhttps://commoncrawl.org/. We use the Arabic subset of the corpus.
A corpus of 60M Arabic tweets used in training AraBERT-T Antoun et al. (2020).
Appendix F Additional Results
We give a description of the Odds Ratio computed for adjectives in LM-generated stories about Arab and Western characters in § 4.2. We also provide additional results on female names.
Let and be the set of adjectives extracted from stories about characters with Arab and Western names respectively. The Odds Ratio Wan et al. (2023); Szumilas (2010) of an adjective is calculated as the odds of it appearing in stories with Western-named characters over its odds of appearing in stories with Arab-named characters:
where is the count of the adjective in stories with Western-named characters, and is its count in ones with Arab-named characters. A larger Odds Ratio reflects more likelihood for an adjective to appear in stories with Western-named characters, while a smaller ratio reflects higher likelihood of appearing in stories with Arab-named characters.
The identified adjectives with stereotypical traits on stories with Arab and Western female names are shown in Figure 10. We notice the association of Arab-named female characters with Traditionalism and Poverty, similar to what was observed with male names. The adjective "generous" appeared frequently in Arab stories as well, reflecting a Benevolent trait. On the other hands, adjectives that were salient in stories about Western-named characters reflect a Likeable and High-Status trait. However, unlike the case of male characters, adjectives describing a Wealthy trait do not appear frequently for stories with female Western-named characters.
F.2 Fairness in NER and Sentiment Analysis
The performance of all fine-tuned BERT-type models on NER tagging of Arab vs. Western entities is shown in Figure 11. We also report results on recognizing author names, where LMs show better performance on recognizing Western authors compared to Arab authors.
F.2.2 Experimental Details
We used a learning rate of 5e-5 and the AdamW Optimizer. We train models for 5 epochs and set the batch size to 8. Fine-tuning was performed on 1 NVIDIA A100 GPU. Since the HARD Elnagar et al. (2018) dataset for Arabic sentiment analysis is originally imbalanced in terms of sentiment labels, we took a random sample of 30k sentences from the dataset balanced across positive/negative/neutral sentiment for our experiment.
We perform Sentiment Analysis and NER for GPT-type models via in-context learning Brown et al. (2020); Min et al. (2022), where models are prompted with 5 randomly sampled demonstrations (5-shots). In the following, we describe how prompting was performed for each task.
The prompt used to predict sentiment with GPT-type models is shown in Table 14, where the model is given an instruction to classify the sentiment of a test sentence, a key mapping labels to sentiments, and 5-shot demonstrations that we randomly sampled from the HARD dataset Elnagar et al. (2018) for each test sentence.
F.3 Text Infilling
We use the recent approach of Wang et al. (2023a) for NER using GPT models, where models are prompted to mark entities using the special tokens @@ and ## (in the format: @@ [entity] ##). We prompt models with 5 randomly sampled demonstrations from ANERCorp Benajiba et al. (2007), where entities were marked with the special tokens. The prompt used in shown in Table 15. The results are reported in Figure 12 for the three entity types of Names, Location, and Authors. The most noticeable discrepancy is observed in location tagging, where models show superior performance on Western location entities. We found JAIS not to perform well on this task, with an F1 score below 10, and hence do not report its results.
We report the CBS scores achieved by LMs for each entity types on the culturally-contextualized prompts from CAMeL-Co in Table F.2.2.
F.3.2 Results on CAMeL-Ag
We report the CBS scores achieved by the models on culturally-agnostic prompts from CAMeL-Ag in Table F.3.2. We observe similar trends to what is seen in the main results of § 4.4. Without any cultural contextualization, models show high CBS scores across entity types, reaching up to 70-80%. Most multilingual models also show higher CBS than monolingual models.
To compare how LMs encode Arab and Western entities, we compute the contextualized embeddings of 50 randomly sampled entities from each entity type, when placed in prompts from CAMeL-Ag. For entities that get tokenized into multiple tokens, we take the average of their embeddings. To obtain a final encoding for each entity, we average its contextualized embeddings across all prompts.
We visualize entity embeddings by projecting them into a 2-dimensional space using t-SNE (Van der Maaten and Hinton, 2008). The results are shown for BERT-type LMs in Figure 13. It appears that most monolingual models (ARBERT, MARBERT, AraBERT, AraBERT-Twi) separate Arab and Western entities into distinctive clusters. In constrast, such distinction is not observed for most multilingual models, especially for XLM-R and mBERT which are trained a wide variety of languages. On the other hand, distinct clusters can still be recognized for the bilingual GigaBERT models which are trained only on English and Arabic. These observations may indicate that multilingual training with a large variety of languages makes it more challenging for LMs to capture distinctions between entities in a specific language.
To verify these observations, we treat Arab and Western entity embeddings for a particular entity type as two distinct clusters in high dimensional space and measure the cluster quality using the Davies-Bouldin Index (DBI) (Davies and Bouldin, 1979). The DBI measures (1) how close items within the same cluster are and (2) how far apart distinct clusters are. Ideally, a good clustering will have tight internal cluster distances and far separation between clusters. Such clustering achieves a DBI closer to 0. Average DBIs across cultural categories for each model are reported in Table F.3.2. The average DBIs of multilingual models are generally higher than monolingual models, with XLM-R achieving the worst clustering quality, supporting the observations in our visualizations. These findings suggest that as models become more capable at multilingual modeling, they could simultaneously lose the cultural distinctiveness of their representations.
G.2 Does English-like grammatical structure incite more Western bias?
We study the effect of having an English-like grammatical structure of the Arabic prompts on the amplification of bias towards Western entities in LMs. In Arabic, subject pronouns can be and are often dropped, as they can be inferred from verb conjugation. In contrast, subject pronouns are typically necessary to convey the subject of a sentence in English; null subjects are rarely allowed. We test whether an English-like grammatical structure contributes to increased preference of Western entities by dropping all first-person pronouns "\setcodeutf8\<أنا>" (I) in the Arabic prompts, whenever applicable, and recomputing the CBS scores. We use prompts from CAMeL-Ag, which we constructed using search queries defined in a pronoun-verb format to facilitate analysis on dropping subject pronouns. The average CBS achieved by LMs before (English-like) and after dropping pronouns in the prompts are shown in Table 13. Author prompts are omitted from this analysis since they do not include pronouns. Nearly all multilingual LMs show a reduction in average CBS when pronouns are dropped, indicating that Arabic prompts which are more grammatically aligned with English sentence structure incite more preference towards Western entities. Half of the monolingual LMs also show a reduction in CBS. This supports our observations in § 5 that some portions of the Arabic pre-training corpora could be translated from English, introducing irrelevant linguistic elements that can contribute to increased bias towards Western entities.