Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

Yang Trista Cao, Anna Sotnikova, Hal Daumé, Rachel Rudinger, Linda Zou

Introduction

Stereotypes are abstract and over-generalized pictures in people’s minds that capture attributes about groups of people in the complex social world (Lippmann, 1965). They influence people’s thoughts and behaviors, and allow people to make predictions beyond their personal experience or information given (Bruner et al., 1957; Wheeler and Petty, 2001). Stereotypes are also entwined with the production of prejudice, discrimination, and in-group favoritism (Stangor, 2014; Jackson, 2011). A long line of research in social psychology has established models of generic dimensions that estimate people’s stereotypes of social groups (Koch et al., 2016; Fiske et al., 2002, i.a.). We build on the Agency Beliefs Communion (ABC) model, which measures stereotypes toward a social group with respect to 16 traits in three dimensions: Agency/Socioeconomic Success, Conservative–Progressive Beliefs, and Communion (§ 2); an analysis of the group “man” across 32 traits (16 opposing dyads) is shown in Figure 1.

Pre-trained language models (LMs) encode correlations between social groups and traits, like associating the group “Muslim” with the trait threatening, or “man” with confident (e.g., Bender et al., 2021; Nozza et al., 2021; Hovy and Yang, 2021). We conduct a systematic study of social stereotypes in contextualized English masked LMs, grounded in group-trait associations from the ABC model. To capture the group-trait associations in the LM, we first assess two previously proposed word association tests and also propose a new measurement: the sensitivity test (SeT) (§ 3).

To evaluate the degree to which two LMs—BERT (Devlin et al., 2019) and RoBERTa (Liu et al., 2019)—align with human stereotype judgments, we design a human study for collecting group-trait judgments (§ 4). We show that our measure, SeT, best aligns with human judgements on group-trait associations and find that, in general, the association from language models have moderate alignment with human judgements.

Finally, with the best-aligned association measurement, we extend the ABC approach to study LM stereotypes on intersectional groups (§ 5.2). Due largely to the difficulty of extending current approaches for measuring stereotypes in LMs to large numbers of groups, most current approaches only study isolated groups, despite the fact that people’s social identities are multifaceted (Ghavami and Peplau, 2013). Because our approach is generalizable to unstudied groups, we take a step towards exploring stereotypes of intersectional identities, finding some correspondence between model behavior and the literature on intersectional stereotypes.

Background and Related Work

People’s impressions of the world and the actions they take are guided by their stereotypes. To systematize this observation, the field of social psychology has proposed models of stereotypes, including traits that can coordinate social behaviors to serve as fundamental dimensions of stereotyping. Some models are designed to focus on social evaluation towards individual persons (Abele and Wojciszke, 2014), ingroup members (Ellemers, 2017; Yzerbyt, 2018), or a small set of outgroups (Fiske et al., 2002); the Agency Beliefs Communion (ABC) model—whose traits are designed to distinguish groups—is suited for a larger set of U.S. social groups (Abele et al., 2020). The ABC model takes a data-driven strategy to select a set of traits by eliminating those that are less effective in capturing stereotypes. The list contains 1616 pairs, where each pair represents two polarities (see Table 1), categorized into three dimensions: agency/socioeconomic success, conservative-progressive beliefs, and communion/warmth.

Ours is far from the first work to assess stereotypes in language models, and has both advantages and disadvantages compared to previous approaches (see Table 2). Past work has generally taken one of two approaches. The first approach tests systems with hand-constructed templates like “The [group] is □\square”, where [group] ranges over social groups (e.g., “woman” or “Hispanic”), and □\square represents a “masked word” and ranges over occupations (“a professor” or “a nurse”) (e.g., Bolukbasi et al., 2016; May et al., 2019) or associations drawn from implicit association tests (IAT) (e.g., pleasant/unpleasant words or career/family-related words) (e.g., Caliskan et al., 2017; Guo and Caliskan, 2021). In Table 2 we refer to these as “unnatural” prompts. The second approach collects more natural sentences containing stereotypes, either by web crawling with crowdworkers annotations for social bias (Sap et al., 2019) or by having crowdworkers directly write stereotyping sentences (Nangia et al., 2020; Nadeem et al., 2020).

In our work, we take the first approach with traits from the ABC model, using prompts. The advantage of this approach is that the templates and the traits are completely controlled and are easy to extend to other social groups. The second approach is harder to control, which also leads to significant annotation challenges (Blodgett et al., 2021). Using natural sentences limits generalizability, as it requires a unique collection of prompts (and embedded traits) for each social group; in contrast, the prompt-based approach easily generalizes to any plausible group, especially when based on a theoretically grounded framework like ABC or IAT.

An advantage of our work is that the ABC traits are more exhaustive in stereotype coverage with verification from social psychological experiments. The ABC model covers three dimensions with 16 traits, which are consensual, spontaneous, and have been tested using expansive range of social groups (Koch et al., 2021). They used a carefully designed data-driven approach to gather people’s fundamental dimensions of social perceptions with as little sampling bias as possible. Thus the resulted 16 traits cover most stereotypes.

Nevertheless, the main trade-off of our approach is that the testing data are not as natural and specific as other approaches. Although we carefully pick and adjust the templates and the form of the social group terms so that the testing sentences are grammatically correct, they are likely not representative of sentences seen in the real world or in the training data of the language models. Further, while our approach has the benefit of near-exhaustive coverage of potential stereotypes, this comes at a cost: the traits we consider are much more high level (e.g., “repellent”) than more fine-grained stereotypes collected by other means (e.g., the angry Black woman stereotype Collins (2002))—this approach therefore trades coverage for specificity.

Measuring Stereotypes in LMs

Our goal is to measure stereotypes in (masked) LMs, and compare them to stereotypes elicited from people. Both the code and the dataset, along with a datasheet (Gebru et al., 2018), are available under a MIT licence at: https://github.com/TristaCao/U.S_Stereotypes. In § 4 we describe our approach for eliciting human judgments of group-trait affinities; here we describe how we measure these in LMs. Previous work has proposed various ways to measure word associations in LMs, including increased log probability score (ILPS) and contextualized embedding association test (CEAT), both of which we summarize below. Finally, we present a new measurement which we call the Sensitivity Test (SeT), which adapts concepts from active learning to the task of measuring a LM’s associations.

quantifies word associations in language models through masked word probabilities. It calculates the association score with a pre-defined template, “[Group] are □\square.” (Kurita et al., 2019), where □\square is a masked token. For example, given a group “Asian” and a trait smart, P(“Asian”,smart)P(\text{``}Asian\text{''},\texttt{smart}) measures the probability of smart given “Asians” by filling in the template. Since this probability is affected by the prior probability of smart, ILPS normalizes this probability by the “prior” probability of the trait given a masked group, as below:

Intuitively, ILPS measures how much each group raises the likelihood of a trait filling in the template. One can easily show that this equivalent to the weight of evidence of the trait in favor of the hypothesis that the group is the target: s(g,t)=woe(g:t ∣ template)s(\texttt{g},\texttt{t})=\textrm{woe}(\texttt{g}:\texttt{t}~{}|~{}\textrm{template}) (Wod, 1985).

Contextualized Embedding Association Test (CEAT)

estimates word associations with word embedding distances (Guo and Caliskan, 2021). Intuitively, CEAT measures whether some groups are closer to certain traits in a latent vector space. CEAT is a function of A,B,X,YA,B,X,Y. Given two sets of target words defining groups X,YX,Y (e.g. Xmale={X_{\textrm{male}}=\{“man”, “father”, …}\}, Yfemale={Y_{\textrm{female}}=\{“woman”, “mother”, …}\}) and two sets of polar traits A,BA,B (e.g. Apleasant={A_{\textrm{pleasant}}=\{ love, peace, …}\}, Bpleasant={B_{\textrm{pleasant}}=\{ evil, nasty, … }\}), CEAT computes the effect sizes of the difference between XX and YY being closer to AA than BB and corresponding p-values. Since contextualized word representations are affected by the contexts around the word, for each word in the four word sets, CEAT randomly samples 10001000 sentences from Reddit, in which the word appears, and uses these to approximate the true effect size as below:

In our setting, since we care about social bias among multiple groups rather than the difference between two groups, we modify the CEAT to calculate the effect size of the distance difference between g with AA and BB for each group as below:

Sensitivity Test (SeT)

is a new approach we propose to measure word association for social bias in language models, inspired by ideas from active learning Beygelzimer et al. (2008). The intuition of SeT is that even though a model assigns the same probability to two different words, the robustness of those two probabilities may be different. For example, both p(\texttt{competent}|\textit{``Blind people are {\square}.''}) and p(\texttt{kind}|\textit{``Men are {\square}''}) might be low. However, the language model may well not have seen many examples with blind people, as opposed to the presumably very large number of examples of men. In this case, a small number of examples may be sufficient to alter the model’s predictions about blind people, while a larger number would be required for men. SeT captures the model’s confidence in a prediction by measuring how much the model weights would have to change in order to change that prediction. Specifically, SeT computes the minimal change to the last-layer of the language model so that a given trait becomes the highest probability trait (over the full vocabulary).

for a fixed margin γ>0\gamma>0, which we set to 11. SeT returns the negative distance as measure of the association between the corresponding group and trait, normalized by a prior akin to ILPS. This optimization problem does not (to our knowledge) admit a closed form solution; we solve it iteratively using the column squishing algorithm (Bittorf et al., 2012; Daumé and Kumar, 2017).

2 Implementation details

We test the above measurements on both BERT and RoBERTa pretrained large models from an open-source HuggingFacehttps://huggingface.co/models library.

Table 3 lists all the individual social groups we cover in this work. We manually construct the list by combining and picking groups from the list of social groups from Sotnikova et al. (2021) and Koch et al. (2016) and also adding social groups we think are stereotyped in U.S. culture.

Traits.

We use the 32 adjectives of the 16 traits from the ABC model (Table 1). For each traits, we calculate the score of its left-side adjective from its right-side adjective: Spowerless-powerful(g)=S(g,powerful)−S(g,powerless)S_{\texttt{powerless-powerful}}(\texttt{g})=S(\texttt{g},\texttt{powerful})-S(\texttt{g},\texttt{powerless}), where SS is one of the scores from § 3.1.In preliminary experiments, when calculating the score for each adjective, we considered including 1-3 additional adjectives by averaging their scores to improve robustness and mitigate ambiguity. The full list is in Appendix Table A7. However, we found that this did not improve correlations, so we reverted to using the 32 adjectives from the ABC model.

Templates.

ILPS and SeT both require templates in calculating scores. We thus carefully construct a list of templates (Table 4) that covers multiple grammatical and semantic variations, inspired by work investigating harmful search automatic suggestions (Hazen et al., 2020). We find that different model structure requires different templates in order to bring up stereotypes that correlate with human data. See § 5 for evidence.

Subwords.

Due to the nature of BERT and RoBERTa’s tokenizers, some of the adjectives are divided into multiple subwords. This is problematic because all the measurements compute their scores at token level. Neither ILPS nor CEAT deals with subwords directly: in their released implementations, they either take the first or the last sub-token of the word. To remedy this, we adjust the ILPS measurement (denoted as ILPS⋆) to properly compute the probability of traits in context using the chain rule across subwords. For SeT, we calculate the sensitivity score for each subword individually and take the maximum SeT score as the SeT score for the word, which effectively computes a lower-bound on how much the model parameters would need to change. We did not modify CEAT’s measurement as it is not clear what is the best way to compute comparable word embeddings for words that consist of multiple subwords.

Human Study

In the previous section, we describe how we compute associations between groups and traits in language models.Approved by our institutional IRB, #1724519-1. In this section, we assess stereotypes of social groups through groups-trait association, like in Figure 1. We adopt this approach because it is widely used to evaluate group stereotypes in social psychology field (Fiske et al., 2002; Koch et al., 2016). It also aligns with Lippmann (1965)’s theory of stereotypes that they are abstract pictures in people’s head. We broadly follow procedures from previous social psychology papers to collect human evaluation on social groups.

We recruit participants from Prolifichttps://www.prolific.co/. Each participant is paid \2.00torateto rate5socialgroupsonsocial groups on16pairsoftraitsandonaverageparticipantsspendaboutpairs of traits and on average participants spend about10minutesonthesurvey.Thisresultsinapayofminutes on the survey. This results in a pay of\12.0012.00 per hour. Maryland’s current minimum wage is \12.20https://www.minimum−wage.org/maryland.First,participantsreadtheconsentform,andiftheyagreetoparticipateinthestudy,theyseethesurvey’sinstructions.Foreachsocialgroup,participantsread"AsviewedbyAmericansociety,(whilemyownopinionsmaydiffer),how[e.g.,powerless,dominant,poor]versus[e.g.,powerful,dominated,wealthy]are<group>?"Theythenrateeachtraitwitha−https://www.minimum-wage.org/maryland. First, participants read the consent form, and if they agree to participate in the study, they see the survey’s instructions. For each social group, participants read "As viewed by American society, (while my own opinions may differ), how [e.g., powerless, dominant, poor] versus [e.g., powerful, dominated, wealthy] are <group>?" They then rate each trait with a -100$ slider scale where two sides are the two dimensions of the trait (e.g. powerless and powerful). Each annotated group is shown on a separate page, and participants cannot go back to previous pages. To avoid social-desirability bias, we explicitly write in the instruction that “we are not interested in your personal beliefs, but rather how you think people in America view these groups.”

Participant Demographics.

At the end of the survey we collect participants’ demographic information, including gender, race, age, education level, type of living area, etc. Our participants represent 2626 states, with 63.3%63.3\% from California, New York, Texas, or Florida; the gender breakdown is 48.2%48.2\% male, 49.6%49.6\% female, and 2.2%2.2\% genderqueer, agender, or questioning; and skew young, with over 96%96\% at most 4040 years old; and with racial demographics that approximately match the U.S. census. For more details on demographics, see Appendix E.

Quality Assurance.

Ensuring annotation quality in a highly subjective task is a challenge, and common approaches in NLP like having questions where we “know” the answer as tests, measuring interannotator agreement, and calibrating reviewers against each other (Paun et al., 2018) do not make sense here. Yet, it is still important to ensure the annotation quality. After much iteration, we include three test questions, and warn the participants at the beginning that there are test questions.

After the first group, participants must name the group they just scored.

After the second, participants must list one trait they just marked high and one marked low.

The fifth (final) group is a repetition of one of the four groups they previously scored.

We discard annotations with incorrect answers to either of the first two questions. For the third test, we compute intra-annotator (self) agreement and discard annotations with accuracy-to-self lower than 80%80\%. For each group we collect 2020 annotations that pass our quality threshold. In total, we collected annotations from 247247 participants, with 133133 passing the quality tests (suggesting that having such tests is important). The 114114 annotations that did not pass tests were excluded from our dataset, but all 247247 participants were paid.

Social groups and traits.

The social groups we used for the human study are highlighted in Table 3. This table contains only single groups used for the model § 3 and human experiments. We collect annotations for 2525 social groups within 55 domains, across all 1616 pairs of traits.

Results

In this section we present results on correlations between human and model stereotypes for individual groups, comparing across different measurements, including our proposed measurement, SeT (§ 5.1). Next, we analyze how model scores change for intersectional social groups. We consider several possible factors that may influence the score changes such as identity order, some domain domination, and consider emergent traits (§ 5.2).

Before we answer the question of how language model stereotype scores align with human stereotypes across the measurements introduced in § 3, we first run a pilot experiment to select the best template(s) for each measurement-model pair from the set of templates in Table 4 (except for CEAT, which does not require templates). We randomly picked four social groups (Asian, Black, Hispanic, immigrant) and five annotations from each group for the pilot. Since our goal is to inspect the alignment between human and model stereotypes, we take the averaged score of the five annotations as “ground truth” and select templates that give the correlation score according to Kendall τ\tau. We limit the selection to at most two templates to avoid overfitting on the pilot data, selected to maximize correlation for each measurement-model pair.

The selected templates and corresponding correlation scores are shown in appendix (Table 5); the score range for weak correlation is 0.100.10 - 0.190.19, moderate 0.200.20 - 0.290.29, and strong 0.300.30 and above Botsch (2011). For a fixed LM, the best templates tend to be similar across all measures: RoBERTa tends to achieve highest correlation with templates like “That [group] is [trait].” while for BERT the preferred templates tend to be “All [group] are [trait].” or “[Group] should be [trait].”

Given the best templates for each measurement-model pair, we measure to what degree language model stereotypes are aligned with human stereotypes with all annotations on 25 social groups. To quantify alignment, we both calculate the Kendall rank correlation coefficient (Kendall’s τ\tau) and the Precision at 3 (P@3). The former indicates the correlation between model and human scores on group-trait associations in terms of the number of swaps required to get the same order. The latter indicates the percentage of the model’s top stereotypes which accord with human’s judgements. For P@3, we also calculate at both the group level and overall with all groups. For each group, we compute its P@3 score by taking the average of the P@3 scores with the top 3 traits (top at one polarity) and the score with the bottom 3 (top at the other polarity) because each trait has two polar adjectives and the group-trait score is calculated with the difference of the two polarities. To calculate the P@3 scores, we binarize the human group-trait scores at a threshold of 5050. The overall P@3 score is the average of the groups’ individual P@3 scores.

The overall scores are in Table 6. We see that in general that RoBERTa contains group-trait associations that are more similar to human judgements than does BERT. Additionally, we see that both ILPS⋆ and SeT have higher P@3 scores than CEAT and ILPS. The RoBERTa model with the SeT measurement approach yields outputs are the most aligned with human’s judgements, with RoBERTa/ILPS⋆ a close second. From its scores, we see that model’s group-trait associations have moderate correlation with human’s judgements. Moreover, in general, two out of the three top ranked group-trait associations from the model agree with human data. See Table A19 for the overall scores of test groups only, where the four pilot groups are excluded, and Appendix B for group level alignment scores.

2 Intersectional Groups in LMs

Intersectionality is a core concept in Black feminism, introduced in the Combahee River Collective Statement in 1977 (1977; 1983), considering the ways in which feminist theory and antiracism need to combine: “Because the intersectional experience is greater than the sum of racism and sexism, any analysis that does not take intersectionality into account cannot sufficiently address the particular manner in which Black women are subordinated.” The concept was applied in law by Crenshaw (1989) to analyze the ways in which U.S. antidiscrimination law fails Black women.

The concept of intersectionality has broadened and, while its boundaries remain contested (e.g., Browne and Misra, 2003), there are a number of core principles that are central (Steinbugler et al., 2006; Zinn and Dill, 1996): (1) social categories and hierarchies are historically contingent, (2) the experience at an intersection is more than the sum of its parts Collins (2002); King (1988), (3) intersections create both oppression and opportunity Bonilla-Silva (1997), (4) individuals may experience both advantage and disadvantage as a result of intersectionality, and (5) these hierarchies impact social structure and social interaction.

Goals and Research Questions.

We aim to understand whether we can measure evidence of intersectional behavior in language models with respect to stereotyping. In particular, we are interested in questions surrounding how language models stereotype people who simultaneously belong to multiple social groups. We will only use the term “intersectionality” when specifically considering cases where (per (3) above) the resulting experience (in this case, stereotyping) is more than the sum of its parts. For example, common U.S. stereotypes for Black women are as “welfare queens” (which may show up as low agency in our traits), while common stereotypes for Black men is as “criminal” (which may show up as low communion) (hooks, 1992; Collins, 2002). To limit our scope, we will only consider pairs of social groups (e.g., cis men), and will refer to the the groups that make up a pair as the component identities (e.g., cis, or men). We aim to answer the following research questions:

When presented with a paired identity, is the language model sensitive to the order in which the component identities appear?

When paired, do certain social categories dominate others in a language model’s predictions?

Can the language model detect stereotypes that belong to an intersectional group (but not to either of the components that make up the pair)?

To answer these questions, we use the SeT measurement with the RoBERTa model (the best performing pair on the single-group experiments) to compute group-trait associations on our paired groups, which are combinations of all the single groups in Table 3. We manually omit the groups that do not logically exist (e.g. “cis non-binary person”, “teenage elderly person”) or are grammatically awkward (e.g. “doctor elderly person”, “immigrant blind person”). Note we include both orders of the single groups in the paired groups when possible (e.g. “Catholic teenager” and “teenage Catholic person”). We then conduct the analysis by computing the correlation between groups’ list of trait scores with Kendall’s τ\tau.

Q1: Identity Order.

Given an paired group with two identities, the language model may not be able to capture both of the identities and may predict stereotypes based only on one of the components. In fact, the average correlation score between a paired group and the most correlated of its components is 0.560.56, which is moderately high. We thus calculate the correlation of trait scores between the paired group and both its first and second component identities (when both orders are possible). In addition, we calculate the correlation of paired groups with reversed identity order (e.g. “Asian teenager” and “teenage Asian person”). The average correlation score between a paired group and its first component is 0.430.43; the correlation score to its second component is 0.460.46, which are quite close. Further, the average correlation score of intersectional groups with reversed identity is 0.690.69, which is moderately high. Taken together, these results indicate that (a) many paired groups have similar group-trait association scores with one of their component identities alone; (b) the order does not matter significantly, but the language model tends to focus slightly more on the second component. The implication of this is that we can expect that the language model may be able to capture intersectional stereotypes.

Q2: Dominant Domains.

Stryker (1980) suggests that people tend to identify themselves with their race/ethnicity identity before other identities, though this is contested and, in some cases, thought to be antithetical to the idea of intersectionality (e.g., Collins, 2002). Prompted by this debate, we ask if there is a hierarchy of the domains that language model picks up on for paired groups. To answer this question, for each identity domain pair, we compute the average correlation score between the paired groups with each of its two component identities, and take the difference of the averaged correlation scores of the two domains. For each domain, we count the domains it dominates (i.e. has score difference ≥0.1\geq 0.1) and is dominated by.

These results show that age and political stance are dominant domains, which is expected as identities within these two domains have strong characteristics that may overwhelm domains they are paired with. On the other end, race and nationality are, generally, dominated domains. It is surprising that the race domain is majorly dominated, contrasting documented literature in human behavior. The full results are shown in Appendix Table A8 as well as detailed scores Table A9.

Q3: Emergent Intersectional Stereotypes.

Finally, we look into emergent stereotypes of paired groups, with the goal of finding intersectional behavior in the language model. To detect intersectional stereotypes, we need to operationalize the notion of the whole being greater than its parts. For a fixed paired group g=(g1,g2)\texttt{g}=(\texttt{g}_{1},\texttt{g}_{2}) (e.g., “trans Democrats”), and a given trait t (e.g., warm), we compute S(g,t)−max⁡{S(g1,t),S(g2,t)}S(\texttt{g},\texttt{t})-\max\{S(\texttt{g}_{1},\texttt{t}),S(\texttt{g}_{2},\texttt{t})\}, where SS is the score from the language model, capturing whether this trait is more associated with the paired group than the maximum of its association with the component identities. (We consider also the reverse, where we look for scores much less than the min.) We might hope to find some well attested intersectional identities from the literature, such as “Black women” have an attitude (low communion) and “White men” are privileged (high agency) (Ghavami and Peplau, 2013).

The top 5050 emergent group-trait associations according to our measure are listed in Table A10. We also see some good examples are: the language model scores “Hispanic unemployed people” as more egotistic than people of the component identities, “Democrat teenagers” as more altruistic, “male doctors” as more benevolent, etc. However, there are also some unexpected patterns; for instance, almost all nationality identities combined with “mechanic” are trustworthy and likeable, and almost all nationality identities combined with “autistic” are egotistic. Looking into the scores themselves, we find that both “mechanic” and “autistic” have low scores on the corresponding traits, and combining them with nationalities raises to about average levels.

Aside from analyzing face validity—which is mixed—we compare the results of our model to the traits that Ghavami and Peplau (2013) found when conducting human studies of race/gender pairs. To do this, we categorize the traits from Ghavami and Peplau (2013) to the ABC dimensionsGhavami and Peplau (2013) covers paired groups combined with race domain and binary genders. The traits they raised span the agency and communion dimensions. and compare with our full list of emergent group-trait associations. Taking their group-trait matches as ground truth, our detection of traits for these race/gender intersectional groups achieves a precision 0.830.83 and recall 0.650.65—better than random guessing (precision 0.720.72, recall 0.500.50) but far from perfect.

Limitations and Ethical Considerations

There are several limitations to our work, which should be taken into account in the interpretation of our results.

First, our results are likely affected by reporting bias and by a defaulting effect where, when people annotate traits for “men”, they may actually have in their head “cis straight white men”, because the defaults go unremarked. This goes both for the human scores (how does a participant conceptualize “men”?) and language model scores (what do sentences containing the word “man” assume given that most language a langauge model has been trained on likely exhibits defaulting?).

Second, our work only focus on assessing stereotypes within language models and not in any deployed system. Though stereotypes from language models may impact the outputs of downstream systems which are built upon these language models, it is not clear how exactly the stereotypes transfer (Cao et al., 2022). Additionally, our work is limited to English and U.S. social stereotypes.

Third, although we followed and built on best practices from social psychology in developing the human study, it nevertheless has some shortcomings. In particular, even after many iterations on wording, it was difficult to phrase the survey questions to encourage people to reporting their true impressions. There is tension between asking a participant what they think—which risks a counfounding potential social desirability bias (Latkin et al., 2017) (people’s tendency to respond in socially acceptable ways)—and asking what they think others think—which led to comments from a few participants that they felt unqualified to speak for others. Asking these questions of participants and collecting the data also raises the possibility of this work inadvertantly reinforcing stereotypes.

Finally, aggregating human judgements into a single number by averaging (or any other statistic) to compare to model predictions risks collapsing a significant amount of information down to a single number. This number cannot distinguish between a weakly held but common stereotype and a strongly held but rare one. Nor can it distinguish between traits where half of annotators say 0 and the other half say 100, from traits where all annotators say 50. These average judgments should be interpreted as not what any single person would say, but an average over people. This limitation is exacerbated by the defaulting effect, where some people may imagine a different prototype for a given group, and other people may imagine another.

Conclusion

In this paper, we measured language model (LM) stereotypes by adopting the ABC stereotype model from social psychology. Comparing to previous work on detecting LM stereotypes, our approach is easy to extend to previously unconsidered groups, grounded in traits proven effective by social psychology, and exhaustively covering the space of possible stereotypes, at the cost of being more abstract than in other NLP work. This yields a different set of trade-offs than previous approaches to measuring stereotypes in LMs.

With the ABC model and data regarding human stereotypes from our human study, we assessed LM stereotypes using three different association measurements, including SeT, a metric we proposed. We showed that LM group-trait stereotypes in general have moderate correlation with human judgements, and that SeT provides correlations that better align with human’s. Based on these results, we extended our analysis to intersectional groups. We found that the LM may be able to capture intersectional stereotypes but is not particularly good on identifying emergent intersectional stereotypes. Our results also show that that, in general, age and political stance are dominant domains in language models, whereas race and nationality are dominated domains. We hope that our work provides insights for future works on measuring and mitigating stereotypes in natural language processing systems, and that the grounding in theories from social psychology has benefits beyond just studying stereotypes.

Acknowledgments

This material is based upon work partially supported by the National Science Foundation under Grant No. 2131508. The authors are also grateful to all the reviewers who have provided helpful suggestions to improve this work, and thank members of the CLIP lab at the University of Maryland for the support on this project. We are grateful to all those who participated in our human study, without whom this research would not have been possible.

References

Appendix A Traits

The full list of traits and respective adjectives is in the Table A7

Appendix B Experiment Results with Single Groups

Table A11 presents the Kendall’s τ\tau correlation scores between model and human at group level, while Table A12 and Table A13 shows the alignment with the precision at 3 scores (former computed with the top 3 traits and latter with the bottom 3 traits).

Appendix C Experiment Results of Intersectional Groups

Table A8 presents the dominating relationship between domains, while Table A9 lists the average correlation scores of the paired group with each of its identities’ domain for each domain pairs.

Table A10 shows the top 5050 emergent group-trait associations.

Appendix D Human study setup

The survey for the collection of associated traits is presented in Figure A2.

Appendix E Annotators demographics

55.4%55.4\% are white, with 50.6%50.6\% male annotators, 40.440.4 female annotators and no annotators who provided another gender. 15.1%15.1\% of annotators are Black, and 25.6%25.6\% are Hispanic with slightly more female annotators 56.4%56.4\%. We provide four tables A14, A15, A16, A17 showing how perceptions of White people, Black people, White men, and White women are different from each other across annotator demographics. We see variations between in-group and out-group annotations. For instance, women see themselves as more powerful than men see women. While overall scores for men and women groups are similar across white and Black annotators. In Table A18, we show correlation scores for all social groups and overall score between the model and Black, white, white female, and white male annotators.