What are the biases in my word embedding?

Nathaniel Swinger, Maria De-Arteaga, Neil Thomas Heffernan, Mark DM Leiserson, Adam Tauman Kalai

Introduction

Bias in data representation is an important element of fairness in Artificially Intelligent systems (Barocas et al., 2017; Caliskan, Bryson, and Narayanan, 2017; Zemel et al., 2013; Dwork et al., 2012). We consider the problem of Unsupervised Bias Enumeration (UBE): discovering biases automatically from an unlabeled data representation. There are multiple reasons why such an algorithm is useful. First, social scientists can use it as a tool to study human bias, as data analysis is increasingly common in social studies of human biases (Garg et al., 2018; Kozlowski, Taddy, and Evans, 2018). Second, finding bias is a natural step in “debiasing” representations (Bolukbasi et al., 2016). Finally, it can help in avoiding systems that perpetuate these biases: problematic biases can raise red flags for engineers, who can choose to not use a representation or watch out for certain biases in downstream applications, while little or no bias can be a useful green light indicating that a representation is usable. While deciding which biases are problematic is ultimately application specific, UBE may be useful in a “fair ML” pipeline.

We design a UBE algorithm for word embeddings, which are commonly used representations of tokens (e.g. words and phrases) that have been found to contain harmful bias (Bolukbasi et al., 2016). Researchers linking these biases to human biases proposed the Word Embedding Association Test (WEAT) (Caliskan, Bryson, and Narayanan, 2017). The WEAT draws its inspiration from the Implicit Association Test (IAT), a widely-used approach to measure human bias (Greenwald, McGhee, and Schwartz, 1998). An IAT T=(X1,A1,X2,A2)\mathcal{T}=(X_{1},A_{1},X_{2},A_{2}) compares two sets of target tokens X1X_{1} and X2X_{2}, such as female vs. male names, and a pair of opposing sets of attribute tokens A1A_{1} and A2A_{2}, such as workplace vs. family-themed words. Average differences in a person’s response times when asked to link tokens that have anti-stereotypical vs. stereotypical relationships have been shown to indicate the strength of association between concepts. Analogously, the WEAT uses vector similarity across pairs of tokens in the sets to measure association strength. As in the case of the IAT, the inputs for a WEAT are sets of tokens T\mathcal{T} predefined by researchers.

Our UBE algorithm takes as input a word embedding and a list of target tokens, and outputs numerous tests T1,T2,…,\mathcal{T}_{1},\mathcal{T}_{2},\ldots, that are found to be statistically significant by a method we introduce for bounding false discovery rates. A crowdsourcing study of tests generated on three publicly-available word embeddings and a list of names from the Social Security Administration confirms that the biases enumerated are largely consistent with human stereotypes. The generated tests capture racial, gender, religious, and age biases, among others. Table 1 shows the name/word associations output by our algorithm that were rated most offensive by crowd workers.

Creating such tests automatically has several advantages. First, it is not feasible to manually author all possible tests of interest. Domain experts normally create such tests, and it is unreasonable to expect them to cover all possible groups, especially if they do not know which groups are represented in their data. For example, a domain expert based on the United States may not think of testing for caste discrimination, hence biases that an embedding may have against certain Indian last names may go unnoticed. Finally, if a word embedding reveals no biases, this is evidence for lack of bias. We test this by running our UBE algorithm on the supposedly debiased embedding of Bolukbasi et al. (2016).

Our approach for UBE leverages two geometric properties of word embeddings, which we call the parallel and cluster properties. The well-known parallel property indicates that differences between two similar token pairs, such as Mary−-John and Queen−-King, are often nearly parallel vectors. This suggests that among tokens in a similar topic or category, those parallel to name differences may represent biases, as was found by Bolukbasi et al. (2016) and Caliskan, Bryson, and Narayanan (2017). The cluster property, which we were previously unaware of, indicates that the (normalized) vectors of names and words cluster into semantically meaningful groups. For names, the clusters capture social structures such as gender, religion, and others. For words, clusters of words include word categories on topics such as food, education, occupations, and sports. We use these properties to design a UBE algorithm that outputs WEATs.

Technical challenges arise around any procedure for enumerating biases. First, the combinatorial explosion of comparisons among multiple groups parallels issues in human IAT studies as aptly described by Bluemke and Friese (2008): “The evaluation of multiple target concepts such as social groups within a multi-ethnic nation (e.g. White vs. Asian Americans, White vs. African Americans, African vs. Asian Americans; Devos and Banaji, 2005) requires numerous pairwise comparisons for a complete picture”. We alleviate this problem, paralleling that work on human IATs, by generalizing the WEAT to nn groups for arbitrary nn. The second problem, for any UBE algorithm, is determining statistical significance to account for multiple hypothesis testing. To do this, we introduce a novel rotational null hypothesis specific to word embeddings. Third, we provide a human evaluation of the biases, contending with the difficulty that many people are unfamiliar with some groups of names.

Beyond word embeddings and IATs, related work in other subjects is worth mention. First, a body of work studies fairness properties of classification and regression algorithms (e.g. Dwork et al., 2012; Kearns et al., 2017). While our work does not concern supervised learning, it is within this work that we find one of our main motivations–the importance of accounting for intersectionality when studying algorithmic biases. In particular, Buolamwini and Gebru (2018) demonstrate accuracy disparities in image classification highlighting the fact that the magnitude of biases against an intersectional group may go unnoticed when only evaluating for each protected feature independently. Finally, while a significant portion of the empirical research on algorithmic fairness has focused on the societal biases that are most pressing in the countries where the majority of researchers currently conducting the work are based, the literature also contains examples of biases that may be of particular importance in other parts of the world (Shankar et al., 2017; Hoque et al., 2017). UBE can aspire to be useful in multiple contexts, and enable the discovery of biases in a way that relies less on enumeration by domain experts.

Definitions

We assume that there is a given set of possible targets X\mathcal{X} and attributes A\mathcal{A}. Henceforth, since in our evaluation all targets are names and all attributes are lower-case words (or phrases), we refer to targets as names and attributes as words. Nonetheless, in principle, the algorithm can be run on any sets of target and attribute tokens. Caliskan, Bryson, and Narayanan (2017) define a WEAT statistic for two equal-sized groups of names X1,X2⊆XX_{1},X_{2}\subseteq\mathcal{X} and words A1,A2⊆AA_{1},A_{2}\subseteq{\mathcal{A}} which can be conveniently written in our notation as,

In studies of human biases, the combinatorial explosion in groups can be avoided by teasing apart Single-Category IATs which assess associations one group at a time (e.g. Karpinski and Steinman, 2006; Penke, Eichstaedt, and Asendorpf, 2006; Bluemke and Friese, 2008). In word embeddings, we define a simple generalization for n≥1n\geq 1, nonempty groups X1,…,XnX_{1},\ldots,X_{n} of arbitrary sizes and words A1,…,AnA_{1},\ldots,A_{n}, as follows:

Note that gg is symmetric with respect to ordering and weights groups equally regardless of size. The definition differs for n=1n=1, otherwise g≡0g\equiv 0.

The following three properties motivate this as a “natural” generalization of WEAT to one or more groups.

For any X1,X2⊆XX_{1},X_{2}\subseteq\mathcal{X} of equal sizes ∣X1∣=∣X2∣|X_{1}|=|X_{2}| and any nonempty A1,A2⊆AA_{1},A_{2}\subseteq\mathcal{A},

For any nonempty sets X⊂XX\subset\mathcal{X}, A⊂AA\subset\mathcal{A}, let their complements sets Xc=X∖XX^{c}=\mathcal{X}\setminus X and Ac=A∖AA^{c}=\mathcal{A}\setminus A. Then,

For any n>1n>1 and nonempty X1,X2,…,Xn⊆XX_{1},X_{2},\ldots,X_{n}\subseteq\mathcal{X} and A1,A2,…,An⊆A‾A_{1},A_{2},\ldots,A_{n}\subseteq\overline{\bm{\mathcal{A}}},

Lemma 1 explains why we call it a generalization: for n=2n=2 and equal-sized name sets, the values are proportional with a factor that only depends on the set size. More generally, gg can accommodate unequal set sizes and n≠2n\neq 2.

Lemma 2 shows that for n=1n=1 group, the definition is proportional the WEAT with the two groups XX vs. all names X\mathcal{X} and words AA vs. A\mathcal{A}. Equivalently, it is proportional to the WEAT between XX and AA and their compliments.

Finally, Lemma 3 gives a decomposition of a WEAT into n2n^{2} single-group WEATs g(Xi,Aj)g(X_{i},A_{j}). In particular, the value of a single multi-group WEAT reflects a combination of the nn association strengths between XiX_{i} and AiA_{i}, and n2n^{2} disassociation strengths between XiX_{i} and AjA_{j}. As discussed on the literature on IATs, a large effect could reflect a strong association between X1X_{1} and A1A_{1} or X2X_{2} and A2A_{2}, a strong disassociation between X1X_{1} and A2A_{2} or X2X_{2} and A1A_{1}, or some combination of these factors. Proofs are deferred to Appendix B.

Unsupervised Bias Enumeration algorithm

The inputs to our UBE algorithm are shown in Table 2. The output is mm WEATs, each with nn groups with associated sets of words and statistical confidences (p-values) in $.EachWEAThaswordsfromasinglecategory,butseveralofthe. Each WEAT has words from a single category, but several of them$ WEATs may yield no significant associations.

At a high level, the algorithm follows a simple structure. It selects nn disjoint groups of names X1,…,Xn⊂XX_{1},\ldots,X_{n}\subset\mathcal{X}, and mm disjoint categories of lower-case words A1,…,Am\mathcal{A}_{1},\ldots,\mathcal{A}_{m}. All WEATs share the same nn name groups, and each WEAT has words from a single category Aj\mathcal{A}_{j}, with tt words associated to each XiX_{i}. Thus the WEATs can be conveniently visualized in a tabular structure.

For convenience, we normalize all word embedding vectors to be unit length. Note that we only compute cosines between them, and the cosine is simply the inner product for unit vectors. We now detail the algorithm’s steps.

We begin with a set of namesWhile the set of names is an input to our system, they could also be extracted from the embedding itself. X\mathcal{X}, e.g., frequent first names from a database. Since word embeddings do not differentiate between words that have the same spelling but different meanings, we first “clean” the given names to remove names such as “May” and “Virginia”, whose embeddings are more reflective of other uses, such as a month or verb and a US state. Our cleaning procedure, detailed in Appendix C, is similar to that of Caliskan, Bryson, and Narayanan (2017).

We then use K-means++ clustering (from scikit-learn, Pedregosa et al., 2011, with default parameters) to cluster the normalized word vectors of the names, yielding groups X1∪…∪Xn=XX_{1}\cup\ldots\cup X_{n}=\mathcal{X}. Finally, we define μ=∑iX‾i/n\mu=\sum_{i}\overline{\bm{X}}_{i}/n.

2 Step 2: Defining word categories

To define categories, we cluster the most frequent MM lower-case tokens in the word embedding into mm clusters using K-means++, yielding clusters of categories A1,…,Am\mathcal{A}_{1},\ldots,\mathcal{A}_{m}. The constant MM is chosen to cover as many recognizable words as possible without introducing too many unrecognizable tokens. As we shall see, categories capture concepts such as occupations, food-related words, and so forth.

A test Tj=(X1,A1j,…,Xn,Anj)\mathcal{T}_{j}=(X_{1},A_{1j},\ldots,X_{n},A_{nj}) is chosen with disjoint Aij⊂AjA_{ij}\subset\mathcal{A}_{j}, each of size t=∣Aij∣t=|A_{ij}|. To ensure disjointness,If multiplicities are desired, the Voronoi sets VijV_{ij} could be omitted, optimizing Aij⊂AjA_{ij}\subset\mathcal{A}_{j} directly. Aj\mathcal{A}_{j} is first partitioned into nn “Voronoi” sets Vij⊆AjV_{ij}\subseteq\mathcal{A}_{j} consisting of the words whose embedding is closest to each corresponding center X‾i\overline{\bm{X}}_{i}, i.e.,

It then outputs AijA_{ij} defined as the tt words maximizing the following:

The more computationally-demanding step is to compute, using Monte Carlo sampling, the nn p-values for Tj\mathcal{T}_{j}, as described next.

4 Step 4: Computing p-values and ordering

To test whether the associations we find are larger than one would find if there was no relationship between the names XiX_{i} and words A\mathcal{A}, we consider the following “rotational null hypothesis”: the words in the embedding are generated through some process in which the alignment between names and words is random. This is formalized by imagining that a random rotation was applied (multiplying by a uniformly Haar random orthogonal matrix UU) to the word embeddings but not to the name embeddings.

Furthermore, since the algorithm outputs many (hundreds) of name/word biases, the Benjamini-Hochberg (1995) procedure is used to determine a critical p-value that guarantees an α\alpha bound on the rate of false discoveries. Finally, to choose an output ordering on significant tests, the mm tests are then sorted by the total scores σij\sigma_{ij} over the pairs determined significant.

Evaluation

To illustrate the performance of the proposed system in discovering associations, we use a database of first names provided by the Social Security Administration (SSA), which contains number of births per year by sex (F/M) (Administration, 2018). Preprocessing details are in Appendix C.

We use three publicly available word embeddings, each with d=300d=300 dimensions and millions of words: w2v, released in 2013 and trained on approximately 100 billion words from Google News (Mikolov et al., 2013), fast, trained on 600 billion words from the Web (Mikolov et al., 2018), and glove, also trained on the Web using the GloVe algorithm (Pennington, Socher, and Manning, 2014).

While it is possible to display the three words in each AijA_{ij}, the hundreds or thousands of names in each XiX_{i} cannot be displayed in the output of the algorithm. Instead, we use a simple greedy heuristic to give five “illustrative” names for each group, which are displayed in the tables in this paper and in our crowdsourcing experiments. The k+1stk+1^{\text{st}} name shown is chosen, given the first kk names, so as to maximize the average similarity of the first k+1k+1 names to that of the entire set XiX_{i}. Hence, the first name is the one whose normalized vector is most central (closest to the cluster mean), the second name is the one which when averaged with the first is as central as possible, and so forth.

The WEATs can be evaluated in terms of the quality of the name groups and also their associations with words. A priori, it was not clear whether clustering name embeddings would yield any name groups or word categories of interest. For all three embeddings we find that the clustering captures latent groups defined in terms of race, age, and gender (we only have binary gender statistics), as illustrated in Table 3 for n=12n=12 clusters. While even a few clusters suffice to capture some demographic differences, more clusters yield much more fine-grained distinctions. For example, with n=12n=12 one cluster is of evidently Israeli names (see column I of table 3), which one might not consider predefining a priori since they are a small minority in the U.S. Table 6 in the Appendix shows demographic composition of clustering for other embeddings. Note that, although we do not have religious statistics for the names, several of the words in the generated associations are religious in nature, suggesting religious biases as well.

Table 7 in the appendix shows the biases found in the “debiased” w2v embedding of Bolukbasi et al. (2016). While the name clusters still exhibit strong binary gender differences, many fewer statistically significant associations were generated for the most gender-polarized clusters.

We solicited ratings on the biases generated by the algorithm from US-based crowd workers on Amazon’s Mechanical Turkhttp://mturk.com platform. The aim is to identify whether the biases found by our UBE algorithm are consistent with (problematic) biases held by society at large. To this end, we asked about society’s stereotypes, not personal beliefs.

We evaluated the top 12 WEATs generated by our UBE algorithm for the three embeddings, considering n=12n=12 first name groups. Our approach was simple: after familiarizing participants with the 12 groups, we showed the (statistically significant) words and name groups of a WEAT and asked them to identify which words would, stereotypically, be most associated with which names group. A bonus was given for ratings that agreed with most other worker’s ratings, incentivizing workers to provide answers that they felt corresponded to widely held stereotypes.

This design was chosen over a simpler one in which WEATs are shown to individuals who are asked whether or not these are stereotypical. The latter design might support confirmation bias as people may interpret words in such a way that confirms whatever stereotypes they are being asked about. For instance, someone may be able to justify associating the color red with almost any group, a posteriori.

Note that the task presented to the workers involved fine-grained distinctions: for each of the top-12 WEATs, at least 18 workers would each be asked to match the significant c≤12c\leq 12 word triples to the cc name groups (each identified by five names each). For example, workers faced the triple of “registered nurse, homemaker, chairwoman” with c=8c=8 groups of names, half of which were majority female, and the most commonly chosen group matched the one generated: “Janice, Jeanette, Lenna, Mattie, Marylynn.” Across the top-12 WEATs over the three embeddings, the mean number of choices cc was 8.1, yet the most commonly chosen group (plurality) agreed with the generated group 65% of the time (see Table 4). This is significantly more than one would expect from chance. The top-12 WEATs generated for w2v are shown in Table 5.

One challenge faced in this process was that, in pilot experiments, a significant fraction of the workers were not familiar with many of the names. To address this challenge, we first administered a qualification exam (common in crowdsourcing) in which each worker was shown 36 random names, 3 from each group, and was offered a bonus for each name they could correctly identify the group from which it was chosen. Only workers whose accuracy was greater than 1/2 (which happened 37% of the time) evaluated the WEATs. Accuracy greater than 50% on a 12-way classification indicates that the groups of names were meaningful and interpretable to many workers.

Finally, we asked 13-15 workers to rate associations on a scale of 1-7 of political incorrectness, with 7 being “politically incorrect, possibly very offensive” and 1 being “politically correct, inoffensive, or just random.” Only those biases for which the most commonly chosen group matched the association identified by the UBE algorithm were included in this experiment. The mean ratings are shown in Table 4 and the terms present in associations deemed most offensive are presented in Table 1.

2 Potential Indirect Biases and Proxies

Naively, one may think that removing names from a dataset will remove all problematic associations. However, as suggested by Bolukbasi et al. (2016), indirect biases are likely to remain. For example, consider the w2v word embedding, in which hostess is closer to volleyball than to cornerback, while cab driver is closer to cornerback than to volleyball. These associations, taken from columns F1 and F11 of Table 5, might serve as a proxy for gender and/or race. For instance, if someone is applying for a job and their profile includes college sports words, such associations encoded in the embedding may lead to racial or gender biases in cases in which there is no professional basis for these associations. In contrast, volunteer being closer to volunteers than recruits may represent a definitional similarity more than a proxy, if we consider proxies to be associations that mainly have predictive power due to their correlation with a protected attribute. While defining proxies is beyond the scope of this work, we do say that Aij,Ai′j,Aij′,Ai′j′A_{ij},A_{i^{\prime}j},A_{ij^{\prime}},A_{i^{\prime}j^{\prime}} is a potential indirect bias if,

One way to interpret this definition is that if the embedding were to match the pair of word sets {Aij,Ai′j}\{A_{ij},A_{i^{\prime}j}\} to the pair of word sets {Aij′,Ai′j′}\{A_{ij^{\prime}},A_{i^{\prime}j^{\prime}}\}, it would align with the way in which they were generated. For example, does the embedding predict that hostess-cab driver better fits volleyball-cornerback or cornerback-volleyball (but this question is asked with sets of t=3t=3 words)? Downstream, this would mean that a replacing a the word cornerback with volleyball on a profile would make it closer to hostess than cab driver

We consider all possible fourtuples of significant associations, such that 1≤i<i′≤n1\leq i<i^{\prime}\leq n and 1≤j<j′≤m1\leq j<j^{\prime}\leq m. In the case of w2v, 99%99\% of 2,713 significant fourtuples lead to potential indirect biases according to eq. (1). This statistic is of 98% of 1,125 fourtuples and 97% of 1,796 fourtuples for the fast and glove embeddings, respectively. Hence, while names allow us to capture biases in the embedding, removing names is unlikely to be sufficient to debias the embedding.

Limitations

Absent clusters show the limitations of our approach and data. For example, even for large nn, no clusters represent demographically significant Asian-American groups. However, if instead of names we use surnames (U.S. Census, Comenetz, 2016), a cluster “Yu, Tamashiro, Heng, Feng, Nakamura, +393” emerges, which is largely Asian according to Census data (see Table 8 in the Appendix). This distinction may reflect naming practices among Asian Americans (Wu, 1999). Similarly, our approach may miss biases against small minorities or other groups whose names are not significantly differentiated. For example, it is not immediately clear to what extent this methodology can capture biases against individuals whose gender identity is non-binary, although interestingly terms associated with transgender individuals were generated and rated as significant and consistent with human biases.

Conclusions and Discussion

We introduce the problem of Unsupervised Bias Enumeration. We propose and evaluate a UBE algorithm that outputs Word Embedding Association Tests. Unlike humans, where implicit tests are necessary to elicit socially unacceptable biases in a straightforward fashion, word embeddings can be directly probed to output hundreds of biases of varying natures, including numerous offensive and socially unacceptable biases.

The racist and sexist associations exposed in publicly available word embeddings raise questions about their widespread use. An important open question is how to reduce these biases.

Acknowledgments. We are grateful to Tarleton Gillespie, other colleagues and the anonymous reviewers for useful feedback.

References

Appendix A Offensive Stereotypes and Derogatory Terms

The authors consulted with colleagues whether to display the offensive terms and stereotypes that emerged from the embedding using our algorithms. First, regarding derogatory terms, people we consulted found the explicit inclusion of some of these terms offensive. We are also sensitive to the fact that, even in investigating them, we are ourselves using them. The terms we bleep-censor in the tables include slurs regarding race, homosexuality, transgender, and mental ability (Bianchi, 2014). In particular, these include three variants on “the n word” (Asim, 2008), shemale, faggot, twink, mentally retarded, and rednecks. It is not obvious that such slurs would be generated given common naming conventions. Nonetheless, many of these terms were in groups of words that matched stereotypes indicated by crowd workers.

Of course, the associations of words and groups are also offensive, but unfortunately, it is impossible to convey the nature of these associations without presenting the words in the tables associated with the groups. In an attempt to soften the effect, we use group letters rather than illustrative names or summary statistics in our tables. While this decreases the transparency, it gives the reader a choice about whether or not to examine the associated names. Some colleagues were taken aback by an initial draft, in which names and associations were displayed in the same table, and it was noted that it that may be especially offensive to individuals whose name appeared on top of a column of offensive stereotypes. For the names, we restrict our selection of names to those that had at least 1,000 occurrences in the data so that the name would not be uniquely identified with any individual.

In addition, we considered withholding the entire tables and merely presenting the rating statistics. However, we decided that, given that our concern in the analysis is uncovering that such troubling associations are being made by these tools, it was important to be clear and unflinching about what we found, and not risk obscuring the very phenomenon in our explanation.

Appendix B Proofs of Lemmas

For n=2n=2, using our X‾\overline{\bm{X}} notation and their assumption ∣X1∣=∣X2∣|X_{1}|=|X_{2}|, simple algebra shows that,

Since μ=(X‾1+X‾2)/2\bm{\mu}=(\overline{\bm{X}}_{1}+\overline{\bm{X}}_{2})/2, we have that X‾1−μ=(X1‾−X‾2)/2=−(X‾2−μ)\overline{\bm{X}}_{1}-\bm{\mu}=(\overline{\bm{X_{1}}}-\overline{\bm{X}}_{2})/2=-(\overline{\bm{X}}_{2}-\bm{\mu}), and:

which when combined with the previous equality establishes the first equation in Lemma 1. ∎

Since we have shown that (X‾1−X‾2)⋅(A‾1−A‾2)=2g(X1,A1,X2,A2)(\overline{\bm{X}}_{1}-\overline{\bm{X}}_{2})\cdot(\overline{\bm{A}}_{1}-\overline{\bm{A}}_{2})=2g(X_{1},A_{1},X_{2},A_{2}) above, we immediately have that g(X,A)=2g(X,A,X,A)g(X,A)=2g(X,A,\mathcal{X},\mathcal{A}). Moreover, simple algebra shows that g(X,A,X,A)g(X,A,\mathcal{X},\mathcal{A}) and g(X,A,Xc,Ac)g(X,A,X^{c},A^{c}) are proportional because X‾−X‾=∣Xc∣∣X∣(X‾−Xc‾)\overline{\bm{X}}-\overline{\bm{\mathcal{X}}}=\frac{|X^{c}|}{|\mathcal{X}|}(\overline{\bm{X}}-\overline{\bm{X^{c}}}) and similarly A‾−A‾=∣Ac∣∣A∣(A‾−Ac‾)\overline{\bm{A}}-\overline{\bm{\mathcal{A}}}=\frac{|A^{c}|}{|\mathcal{A}|}(\overline{\bm{A}}-\overline{\bm{A^{c}}}). ∎

Follows simply from the definition of gg and μ\mu for n≥2n\geq 2 and n=1n=1. ∎

Appendix C Preprocessing names and words

The SSA dataset (Administration, 2018) has partial coverage for earlier years and includes all names with at least 5 births, we use only years 1938-2017 and select only the names that appeared at least 1,000 times, which cover more than 99% of the data by population. From this data, we extract the fraction of female and male births for each name as well as the mean year of birth. Of course, we select only the names appearing in the embedding.

Note that the mean of the fraction of females among our names is significantly greater than 50%, even though the US population is nearly balanced in binary gender demographics. The subtle reason is there is greater variability in female names in the data, whereas the most common names are more often male. That is, the data have fewer predominantly male first names in total with more people being given those names on average. Since we are including each name only once, this increases the female representation in the population.We performed similar experiments on a sample of names drawn according to the population and, while the names are gender balanced, the clusters exhibit less diversity and most often simply are split by gender and age – one can even have an entire cluster solely consisting of people named Michael.

C.2 Preprocessing last names from U.S. Census

A dataset of last names is made publicly available by the Census Bureau of the United States and contains last names occurring at least 100 times in the 2010 census (Comenetz, 2016), broken down by percentage of race, including White, Black, Hispanic, Asian and Pacific Islander, and Native American. Again we filter for names that appear at least 1,000 times and apply the binary classification procedure described in Section 3.1 to clean the data.

C.3 “Cleaning” names

Caliskan, Bryson, and Narayanan (2017) apply a simple procedure in which they remove the 20% of words whose mean similarity to the other names is smallest. We apply a similar but slightly more sophisticated procedure by training an linear Support Vector Machine (scikit-learn’s LinearSVC, Pedregosa et al., 2011, with default parameters) to distinguish the input names from an equal number of non-names chosen randomly from the most frequent 50,000 words in the embedding. We then remove the 20% of names with smallest margin in the direction identified by the linear classifier.

Figure 1 illustrates the effect of cleaning the last names and shows that the names that tend to be removed are those that violate Zipf’s law.

C.4 Preprocessing words

To identify the most frequent MM words in the embedding, we first restrict to tokens that consist only of the 26 lower-case English letters or spaces for embeddings that contain phrases. We also omit lower-case tokens when the upper-case version of the token is more frequent. For instance, the lower-case token “john” is removed because “John” is more frequent.

Appendix D Biases in different lists/embeddings

Table 6 shows the names from other embeddings. Table 7 shows the biases found in the “debiased” w2v embedding of Bolukbasi et al. (2016), while Table 8 show last-name biases generated from the w2v embeddings.