Attenuating Bias in Word Vectors

Sunipa Dev, Jeff Phillips

BIAS IN WORD VECTORS

Word embeddings are an increasingly popular application of neural networks wherein enormous text corpora are taken as input and words therein are mapped to a vector in some high dimensional space. Two commonly used approaches to implement this are WordToVec and GloVe . These word vector representations estimate similarity between words based on the context of their nearby text, or to predict the likelihood of seeing words in the context of another. Richer properties were discovered such as synonym similarity, linear word relationships, and analogies such as man : woman :: king : queen. Their use is now standard in training complex language models.

However, it has been observed that word embeddings are prone to express the bias inherent in the data it is extracted from . Further, Zhao et al. (2017) and Hendricks et al. (2018) show that machine learning algorithms and their output show more bias than the data they are generated from.

Word vector embeddings as used in machine learning towards applications which significantly affect people’s lives, such as to assess credit , predict crime , and other emerging domains such judging loan applications and resumes for jobs or college applications. So it is paramount that efforts are made to identify and if possible to remove bias inherent in them. Or at least, we should attempt minimize the propagation of bias within them. For instance, in using existing word embeddings, Bolukbasi et al. (2016) demonstrated that women and men are associated with different professions, with men associated with leaderships roles and professions like doctor, programmer and women closer to professions like receptionist or nurse. Caliskan et al. (2017) similarly noted how word embeddings show that women are more closely associated with arts than math while it is the opposite for men. They also showed how positive and negative connotations are associated with European-American versus African-American names.

Our work simplifies, quantifies, and fine-tunes these approaches: we show that very simple linear projection of all words based on vectors captured by common names is an effective and general way to significantly reduce bias in word embeddings. More specifically:

We demonstrate that simple linear projection of all word vectors along a bias direction is more effective than the Hard Debiasing of Bolukbasi et al. (2016) which is more complex and also partially relies on crowd sourcing.

We show that these results can be slightly improved by dampening the projection of words which are far from the projection distance.

We examine the bias inherent in the standard word pairs used for debiasing based on gender by randomly flipping or swapping these words in the raw text before creating the embeddings. We show that this alone does not eliminate bias in word embeddings, corroborating that simple language modification is not as effective as repairing the word embeddings themselves.

We show that common names with gender association (e.g., john, amy) often provides a more effective gender subspace to debias along than using gendered words (e.g., he, she).

We demonstrate that names carry other inherent, and sometimes unfavorable, biases associated with race, nationality, and age, which also corresponds with bias subspaces in word embeddings. And that it is effective to use common names to establish these bias directions and remove this bias from word embeddings.

DATA AND NOTATIONS

We set as default the text corpus of a Wikipedia dump (dumps.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2) with 4.57 billion tokens and we extract a GloVe embedding from it in D=300D=300 dimensions per word. We restrict the word vocabulary to the most frequent 100,000100{,}000 words. We also modify the text corpus and extract embeddings from it as described later.

HOW TO ATTENUATE BIAS

Given such a set E\mathcal{E} of equality sets, the bias vector vBv_{B} can be formed as follows . For each Ej={ej+,ej−}E_{j}=\{e_{j}^{+},e_{j}^{-}\} create a vector e⃗i=ei+−ei−\vec{e}_{i}=e_{i}^{+}-e_{i}^{-} between the pairs. Stack these to form a matrix Q=[e⃗1  e⃗2  …  e⃗m]Q=[\vec{e}_{1}\;\vec{e}_{2}\;\ldots\;\vec{e}_{m}], and let vBv_{B} be the top singular vector of QQ. We revisit how to create such a bias direction in Section 4.

Now given a word vector w∈Ww\in W, we can project it to its component along this bias direction vBv_{B} as

The most notable advance towards debiasing embeddings along the gender direction has been by Bolukbasi et al. (2016) in their algorithm called Hard Debiasing (HDHD). It takes a set of words desired to be neutralized, {w1,w2,…,wn}=WN⊂W\{w_{1},w_{2},\ldots,w_{n}\}=W_{N}\subset W, a unit bias subspace vector vBv_{B}, and a set of equality sets E1,E2,…,EmE_{1},E_{2},\ldots,E_{m}.

First, words {w1,w2,…,wn}∈WN\{w_{1},w_{2},\ldots,w_{n}\}\in W_{N} are projected orthogonal to the bias direction and normalized

Second, it corrects the locations of the vectors in the equality sets. Let μj=1∣E∣∑e∈Eje\mu_{j}=\frac{1}{|E|}\sum_{e\in E_{j}}e be the mean of an equality set, and μ=1m∑j=1mμj\mu=\frac{1}{m}\sum_{j=1}^{m}\mu_{j} be the mean of of equality set means. Let νj=μ−μj\nu_{j}=\mu-\mu_{j} be the offset of a particular equality set from the mean. Now each e∈Eje\in E_{j} in each equality set EjE_{j} is first centered using their average and then neutralized as

Intuitively νj\nu_{j} quantifies the amount words in each equality set EjE_{j} differ from each other in directions apart from the gender direction. This is used to center the words in each of these sets.

This renders word pairs such as man and woman as equidistant from the neutral words wi′w_{i}^{\prime} with each word of the pair being centralized and moved to a position opposite the other in the space. This can filter out properties either word gained by being used in some other context, like mankind or humans for the word man.

The word set WN={w1,w2,…,wn}⊂WW_{N}=\{w_{1},w_{2},\ldots,w_{n}\}\subset W which is debiased is obtained in two steps. First it seeds some words as definitionally gendered via crowd sourcing and using dictionary definitions; the complement – ones not selected in this step – are set as neutral. Next, using this seeding an SVM is trained and used to predict among all WW the set of other biased WBW_{B} or neutral words WNW_{N}. This set WNW_{N} is taken as desired to be neutral and is debiased. Thus not all words WW in the vocabulary are debiased in this procedure, only a select set chosen via crowd-sourcing and definitions, and its extrapolation. Also the word vectors in the equality sets are also handled separately. This makes this approach not a fully automatic way to debias the vector embedding.

2 Alternate and Simple Methods

We next present some simple alternatives to HD which are simple and fully automatic. These all assume a bias direction vBv_{B}.

Subtraction. As a simple baseline, for all word vectors ww subtract the gender direction vBv_{B} from ww:

Linear Projection. A better baseline is to project all words w∈Ww\in W orthogonally to the bias vector vBv_{B}.

This enforces that the updated set W′={w′∣w∈W}W^{\prime}=\{w^{\prime}\mid w\in W\} has no component along vBv_{B}, and hence the resulting span is only D−1D-1 dimensions. Reducing the total dimension from say 300300 to 299299 should have minimal effects of expressiveness or generalizability of the word vector embeddings.

Bolukbasi et al. apply this same step to a dictionary definition based extrapolation and crowd-source-chosen set of word pairs WN⊂WW_{N}\subset W. We quantify in Section 5 that this single universal projection step debiases better than HD.

For example, consider the bias as gender, and the equality set with words man and woman. Linear projection will subtract from their word embeddings the proportion that were along the gender direction vBv_{B} learned from a larger set of equality pairs. It will make them close-by but not exactly equal. The word man is used in many extra senses than the word woman; it is used to refer to humankind, to a person in general, and in expressions like “oh man”. In contrast a simpler word pair with fewer word senses, like (he - she) and (him - her), we can expect them to be almost at identical positions in the vector space after debiasing, implying their synonymity.

Thus, this approach uniformly reduces the component of the word along the bias direction without compromising on the differences that words (and word pairs) have.

3 Partial Projection

A potential issue with the simple approaches is that they can significantly change some embedded words which are definitionally biased (e.g., the neutral words WBW_{B} described by Bolukbasi et al. ). [[We note that this may not *actually* be a problem (see Section 5); the change may only be associated with the bias, so removing it would then not change the meaning of those words in any way except the ones we want to avoid.]] However, these intuitively should be words which have correlation with the bias vector, but also are far in the orthogonal direction. In this section we explore how to automatically attenuate the effect of the projection on these words.

This stems from the observation that given a bias direction, the words which are most extreme in this direction (have the largest dot product) sometimes have a reasonable biased context, but some do not. These “false positives” may be large normed vectors which also happen to have a component in the bias direction.

We start with a bias direction vBv_{B} and mean μ\mu derived from equality pairs (defined the same way as in context of HD). Now given a word vector ww we decompose it into key values along two components, illustrated in Figure 1. First, we write its bias component as

This is the difference of ww from μ\mu when both are projected onto the bias direction vBv_{B}.

Second, we write a (residual) orthogonal component

Let η(w)=∥r(w)∥\eta(w)=\|r(w)\| be its value. It is the orthogonal distance from the bias vector vBv_{B}; recall we chose vBv_{B} to pass through the origin, so the choice of μ\mu does not affect this distance.

Now we will maintain the orthogonal component (r(w)r(w), which is in a subspace spanned by D−1D-1 out of DD dimensions) but adjust the bias component β(w)\beta(w) to make it closer to μ\mu. But the adjustment will depend on the magnitude η(w)\eta(w). As a default we set

so all word vectors retain their orthogonal component, but have a fixed and constant bias term. This is functionally equivalent to the Linear Projection approach; the only difference is that instead of having a magnitude along vBv_{B} (and the orthogonal part unchanged), it instead has a magnitude of constant μ\mu along vBv_{B} (and the orthogonal part still unchanged). This adds a constant to every inner product, and a constant offset to any linear projection or classifier. If we are required to work with normalize vectors (we do not recommend this as the vector length captures veracity information about its embedding), we can simple set w′=r(w)/∥r(w)∥w^{\prime}=r(w)/\|r(w)\|.

Given this set-up, we now propose three modifications. In each set

were fif_{i} for i={1,2,3}i=\{1,2,3\} is a function of only the orthogonal value η(w)\eta(w). For the default case f(η)=0f(\eta)=0

Here σ\sigma is a hyperparameter that controls the importance of η\eta; in Section 3.4 we show that we can just set σ=1\sigma=1. x

In Figure 2 we see the regions of the (η,β)(\eta,\beta)-space that the functions ff, f1f_{1} and f2f_{2} consider gendered. ff projects all points onto the y=μy=\mu line. But variants f1f_{1}, f2f_{2}, and f3f_{3} are represented by curves that dampen the bias reduction to different degrees as η\eta increases. Points P1 and P2 have the same dot products with the bias direction but different dot products along the other D−1D-1 dimensions. We can observe the effects of each dampening function as η\eta increases from P1 to P2.

4 SETTING σ=1𝜎1\sigma=1.

To complete the damping functions f1f_{1}, f2f_{2}, and f3f_{3}, we need a value σ\sigma. If σ\sigma is larger, then more word vectors have bias completely removed; if σ\sigma is smaller, than more words are unaffected by the projection. The goal is that words SS which are likely to carry a bias connotation to have little damping (small fif_{i} values) and words TT which are unlikely to carry a bias connotation to have more damping (large fif_{i} values – they are not moved much).

Given sets SS and TT, we can define a gain function

with a regularization term ρ\rho. The gain γ\gamma is large when most bias words in SS have very little damping (small fif_{i}, large 1−fi1-f_{i}), and the opposite is true for the neutral words in TT. We want the neutral words to have large fif_{i} and hence small 1−fi1-f_{i}, so they do not change much.

To define the gain function, we need sets SS and TT; we do so with the bias of interest as gender. The biased set SS is chosen among a set of 10001000 popular names in WW which (based on babynamewizard.com and and SSN databases ) are strongly associated with a gender. The neutral set TT is chosen as the most frequent 10001000 words from WW, after filtering out obviously gendered words like names man and he. We also omit occupation words like doctor and others which may carry unintentional gender bias (these are words we would like to automatically de-bias). The neutral set may not be perfectly non-gendered, but it provides reasonable approximation of all non-gendered words.

We find for an array of choices for ρ\rho (we tried ρ=1\rho=1, ρ=10\rho=10, and ρ=100\rho=100), the value σ=1\sigma=1 approximately maximizes the gain function γi,ρ(σ)\gamma_{i,\rho}(\sigma) for each i∈{1,2,3}i\in\{1,2,3\}. So for hereafter we fix σ=1\sigma=1.

Although these sets SS and TT play a role somewhat similar to the crowd-sourced sets WBW_{B} and WNW_{N} from HD that we hoped to avoid, the role here is much reduced. This is just to verify that a choice of σ=1\sigma=1 is reasonable, and otherwise they are not used.

5 Flipping the Raw Text

Since the embeddings preserve inner products of the data from which it is drawn, we explore if we can make the data itself gender unbiased and then observe how that change shows up in the embedding. Unbiasing a textual corpus completely can be very intricate and complicated since there are a many (sometimes implicit) gender indicators in text. Nonetheless, we propose a simple way of neutralizing bias in textual data by using word pairs E1,E2,…EmE_{1},E_{2},\ldots E_{m}; in particular, when we observe in raw text on part of a word part, we randomly flip it to the other pair. For instance for gendered word pairs (e.g., (he - she)) in a string “he was a doctor” we may flip to “she was a doctor.”

We implement this procedure over the entire input raw text, and try various probabilities of flipping each observed word, focusing on probabilities 0.50.5, 0.750.75 and 1.001.00. The first 0.50.5-flip probability makes each element of a word pair equally likely. The last 1.001.00-flip probability reverses the roles of those word pairs, and 0.750.75-flip probability does something in between. We perform this set of experiments on the default Wikipedia data set and switch between word pairs (say man →\rightarrow woman, she →\rightarrow he, etc), from a list larger that Table 3 consisting of 75 word pairs; see Supplementary Material D.1.

We observe how the proportion along the principal component changes with this flipping in Figure 3. We see that flipping with 0.50.5 somewhat dampens the difference between the different principal components. On the other hand flipping with probability 1.01.0 (and to a lesser extent 0.750.75) exacerbates the gender components rather than dampening it. Now there are two components significantly larger than the others. This indicates this flipping is only addressing part of the explicit bias, but missing some implicit bias, and these effects are now muddled.

We list some gender biased analogies in the default embedding and how they change with each of the methods described in this section in Table 2.

THE BIAS SUBSPACE

We explore ways of detecting and defining the bias subspace vBv_{B} and recovering the most gendered words in the embedding. Recall as default, we use vBv_{B} as the top singular vector of the matrix defined by stacking vectors e⃗i=ei+−ei−\vec{e}_{i}=e_{i}^{+}-e_{i}^{-} of biased word pairs. We primarily focus on gendered bias, using words in Table 1, and show later how to effectively extend to other biases. We discuss this in detail in Supplementary Material C.

The dot product, ⟨vB,w⟩\langle v_{B},w\rangle of the word vectors ww with the gender subspace vBv_{B} is a good indicator of how gendered a word is. The magnitude of the dot product tells us of the length along the gender subspace and the sign tells us whether it is more female or male. Some of the words denoted as most gendered are listed in Table 3.

1 Bias Direction using Names

When listing gendered words by ∣⟨vB,w⟩∣|\langle v_{B},w\rangle|, we observe that many gendered words are names. This indicates the potential to use names as an alternative (and potentially in a more general way) to bootstrap finding the gender direction.

From the top 100K words, we extract the 10 most common male {m1,m2,…,m10}\{m_{1},m_{2},\ldots,m_{10}\} and female {s1,s2,…,s10}\{s_{1},s_{2},\ldots,s_{10}\} names which are not used in ambiguous ways (e.g., not the name hope which could also refer to the sentiment). We pair these 10 names from each category (male, female) randomly and compute the SVD as before. We observe in 4 that the fractional singular values show a similar pattern as with the list of correctly gendered word pairs like (man - woman), (he - she), etc.

But this way of pairing names is quite imprecise. These names are not ‘opposites’ of each other in the sense that word pairs are. So, we modify how we compute vBv_{B} now so that we can better use names to detect the bias in the embedding. The following method gives us this advantage where we do not necessarily need word pairs or equality sets as in Bolukbasi et al. .

where s=110∑isis=\frac{1}{10}\sum_{i}s_{i} and m=110∑imim=\frac{1}{10}\sum_{i}m_{i}.

Using the default Wikipedia dataset, we found that this is a good approximator of the gender subspace defined by the first right singular vector calculated using gendered words from Table 1; there dot product is 0.8090.809. We find similar large dot product scores for other datasets too.

Here too we collect all the most gendered words as per the gender direction vB,namesv_{B,\textsf{names}} determined by these names. Most gendered words returned are similar as using the default vBv_{B}, like occupational words, adjectives, and synonyms for each gender. We find names to express similar classification of words along male - female vectors with homemaker more female and policeman being more male. We illustrate this in more detail in Table 4.

Using that direction, we debias by linear projection. There is a similar shift in analogy results. We see a few examples in Table 5.

QUANTIFYING BIAS

In this section we develop new measures to quantify how much bias has been removed from an embedding, and evaluate the various techniques we have developed for doing so.

As one measure, we use the Word Embedding Association Test (WEAT) test developed by Caliskan et al. (2017) as analogous to the IAT tests to evaluate the association of male and female gendered words with two categories of target words: career oriented words versus family oriented words. We detail WEAT and list the exact words used (as in ) in Supplementary Material B; smaller values are better.

Bolukbasi et al. evaluated embedding bias use a crowsourced judgement of whether an analogy produced by an embedding is biased or not. Our goal was to avoid crowd sourcing, so we propose two more automatic tests to qualitatively and uniformly evaluate an embedding for the presence of gender bias.

A way to evaluate how the neutralization technique affects the embedding is to evaluate how the nearest neighbors change for (a) gendered pairs of words E\mathcal{E} and (b) indirect-bias-affected words such as those associated with sports or occupational words (e.g., football, captain, doctor). We use the gendered word pairs in Table 1 for E\mathcal{E} and the professions list P={p1,p2,…,pk}P=\{p_{1},p_{2},\ldots,p_{k}\} as proposed and used by Bolukbasi et al. https://github.com/tolga-b/debiaswe (see also Supplementary Material D.2) to represent (b).

We transform these similarity vectors to replace each coordinate by its rank order, and compute the Spearman Coefficient (in $,largerisbetter)betweentherankorderofthesimilaritiestowordsin, larger is better) between the rank order of the similarities to words inP$.

Thus, here, we care about the order in which the words in PP occur as neighbors to each word pair rather than the exact distance. The exact distance between each word pair would depend on the usage of each word and thus on all the different dimensions other than the gender subspace too. But the order being relatively the same, as determined using Spearman Coefficient would indicate the dampening of bias in the gender direction (i.e., if doctor by profession is the 2nd closest of all professions to both man and woman, then the embedding has a dampened bias for the word doctor in the gender direction). Neutralization should ideally bring the Spearman coefficient towards 1.

Embedding Quality Test (EQT).

The demonstration by Bolukbasi et al. about the skewed gender roles in embeddings using analogies is what we try to quantify in this test. We attempt to quantify the improvement in analogies with respect to bias in the embeddings.

We use the same sets E\mathcal{E} and PP as in the ECT test. However, for each profession pi∈Pp_{i}\in P we create a list SiS_{i} of their plurals and synonyms from WordNet on NLTK .

For each word pair {ej+,ej−}=Ej∈E\{e_{j}^{+},e_{j}^{-}\}=E_{j}\in\mathcal{E}, and each occupation word pi∈Pp_{i}\in P, we test if the analogy ej+:ej−::pie_{j}^{+}:e_{j}^{-}::p_{i} returns a word from SiS_{i}. If yes, we set Q(Ej,pi)=1Q(E_{j},p_{i})=1, and Q(Ej,pi)=0Q(E_{j},p_{i})=0 otherwise.

Return the average value across all combinations 1∣E∣1k∑Ej∈E∑pi∈PQ(Ej,pi)\frac{1}{|\mathcal{E}|}\frac{1}{k}\sum_{E_{j}\in\mathcal{E}}\sum_{p_{i}\in P}Q(E_{j},p_{i}).

The scores for EQT are typically much smaller than for ECT. We explain two reasons for this.

First, EQT does not check for if the analogy makes relative sense, biased or otherwise. So, “man : woman :: doctor : nurse” is as wrong as “man : woman :: doctor : chair.” This pushes the score down.

Second, synonyms in each set sis_{i} as returned by WordNet on the Natural Language Toolkit, NLTK do not always contain all possible variants of the word. For example, the words psychiatrist and psychologist can be seen as analogous for our purposes here but linguistically are removed enough that WordNet does not put them as synonyms together. Hence, even after debiasing, if the analogy returns “man : woman :: psychiatrist : psychologist“ S1 returns 0. Further, since the data also has several misspelt words, archeologist is not recognized as a synonym or alternative for the word archaeologist. For this too S1 returns a 0.

The first caveat can be side-stepped by restricting the pool of words we search over for the analogous word to be from list PP. But it is debatable if an embedding should be penalized equally for returning both nurse or chair for the analogy “man : woman :: doctor : ?”

This measures the quality of analogies, with better quality having a score closer to 11.

Evaluating embeddings.

We mainly run 44 methods to evaluate our methods WEAT, EQT, and two variants of ECT: ECT (word pairs) uses E\mathcal{E} defined by words in Table 1 and ECT (names) which uses vectors mm and ss derived by gendered names.

We observe in Table 6 that the ECT score increases for all methods in comparison to the non-debiased (the original) word embedding; the exception is flipping with 1.01.0 probability score for ECT (word pairs) and all flipping variants for ECT (names). Flipping does nothing to affect the names, so it is not surprising that it does not improve this score; further indicating that it is challenging to directly fix bias in raw text before creating embeddings. Moreover, HD has the lowest score (of 0.9170.917) whereas projection obtains scores of 0.9960.996 (with vBv_{B}) and 0.9430.943 (with vB,namesv_{B,\textsf{names}}).

EQT is a more challenging test, and the original embedding only achieves a score of 0.1280.128, and HD only obtains 0.1450.145 (that is 12−15%12-15\% of occupation words have their related word as nearest neighbor). On the other hand, projection increases this percentage to 28.3%28.3\% (using vBv_{B}) and 29.1%29.1\% (using vB,namesv_{B,\textsf{names}}). Even subtraction does nearly as well at between 23−27%23-27\%. Generally, the subtraction always performs slightly worse than projection.

For the WEAT test, the original data has a score of 1.6231.623, and this is decreased the most by all forms of flipping, down to about 1.11.1. HD and projection do about the same with HD obtaining a score of 1.2211.221 and projection obtaining 1.2191.219 (with vB,namesv_{B,\textsf{names}}) and 1.2341.234 (with vBv_{B}); values closer to 0 are better (See Supplementary Material B).

In the bottom of Table 6 we also run these approaches on standard similarity and analogy tests for evaluating the quality of embeddings. We use cosine similarity on WordSimilarity-353 (WSim, 353 word pairs) and SimLex-999 (Simlex, 999 word pairs) , each of which evaluates a Spearman coefficient (larger is better). We also use the Google Analogy Dataset using the function 3COSADD which takes in three words which for a part of the analogy and returns the 4th word which fits the analogy the best.

We observe (as expected) that all that debiasing approaches reduce these scores. The largest decrease in scores (between 1%1\% and 10%10\%) is almost always from HD. Flipping at 0.50.5 rate is comparable to HD. And simple linear projection decreases the least (usually only about 1%1\%, except on analogies where it is 7%7\% (with vBv_{B}) or 5%5\% (with vB,namesv_{B,\textsf{names}}).

In Table 7 we also evaluate the damping mechanisms defined by f1f_{1}, f2f_{2}, and f3f_{3}, using vBv_{B}. These are very comparable to simple linear projection (represented by ff). The scores for ECT, EQT, and WEAT are all about the same as simple linear projection, usually slightly worse.

While ECT, EQT and WEAT scores are in a similar range for all of ff, f1f_{1}, f2f_{2}, and f3f_{3}; the dampened approaches f1f_{1}, f2f_{2}, and f3f_{3} performs better on the Google Analogy test. This test set is devoid of bias and is made up of syntactic and semantic analogies. So, a score closer to that of the original, biased embedding, tells us that more structure has been retained by f1f_{1}, f2f_{2} and f3f_{3}. Overall, any of these approaches could be used if a user wants to debias while retaining as much structure as possible, but otherwise linear projection (or ff) is roughly as good as these dampened approaches.

DETECTING OTHER BIAS USING NAMES

We saw so far how projection combined with finding the gender direction using names works well and works as well as projection combined with finding the gender direction using word pairs.

We explore here a way of extending this approach to detect other kinds of bias where we cannot necessarily find good word pairs to indicate a direction, like Table 1 for gender, but where names are known to belong to certain protected demographic groups. For example, there is a divide between names that different racial groups tend to use more. Caliskan et al. use a list of names that are more African-American (AA) versus names that are more European-American (EA) for their analysis of bias. There are similar lists of names that are distinctly and commonly used by different ethnic, racial (e.g., Asian, African-American) and even religious (for e.g., Islamic) groups.

We first try this with two common demographic group divides : Hispanic / European-American and African-American / European-American.

Hispanic and European-American names. Even though we begin with most commonly used Hispanic (H) names (Supplementary Material D), this is tricky as not all names occur as much as European American names and are thus not as well embedded. We use the frequencies from the dataset to guide us in selecting commonly used names that are also most frequent in the Wikipedia dataset. Using the same method as Section 4.1, we determine the direction, vB,namesv_{B,\textsf{names}}, which encodes this racial difference and find the words most commonly aligned with it. Other Hispanic and European-American names are the closest words. But other words like, latino or hispanic also appear to be close, which affirms that we are capturing the right subspace.

African-American and European-American names. We see a similar trend when we use African-American names and European-American names (Figure 5). We use the African-American names used by Caliskan et al. (2017) . We determine the bias direction by using method in Section 4.1.

We plot in Figure 5 a few occupation words along the axes defined by H-EA and AA-EA bias directions, and compare them with those along the male-female axis. The embedding is different among the groups, and likely still generally more subordinate-biased towards Hispanic and African-American names as it was for female. Although footballer is more Hispanic than European-American, while maid is more neutral in the racial bias setting than the gender setting. We see this pattern repeated across embeddings and datasets (see Supplementary Material A).

When we switch the type of bias, we also end up finding different patterns in the embeddings. In the case of both of these racial directions, there is a the split in not just occupation words but other words that are detected as highly associated with the bias subspace. It shows up foremost among the closest words of the subspace of the bias. Here, we find words like drugs and illegal close to the H-EA direction while, close to the AA-EA direction, we retrieve several slang words used to refer to African-Americans. These word associations with each racial group can be detected by the WEAT tests (lower means less bias) using positive and negative words as demonstrated by Caliskan et al. (2017) . We evaluate using the WEAT test before and after linear projection debiasing in Table 8. For each of these tests, we use half of the names in each category for finding the bias direction and the other half for WEAT testing. This selection is done arbitrarily and the scores are averaged over 3 such selections.

More qualitatively, as a result of the dampening of bias, we see that biased words like other names belonging to these specific demographic groups, slang words, colloquial terms like latinos are removed from the closest 10% words. This is beneficial since the distinguishability of demographic characteristics based on names is what shows up in these different ways like in occupational or financial bias.

We observed that names can be masked carriers of age too. Using the database for names through time and extracting the most common names from early 1900s as compared to late 1900s and early 2000s, we find a correlation between these names (see Supplementary Material) and age related words. In Figure 6, we see a clear correlation between age and names. Bias in this case does not show up in professions as clearly as in gender but in terms of association with positive and negative words . We again evaluate using a WEAT test in Table 8, the bias before and after debiasing the embedding.

DISCUSSION

Different types of bias exist in textual data. Some are easier to detect and evaluate. Some are harder to find suitable and frequent indicators for and thus, to dampen. Gendered word pairs and gendered names are frequent enough in textual data to allow us to successfully measure it in different ways and project the word embeddings away from the subspace occupied by gender. Other types of bias don’t always have a list of word pairs to fall back on to do the same. But using names, as we see here, we can measure and detect the different biases anyway and then project the embedding away from. In this work we also see how a weighted variant of propection removes bias while retaining best the inherent structure of the word embedding.

References

Appendix A Bias in different embeddings

We explore here how gender bias is expressed across different embeddings, datasets and embedding mechanisms. Similar patterns are reflected across all as seen in Figure 7.

For this verification of the permeative nature of bias across datasets and embeddings, we use the GloVe embeddings of a Wikipedia dump (dumps.wikimedia.org/enwiki/latest/enwiki-latest-pages-articles.xml.bz2, 4.74.7B tokens) Common Crawl (840840B tokens, 2.22.2M vocab) and Twitter (2727B tokens, 1.21.2M vocab) from https://nlp.stanford.edu/projects/glove/, and the WordToVec embedding of Google News (100100B tokens, 33M vocab) from https://code.google.com/archive/p/word2vec/.

Appendix B WORD EMBEDDING ASSOCIATION TEST

Word Embedding Association Test (WEAT) was defined as an analogue to Implicit Association Test (IAT) by Caliskan et al. . It checks for human like bias associated with words in word embeddings. For example, it found career oriented words (executive, career, etc) more associated with male names and male gendered words (’man’,’boy’ etc) than female names and gendered words and family oriented words (’family’,’home’ etc) more associated with female names and words than male. We list a set of words used for WEAT by Calisan et al. and that we used in our work below.

For two sets of target words X and Y and attribute words A and B, the WEAT test statistic is :

s(X,Y,A,B)=∑x∈Xs(x,A,B)−s(y,A,B)s(X,Y,A,B)=\sum_{x\in X}s(x,A,B)-s(y,A,B)

s(w,A,B)=meana∈Acos(a,w)−meanb∈Bcos(b,w)s(w,A,B)=mean_{a\in A}cos(a,w)-mean_{b\in B}cos(b,w) and, cos(a,b)cos(a,b) is the cosine distance between vector a and b.

This score is normalized by std−devw∈X∪Ys(w,A,B)std-dev_{w\in X\cup Y}s(w,A,B). So, closer to 0 this value is, the less bias or preferential association target word groups have to the attribute word groups.

Here target words are occupation words or career/family oriented words and attributes are male/female words or names.

Career : { executive, management, professional, corporation, salary, office, business, career }

Family : { home, parents, children, family, cousins, marriage, wedding, relatives }

Male names : { john, paul, mike, kevin, steve, greg, jeff, bill }

Female names : { amy, joan, lisa, sarah, diana, kate, ann, donna }

Male words : { male, man, boy, brother, he, him, his, son }

Female words : { female, woman, girl, she, her, hers, daughter }

Appendix C DETECTING THE GENDER DIRECTION

For this, we take a set of gendered word pairs as listed in Table 3. From our default Wikipedia dataset, using the embedded vectors for these word pairs (i.e., (woman - man), (she - he), etc), we create a basis for the subspace FF, of dimension 1010. We then try to understand the distribution of variance in this subspace. To do so, we project the entire dataset onto this subspace FF, and take the SVD. The top chart in Figure 10 shows the singular values of the entire data in this subspace FF. We observe that there is a dominant first singular vector/value which is almost twice the size of the second value. After the this drop, the decay is significantly more gradual. This suggests to use only the top singular vector of FF as the gender subspace, not 22 or more of these vectors.

Now, for any word ww in vocabulary WW of the embedding, we can define wBw_{B} as the part of ww along the gender direction.

Based on the experiments shown in Figure 10, it is justified to take the gender direction as the (normalized) first right singular vector, vBv_{B}, or the full data set data projected onto the subspace FF. Then, the component of a word vector ww along vBv_{B} is simply ⟨w,vB⟩vB\langle w,v_{B}\rangle v_{B}.

Calculating this component when the gender subspace is defined by two or more of the top right singular vectors of VV can be done similarly.

We should note here that the gender subspace defined here passes through the origin. Centering the data and using PCA to define the gender subspace lets the gender subspace not pass through the origin. We see a comparison in the two methods in Section 5 as HDHD uses PCA and we use SVD to define the gender direction.

Appendix D Word Lists

actor actress author authoress bachelor spinster boy girl brave squaw bridegroom bride brother sister conductor conductress count countess czar czarina dad mum daddy mummy duke duchess emperor empress father mother father-in-law mother-in-law fiance fiancee gentleman lady giant giantess god goddess governor matron grandfather grandmother grandson granddaughter he she headmaster headmistress heir heiress hero heroine him her himself herself host hostess hunter huntress husband wife king queen lad lass landlord landlady lord lady male female man woman manager manageress manservant maidservant masseur masseuse master mistress mayor mayoress milkman milkmaid millionaire millionairess monitor monitress monk nun mr mrs murderer murderess nephew niece papa mama poet poetess policeman policewoman postman postwoman postmaster postmistress priest priestess prince princess prophet prophetess proprietor proprietress shepherd shepherdess sir madam son daughter son-in-law daughter-in-law step-father step-mother step-son step-daughter steward stewardess sultan sultana tailor tailoress uncle aunt usher usherette waiter waitress washerman washerwoman widower widow wizard witch

D.2 Occupation Words

detective ambassador coach officer epidemiologist rabbi ballplayer secretary actress manager scientist cardiologist actor industrialist welder biologist undersecretary captain economist politician baron pollster environmentalist photographer mediator character housewife jeweler physicist hitman geologist painter employee stockbroker footballer tycoon dad patrolman chancellor advocate bureaucrat strategist pathologist psychologist campaigner magistrate judge illustrator surgeon nurse missionary stylist solicitor scholar naturalist artist mathematician businesswoman investigator curator soloist servant broadcaster fisherman landlord housekeeper crooner archaeologist teenager councilman attorney choreographer principal parishioner therapist administrator skipper aide chef gangster astronomer educator lawyer midfielder evangelist novelist senator collector goalkeeper singer acquaintance preacher trumpeter colonel trooper understudy paralegal philosopher councilor violinist priest cellist hooker jurist commentator gardener journalist warrior cameraman wrestler hairdresser lawmaker psychiatrist clerk writer handyman broker boss lieutenant neurosurgeon protagonist sculptor nanny teacher homemaker cop planner laborer programmer philanthropist waiter barrister trader swimmer adventurer monk bookkeeper radiologist columnist banker neurologist barber policeman assassin marshal waitress artiste playwright electrician student deputy researcher caretaker ranger lyricist entrepreneur sailor dancer composer president dean comic medic legislator salesman observer pundit maid archbishop firefighter vocalist tutor proprietor restaurateur editor saint butler prosecutor sergeant realtor commissioner narrator conductor historian citizen worker pastor serviceman filmmaker sportswriter poet dentist statesman minister dermatologist technician nun instructor alderman analyst chaplain inventor lifeguard bodyguard bartender surveyor consultant athlete cartoonist negotiator promoter socialite architect mechanic entertainer counselor janitor firebrand sportsman anthropologist performer crusader envoy trucker publicist commander professor critic comedian receptionist financier valedictorian inspector steward confesses bishop shopkeeper ballerina diplomat parliamentarian author sociologist photojournalist guitarist butcher mobster drummer astronaut protester custodian maestro pianist pharmacist chemist pediatrician lecturer foreman cleric musician cabbie fireman farmer headmaster soldier carpenter substitute director cinematographer warden marksman congressman prisoner librarian magician screenwriter provost saxophonist plumber correspondent organist baker doctor constable treasurer superintendent boxer physician infielder businessman protege

D.3 Names used for Gender Bias Detection

Male : { john, william, george, liam, andrew, michael, louis, tony, scott, jackson } Female : { mary, victoria, carolina, maria, anne, kelly, marie, anna, sarah, jane }

D.4 Names used for Racial Bias Detection and Dampening

European American : { brad, brendan, geoffrey, greg, brett, matthew, neil, todd, nancy, amanda, emily, rachel } African American : { darnell, hakim, jermaine, kareem, jamal, leroy, tyrone, rasheed, yvette, malika, latonya, jasmine } Hispanic : { alejandro, pancho, bernardo, pedro, octavio, rodrigo, ricardo, augusto, carmen, katia, marcella , sofia }

D.5 Names used for Age related Bias Detection and Dampening

Aged : { ruth, william, horace, mary, susie, amy, john, henry, edward, elizabeth } Youth : { taylor, jamie, daniel, aubrey, alison, miranda, jacob, arthur, aaron, ethan }