Understanding the Origins of Bias in Word Embeddings
Marc-Etienne Brunet, Colleen Alkalay-Houlihan, Ashton Anderson, Richard Zemel
Introduction
As machine learning algorithms play ever-increasing roles in our lives, there are ever-increasing risks for these algorithms to be systematically biased (Zhao et al. 2018; Zhao et al. 2017; Kleinberg et al. 2016; Dwork et al. 2012; Hardt et al. 2016). An ongoing research effort is showing that machine learning systems can not only reflect human biases in the data they learn from, but also magnify these biases when deployed in practice (Sweeney 2013). With algorithms aiding critical decisions ranging from medical diagnoses to hiring decisions, it is important to understand how these biases are learned from data.
In recent work, researchers have uncovered an illuminating example of bias in machine learning systems: Popular word embedding methods such as word2vec (Mikolov et al. 2013a) and GloVe (Pennington et al. 2014) acquire stereotypical human biases from the text data they are trained on. For example, they disproportionately associate male terms with science terms, and female terms with art terms (Angwin et al. 2016; Caliskan et al. 2017). Deploying these word embedding algorithms in practice, for example in automated translation systems or as hiring aids, thus runs the serious risk of perpetuating problematic biases in important societal contexts. This problem is especially pernicious because these biases can be difficult to detect—for example, word embeddings were in broad industrial use before their stereotypical biases were discovered.
Although the existence of these biases is now established, their origins—how biases are learned from training data—are poorly understood. Ideally, we would like to be able to ascribe how much of the overall embedding bias is due to any particular small subset of the training corpus—for example, an author or single document. Naïvely, this could be done directly by removing the document in question, re-training an embedding on the perturbed corpus, then comparing the bias of the original embedding with the bias of the retrained embedding. The change in bias resulting from this perturbation could then be interpreted as the document’s contribution to the overall bias. But this approach comes at a prohibitive computational cost; completely retraining the embedding for each document is clearly infeasible.
In this work, we develop an efficient and accurate method for solving this problem. Given a word embedding trained on some corpus, and a metric to evaluate bias, our method approximates how removing a small part of the training corpus would affect the resulting bias. We decompose this problem into two main subproblems: measuring how perturbing the training data changes the learned word embedding; and measuring how changing the word embedding affects its bias. Our central technical contributions solve the former subproblem (the latter is straightforward for many bias measures). Our method provides a highly efficient way of understanding the impact of every document in a training corpus on the overall bias of a word embedding; therefore, we can rapidly identify the most bias-influencing documents in the training corpus. These documents may be used to manipulate the word embedding’s bias through highly selective pruning of the training corpus, or they may be analyzed in conjunction with metadata to identify particularly biased subsets of the training data.
We demonstrate the accuracy of our technique with experimental results on both a simplified corpus of Wikipedia articles in broad use (Wikimedia 2018), and on a corpus of New York Times articles from 1987–2007 (Sandhaus 2008). Across a range of experiments, we find that our method’s predictions of how perturbing the input corpus will affect the bias of the embedding are extremely accurate. We study whether our results transfer across embedding methods and bias metrics, and show that our method is much more efficient at identifying bias-inducing documents than other approaches. We also investigate the qualitative properties of the influential documents surfaced by our method. Our results shed light on how bias is distributed throughout the documents in the training corpora, as well as expose interesting underlying issues in a popular bias metric.
Related Work
Word embeddings are compact vector representations of words learned from a training corpus, and are actively deployed in a number of domains. They not only preserve statistical relationships present in the training data, generally placing commonly co-occurring words close to each other, but they also preserve higher-order syntactic and semantic structure, capturing relationships such as Madrid is to Spain as Paris is to France, and Man is to King as Woman is to Queen (Mikolov et al. 2013b). However, they have been shown to also preserve problematic relationships in the training data, such as Man is to Computer Programmer as Woman is to Homemaker (Bolukbasi et al. 2016).
A recent line of work has begun to develop measures to document these biases as well as algorithms to correct for them. Caliskan et al. 2017 introduced the Word Embedding Association Test (WEAT) and used it to show that word embeddings trained on large public corpora (e.g., Wikipedia, Google News) consistently replicate the known human biases measured by the Implicit Association Test (Greenwald et al. 1998). For example, female terms (e.g., “her”, “she”, “woman”) are closer to family and arts terms than they are to career and math terms, whereas the reverse is true for male terms. Bolukbasi et al. 2016 developed algorithms to de-bias word embeddings so that problematic relationships are no longer preserved, but unproblematic relationships remain. We build upon this line of work by developing a methodology to understand the sources of these biases in word embeddings.
Stereotypical biases have been found in other machine learning settings as well. Common training datasets for multilabel object classification and visual semantic role labeling contain gender bias and, moreover, models trained on these biased datasets exhibit greater gender bias than the training datasets (Zhao et al. 2017). Other types of bias, such as racial bias, have also been shown to exist in machine learning applications (Angwin et al. 2016).
Recently, Koh & Liang 2017 proposed a methodology for using influence functions, a technique from robust statistics, to explain the predictions of a black-box model by tracing the learned state of a model back to individual training examples (Cook & Weisberg 1980). Influence functions allow us to efficiently approximate the effect on model parameters of perturbing a training data point. Other efforts to increase the explainability of machine learning models have largely focused on providing visual or textual information to the user as justification for classification or reinforcement learning decisions (Ribeiro et al. 2016; Hendricks et al. 2016; Lomas et al. 2012).
Background
2 Influence Functions
Influence functions offer a way to approximate how a model’s learned optimal parameters will change if the training data is perturbed. We summarize the theory here.
Let be a convex scalar loss function for a learning task, with optimal model parameters of the form in Equation (2) below, where are the training data points and is the point-wise loss.
where is the Hessian of the total loss, and it is assumed . Note that we have extended the equations presented by Koh & Liang 2017 to address multiple perturbations. This is explained in the supplemental materials.
3 The Word Embedding Association Test
The Word Embedding Association Test (WEAT) measures bias in word embeddings (Caliskan et al. 2017). It considers two equal-sized sets , of target words, such as math, algebra, geometry, calculus} and poetry, literature, symphony, sculpture}, and two sets , of attribute words, such as male, man, boy, brother, he} and female, woman, girl, sister, she}.
The similarity of words and in word embedding is measured by the cosine similarity of their vectors, . The differential association of word with the word sets and is measured with:
For a given , the effect size through which we measure bias is:
Where mean and std-dev refer to the arithmetic mean and the sample standard deviation respectively. Note that only depends on the set of word vectors
Methodology
Our technical contributions are twofold. First, we formalize the problem of understanding bias in word embeddings, introducing the concepts of differential bias and bias gradient. Then, we show how the differential bias can be approximated in word embeddings trained using the GloVe algorithm. We address how to approximate the bias gradient in GloVe in the supplemental material.
Which is the incremental contribution of part to the total bias. This value decomposes the total bias, enabling a wide range of analyses (e.g., studying bias across metadata associated with each part).
It is natural to think of as a collection of individual documents, and think of as a single document. Since a word embedding is generally trained on a corpus consisting of a large set of individual documents (e.g., websites, newspaper articles, Wikipedia entries), we use this framing throughout our analysis. Nonetheless, we note that the unit of analysis can take an arbitrary size (e.g., paragraphs, sets of documents), provided that only a relatively small portion of the corpus is removed. Thus our methodology allows an analyst to study how bias varies across documents, groups of documents, or whichever grouping is best suited to the domain.
Co-occurrence perturbations.
Bias Gradient.
If a word embedding is (or can be approximated by) a differentiable function of the co-occurrence matrix , and the bias metric is also differentiable, we can consider the bias gradient:
Where the above equality is obtained using the chain rule.
The bias gradient has the same dimension as the co-occurrence matrix . While is a daunting size, if the bias metric is only affected by a small subset of the words in the vocabulary, as is the case with the WEAT bias metric, the gradient will be very sparse. It may then be feasible to compute and study. Since it “points” in the direction of maximal bias increase, it provides insight into the co-occurrences most affecting bias.
We then rearrange, and apply the chain rule, obtaining:
2 Computing the Differential Bias for GloVe
To overcome the computational barrier of using influence functions in large models, Koh & Liang 2017 use the LiSSA algorithm (Agarwal et al. 2017) to efficiently compute inverse Hessian vector products. They compute influence in roughly time, where is the number of model parameters and is the number of training examples. However, our analysis and initial experimentation showed that this method would still be too slow for our needs. In a typical setup, GloVe simply has too many model parameters (), and most corpora of interest cause to be too large. One of our principal contributions is a simplifying assumption about the behavior of the GloVe loss function around the learned embedding . This simplification causes the Hessian of the loss to be block diagonal, allowing for the rapid and accurate approximation of the differential bias for every document in a corpus.
Tractably approximating influence functions.
and the total loss is then , now in the form of Equation (2).
where the D-dimensional vector given by is:
From Equation (7), we see that the Hessian of the point-wise loss with respect to , (a -dimensional matrix), is extremely sparse, consisting of only a single block in the th diagonal block position. As a result, the Hessian of the total loss, (also a matrix), is block diagonal, with blocks of dimension . Each diagonal block is given by:
which is the Hessian with respect to only word vector of the point-wise loss at .
An efficient algorithm.
Experimentation
Our experimentation has several objectives. First, we test the accuracy of our differential bias approximation. We then compare our method to a simpler count-based baseline. We also test whether the documents which we identify as bias influencing in GloVe embeddings affect bias in word2vec. Finally, we investigate the qualitative properties of the influential documents surfaced by our method. Our results shed light on how bias is distributed throughout the documents in the training corpora, and expose interesting underlying issues in the WEAT bias metric.
We use two corpora in our experiments, each with a different set of GloVe hyperparameters. This first setup consists of a corpus constructed from a Simple English Wikipedia dump (2017-11-03) (Wikimedia 2018) using 75-dimensional word vectors. These dimensions are small by the standards of a typical word embedding, but sufficient to start capturing syntactic and semantic meaning. Performance on the TOP-1 analogies test shipped with the GloVe code base was around 35%, lower than state-of-the-art performance but still clearly capturing significant meaning.
Our second setup is more representative of the academic and commercial contexts in which our technique could be applied. The corpus is constructed from 20 years of New York Times (NYT) articles (Sandhaus 2008), using 200-dimensional vectors. The TOP-1 analogy performance is approximately 54%. The details of these two configurations are tabulated in the supplemental material.
Choice of experimental bias metric.
Throughout our experiments, we consider the effect size of two different WEAT biases as presented by Caliskan et al. 2017. Recall that these metrics have been shown to correlate with known human biases as measured by the Implicit Association Test. In WEAT1, the target word sets are science and arts terms, while the attribute word sets are male and female terms. In WEAT2, the target word sets are musical instruments and weapons, while the attribute word sets are pleasant and unpleasant terms. A full list of the words in these sets can be found in the supplemental material. They are summarized in Table 1. These sets were chosen so as to include one societal bias that would be widely viewed as problematic, and another which would be widely viewed as benign.
2 Testing the Accuracy of our Method
To test the accuracy of our methodology, ideally we would simply remove a single document from a word embedding’s corpus, train a new embedding, and compare the change in bias with our differential bias approximation. However, the cosine similarities between small sets of word vectors in two word embeddings trained on the same corpus can differ considerably simply because of the stochastic nature of the optimization (Antoniak & Mimno 2018). As a result, the WEAT biases vary between training runs. The effect of removing a single document, which is near zero for a typical document, is hidden in this variation. Fixing the random seed is not a practical approach. Many popular word embedding implementations also require limiting training to a single thread to fully eliminate randomness. This would make experimentation prohibitively slow.
In order to obtain measurable changes, we instead remove sets of documents, resulting in larger corpus perturbations. Accuracy is assessed by comparing our method’s predictions to the actual change in bias measured when each document set is removed from the corpus and a new embedding is trained on this perturbed corpus. Furthermore, we make all predictions and assessments using several embeddings, each trained with the same hyperparameters, but differing in their random seeds.
We construct three types of perturbation sets: increase, random, and decrease. The targeted (increase, decrease) perturbation sets are constructed from the documents whose removals were predicted (by our method) to cause the greatest differential bias, e.g., the documents located in the tails of the histograms in Figure 1. The random perturbation sets are simply documents chosen from the corpus uniformly at random. For a more detailed description, please refer to the supplemental material. Most of the code used in the experimentation has been made available online Code at https://github.com/mebrunet/understanding-bias.
Experimental Results.
Here we present a subset of our experimental results, principally from NYT WEAT1 (science vs. arts). Complete sets of results from the four configurations ({NYT, Wiki} {WEAT1, WEAT2}) can be found in the supplemental materials.
The baseline WEAT effect sizes ( 1 std. dev.) are shown in Table 2. It is worth noting that the WEAT2 (weapons vs. instruments) bias was not significant in our Wiki setup. However, our analysis does not require that the bias under consideration fall within any particular range of values.
A histogram of the differential bias of removal for each document in our NYT setup (WEAT1) can be seen in Figure 1. Notice the log scale on the vertical axis, and how the vast majority of documents are predicted to have a very small impact on the differential bias.
We assess the accuracy of our approximations by measuring how they correlate with the ground truth change in bias (as measured by retraining the embedding after removing a subset of the training corpus). Recall these ground truth changes are obtained using several retraining runs with different random seeds. We find extremely strong correlations () in every configuration, for example Figure 2.
We further compare our approximations to the ground truth in Figure 3. We see that while our approximations underestimate the magnitude of the change in effect size when the perturbation causes the bias to invert, relative ranking is nonetheless preserved. There was no apparent change in the TOP-1 analogy performance of the perturbed embeddings.
We ran a Welch’s t-test comparing the perturbed embeddings’ biases with the baseline biases measured in the original (unperturbed) embeddings. For 36 random perturbation sets, only 2 differed significantly () from the baseline. Both of these sets were perturbations of the smaller Wiki corpus and they only caused a significant difference for WEAT2. This is in strong contrast to the 40 targeted perturbation sets, where only 2 did not significantly differ from their respective baselines. In this case, both were from the smallest (10 document) perturbation sets.
3 Comparison to a PPMI Baseline
We have shown that our method can be used to identify bias-influencing documents and accurately approximate the impact of their removal, but how does it compare to a more naive, straightforward approach? The positive point-wise mutual information (PPMI) matrix is a count-based distributed representation commonly used in natural language processing (Levy et al. 2015). We compare the WEAT effect size in our NYT GloVe embeddings versus when measured in the corpus’ PPMI representation (on 2000 randomly generated word sets). As expected, there is a clear correlation (). It is therefore sensible to use the change in PPMI WEAT effect size to predict how the GloVe WEAT effect size will change.
A change in the PPMI representation due to a co-occurrence perturbation (e.g. document removal) can be computed rapidly. This allows us to scan the whole corpus for the most bias influencing documents. However, we find that the documents identified in this way have a much smaller impact on the bias than those identified by our method. For example in our Wiki setup (WEAT1) removing the 10 documents identified as most bias increasing by the PPMI method reduced the WEAT effect size by 4%. In contrast, the 10 identified by our method reduced it by 40%. Further comparisons are tabulated in the supplemental material.
4 Impact on Word2Vec and Other Bias Metrics
The documents identified as influential by our method clearly have a strong impact on the WEAT effect size in GloVe embeddings. Here we explore how those same documents impact the bias in word2vec embeddings, as well as other bias metrics.
We start by training five word2vec emebeddings with comparable hyperparameters We use a CBOW architecture with the same vocabulary, vector dimensions, and window size as our GloVe embeddings. for each perturbation set, and measure how their removals affect the bias. Figure 4 shows how the WEAT effect size changes in GloVe, the PPMI, and word2vec for each set (NYT-WEAT1). We see that while the response is weaker, both the PPMI representation and the word2vec embeddings show a clear change in effect size due to the perturbations. For example, the baseline WEAT effect size in word2vec is in the unperturbed corpus, but after removing decrease-10000 (the 10k most bias contributing documents for GloVe), the effect size drops to . This means we have nearly neutralized the bias in word2vec through the removal of less than of the corpus (and there is no significant change in TOP-1 analogy performance).
We also see a change as measured by other bias metrics in our perturbed GloVe embeddings. The metric proposed by Bolukbasi et al. 2016 involves computing a single dimensional gender subspace using a definitional sets of words. One can then project test words onto this axis and measure how the embedding implicitly genders them. We explore this in our NYT setup by using the WEAT 1 attribute word sets (male, female) to construct a gender axis, then projecting the target words (science, arts) onto it. In Figure 5 we show the baseline projections and compare them to the projections after having removed the 10k most bias increasing and bias decreasing documents. We see a strong response to the perturbations in the expected directions.
5 Qualitative Analysis
We’ve demonstrated that removing the most influential documents identified by our methodology significantly impacts the WEAT, a metric that has been shown to correlate with known human biases. But can the semantic content of these documents be intuitively understood to affect bias?
We comment here on the 50 most bias influencing documents in the New York Times corpus, considering the WEAT 1 bias metric ({male, female}, {science, arts}). This list is included in the supplemental materials. We indeed found that most of these documents could be readily understood to affect the bias in the expected semantic sense. For example, the second most bias decreasing document is entitled “For Women in Astronomy, a Glass Ceiling in the Sky”, which investigates the pay and recognition gap in astronomy. Many of the other bias decreasing documents included interviews with female doctors or scientists.
Correspondingly, the most bias increasing documents consisted mainly of articles describing the work of male engineers and scientists. There were several obituary entries detailing the scientific accomplishments of men, e.g., “Kaj Aage Strand, 93, Astronomer At the U.S. Naval Observatory”. Perhaps the most self-evident example was an article entitled “60 New Members Elected to Academy of Sciences”, a list of almost exclusively male scientists receiving awards.
There were, however, a few examples of articles that seemed like their semantic content should affect the bias inversely to how they were categorized. For example, an article entitled “The Guide”, a guide to events in Long Island, mentions that the group Woman in Science would be hosting an astronomy event, but nonetheless increases the bias. Only 2 or 3 documents seemed altogether unrelated to the bias’ theme.
Surprisingly, some of the most bias influencing articles contained none of the science or arts WEAT terms explicitly, only synonyms (and some of the male or female terms). This shows that the impact of secondary co-occurrences can be very strong. A naive approach to understanding bias may only consider co-occurrences between WEAT words, but our method shows that this would miss some of the most bias influencing documents in the corpus.
Importantly, we also noticed a large portion of the most bias influencing documents dealt with astronomy or contained hers, the rarest words their respective WEAT subsets. Upon further investigation, we found that the log of a word’s frequency is correlated with the extent to which its relative position (among WEAT words) is affected by the perturbation sets . This can be seen in Figure 5. Not surprisingly, our results indicate that the embedded representations of rare words are more sensitive to corpus perturbations. However, this leaves the WEAT metric vulnerable to exploitation through the manipulation of rarer words. The WEAT effect size is an average of cosine-similarities between the embedded representations of four subsets of words. A handful of well chosen documents can significantly alter the embeddings of a few rare words in those subsets. Therefore documents containing the rare words can have a disproportionate impact on the metric. This weakness helps explain how removing a mere of articles can reverse the WEAT effect size in the New York Times, as is shown in Figure 3, decrease-1000.
Conclusion
In this work, we introduce the problem of tracing the origins of bias in word embeddings, and we develop and experimentally validate a methodology to solve it. We conceptualize the problem as measuring the resulting change in bias when we remove a training document (or small subset of the training corpus), and interpret this as the amount of bias contributed by the document to the overall embedding bias. Computing this naively for each training document would be infeasible. We develop an efficient approximation of this differential bias using influence functions and apply it to the GloVe word embedding algorithm. We experimentally validate our approach and find that it very accurately approximates the true change in bias that results from manually removing training documents and retraining. It performs well on tests using Simple Wikipedia and New York Times corpora and two WEAT bias metrics.
Our work represents a new approach to understanding how machine learning algorithms learn biases from training data. Our methodology could be applied to assess how the bias of a set of texts has evolved over time. For example, using publicly available datasets of newspaper articles or books, one could measure how cultural biases as measured by WEAT or other metrics have evolved over time. More broadly, our efficient method for tracing how perturbations in training data affect changes in the bias of the output is a general idea, and could be applied in many other contexts.
References
Supplemental Material
Appendix A Computing the Bias Gradient for GloVe
The bias gradient can be thought of as a matrix indicating the direction of perturbation of the corpus (co-occurrences) that will result in the maximal change in bias.
When the bias metric is only a function of a small subset of the words in the vocabulary, as in the case of WEAT, this can be further simplified to:
Where are the indices of the words used by the bias metric; for WEAT. For the bias metrics we have explored, the first part of this expression, , can be efficiently computed through automatic differentiation. The difficulty lies in finding an expression for . However, in Section 4.2 of the main text we developed an approximation for the learned embedding under corpus (co-occurrence) perturbations in GloVe using influence functions. We can use this same approximation to create an expression for that is differentiable in .
Recall, given the learned optimal GloVe parameters , , , , on co-occurrence matrix , we can approximate the word vectors given a small corpus perturbation as:
evaluated at . Alternatively the Jacobian can simply by obtained using automatic differentiation.
Substituting this result into Equation (9), we get:
Which gives us the full approximation of the Bias Gradient in GloVe.
Appendix B Experimental Setup
Table 3 presents a summary of the corpora and embedding hyperparameters used throughout our experimentation. We list the complete set of words used in each of the two WEATs below.
Appendix C Detailed Experimental Methodology
Here we detail the experimental methodology used to test our method’s accuracy.
We start by training 10 word embeddings using the parameters in Table 3 above, but using different random seeds. These embeddings create a baseline for the unperturbed bias .
II - Approximate the differential bias of each document.
For each WEAT test, we approximate the differential bias of every document in the corpus. We do so with a combination of Equations (8) and (5) of the main text. This step is summarize by Algorithm 1 in the main text. Note that we make the differential bias approximation for each document several times, using the learned parameters , , and from the 10 different baseline embeddings in our different approximations. We then average these approximations for each document, and construct a histogram.
III - Construct perturbation sets.
We perturb the corpus by removing sets of documents. We construct three types of perturbation sets: increase, random, and decrease. The targeted (increase, decrease) perturbation sets are constructed from the documents whose removals were predicted to cause the greatest differential bias (in absolute value), i.e., the documents located in the tails of the histograms. For the Wiki setup we consider the 10, 30, 100, 300, and 1000 most influential documents for each bias, while for the NYT setup we consider the 100, 300, 1000, 3000, and 10,000 most influential. This results in 10 perturbations sets per corpus per bias, for a total of 40.
The random sets are, as their name suggests, drawn uniformly at random from the entire set of documents used in the training corpus. For the Wiki setup we consider 6 sets of 10, 30, 100, 300, and 1000 documents (30 total). Because training times are much longer, we limit this to 6 sets of 10,000 documents for the NYT setup. Therefore we consider a total of 36 random sets.
IV - Approximate the differential bias of each perturbation set.
We then approximate the differential bias of each perturbation set. Note that is not linear in . Therefore determining the differential bias of a perturbation set does not amount to simply summing the differential bias of each document in the set (although in practice we find it to be close). Here we also make 10 approximations, one with each of the different baseline embeddings.
V - Construct ground truth and assess.
Finally, for each perturbation set, we remove the target documents from the corpus, and train 5 new embeddings on this perturbed corpus. We use the same hyperparameters, again varying only the random seed. The bias measured in these embeddings serve as the ground truth for assessment.
Appendix D Additional experimental results
Here we include additional experimental results.
Appendix E Influential Documents - NYT WEAT 1
Appendix F Influence of Mulitple Perturbations
Here we show how we can extend the influence function equations presented by Koh & Liang (2017) to address the case of multiple training point perturbations. We do not intend this to be a rigorous mathematical proof, but rather to provide insight into the logical steps we followed.
First we summarize the derivation in the case of a single train point perturbation. Let be a convex scalar loss function for a learning task, with optimal model parameters of the form in Equation 12 below, where are the training data points and is the point-wise loss.
for which we can compute the first order Taylor series expansion (with respect to ) around . This gives:
where is the Hessian of the total loss.