Quantifying and Reducing Stereotypes in Word Embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, Adam Kalai
Introduction
While word-embeddings encode semantic information they also exhibit hidden biases inherent in the dataset they are trained on. For instance, word embeddings based on w2vNEWS can return biased solutions to analogy puzzles such as father:doctor :: mother:nurse and man:computer programmer :: woman:homemaker. Other publicly available embeddings produce similar results exhibiting gender stereotypes. Moreover, the closest word to the query BLACK MALE returns ASSAULTED while the response to WHITE MALE is ENTITLED TO. This raises serious concerns about their widespread use.
The prejudices and stereotypes in these embeddings reflect biases implicit in the data on which they were trained. The embedding of a word is typically optimized to predict co-occuring words in the corpus. Therefore, if mother and nurse frequently co-occur, then the vectors and also tend to be more similar and encode the gender stereotypes. The use of embeddings in applications can amplify these biases. To illustrate this point, consider Web search where, for example, one recent project has shown that, when carefully combined with existing approaches, word vectors can significantly improve Web page relevance results Nalisnick et al. (2016) (note that this work is a proof of concept – we do not know which, if any, mainstream search engines presently incorporate word embeddings). Consider a researcher seeking a summer intern to work on a machine learning project on deep learning who searches for, say, “linkedin graduate student machine learning neural networks.” Now, a word embedding’s semantic knowledge can improve relevance in the sense that a LinkedIn web page containing terms such as “PhD student,” “embeddings,” and “deep learning,” which are related to but different from the query terms, may be ranked highly in the results. However, word embeddings also rank CS research related terms closer to male names than female names. The consequence would be, between two pages that differed in the names Mary and John but were otherwise identical, the search engine would rank John’s higher than Mary. In this hypothetical example, the usage of word embedding makes it even harder for women to be recognized as computer scientists and would contribute to widening the existing gender gap in computer science. While we focus on gender bias, specifically male/female, our approach may be applied to other types of biases.
We propose two methods to systematically quantify the gender bias in a set of word embeddings. First, we quantify how words, such as those corresponding to professions, are distributed along the direction between embeddings of he and she. Second, we design an algorithm for generating analogy pairs from an embedding given two seed words and we use crowdworkers to quantify whether these embedding analogies reflect stereotypes. Some analogies reflect stereotypes such as he:janitor :: she:housekeeper and he:alcoholism :: she:eating disorders. Finally, others may provoke interesting discussions such as he:realist :: she:feminist and he:injured :: she:victim.
Since biases are cultural, we enlist U.S.-based crowdworkers to identify analogies to judge whether analogies: (a) reflect stereotypes (to understand biases), or (b) are nonsensical (to ensure accuracy). We first establish that biases indeed exist in the embeddings. We then show that, surprisingly, information to distinguish stereotypical associations like female:homemaker from definitional associations like female:sister can often be removed. We propose an approach that, given an embedding and only a handful of words, can reduce the amount of bias present in that embedding without significantly reducing its performance on other benchmarks.
Contributions. (1) We initiate the study of stereotypes and biases in word embeddings. Our work follows a large body of literature on bias in language, but word embeddings are of specific interest because they are commonly used in machine learning and they have simple geometric structures that can be quantified mathematically. (2) We develop two metrics to quantify gender stereotypes in word embeddings based on words associated with professions together with automatically generated analogies which are then scored by the crowd. (3) We develop a new algorithm that reduces gender stereotypes in the embedding using only a handful of training examples while preserving useful properties of the embedding.
Prior work. The body of prior work on bias in language and prejudice in machine learning algorithms is too large to fully cover here. We note that gender stereotypes have been shown to develop in children as young as two years old Turner & Gervai (1995). Statistical analyses of language have shown interesting contrasts between language used to describe men and women, e.g., in recommendation letters Schmader et al. (2007). A number of online systems have been shown to exhibit various biases, such as racial discrimination in the ads presented to users Sweeney (2013). Approaches to modify classification algorithms to define and achieve various notions of fairness have been described in a number of works, see, e.g., Barocas & Selbst (2014); Dwork et al. (2012) and a recent survey Zliobaite (2015).
Implicit stereotypes in word embedding
Stereotyped words. A simple approach to explore how gender stereotypes manifest in embeddings is to quantify which words are closer to he versus she in the embedding space (using other words to capture gender, such as man and woman, gives similar but noisier results due to their multiple meanings). We used a list of 215 common profession names, removing names that are associated with one gender by definition (e.g. waitress, waiter). For each name, , we computed its projection onto the gender axis: . Figure 1 shows the projection of professions on the w2vNEWS embedding (-axis) and on a different embedding trained by GloVe on a dataset of web-crawled texts (-axis). Several professions are closer to the he or she vector and this is consistent across the embeddings, suggesting that embeddings encode gender stereotypes.
Stereotyped analogies. While professions give easily-interpretable insights on embedding stereotypes, we developed a more general method to automatically detect and quantify gender bias in any word embedding. Embeddings have shown to perform well in analogy tasks. Motivated by this, we ask the embedding to generate analogous word pairs for he and she, and use crowd-sourcing to evaluate the degree of stereotype of each pair.
A desired analogy he:she :: : has the following propertiesFor the ease of presentation, we abuse the notation to use to represent a word or a word vector depending on the context.: 1) the direction of - has to align with he-she; 2) and should be semantically similar, i.e. is not too large. Based on this, given a word embedding , we proposed to score analogous pairs by the following formulation:
where is the gender direction and is a threshold for similarity.We explored alternatives including a variation of 3-CosMul Levy & Goldberg (2014) for generating word pairs, and observe that the proposed approach works the best. We observe that setting often works well in practice; this corresponds to requiring that the two words forming the analogy are significantly closer together than two random embedding vectors.
From the embedding, we generated the top analogous pairs with the largest scores. To avoid redundancies, if multiple pairs share the same or , we kept only one pair. Then we employed Amazon Mechanical Turk to evaluate the analogies. For each analogy, such as man:woman :: doctor:nurse, we ask the Turkers two yes/no questions to verify if this pairing makes sense as an analogy and whether it exhibits gender stereotype. Every word pair is judged by 10 Turkers, and we used the number of Turkers that rated this pair as stereotyped to quantify the degree of bias of this analogy. Table 1 shows the most and least stereotypical analogies generated by word2vec on Googlenews. Overall, 21% and 32% analogy judgments were stereotypical and nonsensical, respectively, by the Turkers.
Reducing stereotypes in word embedding
Having demonstrated that word embeddings contain substantial stereotypes in both professions and analogies, we developed a method to reduce these stereotypes while preserving the desirable geometry of the embedding.
Word embeddings are often trained on a large corpus (w2vNEWS is trained on Google news corpus with 100 billion words). As a result, it is impractical and even impossible (the corpus is not publicly accessible) to reduce the stereotypes during the training of the the word vectors. Therefore, we assume that we are given a a set of word vectors and aim to remove stereotypes as a post-processing step.
The goal is to generate a transformation matrix , which has the following properties: • The transformed embeddings are stereotypical-free. That is every column vectors in should be perpendicular to column vectors in (i.e., ). • The transformed embeddings preserve the distances between any two vectors in the matrix .
Let , we can capture these two objectives as the following semi-positive definite programming problem.
where is the Frobenius norm, the first term ensures that the pairwise distances are preserved, and the second term induces the biases to be small on the seed words. The user-specified parameter balances the two terms.
Directly solving this SDP optimization problem is challenging. In practice, the dimension of matrix is in the scale of 400,000 300. The dimensions of the matrices and are , causing computational and memory issues. We conduct singular value decomposition on , such that , where and are orthogonal matrices and is a diagonal matrix.
The last equality follows the fact that is an orthogonal matrix (.)
Here is a matrix and can be solved efficiently. The solution is the debiasing transformation of the word embedding.
To validate our debiasing algorithm, we asked Turkers to suggest words that are likely to reflect gender stereotype (e.g. manager, nurse). We collected 438 such words, of which a random setup of 350 are used for training as the columns of the matrix. The remaining are used for testing. Figure 2 illustrates the results of the algorithm. The blue circles are the 88 gender-stereotype words suggested by the Turkers which form our held-out test set. The green crosses are a random sample of background words that were not suggested to have stereotype. Most of the stereotype words lie close to the line, consistent with them lies near the midpoint between he and she. In contrast the background points were substantially less affected by the debiasing transformation.
We use variances to quantify this result. For each test word (either gender-stereotypical or background) we project it onto the he - she direction. Then we compute the variance of the projections in the original embedding and after the debiasing transformation. For the gender-stereotype test words, the variance in the original embedding is 0.02 and the variance after the transformation is 0.001. For the background words, the variance before and after the transformation was 0.005 and 0.0055 respectively. This demonstrates that the transformation was able to reduce gender stereotype.
Lastly to verify that the debiasing transformation preserves the desirable geometric structure of the embedding, we tested the transformed embedding a several standard benchmarks that measure whether related words have similar embeddings as well as how well the embedding performs in analogy tasks. Table 2 shows the results on the original and the transformed embeddings and the transformation does not negatively impact the performance.