A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector Spaces
Anne Lauscher, Goran Glavaš, Simone Paolo Ponzetto, Ivan Vulić
Introduction
Distributional word vectors have been recently shown to encode prominent human biases related to, e.g., gender or race (?; ?; ?). Such biases are observed across languages and embedding methods (?), both in static and contextualized word embeddings (?). While this issue requires remedy, the finding itself is hardly surprising: we project our biases, in terms of biased word co-occurrences, into the texts we produce. Consequently, this is propagated to embedding models, both static (?; ?; ?) and contextualized (?) alike, by virtue of the distributional hypothesis (?).Borrowing the famous example (?), man will be found more often in the same context with programmer, and woman with homemaker in any sufficiently large corpus. While biases may be useful for diachronic or sociological analyses (?), they (1) raise ethical issues, since biases are amplified by machine learning models using embeddings as input (?), and (2) impede tasks like coreference resolution (?; ?) and abusive language detection (?).
A number of methods for attenuating and eliminating human-like biases in word vectors have been proposed recently (?; ?; ?; ?). While they address the same types of bias – primarily the gender bias – they start from different bias “specifications” and either lack proper empirical evaluation (?) or employ different evaluation procedures, both hindering a direct comparison of the methods’ “debiasing abilities” (?; ?; ?). What is more, the most prominent debiasing models (?; ?) have been criticized for merely masking the bias instead of removing it (?). To resolve inconsistencies in the current debiasing research and evaluation, in this work we propose a general debiasing framework DEBIE (DEBiasing embeddings Implicitly and Explicitly), which operationalizes bias specifications, groups models according to the bias specification type they operate on, and evaluates models’ abilities to remove biases both explicitly and implicitly (?).
We first define two types of bias specifications – implicit and explicit – and propose a method of augmenting bias specifications with the help of embeddings specialized for semantic similarity (?; ?). We then introduce the main contributions of this work as follows. First, we present three novel debiasing models. (1) We adjust the linear projection method of ? (?), an extension of the debiasing model of ? (?), to operate on the augmented implicit bias specifications. (2) We then propose an alternative model that projects the embedding space to itself using the term sets from implicit bias specification as the projection signal. (3) Finally, we propose a simple and effective neural debiasing model, which is, to the best of our knowledge, the first debiasing model that operates on an explicit bias specification. All three models perform post-hoc debiasing: they can be applied to any pretrained distributional word vector space.In contrast, debiasing models like GN-GloVe (?) integrate debiasing constraints into objectives of embedding models like GloVe (?), and thus cannot be directly ported to other embedding models. As another contribution, we combine existing bias metrics with newly proposed ones and assemble an evaluation suite that tests word vectors for explicit biases, implicit biases, and (preservation of) semantic quality. Finally, by coupling the proposed debiasing models with the cross-lingual embedding spaces (?; ?), we facilitate cross-lingual debiasing transfer: we successfully debias embedding spaces in target languages without bias specifications in those languages. We hope that our work will lead to standardization of preprocessing and evaluation procedures in debiasing research and to increased comparability of debiasing models.The code is available at https://github.com/umanlp/DEBIE.
General Debiasing Framework
In what follows, we first formalize two bias specifications – implicit and explicit. We then introduce new debiasing models: two operate on the implicit bias specification and the third on the explicit bias specification. Finally, we show how to debias word embeddings in a variety of target languages via cross-lingual embeddings.
An implicit bias specification consists of two sets of target terms with respect to which a bias is expected to exist in the embedding space. For example, two sets of science and art terms, = and constitute an implicit specification of the gender bias. Strictly speaking, does not specify a bias directly – it merely specifies two categories of concepts for which we implicitly assume that there exists some set of reference terms (e.g., male terms man, father and/or female terms like woman, girl) with respect to which and exhibit differences. Most existing debiasing models (?; ?; ?; ?) operate on , i.e., not requiring reference terms .
An explicit bias specification defines, in addition to sets and , one or more reference attribute sets. We consider an explicit bias specification with a single attribute set, (as employed by our DebiasNet model),The attribute set can be any set of attributes towards which the bias is to be removed. In our experiments, we joined the WEAT test specification attribute sets and . and also with two (opposing) attribute sets, , as used in WEAT tests (?).
The initial bias specification ( or ) commonly contains only a handful of words in each target and attribute set. These are commonly the most representative words of a category (e.g., man, boy, father to represent the category male). However, in order to provide a finer-grained bias specification, we propose to augment each term set with synonyms and semantically similar words of the initial terms. We therefore extract nearest neighbours of initial terms from an embedding space specialized to accentuate true semantic similarity and attenuate other types of semantic association (?; ?; ?, inter alia). For the augmentation process we rely on the recent state-of-the-art similarity specialization method of ? (?): for more details see the original work.
Given or and a similarity-specialized word vector space , we augment each of the term sets in the specification by retrieving the top most (cosine-)similar terms from for each of the initial terms.We discard nearest neighbors initially present in other sets of the same bias specification: e.g., if we retrieve an augmentation candidate woman for an initial term man, woman will not be added to if it exists in (or in -s). Extending bias specification sets using a similarity-specialized word vector space – as opposed to a regular distributional space – reduces the noisy augmentation stemming from semantic relatedness instead of true semantic similarity.We also considered using clean lexical knowledge from WordNet (?) directly, but this resulted in much lower recall as well as less accurate augmentation candidates. Table 1 illustrates the initial bias specification and the corresponding augmentation (showing nearest neighbors, without the initial terms) for one explicitly defined gender bias.
Debiasing Models
We present three debiasing models, two of which operate on and one on the explicit bias specification .
focuses on as a generalization of the linear projection model proposed by ? (?), itself, in turn, an extension of the hard-debiasing model of ? (?).
where denotes a dot product. In other words, the closer the vector is to the global bias direction , the more it is bias-corrected (i.e., the larger portion of is subtracted from ). Vectors orthogonal to the bias direction remain unchanged (zero dot-product with the bias vector ).
Bias-Alignment Model (BAM).
Here, we use pairs to learn the debiasing projection of with respect to itself. Let and be the matrices obtained by stacking (biased) vectors of left and right words of pairs , respectively. We then learn the orthogonal map , where is the singular value decomposition of . Since is orthogonal, the projection is isomorphic to the original space , and thus equally biased. However, the transformation (specified by ) defines the angle and direction of debiasing. We obtain the debiased space by averaging the original space and the projected space :
Explicit Neural Debiasing (DebiasNet).
The final model, dubbed DebiasNet, is a neural model that operates on the explicit bias specification . It is inspired by the work on semantic specialization of word embeddings (?; ?): but instead of using linguistic constraints (e.g., synonyms), we “specialize” the vector space by leveraging debiasing constraints.
Given a biased input space and the specification , we learn a debiasing function that transforms to a debiased space . We aim for the terms from both sets and to be similarly close to the terms from in . For simplicity, we execute DBN as a feed-forward neural network with non-linear activations. The training set for learning the parameters consists of triples . It is obtained as a full Cartesian product . Let , and be the respective vectors of , , and from the input biased space , and let , and be their “debiased” transformations: , , and . For a training instance , we then minimize the following loss function :
refers to the cosine distance. The objective pushes the terms from the two target sets and to be equidistant to the terms from the attribute set . That is, it is designed to specifically remove the explicit bias. By minimizing as the only objective, the model would remove the bias, but it would also destroy the useful semantic information in the input space. We thus couple the objective with the regularization that prevents the debiased vectors to deviate too much from their original estimates:
The final loss is then , with as the regularization weight. The learned function is then applied to the full input space: .
Composing Debiasing Models.
The presented models can be seamlessly composed with one another. For example, given an explicit specification , we can first explicitly debias a distributional space using DebiasNet. We can then apply either GBDD or BAM on the resulting vector space by deriving from (i.e., by considering only and ): e.g., .
Cross-Lingual Transfer of Debiasing
Cross-lingual word embeddings have been shown to be a viable solution for zero-shot language transfer of NLP models (?; ?). Conceptually, given a source language with its monolingual distributional space and a target language with the space , we can apply any model trained on on the instances from , given a matrix that projects to . From the plethora of cross-lingual word embedding models (?; ?; ?, inter alia), we opt for a supervised projection-based model (?) that obtains by solving the Procrustes problem (?) on the set of word translation pairs.Note that we obtain the cross-lingual projection in the similar way as debiasing projection in BAM; but now the aligned matrices contain vectors (each from respective language) corresponding to word translation pairs (not pairs created from bias target sets as in BAM). We select this approach due to its simplicity and competitive zero-shot language transfer performance on other NLP tasks (?). With the cross-lingual projection matrix in place, the debiasing of the space amounts to composing the projection with the debiasing model in : e.g., for GBDD, .
Evaluation and Experimental Setup
We now introduce the metrics for testing different aspects of debiased embedding spaces, and then outline two datasets used in our experiments.
We use three diverse tests to measure the presence of explicit bias, and two tests that focus on the presence of implicit bias. Finally, we test the debiased spaces for their ability to preserve the initial semantic information.
Introduced by ? (?), WEAT tests the embedding space for the presence of an explicit bias defined as =. It computes the differential association between and based on their mean similarity with terms from the attribute sets and :
The association of term is computed as:
The significance of the statistic is computed by comparing with the scores obtained with all permutations , where and are equally sized partitions of . The -value of the test is the probability of . The “amount” of bias, the so-called effect size, is then a normalized measure of separation between association distributions:
where is the mean and is the standard deviation.
Embedding Coherence Test (ECT).
It quantifies the amount of explicit bias = (?). Unlike WEAT, it compares vectors of target sets and (averaged over the constituent terms) with vectors from a single attribute set . ECT first computes the mean vectors for the target sets and : and . Next, for both and it computes the (cosine) similarities with vectors of all . The two resultant vectors of similarity scores, (for ) and (for ) are used to obtain the final ECT score. It is the Spearman’s rank correlation between the rank orders of and – the higher the correlation, the lower the bias.
Bias Analogy Test (BAT).
Based on the observation of (?) that in a biased vector space , ? (?) proposed an analogy-based bias test: Embedding Quality Test (EQT). However, EQT depends on WordNet to extend the bias definition with synonyms and plurals of bias specification terms. In contrast, we propose an alternative Bias Analogy Test (BAT) that relies only on the specification .
We first create all possible biased analogies for . We then create two query vectors from each analogy: and for each 4-tuple . We then rank the vectors in the vector space according to the Euclidean distance with each of the query vectors. In a biased space, we expect the vector to be ranked higher for the query than the vectors of terms from the opposing attribute set (e.g., for a gender-biased space we expect woman to be ranked higher than father or boy for the query man - programmer + homemaker). Also, is expected to be more similar to than vectors of terms . The BAT score is the percentage of cases where: (1) is ranked higher than a term for and (2) is ranked higher than a term for .
Implicit Bias Tests.
? (?) recently suggested that the two sets of target terms can still be clearly distinguished (with KMeans clustering, or in a supervised manner with an SVM classifier) from one another after applying debiasing procedures of (?) and (?). We adopt their approach and test the debiased spaces for the presence of implicit bias by clustering terms from and with KMeans++, and by classifying them using an SVM with the RBF kernel: it is trained on the vectors of terms from the augmentations of target sets. For each debiasing model, we average the clustering and classification scores over independent runs.
Semantic Quality.
Debiasing procedures change the topology of the input vector space; we thus have to verify that debiasing does not occur at the expense of the encoded semantic information. We test the debiased embedding spaces on two standard word similarity/relatedness benchmarks: SimLex-999 (?) and WordSim-353 (?).
Evaluation Datasets
Our proposed framework is versatile as it enables debiasing models to operate on any bias specified in the or format. To demonstrate this, we evaluate the debiasing models from the previous section on two different bias specifications: tests T1 and T8 from the WEAT dataset (?). WEAT tests are given as explicit bias specifications = .
WEAT T8, shown in Table 1, encodes a type of a gender bias in relation to affinities towards science and art. contains terms from the areas of science and technology, whereas contains art terms. Attribute sets contain male () and female () terms. In a gender-biased vector space the scientific targets are expected to be more strongly associated with male attributes, and artistic targets with female terms.
WEAT 1: Flowers vs. Insects.
WEAT T1 specifies another bias type: the difference in sentiment humans attach to insects as opposed to flowers. Target sets contain different flowers () and insect species (), and attribute sets contain universally positive () and negative () terms. The full bias specification of WEAT T1 is available in the supplementary.
XWEAT.
For evaluating the language transfer setup, we use bias specifications in target languages as our test data. We use tests T1 and T8 from XWEAT, created by ? (?) by translating the English (en) WEAT tests to six languages: German (de), Spanish (es), Italian (it), Russian (ru), Croatian (hr), and Turkish (tr).
Preprocessing and Training Setup
Augmented Bias Specifications. We first augment the bias specifications using a similarity-specialized embedding space produced by ? (?)Available at: https://tinyurl.com/y273cuvk. based on the en fastText embeddings (?). For WEAT T8, we augment the target and attribute lists with nearest neighbours of each term. As the initial lists of WEAT T1 are longer than those of T8, we use with T1. We train all debiasing models using bias specifications containing only the augmentation terms (i.e., without the initial bias specification terms); we use the initial terms for testing.
We test the robustness of debiasing models on three different word embedding models trained on Wikipedia: CBOW (?), GloVe (?), and fastText (FT) (?). For cross-lingual transfer, we induce a multilingual space spanning seven languages (en + 6 targets) by projecting FT vectors of each target to the EN space. Following an established procedure (?), we learn projections using automatically compiled translations of the 5K most frequent en words.
Training Setup.
For GBDD and BAM there is a deterministic closed-form solution for any given bias specification. On the other hand, the hyper-parameters of DebiasNet are optimized via grid search and cross-validation on the training set. The final DebiasNet model uses hidden layers with units each and the weight is fixed to 0.2.
Results and Analysis
We first report debiasing results on three en distributional spaces, for the individual models as well as for three composite models: GBDD BAM = GBDD(BAM()), BAM GBDD, and GBDD DebiasNet.BAM and DebiasNet display similar results and so does their composition. For brevity, we thus omit the scores of BAM DebiasNet. We also do not report the scores with DebiasNet GDBB as its scores were similar to its inverse composition GDBB DebiasNet in our preliminary tests. We then show the results for the cross-lingual debiasing transfer. Finally, we analyze the topology of debiased spaces.
Biases of Distributional Spaces. The main results are summarized in Table 2. All three input distributional spaces generally exhibit explicit and implicit biases, with CBOW spaces displaying the lowest biases, both according to the WEAT tests (e.g., the effect size is even insignificant with for the gender bias test T8) and the implicit bias tests of ? (?). Interestingly – according to our BAT test, and despite the original claims and examples from ? (?) – the encoded biases do not reflect strongly in the analogy tests. Nonetheless, our debiasing methods in most test settings manage to affect the input vector spaces by further reducing BAT scores.
While the results vary across the two WEAT tests and evaluation metrics, GBDD emerges as the most robust model on average. It attenuates the explicit bias while being the most successful in removing the bias implicitly: the spaces debiased with GBDD completely confuse the KM clustering and SVM classifier. It also fully retains the useful semantic information: we do not observe drops on SL and WS compared to the input distributional spaces. While GBDD outperforms BAM and DebiasNet (DBN) on average according to ECT and BAT measures, it is not able to fully remove the explicit gender bias (T8) according to the WEAT test.
Despite operating on an implicit specification , BAM removes the explicit biases much better than the implicit ones. DBN seems even better than BAM in removing the explicit biases. This is not a surprise, since DBN is trained on an explicit bias specification. However both DBN and BAM are unsuccessful in removing the implicit biases. Moreover, DBN distorts the input space more than BAM, yielding substantial drops on SL and WS.
The complementarity of debiasing effects between GBDD, and BAM/DBN are confirmed by the performance of their compositions. All composition models robustly remove both explicit and implicit biases, also showing that there is no “one model rules them all” solution to various debiasing aspects. GBDD DBN most effectively removes the biases, but it inherits the undesirable semantic distortions of DBN. On the other hand, BAM GBDD offers solid bias removal while for the most part retaining the semantic quality of the space.
Differences between Evaluation Measures.
The three aspects of evaluation complement each other: they all inform the selection of the most appropriate debiasing model w.r.t. the desired application-specific criteria.E.g., Note that for some bias specifications, one might not want to reduce/remove the implicit bias. WEAT T1 can be seen as an example of such bias: while we may want to make insects similarly good/bad as flowers, we do not want to make them indistinguishable from flowers in the vector space. However, results of WEAT, ECT, and BAT are not always aligned. For example, the CBOW space is unbiased according to the WEAT test, but extremely biased (negative correlation!) according to ECT. In contrast, GloVe vectors are biased according to WEAT but not according to ECT (correlation of ). These findings point to different bias aspects, accentuating the need for multiple, mutually complementary, bias measures.
Cross-Lingual Transfer
The results in the cross-lingual debiasing transfer are shown in Table 3. For brevity, we show only the results on XWEAT T8 (gender bias wrt. science vs. art) and for a subset of evaluation measures (one for each evaluation aspect): WEAT (W), KMeans++ (KM), and SimLex (SL).We provide the full results, with all evaluation measures, and also on the XWEAT T1 test in the supplemental material.,We evaluate word similarities for de, it, ru, and hr on their respective SimLex datasets (?; ?); there is no es and tr SimLex.
We first confirm the results from ? (?): de, it, ru, and hr fastText vectors do not exhibit significant explicit gender bias (wrt. science vs. art), according to the WEAT test. The explicit bias is, however, significant in es and tr distributional vectors. Implicit bias is clearly present in all distributional spaces except ru. Debiasing models display similar properties as before: DBN reduces the explicit bias more effectively than BAM and GBDD, but it semantically distorts the vectors; and only GBDD successfully removes the implicit bias. None of the models fully removes the explicit bias for tr (the lowest bias effect of for BAM GBDD is still significant). We suspect that this is a result of the lower-quality cross-lingual tren projection, which is in line with the bilingual lexicon induction results from ? (?).
For de and it, BAM and DBN invert the direction of the bias: negative WEAT scores mean that sciences are more correlated with female attributes and arts with male attributes. We believe that this is the result of applying a (strong) bias correction learned on a biased en space on the (explicitly) unbiased de and it spaces. The BAM GBDD composition seems most robust in the cross-lingual transfer setting – it successfully removes both the explicit (if they exist) and implicit biases, while preserving useful semantic information (SL). These results indicate that we can attenuate or remove biases in distributional vectors of languages for which (1) we do not require the initial bias specification and (2) we do not even need similarity-specialized word embeddings used to augment the bias specifications for the target language.
Finally, we qualitatively analyze the debiasing effects suggested by evaluation measures. We project the input and the debiased embeddings into 2D with PCA, and show the constellation of words from the initial bias specification of WEAT T8 (Table 1) in Figure 1.We show only the input space and the spaces debiased with GBDD and BAM. We provide similar illustrations for other debiasing models in the supplementary material. In the distributional space, the two target sets (science vs art) are clearly distinguishable from one another (implicit bias), and so are the male and female attributes. The science terms are notably closer to the male terms and art terms to the female terms (explicit bias). The space produced by BAM intertwines the male and female terms and makes the science and art terms roughly equidistant to the gender terms (explicit bias removed), but the science terms are still clearly distinguishable from art terms (implicit bias still present). In the space produced by GBDD, both biases are removed: science and art terms cannot be clearly separated and are roughly equidistant to gender terms.
Conclusion
We have introduced a general framework for debiasing distributional word vector spaces by 1) formalizing the differences between implicit and explicit biases, 2) proposing new debiasing methods that deal with the two different bias specifications, and 3) designing a comprehensive evaluation framework for testing the (often complementary) effects of debiasing. While the proposed framework offers a systematized view on human biases encoded in word embeddings, the main results indicate that our debiasing methods can effectively attenuate biases in arbitrary input distributional spaces and can also be transferred to a variety of target languages.