A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector Spaces

Anne Lauscher, Goran Glavaš, Simone Paolo Ponzetto, Ivan Vulić

Introduction

Distributional word vectors have been recently shown to encode prominent human biases related to, e.g., gender or race (?; ?; ?). Such biases are observed across languages and embedding methods (?), both in static and contextualized word embeddings (?). While this issue requires remedy, the finding itself is hardly surprising: we project our biases, in terms of biased word co-occurrences, into the texts we produce. Consequently, this is propagated to embedding models, both static (?; ?; ?) and contextualized (?) alike, by virtue of the distributional hypothesis (?).Borrowing the famous example (?), man will be found more often in the same context with programmer, and woman with homemaker in any sufficiently large corpus. While biases may be useful for diachronic or sociological analyses (?), they (1) raise ethical issues, since biases are amplified by machine learning models using embeddings as input (?), and (2) impede tasks like coreference resolution (?; ?) and abusive language detection (?).

A number of methods for attenuating and eliminating human-like biases in word vectors have been proposed recently (?; ?; ?; ?). While they address the same types of bias – primarily the gender bias – they start from different bias “specifications” and either lack proper empirical evaluation (?) or employ different evaluation procedures, both hindering a direct comparison of the methods’ “debiasing abilities” (?; ?; ?). What is more, the most prominent debiasing models (?; ?) have been criticized for merely masking the bias instead of removing it (?). To resolve inconsistencies in the current debiasing research and evaluation, in this work we propose a general debiasing framework DEBIE (DEBiasing embeddings Implicitly and Explicitly), which operationalizes bias specifications, groups models according to the bias specification type they operate on, and evaluates models’ abilities to remove biases both explicitly and implicitly (?).

We first define two types of bias specifications – implicit and explicit – and propose a method of augmenting bias specifications with the help of embeddings specialized for semantic similarity (?; ?). We then introduce the main contributions of this work as follows. First, we present three novel debiasing models. (1) We adjust the linear projection method of ? (?), an extension of the debiasing model of ? (?), to operate on the augmented implicit bias specifications. (2) We then propose an alternative model that projects the embedding space to itself using the term sets from implicit bias specification as the projection signal. (3) Finally, we propose a simple and effective neural debiasing model, which is, to the best of our knowledge, the first debiasing model that operates on an explicit bias specification. All three models perform post-hoc debiasing: they can be applied to any pretrained distributional word vector space.In contrast, debiasing models like GN-GloVe (?) integrate debiasing constraints into objectives of embedding models like GloVe (?), and thus cannot be directly ported to other embedding models. As another contribution, we combine existing bias metrics with newly proposed ones and assemble an evaluation suite that tests word vectors for explicit biases, implicit biases, and (preservation of) semantic quality. Finally, by coupling the proposed debiasing models with the cross-lingual embedding spaces (?; ?), we facilitate cross-lingual debiasing transfer: we successfully debias embedding spaces in target languages without bias specifications in those languages. We hope that our work will lead to standardization of preprocessing and evaluation procedures in debiasing research and to increased comparability of debiasing models.The code is available at https://github.com/umanlp/DEBIE.

General Debiasing Framework

In what follows, we first formalize two bias specifications – implicit and explicit. We then introduce new debiasing models: two operate on the implicit bias specification and the third on the explicit bias specification. Finally, we show how to debias word embeddings in a variety of target languages via cross-lingual embeddings.

An implicit bias specification BI=(T1,T2)B_{I}=(T_{1},T_{2}) consists of two sets of target terms with respect to which a bias is expected to exist in the embedding space. For example, two sets of science and art terms, T1T_{1} = {physics,chemistry,experiment}\{\textit{physics},\textit{chemistry},\textit{experiment}\} and T2={poetry,dance,drama}T_{2}=\{\textit{poetry},\textit{dance},\textit{drama}\} constitute an implicit specification of the gender bias. Strictly speaking, BIB_{I} does not specify a bias directly – it merely specifies two categories of concepts for which we implicitly assume that there exists some set of reference terms AA (e.g., male terms man, father and/or female terms like woman, girl) with respect to which T1T_{1} and T2T_{2} exhibit differences. Most existing debiasing models (?; ?; ?; ?) operate on BI=(T1,T2)B_{I}=(T_{1},T_{2}), i.e., not requiring reference terms AA.

An explicit bias specification BEB_{E} defines, in addition to sets T1T_{1} and T2T_{2}, one or more reference attribute sets. We consider an explicit bias specification with a single attribute set, BE=(T1,T2,A)B_{E}=(T_{1},T_{2},A) (as employed by our DebiasNet model),The attribute set AA can be any set of attributes towards which the bias is to be removed. In our experiments, we joined the WEAT test specification attribute sets A1A_{1} and A2A_{2}. and also with two (opposing) attribute sets, BE=(T1,T2,A1,A2)B_{E}=(T_{1},T_{2},A_{1},A_{2}), as used in WEAT tests (?).

The initial bias specification (BIB_{I} or BEB_{E}) commonly contains only a handful of words in each target and attribute set. These are commonly the most representative words of a category (e.g., man, boy, father to represent the category male). However, in order to provide a finer-grained bias specification, we propose to augment each term set with synonyms and semantically similar words of the initial terms. We therefore extract nearest neighbours of initial terms from an embedding space specialized to accentuate true semantic similarity and attenuate other types of semantic association (?; ?; ?, inter alia). For the augmentation process we rely on the recent state-of-the-art similarity specialization method of ? (?): for more details see the original work.

Given BIB_{I} or BEB_{E} and a similarity-specialized word vector space Xsim\mathbf{X_{sim}}, we augment each of the term sets in the specification by retrieving the top kk most (cosine-)similar terms from Xsim\mathbf{X_{sim}} for each of the initial terms.We discard nearest neighbors initially present in other sets of the same bias specification: e.g., if we retrieve an augmentation candidate woman for an initial T1T_{1} term man, woman will not be added to T1T_{1} if it exists in T2T_{2} (or in AA-s). Extending bias specification sets using a similarity-specialized word vector space – as opposed to a regular distributional space – reduces the noisy augmentation stemming from semantic relatedness instead of true semantic similarity.We also considered using clean lexical knowledge from WordNet (?) directly, but this resulted in much lower recall as well as less accurate augmentation candidates. Table 1 illustrates the initial bias specification and the corresponding augmentation (showing k=2k=2 nearest neighbors, without the initial terms) for one explicitly defined gender bias.

Debiasing Models

We present three debiasing models, two of which operate on BI=(T1,T2)B_{I}=(T_{1},T_{2}) and one on the explicit bias specification BE=(T1,T2,A)B_{E}=(T_{1},T_{2},A).

focuses on BIB_{I} as a generalization of the linear projection model proposed by ? (?), itself, in turn, an extension of the hard-debiasing model of ? (?).

where ⟨⋅,⋅⟩\langle\cdot,\cdot\rangle denotes a dot product. In other words, the closer the vector x\mathbf{x} is to the global bias direction b\mathbf{b}, the more it is bias-corrected (i.e., the larger portion of b\mathbf{b} is subtracted from x\mathbf{x}). Vectors orthogonal to the bias direction b\mathbf{b} remain unchanged (zero dot-product with the bias vector b\mathbf{b}).

Bias-Alignment Model (BAM).

Here, we use pairs (t1i,t2j)(t^{i}_{1},t^{j}_{2}) to learn the debiasing projection of X\bm{X} with respect to itself. Let XT1\bm{X}_{T_{1}} and XT2\bm{X}_{T_{2}} be the matrices obtained by stacking (biased) vectors of left and right words of pairs (t1i,t2j)(t^{i}_{1},t^{j}_{2}), respectively. We then learn the orthogonal map WX=UV⊤\bm{W_{X}}=\bm{U}\bm{V}^{\top}, where UΣV⊤\bm{U}\bm{\Sigma}\bm{V}^{\top} is the singular value decomposition of XT2XT1⊤\bm{X}_{T_{2}}\bm{X}_{T_{1}}^{\top}. Since WX\bm{W_{X}} is orthogonal, the projection X′=XWX\bm{X}^{\prime}=\bm{X}\bm{W_{X}} is isomorphic to the original space X\bm{X}, and thus equally biased. However, the transformation (specified by WX\bm{W_{X}}) defines the angle and direction of debiasing. We obtain the debiased space by averaging the original space X\bm{X} and the projected space X′\bm{X}^{\prime}:

Explicit Neural Debiasing (DebiasNet).

The final model, dubbed DebiasNet, is a neural model that operates on the explicit bias specification BEB_{E}. It is inspired by the work on semantic specialization of word embeddings (?; ?): but instead of using linguistic constraints (e.g., synonyms), we “specialize” the vector space by leveraging debiasing constraints.

Given a biased input space X\bm{X} and the specification BE=(T1,T2,A)B_{E}=(T_{1},T_{2},A), we learn a debiasing function DBN(X;θ)\text{DBN}(\bm{X};\mathbf{\theta}) that transforms X\bm{X} to a debiased space X′\bm{X}^{\prime}. We aim for the terms from both sets T1T_{1} and T2T_{2} to be similarly close to the terms from AA in X′\bm{X}^{\prime}. For simplicity, we execute DBN(X;θ)(\bm{X};\mathbf{\theta}) as a feed-forward neural network with non-linear activations. The training set for learning the parameters θ\theta consists of triples (t1∈T1,t2∈T2,a∈A)(t_{1}\in T_{1},t_{2}\in T_{2},a\in A). It is obtained as a full Cartesian product T1×T2×AT_{1}\times T_{2}\times A. Let t1\mathbf{t}_{1}, t2\mathbf{t}_{2} and a\mathbf{a} be the respective vectors of t1t_{1}, t2t_{2}, and aa from the input biased space X\bm{X}, and let t1′\mathbf{t}^{\prime}_{1}, t2′\mathbf{t}^{\prime}_{2} and a′\mathbf{a}^{\prime} be their “debiased” transformations: t1′=DBN(t1;θ)\mathbf{t}^{\prime}_{1}=\text{DBN}(\mathbf{t}_{1};\theta), t2′=DBN(t2;θ)\mathbf{t}^{\prime}_{2}=\text{DBN}(\mathbf{t}_{2};\theta), and a′=DBN(a;θ)\mathbf{a}^{\prime}=\text{DBN}(\mathbf{a};\theta). For a training instance (t1,t2,a)(t_{1},t_{2},a), we then minimize the following loss function LDL_{D}:

cos⁡d(⋅,⋅)\cos_{d}(\cdot,\cdot) refers to the cosine distance. The objective pushes the terms from the two target sets T1T_{1} and T2T_{2} to be equidistant to the terms from the attribute set AA. That is, it is designed to specifically remove the explicit bias. By minimizing LDL_{D} as the only objective, the model would remove the bias, but it would also destroy the useful semantic information in the input space. We thus couple the objective LDL_{D} with the regularization LRL_{R} that prevents the debiased vectors to deviate too much from their original estimates:

The final loss is then J=LD+λLRJ=L_{D}+\lambda L_{R}, with λ\lambda as the regularization weight. The learned function is then applied to the full input space: X′=DBN(X;θ)\bm{X}^{\prime}=\text{DBN}(\bm{X};\theta).

Composing Debiasing Models.

The presented models can be seamlessly composed with one another. For example, given an explicit specification BEB_{E}, we can first explicitly debias a distributional space X\mathbf{X} using DebiasNet. We can then apply either GBDD or BAM on the resulting vector space by deriving BIB_{I} from BEB_{E} (i.e., by considering only T1T_{1} and T2T_{2}): e.g., X′=GBDD(DBN(X))\bm{X}^{\prime}=\text{GBDD}(\text{DBN}(\bm{X})).

Cross-Lingual Transfer of Debiasing

Cross-lingual word embeddings have been shown to be a viable solution for zero-shot language transfer of NLP models (?; ?). Conceptually, given a source language L1L_{1} with its monolingual distributional space XL1\bm{X}_{L1} and a target language L2L_{2} with the space XL2\bm{X}_{L2}, we can apply any L1L1 model trained on XL1\bm{X}_{L1} on the instances from L2L2, given a matrix WCL\bm{W}_{CL} that projects XL2\bm{X}_{L2} to XL1\bm{X}_{L1}. From the plethora of cross-lingual word embedding models (?; ?; ?, inter alia), we opt for a supervised projection-based model (?) that obtains WCL\bm{W}_{CL} by solving the Procrustes problem (?) on the set of word translation pairs.Note that we obtain the cross-lingual projection WCL\bm{W}_{CL} in the similar way as debiasing projection WX\bm{W}_{X} in BAM; but now the aligned matrices contain vectors (each from respective language) corresponding to word translation pairs (not pairs created from bias target sets as in BAM). We select this approach due to its simplicity and competitive zero-shot language transfer performance on other NLP tasks (?). With the cross-lingual projection matrix WCL\bm{W}_{CL} in place, the debiasing of the space XL2\bm{X}_{L2} amounts to composing the projection with the debiasing model in L1L1: e.g., for GBDD, XL2′=GBDDL1(XL2WCL)\bm{X}^{\prime}_{L2}=\text{GBDD}_{L1}(\bm{X}_{L2}\bm{W}_{CL}).

Evaluation and Experimental Setup

We now introduce the metrics for testing different aspects of debiased embedding spaces, and then outline two datasets used in our experiments.

We use three diverse tests to measure the presence of explicit bias, and two tests that focus on the presence of implicit bias. Finally, we test the debiased spaces for their ability to preserve the initial semantic information.

Introduced by ? (?), WEAT tests the embedding space for the presence of an explicit bias defined as BEB_{E}=(T1,T2,A1,A2)(T_{1},T_{2},A_{1},A_{2}). It computes the differential association between T1T_{1} and T2T_{2} based on their mean similarity with terms from the attribute sets A1A_{1} and A2A_{2}:

The association ss of term t∈Tit\in T_{i} is computed as:

The significance of the statistic is computed by comparing s(BE)s(B_{E}) with the scores s(BE∗)s(B^{*}_{E}) obtained with all permutations BE∗=(T1∗,T2∗,A1,A2)B^{*}_{E}=(T^{*}_{1},T^{*}_{2},A_{1},A_{2}), where T1∗T^{*}_{1} and T2∗T^{*}_{2} are equally sized partitions of T1∪T2T_{1}\cup T_{2}. The pp-value of the test is the probability of s(BE∗)>s(BE)s(B^{*}_{E})>s(B_{E}). The “amount” of bias, the so-called effect size, is then a normalized measure of separation between association distributions:

where μ\mu is the mean and σ\sigma is the standard deviation.

Embedding Coherence Test (ECT).

It quantifies the amount of explicit bias BEB_{E}={T1,T2,A}\{T_{1},T_{2},A\} (?). Unlike WEAT, it compares vectors of target sets T1T_{1} and T2T_{2} (averaged over the constituent terms) with vectors from a single attribute set AA. ECT first computes the mean vectors for the target sets T1T_{1} and T2T_{2}: μ1=1∣T1∣∑t1∈T1t1\mathbf{\mu}_{1}=\frac{1}{|T_{1}|}\sum_{t_{1}\in T_{1}}{\mathbf{t}_{1}} and μ2=1∣T2∣∑t2∈T2t2\hskip 5.0pt\mathbf{\mu}_{2}=\frac{1}{|T_{2}|}\sum_{t_{2}\in T_{2}}{\mathbf{t}_{2}}. Next, for both μ1\mathbf{\mu}_{1} and μ1\mathbf{\mu}_{1} it computes the (cosine) similarities with vectors of all a∈A\mathbf{a}\in A. The two resultant vectors of similarity scores, s1\mathbf{s}_{1} (for T1T_{1}) and s2\mathbf{s}_{2} (for T2T_{2}) are used to obtain the final ECT score. It is the Spearman’s rank correlation between the rank orders of s1\mathbf{s}_{1} and s2\mathbf{s}_{2} – the higher the correlation, the lower the bias.

Bias Analogy Test (BAT).

Based on the observation of (?) that in a biased vector space programmer−homemaker≈man−woman\mathit{programmer}-\mathit{homemaker}\approx\mathit{man}-\mathit{woman}, ? (?) proposed an analogy-based bias test: Embedding Quality Test (EQT). However, EQT depends on WordNet to extend the bias definition with synonyms and plurals of bias specification terms. In contrast, we propose an alternative Bias Analogy Test (BAT) that relies only on the specification BE=(T1,T2,A1,A2)B_{E}=(T_{1},T_{2},A_{1},A_{2}).

We first create all possible biased analogies t1−t2≈a1−a2\mathbf{t}_{1}-\mathbf{t}_{2}\approx\mathbf{a}_{1}-\mathbf{a}_{2} for (t1,t2,a1,a2)∈T1×T2×A1×A2(t_{1},t_{2},a_{1},a_{2})\in T_{1}\times T_{2}\times A_{1}\times A_{2}. We then create two query vectors from each analogy: q1=t1−t2+a2\mathbf{q}_{1}=\mathbf{t}_{1}-\mathbf{t}_{2}+\mathbf{a}_{2} and q2=a1−t1+t2\mathbf{q}_{2}=\mathbf{a}_{1}-\mathbf{t}_{1}+\mathbf{t}_{2} for each 4-tuple (t1,t2,a1,a2)(t_{1},t_{2},a_{1},a_{2}). We then rank the vectors in the vector space X\bm{X} according to the Euclidean distance with each of the query vectors. In a biased space, we expect the vector a1\mathbf{a}_{1} to be ranked higher for the query q1\mathbf{q}_{1} than the vectors of terms from the opposing attribute set A2A_{2} (e.g., for a gender-biased space we expect woman to be ranked higher than father or boy for the query man - programmer + homemaker). Also, a2\mathbf{a}_{2} is expected to be more similar to q2\mathbf{q}_{2} than vectors of A1A_{1} terms . The BAT score is the percentage of cases where: (1) a1a_{1} is ranked higher than a term a2′∈A2∖{a2}a^{\prime}_{2}\in A_{2}\setminus\{a_{2}\} for q1\mathbf{q}_{1} and (2) a2a_{2} is ranked higher than a term a1′∈A1∖{a1}a^{\prime}_{1}\in A_{1}\setminus\{a_{1}\} for q2\mathbf{q}_{2}.

Implicit Bias Tests.

? (?) recently suggested that the two sets of target terms can still be clearly distinguished (with KMeans clustering, or in a supervised manner with an SVM classifier) from one another after applying debiasing procedures of (?) and (?). We adopt their approach and test the debiased spaces for the presence of implicit bias by clustering terms from T1T_{1} and T2T_{2} with KMeans++, and by classifying them using an SVM with the RBF kernel: it is trained on the vectors of terms from the augmentations of target sets. For each debiasing model, we average the clustering and classification scores over 2020 independent runs.

Semantic Quality.

Debiasing procedures change the topology of the input vector space; we thus have to verify that debiasing does not occur at the expense of the encoded semantic information. We test the debiased embedding spaces on two standard word similarity/relatedness benchmarks: SimLex-999 (?) and WordSim-353 (?).

Evaluation Datasets

Our proposed framework is versatile as it enables debiasing models to operate on any bias specified in the BIB_{I} or BEB_{E} format. To demonstrate this, we evaluate the debiasing models from the previous section on two different bias specifications: tests T1 and T8 from the WEAT dataset (?). WEAT tests are given as explicit bias specifications BEB_{E}= (T1,T2,A1,A2)(T_{1},T_{2},A_{1},A_{2}).

WEAT T8, shown in Table 1, encodes a type of a gender bias in relation to affinities towards science and art. T1T_{1} contains terms from the areas of science and technology, whereas T2T_{2} contains art terms. Attribute sets contain male (A1A_{1}) and female (A2A_{2}) terms. In a gender-biased vector space the scientific targets are expected to be more strongly associated with male attributes, and artistic targets with female terms.

WEAT 1: Flowers vs. Insects.

WEAT T1 specifies another bias type: the difference in sentiment humans attach to insects as opposed to flowers. Target sets contain different flowers (T1T_{1}) and insect species (T2T_{2}), and attribute sets contain universally positive (A1A_{1}) and negative (A2A_{2}) terms. The full bias specification of WEAT T1 is available in the supplementary.

XWEAT.

For evaluating the language transfer setup, we use bias specifications in target languages as our test data. We use tests T1 and T8 from XWEAT, created by ? (?) by translating the English (en) WEAT tests to six languages: German (de), Spanish (es), Italian (it), Russian (ru), Croatian (hr), and Turkish (tr).

Preprocessing and Training Setup

Augmented Bias Specifications. We first augment the bias specifications using a similarity-specialized embedding space produced by ? (?)Available at: https://tinyurl.com/y273cuvk. based on the en fastText embeddings (?). For WEAT T8, we augment the target and attribute lists with k=4k=4 nearest neighbours of each term. As the initial lists of WEAT T1 are longer than those of T8, we use k=2k=2 with T1. We train all debiasing models using bias specifications containing only the augmentation terms (i.e., without the initial bias specification terms); we use the initial terms for testing.

We test the robustness of debiasing models on three different word embedding models trained on Wikipedia: CBOW (?), GloVe (?), and fastText (FT) (?). For cross-lingual transfer, we induce a multilingual space spanning seven languages (en + 6 targets) by projecting FT vectors of each target to the EN space. Following an established procedure (?), we learn projections WCL\bm{W}_{CL} using automatically compiled translations of the 5K most frequent en words.

Training Setup.

For GBDD and BAM there is a deterministic closed-form solution for any given bias specification. On the other hand, the hyper-parameters of DebiasNet are optimized via grid search and cross-validation on the training set. The final DebiasNet model uses 55 hidden layers with 300300 units each and the weight λ\lambda is fixed to 0.2.

Results and Analysis

We first report debiasing results on three en distributional spaces, for the individual models as well as for three composite models: GBDD ∘\circ BAM = GBDD(BAM(X\bm{X})), BAM ∘\circ GBDD, and GBDD ∘\circ DebiasNet.BAM and DebiasNet display similar results and so does their composition. For brevity, we thus omit the scores of BAM ∘\circ DebiasNet. We also do not report the scores with DebiasNet ∘\circ GDBB as its scores were similar to its inverse composition GDBB ∘\circ DebiasNet in our preliminary tests. We then show the results for the cross-lingual debiasing transfer. Finally, we analyze the topology of debiased spaces.

Biases of Distributional Spaces. The main results are summarized in Table 2. All three input distributional spaces generally exhibit explicit and implicit biases, with CBOW spaces displaying the lowest biases, both according to the WEAT tests (e.g., the effect size is even insignificant with p<0.05p<0.05 for the gender bias test T8) and the implicit bias tests of ? (?). Interestingly – according to our BAT test, and despite the original claims and examples from ? (?) – the encoded biases do not reflect strongly in the analogy tests. Nonetheless, our debiasing methods in most test settings manage to affect the input vector spaces by further reducing BAT scores.

While the results vary across the two WEAT tests and evaluation metrics, GBDD emerges as the most robust model on average. It attenuates the explicit bias while being the most successful in removing the bias implicitly: the spaces debiased with GBDD completely confuse the KM clustering and SVM classifier. It also fully retains the useful semantic information: we do not observe drops on SL and WS compared to the input distributional spaces. While GBDD outperforms BAM and DebiasNet (DBN) on average according to ECT and BAT measures, it is not able to fully remove the explicit gender bias (T8) according to the WEAT test.

Despite operating on an implicit specification BIB_{I}, BAM removes the explicit biases much better than the implicit ones. DBN seems even better than BAM in removing the explicit biases. This is not a surprise, since DBN is trained on an explicit bias specification. However both DBN and BAM are unsuccessful in removing the implicit biases. Moreover, DBN distorts the input space more than BAM, yielding substantial drops on SL and WS.

The complementarity of debiasing effects between GBDD, and BAM/DBN are confirmed by the performance of their compositions. All composition models robustly remove both explicit and implicit biases, also showing that there is no “one model rules them all” solution to various debiasing aspects. GBDD ∘\circ DBN most effectively removes the biases, but it inherits the undesirable semantic distortions of DBN. On the other hand, BAM ∘\circ GBDD offers solid bias removal while for the most part retaining the semantic quality of the space.

Differences between Evaluation Measures.

The three aspects of evaluation complement each other: they all inform the selection of the most appropriate debiasing model w.r.t. the desired application-specific criteria.E.g., Note that for some bias specifications, one might not want to reduce/remove the implicit bias. WEAT T1 can be seen as an example of such bias: while we may want to make insects similarly good/bad as flowers, we do not want to make them indistinguishable from flowers in the vector space. However, results of WEAT, ECT, and BAT are not always aligned. For example, the CBOW space is unbiased according to the WEAT test, but extremely biased (negative correlation!) according to ECT. In contrast, GloVe vectors are biased according to WEAT but not according to ECT (correlation of 0.840.84). These findings point to different bias aspects, accentuating the need for multiple, mutually complementary, bias measures.

Cross-Lingual Transfer

The results in the cross-lingual debiasing transfer are shown in Table 3. For brevity, we show only the results on XWEAT T8 (gender bias wrt. science vs. art) and for a subset of evaluation measures (one for each evaluation aspect): WEAT (W), KMeans++ (KM), and SimLex (SL).We provide the full results, with all evaluation measures, and also on the XWEAT T1 test in the supplemental material.,We evaluate word similarities for de, it, ru, and hr on their respective SimLex datasets (?; ?); there is no es and tr SimLex.

We first confirm the results from ? (?): de, it, ru, and hr fastText vectors do not exhibit significant explicit gender bias (wrt. science vs. art), according to the WEAT test. The explicit bias is, however, significant in es and tr distributional vectors. Implicit bias is clearly present in all distributional spaces except ru. Debiasing models display similar properties as before: DBN reduces the explicit bias more effectively than BAM and GBDD, but it semantically distorts the vectors; and only GBDD successfully removes the implicit bias. None of the models fully removes the explicit bias for tr (the lowest bias effect of 0.990.99 for BAM ∘\circ GBDD is still significant). We suspect that this is a result of the lower-quality cross-lingual tr→\rightarrowen projection, which is in line with the bilingual lexicon induction results from ? (?).

For de and it, BAM and DBN invert the direction of the bias: negative WEAT scores mean that sciences are more correlated with female attributes and arts with male attributes. We believe that this is the result of applying a (strong) bias correction learned on a biased en space on the (explicitly) unbiased de and it spaces. The BAM ∘\circ GBDD composition seems most robust in the cross-lingual transfer setting – it successfully removes both the explicit (if they exist) and implicit biases, while preserving useful semantic information (SL). These results indicate that we can attenuate or remove biases in distributional vectors of languages for which (1) we do not require the initial bias specification and (2) we do not even need similarity-specialized word embeddings used to augment the bias specifications for the target language.

Finally, we qualitatively analyze the debiasing effects suggested by evaluation measures. We project the input and the debiased embeddings into 2D with PCA, and show the constellation of words from the initial bias specification of WEAT T8 (Table 1) in Figure 1.We show only the input space and the spaces debiased with GBDD and BAM. We provide similar illustrations for other debiasing models in the supplementary material. In the distributional space, the two target sets (science vs art) are clearly distinguishable from one another (implicit bias), and so are the male and female attributes. The science terms are notably closer to the male terms and art terms to the female terms (explicit bias). The space produced by BAM intertwines the male and female terms and makes the science and art terms roughly equidistant to the gender terms (explicit bias removed), but the science terms are still clearly distinguishable from art terms (implicit bias still present). In the space produced by GBDD, both biases are removed: science and art terms cannot be clearly separated and are roughly equidistant to gender terms.

Conclusion

We have introduced a general framework for debiasing distributional word vector spaces by 1) formalizing the differences between implicit and explicit biases, 2) proposing new debiasing methods that deal with the two different bias specifications, and 3) designing a comprehensive evaluation framework for testing the (often complementary) effects of debiasing. While the proposed framework offers a systematized view on human biases encoded in word embeddings, the main results indicate that our debiasing methods can effectively attenuate biases in arbitrary input distributional spaces and can also be transferred to a variety of target languages.

References