Sustainable Modular Debiasing of Language Models
Anne Lauscher, Tobias Lüken, Goran Glavaš
Introduction
Recent work has shown that pretrained language models such as ELMo (Peters et al., 2018), BERT (Devlin et al., 2019), or GPT-2 (Radford et al., 2019) tend to exhibit a range of stereotypical societal biases, such as racism and sexism (e.g., Kurita et al., 2019; Dev et al., 2020; Webster et al., 2020; Nangia et al., 2020; Barikeri et al., 2021, inter alia). The reason for this lies in the distributional nature of these models: human-produced corpora on which these models are trained are abundant with stereotypically biased concept co-occurrences (for instance, male terms like man or son appear more often together with certain career terms like doctor or programmer than female terms like women or daughter) and the PLMs models, being trained with language modeling objectives, consequently encode these biased associations in their parameters. While this effect can lend itself to diachronic analysis of societal biases (e.g., Garg et al., 2018; Walter et al., 2021), it represents stereotyping, one of the main types of representational harm (Blodgett et al., 2020) and, if unmitigated, may cause severe ethical issues in various sociotechnical deployment scenarios.
To alleviate this problem and ensure fair language technology, previous work introduced a wide range of bias mitigation methods (e.g., Bordia and Bowman, 2019; Dev et al., 2020; Lauscher et al., 2020a, inter alia). All existing debiasing approaches, however, modify all parameters of the PLMs which has two prominent shortcomings: (1) it comes with a high computational costWhile a full fine-tuning approach to PLM debiasing may still be feasible for moderate-sized PLMs like BERT Devlin et al. (2019), it is prohibitively computationally expensive for giant language models like GPT-3 Brown et al. (2020) or GShard Lepikhin et al. (2020). and (2) can lead to (catastrophic) forgetting (McCloskey and Cohen, 1989; Kirkpatrick et al., 2017) of the useful distributional knowledge obtained during pretraining. For example, Webster et al. (2020) incorporate counterfactual debiasing already into BERT’s pretraining: this implies a debiasing framework in which a separate “debiased BERT” instance needs to be trained from scratch for each individual bias type and specification. In sum, current debiasing procedures designed for pretraining or full fine-tuning of PLMs have a large carbon footprint Strubell et al. (2019) and consequently jeopardize the sustainability Moosavi et al. (2020) of fair representation learning in NLP.
In this work, we move towards more sustainable removal of stereotypical societal biases from pretrained language models. To this end, we propose Adele (Adapter-based DEbiasing of LanguagE Models), a debiasing approach based on the the recently proposed modular adapter framework (Houlsby et al., 2019; Pfeiffer et al., 2020a). In Adele, we inject additional parameters, the so-called adapter layers into the layers of the PLM and incorporate the “debiasing” knowledge only in those parameters, without changing the pretrained knowledge in the PLM. We show that, while being substantially more efficient (i.e., sustainable) than existing state-of-the-art debiasing approaches, Adele is just as effective in bias attenuation.
The contributions of this work are three-fold: (i) we first present Adele, our novel adapter-based framework for parameter-efficient and knowledge-preserving debiasing of PLMs. We combine Adele with one of the most effective debiasing strategies, Counterfactual Data Augmentation (CDA; Zhao et al., 2018), and demonstrate its effectiveness in gender-debiasing of BERT Devlin et al. (2019), the most widely used PLM. (ii) We benchmark Adele in what is arguably the most comprehensive set of bias measures and data sets for both intrinsic and extrinsic evaluation of biases in representation spaces spanned by PLMs. Additionally, we study a previously neglected effect of fairness forgetting present when debiased PLMs are subjected to large-scale downstream training for specific tasks (e.g., natural language inference, NLI); we show that Adele’s modular nature allows to counter this undesirable effect by stacking a dedicated task adapter on top of the debiasing adapter. (iii) Finally, we successfully transfer Adele’s debiasing effects to six other languages in a zero-shot manner, i.e., without relying on any debiasing data in the target languages. We achieve this by training the debiasing adapter stacked on top of the multilingual BERT on the English counterfactually augmented dataset.
Adele: Adapter-Based Debiasing
In this work, we seek to fulfill the following three desiderata: (1) we want to achieve effective debiasing, comparable to that of existing state-of-the-art debiasing methods while (2) keeping the training costs of debiasing significantly lower; and (3) fully preserving the distributional knowledge acquired in the pretraining. To meet all three criteria, we propose debiasing based on the popular adapter modules (Houlsby et al., 2019; Pfeiffer et al., 2020a). Adapters are lightweight neural components designed for parameter-efficient fine-tuning of PLMs, injected into the PLM layers. In downstream fine-tuning, all original PLM parameters are kept frozen and only the adapters are trained. Because adapters have fewer parameters than the original PLM, adapter-based fine-tuning is more computationally efficient. And since fine-tuning does not update the PLM’s original parameters, all distributional knowledge is preserved.
The debiasing adapters could, in principle, be trained using any of the debiasing strategies and training objectives from the literature, e.g., via additional debiasing loss objectives Qian et al. (2019, inter alia); Bordia and Bowman (2019, inter alia); Lauscher et al. (2020a, inter alia) or data-driven approaches such as Counterfactual Data Augmentation (Zhao et al., 2018). For simplicity, we opt for the data-driven CDA approach: it has been shown to offer reliable debiasing performance Zhao et al. (2018); Webster et al. (2020) and, unlike other approaches, it does not require any modifications of the model architecture nor training procedure.
In this work, we employ the simple adapter architecture proposed by Pfeiffer et al. (2021), in which only one adapter module is added to each layer of the pretrained Transformer, after the feed-forward sub-layer. The more widely used architecture of Houlsby et al. (2019) inserts two adapter modules per Transformer layer, with the other adapter injected after the multi-head attention sublayer. We opt for the “Pfeiffer architecture” because in comparison with the “Houlsby architecture” it is more parameter-efficient and has been shown to yield slightly better performance on a wide range of downstream NLP tasks Pfeiffer et al. (2020a, 2021). The output of the adapter, a two-layer feed-forward network, is computed as follows:
In our case, we train the adapters for debiasing: we inject adapter layers into BERT Devlin et al. (2019), freeze the original BERT’s parameters, and run a standard debiasing training procedure – language modeling on counterfactual data (§2.2) – during which we only tune the parameters of the debiasing adapters. At the end of the debiasing training, the debiasing functionality is isolated into the adapter parameters. This not only preserves the distributional knowledge in the Transformer’s original parameters, but also allows for more flexibility and “on-demand” usage of the debiasing functionality in downstream applications. For example, one could train a separate set of debiasing adapters for each bias dimension of interest (e.g., gender, race, religion, sexual orientation) and selectively combine them in downstream tasks, depending on the constraints and requirements of the concrete sociotechnical environment.
2 Counterfactual Augmentation Training
In the context of representation debiasing, counterfactual data augmentation (CDA) refers to the automatic creation of text instances that in some way counter the stereotypical bias present in the representation space. CDA has been successfully used for attenuating a variety of bias types, e.g., gender and race, and in several variants, e.g., with general terms describing dominant and minoritized groups, or with personal names acting as proxies for such groups Zhao et al. (2018); Lu et al. (2020). Most commonly, CDA modifies the training data by replacing terms describing one of the target groups (dominant or minoritized) with terms describing the other group. Let be our training corpus, consisting of sentences and let be a set of term pairings between the dominant and minoritized group (i.e., is a term representing the dominant group, e.g., man, and is a corresponding term representing the minoritized group, e.g., woman). For each sentence and each pair , we check whether either or occur in : if is present, we replace its occurrence with and vice versa. We denote the counterfactual sentence of obtained this way with and the whole counterfactual corpus with . We adopt the so-called two-sided CDA from (Webster et al., 2020): the final corpus for debiasing training consists of both the original and counterfactually created sentences. Finally, we train the debiasing adapter via masked language modeling on the counterfactually augmented corpus . We train sequentially by first exposing the adapter to the original corpus and then to the augmented portion .
Experiments
We showcase Adele for arguably the most explored societal bias – gender bias – and the most widely used PLM, BERT. We profile its debiasing effects with a comprehensive set of intrinsic and downstream (i.e., extrinsic) evaluations.
We test Adele on three intrinsic (BEC-Pro, DisCo, WEAT) and two downstream debiasing benchmarks (Bias-STS-B and Bias-NLI). We now describe each of the benchmarks in more detail.
We intrinsically evaluate Adele on the BEC-Pro data set (Bartl et al., 2020), designed to capture gender bias w.r.t. professions. The data set consists of 2,700 sentence pairs in the format (“m [temp] p”; “f [temp] p”), where m is a male term (e.g., boy, groom), f is a female term (e.g., girl, bride), p is a profession term (e.g., mechanic, doctor), and [temp] is one of the predefined connecting templates, e.g., “is a” or “works as a”.
We measure the bias on BEC-Pro using the bias measure of Kurita et al. (2019). They compute the association between a gender term (male or female) and a profession as:
where is the probability of the PLM generating the target term when only itself is masked, and is the probability of being generated when both and the profession are masked. The bias score is then simply a difference in the association score between the male term and its corresponding female term : . We measure the overall bias on the whole dataset in two complementary ways: (a) by averaging the bias scores across all 2,700 instances ( bias) and (b) by measuring the percentage of instances for which is below some threshold value: we report this score for two different thresholds ( and ).
Bartl et al. (2020) additionally published a German version of the BEC-Pro data set, which we use to evaluate Adele’s zero-shot transfer abilities.
Discovery of Correlations (DisCo).
The second data set for intrinsic debiasing evaluation, DisCo (Webster et al., 2020), also relies on templates (e.g., “[PERSON] studied [BLANK] at college”). For each template, the [PERSON] slot is filled first with a male and then with a female term (e.g., for the pair (John, Veronica), we get John studied [BLANK] at college and Veronica studied [BLANK] at college). Next, for each of the two instances, the model is asked to fill the [BLANK] slot: the goal is to determine the difference in the probability distribution for the masked token, depending on which term is inserted in the [PERSON] slot. While Webster et al. (2020) retrieve the top three most likely terms for the masked position, we retrieve all terms t with the probability .We argue that retrieving more terms from the distribution allows for a more accurate estimate of the bias.
Let and be the candidate sets obtained for the -th instance when filled with a male [PERSON] term and the corresponding female term , respectively. We then compute two different measures. The first is the average fraction of shared candidates between the two sets ():
with as the total number of test instances. Intuitively, a higher average fraction of shared candidates indicates lower bias.
For the second measure, we retrieve the probabilities for all candidates in the union of two sets . We then compute the normalized average absolute probability difference:
We create test instances by collecting most frequent baby names for each gender from the US Social Security name statistics for 2019.https://www.ssa.gov/oact/babynames/limits.html We create pairs (, ) from names at the same frequency rank in the two lists (e.g., Liam and Olivia). Finally, we remove pairs with ambiguous names that may also be used as general concepts (e.g., violet, a color), resulting in final pairs.
Word Embedding Association Test (WEAT).
As the final intrinsic measure, we use the well-known WEAT (Caliskan et al., 2017) test. Developed for detecting biases in static word embedding spaces, it computes the differential association between two target term sets (e.g., male terms) and (e.g., female terms) based on the mean (cosine) similarity of their embeddings with embeddings of terms from two attribute sets (e.g., science terms) and (e.g., art terms):
The association of term or is computed as:
The significance of the statistic is computed with a permutation test in which is compared with the scores where and are equally sized partitions of . We report the effect size, a normalized measure of separation between the association distributions:
where is the mean and is the standard deviation.
Lauscher and Glavaš (2019) created XWEAT by translating some of the original WEAT bias specifications to six target languages: German (de), Spanish (es), Italian (it), Croatian (hr), Russian (ru), and Turkish (tr). We use their translations of the WEAT 7 gender test in the zero-shot debiasing transfer evaluation of adele.
Bias-STS-B.
The first extrinsic measure we use is Bias-STS-B, introduced by Webster et al. (2020), based on the well-known Semantic Textual Similarity-Benchmark (STS-B; Cer et al., 2017), a regression task where models need to predict semantic similarity for pairs of sentences. Webster et al. (2020) adapt STS-B for discovering gender-biased correlations. They start from neutral STS templates and fill them with a gendered term (man, woman) and a profession term from (Rudinger et al., 2018) (e.g., A man is walking vs. A nurse is walking and A woman is walking vs. A nurse is walking). The dataset consists of 16,980 such pairs. As a measure of bias, we compute the average absolute difference between the similarity scores of male and female sentence pairs, with a lower value corresponding to less bias. We couple the bias score with the actual STS task performance score (Pearson correlation with human similarity scores), measured on the STS-B development set.
Bias-NLI.
We select the task of understanding biased natural language inferences (NLI) as the second extrinsic evaluation. To this end, we fine-tune the original BERT as well as our adapter-debiased BERT on the MNLI data set (Williams et al., 2018). For evaluation, we follow Dev et al. (2020), and create a synthetic NLI data set that tests for the gender-occupation bias: it comprises NLI instances for which an unbiased model should not be able to infer anything, i.e., it should predict the neutral class. We use the code of Dev et al. (2020) and, starting from the generic template The
2 Experimental Setup
Aligned with BERT’s pretraining, we carry out the debiasing MLM training on the concatenation of the English Wikipedia and the BookCorpus (Zhu et al., 2015). Since we are only training the parameters of the debiasing adapters, we uniformly subsample the corpus to one third of its original size. We adopt the set of gender term pairs for CDA from Zhao et al. (2018) (e.g., actor-actress, bride-groom)https://github.com/uclanlp/corefBias/tree/master/WinoBias/wino and augment it with three additional pairs: his-her, himself-herself, and male-female, resulting with the total of term pairs. Our final debiasing CDA corpus consists of 105,306,803 sentences.
Models and Baselines.
In all experiments we inject Adele adapters of bottleneck size into the pretrained BERT Base Transformer ( layers, attention heads, hidden size).We implement Adele using the Huggingface tranformers library (Wolf et al., 2020) in combination with the AdapterHub framework (Pfeiffer et al., 2020a). We compare Adele with the debiased BERT Large models released by Webster et al. (2020): (1) ZariCDA is counterfactually pretrained (from scratch); whereas (2) ZariDO was post-hoc MLM-fine-tuned on regular corpora, but with more aggressive dropout rates. In cross-lingual zero-shot transfer experiments, we train Adele on top of multilingual BERT (Devlin et al., 2019) in its base configuration (uncased, layers, hidden size).
Debiasing Training.
We follow the standard MLM procedure for BERT training and mask 15% of the tokens. We then train Adele’s debiasing adapters on our CDA data set for epochs, with a batch size of . We optimize the adapter parameters using the Adam algorithm (Kingma and Ba, 2015), with the constant learning rate of .
Downstream Fine-tuning.
Our two extrinsic evaluations require task-specific fine-tuning on the STS-B and MNLI training datasets, respectively. We couple BERT (with and without Adele adapters) with the standard single-layer feed-forward softmax classifier and fine-tune all parameters in task-specific training.The only exception is the fairness forgetting experiment in §4, in which we freeze both the Transformer and the debiasing adapters and train the dedicated task adapter on top. We optimize the hyperparameters on the respective STS-B and MNLI (matched) development sets. To this end, we search for the optimal number of training epochs in and fix the learning rate to , maximum sequence length to , and batch size to . Like in debiasing training, we use Adam Kingma and Ba (2015) for optimization.
Results and Discussion
Our main monolingual English debiasing results on three intrinsic and two extrinsic benchmarks are summarized in Table 1. The results show that (1) Adele successfully attenuates BERT’s gender bias across the board, and (2) it is, in many cases, more effective in attenuating gender biases than the computationally much more intensive Zari models Webster et al. (2020). In fact, on BEC-Pro and DisCo Adele substantially outperforms both Zari variants.
The results from two extrinsic evaluations – STS and NLI – demonstrate that Adele successfully attenuates the bias, while retaining the high task performance. Zari variants yield slightly better task performance for both STS-B and MNLI: this is expected, as they are instances of the BERT Large Transformer with 336M parameters; in comparison, Adele has only 110M parameters of BERT Base and approx. 885K adapter parameters.Adele adds 884,736 parameters to BERT Base: 12 (layers) 2 (down-projection and up-projection matrix) 768 (hidden size of BERT Base) 48 (bottleneck size ).
According to WEAT evaluation on static embeddings extracted from BERT (§3.1), the original BERT Transformer is only slightly and insignificantly biased. Consequently, Adele inverts the bias in the opposite direction. In Figure 1, we further analyze the WEAT bias effects w.r.t. the subset of BERT layers from which we aggregate the word embeddings. For the original BERT (Figure 1(a)), we obtain the gender unbiased embeddings if we aggregate representations from higher layers (e.g., [5:12], [6:9], or by taking final layer vectors, [12:12]). For Adele, we get the most gender-neutral embeddings by aggregating representations from lower layers (e.g., [0:3] or [1:3]); representations from higher layers (e.g., [6:12]) flip the bias into the opposite direction (blue color). Both Zari models produce embeddings which are relatively unbiased, but ZariCDA still exhibits slight gender bias in higher layer representations. The dropout-based debiasing of ZariDO results in an interesting per-layer-region oscillating gender bias.
Zero-Shot Cross-Lingual Transfer.
We show the results of zero-shot transfer of gender debiasing with Adele (on top of mBERT) on German BEC-Pro in Table 2. On the En BEC-Pro portion Adele is as effective on top of mBERT as it is on top of the En BERT (see Table 1): it reduces mBERT’s bias from to . More importantly, the positive debiasing effect successfully transfers to German: the bias effect on the de portion is reduced from to , despite not using any German data in the training of debiasing adapters. We also see an improvement with respect to the fraction of unbiased instances for both thresholds, expectedly with larger improvements for the more lenient threshold of .
In Table 3, we show the bias effects of static word embeddings, aggregated from layers of mBERT and Adele-debiased mBERT, on the XWEAT gender-bias test 7 for six different target languages. We show the results for two aggregation strategies, including ([0:12]) and excluding ([1:12]) mBERT’s (sub)word embedding layer.
Like BEC-Pro, WEAT confirms that Adele also attenuates the bias in En representations coming from mBERT. The results across the six target languages are somewhat mixed, but overall encouraging: for all significantly biased combinations of languages and layer aggregations from original mBERT ([0:12] – it, ru; [1:12] – hr, ru), Adele successfully reduces the bias. E.g., for it embeddings extracted from all layers ([0:12]), the bias effect size drops from significant to insignificant . In case of already insignificant biases in original mBERT, Adele often further reduces the bias effect size (de, tr) and if not, the bias effects remain insignificant.
We additionally visualize all XWEAT bias effect sizes in the produced embeddings via heatmaps in Figure 2. The intuition we can get from the plots supports our conclusion: for all languages, especially for the source language en and the target language de, the bias gets reduced, which is indicated by the lighter colors throughout all plots.
Fairness Forgetting.
Finally, we investigate whether the debiasing effects persist even after the large-scale fine-tuning in downstream tasks. Webster et al. (2020) report the presence of debiasing effects after STS-B training. With merely 5,749 training instances, however, STS-B is two orders of magnitude smaller than MNLI (392,702 training instances). Here we conduct a study on MNLI, testing for the presence of the gender bias in Bias-NLI after Adele’s exposure to varying amount of MNLI training data. We fully fine-tune BERT Base and BERT (i.e., BERT augmented with debiasing adapters) on MNLI datasets of varying sizes (10K, 25K, 75K, 100K, 150K, and 200K) and measure, for each model, the Bias-NLI net neutral (NN) score as well as the NLI accuracy on the MNLI (matched) development set. For each model and each training set size, we carry out five training runs and report the average scores.
Figure 3 summarizes the results of our fairness forgetting experiment. We report the mean and the 95% confidence interval over the five runs for NN on Bias-NLI and Accuracy (Acc) on the MNLI-m development set. Several interesting observations emerge. First, the NN scores seem to be quite unstable across different runs (wide confidence intervals) for both BERT and Adele, which is surprising given the size of the Bias-NLI test set (1,936,512 instances). This could point to the lack of robustness of the NN measure (Dev et al., 2020) as means for capturing biases in fine-tuned Transformers. Second, after training on smaller datasets (10K), Adele still retains much of its debiasing effect and is much fairer than BERT. With larger NLI training (already at 25K), however, much of its debiasing effect vanishes, although it still seems to be slightly (but consistently) fairer than BERT over time. We dub this effect fairness forgetting and will investigate it further in future work.
Preventing Fairness Forgetting.
Finally, we propose a downstream fine-tuning strategy that can prevent fairness forgetting and which is aligned with the modular debiasing nature of Adele: we (1) inject an additional task-specific adapter (TA) on top of Adele’s debiasing adapter and (2) update only the TA parameters in downstream (MNLI) training. This way, the debiasing knowledge stored in Adele’s debiasing adapters remains intact. Table 4 compares Bias-NLI and MNLI performance of this fairness preserving variant (Adele-TA) against BERT and Adele.
Results strongly suggest that by freezing the debiasing adapters and injecting the additional task adapters, we indeed retain most of the debiasing effects of Adele: according to bias measures, Adele-TA is massively fairer than the fully fine-tuned Adele (e.g., FN score of vs. Adele’s ). Preventing fairness forgetting comes at a tolerable task performance cost: Adele-TA loses 3 points in NLI accuracy compared to fully fine-tuning BERT and Adele for the task.
Related Work
We provide a brief overview of work in two areas which we bridge in this work: debiasing methods and parameter efficient fine-tuning with adapters.
Adapters (Rebuffi et al., 2018) have been introduced to NLP by Houlsby et al. (2019), who demonstrated their effectiveness and efficiency for general language understanding (NLU). Since then, they have been employed for various purposes: apart from NLU, task adapters have been explored for natural language generation (Lin et al., 2020) and machine translation quality estimation (Yang et al., 2020). Other works use language adapters encoding language-specific knowledge, e.g., for machine translation (Philip et al., 2020; Kim et al., 2019) or multilingual parsing (Üstün et al., 2020). Further, adapters have been shown useful in domain adaptation (Pham et al., 2020; Glavaš et al., 2021) and for injection of external knowlege (Wang et al., 2020; Lauscher et al., 2020b). Pfeiffer et al. (2020b) use adapters to learn both language and task representations. Building on top of this, Vidoni et al. (2020) prevent adapters from learning redundant information by introducing orthogonality constraints.
Debiasing Methods.
A recent survey covering research on stereotypical biases in NLP is provided by Blodgett et al. (2020). In the following, we focus on approaches for mitigating biases from PLMs, which are largely inspired by debiasing for static word embeddings (e.g., Bolukbasi et al., 2016; Dev and Phillips, 2019; Lauscher et al., 2020a; Karve et al., 2019, inter alia). While several works propose projection-based debiasing for PLMs (e.g., Dev et al., 2020; Liang et al., 2020; Kaneko and Bollegala, 2021), most of the debiasing approaches require training. Here, some methods rely on debiasing objectives (e.g., Qian et al., 2019; Bordia and Bowman, 2019). In contrast, the debiasing approach we employ in this work, CDA (Zhao et al., 2018), relies on adapting the input data and is more generally applicable. Variants of CDA exist, e.g., Hall Maudslay et al. (2019) use names as bias proxies and substitute instances instead of augmenting the data, whereas Zhao et al. (2019) use CDA at test time to neutralize the models’ biased predictions. Webster et al. (2020) investigate one-sided vs. two-sided CDA for debiasing BERT in pretraining and show dropout to be effective for bias mitigation.
Conclusion
We presented Adele, a novel sustainable and modular approach to debiasing PLMs based on the adapter modules. In contrast to existing computationally demanding debiasing approaches, which debias the entire PLM via full fine-tuning, Adele performs parameter-efficient debiasing by training dedicated debiasing adapters. We extensively evaluated Adele on gender debiasing of BERT, demonstrating its effectiveness on three intrinsic and two extrinsic debiasing benchmarks. Further, applying Adele on top of mBERT, we successfully transfered its debiasing effects to six target languages. Finally, we showed that by combining Adele’s debiasing adapters with task-adapters, we can preserve the representational fairness even after large-scale downstream training. We hope that Adele catalyzes more research efforts towards making fair NLP fairer, i.e., more sustainable and more inclusive (i.e., more multilingual).
Acknowledgments
The work of Anne Lauscher and Goran Glavaš has been supported by the Multi2ConvAI Grant (Mehrsprachige und Domänen-übergreifende Conversational AI) of the Baden-Württemberg Ministry of Economy, Labor, and Housing (KI-Innovation). Additionally, Anne Lauscher has partially received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program (grant agreement No. 949944, INTEGRATOR).
Further Ethical Considerations
In this work, we employed a binary conceptualization of gender due to the plethora of existing bias evaluation tests that are restricted to such a narrow notion of gender available. Our work is of methodological nature (i.e., we do not create additional data sets and text resources), and our primary goal was to demonstrate the bias attenuation effectiveness of our approach based on debiasing adapters: to this end, we relied on the available evaluation data sets from previous work. We fully acknowledge that gender is a spectrum: we fully support the inclusion of all gender identities (nonbinary, gender fluid, polygender, and other) in language technologies and strongly support work on creating resources and data sets for measuring and attenuating harmful stereotypical biases expressed towards all gender identities. Further, we acknowledge the importance of research on the intersectionality (Crenshaw, 1989) of stereotyping, which we did not consider here for similar reasons – lack of training and evaluation data. Our modular adapter-based debiasing approach, Adele, however, is conceptually particularly suitable for addressing complex intersectional biases, and this is something we intend to explore in our future work.
References
Appendix A Code Base
We provide further information and links to all frameworks, code bases, and model checkpoints used in this work in Table 5.
Appendix B Word Pairs
We list all word pairs we employ in our study.
(liam, olivia), (noah, emma), (oliver, ava), (william, sophia), (elijah, isabella), (james, charlotte), (benjamin, amelia), (lucas, mia), (mason, harper), (alexander, abigail), (henry, emily), (jacob, ella), (michael, elizabeth), (daniel, camila), (logan, luna), (jackson, sofia), (sebastian, avery), (jack, mila), (aiden, aria), (owen, scarlett), (samuel, penelope), (matthew, layla), (joseph, chloe), (levi, victoria), (mateo, madison), (david, eleanor), (john, grace), (wyatt, nora), (carter, riley), (julian, zoey), (luke, hannah), (grayson, hazel), (isaac, lily), (jayden, ellie), (gabriel, lillian), (anthony, zoe), (dylan, stella), (leo, aurora), (lincoln, natalie), (jaxon, emilia), (asher, everly), (christopher, leah), (josiah, aubrey), (andrew, willow), (thomas, addison), (joshua, lucy), (ezra, audrey), (hudson, bella), (charles, nova), (isaiah, paisley), (nathan, claire), (adrian, skylar), (christian, isla), (maverick, genesis), (colton, naomi), (elias, elena), (aaron, caroline), (eli, eliana), (landon, anna), (nolan, valentina), (cameron, kennedy), (connor, ivy), (jeremiah, aaliyah), (ezekiel, cora), (easton, kinsley), (miles, hailey), (robert, gabriella), (jameson, allison), (nicholas, gianna), (greyson, serenity), (cooper, samantha), (ian, sarah), (axel, quinn), (jaxson, eva), (dominic, piper), (leonardo, sophie), (luca, sadie), (jordan, josephine), (adam, nevaeh), (xavier, adeline), (jose, arya), (jace, emery), (everett, lydia), (declan, clara), (evan, vivian), (kayden, madeline), (parker, peyton), (wesley, julia), (kai, rylee), (ryan, serena), (jonathan, mandy), (ronald, alice)
General Noun Pairs (Zhao et al., 2018).
(actor, actress), (actors, actresses) (airman, airwoman), (airmen, airwomen), (aunt, uncle), (aunts, uncles) (boy, girl), (boys, girls), (bride, groom), (brides, grooms), (brother, sister), (brothers, sisters), (businessman, businesswoman), (businessmen, businesswomen), (chairman, chairwoman), (chairmen, chairwomen), (chairwomen, chairman) (chick, dude), (chicks, dudes), (dad, mom ), (dads, moms), (daddy, mommy), (daddies, mommies), (daughter, son), (daughters, sons), (father, mother), (fathers, mothers), (female, male), (females, males), (gal, guy), (gals, guys), (granddaughter, grandson), (granddaughters, grandsons), (guy, girl), (guys, girls), (he, she), (herself, himself), (him, her), (his, her), (husband, wife), (husbands, wives), (king, queen ), (kings, queens), (ladies, gentlemen), (lady, gentleman), (lord, lady), (lords, ladies) (ma’am, sir), (man, woman), (men, women), (miss, sir), (mr., mrs.), (ms., mr.), (policeman, policewoman), (prince, princess), (princes, princesses), (spokesman, spokeswoman), (spokesmen, spokeswomen)
Extra Word List (Zhao et al., 2018).
(cowboy, cowgirl), (cowboys, cowgirls), (camerawomen, cameramen), (cameraman, camerawoman), (busboy, busgirl), (busboys, busgirls), (bellboy, bellgirl), (bellboys, bellgirls), (barman, barwoman), (barmen, barwomen), (tailor, seamstress), (tailors, seamstress’), (prince, princess), (princes, princesses), (governor, governess), (governors, governesses), (adultor, adultress), (adultors, adultresses), (god, godess), (gods, godesses), (host, hostess), (hosts, hostesses), (abbot, abbess), (abbots, abbesses), (actor, actress), (actors, actresses), (bachelor, spinster), (bachelors, spinsters), (baron, baroness), (barons, barnoesses), (beau, belle), (beaus, belles), (bridegroom, bride), (bridegrooms, brides), (brother, sister), (brothers, sisters), (duke, duchess), (dukes, duchesses), (emperor, empress), (emperors, empresses), (enchanter, enchantress), (father, mother), (fathers, mothers), (fiance, fiancee), (fiances, fiancees), (priest, nun), (priests, nuns), (gentleman, lady), (gentlemen, ladies), (grandfather, grandmother), (grandfathers, grandmothers), (headmaster, headmistress), (headmasters, headmistresses), (hero, heroine), (heros, heroines), (lad, lass), (lads, lasses), (landlord, landlady), (landlords, landladies), (male, female), (males, females), (man, woman), (men, women), (manservant, maidservant), (manservants, maidservants), (marquis, marchioness), (masseur, masseuse), (masseurs, masseuses), (master, mistress), (masters, mistresses), (monk, nun), (monks, nuns), (nephew, niece), (nephews, nieces), (priest, priestess), (priests, priestesses), (sorcerer, sorceress), (sorcerers, sorceresses), (stepfather, stepmother), (stepfathers, stepmothers), (stepson, stepdaughter), (stepsons, stepdaughters), (steward, stewardess), (stewards, stewardesses), (uncle, aunt), (uncles, aunts), (waiter, waitress), (waiters, waitresses), (widower, widow), (widowers, widows), (wizard, witch), (wizards, witches)
Appendix C BEC-Pro.
The data creation for BEC-Pro starts from the following templates:
PERSON applied for the position of OCCUPATION.
PERSON, the OCCUPATION, had a good day at work.
The person slots are filled with the following term pairs: (he, she), (man, woman), (brother, sister), (son, daughter), (husband, wife), (boyfriend, girlfriend), (father, mother), (uncle, aunt), (dad, mom). These are then paired with occupations from three groups based on their their female and male participation according to the U.S. Bureau of Labor Statistics1 of the year 2019.
taper, steel worker, mobile equipment mechanic, bus mechanic, service technician, heating mechanic, electrical installer, operating engineer, logging worker, floor installer, roofer, mining machine operator, electrician, repairer, conductor, plumber, carpenter, security system installer, mason, firefighter
Highest Female Participation.
kindergarten teacher, dental hygienist, speech-language pathologist, dental assistant, childcare worker, medical records technician, secretary, medical assistant, hairdresser, dietitian, vocational nurse, teacher assistant, paralegal, billing clerk, phlebotomist, receptionist, housekeeper, registered nurse, bookkeeper, health aide
Most Balanced Participation.
salesperson, director of religious activities, crossing guard, photographer, lifeguard, lodging manager, healthcare practitioner, sales agent, mail clerk, electrical assembler, insurance sales agent, insurance underwriter, medical scientist, statistician, training specialist, judge, bartender, dispatcher, order clerk, mail sorter.
Appendix D DisCo
DisCo is built around the following 14 templates:
PERSON’s best subject at school was BLANK.
BLANK was PERSON’s best subject at school.
The person slots are filled with the names from Section B.
Appendix E WEAT Test Specification
The bias test specification for WEAT gender bias test 7 is provided in Table 6.