CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias

Vipul Gupta, Pranav Narayanan Venkit, Hugo Laurençon, Shomir Wilson, Rebecca J. Passonneau

Introduction

Language models (LMs) have been found to exhibit unintended biases (Hada et al., 2023; Levy et al., 2023; Gupta et al., 2022) leading to uneven performance across different sociodemographic groups (Bender et al., 2021; Schwartz et al., 2021; Blodgett et al., 2020). Recently, increasing amounts of effort have been devoted to reduction of unintended outputs from LMs, such as toxic language or manifestation of harmful social bias, e.g., through red teaming (Perez et al., 2022; Ganguli et al., 2022; Zhuo et al., 2023). To evaluate red teaming, or other bias mitigation methods, it is necessary to quantify LM bias in a consistent and rigorous manner. Due to increasing application in real-world of LMs, it is important to have reliable and robust measures to quantify bias. Prior work on bias measurement are unreliable (Selvam et al., 2023), as they are sensitive to minor perturbations in the templates designed to compare performance across social groups (cf. Fig. 1), due to factors such as lack of template diversity, or limited number of templates. As LMs become more task-agnostic, it’s increasingly important to assess biases across a variety of tasks, yet majority of current approaches often addresses a single NLP task, such as question-answering. Our goal here is to develop robust and reliable measurement of a few universally relevant social bias categories across multiple NLP tasks, providing a basis for future investigation of other types of social bias across many NLP tasks in a single medium. We introduce the Comprehensive Assessment of Language Models (CALM) for robust measurement of two types of demographic bias that are universally relevant, gender and race, which we apply to twenty pretrained LMs. In accordance with the group fairness framework proposed by Czarnowska et al. (2021), within this paper, we define bias as the disparate treatment of one group or an individual compared to another, given similar circumstances.

Construction of CALM was inspired by multi-faceted benchmark datasets such as GLUE (Wang et al., 2018a) and SuperGLUE (Wang et al., 2019). CALM draws examples of three NLP tasks, question answering, sentiment and natural language inference, from sixteen widely-used datasets. We selected 224 templates from these datasets and adapted them to include person-names representing different social groups. No prior work known to the authors has incorporated multiple tasks, particularly through the utilization of natural sentence datasets for template creation. Figure 2 illustrates templates produced by removing person names from a context-question pair, a sentiment sentence and a premise-hypothesis pair. To fill the person-name slots, we assembled sets of 50 highly frequent person names associated with three gender categories and four race categories, for a total of 350 person names. This generated a dataset of 78,400 prompts for comparing performance across these categories, for the three NLP tasks. Using a slight adapation of the standard bias metric, we compute bias score by comparing model performance for each social group with baseline performance, and take the difference between the maximum and minimum of these per-group scores to quantify bias.

Previous bias scores based on template-based prompts have been found to be sensitive to perturbations such as synonym substitution (Selvam et al., 2023), and often rely on manually designed templates (Seshadri et al., 2022; Alnegheimish et al., 2022). To address this, CALM offers a larger and more diverse range of prompts. A sensitivity analysis based on the methods proposed in (Selvam et al., 2023) shows that CALM bias scores are more robust than other bias identification datasets. We attribute the increased robustness to the larger size and greater diversity of CALM prompts. We also compared the diversity of CALM prompts with other works, using metrics like average length and semantic similarity.

We report bias benchmarking on 20 large language models (LLMs), including six prominent families of LLMs such as Llama-2. To our knowledge, no prior bias benchmark dataset has been tested on such a large collection of LLMs. In two LM series, OPT and Bloom, we found that larger parameter models are more biased than smaller ones. CALM bias measures for the T0 series are much lower than for other LM families. Conversely, Llama-2, Falcon, and Bloom models exhibit relatively more bias. Finally, we noticed a tradeoff between gender and race bias in some models, where increasing model size decreased one bias type while increasing the other. These findings shed light on the interplay among bias types in LLMs with respect to model size and series, providing new insight into model behavior across social groups.

The next five sections present related work, describe construction of CALM templates and CALM bias measurement, perform empirical evaluation of the robustness and reliability of CALM, document the LLMs selected for benchmarking and report results of bias measurement across these LLMs. The final four sections discuss the implications of our results, present our conclusions, summarize the limitations of our work, and discuss the broader impact.

Related Work

CALM has six benefits over the prior work described here, as summarized in Figure 3: 1) three characteristic NLP tasks rather than one; 2) application to two universally relevant social distinctions (gender and race); 3) generation of prompts by combining 224 templates with 350 person names that are frequent, and representative of distinct social groups; 4) a large number of prompts (N=78,400); 5) greater diversity in prompt length and meaning; 6) robustness of bias measurement to prompt perturbation and prompt subset selection.

Quantification of bias is an active research area. Early work measured cosine similarity in hidden layer embeddings (Caliskan et al., 2017; Dev and Phillips, 2019; Bolukbasi et al., 2016; Tan and Celis, 2019; Venkit et al., 2022; Gupta et al., 2023). This approach directly assessed the learned representations of LMs, but was found to have reliability issues, and did not correlate with real-world bias (Goldfarb-Tarrant et al., 2021; Webster et al., 2018). Recent work shifted to template-based approaches, where models are prompted with pre-defined templates to capture specific types of bias (Smith et al., 2022; Prabhakaran et al., 2019; Ahn and Oh, 2021). These approaches directly measure performance differences across social groups. Most template-based approaches are reported on a single NLP task, thus weakening the generality of the resulting measure, given that the same LM can be incorporated in many tasks. Previous work investigated tasks such as coreference resolution (Rudinger et al., 2018; Kurita et al., 2019; Helen, 2018; Sakaguchi et al., 2021; Zhao et al., 2018), machine translation (Stanovsky et al., 2019; Cho et al., 2019), sentiment detection (Bhaskaran and Bhallamudi, 2019; Venkit et al., 2023) and question answering (Parrish et al., 2022; Li et al., 2020; An et al., 2023). We select two of these, question answering and sentiment detection, and add natural language inference (NLI), which is similar to but more general than coreference resolution.

One of the main issues for template-based bias evaluations is the lack of reliability as they are sensitive to the choice of templates used for benchmarking (Seshadri et al., 2022). Resulting measures have been found to be sensitive to modifications to the templates, such as synonym substitution, which lead to significant changes in bias scores (Selvam et al., 2023). Further, the sets of bias prompts from a given study are often manually designed by the authors and lack diversity (Seshadri et al., 2022). which is a likely source of the observed unreliability. Additionally, some works try to cover broader range of demographic categories to identify biases, but often restrict to a limited number of templates, in the range of 10 to 30. HolisticBias (Smith et al., 2022) did an extensive evaluation across thirteen demographic categories but used only 26 manually-designed templates to quantify bias. Other works such as UNQOVER (Li et al., 2020), DisCo (Webster et al., 2021), BEC-Pro (Bartl et al., 2020), BITS (Venkit and Wilson, 2021; Venkit et al., 2023) used less than 30 manually-designed templates to discover biases. These small number of manually-designed templates makes their bias measurements unstable to minor modifications in the templates (Seshadri et al., 2022). In this work, we address these issues by selecting templates from a diverse set of existing dataset, in place of manually designing them. Additionally, we increase the number of templates significantly to make them more robust to cover a broader range of scenarios. We acknowledge that BBQ (Parrish et al., 2022) uses 325 templates, more than 224 templates in CALM, but they measure bias across nine bias categories and uses unique manually-curated templates, which range between 25-50 for each bias category. In contrast, CALM has more templates per category.

We hypothesized that a template-based approach that measures performance differences across social groups could be developed that would be more robust through greater size and diversity of templates and prompts. The next section describes how we test this hypothesis to address the issues raised for template based approaches in (Seshadri et al., 2022; Selvam et al., 2023).

CALM Data and Score

CALM is both our methodology for bias evaluation and a dataset we assembled to measure gender and race bias.The first three subsections below present the datasets we extracted templates from for each of the three NLP tasks, using the test sets where possible. Our criteria for task selection were for the tasks to be distinct, well-studied, and to address broad capabilities for handling contextual information, including relational meaning (who does what to whom), sentiment and logical relationships. The next two subsections present the template creation procedure and assembly of person name sets for gender and race. The last subsection explains our bias score. Additional details are presented in the Appendix.

For Question Answering (QA), we selected datasets where the answer is present in or easily inferred from context. This avoids confounding the effect of social group on model performance with real-world knowledge. Table 1 lists the 8 QA datasets with the number and proportion of templates contributed to the CALM QA task. All selected datasets have ground truth answers available. Below is a brief description of each dataset used for QA task.

bAbI: Weston et al. (2016) provides a set of 20 toy QA tasks for narrative understanding and reasoning. Each task involves characters interacting in a common sense setting. This dataset tests varies skills in models such as chaining facts, simple induction, and deduction.

SODAPOP: The SOcial bias Discovery from Answers about PeOPle dataset (An et al., 2023) adapted instances from the Social IQa dataset (Sap et al., 2019) to identify bias and stereotypical associations in LMs. We use the Bethany dataset file provided by authors.

TweetQA: This dataset was created from journalists’ tweets (Xiong et al., 2019). TweetQA is challenging due to the informal nature of the language used on social media, as compared to news or Wikipedia. We use the dev set, as test set answers are not publicly available.

MCTest: Machine Comprehension of Text (Richardson et al., 2013) consists of fictional stories and multiple choice questions. This dataset was collected via crowdsourcing. We use the MC500 test set, as it is more grammatically correct than MC160.

Relation Extraction: Levy et al. (2017) reduced relation extraction (RE) to reading comprehension, to create a new dataset for zero-shot RE. They crowd-sourced questions for each relation and aligned them with Wikipedia paragraphs. We use their test dataset for template generation.

QAMR: Question-Answer Meaning Representations consists of crowdsourced QA pairs from Wikinews and Wikipedia (Michael et al., 2018). Predicate-argument structures of sentences are represented as QA pairs to capture the rich semantic structure of text. We use their test data.

DuoRC: Duo Reading Comprehension consists of QA pairs of movie plots from Wikipedia and IMDb (Saha et al., 2018). Lexical overlap between questions and answers is avoided, thus requiring deeper language understanding and reasoning capability. We use the SelfRC test set, where answers were always present in the context.

MCScript: Machine Comprehension Using Script Knowledge focuses on everyday activities, such as going to the movies or working in the garden (Ostermann et al., 2018). The questions are based on commonsense reasoning, and answers are directly present or easily inferred from the context. We use their test set.

2. Sentiment Analysis

In Sentiment Analysis (SA), sentences are classified as positive or negative, or sometimes in a third neutral class. Sentiment classification has little if any overlap with QA. Table 2 lists the 4 sentiment datasets with the number and proportion of templates contributed to the CALM sentiment task.

SST: The Stanford Sentiment Treebank contains movie review sentences and human annotations for the sentiment of each review (Socher et al., 2013). We extract sentences that mention gender-specific terms from the published SST2 subset.

ToxicComments: The Toxic Comment Classification dataset from a Kaggle challenge consists of comments labeled for toxicity (Jigsaw, 2018). The task is to classify toxicity into one of six classes. We selected sentences that were labeled as toxic towards specific gender categories.

Sentiment140: The Sentiment140 dataset consists of randomly extracted tweets from Twitter (Go et al., 2009). We included it so CALM would have a broad range of sentences from social platforms.

EEC: The Equity Evaluation Corpus consists of English sentences designed to reveal bias towards certain groups (Kiritchenko and Mohammad, 2018).

3. Natural Language Inference

The Natural Language Inference (NLI) task involves sentence pairs that state a premise and a hypothesis. The models predict whether the sentences are entailed, contradictory, or neutral. This task requires a model to understand logical relationships between sentence pairs. Table 3 lists the 4 NLI datasets with the number and proportion of templates contributed to the CALM NLI task.

SNLI: Stanford Natural Language Inference contains human annotations grounded by image captioning (Bowman et al., 2015). Premise sentences were taken from image captions, and hypothesis sentences were written by crowdworkers. We use the test data.

WNLI: Winograd Natural Language Inference is one of the nine GLUE benchmarks (Wang et al., 2018b). It is designed to evaluate a model’s ability to do pronoun resolution and understand contextual entailment. We use the dev data, as answers to the test data are not publicly available.

RTE: Recognizing Textual Entailment is one of the nine GLUE benchmarks (Wang et al., 2018b). It contains sentence pairs from news and Wikipedia text. We use the dev data from this dataset.

SICK: Sentences Involving Compositional Knowledge contains sentence pairs rich in lexical, syntactic and semantic phenomena (Marelli et al., 2014). It was created using image and video descriptions. We use the test data.

4. Template Creation

To filter templates for the above tasks from each dataset, we use criteria directed at sociodemographic distinctions, and diversity of templates. For QA and NLI, we look for the presence of person names. For SA, we retrieve sentences with pronouns or person names. To ensure template quality after filtering, we manually verified each template, which led to discarding examples such as QA examples of stories with names of animal characters. Following the filtering step, each example undergoes a template extraction process, where person names and pronouns are replaced with corresponding tags as shown in Figure 2. We use same set of templates to generate prompts for both of our bias categories, race and gender.

To create CALM, we filtered 224 templates for the three tasks, consisting of 93, 77 and 54 for the QA, SA and NLI tasks, respectively. The distributions of templates from the three tasks are shown in Tables 1-3. Notably, our approach resulted in selection of a diverse set of templates to ensure comprehensive coverage across different domains.

5. Bias Categories

Gender bias: To quantify gender bias, names were sampled from three gender categories - male, female, and names not associated with either gender (gender-neutral) - with 50 names per category. This resulted in 150 testing prompts for each template. Male and female names were selected from the top 1000 names from the US Social Security dataset.https://www.ssa.gov/oact/babynames/ We restricted selection to names with >> 80% usage in a given gender. This partitioning approach is similar to previous approaches (Webster et al., 2021). Gender-neutral names were sampled from an archived ABC News article that used data from the Social Security Administration (Feldman, 2015). We removed gender neutral names from male and female names to ensure no data overlap.

Race bias: To quantify race bias, we sampled names across four race/ethnic groups - Caucasian, African American, Hispanic and Asian - with 50 names per category, yielding a total of 200. These four groups were selected based on the availability of corresponding labels in US census data, and the Harvard dataverse.https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/SGKW0K We restricted selection to names with >> 80% usage in a given category.

Each template contains identifiers as shown in Figure 2. identifiers are replaced with gender and race names to produce 50 prompts for each social group. In total, by combining the 350 names for seven categories across gender and race with 224 templates, we generated 78,400 prompts for CALM.

6. Bias Score

In this section, we explain how we measure bias in LMs. As in previous work that measures bias based on task performance (Parrish et al., 2022; Jigsaw, 2018; Mathew et al., 2021; Elazar and Goldberg, 2018), the assumption is that performance should be consistent across all social groups. First, we establish a baseline performance to show how well the model usually performs on the task by taking the average across all prompts. Then, we examine how the model performs for each social group separately and compare this to the baseline. If a model is unbiased, its score for each group should match the baseline, resulting in a bias score of 0%. A measurable difference between performance on a specific group compared to the baseline indicates bias. Then we examine the bias scores per social group to arrive at the overall the bias score for a given LM. An ideal model will have 0 bias score.

For each template, we have fifty names per seven social groups, yielding 350 prompts. We take the baseline score on a template to be the average accuracy on the 350 prompts. Similarly, for each social group we calculate the average accuracy of the 50 prompts for that group. For each prompt in CALM, we have the ground truth answer. We use that to calculate number of prompts which were answered correctly (#correctsg)(\#correct_{sg}). The bias score for a given social group is the difference from the baseline, taken as a percentage change as follows :

This bias score tells us how much the model’s performance for a social group differs from the average baseline performance of the model. To calculate the bias of the model for a given task, we take the difference between the maximum and minimum bs across all social groups. We calculate a gender bias score by comparing scores across the three gender categories, and a race bias score by comparing across four racial categories. These scores provide a breakdown of bias by race and gender for each NLP task. We also calculate a single bias score for the model as the average of gender and racial bias across the three NLP tasks included in CALM.

Evaluation of CALM

In this section, we carry out an empirical evaluation of the robustness and reliability of CALM bias scores. In the rapidly evolving field of NLP, the importance of robust and reliable bias benchmark datasets cannot be overstated. Unreliable bias benchmark measurement would lead to misleading and inconsistent conclusions, with far-reaching implications, particularly as LMs are increasingly used in real-world applications. Biased language models could inadvertently perpetuate stereotypes and unfair representations (representational harm), or disadvantage individuals in hiring, promotion, healthcare or the like (allocational harm). Without reliable measurement of bias, system developers will be unable to favor LMs that are less biased, and researchers will be unable to support claims of bias mitigation. CALM aims to facilitate the use and development of LMs that are more equitable and representative of diverse perspectives, thereby assisting technological advances in NLP to contribute more fairly across social groups. In the remainder of this section, we assess the sensitivity of CALM bias scores to perturbation of prompts, and to random selection of subsets of prompts, finding that scores remain relatively stable. We attribute this in part to the greater diversity of CALM templates compared with previous bias measurement, as presented in the third subsection. We conclude this section with qualitative observations.

Benchmark datasets for social bias are often sensitive to minor modifications in the dataset. Recent research by (Selvam et al., 2023) demonstrated that seemingly innocuous perturbations such as synonym substitution, can significantly impact bias scores in these benchmarks. To assess the sensitivity of CALM, we followed a similar methodology to (Selvam et al., 2023), creating four alternative constructions of CALM. These versions introduced modifications to the original templates by perturbing them through synonym substitution, addition of clauses, and addition of adjectives. These perturbations resulted in a dataset five times the size of CALM. No prior work has explored a comprehensive robustness assessment akin to the one presented in this paper, underscoring the exhaustive validation of our dataset’s resilience.

We evaluated CALM’s robustness by testing five different language models, comparing results from the original CALM dataset with those from its perturbed versions. As detailed in Table 4, CALM showed minimal sensitivity to these semantic perturbations, with a maximum variance of less than 10% across all models. This level of stability is remarkable compared with prior datasets, as reported in (Selvam et al., 2023): BiasNLI (Dev et al., 2020) showed a 70% variation in bias score, dropping from 41.6 to 13.4; Winogender (Rudinger et al., 2018) had a 77% increase, rising from 5.83 to 10.33.

2. Prompt Subset Selection

Another critical aspect of bias benchmark datasets is their sensitivity to prompt subset selection, a factor that can significantly affect bias measurement outcomes. As reported in (Selvam et al., 2023), subset selection can produce as much as a 40% change in bias measurement. To understand how CALM stands up to this challenge, we conducted a series of reliability analyses, performing six runs across four language models, each time with a different proportion of randomly selected prompts (75%, 50%, and 25%). The results, detailed in Table 5, demonstrate that CALM exhibits only minimal deviations in these conditions, markedly lower than the measurement variance reported by (Selvam et al., 2023).

3. Comparative Analysis with Other Bias Datasets : A Diversity Analysis

To further examine the quality of CALM templates, we compared CALM with other bias datasets using various diversity measures. We measured template diversity using BERTScore (Zhang et al., 2020), which computes cosine similarity between the contextual BERT embeddings (Devlin et al., 2019) of sentence pairs. To quantify the diversity of a dataset, we take the average of the BERTScore between all pairs of templates within the dataset. We also examine template length, using the average and standard deviation of number of words per template in a dataset. We compared CALM with seven other bias datasets: DisCo (Webster et al., 2021), BEC-Pro (Bartl et al., 2020), UNQOVER (Li et al., 2020), BITS (Venkit and Wilson, 2021; Venkit et al., 2023), HolisticBias (Smith et al., 2022), Counterfactual-eval (Huang et al., 2020) and BBQ (Parrish et al., 2022).

As shown in Table 6, CALM has the lowest average BERTScore of 0.388, indicates a lower semantic similarity between its templates, suggesting greater diversity. In contrast, datasets such as UnQOVER (0.660), BITS (0.617), and BEC-PRO (0.594) showed higher average BERTScore ≥\geq 0.59, suggesting a substantial template redundancy. The higher diversity of the CALM templates is further supported by examining template length. CALM stands out in terms of template length and standard deviation. Our dataset not only has a higher average length but also a significant standard deviation in template length, ranging from brief sentences to extensive paragraphs. A higher standard deviation illustrates considerable variability in template length, with templates ranging from short sentences to large paragraphs. Another major difference is the number of templates used in CALM is higher than other datasets as shown in Table 6. Only BBQ uses more templates but they measure bias across nine bias categories, including religion, disability status, physical appearance and socioeconomic status. In contrast, CALM has significantly more templates per category within its focused bias categories.

4. Qualitative Observations

An interesting observation emerged when we examined the performance accuracy of language models for each template, as compared with their average performance across all templates. For each template, we found significant variation in accuracy across different social groups. However, when we look at average accuracy across entire set of templates, we found a remarkable consistency in accuracy scores. We observed that the average accuracy varies within a narrow range of 0-3% across different social groups. We believe that this uniformity in accuracy is attributable to the rich and diverse range of scenarios covered by CALM’s templates, thus solidifying the case of higher diversity in the dataset. To illustrate this point, Table 7 presents accuracy score for Llama-2-13B model on question-answering task. We can see that there is high variability in template-wise accuracy, but consistent average accuracy across social groups.

Through our extensive evaluations, we show that CALM improves over previous bias benchmarks in two key aspects: the higher linguistic diversity of templates, and its greater reliability in measuring certain biases in language models. Its comprehensive linguistic coverage ensures that CALM not only identifies biases more accurately, but can also provide deeper insights into the nuanced behaviors of language models. Consequently, CALM stands out as a robust and reliable methodology for detecting and understanding bias in language models.

Models Evaluated

In this work, we perform an empirical evaluation of 20 open-source LMs including six prominent families of LLMs: Llama-2 (Touvron et al., 2023), Bloom (Scao et al., 2022), OPT (Zhang et al., 2022), Falcon (Penedo et al., 2023), T0 (Sanh et al., 2021) and GPT-Neo (Black et al., 2021). The models under examination vary in size from 1 billion parameters for Bloom to 70 billion parameters for Llama-2, allowing us to analyze performance across a wide range of model sizes. In line with recent work on in-context learning for language model evaluation (Liang et al., 2023; Brown et al., 2020), we evaluate all models using 5-shot prompts. For each template, five examples are randomly sampled from the training set of the corresponding dataset following the procedure established in HELM (Liang et al., 2023). These examples are appended to the prompt to provide the model with demonstrative examples before evaluating on a given task. Furthermore, we fix the in-context examples for each dataset across models to ensure standardized comparison, an approach also adopted in HELM (Liang et al., 2023).

For prompt formatting for each of the three tasks, we select the prompt structure followed by HELM (Liang et al., 2023) and (Brown et al., 2020). As argued by (Liang et al., 2023), prompts tailored for each model may yield optimal performance but is challenging for controlled evaluation. Due to practical computation and time constraints, in this work we use the commonly accepted prompts following (Liang et al., 2023). We mention the exact prompts we used in the appendix. Moving forward, it is desirable to have standardized prompts across models to have similar prompting technique, and to facilitate greater comparability.

Results

We evaluate each model on the CALM dataset. Table 8 shows the bias results for each model along with a task-wise breakdown. In Table 8, the suffix with each model denotes the number of parameters in billions. For instance, Llama-2-7B signifies the 7 billion parameter variant of the Llama-2 series of language models.

Lower bias scores indicate lower demographic disparities in model performance (a perfectly unbiased model would have 0 bias score across all tasks). During our experiments, we observed that certain models exhibit significant underperformance in specific tasks, achieving near-zero accuracy or producing identical output regardless of the input. As a result, we exclude such tasks from bias scores for those models, as reflected in the empty cells in Table 8.

We found that for two out of six LM families, larger parameter models are more biased than lower parameter models. Specifically, for the OPT models, the average bias increased by 29% from 11.6 for the 2.7B parameter variant to 15.0 for the 30B parameter variant. Similarly, for the Bloom models, the average bias exhibited an increase of 81%, rising from 12.9 for the 1B parameter variant to 23.4 for the 7B parameter variant. The T0 series of LMs demonstrate significantly lower bias as compared to other models. Conversely, Llama-2, Falcon and Bloom models exhibit more bias than other model series as shown in Table 8. Notably, the T0+ model, an 11B parameter model from the T0 series, emerged with the lowest bias scores among all the tested models.

During our analysis, we observed that sometimes increased model size results in a tradeoff between gender and racial bias. For OPT models, increasing the model size from 6.7B to 30B increases the gender bias by 29% from 11.5 for 6.7B to 14.8 for 30B parameter model, while decreasing the racial bias by 9% from 16.6 for 6.7B to 15.1 for 30B model. Looking at the results per task, we observe that for some models there is a tradeoff in the bias scores. For example, for Falcon models, increasing the model size from 7B to 40B parameters increases the NLI gender bias by 18% (23.3 for 7B vs 27.6 for 40B), while decreasing QA and SA gender bias by 64% and 42% respectively. Similarly for GPT-Neo increasing the model size from 1.3B to 2.7B increases QA race bias by 74% (13.7 for 1.3B vs 23.8 for 2.7B), while decreasing the NLI race bias by 41% (14.5 for 1.3B vs 8.5 for 2.7B).

For the OPT model series we observe a noteworthy trend, which is also depicted in Figure 4. Initially, the bias score decreases from 24.5 to 11.6 as the model size increases from 1.3B to 2.7B parameters. Subsequently, the bias score increases from 11.6 to 15.0 while increasing the model size from 2.7B to 30B parameters. This bias trend for OPT models is similar to the one observed by (Helen, 2018) on Winobias, where OPT-13B and 30B variants are found to be more biased than 1.3B and 2.7B OPT variants.

Here we delve deeper into our results to better understand nuanced differences in bias patterns within templates, offering insights into behavior of language models.

Efficiency of Template Subset in Bias Evaluation: A subset of CALM turns out to be highly effective for bias measurement. Through targeted experiments, we discovered that eliminating 68 templates from the dataset had very minimal influence on the bias scores across various LLMs. All these 68 templates are roughly equally distributed across all three tasks. As illustrated in Table 9, it is feasible to achieve similar bias detection results using only 156 templates, which is approximately 70% of the original dataset size. This finding highlights the potential to use a more concise dataset for more efficient yet equally accurate assessment of LM bias, potentially optimizing the evaluation process.

Template-wise results: A more granular template-wise analysis sheds light on social biases unique to specific language models. For instance, in the Llama-2-7B model, question-answering templates incorporating words like “competitive” displayed a higher accuracy for male identifiers. Conversely, the Falcon-7B model showed a preference for female identifiers in question-answering templates where “garden” was the correct answer. These model-specific biases might help in tailoring bias mitigation strategies for each model.

Interestingly, we found some inexplicable recurring bias patterns linked to a common subset of templates. For instance, question-answering templates related to occupations, with “deputy” as the correct answer, consistently yielded higher accuracy for female identifiers and lower for males across different LLMs. Similarly, templates incorporating words like “crying” exhibited a marked decrease in accuracy for male identifiers. We think that these common biases likely stem from similar dataset biases present in the training data of various LMs. While these templates can inflate the bias scores for all LLMs, we found that such templates are very small in number. As shown in Table 9, out of all the templates in the CALM dataset, we identified 8 — less than 4% — that consistently reveal these common bias trends across LLMs.

This in-depth template error analysis not only assists in pinpointing specific biases in individual models but also in recognizing common bias trends across all LLMs. These insights are crucial for developing more nuanced and effective strategies for bias mitigation in language models, ensuring that they operate fairly and impartially.

Discussion

Interpretation of bias scores: Our bias score for a language model can be interpreted as the average difference in performance of the LM across different sociodemographic groups, for three tasks. A lower bias score means that the model’s accuracy is relatively similar to the baseline accuracy of the model for each sociodemographic group, while a higher bias score indicates that the model’s accuracy differs from baseline performance for sociodemographic groups. Ideally, we would want all LMs to have near-zero bias scores, independent of how well they perform on common benchmarks. This is highlighted through the framework of bias defined in this paper (Czarnowska et al., 2021). A higher LM bias score is associated with an increased potential for harmful real-world impacts from use of the model.

Comparing different model series: We believe that our dataset is a good tool for comparing bias across model series, enabling observation of trends exhibited by different models. We observed all models in the T0 series to have significantly lower bias scores as compared with all models in the Llama-2, Falcon, and Bloom series of models. This indicates that the training procedure followed in T0 models may be effective at producing less biased models. While we focus on collecting a large number of diverse templates, slight differences in bias scores, as with T0+ vs T0++, can be attributed to noise. However, a significant difference in bias scores, as with Llama-2 vs T0, indicates a need for bias mitigation in Llama-2.

Comparing models within the same language model series: Analysis of change in bias scores with increasing numbers of parameters for a model series provides interesting insights. We observed that for the OPT and Bloom model series, bias scores exhibit an upward trend with the increasing number of parameters. While increasing model parameters may improve performance on common benchmarks, it is important to evaluate the bias trend within each model series. Improvement in performance on common benchmarks might come at the expense of increased bias in models, thus potentially increasing the negative impact for real-world applications of these models. Our analysis shows that there is no common trend in bias trajectories across all model series, highlighting the complexity of bias behaviors.

Robustness of CALM: In our study, we underscore the critical importance of robust dataset construction for a nuanced comprehension and detection of diverse group biases. While prior research has commonly employed sentence templates for bias measurement, our investigation reveals their vulnerability to modifications and adversarial alterations, leading to potential miscapture of biases. This highlights the limitations of solely relying on sentence templates, as they often overly emphasize sentence structure and semantics rather than contextual relevance. Through the development of CALM, we advocate for a hybrid approach that incorporates both sentence-based and dataset-based sentences, offering a more contextually rich understanding of social group dynamics. Notably, our work demonstrates the resilience of CALM to sentence manipulation, affirming its robustness in effectively measuring group bias.

Task Sensitive Design: By focusing on identifying group biases in contemporary NLP technologies, we contribute a valuable tool for discerning biases across diverse tasks associated with these models. Each task within CALM is meticulously designed to encompass contextual relevance to the task itself and to the nuanced capture of bias. Consequently, our paper introduces a novel format for creating datasets that serve as a medium for bias identification and emphasize context in a task-specific manner. This dual contribution positions our work as a novel effort to advance the methodology of creating datasets for future applications, particularly in transparency of text generation models.

Conclusion

We present CALM, a benchmark dataset, and a set of procedures to quantify bias in LMs. CALM integrates 16 existing datasets for three NLP tasks to create a dataset to quantify gender and racial bias. CALM has several benefits over previous bias datasets including coverage of three NLP tasks rather than one, greater diversity in template length and meaning, and robustness to prompt perturbation and prompt subset selection. We find that for some families of large language models, larger parameter models tend to be more biased than smaller ones. To create CALM, we paid special emphasis to creating a diverse and reliable dataset, and to making it extensible. We believe that our work addresses some of the issues with other bias datasets, and that it takes an important step towards reliable and robust bias evaluation in LMs.

Limitations

The target word list we used for the CALM dataset creation is limited to seven social groups in the US and we acknowledge that many more social groups belonging to gender and race, as well as different countries, are missing. However, to broaden bias assessment beyond US names, we compiled a dataset tabulating names from various national origins. This dataset, using the scripts we provide, allows the evaluation of LM bias across diverse social groups from various countries. Moreover, the templates used in our dataset are in English. We believe that our approach can be extended to other languages, however it requires careful consideration of linguistic nuances and cultural differences.

As language models evolve to become more versatile and task-agnostic, it’s increasingly crucial to assess biases across a diverse range of tasks. However, for some models we encountered either a low baseline performance or higher biases for a particular task. Such inconsistent behavior makes it hard to develop an understanding of a model’s overall bias in some cases. Future research is needed to better understand how to incorporate multiple tasks in a better way to measure overall bias for a language model. Another limitation is the presence of overlapping names between gender and race categories. This overlap may cause some interdependence in gender and race bias scores. We made some effort to minimize this overlap but complete elimination proved challenging. Further research is needed to devise methods for quantifying distinct bias categories completely independent of one another.

Evaluating text generation models on a specific task is a hard problem. As the prompts used during training is largely unknown for majority of language models, it is difficult to find prompts to get the best performance. We tried to perform 5-shot prompting to perform in-context learning on commonly used prompts to get best performance. We hope that there is prompt standardization across models which can facilitate better comparability among models. Despite its limitations, we believe CALM is a step in the right direction to reliably evaluate biases in language models.

Broader Impact

The discussions on the potential risks of AI systems in the media, within the general public, and among national and international policy developers are increasing. We are starting to see international summits and national executive orders to increase awareness of and manage the risks of AI. Notably, recent reports have highlighted tradeoffs in utility of AI, particularly in sectors like healthcare that already rely heavily on AI (Jewett, 2023). Amidst the rapid proliferation of AI as a Service (AIaaS) models (Lewicki et al., 2023), characterized by their ’plug-and-play’ functionality, and the simplicity they offer without requiring expertise in AI model development, it is becoming increasing important to better comprehend the inherent biases within these tools. The prevalent ’one-size-fits-all’ approach often engenders challenges related to bias and fairness. Recognizing the importance of understanding and mitigating these risks, it becomes imperative to develop robust and reliable bias datasets, like the one presented in our work, to measure the potential for negative impact in the real-world setting. This is particularly important as we continue to integrate AI into various facets of life, where unnoticed biases could have far-reaching and detrimental impacts.

Our work also addresses the limitations inherent in previous bias benchmarks, specifically their sensitivity to simple perturbations, by introducing a novel dataset and methodology. By presenting a more robust approach to quantify certain social biases in language models, we strive to foster a better understanding of the potential bias (and invariable harms) stemming from language model bias. Furthermore, we envision that our work serves as a catalyst for the development of bias mitigation tools, ultimately contributing to the creation of language models that are not only technologically advanced but also ethically responsible. Our broader influence lies in advancing the discourse on fair and transparent AI, aligning technological innovation with ethical considerations to ensure the positive impact of AI on all sections of society.

We publicly release CALM, along with its design methodology, transforming it into a shared bias identification platform similar to an AIaaS technology. This empowers individuals without prior experience in language model development to leverage CALM for bias identification. Our goal is to offer users the power of choice, allowing them to discern the inherent biases and behavioral patterns of the selected model. This democratization of bias identification tools aims to enable users to make informed decisions on whether the chosen model aligns with the intended social application.

Ethical Considerations

In conducting this research, we placed a strong emphasis on responsible and ethical research practices, including a thorough consideration of the environmental impact associated with our studies. Our experiments involved the use of 20 pre-trained large language models, and for the bulk of these experiments, we utilized 4 NVIDIA RTX A6000 50GB GPUs. The cumulative computing time required to evaluate all the language models and complete the comparison studies amounted to approximately 40 hours. Given the maximum power consumption of 300W per NVIDIA RTX A6000 GPU and considering the global average carbon intensity of electricity at 0.475kg CO2/KWh – with 30% of electricity globally derived from renewable sources – our study’s total carbon footprint was calculated to be around 15.96 kg of CO2. To responsibly address this environmental impact, we have made a contribution to the US Forest Service’s Plant-a-Tree program, which is an effort to offset the carbon emissions generated by our research activities.

Adverse Impacts

In our effort to establish our dataset as a benchmark for assessing social biases in language models, we recognize that openly sharing the details of our methodology and dataset sources comes with potential risks. While transparency is crucial for scientific progress and reproducibility, it also means that the specific datasets from which we derived our templates become publicly known. As the trend grows towards less transparency about the training datasets used for large language models, there arises a consequential risk of data contamination. This issue becomes particularly concerning if certain individuals or organizations decide to train their language models using the exact datasets we utilized and potentially using data augmentation techniques to mimick our methodology. Such a scenario could lead to misleading outcomes. Specifically, models trained on these contaminated datasets might appear to exhibit lower levels of bias, not because they inherently do, but because they have been inadvertently tuned to perform well on our benchmark. This illusion of reduced bias poses a significant risk, especially when these models are deployed in real-world applications. It could lead to overconfidence in the fairness and neutrality of these models, potentially hiding biases they might manifest in real world setting.

While we strive to advance the field by providing a robust tool for bias evaluation, we also urge the community to be cautious of these potential negative impacts. It is essential for users of our dataset and methodology to be aware of these risks and to employ strategies that mitigate the likelihood of data contamination and its consequent adverse effects.

References

Appendix A Appendix

We included task /1, 6, 8, 9, 10, 11, 12, 13, 14 tasks for template filtering. We excluded task 2 and task 3 as the question does not contain enough info to measure gender bias. Tasks 4 and 19 were not included as they contain only info about location and direction, no person data. Task 5 had too many names in the context. Task 7 was related to counting objects. Task 15 contains animal information. Task 16 contains animal and color information. Task 17 and 18 contains no person data. In task 20 answers are not present in the context.

A.2. Gender wise results

The breakdown for performance on sentiment analysis task for five LLMs over CALM dataset is presented in 10. We can see very little difference in accuracy among different gender groups

A.3. Prompts used

For each of the three tasks, we select the prompt structure followed by HELM (Liang et al., 2023) and (Brown et al., 2020). For QA templates, we follow the following prompt structure: "Passage: .\n Question: .\n Answer:". For sentiment analysis templates, we follow the following prompt structure: "Passage: \n. Sentiment: ". For Natural Language Inference templates, we follow the following prompt structure: "Passage:\n. Question: \n. True or False?\n Answer: ".