Social Bias Probing: Fairness Benchmarking for Language Models

Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, Isabelle Augenstein

Trigger warning

This paper contains examples of offensive content.

Introduction

The unparalleled ability of language models to generalize from vast corpora is tinged by an inherent reinforcement of societal biases which are not merely encoded within language models’ representations but are also perpetuated to downstream tasks (Blodgett et al., 2021; Stańczak and Augenstein, 2021). These societal biases can manifest in an uneven treatment of different demographic groups – a challenge documented across various studies (Rudinger et al., 2018; Stanovsky et al., 2019; Kiritchenko and Mohammad, 2018; Venkit et al., 2022).

A direct analysis of biases encoded within language models allows to pinpoint the problem at its source, potentially obviating the need for addressing it for every application (Nangia et al., 2020). Therefore, a number of studies have attempted to evaluate societal biases within language models (Nangia et al., 2020; Nadeem et al., 2021; Stańczak et al., 2021; Nozza et al., 2022a). One approach to quantifying societal biases involves adapting small-scale association tests with respect to the stereotypes they encode (Nangia et al., 2020; Nadeem et al., 2021). These association tests restrict the scope of possible analysis to two groups, stereotypical and their anti-stereotypical counterparts. This binary approach not only restricts the breadth of the analysis by overlooking the complex spectrum of gender identities beyond the male–female dichotomy but is also problematic in evaluating other types of societal biases, such as racial biases, where identities span a broad spectrum and there is no singular “ground truth” with respect to stereotypical identity. The nuanced nature of societal biases within language models has thus been largely unexplored.

In response to these limitations, we introduce a novel probing framework, as outlined in Figure 1. The input of our approach consists of a dataset gathering stereotypes and a set of identities belonging to different societal categories: gender, religion, disability, and nationality. First, we combine stereotypes and identities resulting in our probing dataset. Secondly, we assess societal biases across three language modeling architectures in English. We propose perplexity (Jelinek et al., 1977), a measure of a language model’s uncertainty, as a proxy for bias. By evaluating how a language model’s perplexity varies when presented with probes that contain identities belonging to different societal categories, we can infer which identities are considered the most likely. Using the perplexity-based fairness score, we conduct a three-dimensional analysis: by societal category, identity, and stereotype for each of the considered language models. In summary, the contributions of this work are:

We conceptually facilitate fairness benchmarking across multiple identities going beyond the binary approach of a stereotypical and an anti-stereotypical identity.

We deliver SoFa (Social Fairness), a benchmark resource to conduct fairness probing addressing drawbacks and limitations of existing fairness datasets, expanding to a variety of different identities and stereotypes.

We propose a perplexity-based fairness score to measure language models’ associations with various identities.

We study societal biases encoded within three different language modeling architectures along the axes of societal categories, identities, and stereotypes.

A comparative analysis with the popular benchmarks CrowS-Pairs Nangia et al. (2020) and StereoSet Nadeem et al. (2021) reveals marked differences in the overall fairness ranking of the models, suggesting that the scope of biases LMs encode is broader than previously understood. In agreement with recent findings (Bender et al., 2021), we find that larger model variants exhibit a higher degree of bias. Moreover, we expose how identities expressing religions lead to the most pronounced disparate treatments across all models, while the different nationalities appear to induce the least variation compared to the other examined categories, namely, gender and disability.

Related Work

Presenting a recent framing on the fairness of language models, Navigli et al. (2023) define social biasThe term social is employed to characterize bias in relation to the risks and impacts on demographic groups, distinguishing it from other forms of bias, such the statistical one. as the manifestation of “prejudices, stereotypes, and discriminatory attitudes against certain groups of people” through language. Social biases are featured in training datasets and propagated in downstream NLP applications, where it becomes evident when the model exhibits significant errors in classification settings for specific minorities or generates harmful content when prompted with sensitive identities (Nozza et al., 2021).

Recent work Blodgett et al. (2021) has pointed out relevant concerns regarding stereotype framing and data reliability of benchmark collections explicitly designed to analyze biases in language models, such as CrowS-Pairs Nangia et al. (2020) and StereoSet Nadeem et al. (2021). Consequently, the effectiveness and soundness of the resulting fairness auditing is partly comprised. The scores proposed in the contributions presenting the datasets are highly dependent on the form of the resource they propose and therefore they are hardly generalizable to other datasets to conduct a more general comparative analysis. Specifically, Nangia et al. (2020) leverage on pseduo-log likelihood Salazar et al. (2020) based scoring. The score assesses the likelihood of the unaltered tokens based on the modified tokens’ presence. It quantifies the proportion of instances where the LM favors the stereotypical sample (or, vice versa, the anti-stereotypical one). The stereotype score proposed by Nadeem et al. (2021) differs from the former as it allows bias assessment on both masked and autoregressive language models, whereas CrowS-Pairs is limited to the masked ones. Another significant constraint highlighted in both datasets, as pointed out by Pikuliak et al. (2023), is the establishment of a bias score threshold at 50%50\%. It implies that a model displaying a preference for stereotypical associations more than 50%50\% of the time is considered biased, and vice versa. This threshold implies that a model falling below it may be deemed acceptable or, in other words, unbiased. Furthermore, these datasets exhibit limitations regarding their focus and coverage of identity diversity and the number of stereotypes. This limitation stems from their reliance on the binary comparison between two rigid alternatives — the stereotypical and anti-stereotypical associations — which fails to capture the phenomenon’s complexity. Indeed, they do not account for how the model behaves in the presence of other plausible identities associated with the stereotype, and these associations need scrutiny for low probability generation by the model, as they can be harmful regardless of the specific target. Additionally, these approaches do not address situations where associations are implausible, and the model is unlikely to generate them. Therefore, bias measurements using these resources could lead to unrealistic and inaccurate fairness evaluations.

Given the constraints of current solutions, our work introduces a dataset that encompasses a wider range of identities and stereotypes. The contribution relies on a novel framework for probing language models for societal biases. To address the limitations identified in the literature review, we design an original perplexity-based ranking that produces a more nuanced evaluation of fairness.

Social Bias Probing Framework

The proposed Social Bias Probing framework serves as a fine-grained language models’ fairness benchmarking technique. Contrary to the existing fairness assessments, which rely on a dichotomous framework of stereotypical vs. anti-stereotypical associations, our methodology expands beyond this oversimplified binary categorization. Ultimately, our approach enables the comprehensive evaluation of language models by incorporating a diverse array of identities, thus providing a more realistic and rigorous audit of fairness within these systems.

In Figure 1, we present a visual workflow of our approach. We first collect a set of stereotypes and identities leveraging the Social Bias Inference Corpus (SBIC; Sap et al. 2020) and an identity lexicon curated by Czarnowska et al. (2021). At this stage, we develop the new SoFa (Social Fairness) dataset, which encompasses all probes — identity–stereotype combinations (Section 3.1). The final phase of our workflow involves evaluating language models by employing our proposed perplexity-based fairness measure in response to the constructed probes (Section 3.2).

Our approach requires a set of identities from diverse social and demographic groups, alongside an inventory of stereotypes.

We derive stereotypes from the list of implied statements in SBIC, a corpus that collects social media posts having harmful biased implications. The posts are sourced from previously published collections that include English content from Reddit and Twitter, for a total of 44,00044,000 instances. Additionally, the authors draw from two “Hate Sites”, namely Gab and Stormfront. Annotators were asked to label the texts based on a conceptual framework designed to represent implicit biases and offensiveness.We refer to the dataset for an in-depth description (https://maartensap.com/social-bias-frames/index.html).

We emphasize that the choice of SBIC consists of an instantiation of our framework. Our methodology can be applied more broadly to any dataset containing stereotypes directed towards specific categories. As not all instances of the original dataset have an annotation regarding the stereotype implied by the social media comment, we filter it to isolate abusive samples having a stereotype annotated. Since certain stereotypes contain the targeted identity, whereas our goal is creating multiple control probes with different identities, we remove the subjects from the stereotypes, performing a preprocessing to standardize the format of statements (details are documented in Appendix A). Finally, we discard stereotypes with high perplexity scores to remove unlikely instances. We report the details of the preprocessing operations performed on the identities in Appendix A.

Identities

While we could have directly used the identities provided in the SBIC dataset, we chose not to due to their unsuitability from frequent repetitions and varying expressions influenced by individual annotators’ styles. To unify the set of analyzed identities, we deploy the lexicon created by Czarnowska et al. (2021). We map the SBIC dataset target group categories to the identities available in the lexicon (Table 1). Specifically, the categories are: gender, race, culture, disabilities, victim, social, and body. We first define and rename the culture category to include religions and broaden the scope of the race category to encompass nationalities. We then link the categories in the SBIC dataset to those present in the lexicon as follows: gender identities are extracted from the lexicon’s genders and sexual orientations, nationality identities are derived from race and country entries, religion utilizes terms from the religion category, and disabilities are drawn from the disability category. This mapping results in excluding the broader SBIC categories, i.e., victim, social, and body due to the difficulty in aligning the identities with the lexicon categories, and disproving the assumption of invariance in the related statements. The assignment of an identity to a specific category is inherited from the categorizations of the resources adopted. Recognizing that these framings inevitably simplify the complex nuances of the real world is crucial.

Lastly, since using lexica may introduce grammatical errors, we mitigate this by filtering rare identities based on their perplexity scores.

SoFa

To obtain the final probing dataset, we remove duplicated statements and apply lower-case. Finally, each target is concatenated to each statement with respect to their category, creating dataset instances that differ only for the target (Table 2). In Table 1, we report the coverage statistics regarding targeted categories and identities.

2 Fairness Measure

We propose the perplexity (PPL; Jelinek et al. 1977) as a means of intrinsic evaluation of fairness in language models. PPL is defined as the exponentiated average negative log-likelihood of a sequence. More formally, let X=(x0,x1,…,xt)X=(x_{0},x_{1},\dots,x_{t}) be a tokenized sequence, then the perplexity of the sequence is

where log⁡pθ(xd∣x<d)\log p_{\theta}(x_{d}\mid x<d) is the log-likelihood of the ddth token conditioned on the proceeding tokens given a model parametrized with θ\theta.

Our metric leverages PPL to quantify the propensity of a model to produce a given input sentence: a high PPL value suggests that the model deems the input improbable for generation. We identify bias manifestations when a model exhibits low PPL values for statements that contain stereotypes, thus indicating a higher probability of their generation. The purpose of this metric, and more generally our framework, is to provide a fine-grained summary of models’ behaviors from an invariance fairness perspective.

Formally, let C={religion,gender,disability,nationality}\mathcal{C}=\{\textit{religion},\textit{gender},\textit{disability},\\ \textit{nationality}\} be the set of identity categories; we denote elements of C\mathcal{C} as cc. Further, let ii be the identity belonging to a specific category cc, e.g., Catholics and ss be the stereotype belonging to cc, e.g., are all terrorists. We define Pi+sP_{i+s} as a singular probe derived by the concatenation of ii with ss, e.g., Catholics are all terrorists. Finally, let mm be the LM under analysis. The normalized perplexity of a probe is computed as follows:

Since the identities ii are characterized by their own PPL scores, we normalize the PPL of the probe with the PPL of the identity, addressing the risk that certain identities might yield higher PPL scores because they are considered unlikely.

We highlight that the PPL’s scale across different models can significantly differ based on the training data. Consequently, the raw scores do not allow direct comparisons. We facilitate the comparison of the PPL values of model m1m_{1} and model m2m_{2} for a given combination of identity and a stereotype:

We use the base-1010 logarithm of the PPL values generated by each model to analyze more tractable numbers since the range of PPL is [0,inf⁡)[0,\inf). In our formula, kk is a constant and represents the factor that quantifies the scale of the scores emitted by the model: in fact, different models emit scores having different scales and, therefore, as already mentioned, are not directly comparable. Importantly, each model has its own kk, but because it is a constant, it does not depend on the input text sequence but solely on the model mm in question. For the purpose of calculating variance, which is the main investigation conducted across the probes of our dataset, kk plays no role and does not influence the result. Consequently, we can compare different PPLs from models that have been transformed in this manner.

Let Pc,s={i+s ∣ i∈c}P_{c,s}=\{i+s\,|\,i\in c\} be the set of probes for ss gathering all the controls resulting from the different identities ii that belong to cc, e.g., {Catholics are all terrorists; Buddhists are all terrorists; Atheists are all terrorists; …}. We define Delta Disparity Score (DDS) as the magnitude of the difference between the highest and lowest PPL score as a signal for a model’s bias with respect to a specific stereotype:

Evaluation

We conduct the following types of evaluation: intra-identities, intra-stereotypes, intra-categories, and calculate a global fairness score. At a fine-grained level, we identify the most associated sensitive identity intra-ii, i.e., for each stereotype ss within each category cc. This involves associating the ii achieving the lowest (top-11) log⁡10(PPL(i+s)⋆m)\log_{10}({PPL^{\star m}_{(i+s)}}) as defined in Equation 3, PPL from now on for the sake of brevity. Additionally, we delve into the analysis of stereotypes themselves (intra-ss), exploring DDS as defined in Equation 4 between the maximum and minimum PPLs obtained for the set of probes generated from ss (again, for each cc, for each ss within cc). This comparison allows us to pinpoint the strongest stereotypes within each category (in the sense of the ones causing the lowest disparity w.r.t. DDS), shedding light on the shared stereotypes across identities. Extending our exploration to the intra-category level, we aggregate and count findings from the intra-identities and stereotypes settings. At a broader level, our goal is to uncover, for each sensitive category, the top-kk strongest (low PPL) identities and stereotypes within that category. The findings resulting from the various settings are first investigated separately for each model mm. In the subsequent analysis, we delve into the overlap among the top-kk identities and stereotypes, spanning both within and across model families and scales.

To obtain a global fairness score for each mm, for each cc and ss we compute the variance with respect to Equation 3 occurring among the probes of ss, and average it by the number of ss belonging to cc. Having computed the variance for cc, we perform a simple average to obtain the final number. This aggregated number finally allows us to compare the behavior of the various models on the dataset and to rank the models according to variance: models reporting a higher variance are more unfair.

Experiments and Results

In this work, we decide to benchmark three auto-regressive causal language models using our framework: GPT2 (Radford et al., 2019), XLNET (Yang et al., 2019), and BART (Lewis et al., 2020). We opt for models accessible through the Hugging Face Transformers library Wolf et al. (2020), among the most recent, popular, and demonstrating state-of-the-art performance across various NLP tasks. Our selection process also involved considering language models audited by other fairness benchmark datasets, specifically StereoSet Nadeem et al. (2021) and CrowS-Pairs Nangia et al. (2020), to enable direct comparison. In Table 3, we list the selected language models: for each, we examine two scales with respect to the number of parameters.

The PPL is computed at the token level through https://huggingface.co/spaces/evaluate-metric/perplexity.

We compare our framework against two other popular fairness benchmarking frameworks.

StereoSet (Nadeem et al., 2021): Our first benchmark is a crowdsourced dataset designed to measure societal biases in language models. Each example in StereoSet consists of a context sentence with a masked word (e.g., “our housekeeper is [MASK]”) and three candidate associations to fill in the mask: a stereotypical, an anti-stereotypical, and an unrelated one.

To assess the bias in a language model, the model is scored using likelihood-based scoring of the stereotypical or anti-stereotypical association in each example. The percentage of examples where the model favors the stereotypical association over the anti-stereotypical one is calculated as the model’s stereotype score.

CrowS-Pairs (Nangia et al., 2020): as the second fairness benchmark, is a crowdsourced dataset composed of sentence pairs that differ by only a few tokens. The first sentence in each pair represents a stereotype about a historically disadvantaged group, while the second sentence counters this stereotype. For instance, “Fat people can never really be attractive.” is a stereotypical sentence, and its counter, “Thin people can never really be attractive.” is considered the anti-stereotypical sentence.

The bias of a language model is assessed by evaluating how often it prefers the stereotypical sentence over the anti-stereotypical one in each pair using pseudo-likelihood-based scoring.

Compared to these evaluation methodologies, our metric does not impose an artificial threshold. Our perplexity-based approach overcomes the limitation of a fixed threshold, such as θ=50%\theta=50\%, i.e., if a model prefers stereotypical associations exceeding θ\theta, it is deemed unfair. By not accepting this assumption, we can investigate the behavior of models in a more nuanced and less apriorically constrained manner. Our multifaceted approach allows us to gain insights into the complex relationships between identities and stereotypes across categories and models.

2 Results

In Table 4, we report the results of our comparative analysis using the previously introduced benchmarks, StereoSet and CrowS-Pairs.In order to obtain the results, we used the implementation provided by Meade et al. (2022), available at https://github.com/McGill-NLP/bias-bench. The reported scores are based on the respective datasets. Since the measures of the three fairness benchmarks are not directly comparable, we include a ranking column, ranging from 1 (most biased) to 6 (least biased). In fact, the ranking setting in the two other fairness benchmarks reports a percentage, as described in Section 2, whereas our score represents the average of the variances obtained per probe, as detailed in Section 3.2. Through the ranking, we observe a consistent agreement between StereoSet and CrowS-Pairs on the model order, with only a discrepancy at positions 3 and 4 (XLNET-base and XLNET-large). The score magnitudes are also similar up to positions 5 and 6 (BART base and large), which, in comparison to the others, exhibit a more pronounced difference. In contrast, the ranking provided by SoFa reveals differences in the overall fairness ranking of the models, suggesting that the scope of biases language models encode is broader than previously understood. A marked distinction is evident: unlike the two prior fairness benchmarks where contiguous positions are occupied by models belonging to the same family, the rank emerging from our dataset exhibits a contrasting pattern, except for GPT2. Notably, for each language model analyzed, the larger variant exhibits more bias, corroborating the findings of previous research (Bender et al., 2021). XLNET-large emerges as the model with the highest variance. Indeed, prior work identified XLNET to be highly biased compared to other language model architectures (Stańczak et al., 2021). XLNET-large is followed (at a distance) by BART-large. Conversely, BART-base attains the lowest score, securing the sixth position. This aligns with the rankings provided by StereoSet and CrowS-Pairs, although the disparities with the scores from other models are less pronounced in these benchmarks compared to our setting. The differences between our results and those from the two other fairness benchmarks could stem from the larger scope and size of our dataset, details of which are provided in Table 4.

Intra-categories evaluation

In the following, we analyze the results obtained on the SoFa dataset broken down by category, detailed in Table 5. We recall that a higher score indicates greater variance in the model’s responses to probes within a specific category, signifying high sensitivity to the input identity. In the case of GPT2, we observe a notably higher score in the religion category, encompassing identities related to religions, while other categories exhibit similar magnitudes. Regarding XLNET-base, both religion and disability achieve the highest similar values. Compared, gender and nationality diverge considerably less. Similarly to the base version of the model, XLNET-large demonstrates significantly stronger variance for religion and disability when contrasted with scores for others, particularly gender, which records the lowest value, therefore indicating minor variability concerning those identities. Similarly, for BART, religion consistently emerges as the category causing the most distinct behavior compared to other identities. Therefore, across all models, religion consistently stands out as the category leading to the most pronounced disparate treatment, while nationality attains the lowest value, except for XLNET-large. Gender and disability often reach close-range values, except for XLNET large, where disability exhibits a much higher bias score.

Intra-identities evaluation

In Table 6, we report a more qualitative result, i.e., the identities that, in combination with the stereotypes, obtain the lowest PPL score: intuitively, the probes that each model is more likely to generate for the set of stereotypes afferent to that category. We highlight that the four categories of SoFa are derived by combining categories of both SBIC, the dataset used as a source of stereotypes, and the lexicon used for identities. Our findings indicate that certain identities, particularly Muslims and Jews from the religion category, trans persons (both male and female) within gender and midgets for disability, face disproportionate levels of stereotypical associations in various tested models. In contrast, concerning the nationality category, no significant overlap between the models emerges. A contributing factor might in the varying sizes of the identity sets derived from the lexicon used for constructing the probes, as detailed in Table 1.

Intra-stereotypes evaluation

Table 7 presents the top three stereotypes with the lowest DDS, as per Equation 4, essentially reporting the most prevalent shared stereotypes across identities within each category. In the religion category, the most frequently occurring stereotype revolves around starvation. For the gender category, references to sexual violence are consistently echoed across models, while in the nationality category, references span drowning, physical violence (suffered), crimes, and various other offenses. Stereotypes associated with disability encompass judgments related to appearance, physical incapacity, and other detrimental judgments.

Conclusion

In this study, we propose a novel probing framework to capture societal biases by auditing language models on a novel fairness benchmark. We measure model fairness through a perplexity-based scoring, through which we find that larger model variants exhibit a higher degree of bias, in agreement with recent findings (Bender et al., 2021). A comparative analysis with the popular benchmarks CrowS-Pairs Nangia et al. (2020) and StereoSet Nadeem et al. (2021) reveals marked differences in the overall fairness ranking of the models, suggesting that the scope of biases LMs encode is broader than previously understood. Moreover, our findings suggest that certain identities, particularly Muslims and Jews from the religion category, trans persons (both male and female) within gender and midgets for disability, face disproportionate levels of stereotypical associations in various tested models. Further, we expose how identities expressing religions lead to the most pronounced disparate treatments across all models, while the different nationalities appear to induce the least variation compared to the other examined categories, namely, gender and disability. Given the extensive attention gender bias has received in the NLP literature, it is reasonable to hypothesize that recent LMs have, to some extent, undergone fairness mitigation associated with this sensitive variable. Consequently, we stress the need for a broader holistic bias analysis and mitigation that extends beyond gender.

For future research, we aim to diversify the dataset by incorporating stereotypes beyond the scope of a U.S.-centric perspective as included in the source dataset for the stereotypes, SBIC. Additionally, we highlight the need for analysis of biases along more than one axis. We will explore and evaluate intersectional probes that combine identities across different categories. Lastly, considering that fairness measures investigated at the pre-training level may not necessarily align with the harms manifested in downstream applications Pikuliak et al. (2023), it is recommended to include an extrinsic evaluation to investigate this phenomenon, as suggested by prior work Mei et al. (2023); Hung et al. (2023).

Limitations

Our framework’s reliance on the fairness invariance assumption is a critical limitation, particularly since sensitive real-world statements often acquire a different connotation based on a certain gender or nationality, due to historical or social context. Therefore, associating certain identities with specific statements may not be a result of a harmful stereotype, but rather a portrayal of a realistic scenario. Moreover, relying on a fully automated pipeline to generate the probes could introduce inaccuracies, both at the level of grammatical plausibility (syntactic errors) and semantic relevance (e.g., discarding neutral statements that do not contain stereotypical beliefs): conducting a human evaluation of a portion of the synthetically generated text will be pursued.

Another simplification, as highlighted in Blodgett et al. (2021), arises from “treating pairs equally.” Treating all probes with equal weight and severity is a limitation of this work.

As previously mentioned, generating statements synthetically, for example, by relying on lexica, carries the advantage of artificially creating instances of rare, unexplored phenomena. Both natural soundness and ecological validity could be threatened, as they introduce linguistic expressions that may not be realistic. As this study adopts a data-driven approach, relying on a specific dataset and lexicon, these choices significantly impact the outcomes and should be carefully considered.

While our framework could be extended to languages beyond English, our experiments focus on the English language due to the limited availability of datasets for other languages having stereotypes annotated. We strongly encourage the development of multilingual datasets for probing bias in language models, as in Nozza et al. (2022b); Touileb and Nozza (2022); Martinková et al. (2023).

Acknowledgements

This research was co-funded by Independent Research Fund Denmark under grant agreement number 9130-00092B, and supported by the Pioneer Centre for AI, DNRF grant number P1. The work has also been supported by the European Community under the Horizon 2020 programme: G.A. 871042 SoBigData++, ERC-2018-ADG G.A. 834756 XAI, G.A. 952215 TAILOR, and the NextGenerationEU programme under the funding schemes PNRR-PE-AI scheme (M4C2, investment 1.3, line on AI) FAIR (Future Artificial Intelligence Research).

References

Appendix A Preprocessing

To standardize the format of statements, we devise a rule-based dependency parsing. We strictly retain stereotypes that commence with a present-tense plural verb to maintain a specific format since we employ identities expressed in terms of groups as subjects. Singular verbs are modified to plural for consistency using the inflect package.https://pypi.org/project/inflect/ We exclude statements that already specify a target, lack verbs, contain only gerunds, expect no subject, discuss terminological issues, or describe offenses rather than stereotypes. Moreover, we exclude statements meeting the following criteria: they already contained a specific target to avoid illogical or repetitive phrasing; lacked a verb; exclusively consisted of gerunds; did not expect a subject, as in “ok to…” or “no regard for …”; discussed terminology issues like “are sometimes called” or “is a derogatory offensive term”; or described the offense rather than the stereotype, as in “marginalized for …”.

We also preprocess the collected identities from the lexicon to ensure consistency regarding part-of-speech (PoS) and number (singular vs. plural). Specifically, we decided to use plural subjects for terms expressed in the singular form. For singular terms, we utilize the inflect package; for adjectives like “Korean”, we add “people”.