"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, Nanyun Peng
Introduction
LLMs have emerged as helpful tools to facilitate the generation of coherent long texts, enabling various use cases of document generation Sallam (2023); Osmanovic-Thunström et al. (2023); Stokel-Walker (2023); Hallo-Carrasco et al. (2023). Recently, there has been a growing trend to use LLMs in the creation of professional documents, including recommendation letters. The use of ChatGPT for assisting reference letter writing has been a focal point of discussion on social media platformsSee, for example, the discussion on Reddit https://shorturl.at/eqsV6 and reports by major media outletsFor example, see the article published in the Atlantic https://shorturl.at/fINW3..
However, the widespread use of automated writing techniques without careful scrutiny can entail considerable risks. Recent studies have shown that Natural Language Generation (NLG) models are gender biased Sheng et al. (2019, 2020); Dinan et al. (2020); Sheng et al. (2021a); Bender et al. (2021) and therefore pose a risk to harm minorities when used in sensitive applications Sheng et al. (2021b); Ovalle et al. (2023a); Prates et al. (2018). Such biases might also infiltrate the application of automated reference letter generation and cause substantial societal harm, as research in social sciences Madera et al. (2009); Khan et al. (2021) unveiled how biases in professional documents lead to diminished career opportunities for gender minority groups. We posit that inherent gender biases in LLMs manifests in the downstream task of reference letter generation. As an example, Table 1 demonstrates reference letters generated by ChatGPT for candidates with popular male and female names. The model manifests the stereotype of men being agentic (e.g., natural leader) and women being communal (e.g., well-liked member).
In this paper, we systematically investigate gender biases present in reference letters generated by LLMs under two scenarios: (1) Context-Less Generation (CLG), where the model is prompted to produce a letter based solely on simple descriptions of the candidate, and (2) Context-Based Generation (CBG), in which the model is also given the candidate’s personal information and experience in the prompt. CLG reveals inherent biases towards simple gender-associated descriptors, whereas CBG simulates how users typically utilize LLMs to facilitate letter writing. Inspired by social science literature, we investigate aspects of biases in LLM-generated reference letters: (1) bias in lexical content, (2) bias in language style, and (3) hallucination bias. We construct the first comprehensive testbed with metrics and prompt datasets for identifying and quantifying biases in the generated letters. Furthermore, we use the proposed framework to evaluate and unveil significant gender biases in recommendation letters generated by two recently developed LLMs: ChatGPT OpenAI (2022) and Alpaca Taori et al. (2023).
Our findings emphasize a haunting reality: the current state of LLMs is far from being mature when it comes to generating professional documents. We hope to highlight the risk of potential harm when LLMs are employed in such real-world applications: even with the recent transformative technological advancements, current LLMs are still marred by gender biases that can perpetuate societal inequalities. This study also underscores the urgent need for future research to devise techniques that can effectively address and eliminate fairness concerns associated with LLMs.Code and data are available at: https://github.com/uclanlp/biases-llm-reference-letters
Related Work
Social biases in NLP models have been an important field of research. Prior works have defined two major types of harms and biases in NLP models: allocational harms and representational harms Blodgett et al. (2020); Barocas et al. (2017); Crawford (2017). Researchers have studied methods to evaluate and mitigate the two types of biases in Natural Language Understanding (NLU) Bolukbasi et al. (2016); Dev et al. (2022); Dixon et al. (2018); Bordia and Bowman (2019); Zhao et al. (2017, 2018); Sun and Peng (2021) and Natural Language Generation (NLG) tasks Sheng et al. (2019, 2021b); Dinan et al. (2020); Sheng et al. (2021a).
Among previous works, Sun and Peng (2021) proposed to use the Odds Ratio (OR) Szumilas (2010) as a metric to measure gender biases in items with large frequency differences or highest saliency for females and males. Sheng et al. (2019) measured biases in NLG model generations conditioned on certain contexts of interest. Dhamala et al. (2021a) extended the pipeline to use real prompts extracted from Wikipedia. Several approaches Sheng et al. (2020); Gupta et al. (2022); Liu et al. (2021); Cao et al. (2022) studied how to control NLG models for reducing biases. However, it is unclear if they can be applied in closed API-based LLMs, such as ChatGPT.
2 Biases in Professional Documents
Recent studies in NLP fairness Wang et al. (2022a); Ovalle et al. (2023b) point out that some AI fairness works fail to discuss the source of biases investigated, and suggest to consider both social and technical aspects of AI systems. Inspired by this, we ground bias definitions and metrics in our work on related social science research. Previous works in social science Cugno (2020); Madera et al. (2009); Khan et al. (2021); Liu et al. (2009); Madera et al. (2019) have revealed the existence and dangers of gender biases in the language styles of professional documents. Such biases might lead to harmful gender differences in application success rate Madera et al. (2009); Khan et al. (2021). For instance, Madera et al. (2009) observed that biases in gendered language in letters of recommendation result in a higher residency match rate for male applicants. These findings further emphasize the need to study gender biases in LLM-generated professional documents. We categorize major findings in previous literature into types of gender biases in language styles of professional documents: biases in language professionalism, biases in language excellency, and biases in language agency.
Bias in language professionalism states that male candidates are considered more “professional” than females. For instance, Trix and Psenka (2003) revealed the gender schema where women are seen as less capable and less professional than men. Khan et al. (2021) also observed more mentions of personal life in letters for female candidates. Gender biases in this dimension will lead to biased information on candidates’ professionalism, therefore resulting in unfair hiring evaluation.
Bias in language excellency states that male candidates are described using more “excellent” language than female candidates in professional documents Trix and Psenka (2003); Madera et al. (2009, 2019). For instance, Dutt et al. (2016) points out that female applicants are only half as likely than male applicants to receive “excellent” letters. Naturally, gender biases in the level of excellency of language styles will lead to a biased perception of a candidate’s abilities and achievements, creating inequality in hiring evaluation.
Bias in language agency states that women are more likely to be described using communal adjectives in professional documents, such as delightful and compassionate, while men are more likely to be described using “agentic” adjectives, such as leader or exceptional Madera et al. (2009, 2019); Khan et al. (2021). Agentic characteristics include speaking assertively, influencing others, and initiating tasks. Communal characteristics include concerning with the welfare of others, helping others, accepting others’ direction, and maintaining relationships Madera et al. (2009). Since agentic language is generally perceived as being more hirable than communal language style Madera et al. (2009, 2019); Khan et al. (2021), bias in language agency might further lead to biases in hiring decisions.
3 Hallucination Detection
Understanding and detecting hallucinations in LLMs have become an important problem Mündler et al. (2023); Ji et al. (2023); Azamfirei et al. (2023). Previous works on hallucination detection proposed three main types of approaches: Information Extraction-based, Question Answering (QA)-based and Natural Language Inference (NLI)-based approaches. Our study utilizes the NLI-based approach Kryscinski et al. (2020); Maynez et al. (2020); Laban et al. (2022), which uses the original input as context to determine the entailment with the model-generated text. To do this, prior works have proposed document-level NLI and sentence-level NLI approaches. Document-level NLI Maynez et al. (2020); Laban et al. (2022) investigates entailment between full input and generation text. Sentence-level NLI Laban et al. (2022) chunks original and generated texts into sentences and determines entailment between each pair. However, little is known about whether models will propagate or amplify biases in their hallucinated outputs.
Methods
We consider two different settings for reference letter generation tasks. (1) Context-Less Generation (CLG): prompting the model to generate a letter based on minimal information, and (2) Context-Based Generation (CBG): guiding the model to generate a letter by providing contextual information, such as a personal biography. The CLG setting better isolates biases influenced by input information and acts as a lens to examine underlying biases in models. The CBG setting aligns more closely with the application scenarios: it simulates a user scenario where the user would write a short description of themselves and ask the model to generate a recommendation letter accordingly.
2 Bias Definitions
We categorize gender biases in LLM-generated professional documents into two types: Biases in Lexical Content, and Biases in Language Style.
Biases in lexical content can be manifested by harmful differences in the most salient components of LLM-generated professional documents. In this work, we measure biases in lexical context through evaluating biases in word choices. We define biases in word choices to be the salient frequency differences between wordings in male and female documents. We further dissect our analysis into biases in nouns and biases in adjectives.
Odds Ratio Inspired by previous work Sun and Peng (2021), we propose to use Odds Ratio (OR) Szumilas (2010) for qualitative analysis on biases in word choices. Taking analysis on adjectives as an example. Let and be the set of all adjectives in male documents and female documents, respectively. For an adjective , we first count its occurrences in male documents and in female documents . Then, we can calculate OR for adjective to be its odds of existing in the male adjectives list divided by the odds of existing in the female adjectives list:
Larger OR means that an adjective is more likely to exist, or more salient, in male letters than female letters. We then sort adjectives by their OR in descending order, and extract the top and last adjectives, which are the most salient adjectives for males and for females respectively.
2.2 Biases in Language Style
We define biases in language style as significant stylistic differences between LLM-generated documents for different gender groups. For instance, we can say that bias in language style exists if the language in model-generated documents for males is significantly more positive or more formal than that for females. Given two sets of model-generated documents for males and females , we can measure the extent that a given text conforms to a certain language style by a scoring function . Then, we can measure biases in language style through t-testing on language style differences between and . Biases in language style can therefore be mathematically formulate as:
where and represents sample mean and standard deviation. Due to the nature of as a t-test value, a small value of that is lower than the significance threshold indicates the existence of bias. Following the bias aspects in social science that are discussed in Section 2.2, we establish aspects to measure biases in language style: (1) Language Formality, (2) Language Positivity, and (3) Language Agency.
Biases in Language Formality Our method uses language formality as a proxy to reflect the level of language professionalism. We define biases in Language Formality to be statistically significant differences in the percentage of formal sentences in male and female-generated documents. Specifically, we conduct statistical t-tests on the percentage of formal sentences in documents generated for each gender and report the significance of the difference in formality levels.
Biases in Language Positivity Our method uses positive sentiment in language as a proxy to reflect the level of excellency in language. We define biases in Language Positivity to be statistically significant differences in the percentage of sentences with positive sentiments in generated documents for males and females. Similar to analysis for biases in language formality, we use statistical t-testing to construct the quantitative metric.
Biases in Language Agency We propose and study Language Agency as a novel metric for bias evaluation in LLM-generated professional documents. Although widely observed and analyzed in social science literature Cugno (2020); Madera et al. (2009); Khan et al. (2021), biases in language agency have not been defined, discussed or analyzed in the NLP community. We define biases in language agency to be statistically significant differences in the percentage of agentic sentences in generated documents for males and females, and again report the significance of biases using t-testing.
3 Hallucination Bias
In addition to directly analyzing gender biases in model-generated reference letters, we propose to separately study biases in model-hallucinated information for CBG task. Specifically, we want to find out if LLMs tend to hallucinate biased information in their generations, other than factual information provided from the original context. We define Hallucination Bias to be the harmful propagation or amplification of bias levels in model hallucinations.
Hallucination Detection Inspired by previous works Maynez et al. (2020); Laban et al. (2022), we propose and utilize Context-Sentence NLI as a framework for Hallucination Detection. The intuition behind this method is that the source knowledge reference should entail the entirety of any generated information in faithful and hallucination-free generations. Specifically, given a context and a corresponding model generated document , we first split D into sentences as hypotheses. We use the entirety of as the premise and establish premise-hypothesis pairs: Then, we use an NLI model to determine the entailment between each premise-hypothesis pair. Generated sentences in non-entailment pairs are considered as hallucinated information. The detected hallucinated information is then used for hallucination bias evaluation. A visualization of the hallucination detection pipeline is demonstrated in Figure 1.
Hallucination Bias Evaluation In order to measure gender bias propagation and amplification in model hallucinations, we utilize the same quantitative metrics as evaluation of Biases in Language Style: Language Formality, Language Positivity, and Language Agency. Since our goal is to investigate if information in model hallucinations demonstrates the same level or a higher level of gender biases, we conduct statistical t-testing to reveal significant harmful differences in language styles between only the hallucinated content and the full generated document. Taking language formality as an example, we conduct a t-test on the percentage of formal sentences in the detected hallucinated contents and the full generated document, respectively. For male documents, bias propagation exists if the hallucinated information does not demonstrate significant differences in levels of formality, positivity, or agency. Bias amplification exists if the hallucinated information demonstrates significantly higher levels of formality, positivity, or agency than the full document. Similarly, for female documents, bias propagation exists if hallucination is not significantly different in levels of formality, positivity, or agency. Bias amplification exists if hallucinated information is significantly lower in its levels of formality, positivity, or agency than the full document.
Experiments
We conduct bias evaluation experiments on two tasks: Context-Less Generation and Context-Based Generation. In this section, we first briefly introduce the setup of our experiments. Then, we present an in-depth analysis of the method and results for the evaluation on CLG and CBG tasks, respectively. Since CBG’s formulation is closer to real-world use cases of reference letter generation, we place our research focus on CBG task, while conducting a preliminary exploration on CLG biases.
Since experiments on CLG act as a preliminary exploration, we only use ChatGPT as the model for evaluation. To choose the best models for experiments CBG task, we investigate the generation qualities of four LLMs: ChatGPT OpenAI (2022), Alpaca Taori et al. (2023), Vicuna Chiang et al. (2023), and StableLM AI (2023). While ChatGPT can always produce reasonable reference letter generations, other LLMs sometimes fail to do so, outputting unrelated content. In order to only evaluate valid reference letter generations, we define and calculate the generation success rate of LLMs using criteria-based filtering. Details on generation success rate calculation and behavior analysis can be found in Appendix B. After evaluating LLMs’ generation success rates on the task, we choose to conduct further experiments using only ChatGPT and Alpaca for letter generations.
2 Context-Less Generation
Analysis on CLG evaluates biases in model generations when given minimal context information, and acts as a lens to interpret underlying biases in models’ learned distribution.
Prompting Brown et al. (2020); Sun and Lai (2020) steers pre-trained language models with task-specific instructions to generate task outputs without task fine-tuning. In our experiments, we design simple descriptor-based prompts for CLG analysis. We have attached the full list of descriptors in Appendix C.1, which shows the three axes (name/gender, age, and occupation) and corresponding specific descriptors (e.g. Joseph, 20, student) that we iterate through to query model generations. We then formulate the prompt by filling descriptors of each axis in a prompt template, which we have attached in Appendix C.2. Using these descriptors, we generated a total of CLG-based reference letters. Hyperparameter settings for generation can be found in Appendix A.
2.2 Evaluation: Biases in Lexical Content
Since only letters were generated for preliminary CLG analysis, running statistics analysis on biases in lexical content or word choices might lack significance as we calculate OR for one word at a time. To mitigate this issue, we calculate OR for words belonging to gender-stereotypical traits, instead of for single words. Specifically, we implement the traits as lexicon categories: Ability, Standout, Leadership, Masculine, Feminine, Agentic, Communal, Professional, and Personal. Full lists of the lexicon categories can be found in Appendix F.3. An OR score that is greater than indicates higher odds for the trait to appear in generated letters for males, whereas an OR score that is below indicates the opposite.
2.3 Result
Table 2 shows experiment results for biases in lexical content analysis on CLG task, which reveals significant and harmful associations between gender and gender-stereotypical traits. Most male-stereotypical traits – Ability, Standout, Leadership, Masculine, and Agentic – have higher odds of appearing in generated letters for males. Female-stereotypical traits – Feminine, Communal, and Personal – also demonstrate the same trend to have higher odds of appearing in female letters. Evaluation results on CLG unveil significant underlying gender biases in ChatGPT, driving the model to generate reference letters with harmful gender-stereotypical traits.
3 Context-Based Generation
Analysis on CBG evaluates biases in model generations when provided with certain context information. For instance, a user can input personal information such as a biography and prompt the model to generate a full letter.
We utilize personal biographies as context information for CBG task. Specifically, we further preprocess and use WikiBias Sun and Peng (2021), a personal biography dataset with scraped demographic and biographic information from Wikipedia. Our data augmentation pipeline aims at producing an anonymized and gender-balanced biography datasest as context information for reference letter generation to prevent pre-existing biases. Details on preprocessing implementations can be found in Appendix F.1. We denote the biography dataset after preprocessing as WikiBias-Aug, statistics of which can be found in Appendix D.
3.2 Generation
Prompt Design Similar to CLG experiments, we use prompting to obtain LLM-generated professional documents. Different from CLG, CBG provides the model with more context information in the form of personal biographies in the input. Specifically, we use biographies in the pre-processed WikiBias-Aug dataset as contextual information. Templates used to prompt different LLMs are attached in Appendix C.3. Generation hyper-parameter settings can be found in Appendix A.
Generating Reference Letters We verbalize biographies in the WikiBias-Aug dataset with the designed prompt templates and query LLMs with the combined information. Upon filtering out unsuccessful generations with the criterion defined in Section 4.1, we get generations for ChatGPT and successful generations for Alpaca.
3.3 Evaluation: Biases in Lexical Content
Given our aim to investigate biases in nouns and adjectives as lexical content, we first extract words of the two lexical categories in professional documents. To do this, we use the Spacy Python library Honnibal and Montani (2017) to match and extract all nouns and adjectives in the generated documents for males and females. After collecting words in documents, we create a noun dictionary and an adjective dictionary for each gender to further apply the odds ratio analysis.
3.4 Evaluation: Biases in Language Style
In accordance with the definitions of the three types of gender biases in the language style of LLM-generated documents in Section 3.2.2, we implement three corresponding metrics for evaluation.
Biases in Language Formality For evaluation of biases in language formality, we first classify the formality of each sentence in generated letters, and calculate the percentage of formal sentences in each generated document. To do so, we apply an off-the-shelf language formality classifier from the Transformers Library that is fine-tuned on Grammarly’s Yahoo Answers Formality Corpus (GYAFC) Rao and Tetreault (2018). We then conduct statistical t-tests on formality percentages in male and female documents to report significance levels.
Biases in Language Positivity Similarly, for evaluation of biases in language positivity, we calculate and conduct t-tests on the percentage of positive sentences in each generated document for males and females. To do so, we apply an off-the-shelf language sentiment analysis classifier from the Transformers Library that was fine-tuned on the SST-2 dataset Socher et al. (2013).
Language Agency Classifier Along similar lines, for evaluation of biases in language agency, we conduct t-tests on the percentage of agentic sentences in each generated document for males and females. Implementation-wise, since language agency is a novel concept in NLP research, no previous study has explored means to classify agentic and communal language styles in texts. We use ChatGPT to synthesize a language agency classification corpus and use it to fine-tune a transformer-based language agency classification model. Details of the dataset synthesis and classifier training process can be found in Appendix F.2.
3.5 Result
Biases in Lexical Content Table 3 shows results for biases in lexical content on ChatGPT and Alpaca. Specifically, we show the top salient adjectives and nouns for each gender. We first observe that both ChatGPT and Alpaca tend to use gender-stereotypical words in the generated letter (e.g. “respectful” for males and “warm” for females). To produce more interpretable results, we run WEAT score analysis with two sets of gender-stereotypical traits: i) male and female popular names (WEAT (MF)) and ii) career and family-related words (WEAT (CF)), full word lists of which can be found in Appendix F.4. WEAT takes two lists of words (one for male and one for female) and verifies whether they have a smaller embedding distance with female-stereotypical traits or male-stereotypical traits. A positive WEAT score indicates a correlation between female words and female-stereotypical traits, and vice versa. A negative WEAT score indicates that female words are more correlated with male-stereotypical traits, and vice versa. To target words that potentially demonstrate gender stereotypes, we identify and highlight words that could be categorized within the nine lexicon categories in Table 2, and run WEAT test on these identified words. WEAT score result reveals that the most salient words in male and female documents are significantly associated with gender-stereotypical lexicon.
Biases in Language Style Table 4 shows results for biases in language style on ChatGPT and Alpaca. T-testing results reveal gender biases in the language styles of documents generated for both models, showing that male documents are significantly higher than female documents in all three aspects: language formality, positivity, and agency. Interestingly, our experiment results align well with social science findings on biases in language professionalism, language excellency, and language agency for human-written reference letters.
To unravel biases in model-generated letters in a more intuitive way, we manually select a few snippets from ChatGPT’s generations that showcase biases in language agency. Each pair of grouped texts in Table 5 is sampled from the 2 generated letters for male and female candidates with the same original biography information. After preprocessing by gender swapping and name swapping, the original biography was transformed into separate input information for two candidates of opposite genders. We observe that even when provided with the exact same career-related information despite name and gender, ChatGPT still generates reference letters with significantly biased levels of language agency for male and female candidates. When describing female candidates, ChatGPT uses communal phrases such as “great to work with”, “communicates well”, and “kind”. On the contrary, the model tends to describe male candidates as being more agentic, using narratives such as “a standout in the industry” and “a true original”.
4 Hallucination Bias
We use the proposed Context-Sentence NLI framework for hallucination detection. Specifically, we implement an off-the-shelf RoBERTa-Large-based NLI model from the Transformers Library that was fine-tuned on a combination of four NLI datasets: SNLI Bowman et al. (2015), MNLI Williams et al. (2018), FEVER-NLI Thorne et al. (2018), and ANLI (R1, R2, R3) Nie et al. (2020). We then identify bias exacerbation in model hallucination along the same three dimensions as in Section 4.3.4, through t-testing on the percentage of formal, positive, and agentic sentences in the hallucinated content compared to the full generated letter.
4.2 Result
As shown in Table 6, both ChatGPT and Alpaca demonstrate significant hallucination biases in language style. Specifically, ChatGPT hallucinations are significantly more formal and more positive for male candidates, whereas significantly less agentic for female candidates. Alpaca hallucinations are significantly more positive for male candidates, whereas significantly less formal and agentic for females. This reveals significant gender bias propagation and amplification in LLM hallucinations, pointing to the need to further study this harm.
To further unveil hallucination biases in a straightforward way, we also manually select snippets from hallucinated parts in ChatGPT’s generations. Each pair of grouped texts in Table 7 is selected from two generated letters for male and female candidates given the same original biography information. Hallucinations in the female reference letters use communal language, describing the candidate as having an “easygoing nature”, and “is a joy to work with”. Hallucinations in the male reference letters, in contrast, use evidently agentic descriptions of the candidate, such as “natural talent”, with direct mentioning of “professionalism”.
Conclusion and Discussion
Given our findings that gender biases do exist in LLM-generated reference letters, there are many avenues for future work. One of the potential directions is mitigating the identified gender biases in LLM-generated recommendation letters. For instance, an option to mitigate biases is to instill specific rules into the LLM or prompt during generation to prevent outputting biased content. Another direction is to explore broader areas of our problem statement, such as more professional document categories, demographics, and genders, with more language style or lexical content analyses. Lastly, reducing and understanding the biases with hallucinated content and LLM hallucinations is an interesting direction to explore.
The emergence of LLMs such as ChatGPT has brought about novel real-world applications such as reference letter generation. However, fairness issues might arise when users directly use LLM-generated professional documents in professional scenarios. Our study benchmarks and critically analyzes gender bias in LLM-assisted reference letter generation. Specifically, we define and evaluate biases in both Context-Less Generation and Context-Based Generation scenarios. We observe that when given insufficient context, LLMs default to generating content based on gender stereotypes. Even when detailed information about the subject is provided, they tend to employ different word choices and linguistic styles when describing candidates of different genders. What’s more, we find out that LLMs are propagating and even amplifying harmful gender biases in their hallucinations.
We conclude that AI-assisted writing should be employed judiciously to prevent reinforcing gender stereotypes and causing harm to individuals. Furthermore, we wish to stress the importance of building a comprehensive policy of using LLM in real-world scenarios. We also call for further research on detecting and mitigating fairness issues in LLM-generated professional documents, since understanding the underlying biases and ways of reducing them is crucial for minimizing potential harms of future research on LLMs.
Limitations
We identify some limitations of our study. First, due to the limited amount of datasets and previous literature on minority groups and additional backgrounds, our study was only able to consider the binary gender when analyzing biases. We do stress, however, the importance of further extending our study to fairness issues for other gender minority groups as future works. In addition, our study primarily focuses on reference letters to narrow the scope of analysis. We recognize that there’s a large space of professional documents now possible due to the emergence of LLMs, such as resumes, peer evaluations, and so on, and encourage future researchers to explore fairness issues in other categories of professional documents. Additionally, due to cost and compute constraints, we were only able to experiment with the ChatGPT API and 3 other open-source LLMs. Future work can build upon our investigative tools and extend the analysis to more gender and demographic backgrounds, professional document types, and LLMs. We believe in the importance of highlighting the harms of using LLMs for these applications and that these tools act as great writing assistants or first drafts of a document but should be used with caution as biases and harms are evident.
Ethics Statement
The experiments in this study incorporate LLMs that were pre-trained on a wide range of text from the internet and have been shown to learn or amplify biases from this data. In our study, we seek to further explore the ethical considerations of using LLMs within professional documents through the representative task of reference letter generation. Although we were only able to analyze a subset of the representative user base of LLMs, our study uncover noticeable harms and areas of concern when using these LLMs for real-world scenarios. We hope that our study adds an additional layer of caution when using LLMs for generating professional documents, and promotes the equitable and inclusive advancement of these intelligent systems.
Acknowledgements
We thank UCLA-NLP+ members and anonymous reviewers for their invaluable feedback. The work is supported in part by CISCO, NSF 2331966, an Amazon Alexa AI gift award and a Meta SRA. KC was supported as a Sloan Fellow.
References
Appendix A Generation Hyperparameter Settings
We use the default parameters of ChatGPT with OpenAI’s chat completion API, which are "GPT-3.5-Turbo" with temperature, top_p, and n set to 1 and no stop token. For Alpaca, Vicuna and StableLM, we configurate the maximum number of new tokens to be , repetition penalty to be , temperature to be , top p to be , and number of beams to be . All configuration hyper-parameters are selected through parameter tuning experiments to ensure best generation performance of each model.
Appendix B Generation Success Rate Analysis
During reference letter generation, we observe that i) ChatGPT can always produce reasonable reference letters, and ii) other LLMs that we investigate sometimes fail to do so. In the following section, we will first briefly show typical examples of generation failure. Then, we will provide our definition and criteria for successful generations. Finally, we compare Alpaca, Vicuna and StableLM in terms of their generation success rate, and argue that Alpaca significantly outperforms the other two models investigated in reference letter generation task.
Table 8 presents the three types of unsuccessful generations of LLMs: empty content, repetitive content, and task divergence.
B.2 Successful Generation
Taking into considerations the failure types of LLM generations, we define a success generation to be nonempty, non-repetitive and task-following (i.e. generating a recommendation letter instead of other types of text). Therefore, we establish criteria as a vanilla way to implement rule-based unsuccessful generation detection. Specifically, we keep generations that are: i) non-empty, ii) does not contain long continuous strings, and iii) contains the word “recommend”.
B.3 Generation Success Rate
We calculate and report the generation success rate of LLMs in Table 9. Overall, Alpaca achieves significantly higher generation success rate than the other LLMs. Therefore, we chose to conduct further evaluation experiments only with generated letters ChatGPT and Alpaca.
Appendix C Prompt Design
C.2 Prompts for CLG Task
C.3 Prompts for CBG Task
Table 12 shows the prompts that we use to query the generation of reference letters for CBG task.
Appendix D Dataset Statistics: WikiBias-Aug
Table 13 shows statistics of the pre-processed WikiBias-Aug dataset.
Appendix E Sample Reference Letter Generations
Context-Less Generation Please see Table 14 for an example of generated reference letter by ChatGPT under CLG scenario.
Context-Based Generation Please see Table 15 for an example of generated reference letter by ChatGPT under CBG scenario.
E.2 Alpaca
Context-Based Generation Please see Table 16 for an example of generated reference letter by Alpaca under CBG scenario.
Appendix F Experiment Details
Evaluation on CBG-based professional document generation requires a dataset with gender-balanced and anonymized contexts to avoid i) pre-existing gender biases and ii) potential model hallucinations triggered by real demographic information, such as names. To this end, we propose and use a data preprocessing pipeline to produce an anonymized and gender-balanced personal biography datasest as context information in CBG-based reference letter generation, which we denote as WikiBias-Aug. In our work, the preprocessing pipeline was built to augment the WikiBias dataset Sun and Peng (2021), a personal biography dataset with scraped demographic information as well as biographic information from Wikipedia. However, the proposed pipeline can also be extended to augmentation on other biography datasets. Due of the inclusion of only binary gender in the WikiBias dataset, our study is also limited to studying biases within the two genders. More details will be discussed in the Limitation section. In this study, each biography entry of the original WikiBias dataset consist of the personal life and career life sections in the Wikipedia description of the person. In order to utilize personal biographies as contexts in our CBG-based evaluation pipeline, we need to construct a more gender-balanced dataset with a certain level of anonymization. In addition, considering LLMs’ input tokens limit, we would need to design methods to control the overall length of the biographies in each entry. Figure 2 provides an illustration of the preprocessing pipeline. We first iterate through all demographic information in the WikiBias dataset to stack all the 1) female first names, 2) male first names, as well as 3) all last names regardless of gender. Since we have the gender information of the person described in each biography, we use it as the ground truth to categorize names of each gender, without introducing noises in gender-stereotypical names. For each entry of the WikiBias dataset, we first randomly select paragraphs from the personal and career life sections in the biography. Next, we make heuristics-based changes to the sampled biography to output a number of male biographies and a number of female biographies. For constructing the male biography, we randomly select a male first name and a last name from the according stacks, and replace all name mentions in the original biography with the new male name. If the original biography is describing a female, we also make sure to flip all gendered pronouns (e.g. her, she, hers) in the sentence to male pronouns. Similarly, for constructing the female biography, we randomly select a female first name and a last name and replace all name mentions in the original biography with the new female name. We also flip the gendered pronouns if the original biography is describing a male.
F.2 Agentic Classifier Training Details
Given that no prior research in the NLP community has covered a classifier to detect agentic vs communal, we opted to create our classifier and dataset. For this approach, we use ChatGPT to synthetically generate an evenly distributed dataset of 400 unique biographies per category. The initial biography is sampled from Bias in Bios dataset De-Arteaga et al. (2019a), which is sourced from online biographies in the Common Crawl corpus. The dataset also includes metadata across several occupations and gender indicators. We prompt ChatGPT to rephrase this initial biography into two versions: one learning towards agentic language style (e.g. leadership) and another leaning towards communal language style. Additionally, we include few-shot examples of both agentic/communal sentences and definitions inspired from social science literature. Given this synthetic dataset of around 600 samples, we build a BERT classifier given an 80/10/10 train/dev/test split. We performed hyperparameter search and ended up with a learning rate of 2e-5, train for 10 epochs, and have a batch size of 16. After training and saving the best performing checkpoints on the validation samples, our model performs around 90% on our test set. The dataset and model checkpoint will be released once the github is public.
F.3 Full List of Lexicon Categories
Table 17 demonstrates the full lists of the nine lexicon categories investigated.
F.4 Word Lists For WEAT Test
Table 18 demonstrates Gendered word lists used for WEAT testing.