Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation
Adithya V Ganesan, Yash Kumar Lal, August Håkan Nilsson, H. Andrew Schwartz
Introduction
Human-level NLP tasks, rooted in computational social science, focus on the link between social or psychological characteristics and language. Example tasks include personality assessment Mairesse and Walker (2006); Kulkarni et al. (2017); Lynn et al. (2020), demographic estimation Sap et al. (2014); Preotiuc-Pietro and Ungar (2018), and mental health-related tasks Coppersmith et al. (2014); Guntuku et al. (2017); Matero et al. (2019). Although using LMs as embeddings or fine-tuning them for human-level NLP tasks is becoming popular V Ganesan et al. (2021); Butala et al. (2021); Yang et al. (2021b), very little is known about zero-shot performance of LLMs on such tasks. {NoHyper}
In this paper, we test the zero-shot performance of a popular LLM, GPT-3, to perform personality trait estimation. We focus on personality traits because they are considered the fundamental characteristics that distinguish people, persisting across cultures, demographics, and time Costa and McCrae (1992); Costa Jr and McCrae (1996). These characteristics are useful for a wide range of social, economic, and clinical applications such as understanding psychological disorders Khan et al. (2005), choosing content for learning styles Komarraju et al. (2011) or occupations Kern et al. (2019), and delivering personalized treatments for mental health issues Bagby et al. (2016). Focusing on zero-shot evaluation of GPT-3 on these fundamental characteristics forms a strong benchmark for understanding how much and what dimensions of traits GPT-3 encodes out-of-the-box. Further, while fine-tuned LMs have only had mixed success beyond lexical approaches Lynn et al. (2020); Kerz et al. (2022), using zero-shot capable LLMs could help lead to better estimates.
The NLP community has a growing interest in understanding the capabilities and failure modes of LLMs Wei et al. (2022a); Yang et al. (2021c), and we explore questions that surround LLMs in the context of fundamental human traits of personality. Zero-shot performance can depend heavily on the explicit information infused in the prompt Lal et al. (2022). Personality, defined by information in its well-established questionnaire tests, presents new opportunities for information infusion.
Our contributions address: (1) what information about personality is useful for GPT-3, (2) how its performance compares to current SotA, (3) the relation between ordinality of outcome labels with performance and (4) whether GPT-3 predictions stay consistent given similar external knowledge.
Background
Psychological traits are stable individual characteristics associated with behaviors, attitudes, feelings, and habits APA (2023). The “Big 5” is a popular personality model that breaks characteristics into five fundamental dimensions, validated across hundreds of studies across cultures, demographics, and time Costa and McCrae (1992); McCrae and John (1992). The approach is rooted in the lexical hypothesis that the most important traits must be encoded in language Goldberg (1990). We investigate all five factors from this model: openness to experience (OPE– intellectual, imaginative and open-minded), conscientiousness (CON– careful, thorough and organized), extraversion (EXT– energized by social and interpersonal interactions), agreeableness (AGR– friendly, good natured, conflict avoidant) and neuroticism (NEU– less secure, anxious, and depressive).
LLMs like PaLM Chowdhery et al. (2022) have shown significant improvement in performance on various NLP tasks Wei et al. (2022b); Suzgun et al. (2022), even without finetuning. There is a growing body of work investigating one of the ubiquitous LLMs, GPT-3, under different settings Wei et al. (2022a); Shi et al. (2022); Bommarito et al. (2023). Inspired by this, we systematically study the ability of GPT-3 to perform personality assessment under zero-shot setting. Following evidence that incorporating knowledge about the task can improve performance Vu et al. (2020); Yang et al. (2021b); Lal et al. (2022), we evaluate the impact of three different types of knowledge to determine which type improves personality estimation.
Modeling personality traits through natural language has been extensively studied using a wide range of approaches, from simple count-based models Pennebaker and Stone (2003); Golbeck et al. (2011) to complex hierarchical neural networks Read et al. (2010); Yang et al. (2021a). Finetuning LMs has become the mainstream approach for this task only recently V Ganesan et al. (2021). With the advent of GPT-3, zero- or few-shot settings have become the primary approach to leverage LLMs in other NLP applications, but are yet untested for personality estimation.
Dataset
To get a sample of language associated with personality, we followed the paradigm set forth in Jose et al. (2022) whereby consenting participants shared their own Facebook posts along with taking a battery of psychological assessments, including the big five personality test (Donnellan et al., 2006; Kosinski et al., 2013). The dataset comprises of 202 participants with outcomes of interests who had also shared their Facebook posts. First, we filter the data to only include user posts from the last year of data collection Eichstaedt et al. (2018). Next, we only retain users for whom we have exactly 20 Facebook posts, similar to the approach described in other human-level NLP works Lynn et al. (2020); Matero et al. (2021). Finally, we anonymize the data by replacing personable identifiable information using SciPy’s Virtanen et al. (2020) NER model. We also remove phone numbers and email IDs using regular expressions. Finally, we are left with anonymized Facebook posts for 142 users and their associated 5 personality traits. This population (all from US) has a gender ratio of 79:18:3 (female:male:others). The age ranged from 21 to 66 (median=37). The big 5 personality trait scores fall in the continuous range of . We discretize the outcome values into the desired number of bins/classes using a quantile discretizer (in Pandas). We explain why we choose to discretize the outcome values in §4.
Experimental Design
In this work, GPT-3 is evaluated in a zero-shot setting. We frame the problem of personality prediction as classifying the degree (i.e. high/low or high/medium/low) to which a person exhibits a trait. Ideally, because the big 5 are considered continuously valued variables (McCrae and Costa Jr, 1989), one would model as a regression task, but we found this simplification to classification necessary to get any meaningful insights from GPT-3’s zero-shot capability. We also investigate the degradation of performance for tertiary classification instead of binary in §5.
We devise a simple, reasonable prompt (Basic)Examples of all prompts are in Appendix Figure 2. to first estimate the ability of GPT-3 to predict the Big 5 personality traits. Building on this, we investigate whether adding external knowledge about these traits helps the model perform better. We use three types of knowledge: (1) Textbook: a concise definition of these traits from Roccas et al. (2002), (2) Wordlist: frequent and infrequent wordsWe use the wordlist from Schwartz et al. (2013). used by people exhibiting those traits, and (3) ItemDesc: survey itemsSee Appendix Table 7 for detailed item descriptions. (a positive and a negative) users responded to, based on which their personality scores were estimated.
The baseline, Wt-Lex, is a ridge regression model from Park et al. 2015 trained on dimensionally reduced feature set of n-grams and LDA-based topics extracted from Kosinski et al. (2013) Facebook data. The number of parameters in this model is orders of magnitude less than GPT-3. Even complex neural models Lynn et al. (2020) have been unsuccessful to surpass its performance. Wt-Lex also produces predictions in the continuous scale within the range of . In order to make a fair comparison with GPT-3, we perform the quantile discretization described in §3 and calculate Macro F1. We evaluate the predictions using macro F1 scores.
Results
Table 1 shows GPT-3’s performance on different personality traits, with and without knowledge. We find that ItemDesc prompts the best performance with GPT-3 on average. Surprisingly, the model is able to directly use survey items (ItemDesc) to predict EXT and CON the best. Utilizing these is hard since it requires relating abstract concepts described in these survey items to the ecological language in the posts. The top frequent and infrequent words (Wordlist) help model perform the most on AGR, OPE and NEU. We hypothesize that simple, lexical cues are more helpful here since it is easier to draw relations from the surface form in posts. We also note that estimating NEU is difficult for the model, which also is difficult for humans to estimate in zero-acquaintance contexts, Kenny (1994), including estimating neuroticism from Facebook profiles. Overall, GPT-3’s predictions are heavily biased towards predicting individuals to be high openness and low in neuroticism.
We also tried incorporating all types of knowledge into a prompt and found that performance dropped below Basic. However, combining knowledge types involves non-trivial decisions such as the order of knowledge types and its composition. We leave this to future work.
Using ItemDesc, we establish the best possible GPT-3 performance for personality estimation. Although GPT-3’s average performance over all traits is still lower than Wt-Lex, it outperforms the MFC baseline. Prior work V Ganesan et al. (2022); Matero et al. (2022) has shown dimensions of mental health constructs and personality traits being captured through language use patterns in LMs. GPT-3’s performance in zero-shot setting provides reasonable evidence to believe that language patterns associated with these traits are encoded in its embedding space as well.
Analysis
To better understand the utility of GPT-3 for personality estimation, we analyze the effect of (1) problem framing, and (2) effect of survey items. Furthermore, we perform error analysis of GPT-3 to suggest avenues for improvement.
When personality estimation is framed as a binary classification, GPT-3 is worse than SoTA on average in a zero-shot setting. Upon looking closer, we note that it is the best model for 2 out of the 5 traits. However, these observations are made in a simplified two-class setting, whereas the big 5 personality model produces a real valued outcome. In order to assess GPT-3’s practical viability, we prompt it (ItemDesc) to provide more fine-grained predictions by presenting trait estimation as a three-class classification problem.
Table 2 shows that problem framing has a major impact on GPT-3 performance for all traits. Three class framing of the problem is harder than the binary framing which is evident from GPT-3’s drop in performance (0.229) to close to MFC (0.212). This trend indicates that GPT-3 is ineffective in performing more fine-grained prediction tasks and consequently regression, which is the natural way to estimate the Big 5 traits. Clearly, GPT-3 is yet unsuited for fine-grained personality estimation.
Consistency with Survey Items.
The standard questionnaire used to create the dataset had a total of 4 survey items per trait (2 positive and 2 negative). For ItemDesc, we use one positive and one negative item to describe each trait (see Figure 2). To investigate whether GPT-3 performance can be attributed to specific items in the prompt, we perform ItemDesc with all possible combinations of a positive and a negative survey item for all traits.
Table 3 shows that there is no meaningful difference in performance when provided different item combinations. This shows that GPT-3 is not sensitive to the items of the personality questionnaire. This is in line with data in Table 8, which shows that factor loading values Fabrigar and Wegener (2011) of these item combinations have similar powers to distinguish the corresponding traits.
Error Analysis.
Finally, we examine the linguistic variables that account for the errors in GPT-3 and the areas where it excels as compared to a traditional, lexical-based technique Wt-Lex. Figure 1 shows the distributions of SOCIAL words (Tausczik and Pennebaker, 2010) between users that were correctly predicted by only GPT-3 and the users that were either misclassfied by GPT-3 or correctly predicted by Wt-Lex for EXT task. SOCIAL words are better captured by LLMs probably owing to its ability to produce contextualized embeddings. Figure 1 depicts the distributions of AFFECT words between the users that were misclassified only by GPT-3 and the users that were either correctly classified by GPT-3 or Wt-Lex misclassifies for OPE taskWe also looked at the differences in other LIWC categories for EXT and OPE tasks measured using Cohen’s d Diener (2010) and logs odds ratio with informative dirichlet prior Monroe et al. (2008) that offers more explanations for the errors and correctness of GPT-3 in Appendix C..
Conclusion
We performed a systematic investigation of GPT-3’s zero-shot performance on personality estimation. While using a simple prompt did not yield strong performance, injecting knowledge about the traits themselves led to significant improvement. Even so, it falls short of using a strong, extensively-trained, supervised model (Wt-Lex). Further, we find that it is much harder for GPT-3 to provide more fine-grained predictions (when asked to select between 3 labels instead of 2), suggesting that LLMs may not be as capable at making dimensional estimates about personality. Our systematic investigation helps understand GPT-3’s zero-shot capabilities for a human-level NLP task, contextualizing its failure modes and showing avenues for LLM improvements.
Ethics Statement
Our work seeks to advance interdisciplinary NLP-psychology research for understanding human attributes associated with language. This research is intended to inform Computational Social Science researchers about the ability of LLMs to estimate psychological rating scales as well as for LLM researchers to understand types of psychological information that LMs capture. We intend for our work on personality trait assessments to have an impact on social, NLP, and clinical use cases to improve the well-being of people. We strongly condemn malevolent adoption of these technologies for targeted advertising, directed misinformation campaigns, and other malicious acts that could have potential harms on mental health.
If used for clinical practice, we strongly recommend that any use of LLM-based personality estimates be overseen by clinical psychology experts. During trials, models should be extensively tested for their failure mode rates (e.g. False-positive vs False-negative rates), and error disparities Shah et al. (2020).
This interdisciplinary computer science, psychology, and health study had extensive privacy & ethical human subjects research protocols. All procedures were approved by an academic institutional review board. All contributors are certified to perform human subject research, and took steps and precautions while collecting and analyzing data to keep participants protected. The Facebook posts shared by consenting users were anonymized as described in §3 to prevent the participants from being identified.
Limitations
The Big 5 personality trait model measures the fundamental dimensions of human on a continuous scale. This real valued representation preserves more information and is more descriptive of inter-individual differences. While we acknowledge that the binary classification of Big 5 traits fails the purpose of the model, it is a necessary simplification to understand the ability of LLMs to perform personality assessment. Our investigation shows potential to improve the practical utility of LLMs in personality estimation.
Despite the strong results from existing works in support of in-context learning and larger message history for better performance, we were limited by the significant multiplicative cost these experiments entailed, as the GPT-3 API is billed based on token usage. Further, since each user’s post history is typically long, it is infeasible to experiment with all in-context learning options due to GPT-3’s context window size limitation. This is worthy of exploration, to understand the sample efficiency of GPT-3 and the impact of post history on its performance.
Acknowledgement
This work wouldn’t have been possible without the support of AVG’s and YKL’s dear friends, Swanie Juhng, Aakanksha Rajiv Kapoor, Aravind Parthasarathy, Aditya Krishna, Akshay Bharadhwaj, and Somadutta Bhatta, who provided OpenAI API keys for running the experiments. We would also like to extend our gratitude to Matthew Matero, Niranjan Balasubramanian, Harsh Trivedi and Sid Mangalik for providing valuable feedback. AVG, AHN, and HAS were supported in part by NIH grant R01-AA028032 and YKL was supported by DARPA.
References
Appendix A GPT-3
We used a temperature of 0.0 for all the experiments to select the most likely token at each step, as this setting allows for reproducibility.
We restricted the model outputs to just one token. Only “Yes" or “No" are considered valid answers for our binary classification task. For the 3-class classification, “High", “Medium" and “Low" are considered valid answers.
For one data point in the Wordlist EXT experiment, the model output was a newline character instead of Yes/No. By adding another newline to the prompt, we were able to get it to generate an answer (in this case, No). For one data point in the Basic OPE experiment, the model output contained irrelevant tokens instead of High/Medium/Low. By adding another 2 newlines to the prompt, we were able to get it to generate an answer (in this case, High).
A.2 Prompt Design
For our binary classification task, we used the following prompt template:
A user’s posts are concatenated with the most recent post presented at the end to fill the messages field. Options for trait are agreeable, extraverted, open to experiences, neurotic, and conscientious.
For our 3-class problem framing, we used the following prompt template:
Options for trait are agreeableness, extraversion, openness to experiences, neuroticism, and conscientiousness. The different types of knowledge injected into the prompt for each personlity trait can be found in Figure 2.
Appendix B Glossary
We include the survey items from the questionnaires used in the study to collect data from consenting users along with their associated personality trait in Table 7, as well as the categories of language from the LIWC error analysis model in Table 4.
Appendix C Error Analysis
We examine where GPT-3 differs from Wt-Lex: (1) performing better on EXT in Table 5, and (2) predicting OPE worse in Table 6. Results from Table 5 suggest that GPT-3 encodes language categoriesSee Table 4 for details on LIWC categories Tausczik and Pennebaker (2010) highly predictive of EXT such as social processes (SOCIAL), group identification (AFFILIATION), and use of second person pronoun (YOU), all of which have been shown to have strong significant association with this trait Schwartz et al. (2013). GPT-3 can disambiguate common social lexicons occurring in different contexts Burdick et al. (2022) (e.g., "party" in the context of gathering vs political ideology), which count-based lexical models can’t do.
Table 6 indicates that GPT-3 fails for OPE on language reflective of social processes (SOCIAL) and affect (AFFECT). Previous work on lexical correlates of personality showed that these categories are discussed more for users low in openness (Yarkoni, 2010), suggesting (together with our result) that GPT-3 misses the connection between these categories of language and personality. These are areas to improve the human-level capabilities of GPT-3.