Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs

Maarten Sap, Ronan LeBras, Daniel Fried, Yejin Choi

Introduction

With the growing prevalence of AI and NLP systems in everyday social interactions, the need for AI systems with social intelligence and Theory of Mind (ToM), i.e., the ability to infer and reason about the intents, feelings, and mental states of others, becomes increasingly evident Pereira et al. (2016); Langley et al. (2022). For humans, Theory of Mind is a crucial component that enables us to interact and communicate effectively with each other Premack and Woodruff (1978); Apperly (2010). It allows us, for example, to infer that someone likely feels boastful instead of ashamed after winning a wrestling match (Fig. 1; top). In addition, ToM also enables us to reason about people’s mental realities, e.g., if someone was out of the room while a pen was moved, she will likely search for the pen where she last saw it instead of where it was moved to (Fig. 1; bottom).

While humans develop it naturally, ToM and social intelligence remain elusive goals for modern AI systems Choi (2022), including large neural language models (LLMs). With advances in scaling the sizes of models and datasets, these LLMs have proven very impressive at generating human-like language for conversational, summarization, or sentence continuation settings, often with zero to few examples to learn from Brown et al. (2020); Clark et al. (2021); Chowdhery et al. (2022). However, increasing scrutiny has shed light on the shortcomings of these LLMs, showing that they often fall prey to spurious correlational patterns instead of displaying higher-order reasoning Elkins and Chun (2020); Dale (2021); Marcus (2022).

In line with EMNLP 2022’s theme, we examine the open research question of whether and how much LLMs—which are the backbone of most modern NLP systems—exhibit social intelligence and ToM abilities. Using some of the largest English models in existence (GPT-3; Brown et al., 2020), we demonstrate that out-of-the-box LLMs struggle at two types of reasoning abilities that requisites for Theory of Mind (shown in Fig. 1). We argue that these reasoning abilities are necessary but not sufficient for Theory of Mind, and that larger models will likely provide upper bounds on what equivalent-but-smaller models are capable of.

We first assess whether LLMs can reason about social commonsense and emotional intelligence with respect to social interactions (§3), using the SocialIQa benchmark Sap et al. (2019b) illustrated in Fig. 1 (top). Results show our best performing few-shot GPT-3 setup achieving only 55% accuracy, lagging >>30% behind human performance. Furthermore, social reasoning about the protagonists of situations is easier for GPT-3 (5-15% absolute difference) compared to reasoning about other secondary participants.

Second, we measure LLMs’ ability to understand other people’s mental states and realities in short stories (§4). We use the ToMi QA benchmark (illustrated in Fig. 1; bottom; Le et al., 2019), which was inspired by the classic Sally-Ann False Belief Theory of Mind test Baron-Cohen et al. (1985). Here, our results show that GPT-3 models peak at 60% accuracy on questions about participants’ mental states, compared to 90–100% on factual questions.

Our novel insights show that reasoning about social situations and false beliefs still presents a significant challenge for large language models, despite their seemingly impressive performance on tasks that could require social intelligence (e.g., story generation, dialogues). In §5, we first examine these shortcomings; drawing on theories of the pragmatics of language, we speculate that the type of texts in LLMs’ training datasets could substantially limit learning social intelligence. Then, we outline some possible future directions towards socially aware LLMs, reflecting on the feasibility of interactional data selection, person-centric inductive biases, and interaction-based language learning. Our findings suggest that only increasing the scale of LLMs is likely not the most effective way to create socially aware AI systems, challenging a prevalent narrative in AI research Narang and Chowdhery (2022).

Theory of Mind & Large LMs

Social intelligence, Theory of Mind, and commonsense reasoning have been a longstanding but elusive goal of artificial intelligence for decades (Gunning, 2018; Choi, 2022). These reasoning abilities are becoming increasingly necessary as AI assistants are used in situations that require social intelligence and Theory of Mind in order to operate effectively Wang et al. (2007); Dhelim et al. (2021); Langley et al. (2022). For example, new technologies are emerging where AI is used to interact and adapt to users Bickmore and Picard (2005); Jaques (2019), e.g., voice assistants, and tutoring systems; or where AI helps enhance communication between multiple users, e.g., email autocomplete (Chen et al., 2019), AI-assisted counseling Kearns et al. (2020); Allen (2020); Sharma et al. (2021), or facilitated discussion Rosé et al. (2014).

As we move beyond just asking single-turn questions to social and interactive AI assistants, higher-order reasoning becomes necessary McDonald and Pearson (2019). For example, AI systems should be capable of more nuanced understanding, such as ensuring an alarm is on if someone has a job interview the next morning Dhelim et al. (2021), knowing to call for help when an elderly person falls Pollack (2005), inferring personality and intentions in dialogues Mairesse et al. (2007); Wang et al. (2019), reasoning about public commitments Asher and Lascarides (2013), predicting emotional and affective states Litman and Forbes-Riley (2004); Jaques et al. (2020), and incorporating empathy, interlocutor perspective, and social intelligence Kearns et al. (2020); Sharma et al. (2021).

What is Theory of Mind?

Theory of Mind (ToM) describes the ability that we, as humans, have to ascribe and infer the mental states of others, and to predict which likely actions they are going to take Apperly (2010).While Theory of Mind is well developed in most adults Ganaie and Mudasir (2015), reasoning and inference capabilities can be influenced by age, culture, neurodiversity, or developmental disorders Korkmaz (2011). This ability is closely related to (interpersonal) social intelligence Ganaie and Mudasir (2015), which allows us to navigate and understand social situations ranging from simple everyday interactions to complex negotiations Gardner et al. (1995).

Interestingly, the development of Theory of Mind and language seem to happen around similar ages in children Sperber and Wilson (1986); Wellman (1992); Miller (2006); Tauzin and Gergely (2018).The direction of the ToM-language association is still debated de Villiers (2007). Some researchers believe language development enables ToM-like abilities Pyers and Senghas (2009); Rubio-Fernandez (2021). On the other hand, some argue that language develops after ToM since preverbal infants already could possess some level of ToM-like abilities Onishi and Baillargeon (2005); Southgate and Vernetti (2014); Poulin-Dubois and Yott (2018). Theories of the pragmatics of language and communication can frame our understanding of this link Rubio-Fernandez (2021), positing that one needs to reason about an interlocutor’s mental state (ToM) to effectively communicate and understand language Grice (1975); Fernández (2013); Goodman and Frank (2016); Enrici et al. (2019).Most cognitive studies on this subject focus on the English language, which is not representative of the wide variation of language structures, and thus limits the cognitive conclusions one can draw about the link between language and Theory of Mind Blasi et al. (2022).

SocialIQa: Do LLMs have Social Intelligence and Social Commonsense?

A crucial component of Theory-of-Mind is the ability to reason about the intents and reactions of participants of social interactions. To measure this, we use the dev. set of the SocialIQa QA benchmark Sap et al. (2019b), which was designed to probe social and emotional intelligence in various everyday situations. This benchmark covers questions about nine social reasoning dimensions, drawn from the Atomic knowledge graph Sap et al. (2019a).

SocialIQa instances consist of a context, question, and three answer choices, written in English. Each question relates to a specific reasoning dimension from Atomic: six dimensions focus on the pre- and post-conditions of the agent or protagonist of the situation (e.g., needs, intents, reactions, next actions), and three dimensions focus on the post-conditions of other participants involved in the situation (reaction, next action, effect). In total, there are 1954 three-way QA tuples; see Tab. 1 for examples, and Tab. 3 in Appendix A for per-dimension counts.

To probe our language models, we use a kk-shot language probing setup, following Brown et al. (2020). We select the answer that has the highest likelihood under the language model conditioned on the context and question, as described in Appendix C.

To test the limits of what the models can do, we select kk examples that have the same Atomic reasoning dimension as the question at hand, varying kk from 0 to 35 in increments of 5. We use three GPT-3 model sizes: GPT-3-Ada (smallest), and GPT-3-Curie and GPT-3-DaVinci (two largest).

2 SocialIQa Results

Shown in Fig. 2, GPT-3 models perform substantially worse than humans (>30% less) on SocialIQa, We find similar results when using InstructGPT Ouyang et al. (2022) instead of GPT-3-DaVinci. and also worse than models finetuned on the SocialIQa training set (>20%; Lourie et al., 2021).Lourie et al. (2021) achieves 83% on the test set, as shown on the AI2 SocialIQa leaderboard. Although it is not surprising that GPT-3-DaVinci reaches higher accuracies than GPT-3-Ada and GPT-3-Curie, the gains are small, which suggests that increasing model size might not be enough to reach human-level accuracy. These findings are in line with recent BIG-Bench results on SocialIQa with the BIG-G (128B parameters; Srivastava et al., 2022) and PaLM (353B parameters; Chowdhery et al., 2022) LLMs, which lag behind humans with 45% and 73% accuracy, respectively (see Fig. 7 in Appendix A.2).

Focusing on GPT-3-DaVinci, while increasing the number of examples kk improves performance, the differences are marginal after kk=10 examples (only 1% increase from 10 to 35 examples). This suggest that performance either plateaus or follows a logarithmic relationship with increasing number of conditioning examples.

Finally, we examine the differences in GPT-3-DaVinci with respect to which participant is the focus. Shown in Fig. 3, we find that GPT-3-DaVinci performs consistently better on agent-centric questions, compared to other-oriented questions. Shown in the example predictions in Tab. 1, GPT-3-DaVinci often confuses which participant is being asked about. In example (e), after Aubrey babysat for Tracy, GPT-3-DaVinci fails to predict that Tracy will likely want to “let Aubrey know they are appreciated,” and instead mistakenly predicts that Tracy will want to “save up for vacation,” which is what Aubrey would likely do. GPT-3-DaVinci displays a similar participant confusion in example (f) in Tab. 1.

ToMi: Can LLMs Reason about Mental States and Realities?

Another key component of Theory of Mind is the ability to reason about mental states and realities of others, recognizing that they may be different than our own mental states. As a measure of this ability in humans, psychologists developed the Sally Ann false-belief test Wimmer and Perner (1983), in which two people (Sally and Ann) are together in a room with a ball, a basket, and a box, and while Sally is away, Ann moves the ball from the basket to the box. When asked where Sally will look for her ball, Theory of Mind allows us to infer that Sally will look in the basket (where she left the ball), instead of in the box (where the ball is, unbeknownst to Sally).

To measure the false-belief abilities of LLMs, we use the ToMi QA dataset of English Sally-Ann-like stories and questions Le et al. (2019).ToMi is a more challenging version of the rule-based datasets by Nematzadeh et al. (2018) and Grant et al. (2017), as it contains randomly inserted distractor actions that prevent trivial reverse engineering. ToMi stories were created using a stochastic rule-based algorithm that samples two participants, an object of interest, and a set of locations or containers, and weaves together a story that involves an object being moved (see Tab. 2). All questions have two possible answers: the original object location, and the final object location.

We investigate how LLMs answer the ToMi story-question pairs, distinguishing between questions about factual object locations (Fact) and questions about where participants think objects are located (i.e., their mental states; Mind). The Fact questions either ask about the object’s original (Fact-Mem) or final (Fact-Real) location. The Mind questions cover first-order (e.g., “where will Abby look for the object?”; Mind-1st) and second-order beliefs (e.g., “where does James think that Abby will look for the object?”; Mind-2nd). We further distinguish the Mind questions between true belief (Tb) and false belief (Fb), i.e., stories where a participant was present or absent when an object was moved, respectively.

Importantly, answering the Mind questions requires Theory of Mind and reasoning about realities and mental states of participants—regardless of the true- or false-belief setting—whereas Fact questions do not require such ToM. There are a total of 1861 two-way QA pairs in our ToMi probe set, with 519 Fact and 1342 Mind questions (see Tab. 4 in Appendix B for more detailed counts).

We use the kk-shot probing setup to test this ToM component in LLMs, with k∈{2,4,8,16,24}k\in\{2,4,8,16,24\}. We select kk examples of the same reasoning type (i.e., Fact-Mem, Mind-1st, etc.), ensuring a 50-50 split between true- and false-belief examples for the Mind questions. As before, we test GPT-3-Ada, GPT-3-Curie, and GPT-3-DaVinci.

2 ToMi Results

Shown in Fig. 4, our results indicate that GPT-3 models struggle substantially with the ToMi questions related to mental states (Mind), reaching 60% accuracy in the best setup. As expected, the best performance is reached with GPT-3-DaVinci compared to smaller models which do not surpass 55% accuracy; however, as before, the gains from scaling up GPT-3 are very small. Similarly, increasing the number of few-shot examples beyond k=4k=4 does not substantially improve performance, corroborating findings on SocialIQa.

Further examining GPT-3-DaVinci with respect to question types, we show that the model struggles substantially more with questions about mental states (55–60% for k>0k>0) compared to factual questions (90–100% for k>0k>0; Fig. 5; columns). Furthermore, the difference between performance on Mind-Tb and Mind-Fb questions shows an interesting pattern when conditioning on an increasing number of examples kk (Fig. 5; lines): GPT-3-DaVinci’s Mind-Tb accuracy first increases, peaks at k=4k=4, then decreases. This peak seems to be due to the model defaulting to the most recent object location (i.e., the correct Mind-Tb answer), as illustrated in example (e) in Tab. 2. Apparent in Fig. 10 in Appendix B, this recency bias is a phenomenon that has been previously documented in LLMs O’Connor and Andreas (2021). In general, GPT-3-DaVinci’s comparably poor performance for Mind-Tb and Mind-Fb questions at k>8k>8 suggests that it cannot properly answer questions about participants’ mental states and realities.

Discussion: Towards NLP with Neural Theory of Mind

Most humans develop social intelligence and Theory of Mind naturally. However, in this work, we showed that these abilities do not emerge automatically in large-pretrained language models. These shortcomings contrast with the wealth of successes of LLMs at a variety of tasks, including tasks that potentially require social intelligence. For example, GPT-3 has been shown to generate stories with emotional arcs that are virtually indistinguishable from human-written stories Clark et al. (2021). Additionally, recent work has used GPT-3 to generate social commonsense knowledge related to protagonists of situations West et al. (2022). While those findings suggest some level of social and emotional intelligence in LLMs, our explorations highlight the limits of these abilities, and raise the open question: how can we create NLP systems with true social intelligence and Theory of Mind?

To begin answering this question, we first discuss the current LLMs training paradigm (§5.1), drawing from theories of pragmatics to examine why these models are not learning social intelligence efficiently. Then, we outline some possible future directions to bias models towards Theory of Mind (§5.2), through person-centric neural architectures, data selection, and training objectives.

To understand why LLMs are still struggling with social intelligence, we examine LLMs’ training paradigm through the lens of pragmatics. As discussed in §2, pragmatics provides a connection between language development and Theory of Mind Sperber and Wilson (1986); Miller (2006); Tauzin and Gergely (2018): learning to communicate effectively with language requires reasoning about what our interlocutor knows or does not know Grice (1975); Fernández (2013); Goodman and Frank (2016); Enrici et al. (2019). Note here that, in contrast to other work Bender and Koller (2020); Bisk et al. (2020), we do not focus on whether LLMs “understand” language, instead we examine whether LLMs can answer questions about the emotions and mental states of participants of situations.

One major use of language by people is to communicate about relationships and personal experiences Clark and Schaefer (1989); Dunbar (1993). This is fundamentally different from the training data of LLMs, which consists of language found in what we call static texts: documents that are written for a general audience and are relatively self-contained and topically focused (e.g., news articles, books, Wikipedia articles; Gao et al., 2020; Dodge et al., 2021). Such static text is typically written such that readers only require the language itself as input, which they then combine with their world knowledge and commonsense to understand its meaning Graesser et al. (1994).

If AI systems are to learn social intelligence and Theory of Mind, we posit that static text has certain limitations, from a pragmatics lens, outlined below.

Following Grice’s maxim of quantity Grice (1975), static text often avoids redundancy by omitting content that is known by both the author and the reader Clark and Brennan (1991). Also known as reporting bias Gordon and Van Durme (2013); Lucy and Gauthier (2017), this phenomenon likely limits LLMs’ ability to learn social commonsense knowledge from static text.

Lack of communicative intent and alternatives.

A corollary to reporting bias, static text does not provide any direct access to communicative intent (why words were used) or to alternatives (which words were not used, and why). This reasoning about intents, alternatives, and their implications is highly predictive of the pragmatic inferences people draw about their interlocutors Goodman and Frank (2016) — for example, when someone answers Where does Taylor live? with Somewhere in the U.S., it implies that they likely do not know or do not want to share the exact location, since, if they did, they would have been more specific. This poses a likely limitation that LLMs only learn what words are used, but not which words were not used, and why.

Lack of communicative effects.

Language is primarily learned Wells and Bridges (1981); Tomasello et al. (2005) and used Clark (1996) in collaborative and interactive settings Clark and Schaefer (1989), which allow interlocutors to give immediate feedback to each other on whether their language was understood Clark and Krych (2004) or should be adjusted Krauss and Weinheimer (1966), and observe the perlocutionary effects that their language has on their partners Austin (1975). Since static text has no such feedback, LLMs learn from all texts, as if they were all equally understandable by readers.

Centering theory.

At any given time, most text focuses on describing one protagonist and their relation to their surroundings, according to Centering Theory Grosz et al. (1995). As such, main characters and their mental states are more likely to be described, whereas other participants might only be mentioned. Additionally, main characters or protagonists are more likely to be referred to with pronouns, whereas secondary characters with their names.

Thus, a model trained purely on static text might not learn to reason about social intelligence or mental states and realities of different characters of situations; they might not even inherently learn to resolve coreference for multiple characters Sakaguchi et al. (2020). In fact, challenges of coreference resolution could explain why GPT-3 models struggle on SocialIQa which contains questions with pronouns, and centering theory and main character biases in static text could explain why models find non-protagonist questions more challenging. On the other hand, ToMi does not contain any pronouns, and thus requires social intelligence beyond coreference resolution.

2 Future directions towards LLMs with Theory of Mind

While there is no one best path towards LLMs with social intelligence and Theory of Mind, it seems likely that progress will require challenging the standard paradigm of training on static text with the language modeling objective. Based on our findings and the limitations we discussed, we reflect on some possible directions forward.

Perhaps the key is in the data: the knowledge contained in static text might be too limited for models to learn social intelligence, for reasons described in §5.1 Socially grounded text (containing elaborations of communicative intents, character mental states, speaker identities, etc.) could enable more efficient learning of Theory of Mind abilities Bender and Koller (2020); Bisk et al. (2020); Hovy and Yang (2021), similar to how visual groundings can help with learning physical knowledge Zhang et al. (2022a). Examples of such datasets include “Social Stories,” which are devised to help individuals with autism improve their interpersonal skills Gray (1995), or the Story Commonsense Rashkin et al. (2018) and GLUCOSE Mostafazadeh et al. (2020) commonsense-annotated story datasets. Alternatively, perhaps interactional texts, such as dialogues and other datasets that were explicitly created to require reasoning about mental states, could help with neural Theory of Mind Bara et al. (2021).

Nevertheless, the scale of training datasets seems to be crucial for LLMs Kaplan et al. (2020); Chowdhery et al. (2022), which poses a challenge: text datasets rich in social intelligence and interactions are not easily found naturally due to reporting biases, and they are costly to create Rashkin et al. (2018); Mostafazadeh et al. (2020). Promising results on commonsense reasoning suggest a possible hybrid approach: LLMs could be jointly or sequentially trained on static text and commonsense knowledge bases or socially grounded or interactional text (Bosselut et al., 2019; Hwang et al., 2021), first trained on static text and then enhanced for commonsense knowledge via reinforcement learning Zhou et al. (2021).

Person-centric neural inductive biases?

While more socially grounded training data could help, LLMs might also learn social intelligence better if they are designed with person-centric inductive biases and training objectives. Hinting at this, prior work has shown that training entity-centric neural architectures on text with entity coreference information yields more entity-aware LLMs, both in recurrent Henaff et al. (2017); Ji et al. (2017); Yang et al. (2017); Liu et al. (2019) and Transformer-based models Févry et al. (2020); De Cao et al. (2020); Rosset et al. (2020); Zhang et al. (2022c).

However, Theory of Mind and social intelligence require much richer social grounding than coreference chains, which is challenging to obtain for supervised settings, especially at the scale that LLMs require. Thus, unsupervised approaches to adding inductive biases to models could be a promising solution. Future work could look to cognitive science and neuroscience research for possible directions Langley et al. (2022), such as exploring LLMs’ equivalents of human concept cells (i.e., sets of neurons that activate for important people or concepts; Bowers, 2017; Calvo Tapia et al., 2020).

Alternatively, examining the internal or latent representations of LLMs could point to future directions towards inductive biases for neural Theory of Mind. As an example, recent work has found evidence of latent representations of grounded semantics in models trained only on static text Li et al. (2021), which can be tied to real-world grounding with a small amount of additional supervised training Patel and Pavlick (2022). Future work might similarly analyze deep learning models for representations of Theory of Mind, toward augmenting the models with structure or objectives that surface and strengthen these representations.

Interactive and experiential grounding?

It is possible, nevertheless, that socially grounded data and person-centric inductive biases will not suffice. Some researchers have argued that language understanding could only emerge from interactions and experiences Bender and Koller (2020); Bisk et al. (2020). Likely, this applies to Theory of Mind and social intelligence as well, due to lack of communicative intents and alternatives in static text. Future work could explore approaches grounded more explicitly in interaction, intents, and alternatives, e.g., by explicitly predicting possible next steps and learning why predictions were wrong. In fact, promising research has shown that using an interactive learning or multi-agent communication paradigm can enable some Theory of Mind capabilities of models Hawkins et al. (2019); Lazaridou et al. (2020); Zhu et al. (2021); Wang et al. (2022).

However, there are limits to the types of Theory of Mind that can be learned from interactive simulations, which are often task-specific (e.g., describing objects in an image; Lazaridou et al., 2020; Steinert-Threlkeld et al., 2022). Furthermore, models that were trained in interactive simulation settings often struggle to generalize beyond the simulation environment Ludwin-Peery et al. (2021); Mu and Goodman (2021). Based on promising results by Lazaridou et al. (2020); Zhu et al. (2021), future work might create generalizable LLMs with neural Theory of Mind through hybrid approaches that combine pretraining with interactive learning: updating models trained on static text using supervision either from humans Stiennon et al. (2020); Ouyang et al. (2022); Scheurer et al. (2022) or from proxies for human behavior or social environments Ammanabrolu et al. (2022a, b) based on broad coverage LLMs Perez et al. (2022).

Probing and evaluating ToM

While neural Theory of Mind and social intelligence may remain an elusive goal for some time, developing measures of those abilities in systems can be done in tandem. We encourage further research in developing benchmarks that measure specific social abilities in LLMs (e.g., Sap et al., 2019b; Zadeh et al., 2019), especially those that minimize annotation artifacts and spurious correlations Schwartz et al. (2017); Gururangan et al. (2018); Le et al. (2019). Additionally, we encourage further investigations into probing the latent knowledge within LLMs Tenney et al. (2019); Li et al. (2021) or examining how LLMs handle entities and people Onoe et al. (2022); Schuster and Linzen (2022), which could shed light onto better data choices and inductive biases towards neural Theory of Mind and social intelligence.

Conclusion

We explore the open question of whether and how much modern large-scale language models (LLMs) can reason about social intelligence and Theory of Mind. Our results show that out-of-the-box LLMs struggle substantially with these abilities, which we argue are necessary but not sufficient aspects of Theory of Mind. Specifically, GPT-3’s social intelligence as measured by SocialIQa lags behind humans (>30%), and the model struggles to answer ToMi questions about mental states (55-60%) compared to factual questions (90–100%). In light of these shortcomings, we critically examine the large language model pretraining paradigm from a pragmatics-based perspective, and discuss possible directions towards enabling true social intelligence in NLP systems.

We make our preprocessed datasets available at http://maartensap.com/neuralToM.

Limitations

Our work focuses on investigating the Theory of Mind abilities in large pretrained language models, but we focus on accessing GPT-3 Brown et al. (2020) through an API, since we do not have access to some of the larger models out there (PaLM; Chowdhery et al., 2022) nor do we have the computational resources to run an open-source version of GPT-3 (OPT; Zhang et al., 2022b). We hypothesize that results would not be drastically different with such models, based on the low accuracy displayed on SocialIQa in the recently released BIG-Bench experiments Srivastava et al. (2022). Nevertheless, we hope developers of larger LLMs will investigate these ToM abilities to confirm or refute our findings.

We measure the ability to answer questions about people’s mental states using ToMi, which is an automatically constructed corpus of stories involving people, objects, and locations. The automatic nature of the creation process could induce biases and artifacts, such as objects being in locations that are plausible but not typical (e.g., bananas in a closet), which could influence model’s ability to answer questions properly. Based on the near-perfect accuracy on the factual questions, however, this may not be a significant issue. Future work should investigate more naturalistic settings to probe this ability in LLMs.

A potential limitation of our work is that models could latch onto surface patterns and spurious correlations in our two datasets. For example, theoretically, a model prompted with many ToMi examples may be able to reverse-engineer the data creation algorithm to find the solution to each question. However, this would be a bigger limitation if our claims were that LLMs do have social intelligence and Theory of Mind; instead, given that our results show low performance on these tasks even though they are potentially easier due to correlational patterns, this would indicate that LLMs have potentially even less reasoning abilities.

Additionally, while we operationalize our measure of social intelligence and Theory of Mind through two specific tasks, SocialIQa and ToMi, these abilities are much broader. As noted earlier, we view these benchmarks as necessary but not sufficient conditions for LLMs to have ToM; solving the benchmarks does not imply that LLMs have ToM, but LLMs with ToM should be able to solve them. We hope that future research will further investigate other aspects of Theory of Mind abilities in large pretrained LMs, drawing on social science research. For example, future work could make use of the “unexpected content” task Gopnik and Astington (1988) or the “George Washington University Social Intelligence Test” Hunt (1928) to measure the social intelligence of LLMs.

Finally, the focus on English language LLMs and benchmarks for Theory of Mind is another limitation of our work. Echoing recent cognitive science work that argues the need for non-English cognitive science investigations Blasi et al. (2022). Specifically, false-belief abilities are greatly influenced by language structure and grammar Boeg Thomsen et al. (2021); Zhang and Zhou (2022).

AI systems are part of a broader sociotechnical system that also involves individual motivations and societal norms Johnson and Verdicchio (2017). As such, per a contextualist view of AI (instead of utopian or dystopian; Barbour, 1992), we envision AI systems with social intelligence and Theory of Mind being used in ways that enhance human’s lives, autonomy, and agency Chan (2022). In parallel, we strongly support the development and research of policy and regulation, to prevent misuses of AI with social intelligence (Wischmeyer and Rademacher, 2020; Crawford, 2021; Reich et al., 2021).

Acknowledgements

We would like to thank Jack Hessel, Rowan Zellers, Jena D. Hwang, Prithviraj Ammanabrolu for their feedback on preliminary versions of this work, and Anna Jafarpour and Noah Goodman for fruitful cognitive science discussions about the research. We also thank the anonymous reviewers for their thoughtful comments. This research was supported by the Allen Institute for AI and the DARPA MCS program through NIWC Pacific (N66001-19-2-4031).

References

Appendix A SocialIQa Details

We downloaded the SocialIQa training and dev. datasets from the publicly available SocialIQa website.http://maartensap.com/social-iqa/data/socialIQa_v1.4_withDims.tgz This version of the SocialIQa dataset contains the original Atomic dimensions that workers were prompted with to create a question, as well as the correspondence between questions and which character they focus on (agent or other). To ensure consistency, for each context, question, and answer, we normalize the casing to start with a capital letter if the text does not already.

A.2 Further SocialIQa results

In addition to results discussed in §3.2, we report further SocialIQa results here.

We break down the best performing GPT-3-DaVinci (35-shot) setup by reasoning dimension. Shown in Fig. 6, we find that GPT-3-DaVinci struggles most with questions related to what people needed to do before a situation could take place (Need). Conversely, questions related to a situation’s agent’s intent (Intent) and the effect of the situation on the agent (Effect) are seemingly easier for GPT-3-DaVinci. Future work should explore LLMs’s reasoning abilities along each of these dimensions in further detail.

BIG-Bench and PaLM results on SocialIQa.

To further corroborate that LLMs struggle with SocialIQa, we show the performance of the non-publicly available BIG-G Srivastava et al. (2022) and PaLM Chowdhery et al. (2022) LLMs, along with the GPT-3 models, in Fig. 7. Both models are proprietary LLMs developed and tested on the 200+ datasets in BIG-Bench by Google / DeepMind.

While they are not discussed in the main BIG-Bench paper, the SocialIQa results for few-shot settings up to kk=3 for BIG-G and kk=5 for PaLM can be found on the BIG-Bench github website (accessed on 2022-11-10). Plotted in Fig. 7, both the BIG-G and PaLM LLMs lag behind humans with 45% and 73% peak accuracy, respectively.

Appendix B ToMi Details

We generated ToMi stories using the github repository provided by Le et al. (2019). The code generated 5994 training and 5994 dev. stories. From those, we removed the story-question pairs which wrongly answered ToM-requiring questions from an omniscient perspective (i.e., answered Mind-Fb questions from an omniscient perspective instead of the perspective of the character) which we noticed upon manual data inspection.We do not know why these datapoints were generated. After this filtering, 5190 training and 5170 dev. stories remained.

For the final ToMi dev. set, we used stratified sampling to obtain similar numbers of story-question pairs for all types (Fact-Real, Fact-Mem, Mind-1st-Fb, Mind-1st-Tb, Mind-2nd-Fb and Mind-2nd-Tb). The exact counts are shown in Tab. 4. We release our final preprocessed ToMi dev. dataset at http://maartensap.com/neuralToM/ToMi-finalNeuralTOM.csv

B.2 Further ToMi results

Shown in Fig. 8-10, we provide additional results to supplement those in §4.2.

In Fig. 8, we show the different accuracies that GPT-3 models of various sizes, prompted with various number of examples, for ToMi Mind and Fact questions. This plot shows the same accuracies as Fig. 4, with the addition of the Fact accuracies. These results show that in the few-shot prompting setup, GPT-3-Curie and GPT-3-DaVinci can achieve near perfect performance on factual questions about object locations (Fact), but struggle substantially more on questions related to mental states (Mind). Surprisingly, GPT-3-Ada struggles with both factual and mental state questions, possibly due to its smaller size.

Performance by question order.

In Fig. 9, we break the GPT-3-DaVinci performance down by ToM order (i.e., Mind-1st, Mind-2nd). Results show that with a number of examples between 2 and 16, GPT-3-DaVinci performs better on Mind-1st questions (e.g., “Where will Sally look for the ball?”) and struggles more with Mind-2nd questions (e.g., “Where does Ann think that Sally will look for the ball?”). This difference is somewhat diminished but still present for kk=24 few-shot examples. These results somewhat mirror how humans struggle with increasingly higher-order ToM questions Valle et al. (2015).

Recency bias in predictions.

We further examine the results from §4.2, looking at GPT-3-DaVinci’s rate of predicting the location where the object was moved to (i.e., Fact-Real). Shown in Fig. 10, GPT-3-DaVinci accurately learns to almost always predict the last object location for Fact-Fact-Real questions, and almost never for Fact-Fact-Mem locations.

Interestingly, the rates of selecting the last object location for Mind questions follows a concave pattern. This helps shed light onto the concave accuracy pattern seen in Fig. 5 for Mind-Tb (and convex pattern for Mind-Fb). Likely, in the few-shot setting with 2<k<82<k<8, GPT-3-DaVinci defaults to the most recently mentioned object location due to recency bias, which has been previously documented in LLMs O’Connor and Andreas (2021).

Appendix C GPT-3 Access and Probing Details

To probe our language models, we use a kk-shot language probing setup, following Brown et al. (2020). Specifically, we concatenate the context (cc) and question (qq) together with proper punctuation, and assign the model prediction to the answer (aia_{i}, i∈1,2,3i\in{1,2,3}) with the highest conditional likelihood under the language model: arg max⁡ipLM(ai∣c,q,Ck)\operatorname*{arg\,max}_{i}p_{\text{LM}}(a_{i}\mid c,q,\mathcal{C}_{k}) where Ck\mathcal{C}_{k} denotes the kk training examples, for which we provide the context, question, and correct answer concatenated. Note that we explored various probing setups and formats, such as QA-oriented formats and normalizing by marginal likelihood of each answer pLM(a)p_{\text{LM}}(a) (as also explored in Brown et al., 2020), but found very little difference in performance.

Appendix D What About ChatGPT or GPT-4? Effect of Instruction-tuning & RLFH

Our analyses in this paper have focused on large language models that are simply trained on the language modeling objective (e.g., GPT-2, GPT-3). However, in recent years, many improvements in LLMs have come from instruction-finetuning (IFT) and reinforcement learning from human feedback (RLHF), which are the key to the success of ChatGPT (or GPT-3.5; Ouyang et al., 2022) and GPT-4 OpenAI (2023). Despite the opacity of how these models were trained,OpenAI has stated that they will not be releasing any useful details. In their system report for GPT-4, they state: “Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar” OpenAI (2023). the question of whether they have better neural Theory of Mind than LM-objective-only LLMs is still of interest.

We quantify the performance of the newer set of OpenAI models (listed in Table 5) on a randomly selected 400 examples from the SocialIQa and ToMi datasets. Since these models have been increasingly capable, we focus on examining only their performance in zero-shot settings, as a stress-test of their ToM abilities. The newer APIs (GPT-3.5-Turbo, GPT-4) no longer provide the ability to score sequences, and thus, prevent the language modeling probing setup used in §3 and §4. As such, we also compare two types of probing: language modeling as described in §C (LM probing), and multiple-choice probing, where we provide the two or three answer candidates prepended with letter choices (A, B, C) and prompt the model to generate the letter for the correct answer (MC probing).For example, for SocialIQa, we prepend the example with the instructions “You are a multiple-choice answering system that responds with either A, B, or C.”

Shown in Fig. 11, our results show that instruction-tuning and RLHF do indeed improve performance on zero-shot social and emotional intelligence question answering in SocialIQa. With LM probing, supervised instruction-finetuning (GPT-3.5-IFT: 53%) improves performance over vanilla language modelling (GPT-3-DaVinci: 45%) slightly less than reinforcement learning (GPT-3.5-RLHF: 55%) . Notably, multiple-choice probing only improves over random chance (33%) with RLHF models (with 60, 67, and 79% for GPT-3.5-RLHF, GPT-3.5-Turbo, and GPT-4, respectively).

Focusing on the newest models, GPT-3.5-Turbo reaches 67.2% performance, still >>20% below human performance reported in Sap et al. (2019b). However, surprisingly, GPT-4 performance increases to 79.3%, within 10% of human performance on SocialIQa.

Interestingly, however, all models and probing setups perform worse on questions pertaining to non-main characters (Fig. 11(b)), with GPT-4 showing a 7% decrease in accuracy. This suggests that reporting biases due to centering theory, as discussed in §5, may still play a role even for these extreme scale models.

Caveat: it is increasingly likely that models were trained on benchmark data (either due to data collection via the online GUI or API, or by scraping web text without filtering). Indeed, the GPT-4 system report notes that some BIG-Bench test data Srivastava et al. (2022), which contains the SocialIQa dev. dataset, was present in the training data of GPT-4 (OpenAI, 2023, footnote 5). Without more transparency on the training data and possible data contamination, conclusions about models achieving near-human performance cannot be drawn.

D.2 ToMi: Reasoning about Mental States and Realities

Results on ToMi are plotted in Fig. 12, showing that instruction-finetuning and RLHF only somewhat improve the models’ ability to reason about mental states and realities of others. Compared to GPT-3-DaVinci’s performance which is essentially random, both GPT-3.5-IFT and GPT-3.5-RLHF’s performance increases by 17% and 19%, respectively. However, in contrast to SocialIQa, performance of newer models (GPT-3.5-Turbo, GPT-4) does not substantially improve over GPT-3.5 models.

When breaking down the accuracy by question type (Fig. 12(b)), we find that the performance of these models improves only on factual questions (fact), but stays low for questions about mental states of participants (mind). Indeed, GPT-3.5-Turbo reaches the “best” performance on the Mind subset of ToMi with 60% accuracy, surpassing GPT-4’s 59% accuracy.

D.3 Discussion

Based on our new results, it is not clear that the newer generation of models have achieved neural Theory of Mind, corroborating findings by Ullman (2023), Lenci (2023), and Marcus and Davis (2023) and debunking claims of “emergence of ToM in LLMs” by Kosinski (2023) and Bubeck et al. (2023). While models may achieve higher accuracy on social intelligence tasks, their ability to model mental states and realities of others is still very far from humans (only 10% over random chance).

Examining why these instruction-tuned or RLHF models perform somewhat better remains an open question, hindered by the lack of transparency in pretraining and instruction data. Possibly, instruction-tuned models are better able to learn social intelligence due to the more interactional nature of instruction following or dialogue responding compared to static text. However, improvements could solely be due to development and test data leakage as acknowledged by OpenAI OpenAI (2023), calling for the development of better evaluation and probing methods for these ToM abilities. Additionally, approaches such as person-centric inductive biases as well as interactive, experiential, or multimodal grounding could improve their ability to model mental states, as discussed in §5.2.