Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz
Introduction
Theory of Mind (ToM) is the ability to understand that other people have thoughts, beliefs, and emotions that differ from one’s own Wimmer and Perner (1983). As ToM is inherently linked to human cognition, imbuing machines with capabilities that mimic or resemble ToM has the potential to lead to the “ELIZA effect” Weizenbaum (1976), wherein human-like intelligence or even sentience and consciousness is incorrectly ascribed to the machine (e.g., Kosinski, 2023; Bubeck et al., 2023).
In light of these possibly illusory ToM abilities, there is a pressing need to develop robust metrics for assessing Neural-ToM (N-ToM) in machines. This is particularly crucial given the escalating stakes of the debate on the extent to which machines possess ToM-like abilities and the potential ramifications of overblown claims in AI.https://futureoflife.org/open-letter/pause-giant-ai-experiments/ https://amcs-community.org/open-letters/
Two recent papers addressed whether Large Language Models (LLMs; Brown et al., 2020; Bommasani et al., 2021; Zhao et al., 2023) have a ToM, and came to opposite conclusions: Sap et al. (2022) shows they lack this ability and Kosinski (2023) claims this ability has emerged in the newer models spontaneously. The latter was criticized for its flawed methodology Marcus and Davis (2023). Ullman (2023) further showed that simple changes to the ToM questions break LLMs. But to paraphrase the saying, hype gets halfway around the world before rigorous experiments put on their boots; other researchers continue to spread the word about N-ToM, claiming that GPT-4 “has a very advanced level of theory of mind” based on a few anecdotal examples Bubeck et al. (2023).
This paper aims to address the discrepancy and limited scope of previous work (that each tested 2 tasks) by performing an extensive evaluation on 6 tasks targeting various aspects of ToM. We also experiment with different probing methods (i.e., generative QA format vs. probability of answer choices). We find that contemporary LLMs demonstrate certain N-ToM abilities, but these abilities are not robust (§4).
We investigate through a series of experiments the factors influencing performance on N-ToM tasks. We show that LLMs perform worse on datasets that were designed to prevent annotation artifacts. We also enhanced the dataset originally proposed by Kosinski (2023) to incorporate adversarial examples inspired by Ullman (2023). We find that the performance of LLMs decreases for adversarial examples, suggesting that LLMs don’t have robust ToM abilities but rather rely on shallow heuristics (§5).
We summarize these findings and additional insights in §6. In particular, we warn against drawing conclusions from anecdotal examples, testing on a few benchmarks, and using psychological tests designed for humans to test models.The code and data is available at: https://github.com/salavi/Clever_Hans_or_N-ToM
Background: ToM and Clinical Tests
ToM has a long history starting in philosophy Lewis (1966) and later in psychology and cognitive science Premack and Woodruff (1978). ToM involves understanding mental states, beliefs, desires, intentions, and emotions of the self and of others. Clinical psychology tests were developed to test ToM abilities in humans, such as the false belief and faux pas tests detailed here.For a detailed review see Osterhaus and Bosacki (2022).
In a false belief test Wimmer and Perner (1983) the examinee is told a story in which a character in the story is exposed to partial information and therefore mistakenly believes in something that is not true (“false belief”) in contrast to the listener who is exposed to the full story.
A widely used clinical psychology task to assess false belief understanding is the Sally–Anne Test Baron-Cohen et al. (1985) or unexpected transfer. In this test, Sally has a basket, and Anne has a box. Sally puts a marble in her basket and leaves the room. Anne takes the marble out of the basket and puts it in her box. The examinee is asked about first order belief, i.e. where will Sally look for her marble?; about the reality, i.e. where is the marble?; and about their memory, i.e. where was the marble in the beginning?.
The answers are that Sally will look in the basket, where she left the marble. Sally’s belief is false because she is unaware of the marble’s relocation to the box. However, a listener exposed to the entire story knows that the marble is no longer in Sally’s basket and that Sally will look in the wrong place.
In more complex versions, Second Order Belief question would be, where does Anne think Sally will look for her marble?
In a different version of a false belief task, known as the Smarties Test Perner et al. (1987), the protagonist is dealing with unexpected content, i.e., unaware of the actual contents of a container because of false labeling.
2 Faux Pas Test
Faux Pas occurs when “a speaker says something without considering if it is something that the listener might not want to hear or know, and which typically has negative consequences that the speaker never intended” Baron-Cohen et al. (1999). An example of a faux pas situation is when a guest tells their hosts that they “like cakes except for apple pie”, without realizing that the hosts have made an apple pie for them. The complexity of the situation depends not only on the content of the statement (“except for apple pie”) but also on the context in which it was made (e.g., the host had made an apple pie and the guest was unaware). Faux pas is the “uh-oh!” emotion most people would feel when they reveal the reality of the context. In this context, the statement wouldn’t be problematic if the hosts made a cheesecake instead.
In the original test, the subject is told 10 stories that contain faux pas. At the end of each story, the subject is asked 4 questions: detection - In the story did someone say something that they should not have said?; identification - What did they say that they should not have said?; And two questions that differ by story: comprehensive - e.g., Where does the event take place?, and false belief - did they know or remember that?
3 From Human Tests to Machine Tests
Studies have explored the use of NLP techniques to model basic ToM skills. For example, in detecting mental states and emotions Tausczik and Pennebaker (2010); Guntuku et al. (2017); Gordon and Hobbs (2017); Rashkin et al. (2018a, b); Shapira et al. (2021) or by generating a humorous response when the interlocutor is in a playful mood Shani et al. (2022); Shapira et al. (2023a). Recent work is focused around creating datasets testing whether and to what extent models have ToM (see §3). It is important to note that the consequences of the success of these tests do not straightforwardly transfer from humans to models (see §6).
Data
We used all datasets listed in Table 1 in our experiments. Below is a brief description of each dataset. The creation of ToMi’ which is based on ToMi is described immediately after the description of ToMi (§3.1). The creation of Adv-CSFB (§3.2) contains a description of the datasets it is based on.
A set of 100 problems, each describes a short sequence of events involving the characters of the Heider and Simmel (1944) film: two triangles and a circle moving around a box with a hinged opening. The questions require understanding the action sequence and social reasoning, and two answer choices are given.
A large-scale (38k) dataset for commonsense reasoning about social situations. Questions in SocialIQa require reasoning about people’s motivations and mental states, causes and effects. The questions in SocialIQa were crowdsourced along with correct and incorrect answers. Additional distractors were added by using the correct answer for a different question on the same context, using a framework that mitigates stylistic artifacts.
Inspired by the Sally-Anne test, ToMi is an improved iteration of prior datasets Weston et al. (2015); Grant et al. (2017); Nematzadeh et al. (2018), comprising over 1,000 distinct stories and questions regarding memory, reality, and first and second-order false belief. This synthetic dataset was automatically generated for a range of essential objects and actions and was further processed for artifact prevention.See Appendix 8.1 for an example.
ToMi stories are in question-answering format. We randomly sampled 30 stories (each story has 6 questions, 180 questions in total) from the ToMi dataset and modified them to match a sentence completion format with the same meaning.This was done manually by one of the authors. For example the question: “Where does Oliver think that Emma searches for the grapes?”. Was adjusted to the following sentence completion task: “Oliver thinks that Emma searches for the grapes in the”.
This dataset is part of BIG bench Srivastava et al. (2022). It combines ToM with natural language inference. The tests pertain to epistemic mental states Wimmer and Perner (1983) and epistemic logic Hintikka (1962). This is done by using specific verbs related to knowledge and belief: factive (i.e., know, understand, recognize, see, remember, learn), and non-factive (i.e., believe, think, suspect, assume). The dataset contains 3 types of tests: (1) intra-personal tests: reasoning about the mental states of a single agent; (2) inter-personal tests: reasoning about the mental states of multiple agents; and (3) inference reasoning: recognizing that other agents are making inferences (i.e., if X entails Y, and Bob believes that X, then, it is reasonable to conclude that Bob believes Y).
Based on the clinical faux pas test Baron-Cohen et al. (1999), the set contains 44 stories (22 faux pas and 22 equivalent control) with 4 corresponding questions. The stories require both social reasoning skills and detecting false belief. The stories were created by experts and a small part of the stories was created by ChatGPT with rephrasing and fixes by experts.
2 Creation of Adv-CSFB
Inspired by the disagreeing conclusions reached by prior work, we introduce the ADVersarial CommonSense with False-Belief (Adv-CSFB) dataset. Adv-CSFB contains 110 examples of the unexpected contents task and 73 examples of the unexpected transfer task (§2.1). Each manually-created example in the dataset consists of a short paragraph describing two objects and , and is followed by questions pertaining to reality, i.e. whether a certain container contains or , and the protagonist’s belief regarding the content.
The examples in Adv-CSFB are categorized to false belief, i.e. the original examples from ToM-k Kosinski (2023), true belief, and adversarial examples inspired by Ullman (2023).
In the false-belief examples from Kosinski (2023), the protagonist’s belief about the content of the container is different from its actual contents. The examples are variants of the corresponding original tests from psychology, e.g. the unexpected contents examples are variants of the Sally-Anne test. Notably, Kosinski only created false-belief scenarios.
For a more fair evaluation setup, we enhance the unexpected contents task with true belief examples, i.e. in which the protagonist’s belief about the content of the container is the same as its actual contents. We do so by modifying each of the false belief examples such that the label now indicates the true content of the container, . We mention the alternative content in a way that doesn’t change the answer, e.g. Mark walks into the room looking for but finds a bag with labelled as “”. One author of this paper created a variation for each applicable example, which was then verified by another author.
Ullman (2023) showed that LLMs that achieve near-perfect performance on the false belief examples fail to solve a number of adversarial examples where new information is introduced. In particular, LLMs still predict false belief even when new information suggests that the protagonist should know the truth. For example, the LLM predicts that a protagonist looking at a bag full of popcorn that is labelled as “chocolate” believes the bag is full of chocolate, even if the bag is transparent or if the protagonist cannot read. Ullman’s counter examples are sufficient in showing that LLMs did not robustly acquire ToM abilities. To further quantify the LLMs’ abilities, we created up to 4 additional examples for each of the false belief examples, following each of the alterations suggested by Ullman (2023): transparent access, uninformative label, trustworthy testimony, and late labels for the unexpected contents task, and transparent access, inon, trustworthy testimony, and other person for the unexpected transfer task (see Appendix 8.2 for an example for each variation). Again, the examples were created by one author and verified by another.
Experiments & Results
To investigate the ToM abilities of LLMs, we designed experiments that explore various aspects. The first experiment presents a meta-evaluation of 15 LLMs evaluated on multiple ToM-related datasets in a zero-shot manner (§4.1). We then investigate to what extent LLMs are sensitive to the probing method (§4.2).
We examine the performance of 15 different LLMs of different sizes: FlanT5: flan-t5-{small, base, large, xl, xxl} Chung et al. (2022), FlanUl2 Tay et al. (2022), GPT-3 (text-davinci-002, text-davinci-003), GPT-3.5 / ChatGPT (gpt-3.5-turbo-0301), GPT-4 (gpt-4-0314) Brown et al. (2020); Ouyang et al. (2022), and Jurassic2: j2-{jumbo-instruct, grande-instruct, jumbo, grande, large}.https://www.ai21.com/blog/introducing-j2 We provide technical details regarding prompting and decoding parameters in Appendix 8.3.
1 How well do LLMs perform on ToM tasks? Meta-Evaluation
We conducted an evaluation of the performance of 15 LLMs in a zero-shot manner Liu et al. (2021) on all ToM-related datasets considered (§3), and compare to a most-frequent-class (MFC) baseline that always predicts the most frequent answer in each dataset. The summary of the results is presented in Figure 1, and the complete results in Appendix 8.4.
Our findings demonstrate that while some LLMs achieve near perfect accuracies on some datasets (e.g., TriangleCOPA with 96% accuracy by flan-t5-xxl), others datasets remain challenging for LLMs with considerably lower performance. For instance, the best performing LLM on the FauxPasEAI datasets is inferior to a simple most-frequent-class baseline, indicating the difficulty level of these datasets.
Notably, the best LLMs performance seems correlated to the dataset’s age (i.e., the older the dataset, the better the performance). This trend could be attributed to the fact that the increasing sophistication of LLMs is driving the creation of more challenging datasets, prompting researchers to set a higher bar. Another possibility is that LLMs have had more opportunities to train on the older datasets, resulting in better performance (see §8.5).
Based on this meta-evaluation, our results suggest are that while some models exhibit strong ToM abilities on some datasets, no model robustly exhibits ToM on all datasets. These findings are consistent with Sap et al. (2022) and Ullman (2023).
2 How sensitive are LLMs to the probing technique?
We examine the effect of the different probing methods detailed below on LLM performance. Certain techniques have shown to be superior to others (e.g., Wei et al., 2023). However, we argue that to claim that a model has N-ToM abilities, it is essential that it performs well across probing techniques. On one hand, the most efficient method can potentially reveal latent capabilities, while on the other hand, there is a reasonable expectation for LLMs to succeed in the tasks regardless of the probing approach used to extract information.
predicts the option with the highest probability Brown et al. (2020); Sap et al. (2022).
prompts the LLM with the context, question, and answer choices, and asks it to generate the answer in the form of “a, b, c”. This method is applicable for LLMs such as GPT-3.5 and GPT-4 that don’t produce probabilities Hu et al. (2022).
asks the model to first “reason” about the question step-by-step and then give a final answer, which generally contributes to better performance (Wei et al., 2023).We use the zero-shot setup without providing any reasoning examples.
Table 3 shows that the probing techniques influence the LLM performance on both datasets. CoT generally demonstrates enhanced performance, as supported by prior research Camburu et al. (2018); Shwartz et al. (2020); Wei et al. (2023). Nonetheless, there are cases where this trend does not hold, since the reasoning may occasionally result in erroneous conclusions Jung et al. (2022).
Clever Hans vs. Generalized Reasoning
We conducted a series of experiments aimed to enhance our understanding of the factors influencing performance in the context of N-ToM tasks. The research question that guided us was: Do the models that solve the tasks possess a general ability or do they rely on memorization and shallow heuristics (“Clever Hans”; Kavumba et al., 2019)? We detail the experiments and findings below.
ToMi and ToM-k are datasets that examine the unexpected transfer false belief problem. While ToM-k contains only simple positive examples (variants of the original Sally-Annie test), ToMi also contains simple alternations such as omission or duplication of information that create negative examples (see example in Appendix 8.1) and second-order questions.
To ensure a fair comparison between the question answering format of ToMi and the sentence completion format of ToM-k (see the effect of probing methods on performance in §4.2), we adjusted ToMi to match the sentence completion format (details about the adjustments can be found at §3.1). Additionally, we analyzed the results separately for second-order questions in order to facilitate a more accurate comparison with the ToM-k dataset.
Table 4, shows significantly lower scores in ToMi’. The notable discrepancy between the performance of the two datasets suggests that the model’s abilities are not based on generalization. Instead of true understanding of the problem at hand, such as accurately determining one’s exact thoughts, the model might be recognizing patterns from the Sally-Anne story in other ToM-k examples and generating responses based on those patterns. Conversely, the performance on ToMi’ is worse because it is more robust to spurious correlations.
2 Is N-ToM Robust to Adversarial Changes?
To test the robustness of the LLMs’ N-ToM, we test the performance of GPT models on each of the categories in Adv-CSFB (§3.2), using MC-probing. To ensure correct formatting and prevent unintended outputs (e.g., explanation of why the answer is correct), we prepend to the prompt one out-of-domain example from ToMi, which has a similar format. We report the average accuracy of questions 2 and 3, both focusing on an agent’s belief rather than objective truth. Finally, to ensure maximum reproducibility of the results, we set the temperature to 0. Our main finding is that LLMs don’t exhibit robust performance across different categories. In particular, later LLMs excel in some categories while completely failing on others. We provide details below.
Figure 2 illustrates the performance of a range of GPT models on different categories within the unexpected transfer segment of Adv-CSFB. It is evident that both false belief (i.e. the original examples from ToM-k) and trusted testimony (i.e., someone tells the protagonist that the object has been moved) have improved in newer models. GPT-4 achieves 97.5% and 83.3% on the two categories respectively. Nevertheless, there has been a gradual decline in the performance of subsequent models on other categories, such as other person (from 93.8% by davinci-002 to 68.8% by GPT-4), inon (from 71.4% by davinci-002 to 0% by GPT-4), and transparent access (from 66.7% by davinci-002 to 0% by GPT-4).
Figure 3 showcases the performance of the GPT family on various categories within the unexpected contents segment. It becomes apparent that, akin to the unexpected transfer segment, newer models such as GPT-3.5-Turbo and GPT-4 demonstrate improved performance in handling samples that involve false belief and transparent access (i.e., the container is transparent). Furthermore, nearly all models since text-davinci-002 exhibit strong performance on true belief samples. However, both GPT-3.5-Turbo and GPT-4 experience a substantial decline in performance compared to their earlier counterparts when it comes to transparent access, late label (e.g., the protagonist is the one who wrote the label), and uninformative label (i.e., the protagonist can’t read the label).
We regenerated the responses multiple times, consistently obtaining similar results, so we can conclude that the models exhibit confidence in their predictions, even if they are incorrect. It is important to note, however, that the results obtained from LM-probing may slightly differ from MC-probing. In MC-probing, even with our 1-shot setup, the model may produce responses that are not applicable, such as “none of the above” or “both”. This is particularly noticeable in verbose models like GPT-3.5-Turbo and GPT-4. These models tend to be careful to avoid providing incorrect answers and, as a result, generate longer phrases. With that said, as we argue in §4.2, a LLM exhibiting robust N-ToM ability should be able to answer questions correctly regardless of the probing method.
3 Are Spurious Correlations a Trend?
In the previous experiment §5.2, we saw that the datasets contain both difficult and easy questions. Here we show this recurring phenomenon across two ToM datasets, inspired by the analyses in Sap et al. (2022).
Figure 4 describes ToMi accuracies on different question types; ToMi contains questions about facts vs. beliefs (mind), and specifically about true or false beliefs. While GPT-3.5 (the best-performing model) achieves 81% accuracy, on the subset questions “false belief”, it achieves only 46%, close to random performance.
Figure 5 shows the SocialIQa accuracies for questions focusing on the main character vs. others. While GPT-4 (the best-performing model) achieves a total of 79% accuracy score, on the subset questions of “others”, it achieves only 74.5%.
Summary of Findings and Insights
We investigated whether modern LLMs robustly display N-ToM abilities. By quantifying their performance on 6 N-ToM benchmarks, we found that while some datasets have been nearly “solved” (e.g., TriangleCOPA with 96% accuracy by flan-t5-xxl), others remain challenging for LLMs with considerably lower performance (e.g., FauxPas-EAI with 27% accuracy by GPT-4, which is even below the majority baseline). We also created Adv-CSFB, a new ToM benchmark designed to uncover whether LLMs solve ToM questions for the right reasons, or merely rely on surface cues and shallow heuristics.
Our results show that while some datasets have been successfully solved, others remain challenging for LLMs. Thus, models do not have robust N-ToM abilities. These findings are inconsistent with Kosinski (2023), who claimed that ToM has emerged in LLMs as a byproduct of their development, a claim further echoed by Bubeck et al. (2023). We argue that these conclusions were over-generalized based on a specific aspect of ToM and a small number of examples (40 for Kosinski (2023) and 10 for Bubeck et al. (2023)). Following Ullman (2023), we empirically showed that even the best models fail on small variations of the original tasks, proving that even GPT-4 does not display robust N-ToM abilities.
The performance gaps between different question types suggests that LLMs rely on shortcuts, heuristics, and spurious correlations, which often lead them astray. In Adv-CSFB (§5.2), the bad performance on some of the adversarial categories might be partly attributed to reporting bias Gordon and Van Durme (2013); Shwartz and Choi (2020). People don’t share obvious facts Grice (1975), so it is likely that LLMs are biased towards generating surprising rather than unsurprising continuations. In most of these categories, the protagonist belief is the same as the truth, making a boring story.
Furthermore, the newer models such as GPT-3.5 and GPT-4 are trained in addition to the LM objective to follow natural language instructions and generate helpful answers. This might make them cooperative and lead to LLMs assuming that all details are important, rather than that the input is adversarial. For example, they might pay too much attention to the mention of the false label in the unexpected contents task, failing to see that the label doesn’t matter if the person can’t read it or if the container is transparent. The fact that LLMs perform reasonably well on true belief examples (Figure 3) might be attributed to recency bias O’Connor and Andreas (2021), since the correct content is typically the last one to be mentioned.
Finally, we reassess the finding of Sap et al. (2022) that LLMs perform better on predicting the mental states of the main character vs. others (SIQA, §5.1); Sap et al. (2022) suggested that this might be due to centering theory Grosz et al. (1995), according to which texts tends to focus on describing a single protagonist.
The impressive anecdotal examples produced by LLMs in generative settings (e.g., observed with ChatGPT and GPT4 web-demo; Bubeck et al., 2023), tends to captivate non-expert individuals. However, it is important to recognize that these models are specifically designed to generate text that appears high-quality to human observers (Ouyang et al., 2022). This inherent bias in their design can lead to the “ELIZA effect” Weizenbaum (1976); Shapira et al. (2023b), i.e. the human assumption that computer behaviors are analogous to human behaviors. Thus, the illusion that a LLM has acquired human-like N-ToM often says more about the humans reading the text than about the model itself Whang (2023).
Moreover, later models are by design trained to practice “epistemic humility” (i.e., hedge and provide multiple possible answers; Ouyang et al., 2022, p .17). This often leads them to provide rationales for each given answer without committing to actually answering the question. But humans might fall prey to confirmation bias and simply see the right answer and its rational and conclude that the model has gotten it correctly. We thus argue that in order to conclude whether a certain model possesses a certain ability, it is crucial to quantify the performance across multiple large-scale datasets, preferably using an automatic evaluation method.
In clinical psychology, tests designed for humans are carefully constructed and vetted to ensure that they have external and internal validity, i.e., that they are measuring what they aim to measure Frank et al. (2023). While there is evidence that a person’s success in one ToM task can indicate their ToM abilities (e.g., Milligan et al., 2007), this does not necessarily transfer to models. Therefore, it is important to be cautious when drawing conclusions about ToM in models based on their performance on a few tasks Marcus and Davis (2023). In general, when a system succeeds on an instrument designed for humans, we can’t draw the same conclusions as we would for humans (e.g., that they have ToM). Instead, we need to consider other explanations (e.g., that they are relying on heuristics). The same holds in the other direction, when analyzing how models work in order to learn about the human brain.
Relatedly, our results also point to a need for caution when discussing the abilities of machines in relation to concepts referring to human cognition, such as Theory of Mind. While it is common in computer science to use human-related concepts and metaphors for AI systems, we caution readers to interpret “neural ToM” carefully and without aiming to make claims about “AI cognition,” especially since given our propensity for anthropomorphizing non-human animals and computers Epley et al. (2007); Kim and Sundar (2012); our measuring the performance on these benchmarks is not meant as an endorsement of the pursuit of a human-like social intelligence for AI systems.We leave the question of whether LLMs, AI, or even any non-biological entity could develop human-like cognition and Theory of Mind up to philosophers. Instead, in light of the hype around AI and it’s “intelligence,” we sought out to provide a more sober look at the empirical performance of LLMs on tasks related to social intelligence and ToM.
Methodologically, if a model fails at least one ToM task, it does not have ToM in general. Success on one example or task is not a sound proof that a model has ToM. Future work will need to continue to develop benchmarks testing various ToM aspects, and these benchmarks will need to be designed to assess LLMs directly rather than using clinical tests designed for humans.
Additionally, reporting the aggregated performance of LLMs on benchmarks obscures the performance differences across questions of different types and complexities. To overcome this, one approach is to pair a difficult question with an easy question, requiring model to answer both correctly. This methodology resembles the “joint score” employed in FauxPas-EAI, Adv-CSFB, and ToMi. In situations where pairing is challenging, a recommendation for future works is that dataset difficulty could be evaluated by calculating the final score across different splits of the dataset. The difficulty level of the dataset can then be determined based on the lowest score obtained among these splits.
Prior work claimed that ToM abilities emerged as a byproduct of the LLM training Kosinski (2023). We argue that claims about emergence are (i) unfounded, and (ii) unfalsifiable without access to the LLMs’ training data. To make a statement regarding emergent ToM, a careful experiment is needed to ensure that ToM did indeed appear spontaneously and not as a result of other factors such as training on related datasets, exposure to descriptions of clinical tests online, interactions with users, and more.OpenAI acknowledged that GPT-4 was trained on test data from BIGBench (OpenAI, 2023, footnote 5). However, since the data used to train the GPT models is not publicly available, it is impossible to quantify the degree of the potential data leakage.See Appendix 8.5 for an attempt to quantify such data leakage. We echo calls by Dodge et al. (2021) for increased transparency and open-access to the training data of LLMs, which is crucial for scientifically valid and reproducible experiments Rodgers (2023).
Our objective in this study is not to measure benchmark performance or climb leaderboards. It is feasible that techniques such as chain-of-thought prompting (CoT; Wei et al., 2022) would enhance the performance of GPT-4 on tasks where it currently performs poorly. Nevertheless, we need to exercise caution to ensure that the utilization of methods like CoT or others does not excessively guide the models by essentially revealing the task structure to them—just like Clever Hans who appeared proficient in math merely due to subtle hints given by the owner.
Conclusion
Based on our research and replication studies, we conclude that contemporary LLMs demonstrate an enhanced yet limited degree of Theory of Mind abilities. We find that their ToM abilities are often not robust, and in some instances, we identify evidence of their over-reliance on simple heuristics rather than robust generalized reasoning.
Limitations
The datasets used in this study were limited in scope and size; ToM is required in most human interaction, and thus unbounded in scope. In addition, parts of the datasets could be ambiguous, either due to lack of context or inherent ambiguity Plank (2022). Due to this potential ambiguity, some LLMs were safeguarded and refused to answer certain questions; while we attempted to instruct them to respond in the correct format, some LLMs still did not output the right format. This was only an issue for MC-probing, but probability distributions were not available for all LLMs. Future work should investigate how to mitigate this issue via better instructions or methods that map generated answers to multiple choice better (e.g., Niu et al., 2021; Bulian et al., 2022).
Our experiments were conducted with a limited number of LLMs that were accessible at the time of writing, and we did not explore the full spectrum of LLMs that are currently available. Future work could explore the N-ToM abilities displayed by other LLMs, and additionally, explore multimodal models.
Ethical Statement
All the existing and new datasets used in this study are publicly available. The narratives were evaluated by the authors to ensure that they do not contain offensive content.
LLMs may generate offensive content if prompted with certain inputs. However, we used them for evaluation only, with non-offensive inputs, and we did not record their responses.
Acknowledgements
We would like to thank Uri Katz, Royi Rassin, Ori Shapira, Alon Jacoby, Rotem Dror, and Amir DN Cohen for helpful discussions. We thank OpenAI for access to their APIs including GPT-4, and AI21 for the generous budget for using their platform API. This project was partially funded by the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation program, grant agreement No. 802774 (iEXTRACT), the Computer Science Department of Bar-Ilan University, the Vector Institute for AI, Canada CIFAR AI Chairs program, an NSERC discovery grant, DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and AI2.
References
Appendices
Table 5 shows an example from the Tomi dataset. The unexpected transfer test discusses an unexpected (false belief) rather than trivial (true belief) case. ChatGPT solves the more complex task (false belief) while failing on the trivial task, likely due to its exposure to the Salley-Anne task.
2 Ullman’s Variations
Figures 6 and 7 illustrate the variations proposed by Ullman for the examples in ToM-k.
3 Generative LLMs
We provide the technical details regarding the prompts (§8.3.1) and decoding parameters (§8.3.2).
As input to the LLMs, we used (unless written otherwise) an MC-probing setup (§4.2), i.e., concatenation of the original test with all possible answers and an instruction to choose an option. Table 6 exemplifies the prompt for each task.
3.2 Decoding Parameters
A single sample (the first) was selected from each model for the analysis of the stories. We used the hyperparameters detailed below. We chose hyperparameters that minimize randomness and predict the most probable answer (i.e., low temperature, sampling method), and allow for sufficient number of tokens.
Python package transformers implementation (AutoModelForSeq2SeqLM, AutoTokenizer); torch; Generation by generate function; do_sample=True; max_length=50, from_pretrained:google/flan-t5-small, google/flan-t5-base, google/flan-t5-large, google/flan-t5-xl, google/flan-t5-xxl; temperature=0.0001
Python package transformers implementation (T5ForConditionalGeneration, AutoTokenizer); torch; Generation by generate function; do_sample=True; max_length=50; temperature=0.0001
Python package openai model=text-davinci-002, text-davinci-003; Generation by Completion.create function; temperature=0, max_tokens=50
Python package openai model=gpt-3.5-turbo-0301, gpt-4-0314; Generation by ChatCompletion.create function; temperature=0
Python package ai21 model=j2-jumbo-instruct, j2-grande-instruct, j2-jumbo, j2-grande, j2-large; Generation by Completion.execute function; temperature=0, max_tokens=50, topKReturn=0, topP=1, without any panalty
4 Complete Results
Table 7 contains the exhaustive accuracy results for all LLMs on all datasets.
Running the well-organized code provided by Kosinski (2023) we found that task 2 (Unexpected Transfer Task) scored lower than reported for GPT 3.5. Specifically, two samples resulted in clear mispredictions and one sample had borderline predictions that provided the correct answer but in a format that differed from the expected answer (i.e., the first word was not the expected answer). As a result, the score for task 2 was either 85% or 90%, and the average score across the two tasks was either 85% or 87.5%, which is lower than the reported average of 93%.
5 “Emergence” or test data contamination?
We would like to determine whether LLMs generalize or memorize when they solve the ToM tasks Daumé (2017). We explored the possibility that the increase in performance is a result of training on the test data itself. for that purpose we used a second, secret, test set for SocialIQa that was purposefully kept hidden to avoid data contamination and is only available to the original SocialIQa authors as well as through the AI2 leaderboard.https://leaderboard.allenai.org/socialiqa/submissions/public For each test set (i.e., the standard and secret test sets) we randomly sample 11 subsets of 100 questions on which we evaluate gpt3.5-turbo-0301 and gpt-4-0314. Comparing the performance of both models on both test sets samples with a T-test, we found no significant differences, making it inconclusive whether the models were trained on the normal test set or not. As we discuss in Sec 6, this doesn’t mean that ToM has “emerged” in LLMs, since they may have been exposed to training data or similar examples.
6 ToMi’ subsets analysis
Table 8 provides the complete results from the evaluation of GPT-3.5 on the ToMi’ dataset. The same overall conclusion can be drawn from this table as well: although the model can correctly answer simple reading comprehension questions, it doesn’t answer questions that require ToM skill (first and second order) with similar accuracy.
We divided the results into the average score and joint score. The average score is calculated as a simple average on the different types of questions, while the joint score is considers the prediction as correct only if the model answered correctly all the questions from the same story (with a total of 30 stories). The average results emphasize the major gaps between the model’s accuracy on reading comprehension questions to first order questions (“Chloe will look for the boots in the”) and between the first order questions to the second order questions (“Chloe think that Jackson searches for the boots in the”). The joint score reveals that even when the model correctly answers questions about the story, it might still fail to answer more complex questions.