FANToM: A Benchmark for Stress-testing Machine Theory of Mind in Interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, Maarten Sap
Introduction
Existing evaluations for language models’ theory of mind (ToM) – i.e., the ability to understand the mental states (e.g., thoughts, beliefs, and intentions) of others Premack and Woodruff 1978, is primarily focused on using situation descriptions (i.e., narratives) as the target domain Nematzadeh et al. 2018; Le et al. 2019; Sap et al. 2022; Shapira et al. 2023a. However, ToM capabilities play an even more important role in understanding dynamic social interactions, as they form a crucial component of effective communication Frith 1994; Schober 2005. Furthermore, as narratives condense situation information into short texts, reporting biases can cause them to include spurious correlations or surface cues Gordon and Van Durme 2013. These can be exploited by large language models (LLMs) to display illusory ToM -- i.e., a false sense of robust social reasoning by models. We do not believe that current LLMs possess an actual ToM. Please see §8 for further discussions.
In this work, we introduce FANToM, an English benchmark for stress-testing machine ToM in interactions – i.e., conversations. As conversations present interactions in their raw form, they are much less susceptible to reporting biases, and are more aligned with real-world scenarios requiring ToM reasoning. FANToM consists of 10K questions covering 256 multiparty conversations around a certain topic while characters enter and leave the discussion, leading to distinct mental states between characters due to information asymmetry.
The goal of FANToM is to effectively measure how well models can track the belief of multiple characters in conversations where some information may be inaccessible to some participants. For example, in Figure 1, Kailey briefly steps away from the conversation to get a cup of coffee, while the others continue discussing Linda’s new dog. The information exchanged during Kailey’s absence remains unknown to Kailey, and only the information shared after Kailey’s return becomes accessible. We convert factual question-answer pairs to obtain multiple challenging questions about characters’ beliefs concerning the inaccessible information. Our aim is to design questions at different levels that evaluate a model’s capability for a coherent understanding of others’ mental states. In doing so, we are particularly interested in identifying instances of illusory ToM, which we define as situations where a model may answer some questions correctly but fails to answer others that require the same type of ToM reasoning.
The analysis of evaluation results on FANToM reveals several interesting findings (§4): (1) First, existing neural models score significantly lower than humans on individual questions and on the full set of questions by more than 70% on average. (2) While chain-of-thought reasoning (CoT) does improve performance in most models, it does not substantially bridge the gap with human performance. (3) Although our benchmark is not meant for training, we observe that fine-tuning can help models achieve scores higher than human performance on individual question types. However, when it comes to metrics that require coherent responses across multiple question types, the fine-tuned model still significantly underperforms compared to humans. (4) Additionally, we find that models exhibit different error types depending on the format of questions, despite all questions requiring the same underlying reasoning. (5) Moreover, our results indicate that CoT has a selective impact on performance, showing improvement only in specific scenarios.
To the best of our knowledge, FANToM is the first benchmark to introduce conversation-based ToM evaluation for language-based models. Our benchmark design and experiment results yield important insights into the debate around ToM Whang 2023 and the development of artificial general intelligence Metz 2023 in LLMs. We release our benchmark to spark further discussions on evaluating the ToM capabilities of LLMs.
Design Considerations for FANToM
We go over the important design choices that we made when constructing FANToM. Our goal is to incorporate (1) social interactions that necessitate natural theory of mind (ToM) reasoning (§2.1), (2) essential theoretical prerequisites for validating ToM from psychology (§2.2), and (3) empirical findings that must be taken into account when evaluating large language models (§2.3).
To capture the interactive aspect of ToM, we ground our task in natural social interactions – i.e., conversations. By doing so, we gain two key benefits: (1) minimizing reporting bias Gordon and Van Durme 2013 and (2) aligning with real-world scenarios.
Since narratives are condensed descriptions of interactions, the process of deciding what to include or exclude can introduce reporting bias, resulting in artifacts that models exploit. For instance, including “Carlos did not see this, so he does not know currently where the apple is.” in a narrative for ToM evaluation provides a significant clue about the other’s mental state. However, such explicit hints are rarely present in real-world interactions.
Conversations, on the other hand, present interactions in their raw form, without those explicit hints about others’ mental states. During conversations, we reason through the intermediate steps from scratch, thereby grounding the benchmark in conversations enables a more realistic and unbiased assessment of ToM.
2 Meeting Theoretic Requirements
We follow the two important criteria outlined by Quesque and Rossetti 2020 that must be met when designing a task to validate ToM: “non-merging” and “mentalizing”.
(1) “Non-merging”: Evaluation should require the respondent to maintain a distinction between the others’ mental state and its own. For example, suppose someone is asked about the other’s belief regarding the location of the TV remote controller, and both are believing it to be on the sofa. If the respondent answers that the other believes it is on the sofa, it becomes unclear whether the response is based on the respondent’s own belief or the other’s (i.e., merging mental states). Such merging scenario is unsuitable for validating ToM.
Since machines lack emotions or intentions Gros et al. 2022, we exploit information asymmetry when constructing our benchmark to simulate the non-merging mental state scenarios. We design multiparty conversations where specific information is inaccessible to certain characters. While machines do not possess their own point of view, they act as omniscient observers during our evaluation since we provide the entire conversation as input. As a result, the mental states of the model and the character can be regarded as distinct with respect to that information.
(2) “Mentalizing”: Lower-level processes should not be accounted for successful performance of ToM tasks. If a simpler process can explain a phenomenon, it should always be preferred over a more complex one when interpreting the results. For instance, recognizing joy by observing laughter is more of a visual discrimination than reasoning mental representations.
If the correct answer for a ToM task has a high degree of word correlation with a salient part of the given input, it becomes difficult to determine whether the model is accurately ascribing the other’s mental state or simply following a shortcut pattern matching (i.e., the lower-level process). Therefore, such cases should be discouraged when evaluating ToM in neural language models. In FANToM, we create false answers that have high word correlation with the input to verify whether the models can overcome the shortcut pattern matching when reasoning mental states.
3 Seeking Comprehensive Evaluation
Since the performance of LLMs varies significantly based on given prompts Webson and Pavlick 2022, we adopt a series of reiterative questions at various levels for the same input context, including free-form response questions, multiple-choice questions, and straightforward yes or no questions. The inclusion of free-form response questions is important as it aligns with the common usage of LLMs in contrast to multiple-choice questions that are prevalent in existing benchmarks Sakaguchi et al. 2021; Hendrycks et al. 2021. Although their formats are different, all questions in FANToM fundamentally aim to ascertain the same underlying reasoning: “who is aware of the information?” As a result, FANToM enables us to identify illusory ToM instances wherein models deliver accurate responses for one format but struggles to do so for another format.
FANToM Overview
Following the success of previous works Kim et al. 2022; Chen et al. 2023, we automatically construct full conversations using the large language model (LLM) InstructGPT davinci-003 Ouyang et al. 2022. We also generate theory of mind (ToM) question-answer pairs related to the conversation participants’ beliefs using a specially designed pipeline. In preliminary explorations, we find off-the-shelf LLMs struggle with directly generating ToM question-answer pairs for a given conversation. Our pipeline consists of three steps: (1) generate conversations with information asymmetry (§3.1), (2) generate fact question-answer (QA) pairs (§3.2), and (3) construct ToM (e.g., belief) QA pairs from the fact QA pairs (§3.3). We use different evaluation methods for each question types (§3.4), and validate the final dataset (§3.5).
FANToM consists of small talk conversations involving multiple characters, with each conversation centered around a topic (e.g., pets, risk-taking, personal growth). Each topic has several subtopics, e.g. the topic “pets” may include subtopics “breed” and “special moves”. Initially, the conversation begins with two or three characters. As the conversation progresses, characters join and leave the discussion and the conversation’s subtopic changes over time. Conversations include explicit indications of leaving and joining, such as utterances like “Hey guys, I’ll go grab a coffee.” or “Hey, I’m back, what are you guys discussing now?” shown in Figure 1. During the absence of a character, the conversation continues and information is shared among the remaining participants, creating a natural information asymmetry that reflects real-life interactions. After a series of utterances, the character who was absent (re)joins the conversation, unaware of the information that was previously shared with other participants. More details are in Appendix A.1.
Many existing ToM tasks involve some form of asymmetry between characters Braüner et al. 2020. For example, in the Sally-Anne task, Sally does not know that Anne relocated the object, while the observer is aware of the action. In the Smarties task, the character in the story does not know the label changed, whereas the observer is fully aware of this situation. This inherent asymmetry ensures two distinct mental states (i.e., the non-merging criterion; §2.2) to be present during the experiments.
2 Factual Question-Answer (QA) Pairs
The conversations in FANToM include factual question-answer pairs (FactQ) about the inaccessible information—i.e., the information that a specific character is unaware of. An example question would be “What is the breed of Linda’s dog?” in Figure 1. More details are in Appendix A.2.
There are two distinct types of answers for each FactQ: (1) Full Fact A and (2) Limited Fact A. The Full Fact A incorporates the full information in the preceding conversation where the character PersonX was absent. On the other hand, Limited Fact A relies only on the conversation in which PersonX participated. The former answer is based on information that PersonX does not have access to, while the latter answer only takes into account the accessible information for PersonX. For cases where no information was shared regarding the FactQ, the Limited Fact A indicates that no information has been provided. Examples of these answers are in Figure 1. “Linda has a golden retriever.” is based on the preceding conversation where Kailey was absent, while “There is no information on the breed of Linda’s dog” is based on the conversation Kailey participated.
3 ToM QAs based on Fact QAs
For each FactQ, we build six types of ToM QA. Construction details can be found in Appendix A.3.
These questions are created by rephrasing the FactQ to ask beliefs of characters in the conversation. We are particularly interested in PersonX’s belief about the inaccessible information from the previous conversation, in which PersonX did not participate. For example, the FactQ of “What is the breed of Linda’s dog?” in Figure 1 is converted to “What breed would Kailey think Linda’s dog is?” The BeliefQ[Dist.] requires free-form response, while BeliefQ[Choice] provides multiple-choice options for the same question.
The options for BeliefQ are created by rephrasing the Full Fact A and Limited Fact A. For example, the “Linda has a golden retriever.” in Figure 1 is converted to “Kailey believes Linda has a golden retriever.” Since the Full Fact A reflects information that is not accessible to PersonX and the Limited Fact A incorporates only the information accessible to PersonX, we label the converted Full Fact A and Limited Fact A as “Omniscient-view Belief A” and “PersonX-centric Belief A”, respectively.
(3) Answerability Q[List]:
Given the FactQ, we ask models “List all the characters who know the correct answer to this question”. In essence, we are interested in whether the model can identify who among the participants can correctly answer the FactQ. This is a meta-question that necessitates two-step reasoning: first determining the answer itself, and second, identifying the characters who have access to this knowledge.
(4) InfoAccess Q[List]:
Here, we provide the Full Fact A with the FactQ and ask the model “List all the characters who know this information”. Essentially, this question aims to identify the individuals who have knowledge or access to this information. Since the information is explicitly provided to the model, only the second reasoning step of the Answerability Q[List] is required.
(5) Answerability Q[Y/N] and (6) InfoAccess Q[Y/N]:
We ask models to determine, through a simple binary response (yes or no), whether each character is capable of answering the question or knows the information. For example, we ask models “Does David know the correct answer to this question?” and “Does Sally know about this information?” (Figure 1).
4 Evaluation
Each question is provided to the model along with the conversation as input. This makes the model an omniscient observer, having access to all information shared in the conversation. On the other hand, PersonX was absent for a while, thereby an information asymmetry naturally arises between the model and PersonX. Responses that include inaccessible information for PersonX indicate a lack of ToM in the model.
FANToM comprises two types of input conversations: short and full. In the case of short input, the model is provided with the conversation that only includes the part where the specific speaker left and (re)joined, while excluding the other earlier and later parts of the conversation. On the other hand, a full conversation encompasses the entire discussion on the main topic, including all subtopics. As a result, this is significantly longer than the short input.
BeliefQ[Dist.]
When given a belief question regarding PersonX, the model should generate a response that incorporates only the information accessible to PersonX. We use cosine similarity to measure the distance between SentenceBERT Reimers and Gurevych 2019 embeddings of each option and response. A correct response should always be closer to the PersonX-centric Belief A than the Omniscient-view Belief A.
To accurately assess the performance of the response, we also calculate the token F1 score for responses that are considered correct based on the distance metric, following the convention of various QA tasks Rajpurkar et al. 2016; Rajpurkar et al. 2018. When comparing distances in the embedding space, nonsensical responses (e.g., repetition of character names) can be deceptively closer to PersonX-centric Belief A, resulting in misleading accuracy. Therefore, models must score high on both the distance and F1 metrics for the BeliefQ[Dist.].
BeliefQ[Choice]
The model should choose between the Omniscient-view Belief A and the PersonX-centric Belief A. The correct answer is the PersonX-centric Belief A.
Answerability Q[List] and InfoAccess Q[List]
A correct response must include all characters who have access to the answer or information while excluding all characters who do not. No partial marks are assigned.
Answerability Q[Y/N] and InfoAccess Q[Y/N]
The model should respond with “yes” or “true” for all characters who have access to the answer or information, and with “no” or “false” for all characters who do not. More details are in Appendix A.4.
5 Dataset Validation & Statistics
To ensure the quality of our benchmark, we go through a manual validation process for all conversations and question-answer pairs using Amazon Mechanical Turk (MTurk). We conduct validation on the entire conversations in our dataset using 32 annotators who passed a qualification test for assessing conversation coherence. We ask workers to flag conversations that are incoherent or unsafe (e.g., unethical, biased, harmful, dangerous, or offensive). Each conversation is validated by three workers. While 10 conversations received votes for incoherence, none achieved a majority vote indicating they were incoherent. We refine all 10 conversations. As for safety, no conversations were voted as being unsafe. We also request workers to verify the answers provided for BeliefQ[Choice]s. We remove all question sets that were marked as erroneous by the worker (8.6%).
Statistics
FANToM is composed of 256 conversations with 1,415 BeliefQ[Dist.]s and BeliefQ[Choice]s, 703 FactQs, Answerability Q[List]s, and InfoAccess Q[List]s, respectively. Additionally, there are 2,689 Answerability Q[Y/N]s and InfoAccess Q[Y/N]s. Given that the Answerability Q[Y/N]s and InfoAccess Q[Y/N]s iterate over all characters present in the conversations, they have the highest count among all the question types.
The average number of turns in the input context is 13.8 (short conversation), and the average number of words in each turn is 21.9. For reference, the corresponding statistics for ToMi Le et al. 2019 are 4.9 and 4.2, respectively. More statistics can be found in Appendix A.5.
Experiments
We test a total of thirteen recent instruction-tuned neural language models: GPT-4 (OpenAI 2023, gpt-4-0613 and gpt-4-0314;), ChatGPT (OpenAI 2022, gpt-3.5-turbo-0613;), InstructGPT (Ouyang et al. 2022, davinci-003 and curie-001;), Flan-T5-XL and Flan-T5-XXL Chung et al. 2022, Flan-UL2 Tay et al. 2023, Falcon Instruct (Almazrouei et al. 2023, 7B and 40B;), Mistral Instruct 7B Jiang et al. 2023, Zephyr 7B HuggingFace 2023, and Llama-2 Chat 70B Touvron et al. 2023. Descriptions for each model are in Appendix B.
Although our benchmark is not meant for training, we also fine-tune Flan-T5-XL Chung et al. 2022 by randomly splitting FANToM according to the conversation’s main topics. We then test the model on unseen conversation topics. More details can be found in Appendix B.
Human Performance
We also measure human performance by asking graduate students in computer science. We ask BeliefQ[Choice], Answerability Q[List], and InfoAccess Q[List], given a conversation. As it is redundant to ask human testees binary questions when they have already been asked Answerability Q[List] and InfoAccess Q[List], we do not ask Answerability Q[Y/N] and InfoAccess Q[Y/N]. To ensure a fair comparison with the models, we give the same instructions to humans and no other tutorials, examples, or extra instructions were given. Student volunteers solved 32 sets in total.
Metrics
We report accuracy for BeliefQ[Dist.], BeliefQ[Choice], Answerability Q[List], and InfoAccess Q[List]. The weighted F1 scores are reported for Answerability Q[Y/N] and InfoAccess Q[Y/N]. We additionally report the “All” score for the Answerability Q and InfoAccess Q requiring models to be correct on both list-type and binary-type questions. For BeliefQ[Dist.] and FactQ, we also report the token F1 scores to measure the word overlap between the answer and model’s free-form response.
Moreover, we report the All* score which requires the models to answer all six ToM question types (§3.3) in the set correctly for the same information piece in the conversation. This metric aims to measure how well the models show consistent understanding across different types of questions. To compare with human performance, we also report the All score, which only excludes the BeliefQ[Dist.] from the All* score.
1 Results
All the models exhibit scores that are significantly worse than human performance. Table 9 shows the full results of state-of-the-art large language models (LLMs) on FANToM. We break down the table and highlight each discussion point below.
Figure 2 shows the results of a few selected models. We find models perform significantly better on BeliefQ[Choice] compared to Answerability Q[List] and InfoAccess Q[List]. Despite the Answerability Q[List] and InfoAccess Q[List] being pre-requisites for solving BeliefQ[Choice], they are much more challenging for models. Furthermore, models’ performance sharply drops when evaluated for coherent reasoning across multiple question types with the same underlying theory of mind (ToM) reasoning (i.e., All Question Types). These findings suggest that some instances of successful LLM ToM reasoning in FANToM should be interpreted as illusory.
Chain-of-thought and Fine-tuning
Table 1 summarizes the results when we apply zero-shot chain-of-thought (CoT) reasoning or fine-tuning to models. For CoT, we follow Kojima et al. 2022 and use the prompt “let’s think step by step”. We observe an improvement in scores with CoT applied. However, there are still significant score gaps compared to human performance.
We also find fine-tuned Flan-T5 XL still falls short of human performance in metrics that demand consistent accuracy across multiple questions—i.e., the All scores. We find fine-tuning achieves scores comparable with human performance on individual question types (see Table 9). Although our benchmark is not intended for training purposes, developing models with a coherent ToM reasoning remains challenging, even with explicit training on the data.
Comprehending Facts vs. Distinct Beliefs
Figure 3 shows the token F1 scores for FactQ and accuracy for BeliefQ[Dist.]. The token F1 scores for FactQ can be seen as a measure of a model’s basic comprehension capability for interactions. Scoring high in FactQ indicates the model is good at identifying the most relevant information piece to answering the question. Despite its small size, Mistral Instruct 7B shows the strongest performance among the open-source models.
On the other hand, BeliefQ[Dist.] aims to measure a model’s understanding of individual characters’ perspective of a particular information—i.e., belief. To meet the mentalizing criterion (see §2.2), we deliberately design the incorrect answers in BeliefQ[Dist.] to have greater word overlap with the context than correct answers. Also, BeliefQ[Dist.] are rephrased questions inquiring about PersonX’s belief for the facts in FactQ, thereby the two question types share significant word overlap. However, the same information that was used to answer FactQ should not be included in the response for BeliefQ[Dist.] on PersonX as it is from the conversation that PersonX missed. As a result, certain models with higher token F1 scores for FactQ have lower scores for BeliefQ[Dist.] compared to models that perform worse on FactQ (e.g., InstructGPT davinci-003 vs. Llama-2 Chat and Mistral Instruct). This suggests the models lack the ability to comprehend distinct perspectives of individual characters, leading them to reproduce similar responses to FactQ for BeliefQ[Dist.].
Free-Response vs. Choice
We observe a pattern where models score significantly worse in free-response questions than choice questions (BeliefQ[Dist.] vs. BeliefQ[Choice]; Figure 3 and 2). This pattern is consistent for Answerability Q[List] and Answerability Q[Y/N], as well as for InfoAccess Q[List] and InfoAccess Q[Y/N] (see Table 9). However, many of them still achieve scores either below or around 50, which is the random baseline for those binary choice questions.
Reasoning Complexity
Table 2 compares models’ performance between Answerability Q[Y/N] and InfoAccess Q[Y/N]. As Answerability Qs require an additional step of reasoning compared to InfoAccess Qs, models consistently perform worse on Answerability Q[Y/N] compared to InfoAccess Q[Y/N]. However, this pattern is not consistent across models for Answerability Q[List] and InfoAccess Q[List] (see Figure 2). This may be because models significantly struggle with Answerability Q[List] and InfoAccess Q[List], potentially resulting in the absence of meaningful performance patterns.
Short vs. Full Conversations
When a model is provided with the full conversation (Table 9, bottom), its performance noticeably decreases compared to when it is given only the relevant parts of the conversation (Table 9, top). The decrease can be attributed to the model’s need to identify the relevant information within the full conversation, whereas it does not have to do so for the short conversations. This indicates theory of mind reasoning becomes even more challenging for models when it needs to be combined with different types of reasoning (e.g., search).
2 In-depth Analysis
Figure 4 and 5 summarize the error types of Answerability Q and InfoAccess Q for each model with and without chain-of-thought (CoT) reasoning. For list-type questions, models make more errors by including characters who are unaware of the information in the responses, rather than excluding characters who are aware. Interestingly, when CoT is applied, the error of including unaware characters decreases, whereas the error of excluding characters who are aware increases for most models.
In the case of binary questions, false positives and false negatives correspond to including characters who are unaware and excluding characters who are aware in the response for list-type questions, respectively. If the model fails to generate a yes or no response, we mark it as irrelevant. Models tend to exhibit false negative responses more frequently for binary questions compared to list-type questions. Similarly, CoT primarily helps the model in reducing the false positive error rates, but the reduction in false negative error rates is not consistent across models. This suggests that CoT selectively improves reasoning specifically for determining characters who are unaware of the information, rather than characters who are aware.
How accurate and consistent are models’ answers for a given character?
For accuracy, we report the All for Each Character score which is determined by whether the models are able to answer all six types of ToM questions correctly regarding the specific character. For consistency, we measure the ratio of consistent model responses across Answerability Q and InfoAccess Q for each character. Table 3 shows the accuracy and consistency of the models’ responses for each character within the given conversation context. Overall, we observe a pattern where models that score low in accuracy also show low consistency. While CoT generally improves model performance (see Table 9), we find that it does not always lead to improved accuracy and consistency. The decrease in All for Each Character score when CoT is applied suggests that CoT has a selective impact on different question types.
Are there differences in performance in terms of the order of ToM beliefs?
Table 4 presents the results of BeliefQ with respect to different orders of ToM beliefs. Similar to Le et al. 2019, models perform better on the second-order belief questions than those with first-order beliefs. To further investigate the performance on second-order belief questions, we analyze the results based on the cyclic and acyclic patterns in them. The cyclic second-order belief questions inquire about Character 1’s belief regarding Character 2’s belief about Character 1 (e.g., What does Linda think about Kailey’s belief on the breed of Linda’s dog?); while the acyclic second-order questions focus on Character 1’s belief about Character 2’s belief regarding Character 3 (e.g., What does David think about Kailey’s belief on the breed of Linda’s dog?). Models show better performance on the cyclic questions than acyclic ones, which include more characters to track. However, when CoT is applied, the increase in score for acyclic questions is greater than that of cyclic ones, suggesting CoT helps multi-tracking.
Related Work
Many theory of mind (ToM) benchmarks evaluate models on false beliefs with narratives (Grant et al. 2017; Nematzadeh et al. 2018; Le et al. 2019; Gandhi et al. 2023; Zhou et al. 2023). Other works such as Shapira et al. 2023b build benchmarks based on the Faux Pas Test (Baron-Cohen et al. 1999). Also, ToM-related benchmarks focus on reasoning emotions and mental states in narratives (Rashkin et al. 2018; Sap et al. 2019).
Theory of Mind in Large Language Models
Although qualitative assessments might imply a degree of ToM in large language models (Whang 2023, LLMs;), more comprehensive quantitative investigations reveal that they have yet to achieve human-level ToM across various benchmarks (Sap et al. 2022; Shapira et al. 2023a). LLMs struggle to reason ToM robustly Ullman 2023, though their performance can be improved through few-shot samples and chain-of-thought prompting Sap et al. 2022; Moghaddam and Honey 2023 as well as specific inference methods Sclar et al. 2023.
Conclusion & Discussion
We introduced FANToM, a new benchmark for stress-testing theory of mind (ToM) capabilities of neural language models in conversations via question answering. Our benchmark is built upon essential theoretical requisites and empirical considerations required for validating ToM in large language models (LLMs). The conversations in our benchmark involve information asymmetry, with characters joining and leaving the discussion while it continues, to simulate distinct mental states. To identify illusory ToM, we crafted multiple types of challenging belief questions regarding the conversation participants’ mental states by converting factual questions. Our evaluation results show that coherent ToM reasoning is challenging for current LLMs, performing significantly worse than humans even when using chain-of-thought reasoning or fine-tuning.
Although there has been recent debates around whether current LLMs possess ToM capabilities or not Whang 2023, our results indicate that this capacity has not yet emerged in any manner. Previous instances of success on well-known psychology ToM tests may be attributed to exposure during the pretraining phase Ullman 2023. Our work highlights the need for novel interaction-oriented benchmarks that introduce scenarios not encountered during training, and also aligning more closely with real-world use cases as LLMs are increasingly being deployed in interactive settings.
Our results also shed light on a broader issue in neural models – the lack of internal consistency Elazar et al. 2021. We find they often fail to provide consistent answers to questions requiring the same underlying ToM reasoning. To address this concern, future works can explore various directions, such as grounding reasoning in pragmatics Kim et al. 2020, visual information Bisk et al. 2020, or belief graphs Sclar et al. 2023.
Another issue that our work touches upon is the reporting biases inherent in language models. We observed that models often exhibit biases in their responses, showing a tendency to overly rely on the information they are conditioned on, such as preferring answers that have high overlap with the context Sugawara et al. 2018. However, to achieve successful ToM reasoning, it is crucial to distinguish between accessible and inaccessible information for a particular agent, rather than blindly using all information available to the model. One potential approach to mitigate this is to combine pretraining with interactive learning Sap et al. 2022.
In the spirit of encouraging future research in this direction, we make our benchmark publicly available at https://hyunw.kim/fantom.
Limitations
Although FANToM is the first benchmark, to the best of our knowledge, to cover theory of mind (ToM) reasoning in conversational interactions, it is currently limited to small talks on specific topics. Additionally, our benchmark only considers only a single type of relationship between conversation participants, where they do not have prior knowledge of each other. However, social reasoning can become much more dynamic when variables such as relationships (e.g., family, friends, co-workers) are introduced. ToM is essential in all conversational interactions, hence we strongly encourage future works to evaluate ToM in a wider range of diverse conversation scenarios.
Our evaluation solely focuses on language-based models. However, it is important to note that ToM extends beyond a single modality Piaget 1956; Wu and Keysar 2007. For instance, the well-known Sally-Anne test Wimmer and Perner 1983; Baron-Cohen et al. 1985 is typically conducted as a face-to-face experiment, where visual cues affect the performance of the participants. Therefore, interesting future work will involve examining the capabilities of multi-modal models in relation to ToM reasoning.
Lastly, as we generate full conversations with large language models, conversations may contain offensive contents Weidinger et al. 2021. However, we specifically select casual topics for small talks (e.g., pets, personal growth, traveling) to minimize the likelihood of offensive content generation. Also, we manually validate all conversations in our benchmark with crowdworkers from Amazon Mechanical Turk.
Societal and Ethical Considerations
We acknowledge that the term “theory of mind” (ToM) may evoke anthropomorphic connotations regarding AI models. However, we emphasize that the purpose of our work is not to promote anthropomorphism of AI models. Rather, our focus lies in exploring the limitations of existing language models in social reasoning. While the concept of ToM attempts to capture the ability to attribute mental states to oneself and others Premack and Woodruff 1978, it is important to clarify that AI models do not possess subjective consciousness or true understanding of intentions, beliefs, or desires. Our experiment results also demonstrate that current large language models do not exhibit any coherent ToM reasoning; instead, they primarily rely on word correlations.
Acknowledgement
We thank the participants who contributed to the human performance measurement. We also appreciate our colleagues on the Beaker Team at the Allen Institute for AI for helping with the compute infrastructure. This work was supported in part by DARPA MCS program through NIWC Pacific (N66001-19-2-4031). Hyunwoo Kim and Gunhee Kim are supported by the Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No.2019-0-01082, SW StarLab; and No.2022-0-00156, Fundamental research on continual meta-learning for quality enhancement of casual videos and their 3D metaverse transformation). Lastly, we also thank OpenAI, as well as Google Cloud Compute.
References
Appendix A FANToM Construction
Full examples of question sets in FANToM can be found in Table 5 and Table 6.
To create the conversations in our benchmark, we use a predefined set of subtopics for each main topic and employ templates to generate scripts. For example, for the topic “pets” subtopics may include “breed”, “special moves”, and “favorite food”. Following Kim et al. 2022, we use specific speaker prefixes with English names sampled from the Top-1K names in the US SSN database for more natural conversations. We append each utterance with speaker prefixes. We randomly shuffle the subtopics for each topic and generate conversations for each subtopic. We generate the first conversation with the following prompt: “{Character 1}, {Character 2}, ... {Character n} met for the first time at this social event. They are having a conversation on their {topic}. They now discuss {subtopic}.\n{Character 1}:” The initial conversation starts with two or three characters and there can be up to five characters who are participating in the conversation at the same time.
Then, for each subtopic, we randomly select characters to join or leave the conversation. We use the following prompt when a character is selected to leave: “Now, {leaving character} leaves the conversation because of the reason ’{leaving reason}’. They now discuss {subtopic}. Remember to indicate that {leaving character} is leaving the conversation. {Conversation history} \n{leaving character}:”. We use a predefined list of 64 reasons for leaving the conversation. Table 7 shows all reasons for leaving. We append the previous conversation history to the input prompt to make the conversation continue from the previous one.
We use the following prompt when a character is selected to join: “Now {joining character} comes back after leaving the conversation because of the reason {leaving reason}. They now discuss {subtopic}. Remember to indicate that {joining character} is joining the conversation. Do not mention the details in the previous conversations. {Conversation history} \n{joining character}:”.
Extracting the inaccessible information for PersonX
Whenever a character (re)joins the conversation, we extract the inaccessible information by asking GPT-4 (gpt-4-0314) what information was shared in the preceding conversation where the character PersonX did not participate. We provide the previous conversation and the current one as input to GPT-4 with the prompt “What information was shared before PersonX joined, but was not mentioned after PersonX joined?” appended to it. To ease the task, the joining of the character is explicitly denoted by inserting a script between the conversations, as follows: "Previous conversation\n[PersonX joined the conversation]\nCurrent conversation". We observe quality improvements for the output generated by GPT-4 with the inclusion of the hint script. The returned result can be viewed as a conversation summary explicitly covering the previous context.
A.2 Generating Factual QA Pairs
We construct factual question-answer (QA) pairs related to the inaccessible information. First, we generate three non-yes-or-no questions and denote these as “FactQs” and obtain them by prompting GPT-4, given the inaccessible information text. We obtain “FactQs” by prompting GPT-4 with the following: “{inaccessible information}\n\nBased on this, formulate three non-yes-or-no questions that can be answered by this conversation summary.”
Next, we generate two distinct types of answers for each FactQ with GPT-4. (1) First, we generate an answer denoted as “Full Fact A”, which is based on the preceding conversation where PersonX was absent. This answer incorporates the full information by providing GPT-4 with the previous conversation – i.e., the source of the inaccessible information for PersonX. (2) Second, we generate another answer referred to as “Limited Fact A”, which relies only on the conversation where PersonX participated. In this case, we give GPT-4 the PersonX-participating conversation along with the FactQ. We prompt GPT-4 with the following: “{context}\n\nQuestion: {FactQ }\nAnswer:”
A.3 Constructing Belief QAs with Factual QAs
We first convert FactQs into first-order or second-order ToM questions asking about beliefs of characters in the conversation. We are particularly interested in PersonX’s belief or knowledge about the inaccessible information from the previous conversation, in which PersonX did not participate. We prompt GPT-4 with the following: “{FactQ }\n\nConvert this into a theory of mind question asking {character name}’s belief about this.”
Next, we convert the Full Fact As and Limited Fact As into answers about beliefs. Since the Full Fact As reflect information that is not accessible to PersonX and the Limited Fact A incorporates only the information accessible to PersonX, we label the converted Full Fact A and Limited Fact A as “Omniscient-view Belief A” and “PersonX-centric Belief A”, respectively. For the conversion, we prompt GPT-4 with the following format: ‘‘Question: FactQ \n\nAnswer the question using the following sentence. {Full Fact A or Limited Fact A }\nAnswer:’’.
A.4 Evaluation for Answerability Q[Y/N] and InfoAccess Q[Y/N]
We use pattern matching to parse the yes or no answers from model responses. We regard “yes”, “knows”, “does know”, and “true” as responses representing “yes”. Similarly, we regard “no”, “does not know”, “doesn’t know”, and “false” as responses representing “no”.
A.5 Statistics for FANToM
Table 8 compares the basic statistics of FANToM and ToMi Le et al. 2019.
Appendix B Experiments
A total of 11 student volunteers participated in the evaluation. For each question set, we assign a single testee. They solved a total of 32 sets. To ensure a fair comparison, no additional tutorials, examples, or extra instructions were provided beyond what was given to the models.
Baseline models
The GPT models are proprietary models from OpenAI based on the decoder-only transformer architecture. Flan-T5 and Flan-UL2 are open-source (i.e., Apache 2.0) models from Google trained on instruction-phrased datasets. They are based on the encoder-decoder transformer architecture. Falcon Instruct is another open-source (i.e., Apache 2.0) model trained on RedefinedWeb Penedo et al. 2023 and Baize Xu et al. 2023. Llama-2 Chat Touvron et al. 2023 is a fine-tuned 70B large language model, optimized for following user requests in dialogue format. Mistral Instruction Jiang et al. 2023 is a 7B language model fine-tuned to follow instructions, which is reported to surpass the Llama-2 Chat 13B model. Zephyr HuggingFace 2023 is a model based on Mistral, further fine-tuned on UltraChat Ding et al. 2023 and aligned with UltraFeedback Cui et al. 2023.
Results of other models
Table 9 shows the results for other large language models not included in Figure 2. Given the random baseline score is 50 for BeliefQ[Choice], Answerability Q[Y/N], and InfoAccess Q[Y/N], most of the models show low performance on our benchmark.
Fine-tuning details
We fine-tune Flan-T5-XL with learning rate=2e-5 and weight decay=0.01, evaluating per epoch and using early stopping with patience 1 (batch size = 3 for Flan-T5-XL). We observe an increase in validation loss after the first epoch. We also add special tokens before and after the completions to prevent the model from over-generating, which we find in early experiments. We also fine-tune text-curie-001 Ouyang et al. 2022 for two epochs using standard parameters from the OpenAI API.