Language models show human-like content effects on reasoning tasks
Ishita Dasgupta, Andrew K. Lampinen, Stephanie C. Y. Chan, Hannah R. Sheahan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, Felix Hill
Introduction
A hallmark of abstract reasoning is the ability to systematically perform algebraic operations over variables that can be bound to any entity (2, 3): the statement: ‘X is bigger than Y’ logically implies that ‘Y is smaller than X’, no matter the values of X and Y. That is, abstract reasoning is ideally content-independent (2). The capacity for reliable and consistent abstract reasoning is frequently highlighted as a crucial missing component of current AI (4, 5, 6). For example, while large Language Models (LMs) exhibit some impressive emergent behaviors, including some performance on abstract reasoning tasks (7, 8, 9, 10, 11; though cf. 12), they have been criticized for failing to achieve systematic consistency in their abstract reasoning (e.g. 13, 14, 15, 16).
However, humans — arguably the best known instances of general intelligence — are far from perfectly rational abstract reasoners (17, 18, 19). Patterns of biases in human reasoning have been studied across a wide range of tasks and domains (18). Here, we focus in particular on ‘content effects’ — the finding that humans are affected by the semantic content of a logical reasoning problem. In particular, humans reason more readily and more accurately about familiar, believable, or grounded situations, compared to unfamiliar, unbelievable, or abstract ones. For example, when presented with a syllogism like the following:
humans will often classify it as a valid argument. However, when presented with:
humans are much less likely to say it is valid (20, 21, 22) — despite the fact that the arguments above are logically equivalent (both are invalid). Similarly, humans struggle to reason about how to falsify conditional rules involving abstract attributes (1, 23), but reason more readily about logically-equivalent rules grounded in realistic situations (24, 25, 26). This human tendency also extends to other forms of reasoning e.g. probabilistic reasoning, where humans are notably worse when problems do not reflect intuitive expectations (27).
The literature on human cognitive biases is extensive, but many of these biases can be idiosyncratic and context-dependent. For example, even some of the seminal findings in the influential work of Kahneman et al. (18), like ‘base rate neglect’, are sensitive to context and experimental design (28, 29), with several studies demonstrating exactly the opposite effect in a different context (30). However, the content effects on which we focus have been a notably consistent finding and have been documented in humans across different reasoning tasks and domains: deductive and inductive, or logical and probabilistic (31, 1, 32, 33, 20, 27). This ubiquity is notable and makes these effects harder to explain as idiosyncracies. This ubiquitous sensitivity to content is in direct contradiction with the definition of abstract reasoning: that it is independent of content, and speaks directly to longstanding debates over the fundamental nature of human intelligence: are we best described as algebraic symbol-processing systems (2, 34), or emergent connectionist ones (35, 36) whose inferences are grounded in learned semantics? Yet explanations or models of content effects in the psychological sciences often focus on a single (task and content-specific) phenomenon and invoke bespoke mechanisms that only apply to these specific settings (e.g. 25). Could content effects be explained more generally? Could they emerge from simple learning processes over naturalistic data?
In this work, we address these questions, by examining whether language models show this human-like blending of logic with semantic content effects. Language models possess prior knowledge — expectations over the likelihood of particular sequences of tokens — that are shaped by their training. Indeed, the goal of the “pre-train and adapt” or the “foundation models” (37) paradigm is to endow a model with broadly accurate prior knowledge that enables learning a new task rapidly. Thus, language model representations often reflect human semantic cognition; e.g., language models reproduce patterns like association and typicality effects (38, 39), and language model predictions can reproduce human knowledge and beliefs (40, 41, 42, 43). In this work, we explore whether this prior knowledge impacts a language model’s performance in logical reasoning tasks. While a variety of recent works have explored biases and imperfections in language models’ performance (e.g. 13, 14, 15, 44, 16), we focus on the specific question of whether content interacts with logic in these systems as it does in humans. This question has significant implications not only for characterizing LMs, but potentially also for understanding human cognition, by contributing new ways of understanding the balance, interactions, and trade-offs between the abstract and grounded capabilities of a system.
We explore how the content of logical reasoning problems affects the performance of a range of large language models (45, 46, 47). To avoid potential dataset contamination, we create entirely new datasets using designs analogous to those used in prior cognitive work, and we also collect directly-comparable human data with our new stimuli. We find that language models reproduce human content effects across three different logical reasoning tasks (Fig. 1). We first explore a simple Natural Language Inference (NLI) task, and show that models and humans answer fairly reliably, with relatively modest influences of content. We then examine the more challenging task of judging whether a syllogism is a valid argument, and show that models and humans are biased by the believability of the conclusion. We finally consider realistic and abstract/arbitrary versions of the Wason selection task (1) — a task introduced over 50 years ago that demonstrates a failure of systematic human reasoning — and show that models and humans perform better with a realistic framing. Our findings with human participants replicate and extend existing findings in the cognitive literature. We also report novel analyses of item-level effects, and the effect of content and items on continuous measures of model and human responses. We close with a discussion of the implications of these findings for understanding human cognition as well as language models.
In this work, we evaluate content effects on three logical reasoning tasks, which are depicted in Fig. 1. These three tasks involve different types of logical inferences, and different kinds of semantic content. However, these distinct tasks admit a consistent definition of content effects: the extent to which reason is facilitated in situations in which the semantic content supports the correct logical inference, and correspondingly the extent to which reasoning is harmed when semantic content conflicts with the correct logical inference (or, in the Wason tasks, when the content is simply arbitrary). We also evaluate versions of each task where the semantic content is replaced with nonsense non-words, which lack semantic content and thus should neither support nor conflict with reasoning performance. (However, note that in some cases, particularly the Wason tasks, changing to nonsense content requires more substantially altering the kinds of inferences required in the task; see Methods.)
The first task we consider has been studied extensively in the natural language processing literature (48). In the classic Natural Language Inference (NLI) problem, a model receives two sentences, a ‘premise’ and a ‘hypothesis’, and has to classify them based on whether the hypothesis ‘entails’, ‘contradicts’, or ‘is neutral to’ the premise. Traditional datasets for this task were crowd-sourced (49) leading to sentence pairs that don’t strictly follow logical definitions of entailment and contradiction. To make this a more strictly logical task, we follow Dasgupta et al. (50) to generate comparisons (e.g. X is smaller than Y). We then give participants an incomplete inference such as “If puddles are bigger than seas, then…” and give them a forced choice between two possible hypotheses to complete it: “seas are bigger than puddles” and “seas are smaller than puddles.” Note that one of these completions is consistent with real-world semantic beliefs i.e. ‘believable’ while the other is logically consistent with the premise but contradicts real world beliefs. We can then evaluate whether models and humans answer more accurately when the logically correct hypothesis is believable; that is, whether the content affects their logical reasoning.
However, content effects are generally more pronounced in difficult tasks that require extensive logical reasoning (33, 21). We therefore consider two more challenging tasks where human content effects have been observed in prior work.
Syllogisms (51) are a simple argument form in which two true statements necessarily imply a third. For example, the statements “All humans are mortal” and “Socrates is a human” together imply that “Socrates is mortal”. But human syllogistic reasoning is not purely abstract and logical; instead it is affected by our prior beliefs about the contents of the argument (20, 22, 52). For example, Evans et al. (20) showed that if participants were asked to judge whether a syllogism was logically valid or invalid, they were biased by whether the conclusion was consistent with their beliefs. Participants were very likely (90% of the time) to mistakenly say an invalid syllogism was valid if the conclusion was believable, and thus mostly relied on belief rather than abstract reasoning. Participants would also sometimes say that a valid syllogism was invalid if the conclusion was not believable, but this effect was somewhat weaker (but cf. 53). These “belief-bias” effects have been replicated and extended in various subsequent studies (22, 53, 54, 55). We similarly evaluate whether models and humans are more likely to endorse an argument as valid if its conclusion is believable, or to dismiss it as invalid if its conclusion is unbelievable.
The Wason Selection Task (1) is a logic problem that can be challenging even for subjects with substantial education in mathematics or philosophy. Participants are shown four cards, and told a rule such as: “if a card has a ‘D’ on one side, then it has a ‘3’ on the other side.” The four cards respectively show ‘D’, ‘F’, ‘3’, and ‘7’. The participants are then asked which cards they need to flip over to check if the rule is true or false. The correct answer is to flip over the cards showing ‘D’ and ‘7’. However, Wason (1) showed that while most participants correctly chose ‘D’, they were much more likely to choose ‘3’ than ‘7’. That is, the participants should check the contrapositive of the rule (“not 3 implies not D”, which is logically implied), but instead they confuse it with the converse (“3 implies D”, which is not logically implied). This is a classic task in which reasoning according to the rules of formal logic does not come naturally for humans, and thus there is potential for prior beliefs and knowledge to affect reasoning.
Indeed, the difficulty of the Wason task depends upon the content of the problem. Past work has found that if an identical logical structure is instantiated in a common situation, particularly a social rule, participants are much more accurate (56, 24, 25, 26). For example, if participants are told the cards represent people, and the rule is “if they are drinking alcohol, then they must be 21 or older” and the cards show ‘beer’, ‘soda’, ‘25’, ‘16’, then many more participants correctly choose to check the cards showing ‘beer’ and ‘16’. We therefore similarly evaluate whether language models and humans are facilitated in reasoning about realistic rules, compared to more-abstract arbitrary ones. (Note that in our implementations of the Wason task, we forced participants and language models to choose exactly two cards, in order to most closely match answer formats between the humans and language models.)
The extent of content effects on the Wason task are also affected by background knowledge; education in mathematics appears to be associated with improved reasoning in abstract Wason tasks (57, 58). However, even those experienced participants were far from perfect — undergraduate mathematics majors and academic mathematicians achieved less than 50% accuracy at the arbitrary Wason task (57). This illustrates the challenge of abstract logical reasoning, even for experienced humans. As we will see in the next section, many human participants did struggle with the abstract versions of our tasks.
Results
We summarize our primary results in Fig. 2. In each of our three tasks, humans and models show similar levels of accuracy across conditions. Furthermore, humans and models show similar content effects on each task, which we measure as the degree of advantage when reasoning about logical situations that are consistent with real-world relationships or rules. In the simplest Natural Language Inference task, humans and all models show high accuracy and relatively minor effects of content. When judging the validity of syllogisms, both humans and models show more moderate accuracy, and significant advantages when content supports the logical inference. Finally, on the Wason selection task, humans and models show even lower accuracy, and again substantial content effects. We describe each task, and the corresponding results and analyses, in more detail below.
The relatively simple logical reasoning involved in this task means that both humans and models exhibit high performance, and correspondingly show relatively little effect of content on their reasoning (Fig. 3). Specifically, we do not detect a statistically-significant effect of content on accuracy in humans or any of the language models in mixed-effects logistic regressions controlling for the random effect of items (or tests where regressions did not converge due to ceiling effects; all or , all ; see Appx. C.1 for full results). However, we do find a statistically significant relationship between human and model accuracy at the item level (, ; Appx. 27) — even when controlling for condition. Furthermore, as we discuss below, further investigation into the model confidence shows evidence of content effects on this task as well.
Syllogism validity judgements are significantly more challenging than the NLI task above; correspondingly, we find lower accuracy in both humans and language models. Nevertheless, humans and most language models are sensitive to the logical structure of the task. However, we find that both humans and language models are strongly affected by the content of the syllogisms (Fig. 4), as in the past literature on syllogistic belief bias in humans (21). Specifically, if the semantic content supports the logical inference — that is, if the conclusion is believable and the argument is valid, or if the conclusion is unbelievable and the argument is invalid — both humans and all language models tend to answer more accurately (all or , all ; see Appx. C.2 for full results).
Two simple effects contribute to this overall content effect — that belief-consistent conclusions are judged as logically valid and that belief-inconsistent conclusions are judged as logically invalid. As in the past literature, we find that the dominant effect is that humans and models tend to say an argument is valid if the conclusion is belief-consistent, regardless of the actual logical validity. If the conclusion is belief-violating, humans and models do tend to say it is invalid more frequently, but most humans and models are more sensitive to actual logical validity in this case. Specifically, we observe a significant interaction between the content effect and believability in humans, PaLM 2-L, Flan-PaLM 2, and GPT-3.5 (all or , all ); but do not observe a significant interaction in Chinchilla or PaLM 2-M (both , ). Both humans and models appear to show a slight bias towards saying syllogisms with nonsense words are valid, but again with some sensitivity to the actual logical structure.
Furthermore, even when controlling for condition, we observe a significant correlation between item-level accuracy in humans and language models (, ), suggesting shared patterns in the use of lower-level details of the logic or content.
As in the prior human literature, we found that the Wason task was relatively challenging for humans, as well as for language models (Fig. 5). Nevertheless, we observed significant content advantages for the Realistic tasks in humans, and in Chinchilla, PaLM 2-L, and GPT-3.5 (all , all ; Appx. C.3). We only observed marginally significant advantages of realistic rules in PaLM 2-M and Flan-PaLM 2 (both , both ), due to stronger item-level effects in these models (though the item-level variance does not seem particularly unusual; see Appx. B.7.3 for further analysis). Intriguingly, some language models also show better performance at the versions of the tasks with Nonsense nouns compared to the Arbitrary ones, though generally Realistic rules are still easier. We also consider several variations on these rules in Appx. 22.
Our human participants struggled with this task, as in prior research, and did not achieve significantly higher-than-chance performance overall — although their behavior is not random, as we discuss below, where we analyze answer choices in more detail. However, spending longer on logical tasks can improve performance (59, 60), and thus many studies split analyses by response time to isolate participants who spend longer, and therefore show better performance (61, 62). Indeed, we found that human accuracy was significantly associated with response time (, ; Appx. C.3.1). We depict this relationship in Fig. 6. To visualize the the performance of discrete subjects in our Figures 2(c) and 5, we split subjects into ‘slow’ and ‘fast’ groups. The distribution of times taken by subjects is quite skewed, with a long tail. We separate out the top 15% of subjects that take the longest, who spent more than 80 seconds on the problem, as the slow group. These subjects showed above chance performance in the Realistic condition, but still performed near chance in the other conditions. We also dig further into the predictive power of human response times in the other tasks in the following sections (and Appx. B.6.1).
We collected the data for the Wason task in two different experiments; after observing the lower performance in the first sample, we collected a second sample where we offered a performance bonus for this task. We did not observe significant differences in overall performance or content effects between these subsets, so we collapse across them in the main analyses; however, we present results for each experiment and some additional analyses in Appx. B.5.
Language model behavior is frequently sensitive to details of the evaluation. Thus, we performed several experiments to confirm that our results were robust to details of the methods used. We present these results in full in Appx. B.2, but we outline the key experiments here. First, we show that removing the pre-question instructions does not substantially alter the overall results (Appx.B.2.1). Next, we show that our use of the DC-PMI correction for scoring is not the primary driver of content effects (Appx. B.2.2). On the syllogisms tasks, raw likelihood scoring with the instruction prompt yields strong answer biases — several models say every argument is valid irrespective of actual logical validity or content. However, the models that don’t uniformly say valid show content effects as expected. Furthermore, if the instructions before the question are removed, raw likelihood scoring results in less validity bias, and again strong content effects. For the Wason task, raw likelihood scoring actually improves the accuracy of some models; however, again the content effects are as found with the DC-PMI scoring. Thus, although overall model accuracy and response biases change with uncorrected likelihood scoring, the content effects are similar. Finally, we consider few-shot evaluation, and show that giving few-shot examples yields some mild improvements in accuracy (with greater improvement in the simpler tasks), but does not eliminate the content effects (Appx. B.2.4). Together, these results suggest that our findings are not strongly driven by idiosyncratic details of our evaluation, and thus support the robustness of our findings.
While we generally find similar content effects across the various models we evaluate, there are some notable differences among them. First, across tasks the larger models tend to be more accurate overall (e.g., comparing the large vs. the medium variants of PaLM 2); however, this does not necessarily mean they show weaker content effects. While it might be expected that instruction-tuning would affect performance, the instruction-tuned models (Flan-PaLM 2 and GPT-3.5-turbo-instruct) do not show consistent differences in overall accuracy or content effects across tasks compared to the base language models—in particular, Flan-PaLM 2 performs quite similarly to PaLM 2-L overall. (However, there are some more notable differences in the distributions of log-probabilities the instruction-tuned models produce; Appx. B.8.)
On the syllogisms task in particular, there are some noticeable difference among the models. GPT-3.5, and the larger PaLM 2 models, have quite high sensitivity for identifying valid arguments (they generally correctly identify valid arguments) but relatively less specificity (they also consider several invalid conclusions valid). By contrast, PaLM 2-M and Chinchilla models answer more based on content rather than logical validity i.e. regularly judging consistent conclusions as more valid than violating ones, irrespective of their logical validity. The sensitivity to logical structure in the nonsense condition also varies across models – the PaLM models are fairly sensitive, while GPT 3.5 and Chinchilla both having a strong bias toward answering valid to all nonsense propositions irrespective of actual logical validity.
On the Wason task, the main difference of interest is that the PaLM 2 family of models show generally greater accuracy on the Nonsense problems than the other models do, comparable to their performance on the Realistic condition in some cases.
Model confidence is related to content, correctness, and human response times
Language models do not produce a single answer; rather, they produce a probability distribution over the possible answers. This distribution can provide further insight into their processing. For example, the probability assigned to the top answers, relative to the others, can be interpreted as a kind of confidence measure. By this measure, language models are often somewhat calibrated, in the sense that the probability they assign to the top answer approximates the probability that their top answer is correct (e.g. 63). Furthermore, human Response Times (RTs) relate to many similar variables, such as confidence, surprisal, or task difficulty; thus, many prior works have related language model confidence to human response or reaction times for linguistic stimuli (e.g. 64, 65). In this section, we correspondingly analyze how the language model confidence relates to the task content and logic, the correctness of answers, and the human response times.
We summarize these results in Fig. 7. We measure model confidence as the difference in prior-corrected log-probability between the top answer and the second highest—thus, if the model is almost undecided between several answers, this confidence measure will be low, while if the model is placing almost all its probability mass on a single answer, the confidence measure will be high. In mixed-effects regressions predicting model confidence from task variables and average human RTs on the same problem, we find a variety of interesting effects. First, language models tend to be more confident on correct answers (that is, they are somewhat calibrated). Task variables also affect confidence; models are generally less confident when the conclusion violates beliefs, and more confident for the realistic rules on the Wason task.Furthermore, even when controlling for task variables and accuracy, there is a statistically-significant negative association with human response times on the NLI and syllogisms tasks (respectively , ; and , ; Appx. C.4)—that is, models tend to show more confidence on problems where humans likewise respond more rapidly. We visualize this relationship in Fig. 8.
Analyzing components of the Wason responses
Because each answer to the Wason problems involves selecting a pair of cards, we further analyzed the individual cards chosen. The card options presented each problem are designed so that two cards respectively match and violate the antecedent, and similarly for the consequent. The correct answer is to choose one card for the antecedent and one for the consequent; more precisely, the card for which the antecedent is true (AT), and the card for which the consequent is false (CF). In Fig. 9 we examine human and model choices; we quantitatively analyze these choices using a multinomial logistic regression model in Appx. C.3.2.
Even in conditions when performance is close to chance, behavior is generally not random. As in prior work, humans do not consistently choose the correct answer (AT, CF). Instead, humans tend to exhibit the matching bias; that is, they tend to choose each of the two cards that match each component of the rule (AT, CT). However, in the Realistic condition, slow humans answer correctly somewhat more frequently. Humans also exhibit errors besides the matching bias; including an increased rate of choosing the two cards corresponding to a single component of the rule — either both antecedent cards, or both consequent cards. Language models tend to give more correct responses than humans, and to show facilitation in the realistic rules compared to arbitrary ones. Relative to humans, language models show fewer matching errors, fewer errors of choosing two cards from the same rule component, but more errors of choosing the antecedent false options. These differences in error patterns may indicate differences between the response processes engaged by the models and humans. (Note, however, that while the models accuracies do not change too substantially with alternate scoring methods, the particular errors the models make are somewhat sensitive to scoring method — without the DC-PMI correction the model errors more closely approximate the human ones in some cases; Appx. B.2.3.)
Discussion
Humans are imperfect reasoners. We reason most effectively about entities and situations that are consistent with our understanding of the world. Even in these familiar cases, we often make mistakes. Our experiments show that language models mirror these patterns of behavior. Language models likewise perform imperfectly on logical reasoning tasks, but this performance depends on content and context. Most notably, such models often fail in situations where humans fail — when stimuli become too abstract or conflict with prior expectations about the world.
Beyond these simple parallels in accuracy across different conditions and items, we also observed more subtle parallels in language model confidence. The model’s confidence tends to be higher for correct answers, and for cases where prior expectations about the content are consistent with the logical structure. Even when controlling for these effects, model confidence is related to human response times. Thus, language models reflect human content effects on reasoning at multiple levels. Furthermore, these core results are generally robust across different language models with different training and tuning paradigms, different prompts, etc., suggesting that they are a fairly general phenomenon of predictive models that learn from human-generated text.
Since Brown et al. (7) showed that large language models could perform moderately well on some reasoning tasks, there has been a growing interest in language model reasoning (66). Typical methods focus on prompting for sequential reasoning (9, 67, 10), altering task framing (68, 69) or iteratively sampling answers (70).
In response, some researchers have questioned whether these language model abilities qualify as “reasoning”. The fact that language models sometimes rely on “simple heuristics” (15), or reason more accurately about frequently-occurring numbers (14), have been cited to “rais[e] questions on the extent to which these models are actually reasoning” (ibid, emphasis ours). The implicit assumption in these critiques is that reasoning should be a purely algebraic, syntactic computations over symbols from which “all meaning had been purged” (2; cf. 34). In this work, we emphasize how both humans and language models rely on content when answering reasoning problems — using simple heuristics in some contexts, and answering more accurately about frequently-occurring situations (71, 28). Thus, abstract reasoning may be a graded, content-sensitive capacity in both humans and models.
The idea that humans possess dual reasoning systems — an implicit, intuitive system “system 1”, and an explicit reasoning “system 2’ — was motivated in large part by belief bias and Wason task effects (72, 73, 74). The dual system idea has more recently become popular (75, 76), including in machine learning (e.g. 77). It is often claimed that current ML (including large language models) behave like system 1, and that we need to augment this with a classically-symbolic process to get system 2 behaviour (e.g. 78). These calls to action usually advocate for an explicit duality; with a neural network based system providing the system 1 and a system with more explicit symbolic or otherwise structured system being the system 2.
Our results show that a unitary system — a large transformer language model — can mirror this dual behavior in humans, demonstrating both biased and consistent reasoning depending on the context and task difficulty. In the NLI tasks, a few examples takes Chinchilla from highly content-biased performance to near ceiling performance, and even a simple instructional prompt can substantially reduce bias. These findings integrate with prior works showing that language models can be prompted to exhibit sequential reasoning, and thereby improve their performance in domains like mathematics (9, 67, 10).
These observations suggest the possibility that the unitary language model may have implicitly learned a context-dependent control mechanism that arbitrates between conflicting responses (such as more intuitive answers vs. logically correct ones). This perspective suggests several possible directions for future research. First, it would be interesting to seek out mechanistic evidence of such as conflict-arbitration process within language models. Furthermore, it suggests that augmenting language models with a second system might not be necessary to achieve relatively reliable performance. Instead, it might be sufficient to further develop the control mechanisms within these models by altering their context and training, as we discuss below.
From a human cognitive neuroscience perspective, these issues are more complex. The idea of context-dependent arbitration between conflicting responses has been influential in the literature on human cognitive control (79, 80), and has been implicated in humans reasoning successfully in tasks that require following novel, arbitrary reasoning procedures or over-riding pre-existing response tendencies (81, 82). However, these control processes are generally believed to principally reside in frontal regions outside the language areas, or in a network of control-related brain areas that interface with the language regions and other domain-specific brain areas but is not housed therein (81). Thus, understanding the full detail of human cognition in such language-based logical tasks may require incorporating a control network into the architecture more explicitly. Nevertheless, our results with language models suggest that this controller could be more intertwined with the statistical inference system than it would be in a classic dual-systems model; moreover, that the controller does not need to be implemented as a classical symbol system to achieve human-competitive logical reasoning performance.
Deep learning models are increasingly used as models of neural processing in biological systems (e.g. 83, 84), as they often develop similar patterns of representation. These findings have led to proposals that deep learning models capture mechanistic details of neural processing at an appropriate level of description (85, 86), despite the fact that aspects of their information processing clearly differ from biological systems. More recently, large language models have been similarly shown to accurately predict neural representations in the human language system — large language models “predict nearly 100% of the explainable variance in neural responses to sentences” (87; see also 88, 89). Language models also predict low-level behavioral phenomena; e.g. surprisal predicts reading time (64, 90). In the context of these works, our observation of behavioral similarities in reasoning patterns between humans and language models raise important questions about possible similarities of the underlying reasoning processes between humans and language models, and the extent of overlap between neural mechanisms for language and reasoning in humans. This is particularly exciting because neural models for these phenomena in humans are currently lacking or incomplete. Indeed, even prior high-level explanations of these phenomena have often focused on only a single task, such as explaining only the Wason task content effects with appeals to evolved social-reasoning mechanisms (25). Our results suggest that there could be a more general explanation.
Various accounts of human cognitive biases frame them as ‘normative’ according to some objective. Some explain biases as the application of processes — such as information gathering or pragmatics — that are broadly rational under a different model of the world (e.g. 74, 52). Others interpret them as a rational adaptation to reasoning under constraints such as limited memory or time (e.g. 91, 92, 93) — where content effects actually support fast and effective reasoning in commonly encountered tasks (71, 28). Our results show that content effects can emerge from simply training a large transformer to imitate language produced by human culture, without explicitly incorporating any human-specific internal mechanisms.
This observation suggests two possible origins for these content effects. First, the content effects could be directly learned from the humans that generated the data used to train the language models. Under this hypothesis, poor logical inferences about nonsense or belief-violating premises come from copying the incorrect inferences made by humans about these premises. Since humans also learn substantially from other humans and the cultures in which we are immersed, it is plausible that both humans and language models could acquire some of these reasoning patterns by imitation.
The other possibility is that (like humans) the model’s exposure to the world reflects semantic truths and beliefs and that language models and humans both converge on these content biases that reflect this semantic content for more task-oriented reasons: because it helps humans to draw more accurate inferences in the situations they encounter (which are mostly familiar and believable), and helps language models to more accurately predict the (mostly believable) text that they encounter. In either case, humans and models acquire surprisingly similar patterns of behavior, from seemingly very different architectures, experiences, and training objectives. A promising direction for future enquiry would be to causally manipulate features of language model’s training objective and experience, to explore which features contribute to the emergence of content biases in language models. These investigations could offer insights into the origins of human patterns of reasoning, and into what data we should use to train language models.
The language model response patterns do not perfectly match all aspects of the human data. For example, on the Wason task several models outperform humans on the Nonsense condition, and the error patterns on the Wason tasks are somewhat different than those observed in humans (although human error patterns also vary across populations; 57, 58). Similarly, not all models show the significant interaction between believability and validity on the syllogism tasks that humans do (20), although it is present in most models (and the human interaction similarly may not appear in all cases; 53). Various factors could contribute to differences between model and human behaviors.
First, while we attempted to align our evaluation of humans and models as closely as possible (cf. 69), it is difficult to do so perfectly. In some cases, such as the Wason task, differences in the form of the answer are unavoidable — humans had to select answers individually by clicking on cards to select them, and then clicking continue, while models had to jointly output both answers in text, without a chance to revise their answer before continuing. Moreover, it is difficult to know how to prompt a language model in order to evaluate a particular task. Language model training blends many tasks into a homogeneous soup, which makes controlling the model difficult. For example, presenting task instructions might not actually lead to better performance (cf. 94). Similarly, presenting negative examples can help humans learn, but is generally detrimental to model performance (e.g. 95) — presumably because the model infers that the task is to sometimes output wrong answers, while humans might understand the communicative intent behind the use of negative examples. Thus, while we tried to match instructions between humans and models, it is possible that idiosyncratic details of our task framing may have caused the model to infer the task incorrectly. To minimize this risk, we tried various different prompting strategies, and where we varied these details we generally observed similar overall effects. Nevertheless, it is possible that some aspect of the problem instructions or framing contributes to the response patterns.
More fundamentally, language models do not directly experience the situations to which language refers (96); grounded experience (for instance the capacity to simulate the physical turning of cards on a table) presumably underpins some human beliefs and reasoning. Furthermore, humans sometimes use physical or motor processes such as gesture to support logical reasoning (97, 98). Finally, language models experience language passively, while humans experience language as an active, conventional system for social communication (e.g. 99); active participation may be key to understanding meaning as humans do (36, 100). Some differences between language models and humans may therefore stem from differences between the rich, grounded, interactive experience of humans and the impoverished experience of the models.
If language models exhibit some of the same reasoning biases as humans could some of the factors that reduce content dependency in human reasoning be applied to make these models less content-dependent? In humans, formal education is associated with an improved ability to reason logically and consistently (101, 102, 103, 57, 58, 104). However, causal evidence is scarce, because years of education are difficult to experimentally manipulate; thus the association may be partly due to selection effects, e.g. continuing in formal education might be more likely in individuals with stronger prior abilities. Nevertheless, the association with formal education raises an intriguing question: could language models learn to reason more reliably with targeted formal education?
Several recent results suggest that this may indeed be a promising direction. Pretraining on synthetic logical reasoning tasks can improve model performance on reasoning and mathematics problems (105, 106). In some cases language models can either be prompted or can learn to verify, correct, or debias their own outputs (107, 108, 109, 63). Finally, language model reasoning can be bootstrapped through iterated fine-tuning on successful instances (110). These results suggest the possibility that a model trained with instructions to perform logical reasoning, and to check and correct the results of its work, might move closer to the logical reasoning capabilities of formally-educated humans. Perhaps logical reasoning is a graded competency that is supported by a range of different environmental and educational factors (36, 111), rather than a core ability that must be built in to an intelligent system.
In addition to the limitations noted above — such as the challenges of perfectly aligning comparisons between humans and language models — there are several other limitations to our work. First, our human participants exhibited relatively low performance on the Wason task. However, as noted above, there are well-known individual differences in these effects that are associated with factors like depth of mathematical education. We were unfortunately unable to examine these effects in our data, but in future work it would be interesting to explicitly explore how educational factors affect performance on the more challenging Wason conditions, as well as more general patterns like the relationship between model confidence and human response time. Furthermore while our experiments suggest that content effects in reasoning can emerge from predictive learning on naturalistic data, they do not ascertain precisely which aspects of the large language model training datasets contribute to this learning. Other research has used controlled training data distributions to systematically investigate the origin of language model capabilities (112, 113); it would be an interesting future direction to apply analogous methods to investigate the origin of content effects.
Materials and Methods
While many of these tasks have been extensively studied in cognitive science, the stimuli used in cognitive experiments are often online in articles and course materials, and thus may be present in the training data of large language models, which could compromise results (e.g. 114, 115). To reduce these concerns, we generate new datasets, by following the design approaches used in prior work. We briefly outline this process here; see Appx. A.1 for full details.
For each of the three tasks above, we generate multiple versions of the task stimuli. Throughout, the logical structure of the stimuli remains fixed, we simply manipulate the entities over which this logic operates (Fig. 1). We generate propositions that are: Consistent with human beliefs and knowledge (e.g. ants are smaller than whales). Violate beliefs by inverting the consistent statements (e.g. whales are smaller than ants). Nonsense tasks about which the model should not have strong beliefs, by swapping the entities out for nonsense words (e.g. kleegs are smaller than feps).
For the Wason tasks, we slightly alter our approach to fit the different character of the tasks. We generate questions with: Realistic rules involving plausible relationships (e.g. “if the passengers are traveling outside the US, then they must have shown a passport”). Arbitrary rules (e.g. “if the cards have a plural word, then they have a positive emotion”). Nonsense rules relating nonsense words (“if the cards have more bem, then they have less stope”). Note that for the Wason task, this change alters the kinds of inferences that need to be made; while for the basic Wason task, matching each card to the antecedent or consequent is nontrivial (e.g. realizing that “shoes” is a plural word, not singular), it is difficult to match these inferences with Nonsense words that have no prior associations. As shown in the examples, we use a format where the cards either have more or less of a nonsense attribute, which makes the inferences perhaps more direct than other conditions (although models perform similarly on the basic inferences across conditions; Appx. 21).
In Appx. B.1 we validate the semantic content of our datasets, by showing that participants find the propositions and rules from our Consistent and Realistic stimuli much more plausible than those from other conditions.
We attempted to create these datasets in a way that could be presented to the humans and language models in precisely the same manner (for example, prefacing the problems with the same instructions for both the humans and the models).N.B. this required adapting some of the problem formats and prompts compared to an earlier preprint of this paper that did not evaluate humans; see Appx. A.1.4.
We evaluate several different families of language models. First, we evaluate several base LMs that are trained only on language modeling: including Chinchilla (45) a large model (with 70 billion parameters) trained on causal language modeling, and PaLM 2-M and -L (47), which are trained on a mixture of language modeling and infilling objectives (116). We also evaluate two instruction-tuned models: Flan-PaLM 2 (an instruction-tuned version of Palm 2-L), and GPT-3.5-turbo-instruct (46), which we generally refer to as GPT-3.5 for brevity.We fortuitously performed this evaluation during the short window of time in which scoring was available on GPT-3.5-turbo-instruct. We observe broadly similar content effects across all types of models, suggesting that these effects are not too strongly affected by the particular training objective, or by standard instruction-tuning.
For each task, we present the model with brief instructions that approximate the relevant portions of the human instructions. We then present the question, which ends with “Answer:” and assess the model by evaluating the likelihood of continuing this prompt with each of a set of possible answers. We apply the DC-PMI correction proposed by Holtzman et al. (117) — i.e., we measure the change in likelihood of each answer in the context of the question relative to a baseline context, and choosing the answer that has the largest increase in likelihood in context. This scoring approach is intended to reduce the possibility that the model would simply phrase the answer differently than the available choices; for example, answering “this is not a valid argument” rather than “this argument is invalid”. This approach can also be interpreted as correcting for the prior over utterances. For the NLI task, however, the direct answer format means that the DC-PMI correction would therefore control for the very bias we are trying to measure. Thus, for the NLI task we simply choose the answer that receives the maximum likelihood among the set of possible answers. We also report syllogism and Wason results with maximum likelihood scoring in Appx. B.2.2; while overall accuracy changes (usually decreases, but with some exceptions), the direction of content effects is generally preserved under alternative scoring methods.
The human experiments were conducted in 2023 using an online crowd-sourcing platform, and recruiting only participants from the UK who spoke English as a first language, and who had over a 95% approval rate. We did not further restrict participation. We offered pay of £2.50 for our task. Our intent was to pay at a rate exceeding £15/h, and we exceeded this target, as most participants completed the task in less than 10 minutes.
The human participants were first presented with a consent form detailing the experiment and their ability to withdraw at any time. If they consented to participate, they then proceeded to an instructions page. After the instructions they were presented with one question from each of our three tasks, one at a time. Each participant saw the tasks in a randomized order, and with randomized conditions. Subsequently, the participants were presented with three rating questions, rating the believability of a rule from one of the Wason tasks (not the one they had completed), and their degree of agreement with a concluding proposition from a syllogism task, and a concluding proposition from the NLI task. In each case, the ratings were provided on a continuous scale from 0 to 100 (with 50% indicated as neither agree nor disagree). On each question and rating, the participants had to answer in less than a time limit of 5 minutes (to ensure they were not abandoning the task entirely). This time limit was reset on the next question. See Appx. A.3 for further detail on the experimental methods.
We first collected a dataset of responses from 625 participants. After observing the low accuracy in the Wason tasks, we collected an additional dataset from 360 participants in which we offered an additional performance bonus of £0.50 for answering the Wason question correctly, to motivate subjects. In this replication, we collected data only on the Realistic and Abstract Wason conditions. In our main analyses, we collapse across these two subsets, but we present the results for each experiment separately in Appx. 23.
Due to infrastructure restrictions in the framework used to create the human tasks, we assigned participants to conditions and items randomly rather than with precise balancing. Furthermore, a few participants timed out on some questions, and there were a handful of instances of data not saving properly due to server issues. Thus, the exact number of participants for which we have data varies slightly from task to task and item to item.
Our main analyses are quantified with mixed-effects logistic regression models that include task condition variables as predictors, and control for random effects of items, and, where applicable, models. The key results of these models are reported in the main text. The full model specifications and full results are provided in Appx. C.
References
Appendix A Supplemental methods
As noted in the main text, we generated new datasets for each task to avoid problems with training data contamination. In this section we present further details of dataset generation.
In the absence of existing cognitive literature on generating belief-aligned stimuli for this task, we used a larger language model (Gopher, 280B parameters, from 13) to generate 100 comparison statements automatically, by prompting it with 6 comparisons that are true in the real world. The exact prompt used was:
The following are 100 examples of comparisons:1. mountains are bigger than hills2. adults are bigger than children3. grandparents are older than babies4. volcanoes are more dangerous than cities5. cats are softer than lizards
We prompted the LLM multiple times, until we had generated 100 comparisons that fulfilled the desired criteria. The prompt completions were generated using nucleus sampling (118) with a probability mass of 0.8 and a temperature of 1. We filtered out comparisons that were not of the form “[entity] is/are [comparison] than [other entity]”. We then filtered these comparisons manually to remove false and subjective ones, so the comparisons all respect real-world facts. An example of the generated comparisons includes “puddles are smaller than seas”.
We generated a natural inference task derived from these comparison sentences as follows. We began with the consistent version, by taking the the raw output from the LM, “puddles are smaller than seas” as the hypothesis and formulating a premise “seas are bigger than puddles” such that the generated hypothesis is logically valid. We then combine the premise and hypothesis into a prompt and continuations. For example:
If seas are bigger than puddles, then puddles areA. smaller than seasB. bigger than seaswhere the logically correct (A) response matches real-world beliefs (that ‘puddles are smaller than seas’). Similarly, we can also generate a violate version of the task where the logical response violates these beliefs. For example,
If seas are smaller than puddles, then puddles areA. smaller than seasB. bigger than seashere the correct answer, (B), violates the LM’s prior beliefs. Finally, to generate a nonsense version of the task, we simply replace the nouns (‘seas’ and ‘puddles’) with nonsense words. For example:
If vuffs are smaller than feps, then feps areA. smaller than vuffsB. bigger than vuffsHere the logical conclusion is B. For each of these task variations, we evaluate the log probability the language model places on the two options and choose higher likelihood one as its prediction.
A.1.2 Syllogisms data generation
We generated a new set of problems for syllogistic reasoning. Following the approach of Evans et al. (20), in which the syllogisms were written based on the researchers’ intuitions of believability, we hand-authored these problems based on beliefs that seemed plausible to the authors. See Fig. 10(b) for an example problem. We built the dataset from clusters of 4 arguments that use the same three entities, in a combination of valid/invalid, and belief-consistent/violate. For example, in Fig. 11 we present a full cluster of arguments about reptiles, animals, and flowers.
By creating the arguments in this way, we ensure that the low-level properties (such as the particular entities referred to in an argument) are approximately balanced across the relevant conditions. In total there are twelve clusters. We avoided using the particular negative form (“some X are not Y”) to avoid substantial negation, which complicates behavior both for language models and humans (cf. 119, 120). We then sampled an identical set of nonsense arguments by simply replacing the entities in realistic arguments with nonsense words.
We present the arguments to the model, and give a forced choice between “The argument is valid.” or “The argument is invalid.” Where example shots are used, they are sampled from distinct clusters, and are separated by a blank line. We also tried some minor variations in preliminary experiments (such as changing the prompt or prefixing the conclusion with “Therefore:” or omitting the prefix before the conclusion), but observed qualitatively similar results so we omit them here.
A.1.3 Wason data generation
As above, we generated a new dataset of Wason problems to avoid potential for dataset contamination (see Fig. 10(c) for an example). The final response in a Wason task does not involve a declarative statement (unlike completing a comparison as in NLI), so answers do not directly ‘violate’ beliefs. Rather, in the cognitive science literature, the key factor affecting human performance is whether the entities are ‘realistic’ and follow ‘realistic’ rules (such as people following social norms) or consist of arbitrary relationships between abstract entities such as letters and numbers. We therefore study the effect of realistic and arbitrary scenarios in the language models.
We created 12 realistic rules and 12 arbitrary rules. Each rule appears with four instances, respectively matching and violating the antecedent and consequent. Each realistic rule is augmented with one sentence of context for the rule, and the cards are explained to represent the entities in the context. The model is presented with the context, the rule, and is asked which of the following instances it needs to flip over, then the instances. The model is then given a forced choice between sentences of the form “You need to flip over the ones showing “X” and “Y”.” for all subsets of two items from the instances. There are two choices offered for each pair, in both of the possible orders, to eliminate possible biases if the model prefers one ordering or another. (Recall that the model scores each answer independently; it does not see all answers at once.)
See Figs. 13 and 14 for the realistic and arbitrary rules and instances used — but note that problems were presented to the model with more context and structure, see Fig. 10(c) for an example. We demonstrate in Appx. B.3 that the difficulty of basic inferences about the propositions involved in each rule type is similar across conditions.
We also created 12 rules using nonsense words. Incorporating nonsense words is less straightforward in the Wason case than in the other tasks, as the model needs to be able to reason about whether instances match the antecedent and consequent of the rule. We therefore use nonsense rules of the form “If the cards have less gluff, then they have more caft” with instances being more/less gluff/caft. The more/less framing makes the instances roughly the same length regardless of rule type, and avoids using negation which might confound results (119).
Finally, we created two types of control rules based on the realistic rules, which we present here. First, we created shuffled realistic rules by combining the antecedents and consequents of different realistic rules, while ensuring that there is no obvious rationale for the rule. For example, one shuffled-realistic rule is “If they are doctors, then they must have a parachute.”
We then created violate-realistic rules by taking each realistic rule and reversing its consequent. For example, the realistic rule “If the clients are skydiving, then they must have a parachute” is transformed to the violate rule “If the clients are skydiving, then they must have a wetsuit”, but “parachute” is still included among the cards. The violate condition is designed to make the rule especially implausible in context of the examples (viz. requiring the item that is not a parachute to skydive), while the rule in the shuffled condition is somewhat more arbitrary/belief neutral.
To rule out a possible specific effect of cards (which were used in the original tasks) we also sampled versions of each problem with sheets of paper or coins, but results are similar so we collapse across these conditions in the main analyses.
A.1.4 Differences from an earlier preprint of this paper:
Readers of an earlier preprint of this paper (https://arxiv.org/abs/2207.07051v1) may notice some differences in task format and performance of Chinchilla, especially on the NLI task. These differences are due to our attempts to adapt the tasks in order to present them to human participants. In order to align comparisons between humans and the language models (cf. 69), we then ported the human-oriented changes back into the format used for language model evaluation.
For example, in the original paper we did not show the model the two possible choices for the NLI task; we simply evaluated the model’s likelihood of each continuation. However, because we presented the tasks multiple choice to the humans in multiple choice format, we showed them the two possible answers. Thus, in the current version of the paper we also included the two answer choices in the prompt when evaluating the language models, followed by “Answer:”, and only then evaluate the model (see Fig. 10(a)). Likewise, in the original version of the paper we did not provide instructions before the tasks; in this version we attempted to match the relevant portions of the human instructions.
These changes mean that the results in the current version of the paper cannot be directly compared to the results in the earlier version.
A.2 Evaluation
DC-PMI correction: We use the DC-PMI correction (117) for the syllogisms and Wason tasks; i.e., we choose an answer from the set of possible answers () as follows:
Where the baseline prompt is the task instruction prompt, followed by “Answer:” and denotes the model’s evaluated likelihood of continuation after prompt .
Instruction prompt: We prefixed each question with a two-part instruction prompt that attempted to match the human generic and task-specific instructions (see below). We began each of these prompts with the performance relevant generic instructions that preceded our human experiment:
In this task, you will have to answer a series of questions. You will have to choose the best answer to complete a sentence, paragraph, or question. Please answer them to the best of your ability.\n\n
After two linebreaks, a task-specific instruction was appended:
Please choose the best completion for the following sentence:\n
Please assume that the first two sentences in the argument are true. Determine whether the argument is valid, that is, whether the conclusion follows from the first two sentences:\n’
Please answer the following question carefully:\n
Finally, the question was appended to this prompt.
A.3 Human experiments
The exact text seen by the participants before each question was as follows:
Appendix B Supplemental analyses
In order to assess the validity of our new datasets, we collected believability ratings from each subject, after they had completed the three tasks tasks, on one stimulus from each task type (not the version they had seen). Specifically, we asked the participants how believable a Wason rule was, and how much they agreed or disagreed with a proposition. In Fig. 15 we show that participants found Consistent and Realistic stimuli much more believable than those in other conditions.
B.2 Robustnesss of the main language model results to raw-likelihood scoring and few-shot prompting
In this section, we show that the content effects we observe are robust to various manipulations of the evaluation context.
In the main text experiments, we provided models with an instruction prompt that roughly matched the human instructions (cf. 69). However, it is unclear how substantial a role this prompt played in performance, and human-likeness of the content effects. In Fig. 16 we show performance of a subset of the models when removing this instruction prompt; in most cases, results are similar, with a few notable exceptions. In particular Chinchilla shows much stronger content effects on the NLI tasks without instructions.
B.2.2 Using raw likelihoods rather than Domain-Conditional PMI on the Syllogisms and Wason tasks
In the main results for the Syllogisms and Wason tasks, we scored the model using the Domain-Conditional PMI (117). However, it is also common to score language models using raw likelihood comparisons. Would we observe the same content effects in that case?
In Fig. 17 we show the results of raw-likelihood scoring. On the syllogisms tasks, this scoring method results in substantially more answer bias — several of the models say valid in response to every problem, regardless of the content or logical structure. Thus, performance is much worse overall. However, for the models that do show any variability with content, the content effects are broadly similar to those observed in the main text: the models are more likely to say an argument is valid if the conclusion is belief-consistent than if the conclusion violates beliefs. Furthermore, if the instruction prompt is removed, the bias is substantially reduced, and stronger content effects are revealed.
In the Wason tasks, the effects on accuracy are more complex. While some models perform worse without the prior correction (e.g. Chinchilla), others perform much better. In particular, PaLM-2 L achieves over 75% performance in every condition (including Arbitrary and Nonsense). However, all models that perform above chance show the same content effects observed in the main text: better performance on Realistic than Arbitrary rules. (In Appx. B.2.3 we also explore the effect of scoring with raw likelihoods on the individual card choices on the Wason task.)
B.2.3 Effects of scoring method and answer order on the Wason answer choices
In Fig. 18 we show the effect of scoring method (DC-PMI vs. raw likelihoods) and the order in which the cards were presented (antecedent cards first or consequent cards first) on the models’ answer choices on the Wason task. Scoring method does affect the error distribution fairly substantially, even where accuracy is similar; answer order has smaller effects.
B.2.4 Few-shot prompting of Chinchilla
In all the main text experiments, we evaluated the models zero-shot, with only instructions. However, language model performance is generally improved by few-shot prompting (e.g. 7). We therefore evaluated whether few shot prompting with different kinds of prompt examples would alter the content effects we observed. (Note that, for computational reasons, we restrict these analyses to the Chinchilla model.) When we present a few-shot prompt of examples of the task to the model, the examples are presented with correct answers, and each example (as well as the final probe) is separated from the previous example by a single blank line.
In Fig. 19 we show 5-shot prompting results for Chinchilla on the Syllogisms tasks. Content effects are slightly weaker than without the examples, but remain robust.
In Fig. 20 we show 5-shot prompting results for Chinchilla on the Wason selection tasks. Content effects are exaggerated with the 5-shot prompts, because the model improves noticeably at Realistic rules, but improves less (if at all) on Arbitrary ones. We also see a noticeable effect of the type of examples used in the prompt, with Realistic examples offering optimal benefits.
B.3 The Wason rule propositions have similar difficulty across conditions
One possible confounding explanation for our Wason results would be that the base propositions that form the antecedents and consequents of the rules have different difficulty across conditions—this could potentially explain why the realistic rules and shuffled realistic rules are both easier than abstract or nonsesnse ones. To investigate this possibility, we tested the difficulty of identifying which of the options on the cards matched the corresponding proposition. Specifically, for the antecedent of the rule “if the workers work as a doctor then they must have received an MD” we prompted Chinchilla with a question like:
Which choice better matches "work as a doctor"?choice: surgeonchoice: janitorAnswer:And then gave a two-alternative forced choice between ‘surgeon’ and ‘janitor’. To avoid order biases, we repeated this process for both possible answer choice orderings in the prompt, and then aggregated likelihoods across these and chose the highest-likelihood answer.
By this metric, we find that there are no substantial differences in difficulty across the rule types (Fig. 21)—in fact, arbitrary rule premises are numerically slightly easier, though the differences are not significant. Thus, the effects we observed are not likely to be explained by the base difficulty of verifying the component propositions.
B.4 Additional recombined realistic conditions for the Wason tasks
The Wason task rules can be realistic or unrealistic in multiple ways. For example, the component propositions can be realistic even if the relationship between them is not. We therefore generate two variations on realistic rules: Shuffled realistic rules, which combine realistic components in nonsensical ways (e.g. “if the passengers are traveling outside the US, then they must have received an MD”). Violate realistic rules, which directly violate the expected relationship (e.g. “if the passengers are flying outside the US, then they must have shown a drivers license [not a passport]”).
We also evaluated models and humans on these rules. For shuffled rules, results are well above chance. Surprisingly, one family of models (PaLM 2) even perform better at shuffled realistic than realistic rules. For violate rules, by contrast, performance is generally close to chance. It appears that the model reasons more accurately about rules formed from realistic propositions, particularly if the relationships between propositions in the rule are also realistic, but even to some degree if they are shuffled in nonsensical ways that do not directly violate expectations. However, if the rules strongly violate beliefs, performance is low. Humans generally perform poorly on either rule variant.
B.5 Human performance on the Wason tasks, in our original sample & replication
As mentioned in the main text, after collecting our original sample on the Wason task, we recruited an additional set of participants to whom we offered a performance bonus on this task, in an attempt to increase performance. We present the results broken down by sample in Fig. 23. We performed mixed-effects logistic regressions (Table 1) to test for an improvement in performance in the sample with a performance bonus; this effect was marginally significant. However, performance remains low overall, and we do not observe a significant difference in the content effect.
B.6 Human response time distributions on the Wason tasks
In Fig. 24 we show the distribution of response times for humans in the Wason tasks. There is a mean difference in response times, with participants spending about 12 seconds longer on Realistic questions on average. This difference may be due to the time needed to read the extra sentences giving the realistic context, or to the participants engaging more deeply with the problems that seem more sensible. However, in Appx. C.3.1 we show that this difference alone does not explain the advantage of the Realistic conditions.
Given the strong effect of response time on Wason task performance, we also analyzed the effects on the NLI and Syllogism tasks (Figs. 25 & 25; Table 2). In these tasks we do not see clear effects, though there are hints of an interesting potential interaction in the syllogisms task.
B.7 Item-level effects
In this section, we perform item level analyses for each task.
First, for the NLI task, we plot the item-level correlations in accuracy in Fig. 27. Surprisingly (given the close-to-ceiling performance), we find that Human success rates are significantly predictive of LM success rates, even when controlling for condition (Table 3).
B.7.2 Syllogisms
For the Syllogisms task, we plot the item-level correlations in accuracy in Fig. 27. We again find a significant relationship between human success rates and language model success (, when controlling for task variables; Table 4).
B.7.3 Wason
For the Wason task, we plot the item-level correlations in accuracy in Fig. 27. Perhaps because human performance is low overall, we do not observe a significant relationship between human success rates and language model success (Table 5).
Due to the item-level effects observed in some of the main regressions, we also plot performance of each model or human group on each of the Wason rules in Fig. 30. Overall, the variability seems mostly as expected. However, there are some interesting patterns, including one arbitrary rule that most subject perform well on. That particular rule is:
The rule is that if the cards have a French word then they must have a positive number.chapeau / sombrero / 4 / -1
It is not particularly apparent to us why this rule might be easier.
B.8 Model answer log-probability distributions
In this section we plot the log-probability distributions of the models on the different tasks (Figs. 31, 32, 33). There are a variety of interesting effects of task variables, and some striking differences among the models.
For example, the instruction-tuned models (Flan-PaLM 2 and GPT-3.5) have numerically much greater magnitude log-probabilities to the answers, especially GPT-3.5. This may be an artifact of the tuning process. Furthermore, the larger models tend to show clearer separation between the chosen answer and the others (e.g., comparing PaLM 2-L to -M).
B.9 Chinchilla can identify the valid conclusion of a syllogism from among all possible conclusions with high accuracy
In Fig. 34 we show the accuracy of Chinchilla when choosing from among all possible predicates containing one of the quantifiers used and two of the entities appearing in the premises of the syllogism. The model exhibits high accuracy across conditions, and relatively little bias (though bias increases few shot). This observation is reminiscent of the finding of Trippas et al. (54) that humans exhibit less bias when making a forced choice among two possible arguments (one valid and one invalid) rather than deciding if a single syllogism is valid or invalid.
Note that in this case scoring with the Domain-Conditional PMI (117)—which we used for the main Syllogisms and Wason results—produces much lower accuracy than the raw likelihoods, and minor differences in bias. The patterns are qualitatively similar with or without the correction, but accuracy is lower without (around 35-40%) regardless of belief consistency.
Appendix C Statistical analyses
In this section, we provide the full results for all statistical analyses reported in the main text. We generally report results from mixed-effects logistic regressions, controlling for the random effects of the different stimuli used.Unless otherwise noted, we conservatively approximate the degrees of freedom for all -tests by treating all random effects as though they were fixed effects (i.e. by subtracting the number of levels of each random variable from the residual degrees of freedom), rather than using a variance-based approximation.
We report statistical analyses of content effects on the NLI tasks for humans and all models in Tables 6-11. We generally fit mixed effects logistic regressions, but the regressions for PaLM 2-L and Flan-PaLM 2 failed to converge due to ceiling effects. We therefore also report tests of the difference in correct responses across conditions. In all cases, we do not find a significant content effect on this simple task.
C.2 Syllogisms
We report mixed effects logistic regressions for humans and all models in Tables 12-17. We analyze these results using a variable which corresponds to the main content effect (logic_belief_consistent), which is 1 when the logical answer matches the believability of the conclusion — i.e. when the argument is valid and the conclusion is believable, or the argument is invalid and the conclusion is unbelievable — and 0 when there is a mismatch. This measure corresponds to the difference score reported in Fig. 2(b). We ran three nested models for humans and each language model — one regression only incorporating the content effect predictor (whether the logic matches the consistency), another adding consistency condition, and a third adding the interaction of the two.
response_correct ~ logic_belief_consistent + (1 | syllogism_name)response_correct ~ logic_belief_consistent + consistent_plottable_f + (1 | syllogism_name)response_correct ~ logic_belief_consistent * consistent_plottable_f + (1 | syllogism_name)For humans and each language model, we report the best-fitting regression, measured by the BIC (and omitting models which failed to converge). However, since the interaction effect is theoretically interesting (e.g. 53), and several of the interaction models fail to converge, we also report two-way tests of the interactions for each model. For PaLM 2-L all regressions failed to converge due to ceiling effects; thus we also report a test of the content effect for this model only. All models show a significant content effect; all except Chinchilla and PaLM 2-M show a significant interaction.
C.3 Wason
We report mixed-effects logistic regressions for humans (both all humans, and the fast and slow groups individually) and all models in Tables 18-25. We observe a significant effect of content in most cases. However, the fast humans alone do not show a significant content effect. Furthermore, the content effects in PaLM 2-M and Flan-PaLM 2 are only marginally significant, due to high item level variance.
In Appx. C.3.1 we further analyze the human data whil incorporating response time in the regression.
Here we present two regression analyses of the human results that incorporate the response time. In Table 26 we show a mixed-effects logistic regression controlling for log response time; the content effect remains significant. Thus, the content effects are not solely driven by the differences in response time noted above (Appx. 24).
However, it is also possible to conceive of the shift in response time as a part of the content effect. We can analyze the data this way by -scoring response time within each condition; thus, the effect of the mean difference in response time will be included in the condition predictor. We present these results in Table 27. Both content and -scored response time remain significant predictors of success.
C.3.2 Multinomial regression of the response patterns on the Wason tasks
In Table 28 we present the results of a multinomial logistic regression predicting which of the six possible subsets of answers the humans and language models chose on the Wason task. This regression quantitatively supports the claim that the behavior is nonrandom, and more generally quantifies the qualitative observations of response patterns made in the main text.
C.4 Response time and model log-probability differences
In this section we present the mixed-effects linear regressions comparing human response times and model log-probabilities on the NLI and syllogisms tasks, in Tables 29 and 30, respectively. In order to make these comparisons, we breakdown each problem into cases where both humans and models got it correct, and cases where both got it wrong, and only compare log-probabilities and response times within these cases. This breakdown is necessary to control for accuracy in these models, as it is a significantly related to both response times and log-probabilities. Note, however, that this means that problems where a model answered correctly but humans never answered correctly, or vice versa, are omitted.
In both tasks, we see significant effects of the content on the model log-probability differences; even controlling for these we see significant relationships to the human response times, such that on items on which the humans respond more slowly, the models show smaller differences in log-probabilities.