Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks

Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim

Introduction

The striking empirical successes of language models (LMs) suggest that next-word prediction at scale may be a viable approach for distilling the knowledge embedded in large-scale text corpora into general-purpose interactive agents. LMs obtain impressive results on various NLP benchmarks (OpenAI, 2023; Anil et al., 2023; Anthropic, 2023; i.a.) and display surprising abilities that suggest a nontrivial understanding of the world (Bubeck et al., 2023). They have been shown to pass professional exams (Kung et al., 2023; Nori et al., 2023; Terwiesch, 2023; i.a.), exceed state-of-the-art methods on many traditional benchmarks (Sun et al., 2023; Sobania et al., 2023; Zhang et al., 2023a; Dhingra et al., 2023; i.a.), and surpass human performance on tasks that require seemingly nontrivial reasoning (Chowdhery et al., 2022; Hoffmann et al., 2022; Malinka et al., 2023; Guo et al., 2023; i.a.).

Ideally, we expect a general-purpose LM to be able to generalize not only to unseen instances of known tasks, but to new tasks. Humans, for example, can transfer their knowledge to new instances and also flexibly adapt to novel tasks (Singley and Anderson, 1989). To what extent does the performance of current LMs derive from their ability to deploy task-general reasoning skills, versus their ability to recognize and recall specific tasks seen frequently in pre-training?

Past work has focused on instance-level generalization, but this is often complicated by data contamination issues (Dodge et al., 2021; Magar and Schwartz, 2022; i.a.). In this work, we are interested in the models’ generalizability to new task variants, which has been less systematically studied for LMs (though see Li et al. (2022), Mishra et al. (2022), and Wang et al. (2022b)).

We propose to measure such task-level generalizability by taking tasks on which LMs perform well, and altering the conditions or rules under which these tasks are performed. The general reasoning procedure for these tasks remains the same under the new conditions, but the specific input-output mappings are changed. We call the new tasks counterfactual tasks, as they deviate from the default, generally assumed conditions for these tasks. Figure 1 shows examples: in the top left, default arithmetic is performed in base-10, while counterfactual arithmetic is performed in base 9. If models implement a general and transferable task-solving procedure, we expect comparable performance on counterfactual and default tasks; if they employ procedures tailored to default task conditions, we expect a drop in the counterfactual performance.

We design a suite of 11 counterfactual evaluation tasks to measure an LM’s flexibility to adapt to new task variants across multiple categories and domains, as summarized in Figure 1. In each, the original task under the default conditions and its counterfactual variants share the same reasoning procedure but differ in their input-output mappings. We consider traditional NLP tasks such as deductive reasoning, non-language tasks that are nonetheless commonly evaluated such as code generation, as well as non-standard tasks such as drawing and spatial reasoning. The latter extralinguistic tasks test whether LMs are able to learn conceptual structures that mirror the structure of the non-linguistic world, which has been suggested by recent work (Abdou et al., 2021; Ilharco et al., 2021; Patel and Pavlick, 2022; Li et al., 2023a; Bubeck et al., 2023; Søgaard, 2023; i.a.).

We evaluate the performance of GPT-4 (OpenAI, 2023), GPT-3.5, Claude (Anthropic, 2023), and PaLM-2 (Anil et al., 2023) on tasks under both the default and counterfactual conditions. We observe above-random counterfactual performance for most tasks, indicating some degree of task generalizability. However, their performance on counterfactual task variants consistently and substantially degrades relative to the performance on the default settings. This suggests that these models’ ability on these tasks is supported at least in part by non-transferable, default-condition-specific behaviors rather than abstract, generalizable reasoning skills.

These results also reveal several surprising relations between model behavior on default and counterfactual tasks (§5), including correlations between default and counterfactual performance, varying effectiveness of zero-shot chain-of-thought prompting Kojima et al. (2023), and interactions between task- and instance-level frequency effects. Overall, we find that small variations on the default instantiations of tasks are challenging for models, and thus the success of existing LMs should not be fully attributed to fully general capacity for the target task.

Counterfactual Tasks

We informally conceptualize each task as a function fw:X→Yf_{w}:X\to Y that maps an input x∈Xx\in X under a world model w∈Ww\in W to an output y∈Yy\in Y. World models encapsulate the conditions under which function evaluation takes place. For example, in Python programming, ww might specify assumptions of Python such as indexing and operator precedence; in arithmetic, ww could represent the set of conditions required for an arithmetic operation, such as the number base. We refer to the set of assumed default conditions, including but not limited to the base’s being 10, as the default world, or wdefaultw^{\text{default}}. Intuitively, for any task, wdefaultw^{\text{default}} corresponds to the set of conditions underlying the majority of task instances in text corpora.This data-generating process can be described by the following generative model, P(y ∣ x,w)P(x ∣ w)P(w)P(y\,|\,x,w)P(x\,|\,w)P(w). From the perspective of causal inference, our counterfactual framework can be informally seen as performing a do⁡\operatorname{do}-operator on this graph Pearl (2009).

Traditional evaluations of machine learning models assess how closely a model’s learned hypothesis hh estimates fwf_{w} by independently sampling training and test sets from the population distribution Dfw\mathcal{D}_{f_{w}}, and only exposing the model to the training set for learning hh. However, in datasets of scraped web text, these evaluations are subject to potential data contamination issues (Brown et al., 2020; Dodge et al., 2021; Magar and Schwartz, 2022; i.a.). These issues may be more severe in recent LMs: the ever-growing pretraining datasets potentially expose the models to more evaluation instances, and the increasing sizes of recent LMs give them more ability to memorize these instances (Carlini et al., 2020; Magar and Schwartz, 2022).

We hence consider another dimension of generalization; generalization to new task variants in counterfactual worlds wcfw^{\text{cf}}, instead of new inputs xx. This allows us to measure the extent to which a model’s fwdefaultf_{w^{\text{default}}} performance is specific to wdefaultw^{\text{default}} or attributable to a general implementation of the task ff.This setup is reminiscent of intensional models of natural language semantics (Heim and Kratzer, 1998, §12; Von Fintel and Heim, 2011), where ff is analogous to the denotation function ⟦⋅⟧\llbracket\cdot\rrbracket, xx to its input, and yy to its output. By default, the denotation is evaluated under the real world, extensionally, but when a different possible world is specified instead, we expect a competent system to adjust the evaluation accordingly. For arithmetic, a possible wcfw^{\text{cf}} would be the same as wdefaultw^{\text{default}} but assuming a base other than base-10. We expect a model with general arithmetic ability to perform similarly in other bases.

We emphasize that our goal is not to find counterfactual world models that are completely outside the realm of human experience. Base-9 addition, for example, is not a novel concept. Nor do we aim to guarantee that counterfactual world models are unobserved in a pretraining corpus. Instead, counterfactuals are simply defined as variations on the default conditions for a task.

Concretely, we assess an LM’s task performance with 0-shot prompting. We specify the task ff, the test instance xx, and the world model ww in a prompt, parse the LM’s output, and compare it to the ground-truth label. We denote the LM’s implementation of fwf_{w} for a given instance xx to be,

where the arg max⁡\operatorname*{arg\,max} is computed with an approximate decoding procedure and prompt⁡f\operatorname{prompt}_{f} and prompt⁡w\operatorname{prompt}_{w} are prompt templates that describe tasks and world models respectively. For each task, we devise one or more wcfw^{\text{cf}} that deviate from the default world (i.e., the default task conditions). We evaluate both h(f,wdefault,x)h(f,w^{\text{default}},x) and h(f,wcf,x)h(f,w^{\text{cf}},x) via task-specific metrics. If we control fw(x)f_{w}(x) to be similarly hard for both wdefaultw^{\text{default}} and wcfw^{\text{cf}}, we can attribute the performance difference to an LM overfitting to the default instantiation of the task.

One potential confounder is that an LM may be failing at a particular counterfactual task by failing to understand the prompt component that specifies the counterfactual conditions, i.e., prompt⁡w(wcf)\operatorname{prompt}_{w}(w^{\text{cf}}). That is, an LM might still be reasoning in wdefaultw^{\text{default}} and completely ignore the instructions. While this would still be a failure of the LM, it does not necessarily represent a failure to perform the counterfactual task variant. We control for this by designing task-specific counterfactual comprehension checks (CCCs) that test an LM’s surface understanding of the specified counterfactual world.

For each (default, counterfactual) task pair, we introduce another control task gwg_{w} with input x′x^{\prime} and output y′y^{\prime} that is much simpler than fwf_{w} but still allows for the discrimination of wdefaultw^{\text{default}} from wcfw^{\text{cf}} (i.e., gwcf(x′)≠gwdefault(x′)g_{w^{\text{cf}}}(x^{\prime})\neq g_{w^{\text{default}}}(x^{\prime})). A high performance of PLM(y′ ∣ prompt⁡g(g,x′),prompt⁡w(wcf))P_{\text{LM}}(y^{\prime}\,|\,\operatorname{prompt}_{g}(g,x^{\prime}),\operatorname{prompt}_{w}(w^{\text{cf}})) would indicate that prompt⁡w\operatorname{prompt}_{w} is effective at making the LM perform a task in wcfw^{\text{cf}}. In the arithmetic example, for a base-9 counterfactual world, we use the same prompt⁡w(base-9)\operatorname{prompt}_{w}(\texttt{base-9}) to specify the counterfactual world, and check that it facilitates an understanding of w=base-9w=\texttt{base-9} by asking what the next integer after x′x^{\prime} is. If, for example, it consistently carries over digits greater than 8 and does not carry over otherwise, this would show the effectiveness of prompt⁡w(base-9)\operatorname{prompt}_{w}(\texttt{base-9}). Our CCC designs are heuristic: as with control tasks in the probing literature (Hewitt and Liang, 2019), we rely on intuition to craft a gwg_{w} that is “simpler” than fwf_{w}.In this formulation, LM queries for CCC are separate from the main task queries. For some tasks, it is more natural to query about the task and CCC jointly in the same prompt, i.e., PLM(y,y′∣prompt⁡f(f,x),prompt⁡g(g,x′),prompt⁡w(wcf))P_{\text{LM}}(y,y^{\prime}|\operatorname{prompt}_{f}(f,x),\operatorname{prompt}_{g}(g,x^{\prime}),\operatorname{prompt}_{w}(w^{\text{cf}})). We use this formulation instead for those tasks.

Tasks

In this section, we give a quick overview of the tasks we consider. See §A for the full description of each task and §B for all the prompts used.

Modern LMs have been shown to possess basic numerical reasoning abilities (Lewkowycz et al., 2022), with even GPT-3 reporting near-perfect accuracy for two-digit additions (Brown et al., 2020). On the other hand, Razeghi et al. (2022) find that LMs perform significantly better on operations involving numbers that occur more frequently in the pretraining data, and Li et al. (2023d) show that symbol replacement affects the mathematical ability of BERT (Devlin et al., 2019)-like models; both findings point to overfitting and memorization effects. We consider the same two-digit addition task, the simplest arithmetic task in Brown et al. (2020), but inspect a model’s accuracy in different bases. We use base-8, 9, 11, and 16 as the counterfactual setup which are natural generalizations to base-10 arithmetic. These bases were chosen to control for task difficulty (see §7.1 for a discussion) and also to test for how relatively uncommon (9 & 11) and common (8 & 16) bases affect performance (see §5.1 for an analysis). To ensure the model understands the different bases, the CCC evaluates the successor relation under each base.

2 Programming

Even without explicit pretraining on large amounts of code, LMs have been found to possess decent coding ability (Brown et al., 2020). The inclusion of large code corpora in LM pretraining (Gao et al., 2021; Chowdhery et al., 2022; Touvron et al., 2023; i.a.) further improves this capability in recent LMs, with ChatGPT sometimes outperforming state-of-the-art approaches for bug fixing (Sobania et al., 2023). Nevertheless, Miceli-Barone et al. (2023) show that GPT-3 and related models are fragile under identifier swaps in programs, suggesting that these models may only possess a shallow understanding of code. Here, we inspect an LM’s programming ability through a deeper counterfactual perturbation: contrary to the traditional 0-based indexing in Python, we instruct the LM to evaluate or generate programs under a fictional language, ThonPy, that uses 1-based indexing but is otherwise identical to Python. 1-based indexing is a common assumption for other programming languages such as MATLAB and R and hence provides a fair testbed. We evaluate the LM’s performance using the HumanEval dataset (Chen et al., 2021). The CCC here involves the same program execution task but on much simpler inputs, such as simple list indexing, that do not involve deeper reasoning.

3 Basic Syntactic Reasoning

Mahowald et al. (2023) distinguish between two types of LM capabilities: formal competence that encompasses the knowledge of language, and functional competence which involves using language, potentially combined with extralinguistic capacities, to interact with the world. While the other tasks we investigate in this paper assess a model’s functional competence, we also include an evaluation on formal competence. We revisit the attested syntactic knowledge of LMs (Yu et al., 2020; Linzen and Baroni, 2021; Ettinger, 2020; Pimentel and Cotterell, 2021; Belinkov, 2022; Lasri et al., 2022; i.a.) by considering a meta-linguistic task (Beguš et al., 2023; Hu and Levy, 2023; i.a.): evaluating LMs in synthetic versions of English with different word orders from English’s subject-verb-object (SVO) ordering. We ask the LM to identify the main subject and the main verb of a sentence under both the original and counterfactual orders, where the latter is obtained from manipulating dependency trees (Ravfogel et al., 2019). The CCC requires the model to revert simple reordered sentences to the original SVO ordering, equivalent to identifying these elements in a sentence.

4 Natural Language Reasoning with First-Order Logic

We next consider a deductive reasoning task that is still based on natural language. Logical reasoning is a prerequisite ability for many complex tasks (McCarthy, 1959) and has been the focus of much recent work (Clark et al., 2020; Tafjord et al., 2021; Saparov and Mitchell, 2022; Saparov and He, 2023; i.a.). Nevertheless, LMs struggle with reasoning with premises that are inconsistent with common sense (Dasgupta et al., 2022; Yu et al., 2023; Tang et al., 2023). Here, we undertake a similar study from the perspective of counterfactual analysis to disentangle the effect of common sense from a model’s actual logical reasoning capability.

Following prior work, we evaluate in an entailment format and ask LMs if a series of premises entails a conclusion. We use the FOLIO dataset (Han et al., 2022) most of whose premises are consistent with common sense, and manually rewrite them to violate common sense. We study if LM performance is affected by the truthfulness of the premises under which they operate. The CCC directly asks the model if the original or post-rewrite premise is true, when presented both as options.

5 Spatial Reasoning

A major debate around LMs is whether grounded representations of meaning can be learned from form alone (Bender and Koller, 2020; Piantadosi and Hill, 2022; Mollo and Millière, 2023). Studies have shown that LMs can learn meaningful world representations through text-only training (Abdou et al., 2021; Li et al., 2023c; Jin and Rinard, 2023). In particular, Patel and Pavlick (2022) find that LMs learn representations of cardinal directions that can be aligned to grounded conceptual spaces with few-shot demonstrations.

We similarly investigate an understanding of cardinal directions, but instead of evaluating whether a model can induce structured conceptual spaces, we ask if it can apply conceptual spaces to reason about the locations of objects. Specifically, we ask an LM for the coordinates of objects whose positions are described using cardinal directions, under a conventional 2D coordinate system (e.g., where east corresponds to (1,0)(1,0)) versus coordinate systems with swapped, rotated, and randomly permuted axes. We expect a robust representation to not be sensitive to such transformations. The CCC involves asking the model to directly output the counterfactual cardinal directions.

6 Drawing

Despite being trained on only textual data, LMs have been shown to be able to structure their representations of perceptual concepts such as size and color (Abdou et al., 2021; Patel and Pavlick, 2022; Zhang et al., 2020; Ilharco et al., 2021; i.a.) in a way that credibly mirrors the physical world. Recent LMs can even generate plausible drawings of objects using code such as TikZ and SVG Bubeck et al. (2023); Zhang et al. (2023c). We evaluate the visual understanding of LMs by asking them to generate code for drawing various objects in the Processing language. Psychological studies have shown that humans have the ability to rotate mental representations of objects Shepard and Metzler (1971); Vandenberg and Kuse (1978). For the counterfactual settings, we similarly ask the LM to generate code that draws the same object, but rotated or vertically flipped. We disallow the use of functions such as rotate to prevent shortcut solutions (see §7.2 for further discussion). As with the spatial reasoning task (§3.5), an ideal model should be robust to these settings. For the CCC, we ask the model to draw a straight line at the top of the canvas in addition to the object; a flipped/rotated line thus signifies an understanding of the transformations.

7 Music

Recent work has shown the potential of large-scale models for music infilling (Huang et al., 2019a, b) and generation (Agostinelli et al., 2023; Copet et al., 2023; Ren et al., 2020). Bubeck et al. (2023) show that even a text-only LM with no music-specific pretraining exhibits some musical abilities, including understanding musical structure and manipulating melodies. We investigate the extent of LMs’ musical abilities through two tasks.

In the chord placement task, we evaluate whether LMs can provide the correct chord fret placements for string instruments with standard or altered string tunings. The altered tunings, known as scordatura, are typical in music and are used to evoke a specific sound or effect (e.g., enabling heavier, deeper sound in metal music). We evaluate LMs using an existing databasehttps://github.com/tombatossals/chords-db that includes chords for guitar and ukulele. In the counterfactual setting, we instruct LMs to provide fret placements for a special guitar/ukulele where one or two of the strings are altered. For guitar, we include drop-D tuning, a popular alternative guitar tuning that allows us to investigate whether the frequency of counterfactual tunings affects results (see §5.1). To check whether the model has understood the tunings, we ask for the first three notes on each string (including open string) as the CCC.

In the note retreival task, we evaluate whether LMs can retrieve notes from famous melodies (e.g., “Twinkle Twinkle Little Star”). The process of re-writing melodies in different keys, referred to as “transposition,” is common in music (e.g., to accommodate the ranges of different singers or instruments). We evaluate LMs’ musical abilities under transpositions by prompting them to retrieve the nn-th note in a melody in either its canonical key (default setting) or a different key (counterfactual setting). We ask the LMs to retrieve the nn-th note of the scale of the given key as the CCC.

8 Chess

Chess playing has long been regarded as a testbed for AI Silver et al. (2017); Tomasev et al. (2020), and modern LMs have exhibited abilities that imply an understanding of chess rules Srivastava et al. (2023); Du et al. (2023). We test this understanding by asking for the legality of a 4-move opening. In the counterfactual setting, we swap the initial positions of knights and bishops—a setup present in a real-world chess variant “Chess 960”—and similarly ask LMs for opening legality under this new starting configuration.A conceptually similar analysis was performed in Li et al. (2023c) for the game of Othello. We ask for the starting positions of the knights and the bishops as the CCC.

9 SET Game

SET is a popular card game where each card has 4 attributes with 3 different values for each attribute:

In each round, a player finds a SET of 3 cards in a 12-card board whose values for each attribute are either all the same or all unique. This game has been thoroughly studied in computer science, from the perspective of coding theory and combinatorics (Davis and Maclagan, 2003), linear algebra (Coleman and Hartshorn, 2012), and complexity theory (Chaudhuri et al., 2003). We suspect this popularity makes it susceptible to overfitting by LMs and investigate this possibility. We ask the LM to identify the card on a board that completes a 3-card SET with two given cards. In the counterfactual setup, we invert the rule for the number attribute, requiring its value to be mixed, in other words, neither all the same nor all unique. For the CCC, we ask the model for the validty of a SET under the original rule and the counterfactual rule.

Results

For each task, we evaluate GPT-4 (gpt-4-0314; OpenAI, 2023), GPT-3.5 (gpt-3.5-turbo-0301), Claude (claude-v1.3; Anthropic, 2023), and PaLM-2 (text-bison-001; Anil et al., 2023). As these are closed-source models, we do not have any information regarding their size, architecture, and pretaining details.We also explored open-source models in preliminary experiments, but found that they possess unsatisfactory instruction-following ability, to the point that often their output cannot be meaningfully parsed into a prediction. We therefore do not include these models. We note that the largest PaLM model is not publicly accessible, and we can only test the second-largest version. For each task, we experiment with both with and without encouraging the model to reason step by step, by adding the phrase “Let’s think step by step.” in our prompts (Kojima et al., 2023; Reynolds and McDonell, 2021). Following Kojima et al. (2023), we refer to this step-by-step setup as zero-shot chain-of-thought prompting (0-CoT; Nye et al., 2021; Wei et al., 2022). We include all prompts in §B.

Figures 2 and 3 show our results. §C contains the numeric version. We see a consistent pattern where LMs perform substantially worse on the counterfactual task variants, both with and without 0-shot CoT. For most cases, LMs exhibit an above-random counterfactual performance, suggesting some degree of the targeted ability. However, when the CCC accuracy is high, usually the case for GPT-4 and in select settings for other models too, the default-counterfactual gaps demonstrate limitations in the abstract capacity to solve the target task. When the CCC accuracy is lower, the failure of counterfactual world comprehension would be a confounder to this conclusion, but often the gaps are so large (sometimes even dropping from near-perfect to near-zero, such as for arithmetic) that they are nonetheless strongly indicative of non-transferable, default condition-specific implementations of the original task. The fact that the LMs sometimes cannot evaluate the CCC well under the counterfactual conditions, but can do so under the default conditions (e.g., for arithmetic, programming, drawing, etc.) itself also points to overfitting to the latter.

Analysis

We now investigate how a variety of factors affect the default and counterfactual performance trends that we observed in §4. Unless otherwise specified, we only consider GPT-4 with 0-shot CoT, which has the strongest performance in our results above.

Our counterfactual worlds are not designed to be completely alien to the LMs but only less common than the assumed default case. In this sense, the counterfactual-ness of these worlds is relative, and here we take a more nuanced look at how the commonness of these counterfactual conditions affects the default-counterfactual performance gap. For example, in the arithmetic task, all models perform better in bases 8 and 16, likely due to their relative abundance compared to bases 9 and 11. In spatial reasoning, the smallest counterfactual performance degradation is usually from when the north and south directions are swapped—even exceeding the default task performance for PaLM-2—potentially because some programming libraries use an inverted yy-axis, such as matplotlib (Python), ggplot (R), and D3 (JavaScript) (see §A.5). For chord fingering, the common alternative drop-D tuning of guitars leads to the highest counterfactual performance. These correlations between the counterfactual performance and the commonness of the counterfactual worlds paint a more fine-grained picture than a binary default versus counterfactual distinction and point to a memorization-like effect where the models perform better under more common conditions.

2 Proximity between Default and Counterfactual Conditions

Another axis along which the counterfactual worlds differ is in their proximity to the default conditions. For example, for the different arithmetic bases, bases 9 and 11 are closer to base 10, but less common than bases 8 and 16. While the default-counterfactual gap is most affected by commonness for the arithmetic task, for the guitar and ukulele tunings (other than the drop-D tuning), the LM performance generally decreases monotonically with the distance from the original tunings.

The FOLIO dataset (Han et al., 2022) enables another analysis of how proximity to the default conditions affects the model performance, without performing counterfactual perturbations. This dataset was constructed to mostly follow common sense, i.e., containing premises and conclusions that are deemed true in the real world. However, this is not always the case, with premises such as “John can make meals which are popular at the party,” whose factuality cannot be determined alone.

We evaluate how the distance between the world state described by the premises and the belief state of LMs influences LM performance by training a predictive model given features approximating this distance. For each test instance, we ask the LMs whether the premises and conclusion are true, false, or uncertain. We train a logistic regression model to predict LM correctness on each test instance, using as features the total number of premises in an input, the proportion of the premises that are true/false/uncertain, as encoded by the LM, as well as whether the LM-predicted truthfulness of the conclusion matches the label of the instance (that is, a feature that predicts the entailment/neutral/contradiction label of the instance from the truthfulness of the conclusion alone, ignoring premises).

Figure 5 shows the learned coefficients of these features, as well as their 95% confidence interval bootstrapping with 1,000 iterations (Efron and Tibshirani, 1993). Ideally, a robust model should predict solely based on symbolic deduction and extralinguistic truthfulness information should not affect its accuracy. In other words, these features should all have coefficients 0 and have no predictive power with respect to the model’s correctness. However, all LMs predict more correctly with more realistic (true) premises, and when the conclusion’s LM-predicted truthfulness matches the label. On the other hand, they tend to perform worse when there are more false or uncertain premises. Most of these trends are statistically significant. This means that the reasoning ability of LMs is affected by the distance between the real world (as believed by the LMs) which is the default condition and the world state under which reasoning is required.

Overall, these results show that LMs tend to perform better on task variants that are closer to the default instantiation of a task.

3 Relationship between Default vs. Counterfactual Performance

Recalling our formalization hLM(f,w,x)h_{\text{LM}}(f,w,x) in §2, the previous two subsections analyzed how the commonness of ww and its proximity to wdefaultw^{\text{default}} affect the observed patterns. We now explore how the counterfactual performance correlates with the default task performance by varying the other three elements in this formalization: the task ff, the input xx, and the LM.

We first consider different task variants. For arithmetic, beyond 2-digit addition, we also measure GPT-4’s 3- and 4-digit addition performance (Figure 4(a)).See Dziri et al. (2023) for a related analysis but for multiplications. For note retrieval from melodies, we use the index of the inquired note as the proxy for difficulty (Figure 4(b)).This is not a new task variant as compared to the setup in §3.7, but rather a decomposition of our original results. For SET, while our original task shows two cards and asks a model to find the missing one from a 3-card SET, we change the task to instead only show one or none of the cards in a SET, while still requiring the model to identify the SET on a board (Figure 4(c)). For all these task variants, we see a strong correlation between the original and the counterfactual world performance.

We also see this effect when breaking down results by test instances. In Figure 4(d), we separate the different chord types, and observe that the default task performance correlates with the counterfactual performance. Similarly, reexamining our main results in Figures 2 and 3, for most tasks, stronger models under default conditions are also stronger models under counterfactual conditions, and vice versa. Overall, these correlations mean that the default task performance can be a good indicator of its counterfactual performance, and hence we should not discount the utility of traditional evaluations.Though these correlations are not necessarily causal.

Furthermore, despite our evidence of LMs’ overfitting to the default task conditions, these correlations also signify some degree of reasoning that is transferable between the default and counterfactual worlds. This highlights that the question in our title, “Reasoning or Reciting?”, is not a dichotomy, but rather they can co-exist in a continuum. For example, revisiting the arithmetic results with more digits (Figure 4(a)), in addition to the default-counterfactual correlation, we also see an effect of memorization: the base-10 performance decreases much more slowly than the other bases. When the input-output mappings are memorized, increased complexity would not affect the default task accuracy much; but when the counterfactual instances are not memorized, the task complexity should inversely correlate with model performance.

Occasionally, this default-counterfactual correlation trend is reversed. In the spatial reasoning task, for example, GPT-4 achieves the best accuracy under default conditions with 0-shot CoT, but it also suffers from the largest counterfactual performance degradation. PaLM-2 performs worse under default conditions, but is the most robust to counterfactual perturbations. An obvious possible explanation is that these models could be trained on different data, and are hence familiar with different conditions. Nevertheless, McKenzie et al. (2023), who found a similar trend but with respect to pretraining FLOPs and termed it “inverse scaling,” also provided a memorization-based explanation: they observed that when a task contradicts with pretraining texts, similar to how our counterfactual conditions deviate from the default conditions in pretraining, larger LMs tend to rely on the pretraining text and, in turn, fail at the contradictory task.

4 0-Shot Chain-of-Thought Prompting

Consistent with prior findings (Chen et al., 2022; Dasgupta et al., 2022; i.a.), we generally observe 0-shot CoT to be helpful for most cases. There are, however, exceptions. For example, 0-shot CoT substantially hurts PaLM-2’s addition performance in base-10 and 16, and consistently degrades GPT-4 and GPT-3.5’s chord-playing performance for the default tuning. This may be due to a model pragmatically inferring that a task is more difficult than it actually is when explicitly asked to “think step by step”, and this “overthinking” on simple tasks could lead to mistakes (Kojima et al., 2023). It is also possible that these are due to memorization: the model could have memorized the specific input-output mapping of a task, without understanding how to derive the output from the input, and when explicitly instructed to spell out that process, it makes more errors (Zhang et al., 2023b).

5 Few-shot Demonstrations

We study if additional demonstration examples using in-context learning (Brown et al., 2020) bridges the default-counterfactual gap. For the arithmetic task, we construct few-shot CoT prompts (Nye et al., 2021; Wei et al., 2022) and prepend up to 16 samples. As shown in Figure 6, while the gap is reduced, it is still substantial for bases 9, 11, and 16. Moreover, the accuracy improvement with more demonstrations plateaus towards 16-shot, suggesting that the default-counterfactual gap is unlikely to be eliminated by simply adding more demonstrations (at least for arithmetic).An interesting pattern is that bases 11 and 16 suffer from 1-shot demonstration than 0-shot. We hypothesize that this may be due to these being the two bases with letter digits.

6 Qualitative Analysis of Drawing Results

We conduct a qualitative error analysis on the drawing task and show some examples in Figure 7. We first note that GPT-4 successfully passes the CCC for these cases (see §3.6; but not displayed here), indicating that it understands the flip/rotation instructions. However, the objects in the counterfactual worlds are often not flipped or rotated. Even when they are transformed appropriately, the resulting drawing is often simplified or of worse quality (e.g., Unicorn, Cake). We also observed much more syntactically invalid programs in the counterfactual cases for GPT-3.5.On average, the number of parseable programs generated by GPT-3.5 drops from 99% in the default condition to 62%, 71%, and 75% for the vertically flipped, 90° rotated, and 180° rotated settings, respectively. These results indicate that even when a model can perform a task in the counterfactual setup, its capabilities are reduced.

Discussion

Do humans also perform worse with unfamiliar counterfactual conditions? It is possible that humans may have lower performance under the counterfactual conditions with a fixed time budget, but not necessarily when given ample time to reason and revise. Analogous to the classic competence/performance distinction in linguistics (Chomsky, 1965, §1.1), we hypothesize that humans have the competence to generalize to new task conditions, even though it may sometimes require sufficient execution budget to realize it as robust performance.It is arguable if our evaluation setting provides sufficient execution budget (Lampinen, 2023). Our in-context learning experiment (§5.5) may be thought of as increasing this budget, and yet the default-counterfactual gap is still sizeable there. In fact, there is increasing evidence from cognitive science that human reasoning is scaffolded by rich causal models of the world (Pearl, 1988; Lake et al., 2017; Ullman and Tenenbaum, 2020; Wong et al., 2023), and that humans can intervene on these models to perform rapid and flexible counterfactual simulations (Lagnado et al., 2013; Gerstenberg et al., 2017, 2021). However, stepping back, replicating or modeling human intelligence need not be a main goal of LMs in the first place, and human behavior is largely orthogonal to the desiderata we set for these models.

Is task-specific reasoning bad? It is not necessarily bad when solving familiar tasks, but an ideal system should also possess general reasoning abilities that, when prompted, can be used to generalize to novel situations. Our point is that memorization is an often-overlooked confounding factor in interpreting LMs’ reasoning abilities.

Why do we care about counterfactual worlds? Wouldn’t a model for only the default task instantiation be nonetheless useful? It is certainly true that such a model would still be useful. However, many of the counterfactual worlds that we investigate actually are not very distant so that model performance under them still bears utility. For example, addition in different bases is certainly useful for many applications. More generally, we are necessarily interested in the counterfactual tasks themselves; we are only interested in them insofar as performance on these tasks can serve as a measurable proxy for the generalizability of these models and their underlying reasoning capabilities.

Aren’t the observed trends trivial? The default task variant is likely the most frequent during pretraining, so of course an LM performs better under it. Indeed, our results parallel the classic train-test gap in machine learning. However, an ideal learner with the right inductive biases should be able to structure their internal parameters and representations to implement general-purpose abstractions (e.g., the concept of addition), and use these abstractions to generalize to counterfactual conditions, analogous to physicists using mathematical abstractions to make predictions about universes that are substantially different from our own, or more generally to humans who can generalize to new stimuli in cognitive science studies (Lagnado et al., 2013; Gerstenberg et al., 2017, 2021). Our study indicates that LMs trained on large text corpora, remarkable as they may be, are still quite susceptible to overfitting with frequency effects.

Can some more carefully designed prompts eliminate the default-counterfactual gap? This is always a possibility, and one that we can never tractably rule out. Nevertheless, given the consistent patterns across our tasks, we believe that a prompt that completely bridges the default-counterfactual gap is unlikely. Our in-context learning experiment (§5.5) shows that this gap could be reduced, but not fully removed, by more informative prompts. It would be interesting to apply more advanced prompting techniques (Wang et al., 2023a, 2022a; Yao et al., 2023; Sordoni et al., 2023; i.a.) to our counterfactual tasks. We considered 0-shot chain-of-thought in this work, which did not fully bridge the default-counterfactual gap, but we leave the exploration of these more recent prompting techniques to future work.

Limitations

Despite our attempt to devise novel counterfactual conditions to gauge an LM’s “true” reasoning ability, it may not be precisely reflected by the counterfactual performance due to several factors.

For our main evaluations, we aim to construct counterfactual tasks that have the same difficulty as the default variants so that task difficulty does not confound our comparisons. This is not always possible—in fact, an objective difficulty measure may not even exist. One could, for example, argue that base-11 addition is harder than base-10 because it requires reasoning with one additional digit, or base-9 is harder than base-10 because on average the sums would consist of more digits.

Retrieving notes in melodies in different keys faces a similar issue: one way of retrieving a note in a melody in an uncommon key would be to first retrieve it in a canonical key and then transpose it to the desired key. With this strategy, the counterfactual task consists of 2 steps and is harder than (and requires first) completing the 1-step original task. This strategy is not the only way to solve the task: an alternate one would be to recall a melody as a series of abstract relations in a scale and directly map them onto notes in a target key. However, the 2-step process is a natural one that is often employed by musicians. The counterfactual setup thus introduces a confounder: low performance may be driven by the increased difficulty of the counterfactual task, rather than overfitting to melodies in their canonical keys, if models are employing two-step strategy. However, since both strategies are available to models and we do not prompt them to use a particular one, reliance on this two-step strategy may itself be indicative of overfitting to the original canonical keys.

2 Overestimation

We can never be certain of how rarely particular counterfactual conditions are encountered during pretraining. It is quite likely that there is text online that, for example, draws rotated versions of various objects used in our study. Consequently, the effect of overfitting could also manifest in our counterfactual conditions, and the default-counterfactual gap could actually be larger for some genuinely unseen conditions.

We also distinguish between two types of counterfactual perturbations. One type fundamentally affects the operation of the world model and necessitates an understanding of the counterfactual world to perform the task in it (e.g., arithmetic base or 1-based indexingIt may be tempting to consider a simple replacement strategy [i]→[i-1] to map back to 0-based indexing. But this does not work for the dictionary type. There are other complications; see Table 2.). On the other hand, some perturbations are more superficial and may admit a shortcut where the model first figures out a simple mapping of the input back to the default conditions and performs the task (potentially leveraging instance-level memorization) under those. In some of our tasks, this mapping may be simple, such as the word replacements in the natural language logical reasoning taskTo be more concrete, imagine that a model memorizes an instance with nine premises on dogs involving complex logical relationships, and that it entails a given conclusion. For the counterfactual instance, we replace the word “dogs” with another object, say “headphones,” to make the premises no longer factually true. Instead of performing the reasoning over premises with headphones such as how they are, counterfactually, the cutest creatures, a model could identify the mapping “dogs” →\to “headphones”, revert it (i.e., replace all “headphones” back to “dogs”), and perform the task under the default common-sense-complying conditions. (§3.4) and the transformation functions for the drawing task (§3.6), which could potentially be exploited by the models. We explicitly disallow this in our prompt for the drawing task (Table 7) but did not identify a good way to forbid this for logical reasoning, potentially accounting for its generally high counterfactual performance.

Finally, we reiterate from §4 that a non-perfect CCC accuracy does not allow us to perfectly tease apart counterfactual performance and a failure of counterfactual condition comprehension. But often the default-counterfactual gap is so prominent that it is still strongly suggestive of overfitting to the default conditions. Also, recall from §2 that the CCC itself is also a nontrivial task. For ThonPy, for example, the CCC also involves program evaluation, albeit with simpler statements that involve less reasoning, such as print("qrstu"). We do not see an easy way to introduce ThonPy CCC that is entirely disentangled from program evaluation. This conflation would result in the CCC accuracy’s being lower than what would reflect the model’s understanding of the counterfactual conditions.

Related Work

Much prior work has investigated the extent to which LMs acquire a grounded understanding of the world through text-only training (Piantadosi and Hill, 2022; Zhang et al., 2020; Ilharco et al., 2021; Li et al., 2021; i.a.). These works have generally found that conceptual structures of certain concepts (e.g., color, size) often plausibly mirror those of the grounded world (Abdou et al., 2021; Patel and Pavlick, 2022; Mollo and Millière, 2023). As in our study, these studies are a test of generalization—such structures would not manifest if the concepts were memorized in a one-hot-like manner. But our evaluation differs in that it targets the reasoning process instead of the generalization to new concepts or conceptual structures (Kondo et al., 2023). While prior work identified that the latter is embedded in LMs, we found that they do not fully learn the former.

Our counterfactual perturbations can be informally viewed as interventions under a causal inference framework (Pearl, 2009). This relationship has been explored in machine learning and NLP for commonsense reasoning (Kıcıman et al., 2023), interpretability (Elazar et al., 2021; Geiger et al., 2021, 2022), spurious correlation detection (Veitch et al., 2021; Eisenstein, 2022), and fairness (Kusner et al., 2017; Nabi and Shpitser, 2018). Under this perspective, the failure of generalization to counterfactual worlds that we observe in LMs can be viewed as a failure to robustly learn the causal effects of world states on our evaluated tasks.

“Counterfactuals” is an informally-used term in NLP and has been used to refer to different types of perturbations. One line of work concerns counterfactuals to a certain event or situation that is still licensed in a default world model (Qin et al., 2019, 2020; Yang et al., 2020; Frohberg and Binder, 2022; i.a.), in contrast to our counterfactual world states that deviate from the default. Qin et al. (2019) and Frohberg and Binder (2022) found that GPT-3 and earlier models struggle with consistently reasoning under this type of counterfactual conditions, while Kıcıman et al. (2023) observed more recent LMs to achieve higher counterfactual reasoning accuracy. Another body of work examines the robustness of model predictions using counterfactual data (Kaushik et al., 2020, 2021; Gardner et al., 2020). More similar to our study, Li et al. (2023b) showed that while the LMs they investigated seem to be able to perform some reasoning in counterfactual worlds, this is largely affected by superficial lexical cues. Our results reveal that more recent LMs still exhibit such difficulties.

Conclusion

Through our counterfactual evaluation on 11 tasks, we identified a consistent and substantial degradation of LM performance under counterfactual conditions. We attribute this gap to overfitting to the default task variants, and thus encourage future LM analyses to explicitly consider abstract task ability as detached from observed task performance, especially when these evaluated task variants might exist in abundance in the LM pretraining corpora.

Acknowledgments

We thank, alphabetically, Alex Gu, Alisa Liu, Belinda Li, Chenghao Yang, Han Guo, Hao Peng, Heyun Li, Jesse Dodge, Pratyusha Sharma, Tiwa Eisape, and Yizhong Wang for helpful discussions and feedback for this work. We are also grateful to Simeng Han for providing us with an updated version of the FOLIO dataset. Our drawing evaluation would not have been possible without our annotators Alex Hu, Ananya Harsh Jha, Belinda Li, Erjia Cao, Ha-na Park, Huirong Wen, Jiangjie Chen, Kabir Swain, Ka Wai Chan, Lucy Li, Simran Swain, Tejas Srinivasan, Tianyu Liu, Yue Bai, Yutaro Yamada, and Ziwei Wei. Zhaofeng would like to thank Jiamin Zhang for the guitar lessons, which were short but helpful for the relevant components of this paper. Figure 1 uses icons from flaticon.com. This study was supported by funds from the MIT–IBM Watson AI Lab, the MIT Quest for Intelligence, and the National Science Foundation under grants IIS-2212310 and IIS-2238240.

References

Appendix A Full Setups

Unless otherwise specified, we use temperature=0 when sampling from the LMs.

We randomly sample 1,000 two-digit addition expressions and evaluate them in bases 8, 9, 10, 11, and 16. Each base is sampled separately—for bases other than base-10, we make sure all expressions evaluate to a different result in that base compared to base-10 so that these expressions discriminate between the bases. To ensure the LMs understand these bases, we design the CCC to ask the model what the number following a given number is. We want the model to know when to carry over and when not to, so we take the 100 smallest numbers in the given basis that ends with the maximum digit in that base, and 100 that end with 0.

A.2 Programming

We use the HumanEval dataset (Chen et al., 2021) which has short Python programs and is commonly used to assess the coding ability of LMs (Bai et al., 2022; Xu et al., 2022; Wang et al., 2023b; i.a.). It was designed as a code-generation dataset, where a model writes a function from a specification and is evaluated against test cases with input-output pairs. Different from our other tasks, we follow prior work (Touvron et al., 2023; Wang et al., 2023b) and (1) use temperature 0.1 when evaluating pass@1 and 0.8 for pass@10, (2) sample 50 responses, and (3) only evaluate without 0-shot CoT. While the original work Chen et al. (2021) recommended sampling 200 responses, this is very expensive, and we follow Wang et al. (2023b) and only sample 50. In Figure 2, we only show the performance on the subset of HumanEval where a 1-based execution of the ground-truth program fails the unit tests. These are the instances that distinguish between 0- and 1-based indexing. We also report results on the full HumanEval dataset in Table 21.

We also consider another setup—code execution, where we give the LM the ground-truth program and ask the LM for the output of the test cases given the input. We remove four programs in HumanEval that are not compatible with this format (ID: 32, 38, 50, and 53), only for this execution task. Because the program would have a different functionality under 1-based indexing, we remove the docstring that is the function description, and also rename the function to the uninformative function, to avoid confusing the LM. Some programs also become invalid under 1-based indexing, specifically, those that perform any indexing using 0. We remove all test cases that involve indexing with 0 and programs that do not have test cases left after this removal. 150 programs and 969 test cases remain. Some of these test cases may not distinguish between 0- and 1-based indexing. So for our main task (i.e., not CCC), we only consider test cases whose outputs are different under 0- vs. 1-based indexing, and there are 113 of them.

Because we use the same prompt to indicate the counterfactual conditions for both code generation and execution, and because we want to maintain comparability with prior work for the former, we only include CCC in the execution setup. We believe they reflect the LMs’ understanding of 1-based indexing in the generation setup too. We ask the LM for the output of simple tests about 1-based indexing such as "qrstu" and "qrs"[:2]. They do not require sophisticated reasoning under the counterfactual conditions and yet are sufficient to discriminate between the default and the counterfactual conditions. We append 5 such checks after each of the 150 programs, totaling 750 CCC.

For the execution task, we do not consider PaLM-2, because it only has a maximum of 1,024 output context length and leads to truncated, unparseable results for most test instances, especially under 0-shot CoT.

A.3 Basic Syntactic Reasoning

We follow Ravfogel et al. (2019) and create synthetic variants of English with all six orderings of the subject, verb, and object. Given a dependency tree of a regular English sentence, we alter the order of subject and object nodes with respect to the corresponding verb. The subtrees rooted at subject or object nodes are moved as a whole, whereas other non-core dependent nodes (e.g., prepositional phrases) are kept in the original positions. We use 100 sentences from English Penn Treebank Marcus et al. (1993), and convert the original phrase-structure trees into Universal Dependencies Nivre et al. (2016) using the Stanford converter Schuster and Manning (2016).

Our task is to identify the main verb and the main subject of a sentence. We only choose sentences where the main subject contains a single word. Ravfogel et al. (2019)’s data generation procedure sometimes results in sentences in the SVO order to be unnatural English sentences. To eliminate this complexity, we retain only sentences whose SVO variant according to Ravfogel et al. (2019)’s data generation procedure is identical to the original English sentence.

We designed the CCC to assess how well LMs understand the instruction that explains the difference of word orders in the counterfactual settings. We synthetically generate 100 simple three-word sentences (e.g., “anna saw john”) in five counterfactual word orders (e.g., “anna john saw” in SOV), and ask LMs to reconstruct the original English sentences in SVO order. Conceptually, this is equivalent to asking the model to identify the subject, verb, and object in the perturbed order, but using a format that is more familiar to the LM.

To generate the simple sentences for the CCC, we designed a simple context-free grammar where the subject and the object are sampled from the vocabulary of person names, and the verb is sampled from the set {\{saw, loves, calls, knows, sees}\}. A key feature of the sentences generated from this approach is their retained plausibility when the subject and object are interchanged. This means that given a counterfactual sentence (e.g., “anna john saw”), there are two natural English sentences as candidates for reconstruction (i.e., “anna saw john” and “john saw anna”). Due to this inherent ambiguity, LMs cannot default to the heuristic of treating the synthetic sentence as bag-of-words and then reconstructing the most natural ordering of those words in real English. The random baseline chooses a random noun as the main subject and a random verb as the main verb.

The results for this task are shown in Table 22. Generally, the models pass our crafted CCC challenge with decent accuracy, but we observed that, in a few cases, the LMs are confused by the reconstruction ambiguity explained above. GPT-3.5 and Claude fail in the OVS settings where they often directly copy the original sentence—e.g., instead of reconstructing “anna saw john” to “john saw anna”, they simply copy the original sentence “anna saw john” as the output. Similarly, PaLM-2 often incorrectly reverses the subject and object in the SOV and VSO settings—e.g., instead of reconstructing “calls tom lucas” to “tom calls lucas”, it outputs “lucas calls tom”.

A.4 Natural Language Reasoning with First-Order Logic

We use the FOLIO dataset (Han et al., 2022) that contains premises most of which are consistent with common sense and are hence amenable to our counterfactual study. We use the full dataset, combining the training and development sets for a total of 1,204 instances, for the logistic regression analysis in §5.1. But for our counterfactual study, automatically altering the premises to violate common sense is not trivial, so one author manually rewrote the premises of a subset of 81 instances to be counterfactual, and another author verified the rewrite. Considering the analysis in §5.1, we chose this subset by including every instance with premises all of which GPT-4 believes to be true and whose conclusion whose GPT-4-believed truth value matches the entailment label.

We explicitly instruct the model to use no common sense or world knowledge (§B), thereby requiring symbolic reasoning. For the CCC, we ask the model if the unaltered or the altered premise is true, when both are presented as options, and expect the latter.

While the FOLIO dataset has a public release, the authors have made subsequent updates which, at the time of this paper, have not been made public. We hence do not release the LM interaction data for this task, and use a fictional example in Table 5.

A.5 Spatial Reasoning

We ask the LM for the coordinates of objects in a room. We randomly sample 100 rooms, each with 3 different objects placed in 3 different cardinal directions specified using unit vectors (out of north (0,−1)(0,-1), south (0,1)(0,1), east (1,0)(1,0), and west (−1,0)(-1,0) as the default conditions). Though using a downward-facing yy-axis as the default condition may be counter-intuitive, it is natural when drawing top-to-bottom and is the convention in most image processing libraries such as OpenCV (Python), Pillow (Python), and Processing (Java, JavaScript, Python), as well as graphic design applications such as Adobe Illustrator. We believe this system is the most often encountered during LM pretraining. However, other libraries with an upward-facing yy-axis also exist, such as matplotlib (Python), ggplot (R), and D3 (JavaScript).

For the counterfactual setting, we alter the direction–unit vector mapping, and ask for the object coordinates in the new system. We consider two direction-swapped worlds (north-south and east-west), three rotated worlds (by 90°, 180°, and 270°), and a randomly permuted world. We evaluate the relative positions of objects and report the instance-level accuracy that requires all 3 objects in a room to be located correctly as the main metric. The random accuracy is around 16.7%.When not considering cases where objects are placed in the same line, there are 24 permutations for placing 3 objects in 4 different directions, of which 4 can be considered correct. We also report the object-level accuracy in Table 24. As the CCC, we make sure that the LM understands the permuted world by asking it to also specify the coordinates of the unit vectors representing the 4 cardinal directions in the output.

A.6 Drawing

We choose 100 objects from five Emojihttps://getemoji.com categories: activity, travel & places, animals & nature, food & drink, and objects. Since LMs cannot generate images at the pixel level, we use code as an intermediate abstraction for sketch generation. We do our best to select objects that are easy to draw using code, verified by multiple authors. We consider the Processing language for our experiment which supports a variety of shapes and colors and is widely used in visualization. Our initial experiments found this language to achieve the best drawing performance compared to other graphics and image processing frameworks, including TikZ, SVG, and matplotlib.

For the counterfactual settings, we ask the LMs to draw the same object, but vertically flipped, or rotated by 90°or 180°. We also ask the LMs to avoid using any transformation functions such as rotate and scale to avoid shortcuts. Before our quantitative evaluation, we flip/rotate back the generated drawing.

We use human evaluation by asking human annotators to determine whether the drawing matches the object. We instruct the annotators to consider orientation as part of correctness and for objects that have a canonical orientation, they must be drawn in that orientation. We average the results over 4 annotators. We also show a breakdown of accuracy depending on whether an object has a canonical orientation or not, as judged by the annotators, in Table 26. In addition, we consider multi-class classification accuracy using CLIP Radford et al. (2021) as an automatic metric, where we ask CLIP to classify the drawing into our 100 categories in a 0-shot fashion. We include the CLIP multi-class classification accuracy in Table 25. We note that the accuracy of the CLIP model for our setup is not guaranteed: first, our generated sketches may be distributionally different from the predominantly photorealistic images in CLIP’s training data; also, CLIP might be insensitive to the object’s orientation, but that distinguishes between our default and counterfactual settings. Therefore, to verify the reliability of this automatic evaluation, we randomly sample 10 objects for each model and for each default/counterfactual setting, and perform human evaluation on the 240 generated images. We find that CLIP’s judgment aligns with human annotators’ 84% of the time, suggesting the reliability of this evaluation.

For this task, we do not consider PaLM-2 due to its limited context length. Our preliminary experiments also found PaLM-2 to struggle in generating parseable Processing code, even in the default setting.

We construct the CCC baseline by requiring the LMs to additionally draw a line at the top of the figure and flip/rotate it as well. A successful flipping/rotation of the line, as judged by the annotators and verified in the generated code if necessary, demonstrates an understanding of the counterfactual world.

A.7 Music

We measure LMs’ abilities to give correct fret placements for ukulele and guitar chords in an existing databasehttps://github.com/tombatossals/chords-dbWe heuristically filter out incorrect datapoints by filtering out chords that either have the wrong number of notes or lack the root note.. We include the following kinds of chords from the database: sus2 (suspended second chord), sus4 (suspended fourth chord), min triad (minor triad), maj triad (major triad), dim7 (diminished seventh chord), aug7 (augmented seventh chord), maj7 (major seventh chord), min7 (minor seventh chord), dom7 (dominant seventh chord), 5 (fifth interval), and 6 (sixth chord).

In the counterfactual setting, we instruct LMs to provide fret placements for a “special” ukulele or guitar where one of the strings is altered. We experiment with perturbations of different sizes: For guitar, we experiment with one-string changes by one note (EADGBE →\rightarrow EBDGBE; EADGBE →\rightarrow FADGBE), one-string changes by two notes (→\rightarrow ECDGBE), and two string changes (→\rightarrow ECFGBE). We also experiment with a one-string change that corresponds to a common alternate tuning of a guitar called drop-D tuning (→\rightarrow DADGBE). For ukulele, we experiment with one-string changes by one note (GCEA →\rightarrow FCEA; →\rightarrow ACEA), one-string change by two notes (→\rightarrow BCEA), and two-string changes by two notes (→\rightarrow BEEA). The generated fret placements for a chord are considered correct if all and only the notes in the corresponding chord (e.g., C, E, G for a C major triad) are produced, irrespective of order.

As the CCC, we assess LMs’ understanding of the given instrument’s strings by asking them to identify what notes a given sequence of frets corresponds to; for the CCC, the sequences are either all fret 0, all fret 1, or all fret 2. We compute CCC accuracy at the fret level (as opposed to the sequence level).

A.7.2 Retrieving Notes of Famous Melodies

For 8 famous melodies, we prompt LMs to retrieve the nn-th note in the melody, where nn is between 1 and 7 (inclusive). In the counterfactual setting, we prompt the LM to do the same but in a different key. The list of melodies and keys we experiment with is below.

We use C Major as the key for songs as the default condition given its popularity for famous melodies like children’s songs. We use other keys as the counterfactual keys.We note that some songs may have multiple canonical keys (e.g., “Twinkle Twinkle Little Star” is also frequently performed in keys like G major or D Major.) In some initial exploration, we validated that C Major was at least one of the canonical keys for the melodies chosen, both by verifying that popular sheet music for these songs was written in C Major, and by asking GPT-3.5 to generate the melodies in an unspecified key and verifying that the generated key was C Major.

As the CCC, we assess LMs’ understanding of the given keys by asking them to retrieve the nn-th note of the scale of the given key.

Twinkle Twinkle Little Star, Mary Had a Little Lamb, Happy Birthday to You, Somewhere Over the Rainbow, Row Row Row Your Boat, Old Macdonald Had a Farm, Itsy Bitsy Spider, London Bridge is Falling Down.

B# major, C# major, Db major, D major, D# major, Eb major, Fb major, E major, E# major, F major, F# major, Gb major, G major, G# major, Ab major, A major, A# major, Bb major, Cb major, B major.

A.8 Chess

We evaluate an LM’s ability to understand chess rules by checking if it can determine whether a 4-move opening follows the rules of chess or not. In the counterfactual setting, we swap the position of bishops and knights on the board and evaluate the same task. For each setting, we randomly sample 400 unique chess openings via a procedural generation algorithm: 200 are legal for the default setting but not for the counterfactual setting, and vice versa for the other 200, ensuring a more balanced and fair classification problem. We represent the moves as the LM input using the PGN format, the standard for chess moves description.

For the CCC, we ask an LM for the starting positions of the four knights and four bishops on the board to make sure it understands the new initial board. For both the default and counterfactual settings, we ask for the positions of white knights, white bishops, black knights, and black bishops, totaling 8 pieces, and evaluate using accuracy. Since concluding the effectiveness of our counterfactual prompt using merely 8 CCC may not be statistically significant, we sample 15 LM responses using temperature=0.1 for asking about each piece.

A.9 SET Game

We synthetically generate SET boards, consisting of 12 cards, each with exactly one 3-card SET that satisfies the game rules in §3.9. We represent each card with a string representation, e.g., (3|open|red|diamond). In preliminary experiments, we tried to ask the LMs to find the SET directly, but found that they cannot perform this task well (see Figure 4(c), “Number of Cards to Find”=3=3). Therefore, in our main evaluation, we expose 2 cards in the SET and ask the LM to identify the missing one that completes the SET.

In the counterfactual setting, we invert the rule for the number attribute to require that two cards in the SET should have the same number but the other card should be different. For the CCC, we ask the model to verify the validity of a given SET instead of finding it. In each CCC instance, we either give a valid SET from the board, or 3 randomly sampled cards that do not constitute a valid SET. We ask the model to classify whether the given combination is valid or invalid. We note that our counterfactual perturbation ensures that the each SET cannot be simultaneously valid in the default setting and the counterfactual setting, and hence this CCC is discriminative between the two settings.

Appendix B Prompts

We provide the exact prompts that we used to query the LMs in Tables 1 to 17. For clarity, we give a concrete prompt that embeds a test instance, rather than a template. We explain minor design decisions in the respective captions. We do not use the system message field for any model.

Appendix C Raw Results

We show the numeric results in Tables 18 to 34.