Abductive Commonsense Reasoning

Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen-tau Yih, Yejin Choi

Introduction

The brain is an abduction machine, continuously trying to prove abductively that the observables in its environment constitute a coherent situation. – Jerry Hobbs, ACL 2013 Lifetime Achievement AwardThe full transcript of his award speech is available at https://www.mitpressjournals.org/doi/full/10.1162/COLI_a_00171

Abductive reasoning is inference to the most plausible explanation for incomplete observations (Peirce, 1965a). Figure 1 illustrates an example. Given the incomplete observations about the world that O1O_{1}: “Jenny cleaned her house and went to work, leaving the window just a crack open.” and sometime later O2O_{2}: “When Jenny returned home, she saw her house was a mess.”, we can hypothesize different potential explanations and reason about which is the most likely. We can readily rule out H3H_{3} since it fails to justify the observation O2O_{2}. While H1H_{1} and H2H_{2} are both plausible, the most likely explanation based on commonsense is H1H_{1} as H2H_{2} is somewhat implausible given O1O_{1}.

One crucial observation Peirce makes about abductive reasoning is that abduction is “the only logical operation which introduces any new ideas”, which contrasts with other types of inference such as entailment, that focuses on inferring only such information that is already provided in the premise. Abductive reasoning has long been considered to be at the core of understanding narratives (Hobbs et al., 1988), reading between the lines (Norvig, 1987; Charniak & Shimony, 1990), reasoning about everyday situations (Peirce, 1965b; Andersen, 1973), and counterfactual reasoning (Pearl, 2002; Pearl & Mackenzie, 2018). Despite the broad recognition of its importance, however, the study of abductive reasoning in narrative text has very rarely appeared in the NLP literature, in large part because most previous work on abductive reasoning has focused on formal logic, which has proven to be too rigid to generalize to the full complexity of natural language.

In this paper, we present the first study to investigate the viability of language-based abductive reasoning. This shift from logic-based to language-based reasoning draws inspirations from a significant body of work on language-based entailment (Bowman et al., 2015; Williams et al., 2018b), language-based logic (Lakoff, 1970; MacCartney & Manning, 2007), and language-based commonsense reasoning (Mostafazadeh et al., 2016; Zellers et al., 2018). In particular, we investigate the use of natural language as the representation medium, and probe deep neural models on language-based abductive reasoning.

More concretely, we propose Abductive Natural Language Inference (α\alphaNLI) and Abductive Natural Language Generation (α\alphaNLG) as two novel reasoning tasks in narrative contexts.α\alphaNLI and α\alphaNLG are pronounced as alpha-NLI and alpha-NLG, respectively We formulate α\alphaNLI as a multiple-choice task to support easy and reliable automatic evaluation: given a context, the task is to choose the more likely explanation from a given pair of hypotheses choices. We also introduce a new challenge dataset, ART, that consists of 20K narratives accompanied by over 200K explanatory hypothesis. ART: Abductive Reasoning in narrative Text.Data available to download at http://abductivecommonsense.xyz We then establish comprehensive baseline performance based on state-of-the-art NLI and language models. The best baseline for α\alphaNLI based on BERT achieves 68.9% accuracy, with a considerable gap compared to human performance of 91.4%(§5.2). The best generative model, based on GPT2, performs well below human performance on the α\alphaNLG task (§5.2). Our analysis leads to insights into the types of reasoning that deep pre-trained language models fail to perform — despite their strong performance on the closely related but different task of entailment NLI — pointing to future research directions.

Task Definition

We formulate α\alphaNLI as multiple choice problems consisting of a pair of observations as context and a pair of hypothesis choices. Each instance in ART is defined as follows:

O1O_{1}: The observation at time t1t_{1}.

O2O_{2}: The observation at time t2>t1t_{2}>t_{1}.

h+h^{+}: A plausible hypothesis that explains the two observations O1O_{1} and O2O_{2}.

h−h^{-}: An implausible (or less plausible) hypothesis for observations O1O_{1} and O2O_{2}.

Given the observations and a pair of hypotheses, the α\alphaNLI task is to select the most plausible explanation (hypothesis).

Abductive Natural Language Generation

α\alphaNLG is the task of generating a valid hypothesis h+h^{+} given the two observations O1O_{1} and O2O_{2}. Formally, the task requires to maximize P(<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><msup><mi>h</mi><molspace="0em"rspace="0em">+</mo></msup></mrow><annotationencoding="application/x−tex">h+</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.8213em;"></span><spanclass="mord"><spanclass="mordmathnormal">h</span><spanclass="msupsub"><spanclass="vlist−t"><spanclass="vlist−r"><spanclass="vlist"style="height:0.8213em;"><spanstyle="top:−3.113em;margin−right:0.05em;"><spanclass="pstrut"style="height:2.7em;"></span><spanclass="sizingreset−size6size3mtight"><spanclass="mordmtight"><spanclass="mordmtight">+</span></span></span></span></span></span></span></span></span></span></span></span></span>∣P(<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><msup><mi>h</mi><mo lspace="0em" rspace="0em">+</mo></msup></mrow><annotation encoding="application/x-tex">h^{+}</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8213em;"></span><span class="mord"><span class="mord mathnormal">h</span><span class="msupsub"><span class="vlist-t"><span class="vlist-r"><span class="vlist" style="height:0.8213em;"><span style="top:-3.113em;margin-right:0.05em;"><span class="pstrut" style="height:2.7em;"></span><span class="sizing reset-size6 size3 mtight"><span class="mord mtight"><span class="mord mtight">+</span></span></span></span></span></span></span></span></span></span></span></span></span>|O1O_{1}, O2O_{2})).

Models for Abductive Commonsense Reasoning

A distinct feature of the α\alphaNLI task is that it requires jointly considering all available observations and their commonsense implications, to identify the correct hypothesis. Formally, the α\alphaNLI task is to select the hypothesis h∗h^{*} that is most probable given the observations.

Rewriting the objective using Bayes Rule conditioned on O1O_{1}, we have:

We formulate a set of probabilistic models for α\alphaNLI that make various independence assumptions on Equation 2 – starting from a simple baseline that ignores the observations entirely, and building up to a fully joint model. These models are depicted as Bayesian Networks in Figure 2.

Hypothesis Only:

Our simplest model makes the strong assumption that the hypothesis is entirely independent of both observations, i.e. (H⊥O1,O2)(H\perp O_{1},O_{2}), in which case we simply aim to maximize the marginal P(H)P(H).

First (or Second) Observation Only:

Our next two models make weaker assumptions: that the hypothesis depends on only one of the first O1O_{1} or second O2O_{2} observation.

Linear Chain:

Our next model uses both observations, but considers each observation’s influence on the hypothesis independently, i.e. it does not combine information across the observations. Formally, the model assumes that the three variables ⟨O1,H,O2⟩\langle O_{1},H,O_{2}\rangle form a linear Markov chain, where the second observation is conditionally independent of the first, given the hypothesis (i.e. (O1⊥O2∣H)(O_{1}\perp O_{2}|H)). Under this assumption, we aim to maximize a somewhat simpler objective than Equation 2:

Fully Connected:

Finally, our most sophisticated model jointly models all three random variables as in Equation 2, and can in principle combine information across both observations to choose the correct hypothesis.

To help illustrate the subtle distinction between how the Linear Chain and Fully Connected models consider both observations, consider the following example. Let observation O1O_{1}: “Carl went to the store desperately searching for flour tortillas for a recipe.” and O2O_{2}: “Carl left the store very frustrated.”. Then consider two distinct hypotheses, an incorrect h1h^{1}: “The cashier was rude” and the correct h2h^{2}: “The store had corn tortillas, but not flour ones.”. For this example, a Linear Chain model could arrive at the wrong answer, because it reasons about the observations separately—taking O1O_{1} in isolation, both h1h^{1} and h2h^{2} seem plausible next events, albeit each a priori unlikely. And for O2O_{2} in isolation—i.e. in the absence of O1O_{1}, as for a randomly drawn shopper—the h1h^{1} explanation of a rude cashier seems a much more plausible explanation of Carl’s frustration than are the details of the store’s tortilla selection. Combining these two separate factors leads the Linear Chain to select h1h^{1} as the more plausible explanation. It is only by reasoning about Carl’s goal in O1O_{1} jointly with his frustration in O2O_{2}, as in the Fully Connected model, that we arrive at the correct answer h2h^{2} as the more plausible explanation.

In our experiments, we encode the different independence assumptions in the best performing neural network model. For the hypothesis-only and single observation models, we can enforce the independencies by simply restricting the inputs of the model to only the relevant variables. On the other hand, the Linear Chain model takes all three variables as input, but we restrict the form of the model to enforce the conditional independence. Specifically, we learn a discriminative classifier:

where ϕ\phi and ϕ′\phi^{\prime} are neural networks that produce scalar values.

2 Abductive Natural Language Generation

Given h+h^{+}={w1h…wlh}=\{w^{h}_{1}\ldots w^{h}_{l}\}, O1O_{1}={w1o1…wmo1}\{w^{o1}_{1}\ldots w^{o1}_{m}\} and O2O_{2}={w1o2…wno2}\{w^{o2}_{1}\ldots w^{o2}_{n}\} as sequences of tokens, the α\alphaNLG task can be modeled as P(\text{h^{+}}|\text{O_{1}{}, \text{O_{2}{}}})=\prod P(w^{h}_{i}|w^{h}_{<i},w^{o1}_{1}\ldots w^{o1}_{m},w^{o2}_{1}\ldots w^{o2}_{n}) Optionally, the model can also be conditioned on background knowledge K\mathcal{K}. Parameterized models can then be trained to minimize the negative log-likelihood over instances in ART:

ART Dataset: Abductive Reasoning in Narrative Text

ART is the first large-scale benchmark dataset for studying abductive reasoning in narrative texts. It consists of ∼{\sim}20K narrative contexts (pairs of observations ⟨\langle{}O1O_{1}, O2O_{2}⟩\rangle{}) with over 200K explanatory hypotheses. Table 6 in the Appendix summarizes corpus-level statistics of the ART dataset.We will publicly release the ART dataset upon acceptance. Figure 4 shows some illustrative examples from ART (dev split). The best model based on BERT fails to correctly predict the first two dev examples.

The pairs O1O_{1}, O2O_{2} in ART are drawn from the ROCStories dataset (Mostafazadeh et al., 2016). ROCStories is a large collection of short, manually curated five-sentence stories. It was designed to have a clear beginning and ending for each story, which naturally map to the first (O1O_{1}) and second (O2O_{2}) observations in ART.

Collecting Hypotheses Options:

We crowdsourced the plausible and implausible hypotheses options on Amazon Mechanical Turk (AMT) in two separate tasksBoth crowdsourcing tasks are complex and require creative writing. Along with the ART dataset, we will publicly release templates and the full set of instructions for all crowdsourcing tasks to facilitate future data collection and research in this direction.:

Plausible Hypothesis Options: We presented O1O_{1} and O2O_{2} as narrative context to crowdworkers who were prompted to fill in “What happened in-between?” in natural language. The design of the task motivates the use of abductive reasoning to hypothesize likely explanations for the two given observations.

Implausible Hypothesis Options: In this task, we presented workers with observations O1O_{1}, O2O_{2} and one plausible hypothesis option h+h^{+} ∈H+\in{}\mathcal{H}^{+} collected from the previous task. Crowdworkers were instructed to make minimal edits (up to 5 words) to a given h+h^{+} to create implausible hypothesis variations for each plausible hypothesis.

A significant challenge in creating datasets is avoiding annotation artifacts – unintentional patterns in the data that leak information about the target label – that several recent studies (Gururangan et al., 2018; Poliak et al., 2018; Tsuchiya, 2018) have reported on crowdsourced datasets . To tackle this challenge, we collect multiple plausible and implausible hypotheses for each ⟨\langle{}O1O_{1}, O2O_{2}⟩\rangle{} pair (as described above) and then apply an adversarial filtering algorithm to retain one challenging pair of hypotheses that are hard to distinguish between. We describe our algorithm in detail in Appendix A.5. While our final dataset uses BERT as the adversary, preliminary experiments that used GPT as an adversary resulted in similar drops in performance of all models, including all BERT variants. We compare the results of the two adversaries in Table 1.

Experiments and Results

We now present our evaluation of finetuned state-of-the-art pre-trained language models on the ART dataset, and several other baseline systems for both α\alphaNLI and α\alphaNLG. Since α\alphaNLI is framed as a binary classification problem, we choose accuracy as our primary metric. For α\alphaNLG, we report performance on automated metrics such as BLEU (Papineni et al., 2002), CIDEr (Vedantam et al., 2015), METEOR (Banerjee & Lavie, 2005) and also report human evaluation results.

Despite strong performance on several other NLP benchmark datasets, the best baseline model based on BERT achieves an accuracy of just 68.9% on ART compared to human performance of 91.4%. The large gap between human performance and that of the best system provides significant scope for development of more sophisticated abductive reasoning models. Our experiments show that introducing the additional independence assumptions described in Section 3.1 over the fully connected model tends to degrade system performance (see Table 1) in general.

We compute human performance using AMT. Each instance (two observations and two hypothesis choices) is shown to three workers who were prompted to choose the more plausible hypothesis choice.Additional crowdsourcing details in the Appendix A.1 We compute majority vote on the labels assigned which leads to a human accuracy of 91.4% on the ART test set.

Baselines

We include baselines that rely on simple features to verify that ART is not trivially solvable due to noticeable annotation artifacts, observed in several crowdsourced datasets. The accuracies of all simple baselines are close to chance-performance on the task – indicating that the dataset is free of simple annotation artifacts.

A model for the related but distinct task of entailment NLI (e.g. SNLI) forms a natural baseline for α\alphaNLI. We re-train the ESIM+ELMo (Chen et al., 2017; Peters et al., 2018) model as its performance on entailment NLI (88.9%88.9\%) is close to state-of-the-art models (excluding pre-trained language models). This model only achieves an accuracy of 58.8%58.8\% highlighting that performing well on ART requires models to go far beyond the linguistic notion of entailment.

Pre-trained Language Models

BERT (Devlin et al., 2018) and GPT (Radford, 2018) have recently been shown to achieve state-of-the-art results on several NLP benchmarks (Wang et al., 2018). We finetune both BERT-Large and GPT as suggested in previous work and we present each instance in their natural narrative order. BERT-ft (fully connected) is the best performing model achieving 68.9% accuracy, compared to GPT’s 63.1%63.1\%.The input format for the GPT model and BERT variants is described in the Appendix A.4. Our AF approach was able to reduce BERT performance from over 88%88\% by 2020 points.

Learning Curve and Dataset Size

While there is enough scope for considerably scaling up the dataset based on ROCStories, the learning curve in Figure 5 shows that the performance of the best model plateaus after ∼10,000{\sim}10,000 instances. In addition, there is still a wide gap (∼23%{\sim}23\%) between the performance of the best model and human performance.

GPT Adversary

Table 1 also includes results of our experiments where GPT was used as the adversary. Notably, in this case, adversarially filtering the dataset brings down GPT performance under 53%. On the other hand, the best BERT model, that encodes the fully connected bayesian network performs significantly better than the BERT model that encodes the linear chain assumptions – 72% compared to 65%. Therefore, we use the BERT fully connected model as the adversary in ART. The gap between the linear chain and fully connected BERT models diminishes when BERT is used as an adversary – in spite of being a more powerful model – which indicates that adversarial filtering disproportionately impacts the model used as the adversary. However, the dataset also becomes more difficult for the other models that were not used as adversaries. For example, before any filtering, BERT scores 88%88\% and OpenGPT gets 80%80\%, which is much higher than either model achieves in Table 1 when the other model is used for filtering. This result is a reasonable indicator, albeit not a guarantee, that ART will remain challenging for new models released in the future.

2 Abductive Natural Language Generation

As described in Equation 4, we train GPT2 conditioned on the tokens of the two observations O1O_{1} and O2O_{2}. Both observations are enclosed with field-specific tags. ATOMIC (Sap et al., 2019), a repository of inferential if-then knowledge is a natural source of background commonsense required to reason about narrative contexts in ART. Yet, there is no straightforward way to include such knowledge into a neural model as ATOMIC’s nodes are not canonicalized and are represented as short phrases of text. Thus, we rely on COMeT – a transformer model trained on ATOMIC that generates nine commonsense inferences of events in natural language.Please see Appendix A.6 for a full list of the nine relations. Specifically, we experiment with two ways of integrating information from COMeT in GPT2: (i) as textual phrases, and (ii) as embeddings.

Figure 3 shows how we integrate COMeT representations. Concretely, after the input tokens are embedded by the word-embedding layer, we append eighteen (corresponding to nine relations for each observation) embeddings to the sequence before passing through the layers of the Transformer architecture. This allows the model to learn each token’s representation while attending to the COMeT embeddings – effectively integrating background commonsense knowledge into a language model.We describe the format of input for each model in Appendix A.7.

Discussion

Table 2 reports results on the α\alphaNLG task. Among automatic metrics, we report BLEU-4 (Papineni et al., 2002), METEOR (Banerjee & Lavie, 2005), ROUGE (Lin, 2004), CIDEr (Vedantam et al., 2015) and BERT-Score (Zhang et al., 2019) (with the bert-base-uncased model). We establish human performance through crowdsourcing on AMT. Crowdworkers are shown pairs of observations and a generated hypothesis and asked to label whether the hypothesis explains the given observations. The last column reports the human evaluation score. The last row reports the score of a held-out human-written hypothesis and serves as a ceiling for model performance. Human-written hypotheses are found to be correct for 96% of instances, while our best generative models, even when enhanced with background commonsense knowledge, only achieve 45% – indicating that the α\alphaNLG generation task is especially challenging for current state-of-the-art text generators.

Analysis

We investigate the categories of commonsense-based abductive reasoning that are challenging for current systems and the ones where the best model over-performs. While there have been previous attempts to categorize commonsense knowledge required for entailment (LoBue & Yates, 2011; Clark et al., 2007), crowdsourcing this task at scale with high fidelity and high agreement across annotators remains challenging. Instead, we aim to probe the model with soft categories identified by matching lists of category-specific keywords to the hypothesis choices.

Table 3 shows the accuracy of the best model (BERT-ft) across various categories of commonsense knowledge. BERT-ft significantly underperforms on instances involving Numerical (56.8%56.8\%) and Spatial (65.4%65.4\%) commonsense. These two categories include reasoning about numerical quantities and the spatial location of agents and objects, and highlight some of the limitations of the language models. In contrast, it significantly overperforms on the Emotional category (72.6%72.6\%) where the hypotheses exhibit strong textual cues about emotions and sentiments.

Implausible transitions

A model for an instance of the ART dataset should discard implausible hypotheses in the context of the two given observations. In narrative contexts, there are three main reasons for an implausible hypothesis to be labeled as such:

O_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>↛</mo></mrow><annotation encoding="application/x-tex">\not\to</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">→</span></span></span></span></span>h^{-}: h−h^{-} is unlikely to follow after the first observation O1O_{1}.

h^{-}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>↛</mo></mrow><annotation encoding="application/x-tex">\not\to</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">→</span></span></span></span></span>O_{2}: h−h^{-} is plausible after O1O_{1} but unlikely to precede the second observation O2O_{2}.

Plausible: ⟨\langleO1O_{1}, h−h^{-}, O2O_{2}⟩\rangle is a coherent narrative and forms a plausible alternative, but it is less plausible than ⟨\langleO1O_{1}, h+h^{+}, O2O_{2}⟩\rangle.

We analyze the prevalence of each of these reasons in ART. We design a crowdsourcing task in which we show the implausible option along with the narrative context ⟨\langleO1O_{1}, O2O_{2}⟩\rangle and get labels for which transition (O_{1}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>↛</mo></mrow><annotation encoding="application/x-tex">\not\to</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">→</span></span></span></span></span>h^{-}, h^{-}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>↛</mo></mrow><annotation encoding="application/x-tex">\not\to</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">→</span></span></span></span></span>O_{2} or neither) in the narrative chain is broken. Table 4 shows the proportion of each category from a subset of 1,0001,000 instances from the test set. While h^{-}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>↛</mo></mrow><annotation encoding="application/x-tex">\not\to</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="mrel"><span class="mord vbox"><span class="thinbox"><span class="rlap"><span class="strut" style="height:0.8889em;vertical-align:-0.1944em;"></span><span class="inner"><span class="mord"><span class="mrel"></span></span></span><span class="fix"></span></span></span></span></span><span class="mspace nobreak"></span><span class="mrel">→</span></span></span></span></span>O_{2} accounts for almost half of the implausible transitions in ART, all three categories are substantially present in the dataset. BERT performance on each of these categories indicates that the model finds it particularly hard when the narrative created by the incorrect hypothesis is plausible, but less plausible than the correct hypothesis. On that subset of the test set, the fully connected model performs better than the linear chain model where it is important to consider both observations jointly to arrive at the more likely hypothesis.

2 α𝛼\alphaNLG

Figure 6 shows some examples of generations from the trained models compared to human-written generations. The example on the left is an example of an instance that only humans could get correct, while for the one on the right, COMeT-Emb+GPT2also generates the correct explanation for the observations.

Transfer Learning from ART

ART contains a large number of questions for the novel abductive reasoning task. In addition to serving as a benchmark, we investigate if ART can be used as a resource to boost performance on other commonsense tasks. We apply transfer learning by first training a model on ART, and subsequently training on four target datasets – WinoGrande Sakaguchi et al. (2020), WSC Levesque et al. (2011), DPR Rahman & Ng (2012) and HellaSwag Zellers et al. (2019). We show that compared to a model that is only trained on the target dataset, a model that is sequentially trained on ART first and then on the target dataset can perform better. In particular, pre-training on ART consistently improves performance on related datasets when they have relatively few training examples.

On the other hand, for target datasets with large amounts of training data, pre-training on ART does not provide a significant improvement.

Related Work

Since abduction is fundamentally concerned with plausible chains of cause-and-effect, our work draws inspiration from previous works that deal with narratives such as script learning (Schank & Abelson, 1975) and the narrative cloze test (Chambers & Jurafsky, 2009; Jans et al., 2012; Pichotta & Mooney, 2014; Rudinger et al., 2015). Rather than learning prototypical scripts or narrative chains, we instead reason about the most plausible events conditioned on observations. We make use of the ROCStories dataset (Mostafazadeh et al., 2016), which was specifically designed for the narrative cloze task. But, instead of reasoning about plausible event sequences, our task requires reasoning about plausible explanations for narrative omissions.

Entailment vs. Abductive Reasoning

The formulation of α\alphaNLI is closely related to entailment NLI, but there are two critical distinctions that make abductive reasoning uniquely challenging. First, abduction requires reasoning about commonsense implications of observations (e.g., if we observe that the “grass is wet”, a likely hypothesis is that “it rained earlier”) which go beyond the linguistic notion of entailment (also noted by Josephson (2000)). Second, abduction requires non-monotonic reasoning about a set of commonsense implications collectively, to check the potential contradictions against multiple observations and to compare the level of plausibility of different hypotheses. This makes abductive reasoning distinctly challenging compared to other forms of reasoning such as induction and deduction (Shank, 1998). Perhaps more importantly, abduction is closely related to the kind of reasoning humans perform in everyday situations, where information is incomplete and definite inferences cannot be made.

Generative Language Modeling

Recent advancements in the development of large-scale pre-trained language models (Radford, 2018; Devlin et al., 2018; Radford et al., 2019) have improved the quality and coherence of generated language. Although these models have shown to generate reasonably coherent text when condition on a sequence of text, our experiments highlight the limitations of these models to 1) generate language non-monotonically and 2) adhere to commonsense knowledge. We attempt to overcome these limitations with the incorporation of a generative commonsense model during hypothesis generation.

Related Datasets

Our new resource ART complements ongoing efforts in building resources for natural language inference (Dagan et al., 2006; MacCartney & Manning, 2009; Bowman et al., 2015; Williams et al., 2018a; Camburu et al., 2018). Existing datasets have mostly focused on textual entailment in a deductive reasoning set-up (Bowman et al., 2015; Williams et al., 2018a) and making inferences about plausible events (Maslan et al., 2015; Zhang et al., 2017). In their typical setting, these datasets require a system to deduce the logically entailed consequences of a given premise. In contrast, the nature of abduction requires the use of commonsense reasoning capabilities, with less focus on lexical entailment. While abductive reasoning has been applied to entailment datasets (Raina et al., 2005), they have been applied in a logical theorem-proving framework as an intermediate step to perform textual entailment – a fundamentally different task than α\alphaNLI.

Conclusion

We present the first study that investigates the viability of language-based abductive reasoning. We conceptualize and introduce Abductive Natural Language Inference (α\alphaNLI) – a novel task focused on abductive reasoning in narrative contexts. The task is formulated as a multiple-choice question-answering problem. We also introduce Abductive Natural Language Generation (α\alphaNLG) – a novel task that requires machines to generate plausible hypotheses for given observations. To support these tasks, we create and introduce a new challenge dataset, ART, which consists of 20,000 commonsense narratives accompanied with over 200,000 explanatory hypotheses. In our experiments, we establish comprehensive baseline performance on this new task based on state-of-the-art NLI and language models, which leads to 68.9% accuracy with a considerable gap with human performance (91.4%). The α\alphaNLG task is significantly harder – while humans can write a valid explanation 96% of times, the best generator models can only achieve 45%. Our analysis leads to new insights into the types of reasoning that deep pre-trained language models fail to perform – despite their strong performance on the closely related but different task of entailment NLI – pointing to interesting avenues for future research. We hope that ART will serve as a challenging benchmark for future research in language-based abductive reasoning and the α\alphaNLI and α\alphaNLG tasks will encourage representation learning that enables complex reasoning capabilities in AI systems.

Acknowledgments

We thank the anonymous reviewers for their insightful feedback. This research was supported in part by NSF (IIS-1524371), the National Science Foundation Graduate Research Fellowship under Grant No. DGE 1256082, DARPA CwC through ARO (W911NF15-1- 0543), DARPA MCS program through NIWC Pacific (N66001-19-2-4031), and the Allen Institute for AI. Computations on beaker.org were supported in part by credits from Google Cloud.

References

Appendix A Appendices

We describe the crowdsourcing details of our data collection method.

In this task, participants were presented an incomplete three-part story, which consisted of the first observation (O1O_{1}) and the second observation (O2O_{2}) of the story. They were then asked to complete the story by writing a probable middle sentence that explains why the second observation should follow after the first one. We instructed participants to make sure that the plausible middle sentence (1) is short (fewer than 10 words) and (2) simple as if narrating to a child, (3) avoids introducing any extraneous information, and (4) uses names instead of pronouns (e.g., he/she) wherever possible.

All participants were required to meet the following qualification requirements: (1) their location is in the US, (2) HIT approval rate is greater than 95(%), and (3) Number of HITs approved is greater than 5,000. The reward of this task was set to be 0.07perquestion(0.07 per question (14/hour in average), and each HIT was assigned to five different workers (i.e., 5-way redundancy).

Task 2 - Implausible Hypothesis Options

In this task, participants were presented a three-part story, which consisted of the first observation (O1O_{1}), a middle sentence (h+h^{+}) collected in Task 1, and the second observation (O2O_{2}) of the story. They were then asked to rewrite the middle sentence (h+h^{+}) with minimal changes, so that the story becomes unlikely, implausible or inconsistent (h−h^{-}). We asked participants to add or remove at most four words to h+h^{+}, while ensuring that the new middle sentence is grammatical. In addition, we asked them to stick to the context in the given story. For example, if the story talks about “doctors”, they are welcome to talk about “health” or “diagnosis”, but not mention “aliens”. Finally, we also asked workers to verify if the given middle (h+h^{+}) makes a plausible story, in order to confirm the plausibility of h+h^{+}collected in Task 1.

With respect to this task’s qualification, participants were required to fulfill the following requirements: (1) their location is the US or Canada, (2) HIT approval rate is greater than or equal to 99(%), and (3) number of HITs approved is greater than or equal to 10,00010,000. Participants were paid 0.1perquestion(0.1 per question (14/hour in average), and each HIT was assigned to three different participants (i.e., 3-way redundancy).

Task 3 - α𝛼\alphaNLI Human Performance

Human performance was evaluated by asking participants to answer the α\alphaNLI questions. Given a narrative context ⟨\langleO1O_{1}, O2O_{2}⟩\rangle and two hypotheses, they were asked to choose the more plausible hypothesis. They were also allowed to choose “None of the above“ when neither hypothesis was deemed plausible.

We asked each question to seven participants with the following qualification requirements: (1) their location is either in the US, UK, or Canada, (2) HIT approval rate is greater than 98(%), (3) Number of HITs approved is greater than 10,00010,000. The reward was set to $0.05 per HIT. We took the majority vote among the seven participants for every question to compute human performance.

A.2 ART Data Statistics

Table 6 shows some statistics of the ART dataset.

A.3 Fine-tuning BERT

We fine-tuned the BERT model using a grid search with the following set of hyper-parameters:

learning rate: {\{1e-5, 2e-5, 3e-5, 5e-5}\}

The warmup proportion was set to 0.20.2, and cross-entropy was used for computing the loss. The best performance was obtained with a batch size of 44, learning rate of 5e-5, and number of epochs equal to 1010. Table 7 describes the input format for GPT and BERT (and its variants).

A.4 Baselines

The SVM classifier is trained on simple features like word length, overlap and sentiment features to select one of the two hypothesis choices. The bag-of-words baseline computes the average of GloVe (Pennington et al., 2014) embeddings for words in each sentence to form sentence embeddings. The sentence embeddings in a story (two observations and a hypothesis option) are concatenated and passed through fully-connected layers to produce a score for each hypothesis. The accuracies of both baselines are close to 50% (SVM: 50.6; BOW: 50.5).

Specifically, we train an SVM classifier and a bag-of-words model using GLoVE embeddings. Both models achieve accuracies close to 50%. An Infersent (Conneau et al., 2017) baseline that uses sentences embedded by max-pooling over Bi-LSTM token representations achieves only 50.8% accuracy.

A.5 Adversarial Filtering of Hypotheses Choices

Given an observation pair and sets of plausible and implausible hypotheses ⟨\langleO1O_{1}, O2O_{2}, H+\mathcal{H}^{+}, H−⟩\mathcal{H}^{-}\rangle, our adversarial filtering algorithm selects one plausible and one implausible hypothesis ⟨\langleO1O_{1}, O2O_{2}, h+h^{+}, h−h^{-} ⟩\rangle such that h+h^{+} and h−h^{-} are hard to distinguish between. We make three key improvements over the previously proposed Adversarial Filtering (AF) approach in Zellers et al. (2018). First, Instead of a single positive sample, we exploit a pool H+\mathcal{H}^{+} of positive samples to choose from (i.e. plausible hypotheses). Second, Instead of machine generated distractors, the pool H−\mathcal{H}^{-} of negative samples (i.e. implausible hypotheses) is human-generated. Thus, the distractors share stylistic features of the positive samples as well as that of the context (i.e. observations O1O_{1} and O2O_{2}) – making the negative samples harder to distinguish from positive samples. Finally, We use BERT (Devlin et al., 2018) as the adversary and introduce a temperature parameter that controls the maximum number of instances that can be modified in each iteration of AF. In later iterations, fewer instances get modified resulting in a smoother convergence of the AF algorithm (described in more detail below).

Algorithm 1 provides a formal description of our approach. In each iteration ii, we train an adversarial model MiM_{i} on a random subset Ti\mathcal{T}_{i} of the data and update the validation set Vi\mathcal{V}_{i} to make it more challenging for MiM_{i}. For a pair (hk+,hk−)(h^{+}_{k},h^{-}_{k}) of plausible and implausible hypotheses for an instance kk, we denote δ=ΔMi(hk+,hk−)\delta{}=\Delta_{M_{i}}(h^{+}_{k},h^{-}_{k}) the difference in the model evaluation of hk+h^{+}_{k} and hk−h^{-}_{k}. A positive value of δ\delta{} indicates that the model MiM_{i} favors the plausible hypothesis hk+h^{+}_{k} over the implausible one hk−h^{-}_{k}. With probability tit_{i}, we update instance kk that MiM_{i} gets correct with a pair (h+,h−)∈Hk+×Hk−(h^{+},h^{-})\in\mathcal{H}^{+}_{k}\times\mathcal{H}^{-}_{k} of hypotheses that reduces the value of δ\delta{}, where Hk+\mathcal{H}^{+}_{k} (resp. Hk−\mathcal{H}^{-}_{k}) is the pool of plausible (resp. implausible) hypotheses for instance kk .

We ran AF for 5050 iterations and the temperature tit_{i} follows a sigmoid function, parameterized by the iteration number, between ts=1.0t_{s}=1.0 and te=0.2t_{e}=0.2. Our final dataset, ART, is generated using BERT as the adversary in Algorithm 1.

A.6 ATOMIC Relations

ATOMIC (Sap et al., 2019) represents commonsense knowledge as a graph with events are nodes and the following nine relations as edges:

xNeed: What does X need to do before the event?

xEffect: What effects does the event have on X?

xWant: What would X likely want to do after the event?

xReaction: How does X feel after the event?

oReact: How do others’ feel after the event?

oWant: What would others likely want to do after the event?

oEffect: What effects does the event have on others?

A.7 Generation Models Input Format

Table 8 describes the format of input to each variation of the generative model evaluated.