Improved Logical Reasoning of Language Models via Differentiable Symbolic Programming

Hanlin Zhang, Jiani Huang, Ziyang Li, Mayur Naik, Eric Xing

Introduction

Complex applications in natural language processing involve dealing with two separate challenges. On one hand, there is the richness, nuances, and extensive vocabulary of natural language. On the other hand, one needs logical connectives, long reasoning chains, and domain-specific knowledge to draw logical conclusions. The systems handling these two challenges are complementary to each other and are likened to psychologist Daniel Kahneman’s human “system 1” and “system 2” Kahneman (2011): while the former makes fast and intuitive decisions, akin to neural networks, the latter thinks more rigorously and methodically. Considering LMs as “system 1” and symbolic reasoners as “system 2”, we summarize their respective advantages in Table 1.

Although pre-trained LMs have demonstrated remarkable predictive performance, making them an effective “system 1”, they fall short when asked to perform consistent logical reasoning (Kassner et al., 2020; Helwe et al., 2021; Creswell et al., 2022), which usually requires “system 2”. In part, this is because LMs largely lack capabilities of systematic generalization (Elazar et al., 2021; Hase et al., 2021; Valmeekam et al., 2022).

In this work, we seek to incorporate deductive logical reasoning with LMs. Our approach has the same key objectives as neuro-symbolic programming (Chaudhuri et al., 2021): compositionality, consistency, interpretability, and easy integration of prior knowledge. We present DSR-LM, which tightly integrates a differentiable symbolic reasoning module with pre-trained LMs in an end-to-end fashion. With DSR-LM, the underlying LMs govern the perception of natural language and are fine-tuned to extract relational triplets with only weak supervision. To overcome a common limitation of symbolic reasoning systems, the reliance on human-crafted logic rules (Huang et al., 2021; Nye et al., 2021), we adapt DSR-LM to induce and fine-tune rules automatically. Further, DSR-LM allows incorporation of semantic loss obtained by logical integrity constraints given as prior knowledge, which substantially helps the robustness.

We conduct extensive experiments showing that DSR-LM can consistently improve the logical reasoning capability upon pre-trained LMs. Even if DSR-LM uses a RoBERTa backbone with much less parameters and does not explicitly take triplets as supervision, it can still outperform various baselines by large margins. Moreover, we show that DSR-LM can induce logic rules that are amenable to human understanding to explain decisions given only higher-order predicates. As generalization over long-range dependencies is a significant weakness of transformer-based language models (Lake and Baroni, 2018; Tay et al., 2020), we highlight that in systematic, long-context scenarios, where most pre-trained or neural approaches fail to generalize compositionally, DSR-LM can still achieve considerable performance gains.

Related Work

Logical reasoning with LMs. Pre-trained LMs have been shown to struggle with logical reasoning over factual knowledge (Kassner et al., 2020; Helwe et al., 2021; Talmor et al., 2020a). There is encouraging recent progress in using transformers for reasoning tasks (Zhou et al., 2020; Clark et al., 2021; Wei et al., 2022; Chowdhery et al., 2022; Zelikman et al., 2022) but these approaches usually require a significant amount of computation for re-training or human annotations on reasoning provenance (Camburu et al., 2018; Zhou et al., 2020; Nye et al., 2021; Wei et al., 2022). Moreover, their entangled nature with natural language makes it fundamentally hard to achieve robust inference over factual knowledge (Greff et al., 2020; Saparov and He, 2022; Zhang et al., 2022).

There are other obvious remedies for LMs’ poor reasoning capability. Ensuring that the training corpus contains a sufficient amount of exemplary episodes of sound reasoning reduces the dependency on normative biases and annotation artifacts (Talmor et al., 2020b; Betz et al., 2020; Hase et al., 2021). Heuristics like data augmentation are also shown to be effective (Talmor et al., 2020b). But the above works require significant efforts for crowdsourcing and auditing training data. Our method handily encodes a few prototypes/templates of logic rules and is thus more efficient in terms of human effort. Moreover, our goal is fundamentally different from theirs in investigating the tight integration of neural and symbolic models in an end-to-end manner.

Neuro-symbolic reasoning. Neuro-symbolic approaches are proposed to integrate the perception of deep neural components and the reasoning of symbolic components. Representative works can be briefly categorized into regularization (Xu et al., 2018), program synthesis (Mao et al., 2018), and proof-guided probabilistic programming (Evans and Grefenstette, 2018; Rocktäschel and Riedel, 2017; Manhaeve et al., 2018; Zhang et al., 2019; Huang et al., 2021). To improve compositionality of LMs, previous works propose to parameterize grammatical rules (Kim, 2021; Shaw et al., 2021) but show that those hybrid models are inefficient and usually underperform neural counterparts. In contrast to the above works, DSR-LM focuses on improving LMs’ reasoning over logical propositions with tight integration of their pre-trained knowledge in a scalable and automated way.

Methodology

Each question answering (QA) example in the dataset is a triplet containing input text xx, query qq, and the answer yy. Figure 1 shows an instance that we will use as our running example. The input text xx is a natural language passage within which there will be a set of entities, possibly referenced by 3rd person pronouns. The sentences hint at the relationships between entities. For example, “Dorothy went to her brother Rich’s birthday party” implies that Rich is Dorothy’s brother and Dorothy is Rich’s sister. The query qq is a tuple of two entities, representing the people with whom we want to infer the relation. The expected relation is stored in the answer yy, which will be one of a confined set of possible relations R\mathcal{R}, allowing us to treat the whole problem as an ∣R∣|\mathcal{R}|-way classification problem. We focus only on the problems where the desired relation is not explicitly stated in the context but need to be deduced through a sequence of reasoning.

2 Methodology Overview

The design of DSR-LM concerns tightly integrating a perceptive model for relation extraction with a symbolic engine for logical reasoning. While we apply LMs for low-level perception and relation extraction, we employ a symbolic reasoning module to consistently and logically reason about the extracted relations. With a recent surge in neuro-symbolic methods, reasoning engines are made differentiable, allowing us to differentiate through the logical reasoning process. In particular, we employ Scallop Huang et al. (2021) as our reasoning engine. We propose two add-ons to the existing neuro-symbolic methodology. First, some rules used for logical deduction are initialized using language models and further tuned by our end-to-end pipeline, alleviating human efforts. Secondly, we employ integrity constraints on the extracted relation graphs and the logical rules, to improve the logical consistency of LMs and the learned rules.

Based on this design, we formalize our method as follows. We adopt pretrained LMs to build relation extractors, denoted Mθ\mathcal{M}_{\theta}, which take in the natural language input xx and return a set of probabilistic relational symbols r\mathbf{r}. Next, we employ a differentiable deductive reasoning program, Pϕ\mathcal{P}_{\phi}, where ϕ\phi represents the weights of the learned logic rules. It takes as input the probabilistic relational symbols and the query qq and returns a distribution over R\mathcal{R} as the output y^\hat{y}. Overall, the deductive model is written as

Additionally, we have the semantic loss (sl) derived by another symbolic program Psl\mathcal{P}_{\texttt{sl}} computing the probability of violating the integrity constraints:

Combined, we aim to minimize the objective JJ over training set D\mathcal{D} with loss function L\mathcal{L}:

where w1w_{1} and w2w_{2} are tunable hyper-parameters to balance the deduction loss and semantic loss.

3 Relation Extraction

Since pre-trained LMs have strong pattern recognition capabilities for tasks like Named-Entity-Recognition (NER) and Relation Extraction (RE) (Tenney et al., 2019; Soares et al., 2019), we adopt them as our neural components in DSR-LM. To ensure that LMs take in strings of similar length, we divide the whole context into multiple windows. The goal is to extract the relations between every pair of entities in each windowed context. Concretely, our relation extractor Mθ\mathcal{M}_{\theta} comprises three components: 1) a Named-Entity Recognizer (NER) to obtain the entities in the input text, 2) a pre-trained language model, to be fine-tuned, that converts windowed text into embeddings, and 3) a classifier that takes in the embedding of entities and predicts the relationship between them. The set of parameters θ\theta contains the parameters of both the LM and the classifier.

4 Differentiable Symbolic Inference

The symbolic inference modules Pϕ\mathcal{P}_{\phi} and Psl\mathcal{P}_{\texttt{sl}} are responsible for processing the extracted relations to deduce 1) an expected output relation in R\mathcal{R}, and 2) a semantic loss encoding the probability of constraint violation. There are two main objectives for these modules. First, they need to logically reason about the output relation and the semantic loss based on the extracted relational symbols r\mathbf{r}, the query qq, and the rule weights ϕ\phi. Second, they need to compute the gradients of y^\hat{y} and lsll_{\texttt{sl}} with respect to θ\theta and ϕ\phi, namely ∂y^∂θ\frac{\partial\hat{y}}{\partial\theta}, ∂y^∂ϕ\frac{\partial\hat{y}}{\partial\phi}, ∂lsl∂ϕ\frac{\partial l_{\texttt{sl}}}{\partial\phi}, and ∂lsl∂θ\frac{\partial l_{\texttt{sl}}}{\partial\theta}, in order for the fine-tuning and rule learning to happen.

Logic rules can be applied to known facts to deduce new ones. For example, below is a horn clause, which reads “if bb is aa’s brother and cc is bb’s daughter, then cc is aa’s niece”:

Note that the structure of the above rule can be captured by a higher-order logical predicate called “composite” (abbreviated as comp). This allows us to express many other similarly structured rules with ease. For instance, we can have comp(brother, daughter, niece) and comp(father, mother, grandmother). With this set of rules, we may derive more facts based on known kinship relations. In fact, composition is the only kind of rule we need for kinship reasoning. In general, there are many other useful higher-order predicates to reason over knowledge bases, which we list out in Table 2.

Probability propagation.

We seek to have the deduced facts to also be associated with probabilities computed using probabilities predicted by the underlying relation extractor Mθ\mathcal{M}_{\theta}. This is achieved by allowing the propagation of probabilities. For example, we have the proof tree with probabilities:

In practice, there could be multiple steps in the proof tree (multi-hop) and one fact can be derived by multiple proof trees. We employ the inference algorithms based on approximated weighted model counting (WMC) presented in Manhaeve et al. (2018) to account for probabilistic inference under complex scenarios. Since the WMC procedure is augmented for differentiation, we can obtain the gradient ∂y^∂r\frac{\partial\hat{y}}{\partial\mathbf{r}}. From here, we can obtain ∂y^∂θ=∂y^∂r∂r∂θ\frac{\partial\hat{y}}{\partial\theta}=\frac{\partial\hat{y}}{\partial\mathbf{r}}\frac{\partial\mathbf{r}}{\partial\theta}, where the second part can be automatically derived from differentiating Mθ\mathcal{M}_{\theta}.

Rule learning.

Hand-crafted rules could be expensive or even impossible to obtain. To alleviate this issue, DSR-LM applies LMs to help automatically extract rules, and further utilizes the differentiable pipeline to fine-tune the rules. Each rule such as comp(brother, daughter, niece) is attached a weight, initialized by prompting an underlying LM. For example, the prompt we use for extracting comp(rr,pp,qq) is “one’s rr’s pp is their <qq:mask>”. Given that the relations r,p,q∈Rr,p,q\in\mathcal{R}, DSR-LM automatically enumerates rr and pp from R\mathcal{R} while querying for LM to unmask the value of qq. LM then returns a distribution of words, which we take an intersection with R\mathcal{R}. The probabilities combined form the initial rule weights ϕ\phi. This type of rule extraction strategy is different from existing approaches in inductive logic programming since we are exploiting LMs for existing knowledge about relationships.

Note that LMs often make simple mistakes answering such prompt. In fact, with the above prompt, even GPT-3 can only produce 62% of composition rules correctly. While we can edit prompt to include few-shot examples, in this work we consider fine-tuning such rule weights ϕ\phi within our differentiable reasoning pipeline. The gradient with respect to ϕ\phi is also derived with the WMC procedure, giving us ∂y^∂ϕ\frac{\partial\hat{y}}{\partial\phi}. In practice, we use two optimizers with different hyper-parameters to update the rule weights ϕ\phi and the underlying model parameter θ\theta, in order to account for optimizing different types of weights.

Semantic loss and integrity constraints.

In general, learning with weak supervision label is hard, not to mention that the deductive rules are learnt as well. We thereby introduce an additional semantic loss during training. Here, semantic loss is derived by a set of integrity constraints used to regularize the predicted entity-relation graph as well as the learnt logic rules. In particular, we consider rules that detect violations of integrity constraints. For example, “if A is B’s father, then B should be A’s son or daughter” is an integrity constraint for relation extractor—if the model predicts a father relationship between A and B, then it should also predict a son or daughter relationship between B and A. Encoded in first order logic, it is

Through differentiable reasoning, we evaluate the probability of such constraint being violated, yielding our expected semantic loss. In practice, arbitrary number of constraints can be included, though too many interleaving ones could hinder learning.

Experiments

We evaluate DSR-LM on both CLUTRR and DBpedia-INF. We show that DSR-LM has accurate and generalizable long-range reasoning capability.

CLUTRR (Sinha et al., 2019) consists of kinship reasoning questions. Given a context that describes a family’s routine activity, the goal is to deduce the relationship between two family members that is not explicitly mentioned in the story. Although the dataset is synthetic, the sentences are crowd-sourced and hence there is a considerable amount of naturalness inside the dataset. The family kinship graph is synthetic and the names of the family members are randomized. For ablation study, we manually crafted 92 kinship composition rules as an external symbolic knowledge base. This yields the following symbolic information for each datapoint: 1) the full kinship graph corresponding to the story, 2) the symbolic knowledge base (KB), and 3) a query representing the question. The CLUTRR dataset is divided into different difficulties measured by kk, the number of facts used in the reasoning chain. For training, we only have 10K data points with 5K k=2k=2 and another 5K k=3k=3, meaning that we can only receive supervision on data with short reasoning chains. The test set, on the other hand, contains 1.1K examples with k∈{2,…,10}k\in\{2,\dots,10\}.

DBpedia-INF is a curated subset of the evaluation dataset used in RuleBert Saeed et al. (2021). Similar to CLUTRR, it is generated synthetically to test the reasoning capability of LMs. Given a synthetic passage describing the relation between entities, and soft deductive logic rules, we aim to deduce the relationship between any two entities. The symbolic program of DBpedia-INF consists of 26 predicates, 161 soft rules mined from DBpedia, and 16 rules defining the negation and symmetricity between the predicates. The difficulty of the questions is represented in terms of reasoning length from k∈{0,…,5}k\in\{0,\dots,5\}.A length of 0 means that the hypothesis can be verified using the facts alone without using any rules. Larger kk implies harder question. Compared to the exact dataset used in Rulebert, we clean it in order to ensure the question-answer pairs are logically consistent and probabilistically correct.

2 Experimental Setup

We employ Scallop Huang et al. (2021) as the differentiable symbolic inference module. We show the program used for CLUTRR reasoning task in Figure 2. It comprises relation type declarations, deductive rules for kinship reasoning, and integrity constraints for computing semantic loss (attached in the Appendix). The program used for DBpedia-INF is written in a similar manner with additional high-order predicates listed in Table 2.

Pre-trained LMs for fine-tuning.

We used the HuggingFace (Wolf et al., 2019) pre-trained w2v-google-news-300, RoBERTa-base, and DeBERTa-base as the pretrained language models. We finetune RoBERTa-base and DeBERTa-base during training with binary cross entropy loss. Our relation extraction module is implemented by adding an MLP classifier after the LM, accepting a concatenation of the embedding of the two entities and the embedding of the whole windowed context.

Our model.

Our main model, DSR-LM, uses RoBERTa as the underlying LM. The relation classifier is a 2-layer fully connected MLP. For training, we initialize ϕ\phi by prompting the LM. To accelerate the learning process, we use multinomial sampling to retrieve 150 rules for symbolic reasoning. During testing, we will instead pick the top 150 rules. We use two Adam optimizer to update θ\theta and ϕ\phi, with learning rate 10−510^{-5} and 10−210^{-2} respectively.

For ablation studies, we present a few other models. First, we ablate on back-bone LMs. Specifically, we have DSR-LM-DeBERTa which uses DeBERTa as the back-bone LM. DSR-w2v-BiLSTM, on the other hand, uses as back-bone the word2vec Mikolov et al. (2013) model for word embedding and BiLSTM Huang et al. (2015) for sequential encoding. For DSR-LM-with-Manual-Rule we treat the logic rules as given, meaning that we provide 92 composition rules for CLUTRR and around 180 rules for DBpedia-INF. In this case, we set ground truth rules to have 1.01.0 weight and therefore ϕ\phi is not learnt. Then, we have DSR-LM-without-IC which does not have integrity constraints and semantic loss. Lastly, we have DSR-without-LM that takes ground truth structured entity relation graph as input. This way, we do not need the underlying relation extractor and only ϕ\phi needs to be learned.

Baselines.

We compare DSR-LM with a spectrum of baselines from purely neural to logically structured. The baselines include pretrained large language models (BERT (Kenton and Toutanova, 2019) and RoBERTa (Liu et al., 2019)), non-LM counterparts (BiLSTM (Hochreiter and Schmidhuber, 1997; Cho et al., 2014) and BERT-LSTM), structured models (GAT (Veličković et al., 2018), RN (Santoro et al., 2017), and MAC (Hudson and Manning, 2018)), and other neuro-symbolic models (CTP Minervini et al. (2020), RuleBert (Saeed et al., 2021)). The structured models include those models with relational inductive biases, while the neuro-symbolic model uses logic constraints.

Baseline setup.

We highlight a few baselines we include for completeness but are treated as unfair comparison to us: GAT, CTP, and GPT-3 variants. All baselines other than GAT and CTP take as input natural language stories and the question to produce the corresponding answer. GAT and CTP, on the contrary, takes entity relation graph rather than natural language during training and testing.

The model sizes are different across baselines as well. Model size generally depends on two parts, the backbone pre-trained LM, and the classification network built upon the LM. GPT-3 contains 175B parameters, and RoBERTa uses 123M parameters. The classification model of our method has 2.97M parameters (assuming using embeddings from RoBERTa). With extra 10K parameters for rule weights, our DSR-LM framework has around 127M parameters.

For GPT-3 variants, we conduct experiments on CLUTRR with GPT-3 under the Zero-Shot (GPT-3 ZS), GPT-3 Fine-Tuned (GPT-3 FT), and Few(5)-Shot (GPT-3 5S) (Brown et al., 2020), as well as Zero-Shot-CoT (GPT-3 ZS-CoT) (Kojima et al., 2022a) settings. For fair comparison, we also include the ground truth kinship composition knowledge in GPT-3 zero shot (GPT-3 ZS w/ Rule), and 5 shot (GPT-3 5S w/ Rule). We include the prompts we used and additional details in Appendix A.

3 Experimental Results

We evaluate DSR-LM and baselines on both CLUTRR and DBpedia-INF, as reported in Figure 3 and Table 3.

In the CLUTRR experiment, DSR-LM achieves the best performance among all the models (Figure 3). Next, we examine how models trained on stories generated from clauses of length k≤3k\leq 3 and evaluated on stories generated from larger clauses of length k≥4k\geq 4. A fine-grained generalizability study reveals that although all models’ performances decline as the reasoning length of the test sequence increases, pure neural-based models decrease the fastest (Figure 4(a) and 4(b)). It manifests the systematic issue that language models alone are still not robust for length generalization (Lake and Baroni, 2018). On the other hand, the performance of DSR-LM decreases much slower as test reasoning length increases and outperforms all the baselines when k≥4k\geq 4.

In the DBpedia-INF experiment, DSR-LM outperforms RuleBert by 37% in terms of overall performance (Table 3), showing that DSR-LM has much more robust generalization. Recall that RuleBert aims to improve the logical reasoning of LMs by straightforward fine-tuning with soft rules and facts. Our results show that augmenting data alone for fine-tuning do not effectively improve systematicity. Meanwhile, DSR-LM imbues reasoning inductive biases throughout training and learns useful rules to generalize to longer reasoning lengths.

Learning interpretable logic rules.

DSR-LM is capable of producing explicit logic rules as part of the learning process. For presentation, we show the top-10 rules learnt from DSR-LM model in Table 4. We compare the top-92 most likely prompted and fine-tuned rules against the 92 hand-crafted rules, and 70 of them match. Additionally, we find that our rule weight fine-tuning helps correct 11 of the incorrect rules produced by LM. Through this qualitative analysis, it is clear that DSR-LM provides an interface to probe and interpret the intermediate steps, enhancing the interpretability.

GPT-3 variants are inferior in long-range reasoning.

Interestingly, ZS scores 28.6% accuracy on CLUTRR while ZS-CoT scores 25.6%, suggesting that the chain-of-thought prompting might not work in long-range reasoning (Figure 3). In fact, there are many cases where GPT-3 favors complication over simplicity: GPT-3 frequently answers “stepdaughter”, “stepmother”, and “adopted son”, while the real answers are simply “daughter”, “mother”, and “son”. Additionally, GPT-3 could derive the correct result for the wrong reason, e.g. “Jeffrey is Gabrielle’s son, which would make William her grandson, and Jeffrey’s brother.” While we count the final answer to be correct (William is Jeffrey’s brother), there is a clear inconsistency in the reasoning chain: William cannot be Gabrielle’s grandson and Jeffrey’s brother simultaneously, given that Jeffrey is Gabrielle’s son. Lastly, we observe that, both GPT-3 FT and many other methods have an accuracy drop as kk becomes larger (Figure 4(b)), ZS and ZS-CoT stay relatively consistent, suggesting that the size of context and the reasoning chain may have a low impact on GPT-3’s performance.

4 Analyses and Ablation Studies

Since DSR-LM has a model agnostic architecture, we study how the choice of different LMs impacts the reasoning performance. As shown in Table 5, the two transformer-based models have on-par performance and outperform the word2vec one. However, note that the word2vec-based model still has better performance than all other baselines. Besides higher final accuracy, the pre-trained transformer-based language model also accelerates the training process. Both DSR-LM-RoBERTa and DSR-LM-DeBERTa reach their best performance within 20 epochs, while it takes DSR-w2v-BiLSTM 40 epochs to peak.

Incorporate domain knowledge.

DSR-LM allows injecting domain specific knowledge. In DSR-LM-with-Rule, we manually crafted 92 rules for kinship reasoning to replace the learnt rules. As shown in Table 6, it obtained a 0.36% performance gain over DSR-LM. The fact that the improvement is marginal implies our method extracts useful rules to obtain on-par performance with manually crafted ones. DSR-LM-without-IC, our model without integrity constraints specified on predicted relations and rules, performs worse than DSR-LM, suggesting that logical integrity constraints are essential component for improving the model robustness.

The impact of the relation extractor.

To understand what causes the failure case of DSR-LM, we study the performance of our relation classification model separately. We isolate the trained relation extractor and found that it reaches 84.69% accuracy on the single relation classification task. For comparison, we train a relation extractor using all the intermediate labels in the training dataset, and it reaches 85.32% accuracy. It shows that even using only weak supervision (i.e., the final answers to multi-hop questions), our approach can reach on-par performance as supervised relation extraction.

Reasoning over structured KBs.

To understand the rule learning capability of our approach, we design our ablation model DSR-without-LM to take as input ground-truth KBs instead of natural language. In this case, rule weights are not initialized by LM but randomized. As shown in Table 7, our model outperforms GAT and CTP which also operates on structured KBs. It demonstrates that our differentiable rule learning paradigm learns rules to reason about KBs consistently.

Failure cases of DSR-LM.

We showcase in Appendix Table 8 that even state-of-the-art large LMs are prone to logical fallacies. On the other hand, the failure case of our method usually occurs in the stage of relation extraction. For example, for the following sentence “Christopher and Guillermina are having a father-daughter dance”, our RoBERTa based relation extractor fails to recognize the father-daughter relationship but rather thinks C and G have a husband-wife relationship. We require most of the relation extraction to be correct in order to avoid cascading error. As the error rate on individual relation extraction accumulates, it leads to the observed drop in accuracy as kk becomes larger.

Concluding Remarks

We investigate how to improve LMs’ logical reasoning capability using differentiable symbolic reasoning. Through extensive experiments, we demonstrate the effectiveness of DSR-LM over challenging scenarios where widely deployed large LMs fail to reason reliably. We hope our work can lay the groundwork for exploring neuro-symbolic programming techniques to improve the robustness of LMs on reasoning problems.

References

Appendix A Implementation Details

Hardware. We perform all the experiments on a server with two 20-core Intel Xeon CPUs, four GeForce RTX 2080 Ti GPUs, and 768 GB RAM.

Reasoner details. The learning of rules and the fine-tuning of the underlying LM should happen separately with different learning rates – fine-tuning LM is an intricate process that requires a very small learning rate, while rules should be learned with larger learning rates since gradients are directly back-propagated onto the weights. This can be realized by employing two separate optimizers, one for fine-tuning and the other for rule learning. During training time, we rotate training the two parts by toggling one and the other optimizer for every 10 batches of data points.

Rule learning training setup. For rule learning, we can initialize the transitivity tensor using the language model provided composite rules. Since the CLUTRR dataset consists of 2020 different relations and a transitivity relationship is defined over 33 relations, there are 8K possible transitivity facts over these relations. Specifically, we give every predicted composite rule by the GPT with a 0.50.5 weight, while initializing the other rules with a range such as [0,0.1][0,0.1], since otherwise, an insensible transitive fact may be getting a random high weight while it effectively does nothing for reasoning. The learning process encourages the rules that yield the correct query result and suppresses the rules that lead to wrong answers. To avoid the exponential blow-up caused by injecting all the 8K rules in the reasoning engine, we sample 200200 rules according to their weights during the training time and deterministically use the top 200200 learned rules during the test time. For the QA-No-Rule setup, the confidence score of rules, the MLP classifier for relation extraction, and the underlying LM are learned and updated simultaneously during training. To account for their difference, we employ two Adam optimizers ARLA_{\text{RL}} and AREA_{\text{RE}}. AREA_{\text{RE}} is used for optimizing models for relation extraction, and thus will take as parameters the MLP classifier and the underlying LM. It has a low learning rate 0.000010.00001 since it needs to fine-tune LMs. ARLA_{\text{RL}}, on the other hand, will take as a parameter the confidence score tensor for the transitive rules, and is set to have a higher learning rate of 0.0010.001. For the integrity constraints, we set the result integrity violation loss with the weight 0.10.1, and set the rule integrity constraint violation loss with the weight 0.010.01. We set the batch size to 1616 and train for 2020 epochs.

To obtain the initial rule weights for the composition rule in our CLUTRR experiment, the prompt we use is “Mary’s P’s Q is her .” where P and Q are enumerations of all possible relationships, and the unmasked value is treated as the answer R, producing composite(P, Q, R). For the other rule templates we used, the prompts are

transitive: “is R’s R one’s R? ”; the probability of the unmasked word being “yes” is treated the rule weight for transitive(R).

symmetric: “does A is R of B means B is R of A? ”; the probability of the unmasked word being “yes” is treated the rule weight for symmetric(R).

inverse: “A is R of B means B is of A”; the unmasked value is treated as the answer P, producing inverse(R, P).

implies: “does R imply P? ”; the probability of unmasked value being “yes” is treated as the rule weight for implies(R, P).

GPT-3 Prompt Setups. For Zero-Shot, we use the prompt “So BB is AA’s:” for the query pair (A,B)(A,B) to ask GPT-3 to complete the relationship between AA and BB. We pick the phrase in the first line or before the first period from the completed text and compare it directly with the ground truth relation. For the Few(5)-Shot setting, we randomly select 5 examples from the training dataset used for other models (k∈k\in) to serve as examples. We use the same prompt for Few-Shot and Fine-Tuned as the Zero-Shot and the automated GPT-3 fine-tuning setup for our training dataset, trained for 4 epochs. To add in the transitive KB, we simply include 92 hand-crafted rules in natural language as a part of the prompt, and we performed Zero-shot with KB, and Few(5)-shot with KB experiments. For the Zero-Shot-CoT setting, we use the prompt “Who is BB to AA? Let’s think step by step” to suggest GPT-3 to auto-complete while working out a reasoning chain. Under this setup, it is impossible to compare the answer to the ground truth automatically. Therefore, we manually check through the whole test dataset of CLUTRR.

Appendix B Additional Experimental Results

In Table 8, we showcase the failure cases of large LMs for logical inference, where Zero-shot-CoT denotes zero-shot chain-of-thoughts (Kojima et al., 2022b).