Rethinking with Retrieval: Faithful Large Language Model Inference
Hangfeng He, Hongming Zhang, Dan Roth
Introduction
Large language models (LLMs) have shown exceptional performance across various tasks through in-context learning without task-specific training or fine-tuning Brown et al. (2020); Chowdhery et al. (2022); Zhang et al. (2022); Ouyang et al. (2022). Recent progress in prompting Wei et al. (2022); Zhou et al. (2022); Kojima et al. (2022) and decoding Wang et al. (2022) has made it feasible for LLMs to tackle tasks that demand complex reasoning.
However, the knowledge stored in LLMs might inevitably be incomplete, out-of-date, or incorrect. As a result, external sources of knowledge, such as Wikipedia, may be essential for the successful deployment of LLMs for real-world applications. Previously, people tried to utilize knowledge for smaller language models (LMs), such as T5 Raffel et al. (2020), BERT Devlin et al. (2019), and RoBERTa Liu et al. (2019). However, these methods often require additional training or fine-tuning, which can be costly and thus impractical for LLMs.
In this paper, we present a post-processing approach called rethinking with retrieval (RR) for utilizing external knowledge in LLMs. Our method begins by using the chain-of-thought (CoT) prompting method Wei et al. (2022) to generate a diverse set of reasoning paths, as described in Wang et al. (2022). We then use each reasoning step in those paths to retrieve relevant external knowledge, which enables RR to provide more faithful explanations and more accurate predictions, as illustrated in Figure 1.
We evaluate the effectiveness of our proposed method, RR, on three complex reasoning tasks: commonsense reasoning, temporal reasoning, and tabular reasoning, using GPT-3 175B Brown et al. (2020) and different external knowledge sources: Wikipedia, Wikidata Vrandečić and Krötzsch (2014), WordNet Miller (1995), and Conceptnet Speer et al. (2017). The results demonstrate that RR consistently outperforms all baselines on all three tasks without requiring additional training or fine-tuning, indicating the superiority of our approach in leveraging external knowledge to enhance the performance of LLMs.
Related Work
Retrieval-enhanced LMs have received significant attention as a means of improving performance through the incorporation of external knowledge. For example, the k-most similar training contexts can be retrieved to improve the estimation of the next word distribution in both the training stage Borgeaud et al. (2021) and the inference stage Khandelwal et al. (2020). Furthermore, search query generators have been adopted to generate search queries for search engines to retrieve relevant documents Komeili et al. (2022); Shuster et al. (2022); Thoppilan et al. (2022). Other approaches have utilized retrieved documents as the additional context in generation tasks Joshi et al. (2020); Guu et al. (2020); Lewis et al. (2020). Nakano et al. (2021) instead use human feedback in a text-based web-browsing environment. Among these previous works, Khandelwal et al. (2020) is most closely related to our approach. However, they focus on improving local inference by using the nearest neighbor datastore constructed from training data, whereas we focus on conducting faithful inference using external knowledge. In contrast to other aforementioned approaches, which require training or fine-tuning to incorporate retrieved knowledge, we propose a post-processing method for leveraging retrieved knowledge without additional training or fine-tuning.
Incorporating external knowledge into LMs.
Significant effort has been devoted to leveraging external knowledge to improve the reasoning ability of LMs. Previous work has incorporated external knowledge sources such as WordNet Miller (1995) and ConceptNet Speer et al. (2017) to enhance LMs for tabular reasoning tasks Neeraja et al. (2021); Varun et al. (2022). Explicit rules have also been added to inputs to improve reasoning ability over implicit knowledge Talmor et al. (2020). In addition, explicit knowledge from Wikidata Vrandečić and Krötzsch (2014) and implicit knowledge in LLMs have been integrated into a transformer Vaswani et al. (2017) for visual question answering Gui et al. (2021). Nye et al. (2021) instead introduces a symbolic reasoning module to improve coherence and consistency in LLMs. Among these previous works, Nye et al. (2021) is the most relevant to our approach. Still, they focus on incorporating logical constraints to improve coherence and consistency, whereas we aim to improve the faithfulness of explanations through the use of external knowledge. In contrast to other aforementioned approaches that incorporate external knowledge before generation and require additional training or fine-tuning, our proposal leverages external knowledge in a post-processing manner to enhance LMs without additional training or fine-tuning.
Uncovering latent Knowledge in LLMs.
There has been a line of work exploring the knowledge hidden within LLMs for reasoning. This has included the use of careful prompting to encourage LLMs to generate explanations in the reasoning process, such as through chain of thought prompting in few-shot Wei et al. (2022) or zero-shot Kojima et al. (2022) learning, or through the use of scratchpads for intermediate computation Nye et al. (2022). In addition, various methods based on sampling a diverse set of reasoning paths in LLMs have been proposed, including training verifiers to judge the correctness of model completions Cobbe et al. (2021), calibrating model predictions based on the reliability of the explanations Ye and Durrett (2022), and promoting self-consistency over diverse reasoning paths Wang et al. (2022). Zelikman et al. (2022) instead iteratively bootstrap the ability of LLMs to generate high-quality rationales from a few initial examples. Liu et al. (2022) further propose generating knowledge from LLMs, which is then used as additional input to improve commonsense reasoning. In contrast to this line of work, our proposal focuses on leveraging external knowledge to enhance LLMs, while they aim to explore the knowledge hidden within LLMs.
Rethinking with Retrieval
LLMs have been shown to generate incorrect supporting facts from time to time, even when they accurately capture the perspective needed to answer a question. This phenomenon highlights intrinsic issues in the way LLMs store and retrieve knowledge, including (1) the presence of out-of-date, incorrect, or missing relevant knowledge in the pre-training corpus; (2) incorrect memorization of relevant knowledge during pre-training; and (3) incorrect retrieval of relevant knowledge during the inference stage. To address these issues, we propose the use of RR, which leverages external knowledge through the retrieval of relevant information based on decomposed reasoning steps.
Given a query , we utilize chain-of-thought prompting to generate a diverse set of reasoning paths , where each reasoning path consists of an explanation followed by a prediction . After that, we retrieve relevant knowledge from a suitable knowledge base to support the explanation in each reasoning path, and select the prediction that is most faithful to this knowledge. To better illustrate our proposal, we use “Did Aristotle use a laptop?” as a running example in this work.
Chain-of-thought prompting.
In contrast to standard prompting, CoT prompting Wei et al. (2022) includes demonstrations of step-by-step reasoning examples in the prompt to produce a series of short sentences that capture the reasoning process. For instance, given the question “Did Aristotle use a laptop?”, CoT prompting aims to generate the complete reasoning path “Aristotle died in 322 BC. The first laptop was invented in 1980. Thus, Aristotle did not use a laptop. So the answer is no.” rather than simply outputs “No.” Empirical results show that CoT prompting significantly improves the performance of LLMs on many multi-step reasoning tasks. Therefore, we adopt CoT prompting to obtain both explanation and prediction for the query .
Sampling diverse reasoning paths.
Similar to Wang et al. (2022), we sample a diverse set of reasoning paths rather than only considering the greedy path as in Wei et al. (2022). For the question “Did Aristotle use a laptop?”, the potential reasoning paths can be as follows:
Aristotle died in 2000. The first laptop was invented in 1980. Thus, Aristotle used a laptop. So the answer is yes.
Aristotle died in 322BC. The first laptop was invented in 2000. Thus, Aristotle did not use a laptop. So the answer is no.
Aristotle died in 322BC. The first laptop was invented in 1980. Thus, Aristotle did not use a laptop. So the answer is no.
Knowledge retrieval.
Different knowledge bases can be used to address different tasks. For example, to address the question “Did Aristotle use a laptop?”, we can use Wikipedia as the external knowledge base . Information retrieval techniques can be applied to retrieve the relevant knowledge from Wikipedia based on the decomposed reasoning steps. Ideally, we would obtain the following two paragraphs from Wikipedia for this question:
Aristotle (384–322 BC) was a Greek philosopher and polymath during the Classical period in Ancient Greece. …
The Epson HX-20, the first laptop computer, was invented in 1980. …
Faithful inference.
The faithfulness of each reasoning path can be estimated using a function , which is based on relevant knowledge retrieved from the knowledge base . The final prediction is obtained through the application of the following inference procedureNote that this is the basic version of faithful inference, and further variations can be found in Section 5.3.:
where denotes the corresponding prediction in the reasoning path . This inference procedure is designed to identify the most faithful prediction to the knowledge base among all predictions in the reasoning paths. For instance, in the running example, given reasoning paths and the retrieved knowledge , the above inference procedure would output the prediction “So the answer is no.”, as it is supported by both and and has a higher faithfulness score compared to the prediction “So the answer is yes.”, which is only supported by .
Experiments
In this section, we present the evaluation of our proposed method, RR, on three complex reasoning tasks: commonsense reasoning, temporal reasoning, and tabular reasoning.
In our experiments, we consider GPT-3 with standard zero-shot/few-shot prompting as baselines, following the approach described in Brown et al. (2020), in which zero or few in-context exemplars of input-output pairs are provided in the prompt.
Chain-of-thought prompting.
In addition to the standard zero-shot/few-shot prompting, we also consider GPT-3 with the CoT prompting proposed in Wei et al. (2022) as a baseline in our experiments. This approach involves feeding LLMs step-by-step reasoning examples instead of standard input-output examples.
Self-consistency.
In addition, we also consider self-consistency Wang et al. (2022) as a baseline in our experiments. This approach, proposed as an alternative to the naive greedy decoding used in CoT prompting Wei et al. (2022), involves sampling a diverse set of reasoning paths and selecting the most consistent answer by marginalizing the sampled paths.
2 Commonsense Reasoning
For commonsense reasoning, we consider the StrategyQA dataset Geva et al. (2021), which includes questions that require implicit reasoning strategies. For example, the question “Did Aristotle use a laptop?” requires implicit decomposition into reasoning steps, while the question “Was Aristotle alive when the laptop was invented?” explicitly specifies the reasoning process. The StrategyQA dataset includes training examples, each consisting of a question (Q), a yes/no answer (A), a decomposition (D), evidence paragraphs (E), and supporting facts (F). On average, each question requires about reasoning steps and evidence paragraphs. In addition, a development set is constructed by randomly sampling of the training examples (i.e., examples). The answer distribution is roughly balanced, with approximately "yes" questions in both the training and development sets. Unless otherwise specified, the models are evaluated on the development setAs the annotations for the test set are not publicly available, we use the development set for evaluation. This allows us to perform a more comprehensive analysis. for StrategyQA.
Implementation details.
In this part, we utilize Wikipedia as the external knowledge base . For each sentence in the explanation of every reasoning path, we first apply BM25 Robertson et al. (2009) to retrieve the top 10 most relevant paragraphs from Wikipedia. In particular, we use the re-implementation of the sparse retrieval BM25We also experimented with DPR and BM25+DPR, and found that BM25 outperformed these methods in our experiments. More details can be found in Appendix A.3. in Karpukhin et al. (2020) from Pyserini Lin et al. (2021). Subsequently, we use the pre-trained MPNet model Song et al. (2020) to select the most similar paragraph based on the cosine similarity between the sentence embeddings of the retrieved paragraph and the sentence. We then employ a pre-trained natural language inference (NLI) model Nie et al. (2020) to obtain the entailment and contradiction scores for the sentence, treating the most similar paragraph as the premise. The faithfulness of each reasoning path is then calculated using based on the entailment scores, contradiction scores, and MPNet similarities of all sentences in the explanation of the reasoning path. The final prediction for each question is obtained through faithful inference (Equation 1). More details about can be found in Appendix A.2.
3 Temporal Reasoning
In this experiment, we use the TempQuestions dataset Jia et al. (2018) to investigate temporal reasoning. This dataset includes temporal questions that are divided into four classes: explicit temporal, implicit temporal, temporal answer, and ordinal constraints. The questions are paired with their answers from Freebase Bollacker et al. (2008). To examine the most challenging aspect of temporal reasoning, we focus on the set of implicit temporal questions, which contain implicit temporal expressions, including free-text temporal expressions. For example, the question “who was governor of oregon when shanghai noon was released?” is an implicit temporal question. To facilitate our analysis, we only consider questions with a single answer, resulting in a total of examples. Of these examples, the first are used for prompting, and the remaining are used for evaluation.
Implementation details.
In this part, we utilize Wikidata Vrandečić and Krötzsch (2014) as the external knowledge base , as it is the largest publicly available knowledge graph, and the data from Freebase has been migrated to Wikidata. To incorporate this knowledge into our system, we apply an entity linking systemWe use the spacy entity linker: https://pypi.org/project/spacy-entity-linker/. to each sentence in the explanation of each reasoning path to identify the corresponding Wikidata pages for all entities in the sentence. Next, we extract all temporal relations from these relevant Wikidata pages and use templates to convert these temporal relations into sentences. This step generates a set of relevant knowledge sentences for each sentence in the explanation of each reasoning path. The final prediction is then obtained by applying the procedure described in Section 4.2, in which the retrieved paragraphs are replaced with the relevant knowledge sentences from the current part.
4 Tabular Reasoning
We consider the INFOTABS dataset Gupta et al. (2020) for tabular reasoning, which consists of human-written textual hypotheses based on premises in the form of tables extracted from unique Wikipedia info-boxes. We focus on the development set, which includes hypotheses based on tables, and only consider entailed and contradictory hypotheses as it is tricky to write CoT demonstrations for neutral hypotheses. This results in a total of hypotheses based on tables for evaluation, with an equal number of entailed and contradictory hypotheses.
Implementation details.
In this part, we utilize WordNet Miller (1995) and ConceptNet Speer et al. (2017) as external knowledge bases. To convert tables into textual premises, we follow the same technique as in Varun et al. (2022). For each premise-hypothesis pair, we follow the procedure outlined in Varun et al. (2022) to retrieve relevant word relation triples that connect the premise and hypothesis words, such as “married” “spouse”. These triples are then converted into sentences using some simple templates. The resulting sentences, along with the textual premises from the tables, serve as relevant knowledge for each sentence in the explanation of each reasoning path. To obtain the final prediction, the procedure described in Section 4.2 is applied, whereby the retrieved paragraphs in Section 4.2 are replaced with the relevant knowledge from the current part.
5 Evaluation
In all experiments, we utilize GPT-3 text-davinci-002 unless otherwise stated. The maximum number of tokens for generation during completion is set to . For zero-shot, few-shot, and chain-of-thought prompting, the temperature is fixed at . For self-consistency and rethinking with retrieval, we randomly sample outputsFor commonsense reasoning, we sample outputs, as we have found that odd numbers of outputs tend to yield better voting performance for self-consistency on StrategyQA. with temperature . Detailed prompts can be found in Appendix A.1. We evaluate the performance of different methods on commonsense and tabular reasoning using accuracy, and on temporal reasoning using the exact match metric as defined in Rajpurkar et al. (2016).
Results.
As shown in Table 1, our proposed method, rethinking with retrieval, consistently outperforms all baselines on all three reasoning tasks without requiring additional training or fine-tuning. The results highlight the effectiveness of our approach in leveraging external knowledge to improve the performance of LLMs.
Analysis
In this section, we perform a thorough analysis to gain a deeper understanding of RR.
In this subsection, we present an analysis of GPT-3 with CoT prompting on the StrategyQA dataset. Upon closer examination of the outputs of GPT-3, we observed that it can provide reasonable explanations and correct predictions for a number of questions. For example, when given the question “Will the Albany in Georgia reach a hundred thousand occupants before the one in New York?”, GPT-3 produced the following output:
The Albany in New York has a population of about 98,000. The Albany in Georgia has a population of about 77,000. Thus, the Albany in New York is more populous than the Albany in Georgia. So the answer is no.
The above output consists of three components: (1) supporting facts (in cyan) that are based on a particular perspective, (2) chaining arguments (in orange), and (3) a prediction (in green). Components (1) and (2) contribute to the explanation. Overall, the output exhibits a high level of quality. However, we also observed that GPT-3 may occasionally produce incorrect supporting facts for its explanations or make incorrect inferences for its predictions, despite generally being able to identify suitable perspectives.
As shown in Table 2, GPT-3 provides the incorrect supporting fact for Lil Jon’s top-ranked Billboard song, stating that it was “Get Low” instead of the correct answer, “Yeah”. However, it does have the correct perspective on how to answer the question, “Was Lil Jon’s top ranked Billboard song a collaboration with a member of The Lox?”.
Wrong inference.
As shown in Table 2, GPT-3 makes an incorrect inference, stating that the top of Mount Fuji “would not stick out” of the Sea of Japan, rather than the correct answer, “would stick out”. However, it does provide correct supporting facts based on the appropriate perspective for the question, “Would the top of Mount Fuji stick out of the Sea of Japan?”.
2 Ablation Study
In our proposed method, we retrieve relevant external knowledge based on the decomposed reasoning steps rather than the original query. To further investigate the impact of this choice, we conducted additional experiments in which we used the original query for knowledge retrieval while keeping other aspects of our method unchanged. As shown in Table 3, the results for these experiments are poor for both commonsense and temporal reasoning, indicating the importance of using decomposition-based retrieval in our approach.
The impact of different types of knowledge.
For tabular reasoning, we use both external knowledge (WordNet and ConceptNet) and background knowledge (tables) in our experiments. In this section, we further examine the effect of different types of knowledge on the performance of our proposed method. As shown in Table 4, the additional improvement gained by incorporating Wikidata and ConceptNet in addition to tables is limited, indicating that GPT-3 already captures many word-level relations in these external knowledge sources. In addition, the observed significant improvement in tabular reasoning from using tables alone suggests that our proposed method can also effectively leverage background knowledge.
3 Variations of the Proposed Approach
In Section 3, we present a basic version of our proposal for taking advantage of external knowledge. Our basic approach involves weighting outputs as individual units and using a voting mechanism to select the best-supported prediction. We can also directly choose the best-supported output, which includes both an explanation and a prediction, without using voting. For example, in the running example of “Did Aristotle use a laptop?” (see more in Section 3), the third reasoning path is the output most supported by the knowledge paragraphs and .
Variant I: Fact selection.
The first variant of our approach involves selecting facts from the outputs of LLMs based on external knowledge. For example, consider the running example of “Did Aristotle use a laptop?”, where we only have access to the first two reasoning paths, and . In this case, the first sentence in and the second sentence in are supported by knowledge and , respectively. Therefore, the first variant would output the first sentence in and the second sentence in as the supporting facts.
Variant II: Fact generation.
The second variant of our approach involves generating facts based on both the outputs of LLMs and external knowledge. For example, consider the running example of “Did Aristotle use a laptop?”, where we only have access to the first reasoning path . The second sentence in is supported by the second knowledge paragraph . However, the first sentence is not supported by any evidence paragraphs. We can generate questions about the first sentence, such as “When did Aristotle die?” and use the first knowledge paragraph to generate a new fact: “Aristotle died in 322BC.”. As a result, the second variant would output the generated fact “Aristotle died in 322 BC.” and the second sentence in as the supporting facts.
Inference with supporting facts.
For the two variants of our approach, we only have the supporting facts and need to perform a final inference step to obtain the corresponding prediction. One option for this inference is to use LLMs, but they can be costly Brown et al. (2020) or difficult to use Zhang et al. (2022). An alternative is to use an off-the-shelf model for inference with supporting facts, such as UnifiedQA Khashabi et al. (2020, 2022). As discussed in Appendix A.5, UnifiedQA is more robust to noisy supporting facts than GPT-3. We thus use the second version of UnifiedQA, UnifiedQA-v2 Khashabi et al. (2022), for the final step of inference.
Experimental settings.
In this part, we focus on commonsense reasoning and use the evidence paragraphs provided in StrategyQA as the relevant knowledge, rather than the retrieved paragraphs discussed in Section 4.2. To evaluate the quality of the explanations, we adopt the best metric for factual consistency evaluation in Honovich et al. (2022). For simplicity, we use the pre-trained NLI model released by Nie et al. (2020) to compute the NLI-based metric, rather than fine-tuning T5-11B Raffel et al. (2020) ourselves. The implementation details of the two variants can be found in Appendix A.4.
Results.
Table 5 illustrates that the fact selection and fact generation variants of our proposal improve the faithfulness of the supporting facts in explanations, leading to increased prediction accuracy compared to the basic approach without voting. Across all variations of our proposal, we observe significant improvements in both prediction accuracy and the faithfulness of explanations when compared to the CoT prompting baseline.
The incorporation of a voting mechanism leads to an increased prediction accuracy of for the basic approach. Comparison with the performance (i.e., ) of the same approach using retrieved paragraphs rather than evidence paragraphs in Table 1 demonstrates that retrieved paragraphs are also effective for our proposal, as both significantly outperform the voting baseline, self-consistency (i.e., ), as shown in Table 1.
It is noteworthy that UnifiedQA performs poorly on StrategyQA, achieving an accuracy of only . However, when provided with gold supporting facts in StrategyQA, UnifiedQA demonstrates excellent performance with an accuracy of . This suggests that UnifiedQA is suitable for last-step inference, but not effective for answering questions in StrategyQA.
4 Impact of the Size of LMs
In this subsection, we examine the effect of the size of LMs on the performance of our proposed method, specifically in the context of the fact generation variant. We compare the performance of our method using various sizes of OPT models Zhang et al. (2022) in addition to GPT-3 (175B) using the same experimental setup as in Section 5.3. As shown in Figure 2, our proposed method (Variant II) consistently outperforms CoT prompting in terms of both prediction accuracy and the faithfulness of explanations, even when using smaller LMs.
Conclusion
In conclusion, the proposed approach is a promising solution for utilizing external knowledge to assist LLMs. Unlike traditional methods, RR does not require additional training or fine-tuning, making it a lightweight and feasible option for LLMs. Through extensive experiments on three reasoning tasks using GPT-3, we have shown that RR is able to produce more faithful explanations and improve the performance of LLMs. In the future, we plan to investigate various variations of RR to enhance its effectiveness and efficiency in augmenting LLMs with external knowledge.
References
Appendix A Appendix
In this section, we provide additional details on our experimental setup. Further information can be found in our code.
We adopt the same CoT prompt for commonsense reasoning (i.e., StrategyQA) as those presented in Wei et al. (2022). The CoT prompt for temporal reasoning is provided in Table 6. For tabular reasoning, we adopt the method of Brown et al. (2020) for converting NLI into QA for RTE Dagan et al. (2005), and randomly sample examples from the training data to construct the prompt, as shown in Table 8. The few-shot prompt utilizes the same exemplars as the CoT prompt and does not involve CoT reasoning processes.
A.2 Description of Faithfulness Functions
For a sentence , we denote its MPNet similarity, entailment score, and contradiction score as , , and , respectively. In our experiments, the corresponding thresholds for these scores are , , and . Given the entailment scores, contradiction scores, and MPNet similarities of all supporting facts (denoted as ) in the explanation of a reasoning path , different faithfulness functions can be adopted in different settings as follows:
In Section 4, we employ function (1) for commonsense and tabular reasoning. For temporal reasoning, we use function (2) as the distinct nature of sentences converted from temporal relations leads to unreliable contradiction scores. In Sections 5.3-5.4, we use function (3) for commonsense reasoning with evidence paragraphs, as the high quality of the relevant knowledge negates the need for the complementary use of the MPNet similarity to improve the entailment score.
A.3 Comparison of Retrieval Systems
For commonsense reasoning, we utilized different retrieval systems in Karpukhin et al. (2020) to retrieve relevant paragraphs from Wikipedia. The performance of BM25, DPR, and BM25+DPR were , , and , respectively, indicating that BM25 is the best choice in our case.
A.4 Implementation Details for the Two Variants of RR
In this work, we utilize the information present in the top-ranked output produced by our basic approach as a guide. To this end, we apply a greedy clustering algorithm to group the sentences from all outputs into distinct topic categories based on the cosine similarity of their MPNet sentence embeddings. For each fact in the top-ranked output of our basic approach, we identify the fact with the highest faithfulness within the same topic group and replace it in the output. The faithfulness of a fact is calculated using the function by replacing the supporting facts with a single fact.
Fact generation implementation details.
In this part, we generate questions for the named entities present in each fact of the top-ranked output produced by our basic approach, and retrieve the corresponding answers from the evidence paragraphs using UnifiedQA. We employ the question generation model described in Deutsch et al. (2021), which has been shown to be more extractive compared to other models as demonstrated in Fabbri et al. (2021). We adopt the question filtering approach proposed in Honovich et al. (2021) using an off-the-shelf extractive QA model (ktrapeznikov/albert-xlarge-v2-squad-v2 from Hugging Face Wolf et al. (2020)). We then use an off-the-shelf model (MarkS/bart-base-qa2d from Hugging Face) to convert the generated QA pairs into declarative sentences. We apply simple rules based on the entailment and contradiction scores of the selected facts from the fact selection variant and the generated declarative sentences to obtain the final generated facts.
A.5 Comparison of Different Inference Methods with Supporting Facts
In our experiments, we utilize UnifiedQA for the final step of inference in both variants. However, it is worth noting that GPT-3 could also be used for this purpose. As shown in Table 7, we observe that UnifiedQA performs better at inference with generated facts, while GPT-3 with CoT prompting performs better with empty or gold facts. This suggests that UnifiedQA is more robust to noisy inputs compared to GPT-3. Additionally, both UnifiedQA and GPT-3 with CoT prompting significantly outperform GPT-3 with zero-shot prompting, indicating that the CoT prompting is also beneficial for the final step of inference.