Answering Questions by Meta-Reasoning over Multiple Chains of Thought
Ori Yoran, Tomer Wolfson, Ben Bogin, Uri Katz, Daniel Deutch, Jonathan Berant
Introduction
In chain-of-thought (CoT) prompting, a large language model Brown et al. (2020); Chowdhery et al. (2022); Kadavath et al. (2022); Touvron et al. (2023) is prompted to generate its answer following a step-by-step explanation Wei et al. (2022); Nye et al. (2022). CoT prompting has been shown to dramatically improve performance on reasoning-heavy tasks Kojima et al. (2022); Zhou et al. (2022). Furthermore, Wang et al. (2023) showed that sampling multiple chains of thought and returning their majority output further improves accuracy, a method which they term self-consistency (SC).
While SC leads to performance gains, it also has several shortcomings. First, when the space of possible outputs is large Kalyan et al. (2021), each reasoning chain may lead to a different output, in which case no significant majority will be formed. Second, focusing exclusively on the final output discards relevant information that is present in the intermediate reasoning steps. Consider answering the question “Did Brad Peyton need to know about seismology?” (Fig. 1). Reasoning chain #1 leads to an incorrect answer (“No”), but its steps provide useful information. For example, the intermediate question, and following answer, on “What is seismology?” constitute an important fact that is absent from the other two chains. Last, using SC jointly with chain-of-thought prompting reduces interpretability, as there is no single reasoning chain that can be considered as an explanation.
In this work, we propose Multi-Chain Reasoning (MCR), where we prompt a large language model (LLM) to meta-reason across multiple reasoning chains and produce a final answer, alongside an explanation. Unlike prior work, sampled reasoning chains are used not for their predictions (as in SC) but as a means to collect pieces of evidence from multiple chains. Fig. 1 illustrates MCR compared to SC. While both methods rely on sampling multiple reasoning chains, SC returns the majority answer, “No” (grey box, bottom right). By contrast, MCR concatenates the intermediate steps from each chain (blue boxes, top left) into a unified context, which is passed, along with the original question, to a meta-reasoner model. The meta-reasoner is a separate LLM, prompted to meta-reason on multiple reasoning chains and produce a final answer along with an explanation (pink box, bottom left). By reasoning on multiple reasoning chains, MCR is able to mitigate the aforementioned drawbacks – it combines facts from multiple chains to produce the correct final answer, with an explanation of the answer’s validity.
MCR has three main components (§3). To generate reasoning chains we use two components, a decomposition model and a retriever which jointly generate the chain (Fig. 2), similar to prior work Press et al. (2022); Trivedi et al. (2022a). These chains are then concatenated into a unified multi-chain context which is fed to the aforementioned meta-reasoner. Fig. 1 highlights the ability of the meta-reasoner to combine facts from different reasoning chains (intermediate answers in pink). The output explanation combines facts from each of the three chains: (1) “Seismology is the study of earthquakes”; (2) “San Andreas is a film…”; (3) “Brad Peyton is a film director, writer…”. SC (in grey) errs due to only using the answers, while the meta-reasoner reads entire reasoning chains, and is able to correctly answer the question.
We evaluate MCR on a wide range of challenging multi-hop question answering (QA) datasets, in an open-domain setting. The datasets can be categorized into two types of tasks: implicit reasoning tasks, where reasoning steps are implicit given the question text and need to be inferred using a strategy Tafjord et al. (2019); Geva et al. (2021); Kalyan et al. (2021); explicit reasoning tasks, where a single reasoning strategy exists and can be directly inferred given the language of the question Yang et al. (2018); Welbl et al. (2018); Press et al. (2022); Aly et al. (2021). As our baselines, we compare MCR to SC, as well as to variants of Self-Ask Press et al. (2022) and CoT augmented with retrieval, following Trivedi et al. (2022a). Our results show MCR consistently outperforms all other baselines, in particular, beating SC by up to 5.7%, while using the same reasoning chains (§4).
We analyze the qualities of MCR in §5, by manually scoring its generated explanations and estimating their accuracy. Our analysis shows that MCR generates high quality explanations for over 82% of examples, while fewer than 3% are unhelpful. To conclude, our main contributions are:
We introduce the MCR method for meta-reasoning on multiple chains-of-thought.
We show that MCR outperforms all baselines, including self-consistency, on all 7 multi-hop open-domain QA benchmarks.
We analyze MCR for its explanation quality and its multi-chain reasoning capabilities.
Our data and codebase are publicly available.https://github.com/oriyor/reasoning-on-cots
Background
Recently, there has been a surge of interest in answering multi-hop questions through few-shot prompting of LLMs Wei et al. (2022); Nye et al. (2022); Yao et al. (2022). The majority of these works follow a common standard: First, given a question, plan a step-by-step reasoning chain to derive the answer and solve all intermediate steps, aided by a retriever to minimize model hallucination Khot et al. (2023); Press et al. (2022); Yao et al. (2022); Lazaridou et al. (2023); Trivedi et al. (2022a); Khattab et al. (2022). Then, incorporate multiple reasoning chains with answers to derive the final answer Wang et al. (2023); Li et al. (2022). In our work, we follow this template and focus on the latter part. However, our meta-reasoning approach differs from prior work by reasoning on multiple reasoning chains. Namely, we use multiple chains to collect relevant evidence for question answering.
Method
We present a method for answering questions by meta-reasoning on multiple reasoning chains. Our focus is on open-domain QA, where the input is a question , and the evidence to answer it is found in one or more sentences in a corpus . When answering requires multiple reasoning steps, it can be expressed by a reasoning chain, denoted by . The reasoning chain is a list of one or more intermediate question-evidence-answer triples . Evidence is a sentence that is relevant to answering the intermediate question .
Fig. 2 describes our approach when answering “How many ants would fit into The Shard?”. First, we use a prompted LLM to generate multiple reasoning chains, (steps 1-2). Each is generated by interleaving generated intermediate questions with retrieved contexts (§3.1). Our main contribution is step 3: We introduce a second LLM that is prompted to meta-reason on multiple reasoning chains, collecting evidence facts as its explanation and generating the final answer (§3.2).
Given a question , we generate its reasoning chain using: (1) a decomposition model, and (2) a retriever component. Our reasoning chain generation process is largely based on prior work Press et al. (2022); Trivedi et al. (2022a), discussed in §2. Fig. 3 describes the interleaving of decomposition and retrieval. At each step, the decomposition model generates an intermediate question , based on the original question and the previous reasoning steps. Then, the retriever uses to retrieve relevant evidence . We feed and to the decomposition model (along with the previous steps) to generate intermediate answer . During answer generation, we prepend intermediate evidence sentences to the beginning of the chain rather than interleaving them, as it improves the accuracy for all baselines. For decomposition prompts, see §D.
2 Reasoning over Reasoning Chains
The meta-reasoner module is the core contribution of MCR. Instead of sampling multiple chains for their predicted answers Wang et al. (2023), we utilize them for context generation. This context is fed to a prompted LLM to read the generated chains and reason over them to return the answer.
In §3.1, we defined a reasoning chain as a list of triples. We first sample multiple chains and use all of their intermediate question-answer pairs as our multi-chain context (a variant using question-evidence pairs is described in §B.4). Fig. 2 presents the multi-chain context of the three sampled chains (lower pink box). Next, the multi-chain context and the original question are input to the meta-reasoner. This model is an LLM, few-shot prompted for QA over a multi-chain context. Fig. 4 presents one exemplar from the meta-reasoner prompt for the Feverous dataset (full prompts in §D). We instruct the LLM to “answer the question step-by-step” given its multi-chain context, where each line describes a pair from one of the sampled chains. Next, we append the question and a step-by-step reasoning chain followed by the final answer. This last chain serves as the explanation for solving the question. The meta-reasoner is prompted with 6-10 exemplars, based on the dataset (§4.1).
Providing the meta-reasoner with multiple chains allows it to combine and aggregate facts across chains. Moreover, the model needs to extract the most relevant facts in the chains to serve as its explanation. This enables MCR to be both more accurate and more interpretable than past multi-chain approaches (as we analyze in §5).
Experiments
We compare MCR to existing methods on 7 multi-hop QA benchmarks. These cover a wide range of reasoning skills, including commonsense, composition, comparison and fact verification. MCR consistently outperforms existing approaches on all benchmarks, when experimenting with two different LLMs and retrievers. Our setting is described in §4.1 and we discuss our main results in §4.2.
As our focus is on multi-hop questions (in an open-domain setting), all datasets require multiple reasoning steps. Following prior work Khattab et al. (2022); Trivedi et al. (2022a) and to limit the cost of model API calls, we evaluate on 500-1000 random examples from the development set of each dataset.We use the entire development set for Quartz and Bamboogle, since they include less than 500 examples. For Fermi we use all 286 “Real Fermi Problems” in its train and development sets. Exact numbers are listed in Tab. 2. We also evaluate on the official test sets of StrategyQA and Fermi, as they target implicit reasoning, have multiple valid strategies, and their test set evaluation cost is reasonable. For all datasets, we make sure that no evaluation questions appear in any of our prompts. Tab. 1 has example questions from each dataset. Our multi-hop QA benchmarks can be categorized based on their required reasoning skills:
Implicit Reasoning: Questions that entail implicit reasoning steps Geva et al. (2021). The reasoning steps for solving it cannot be explicitly derived from the language of the question and require commonsense or arithmetic reasoning. Such questions may have multiple valid reasoning chains. We evaluate on: StrategyQA Geva et al. (2021), Fermi Kalyan et al. (2021) and Quartz Tafjord et al. (2019).
Explicit Reasoning: Multi-hop questions where the reasoning steps are explicitly expressed in the language of the question (composition, comparison). These include HotpotQA Yang et al. (2018), 2WikiMQA Welbl et al. (2018) and Bamboogle Press et al. (2022). We also evaluate on Feverous Aly et al. (2021), a fact verification dataset where claims require verifying multiple facts, and evidence may be either in sentences, tables or both.
For evaluation, we use F1 to compare predicted and gold answers for all explicit reasoning datasets and exact-match for the binary-choice datasets. In Fermi, we use the official order-of-magnitude evaluation by Kalyan et al. (2021). We provide additional technical details on evaluation in §A.
1.2 Models
Our main models and baselines are all retrieval-augmented instances of code-davinci-002, prompted with in-context learning exemplars Brown et al. (2020). In §4.3, we include additional experiments with the open-source Vicuna-13B Chiang et al. (2023) LLM. Prompt exemplars are formatted as described in §3.2. The number of exemplars varies from 6-12 between datasets. Decomposition prompt exemplars are based on random examples from the train and development sets, coupled with their gold reasoning chain. For the meta-reasoner exemplars, we use reasoning chains sampled from the decomposition model as the multi-chain context. We ensure that the answer can be inferred using the sampled chains and add an explanation before the final answer, as shown in Fig. 4. For the binary-choice datasets, StrategyQA, Quartz, and Feverous, the prompt contains an equal number of exemplars from each label. For additional details regarding the full prompts, length statistics and robustness to a different choice of prompts, please refer to §D.
We experiment with two variants of the meta-reasoner to measure the effect of reasoning on more than a single chain.
MCR: The meta-reasoner is given five reasoning chains as its multi-chain context (§3.2). We decode one chain with greedy decoding, and sample another four reasoning chains with temperature .Like Wang et al. (2023), we observe that greedy-decoded chains have higher accuracy compared to the other chains. This enables the meta-reasoner to review different pieces of evidence when answering the full question (§5).
SCR: Single-Chain Reasoning (SCR) serves as an ablation for the effect of the multi-chain context. In SCR, the meta-reasoner is given the same prompt as MCR aside from having only the greedy-decoded chain in its context. This disentangles the effect of using multiple chains from the effect of having an LLM that is separate from the decomposition model to generate the final answer.
SA: Self-Ask Press et al. (2022) returns the answer of a single reasoning chain, that was generated with greedy decoding.
SC: Self-Consistency serves as a baseline which incorporates multiple reasoning chains Wang et al. (2023). It returns the majority answer based on multiple chains sampled from the decomposition model. We experiment with variants of 3, 5 and 15 sampled chains (SC, SC and SC), in line with prior work Wang et al. (2023); Khattab et al. (2022); Sun et al. (2023). As in MCR, we use the chain generated with greedy decoding along with additional chains sampled with .
Similar to Press et al. (2022); Lazaridou et al. (2023); Paranjape et al. (2023), our models and baselines use a retriever based on Google Search, via the SerpAPI service.https://serpapi.com/ However, we also include results using an open-source retriever Khattab and Zaharia (2020) in §4.3. As most of our datasets contain evidence from Wikipedia (§4.1.1), we consider it as our retrieval corpus. Therefore, we format search queries as “en.wikipedia.org ”, with the Wikipedia domain preceding the intermediate question. We return the top-1 evidence retrieved by Google. Retrieved evidence may be either sentences or parsed lists. Following Trivedi et al. (2022a), we also retrieve evidence for the original question . Last, all retrieved evidence sentences are prepended to the decomposition (§3.1). Additional implementation details about our retrieval and MCR are described in §B.1 and §B.2.
2 Main Results
Next, we report our evaluation results. Overall, MCR outperforms our baselines on all 7 datasets.
Tab. 2 presents the results for all 7 multi-hop datasets (evaluation described in §4.1.1). We evaluate both SC and MCR using five reasoning chains. In addition, we list an oracle score which uses the best answer out of all five chains. MCR outperforms all baselines on all of the benchmarks, beating SC on StrategyQA (+1.4%), Fermi (+0.6%), Quartz (+4.0%), HotpotQA (+5.7%), 2WikiMQA (+2.5%), Bamboogle (+1.5%) and Feverous (+1.5%).
We measure the gains of MCR and SC when adding reasoning chains. As extending MCR is bounded by context length,code-davinci-002 context is capped at 8,001 tokens. we follow a straightforward approach and perform self-consistency on three MCR runs. We compare this model, MCRSC, which used 15 reasoning chains (5 for each MCR run), to SC. Tab. 3 shows that MCRSC consistently outperforms SC. Furthermore, though MCR uses only 5 reasoning chains, it beats SC on all datasets, save StrategyQA. Fig. 5 plots, for each dataset, the effect that adding more reasoning chains has on meta-reasoning performance. It presents the results with 1 chain (SCR), 5 chains (MCR) and 15 reasoning chains (MCRSC).
We evaluate our models on the official test sets of StrategyQAhttps://leaderboard.allenai.org/strategyqa and Fermi, which include 490 and 558 examples respectively. The results in Tab. 4 show that on StrategyQA MCR consistently beats SC, when using the same number of reasoning chains. In Fermi, both methods perform similarly.
Previously, we established the advantages of meta-reasoning over multiple reasoning chains. While an apples-to-apples comparison with other recent approaches is impossible due to fundamental differences in the experimental setup (see §B.3), it serves as a rough measuring stick for the robustness of MCR across different tasks. In §B.3, Tab. 8 we compare MCR to five recent CoT-based approaches for multi-hop QA. MCR performance is comparable with the best results on all datasets (shared between these works), showcasing its robustness.
3 Open-source Models
To further examine MCR’s performance (§4.2) and for better reproducibility, we experiment with an additional open-source retriever and LLM. As our retriever, we use ColBERTv2 Santhanam et al. (2022) over the 2018 Wikipedia dump from Karpukhin et al. (2020). In addition to code-davinci-002, we experiment with Vicuna-13B Chiang et al. (2023), a 13-billion parameters model shown to outperform LLMs like LLaMA and Alpaca Touvron et al. (2023); Taori et al. (2023). We use the same prompts as in code-davinci-002, trimmed to fit a 2,048 tokens context length.
We report the full results of the open-source ColBERTv2 retriever with code-davinci-002 and Vicuna-13B in Tab. 5. In addition, we provide results of open-source models when reasoning over 15 reasoning chains in Tab. 6. For code-davinci-002, substituting Google Search with ColBERTv2 exhibits the same trend as in Tab. 2, albeit a slight decrease in performance. MCR outperforms all other baselines, beating SC on StrategyQA (+2.3%), Fermi (+3.4%), Quartz (+3.9%), HotpotQA (+3.5%), 2WikiMQA (+1.2%), Bamboogle (+3.6%) and Feverous (+1.4%). Unsurprisingly, results sharply decrease when evaluating the smaller Vicuna-13B with ColBERTv2. The comparison between MCR and SCR suggests that reasoning over multiple chains is a challenge for the weaker Vicuna-13B model. For example, it generates open-ended answers such as “Unknown” or “It depends” for over of the questions in StrategyQA. This suggests that meta-reasoning over multiple chains has greater gains (compared to SCR) when both the decomposition model and meta-reasoner are larger LLMs.
However, even on Vicuna-13B, MCR still outperforms all baselines on 5 datasets and beats SC on all 7 of them: StrategyQA (+0.5%), Fermi (+4.6%), Quartz (+3.6%), HotpotQA (+6.5%), 2WikiMQA (+0.3%), Bamboogle (+3.0%) and Feverous (+1.3%). When evaluating with 15 reasoning chains, in Tab. 6, MCRSC continually beats SC.
Analysis
Next, we measure the importance of incorporating multiple reasoning chains in MCR and qualitatively assess its output.
In §4.2 we observed that MCR consistently outperforms single-chain reasoning (SCR). We wish to prove that this advantage lies in cases where the meta-reasoner uses additional chains. To this end, we sort examples based on the similarity of their greedy-decoded chain to the MCR explanation (details in §C.1). Lower similarity indicates less reliance of MCR on the greedy chain. Fig. 6 presents an example where the MCR explanation (pink box) includes relevant facts from a chain other than the greedy one (additional examples in §C.2). Results in Fig. 7 empirically demonstrate that on StrategyQA, MCR gains over SCR are highest when MCR explanations are less similar to the greedy chain. We observe this trend in all datasets (§C.1), serving as further evidence for MCR’s strengths.
In addition to choosing between reasoning chains, an interesting property of the meta-reasoner is that it can combine facts from different chains. We estimate the prevalence of this phenomenon on the implicit datasets, StrategyQA and Fermi, which are more challenging. Given an example, we automatically check if its meta-reasoner explanation is the result of combining chains. We examine if one of the output sentences appears in exactly one chain, while another sentence is absent from that chain and is part of a different chain. We consider sentences as similar if their ROUGE-1 precision is above 0.8, and distinct if it is below 0.2. Overall, in 20% of StrategyQA examples and 25% of Fermi, the MCR explanation results from combining reasoning chains. From a manual analysis of 50 such examples for each dataset, we observe that these multi-chain explanations are better than any individual reasoning chain in of cases (see examples in §C.2, Fig. 10). For the remaining , the reasoning expressed in the resulting combination is a paraphrase of an individual chain.
The meta-reasoner is prompted to generate an explanation alongside the final answer (§3.2). Inspired by past work Pruthi et al. (2022), we test the quality of the MCR explanations. Four of the authors manually reviewed 600 random examples, 100 per dataset (sans Feverous §B.2) and scored their meta-reasoner explanations. Each explanation is scored as either 1 (irrelevant), 2 (partially relevant) or 3 (highly relevant), based on its relevance to answering the question. We find the explanation is highly relevant in 82% of the cases (87% excluding Fermi, which is the most challenging), and is irrelevant in less than 3%.
Next, we evaluate the faithfulness of explanations Jacovi and Goldberg (2020), namely, whether a person provided only with the question and MCR explanation would answer the same as the model. Our focus was on examples with quality explanations (score 3), since they are answerable given the explanation. We answered each question based on the model’s explanation. In 90% of cases (95% excluding Fermi), the MCR predictions matched our own, highlighting the faithfulness of its explanations. We attribute part of the gap between human and MCR predictions to implicit reasoning tasks, where humans lead by five points, on average. For the full results, see §C.3.
We manually analyzed 700 errors by MCR (100 per dataset). We consider the following categories: Valid predictions where the generated answer is accurate or the original question is ambiguous; Decomposition errors where no chain has the necessary reasoning steps to answer the question; Retrieval errors where the retrieved contexts were irrelevant, leading the model to hallucinate; Explanation errors where MCR generates a wrong explanation while a correct one is present in the multi-chain context; Answer errors are when the MCR explanation is correct, but the answer is not; Contradicting facts are cases where MCR errs due to contrasting statements appearing in the multi-chain context.
Tab. 7 lists the prevalence of the error categories per dataset. In four datasets, over 20% of errors appear to be valid predictions, labeled as incorrect due to ambiguous questions, outdated answers or dataset errors. Decomposition is a challenge in the implicit datasets, StrategyQA and Fermi, with more than 24% of errors. Comparing errors on different reasoning datasets (excluding valid examples): Explanation and Answer errors are 50% on implicit reasoning datasets compared to 23% on explicit reasoning ones; Retrieval errors are more prevalent in explicit reasoning tasks with 66% of errors being due to Retrieval or Contradicting facts, compared to 30% in implicit datasets. Additional technical details on our analysis are in §C.4.
Related Work
For a thorough survey on LLM reasoning see Lu et al. (2022); Huang and Chang (2022); Qiao et al. (2022). A slew of recent works have focused on eliciting multi-step reasoning in LLMs, including scratchpads Nye et al. (2022), chain-of-thought prompting Wei et al. (2022); Zhou et al. (2022), learned verifiers Cobbe et al. (2021), selection-inference Creswell et al. (2022) and bootstrapping Zelikman et al. (2022).
Self-consistency Wang et al. (2023); Fu et al. (2022) selects the majority answer across multiple chains, outperforming learned verifiers and “sample-and-rank” approaches Adiwardana et al. (2020); Freitas et al. (2020). Li et al. (2022) further improve SC by increasing chains’ diversity and introducing a trained verifier. Tafjord et al. (2022) over-samples chains and verifies them using a natural language inference model on intermediate steps, while He et al. (2022) re-rank chains based on intermediate retrieved evidence. In addition, meta-reasoning is closely tied to self-reflection in LLMs, which is becoming increasingly important in using the LLM to review multiple strategies Yao et al. (2023); Shinn et al. (2023); Madaan et al. (2023).
Recent works proposed revising LLM-generated texts by using retrieved sentences Gao et al. (2022) or model-generated feedback Madaan et al. (2023); Chen et al. (2023); Paul et al. (2023). MCR similarly reviews LLM-generated reasoning chains however, its focus is meta-reasoning on multiple chains.
Significant QA research has been dedicated to reasoning over multiple facts retrieved from an underlying corpus. Such tasks include multi-step questions that require explicit reasoning Talmor and Berant (2018); Welbl et al. (2018); Wolfson et al. (2020); Trivedi et al. (2022b), implicit reasoning Geva et al. (2021) and multi-modal capabilities Talmor et al. (2021).
Recent works also target retrieval-augmented LLMs, prompted to solve open-domain questions Lazaridou et al. (2023); Khattab et al. (2022); Trivedi et al. (2022a); Ram et al. (2023); Yoran et al. (2023).
Conclusion
This work introduces MCR for meta-reasoning over multiple reasoning chains. We evaluate MCR on 7 datasets for multi-hop QA that require both implicit and explicit reasoning in an open-domain setting and show that it outperforms previous approaches on all evaluation benchmarks.
Limitations
In this work we introduce a meta-reasoner model to reason over multiple reasoning chains. While we opt for a prompted LLM as our meta-reasoner, we do not experiment with a fine-tuned meta-reasoning model. For the meta-reasoner context, we experiment with variants which include either generated QA pairs or retrieved evidence sentences. We leave further improvements to the meta-reasoner context as future work. Due to the inference costs of current state-of-the-art LLMs we evaluate on the code-davinci-002 model, similar to prior work Trivedi et al. (2022a); Wang et al. (2023). However, to improve the reproducibility of our work we also provide results with an open-source LLM Chiang et al. (2023) and retriever Khattab and Zaharia (2020).
Acknowledgements
We would like to thank Harsh Trivedi, Ofir Press, Mor Geva, Peter Clark and Ashish Sabharwal for their feedback and insightful comments. We thank SerpAPI for their support by granting us an academic discount. This research was partially supported by the Yandex Initiative for Machine Learning and the European Research Council (ERC) under the European Union Horizons 2020 research and innovation programme (grant ERC DELPHI 802800). This work was completed in partial fulfillment of the Ph.D. of Ori Yoran and the Ph.D. of Tomer Wolfson.
References
Appendix A Evaluation
As we prompt LLMs to generate answers, a potential outcome is for the model to abstain from answering the question, by generating Unknown as its answer. Additional cases are when the model generates an end-of-sequence token without any final answer. In the binary-choice datasets, StrategyQA, Quartz and Feverous, we assign a score of 0.5 to such examples, thereby simulating a random guess. When submitting predictions to the StrategyQA test set, we identify cases of model abstains or null predictions beforehand. For these examples, we assign a label of either Yes or No at random. In datasets with open-ended answers, we assign a score of 0 when the predicted answer is either Unknown or null. To make Self-Ask a stronger baseline, when the greedy decoded chain has a null answer, we randomly choose a prediction from one of the other chains. For SC, we do not consider predictions from chains where answers are Unknown or null.
A.2 Fermi
The Fermi dataset requires approximating numeric answers for open-ended questions. Example questions are shown in Tab. 1 and Fig. 2. When providing a Fermi question to our models and baselines we also add the gold answers measure units (e.g. meters, cubes, litres, etc.). While this additional input helps the model, we note that we provide it to all our baselines for a fair comparison with MCR. Nevertheless, even when given the gold units, predicting the final answers to Fermi problems remains highly challenging.
Appendix B Models
For our retrieval, we use the Google Search Engine, via SerpAPI, and return the top-1 retrieved result as an evidence snippet. Snippets can include answer-boxes and tables.https://serpapi.com/organic-results We prepend the page title to the beginning of the snippet, as shown in Fig. 8.
B.2 Implementation Details
We describe the design choices made in our MCR model, such as preforming retrieval on the original question and a variant of the meta-reasoner prompt for Feverous. Due to cost limitations, we evaluate our design choices at a smaller scale and avoid running an exhaustive grid search.
We follow past work Trivedi et al. (2022a) by incorporating retrieved evidence for the original question in addition to evidence retrieved for the intermediate steps (§3.1). This has a positive or negligible effect on most datasets however, it dramatically decreases the results of all models on the Fermi task. Results drop for SA (38.30.7 to 34.70.5), SC (38.30.8 to 34.40.3), SCR (38.10.8 to 34.40.8) and MCR (38.90.8 to 37.00.7). Therefore, our models are run without original question retrieval when evaluated on Fermi. Interestingly, while all models perform roughly the same without original question retrieval, MCR appears better by 2 points when evidence for the original question is used. We hypothesize that it might be due to MCR being somewhat more robust to the addition of irrelevant evidence.
As described in §3.2, the meta-reasoner generates an explanation which precedes the final answer. Feverous is distinct from all other datasets as it require verification of multiple facts in order to verify or disprove a complex statement. When a statement is false, we list one or more of its false intermediate facts along with its correction. For example, in Fig. 4 we list that Robert Broderip lived in Bristol, not London. When prompting the meta-reasoner to list both true and false intermediate facts, we observed a decrease in performance for both MCR (69.41.0 to 66.40.7) and SCR (65.10.4 to 62.90.3). We hypothesize that repeating multiple true facts excessively prompts the model to predict the label “Yes” in cases where most (but not all) of the intermediate facts are correct.
B.3 Empirical Comparison to Recent Approaches
In Tab. 8, we compare MCR to recent CoT-based approaches for multi-hop reasoning. An apples-to-apples comparison is not possible, as these methods do not evaluate on all 7 of our datasets and use varying samples of 500-1,000 dev examples for their evaluation. Moreover, different methods use different retrieval corpora, hyperparameters, prompts and LLMs. Nevertheless, we argue that a direct comparison serves as a measuring stick for MCR’s robustness across multiple datasets, compared to similar solutions.
Evaluation differences include the retrieval corpora, as both IR-CoT and DSP use the official Wikipedia dump provided with the HotpotQA dataset Yang et al. (2018). Our retrieved evidence are from an updated version of Wikipedia, via Google Search. Since certain facts may change over time, this could potentially explain the high percentage of MCR predictions labeled as valid in our error analysis (§5).
We emphasize that our focus is on highlighting the potential of reasoning on reasoning chains. MCR is a method aimed at improving models which generate reasoning chains. Compared to SC, we observe that MCR further boosts the underlying SA model. While task-specific improvements are possible, they are orthogonal to our work.
B.4 Reasoning on Retrieved Evidence
The meta-reasoner answers questions given a multi-chain context of question-answer (, ) pairs, extracted from multiple reasoning chains (§3.2). We experiment with an alternative multi-chain context, comprised of questions and retrieved evidence (, ) (§3.1). This setting resembles past work Trivedi et al. (2022a) however, our sentences are intermediate evidence from multiple reasoning chains, not just the greedy-decoded chain. We compare these variants, MCR-Ev and SCR-Ev, to MCR and SCR that reason on QA pairs. Tab. 9 shows that meta-reasoning on retrieved evidence is less effective. The gap is more evident in implicit reasoning tasks, perhaps due to retrieved evidence being less relevant on average. Example prompts for MCR-Ev and SCR-Ev are listed in §D.
Appendix C Analysis
In §5, we have shown that the advantage of MCR over SCR lies in examples where the meta-reasoner uses chains other than the one generated through greedy decoding. In Fig. 9 we provide the results for all other datasets, in addition to the StrategyQA results in Fig. 7. The similar trend among all datasets is that in examples with lower similarity to the greedy chain, MCR gains over SCR are higher.
The similarity between the meta-reasoner explanation and the greedy decoded reasoning chain is defined as follows: We calculate the ROUGE-1-precision Lin (2004) between the explanation and the chain. Low, Medium, and High are based on thresholds of , , and respectively, with the Identical category indicating an exact match.
C.2 Combining Reasoning Chains
Fig. 10 provides additional examples for combining facts between multiple reasoning chains.
C.3 Explanation Quality Analysis
We provide additional details on the annotation for the scoring meta-reasoner explanations. The annotation was performed by 4 graduate students that are authors of this paper. The annotators were presented with a question and an explanation, and asked to perform two tasks: (a) score the explanation for its quality and (b) answer the question based on the meta-reasoner explanation. We provide the full instructions shown to the annotators in Fig. 11 and the full results in Tab. 10.
C.4 Error Analysis
We provide additional details regarding our error analysis (§5). In less than 5%, we encountered grammatically bad questions which we were unable to comprehend and were therefor discarded from our analysis. For example the HotpotQA question: “What does the goddess associated with the goddess Frigg consists of what tales?”
The input to our meta-reasoner model is a context comprised of pairs, generated by the decomposition model. As the decomposition model is an LLM that is conditioned on retrieved evidence (and prior decomposition steps) it may hallucinate false intermediate answers. In cases of such hallucinations we distinguish between two error types, based on the relevant component. First, Retrieval errors are cases where no relevant information was retrieved, leading to the decomposition model hallucinating an incorrect , passed on to the meta-reasoner’s context. Second, we treat cases where relevant evidence was retrieved, but the decomposition model ignored it and hallucinated an incorrect as Decomposition errors.
Errors stemming from Contradicting Facts, are cases where the meta-reasoner context contains two contradicting facts, one accurate while the other was hallucinated by the decomposition model. For example, Fig. 12 displays an example where the context has contradicting facts on who was the father of Eliezer Ben-Yehuda. When the meta-reasoner has contradicting facts, it is expected to select the correct fact, based on the knowledge encoded in its parameters. Addressing such errors in future work could rely on refining generated text with methods such as RARR Gao et al. (2022).
As our error classes mainly match the MCR components, this error breakdown could potentially help to guide future improvements.
Appendix D Prompts
We provide example prompts for our models for one explicit dataset (2WikiMQA, decomposition: Fig. 13, MCR/SCR: Fig. 15, MCR-Ev/SCR-Ev: Fig. 17) and one implicit dataset (StrategyQA, decomposition: Fig. 14, MCR/SCR: Fig. 16, MCR-Ev/SCR-Ev:Fig. 18). All of our prompts will be released along with our codebase. We use random examples and spend minimal effort on prompt engineering. The number of exemplars varies slightly between dataset and model, with the exact numbers listed in Tab. 11.
D.2 Prompt Statistics
In Tab. 12 we provide statistics of the sequence lengths for all of our models, which include all the decomposition prompts, output decomposition sequences, retrieved evidence and the meta-reasoning prompts. The statistics are for our decomposition model (used by all of our baselines), as well as for the meta-reasoning prompts (used by SCR and MCR). Note that generating a single reasoning chain requires multiple LLM calls, one for each decomposition step. Therefore, a single decomposition generation is generally longer than applying one additional meta-reasoning step.
Results are averaged over multiple runs, corresponding to the results in Tab. 2. Sequence lengths in Tab. 12 correspond to the number of tokens provided by the code-davinci-002 tokenizer.
D.3 Robustness to Choice of Prompt
We empirically measure our method’s sensitivity to the prompt of choice. To this end, we randomly sampled new exemplars for both our decomposition and meta-reasoning prompts for StrategyQA and HotpotQA. When using different random exemplars, we observe that MCR still outperforms all baselines. Even though decomposition performance (SA) is more affected by the set of exemplars, the performance trend remains the same, with MCR being on top. Tab. 13 lists the experiment results, evaluated on 500 examples from each dataset. We also provide the original prompt results in parenthesis (averaged over 3 runs).