Automatic Chain of Thought Prompting in Large Language Models
Zhuosheng Zhang, Aston Zhang, Mu Li, Alex Smola
Introduction
Large language models (LLMs) (Brown et al. 2020; Thoppilan et al. 2022; Rae et al. 2021; Chowdhery et al. 2022) have performed impressively on complex reasoning tasks by decomposing the multi-step problems into intermediate steps before producing the answer. This reasoning process is elicited by a very recent technique: chain-of-thought (CoT) prompting (Wei et al. 2022a).
CoT prompting can be categorized into two major paradigms. One adds a single prompt like “Let’s think step by step” after the test question to facilitate the reasoning chains in LLMs (Kojima et al. 2022). Since this prompting paradigm is task-agnostic and does not need input-output demonstrations, it is called Zero-Shot-CoT (left of Figure 1). With Zero-Shot-CoT, LLMs have shown to be decent zero-shot reasoners. The other paradigm is few-shot prompting with manual reasoning demonstrations one by one (Wei et al. 2022a). Each demonstration has a question and a reasoning chain. A reasoning chain is composed of a rationale (a series of intermediate reasoning steps) and an expected answer. With all the demonstrations being manually designed, this paradigm is referred to as Manual-CoT (right of Figure 1).
In practice, Manual-CoT has obtained stronger performance than Zero-Shot-CoT (Wei et al. 2022a; Kojima et al. 2022). However, this superior performance hinges on the hand-drafting of effective demonstrations. Specifically, the hand-drafting involves nontrivial efforts in designs of both questions and their reasoning chains for demonstrations. Moreover, human efforts for designing task-specific demonstrations are even more: different tasks, such as arithmetic (Roy and Roth 2015) and commonsense reasoning (Talmor et al. 2019), require different ways of demonstrations.
To eliminate such manual designs, we advocate another Auto-CoT paradigm to automatically construct demonstrations with questions and reasoning chains. Specifically, Auto-CoT leverages LLMs with the “Let’s think step by step” prompt to generate reasoning chains for demonstrations one by one, i.e., let’s think not just step by step, but also one by one. However, we find that this challenge cannot be effectively addressed by simple solutions. For example, given a test question of a dataset, retrieving semantically similar questions and invoking Zero-Shot-CoT to generate reasoning chains will fail. Although LLMs are decent zero-shot reasoners, they are not perfect: Zero-Shot-CoT can still make mistakes in reasoning chains.
To mitigate the effect of reasoning chain mistakes from Zero-Shot-CoT, our analysis shows that diversity of demonstration questions is the key. Based on this insight, we propose an Auto-CoT method to automatically construct demonstrations. Auto-CoT consists of two main steps. First, partition questions of a given dataset into a few clusters. Second, select a representative question from each cluster and generate its reasoning chain using Zero-Shot-CoT with simple heuristics.
We evaluate Auto-CoT on ten benchmark reasoning tasks including: (i) arithmetic reasoning (MultiArith (Roy and Roth 2015), GSM8K (Cobbe et al. 2021), AQUA-RAT (Ling et al. 2017), SVAMP (Patel et al. 2021)); (ii) commonsense reasoning (CSQA (Talmor et al. 2019), StrategyQA (Geva et al. 2021)); (iii) symbolic reasoning (Last Letter Concatenation, Coin Flip) (Wei et al. 2022a). Experimental results show that with GPT-3, Auto-CoT consistently matches or exceeds the performance of Manual-CoT that requires manual designs. This indicates that LLMs can perform CoT reasoning by automatically constructing demonstrations.
Related Work
This section reviews two lines of research that form the basis of this work: chain-of-thought (CoT) prompting for multi-step reasoning and in-context learning for inducing LLMs to learn from demonstrations.
CoT prompting is a gradient-free technique of inducing LLMs to produce intermediate reasoning steps that lead to the final answer. Wei et al. 2022a formally studied the topic of CoT prompting in language models. This technique elicits LLMs to generate a coherent series of intermediate reasoning steps that lead to the final answer to a question. Studies have shown that LLMs can perform CoT reasoning with zero-shot prompting (Zero-Shot-CoT) (Kojima et al. 2022) or manually written few-shot demonstrations (Manual-CoT) (Wei et al. 2022a).
Kojima et al. 2022 showed that LLMs are decent zero-shot reasoners whose generated rationales have already reflected the CoT reasoning. This finding inspires our work to leverage the self-generated rationales for demonstrations. Generating rationales by LLMs was shown to be practical in a recent work (Zelikman et al. 2022). In their work, an LLM is prompted to generate rationales and those rationales that lead to the correct answer are selected. The selection requires a training dataset of questions with annotated answers. In contrast, our work considers a more challenging scenario where only a set of test questions are given (without a training dataset), following CoT prompting studies by Wei et al. 2022a and Kojima et al. 2022.
Manual-CoT.
Manual-CoT achieves stronger performance by eliciting the CoT reasoning ability with effective manual demonstrations. The demonstrations for the reasoning process are manually designed. However, the human efforts in designs of both questions and their reasoning chains are nontrivial. Instead of addressing this limitation, recent studies mainly focus on hand-crafting more complex demonstrations or leveraging ensemble-like methods. One trend is problem decomposition. In least-to-most prompting (Zhou et al. 2022), complex problems are reduced to sub-problems, and then the sub-problems are solved sequentially. The other trend is to vote over multiple reasoning paths for a test question. Wang et al. 2022a introduced a self-consistency decoding strategy to sample multiple outputs of LLMs and then took a majority over the final answers. Wang et al. 2022b and Li et al. 2022 introduced randomness in the input space to produce more diverse outputs for voting. They used manually-designed demonstrations as the seed set and generated additional rationales: leave one question from the seed set and use the remaining demonstrations to generate rationales for this question by the LLM. Unlike the aforementioned research lines that rely on manually-designed demonstrations, our work intends to eliminate manual designs with competitive performance.
2 In-Context Learning
CoT prompting is closely related to in-context learning (ICL) (Radford et al. 2019; Brown et al. 2020). ICL enables LLMs to perform a target task by feeding a few prompted examples as part of the input. Without gradient update, ICL allows a single model to perform various tasks universally. There are various research lines to improve the performance of ICL: (i) retrieving related demonstrations to the test instance where the popular practice is dynamically retrieving related training examples for a given test input (Rubin et al. 2022; Su et al. 2022); (ii) augmenting with fine-grained information, such as incorporating task instruction (Mishra et al. 2022; Wei et al. 2022b; Sanh et al. 2022); (iii) manipulating output probabilities of LLMs instead of directly computing the likelihood of target labels (Holtzman et al. 2021; Zhao et al. 2021; Min et al. 2022a).
Despite the success of ICL, studies (Liu et al. 2022a; Lu et al. 2022) have shown that the strength of ICL may vary widely depending on the choice of in-context demonstrations (Liu et al. 2022b). In detail, the formatting of the prompt, such as wording or order of demonstrations, may lead to performance fluctuations (Webson and Pavlick 2022; Zhao et al. 2021). A recent work (Min et al. 2022b) even questioned the necessity of ground-truth input-output mapping: using incorrect labels in the examples only marginally lowers the performance. However, the existing analysis of ICL is mainly based on standard classification and multi-choice datasets that only have simple output> mappings. We discover that those findings may not be applicable to the CoT prompting scenario with more complex rationaleoutput> mappings. For example, mistakes in either the rationale> mapping or the
Challenge of Auto-CoT
As just discussed, the performance of ICL hinges on hand-crafted demonstrations. As reported in Manual-CoT (Wei et al. 2022a), using demonstrations written by different annotators brings up to 28.2% accuracy disparity in a symbolic reasoning task, while changing the order of demonstrations results in less than 2% changes in most tasks. This suggests that the key challenge of Auto-CoT lies in automatically constructing demonstrations with good questions and their reasoning chains.
Recall that Manual-CoT hand-crafts a few (e.g., 8) questions in demonstrations. With similarity-based retrieval methods being widely adopted for prompting LLMs (Rubin et al. 2022; Su et al. 2022), a promising candidate solution is to sample demonstration questions using similarity-based retrieval. We follow the more challenging assumption in CoT studies (Wei et al. 2022a; Kojima et al. 2022) that only a set of test questions are given (without a training dataset). Following Liu et al. 2022a, we use Sentence-BERT (Reimers and Gurevych 2019) to encode questions. For each question in a test dataset, we sample demonstration questions () from the rest of the questions. We design a Retrieval-Q-CoT method to retrieve the top- (e.g., ) similar questions based on cosine similarity. To compare with this similarity-based method, we also test a relatively more diversity-based method: Random-Q-CoT, which randomly samples other test questions for each test question.
Both Retrieval-Q-CoT and Random-Q-CoT invoke Zero-Shot-CoT (Kojima et al. 2022) to generate the reasoning chain (rationale and answer) for each sampled question , as LLMs are decent zero-shot reasoners (Kojima et al. 2022). We use GPT-3 (Brown et al. 2020) with 175B parameters (text-davinci-002) for the LLM unless otherwise stated. On a high level, both Retrieval-Q-CoT and Random-Q-CoT take the concatenation of pairs () and as input to predict the reasoning chain for , which contains the answer in the end (like right of Figure 1).
To our surprise, Retrieval-Q-CoT underperforms Random-Q-CoT on the arithmetic dataset MultiArith (Roy and Roth 2015) (Table 1). Note that the retrieval methods were originally proposed in tasks with annotated labels (Rubin et al. 2022; Su et al. 2022), however, invoking Zero-Shot-CoT does not guarantee entirely correct reasoning chains. Thus, we hypothesize that the inferior performance of Retrieval-Q-CoT is caused by incorrect reasoning chains by Zero-Shot-CoT. To test this hypothesis, we experiment with Retrieval-Q-CoT on two other datasets GSM8K (Cobbe et al. 2021) and AQuA (Ling et al. 2017) that have training sets with annotated reasoning chains. The results are shown with in Table 1. Under the setting with annotated reasoning chains, Retrieval-Q-CoT even outperforms Manual-CoT. The result indicates that Retrieval-Q-CoT is effective when human annotations are available.
Although human annotations are useful, such manual efforts are nontrivial. However, automatically generating reasoning chains via Zero-Shot-CoT underperforms Manual-CoT, especially when the challenge of question sampling is not addressed. To design more effective Auto-CoT, we need to understand its challenge better.
Since Retrieval-Q-CoT uses a few prompting demonstrations like in Manual-CoT, Retrieval-Q-CoT is expected to perform competitively as well. However, reasoning chains (both rationales and answers) in Retrieval-Q-CoT are generated by Zero-Shot-CoT: they may have mistakes that lead to wrong answers. Let us simply call demonstrations with wrong answers as wrong demonstrations. Intuitively, after similar questions to a test question are retrieved, wrong demonstrations caused by Zero-Shot-CoT may mislead the same LLM to reason similarly with a wrong answer (e.g., replicating mistakes) for the test question. We refer to this phenomenon as misleading by similarity. We will investigate whether misleading by similarity contributes to the inferior performance of Retrieval-Q-CoT.
To begin with, we invoke Zero-Shot-CoT on all the 600 questions from the MultiArith dataset. Among them, we collect those 128 questions (denoted as ) where Zero-Shot-CoT generates wrong answers (error rate: ). As we mentioned, with extra demonstrations, Retrieval-Q-CoT and Random-Q-CoT are expected to perform more competitively than Zero-Shot-CoT. Among where Zero-Shot-CoT fails, we call those where Retrieval-Q-CoT or Random-Q-CoT still fail as their unresolved questions. We divide the number of unresolved questions by 128 (number of questions in ) to calculate the unresolving rate. A higher unresolving rate means that a method more likely still makes mistakes like Zero-Shot-CoT. Figure 2 shows that the unresolving rate of Retrieval-Q-CoT (46.9%) is much higher than Random-Q-CoT (25.8%). It indicates that with similar questions being sampled for test questions, Retrieval-Q-CoT is negatively affected by misleading by similarity.
To show that unresolved questions of Retrieval-Q-CoT tend to be similar, we present a case study in Table 2. In the left part, the retrieved demonstration questions are similar to the test question and ask “how long will it take him to cook the rest?” The reasoning chains generated by Zero-Shot-CoT produce answers regarding “the total of” instead of “the rest”. Following the demonstrations, Retrieval-Q-CoT also fails by misunderstanding the meaning of “the rest”. In contrast, Random-Q-CoT correctly understands “the rest” better without making similar mistakes in the demonstrations, thanks to relatively more diverse (random) demonstrations.
2 Errors Frequently Fall into the Same Cluster
Motivated by the observations in Table 2, we use -means to partition all the 600 test questions into clusters, where each cluster contains similar questions. We use Sentence-BERT (Reimers and Gurevych 2019) to encode questions and apply -means for clustering. With these clusters and reasoning chains generated by Zero-Shot-CoT (in Section 3.1), now we are curious if certain clusters contain questions where Zero-Shot-CoT frequently fails. Thus, we calculate the error rate (questions with wrong Zero-Shot-CoT answers / total questions) for each cluster.
As shown in Figure 3, there exists a cluster (Cluster 2) with frequent Zero-Shot-CoT errors (52.3%). The phenomenon could be generic as Zero-Shot-CoT may lack some skills to solve some common problems in target tasks. We observe similar phenomena when changing the cluster number or using other datasets (Appendix A.2). For convenience of descriptions, let us call the cluster with the highest error rate as the frequent-error cluster (e.g., Cluster 2 in Figure 3). Therefore, the imperfect nature of generated reasoning chains in a zero-shot fashion poses risks of retrieving multiple similar questions inside a frequent-error cluster by using similarity-based methods. For the test question in the frequent-error cluster, Retrieval-Q-CoT more easily constructs demonstrations with multiple similar mistakes. As a result, Retrieval-Q-CoT often makes similar mistakes like Zero-Shot-CoT, reiterated by its higher unresolving rate in Figure 2.
3 Diversity May Mitigate Misleading by Similarity
The analysis so far compellingly shows that LLMs are still not perfect zero-shot reasoners; thus, we aim to mitigate the effect of their Zero-Shot-CoT errors, especially to mitigate misleading by similarity in the design of Auto-CoT.
As we will show later (Section 5.5), presenting a small portion of mistakes (e.g., 1 or 2 wrong demonstrations out of 8) would not harm the overall reasoning performance for test questions. Suppose that questions of all the wrong demonstrations fall into the same frequent-error cluster; then sampling one question from every different cluster will lead to a higher than chance to construct all the 8 correct demonstrations. Since different clusters reflect diverse semantics of the questions, this clustering-based sampling method can be considered as diversity-based, which is in sharp contrast to similarity-based Retrieval-Q-CoT. On one hand, sampling questions with diversity may mitigate the effect of misleading by similarity (Section 3.1). On the other hand, if we took each demonstration as a kind of skill, diverse demonstrations seem to cover more alternative skills for solving target questions: even though there still exists a small portion (e.g., ) of mistakes in the demonstrations, the performance will not be negatively affected (to be shown in Figure 5.5).
Nevertheless, the clustering-based sampling method may still construct a small portion of wrong demonstrations, such as from questions in the frequent-error cluster. As we will show later, some of these wrong demonstrations may be eliminated with heuristics. For example, wrong demonstrations often come with long questions and long rationales. Using simple and generic heuristics, such as only considering shorter questions with shorter rationales, further helps mitigate the effect of imperfect Zero-Shot-CoT capabilities (Appendix C.2).
Auto-CoT: Automatic Chain-of-Thought Prompting
Based on the observations and considerations in Section 3, we propose an Auto-CoT method to construct demonstrations with questions and reasoning chains automatically. Auto-CoT consists of two main stages: (i) question clustering: partition questions of a given dataset into a few clusters; (ii) demonstration sampling: select a representative question from each cluster and generate its reasoning chain using Zero-Shot-CoT with simple heuristics. The overall procedure is illustrated in Figure 4.
Since diversity-based clustering may mitigate misleading by similarity (Section 3.3), we perform cluster analysis for a given set of questions . We first compute a vector representation for each question in by Sentence-BERT (Reimers and Gurevych 2019). The contextualized vectors are averaged to form a fix-sized question representation. Then, the question representations are processed by the -means clustering algorithm to produce clusters of questions. For questions in each cluster , sort them into a list in the ascending order of the distance to the center of cluster . This question clustering stage is summarized in Algorithm 1.
2 Demonstration Sampling
In the second stage, we need to generate reasoning chains for those sampled questions and then sample demonstrations that satisfy our selection criteria.
More concretely, we construct a demonstration (concatenation of a question, a rationale, and an answer) for each cluster (). For cluster , we iterate over questions in the sorted list (obtained by Algorithm 1) until satisfying our selection criteria. In other words, a question that is closer to the center of cluster is considered earlier. Say that the -th closest question is being considered. A prompted input is formulated as: [Q: . A: [P]], where [P] is a single prompt “Let’s think step to step”. This formed input is fed into an LLM using Zero-Shot-CoT (Kojima et al. 2022) to output the reasoning chain consisting of the rationale and the extracted answer . Then, a candidate demonstration for the -th cluster is constructed by concatenating the question, rationale, and answer: .
Similar to the criteria of the hand-crafting demonstrations in Wei et al. 2022a, our selection criteria follow simple heuristics to encourage sampling simpler questions and rationales: set the selected demonstration as if it has a question with no more than tokens and a rationale with no more than reasoning steps. Because Zero-Shot-CoT often uses “n” for separating the reasoning steps, the rule can be easily implemented by counting the “n” tokens in the generated rationales.
As summarized in Algorithm 2, after demonstration sampling for all the clusters, there will be constructed demonstrations . The constructed demonstrations are used to augment a test question for in-context learning. Specifically, the input is the concatenation of all the demonstrations followed by [Q: . A: [P]]. This input is fed to LLMs to obtain the reasoning chain with the answer in the end for (right of Figure 4).
Experiments
We briefly describe the experimental setup and present main experimental results. More experimental details and results can be found in the appendices.
Our method is evaluated on ten benchmark datasets from three categories of reasoning tasks: (i) arithmetic reasoning (MultiArith (Roy and Roth 2015), GSM8K (Cobbe et al. 2021), AddSub (Hosseini et al. 2014), AQUA-RAT (Ling et al. 2017), SingleEq (Koncel-Kedziorski et al. 2015), SVAMP (Patel et al. 2021)); (ii) commonsense reasoning (CSQA (Talmor et al. 2019), StrategyQA (Geva et al. 2021)); (iii) symbolic reasoning (Last Letter Concatenation, Coin Flip) (Wei et al. 2022a).
Implementation.
We use the public GPT-3 (Brown et al. 2020) of the text-davinci-002 version with 175B parameters for the LLM (Ouyang et al. 2022) unless otherwise stated. We select this LLM because it has the strongest CoT reasoning performance among public LLMs, as reported in Kojima et al. 2022 and Wei et al. 2022a. We also evaluate the Codex model (Chen et al. 2021) (code-davinci-002) as the LLM. Following Wei et al. 2022a, the number of demonstrations is 8 except for AQuA and Letter (4), CSQA (7), and StrategyQA (6).
Baselines.
We compare our methods with four baseline methods: Zero-Shot (Kojima et al. 2022), Zero-Shot-CoT (Kojima et al. 2022), Few-Shot (Wei et al. 2022a), and Manual-CoT (Wei et al. 2022a). Zero-Shot-CoT and Manual-CoT are illustrated in Figure 1. The Zero-Shot baseline concatenates a test question with the prompt “The answer is” as the LLM input. The Few-Shot baseline has the same LLM input as Manual-CoT except for removed rationales from all the demonstrations.
2 Competitive Performance of Auto-CoT on Ten Datasets
Table 3 compares accuracy on ten datasets from three categories of reasoning tasks. The Zero-Shot and Zero-Shot-CoT results are taken from Kojima et al. 2022, the Few-Shot and Manual-CoT results are taken from Wei et al. 2022a, and the Auto-CoT results are averaged over three random runs. Overall, Auto-CoT consistently matches or exceeds the performance of the CoT paradigm that requires manual designs of demonstrations. Due to the cost of manual designs, Manual-CoT may design the same demonstrations for multiple datasets (e.g., of the arithmetic datasets). In contrast, Auto-CoT is more flexible and task-adaptive: every single dataset gets its own demonstrations that are automatically constructed.
3 Visualization of Question Clustering
Figure 5 visualizes question clustering (with PCA projection) in ten datasets. The illustration indicates that there exist generic patterns, where different patterns may be characterized by questions from different clusters. We present the constructed demonstrations of Auto-CoT in Appendix D.
4 General Effectiveness Using the Codex LLM
To evaluate the general effectiveness of Auto-CoT using different LLMs, here we change the LLM to the Codex model (Chen et al. 2021). As in Table 4, the Codex LLM leads to performance improvement for Manual-CoT when compared with Table 3 that uses the GPT-3 (text-davinci-002) LLM. Nonetheless, using the Codex LLM, the overall performance of Auto-CoT is still competitive compared to Manual-CoT, providing additional empirical evidence for the effectiveness of Auto-CoT.
5 Effect of Wrong Demonstrations
Recall our discussions in Section 3.3 that there can be wrong demonstrations (whose answers are wrong). To see if diversity mitigates this effect, we design an In-Cluster Sampling baseline that constructs demonstrations by randomly sampling questions from the same cluster that contains a test question. Figure 5.5 compares accuracy with varying amounts of wrong demonstrations on MultiArith. Compared with In-Cluster Sampling, Auto-CoT (using diversity-based clustering) is less affected by wrong demonstrations: its performance still does not degrade significantly even when presented with 50% wrong demonstrations.
6 More Challenging Streaming Setting
CoT studies commonly assume that a full dataset with test questions is given (Wei et al. 2022a; Kojima et al. 2022). Based on the given dataset, Auto-CoT samples questions to construct the demonstrations. Nonetheless, now we consider a more challenging streaming setting where a small batch of test questions (say questions) arrive at a time like in data streams.
To address this challenge, we extend Auto-CoT to a bootstrapping version Auto-CoT*: (i) Initialize an empty set ; (ii) When batch of questions arrive, invoke Zero-Shot-CoT (no clustering due to small ) for each to obtain its reasoning chain . Add question-chain pairs to and call the new set ; (iii) When batch () of questions arrive, construct demonstrations with existing questions and reasoning chains in (like Auto-CoT) and use the demonstrations for in-context reasoning for each . Add question-chain pairs to and call the new set .
Figure 5.5 compares the accuracy on MultiArith at each batch () in this streaming setting (extended version: Figure 11 in the Appendix). As expected, for batch , Auto-CoT* and Zero-Shot-CoT obtain equal accuracy. From batch , Auto-CoT* performs comparably with Manual-CoT. This result indicates that our method is still effective in the more challenging streaming setting.
Conclusion
LLMs have shown reasoning capabilities with CoT prompting. The superior performance of Manual-CoT hinges on the hand-crafting of demonstrations. To eliminate such manual designs, we proposed Auto-CoT to automatically construct demonstrations. It samples questions with diversity and generates reasoning chains to construct demonstrations. Experimental results on ten public benchmark reasoning datasets showed that with GPT-3, Auto-CoT consistently matches or exceeds the performance of the CoT paradigm that requires manual designs of demonstrations.
References
Appendix A Extended analysis for the challenge of Auto-CoT
A demonstration is a triple composed by
In contrast, shuffling either rationales or answers reduces the accuracy significantly (). The observation indicates that the rationale-answer consistency is critical. This kind of mismatch actually happens in Zero-Shot-CoT. An example is shown in Table 6. Using such demonstrations might teach the model illusion—predicting answers without basis.
A.2 Observation of frequent-error clusters
To verify if Zero-Shot-CoT fails at similar problems, we cluster the questions into a few clusters and calculate the error rate of the answers to the questions in each cluster. As shown in Figure 8, the mistakes tend to gather in one or more clusters across different datasets. We observe a similar phenomenon when the cluster number changes, as shown in Figure 9. The phenomenon has shown to be generic that Zero-Shot-CoT may lack some skills to solve some common problems in target tasks. We call the cluster with the highest error rate as a frequent-error cluster. Therefore, the imperfect nature of generated reasoning chains poses risks of retrieving a set of similar questions inside the frequent-error cluster for similarity-based retrieval.
Appendix B Experimental Details
Our method is evaluated on ten benchmark datasets that cover arithmetic reasoning, commonsense reasoning, and symbolic reasoning tasks. The statistics of the datasets are shown in Table 7.
For arithmetic reasoning, we consider the following six datasets: (i) MultiArith [Roy and Roth 2015], (ii) GSM8K [Cobbe et al. 2021], (iii) AddSub [Hosseini et al. 2014], (iv) AQUA [Ling et al. 2017], (v) SingleEq [Koncel-Kedziorski et al. 2015], and (vi) SVAMP [Patel et al. 2021]. The first three are from the classic Math World Problem Repository [Koncel-Kedziorski et al. 2016], and the last three are from more recent benchmarks.
Commonsense Reasoning.
For commonsense reasoning, we use (i) CommonsenseQA (CSQA) [Talmor et al. 2019] and (ii) StrategyQA [Geva et al. 2021]. CommonsenseQA asks questions with complex semantics that often require reasoning based on prior knowledge [Talmor et al. 2019]. StrategyQA requires models to infer an implicit multi-hop reasoning to answer questions [Geva et al. 2021].
Symbolic Reasoning.
For symbolic reasoning, we use (i) Last Letter Concatenation [Wei et al. 2022a] and (ii) Coin Flip tasks [Wei et al. 2022a]. Last letter Concatenation requires the model to concatenate the last letters of each word. The goal of Coin Flip is to answer whether a coin is still heads up after people either flip or do not flip the coin.
B.2 Implementation Details
We use GPT-3 [Brown et al. 2020] of the text-davinci-002 version with 175B parameters for the LLM [Ouyang et al. 2022] unless otherwise stated. We select the model because it is public and is widely used to assess the ability of CoT reasoning in LLMs [Wei et al. 2022a, Kojima et al. 2022]. The model is accessed via the OpenAI API. Our experiments are run between July-2022 and September-2022 by using OpenAI API. Greedy decoding is used to generate the output. We set max_tokens = 256 and temperature = 0. Following Wei et al. 2022a, the number of demonstrations used for in-context learning is 8 in most tasks, except for 4 in AQuA and Last Letter Concatenation, 7 in CSQA, and 6 in StrategyQA.
Appendix C Analysis
We compare different ways of sorting questions in each cluster, including: (i) minimal distance to the cluster center (In-Cluster Min Dist, as adopted in Auto-CoT), (ii) maximal distance to the cluster center (In-Cluster Max Dist), and (iii) random sampling inside the cluster (In-Cluster Random). To alleviate the influence of wrong demonstrations, we only sample the demonstrations with correct answers for this analysis.
Comparing the results in Table 8, we see that the demonstrations are generally better if they are closer to the cluster center.
C.2 Effectiveness of the simple heuristics
In Section 4, we apply simple heuristics to encourage the model to use simple and accurate demonstrations. Similar to the criteria of the hand-crafting demonstrations in Wei et al. 2022a, our selection criteria follow simple heuristics to encourage sampling simpler questions and rationales: set the selected demonstration as if it has a question with no more than tokens and a rationale with no more than reasoning steps. Because Zero-Shot-CoT often uses “n” for separating the reasoning steps, the rule can be easily implemented by counting the “n” tokens in the generated rationales. For arithmetic reasoning tasks except for AQuA (because it is a multiple-choice problem), we require that is not empty and appears in In arithmetic reasoning tasks, rationales often infer their answers at the last few tokens, as demonstrations examples shown in Appendix D. to mitigate the risk of rationale-answer mismatches (as we find that such mistakes are harmful in Appendix A.1). If the question, rationale, and the answer satisfy the conditions above, then a candidate demonstration for the -th cluster is constructed by concatenating the question, rationale, and answer: .
We run the demonstration construction process three times before and after using simple heuristics to quantify its effect. Table 9 shows the comparison. The simple heuristics reduce the average number of wrong rationales in constructing demonstrations. Figure 10 further depicts the error rate with and without the simple heuristics. The error rate is computed by the average number of wrong rationales divided by the number of demonstrations. We see that our method can keep the error rate below 20% in most tasks ().
C.3 Extended: More Challenging Streaming Setting
In Section 5.6, we discussed the application of Auto-CoT in a more challenging streaming setting where a small batch of test questions (say questions) arrive at a time like in data streams.
Due to page space limits, we only showed the results of the first 10 batches (300 test questions in total) in Section 5.6. In Figure 11, we illustrate the accuracy of each batch on all the 600 test questions in MultiArith. As expected, for batch , Auto-CoT* and Zero-Shot-CoT obtain equal accuracy. From batch , Auto-CoT* quickly performs comparably with Manual-CoT. This result indicates that our method is still effective in the more challenging streaming setting.