Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
Yihe Deng, Weitong Zhang, Zixiang Chen, Quanquan Gu
Introduction
Misunderstandings in interpersonal communications often arise when individuals, shaped by distinct subjective experiences, interpret the same message differently. In social science, such phenomena can be attributed to cognitive biases in frames in thought (Druckman, 2001). A frame represents an individual’s scheme of interpretation, enabling their understanding and response to an input (Erving, 1974). A single message, framed in different ways, can lead individuals to different conclusions. People habitually project their frames onto their received information, and only shift these frames when incongruence arises. Recently, Large Language Models (LLMs), such as the GPT series (Radford et al., 2019; Brown et al., 2020; OpenAI, 2023), have witnessed a surge in popularity due to their profound impact on various real-world applications, including question answering (Lu et al., 2023), code generation (Poesia et al., 2022), and conversational agents (Bozkurt, 2023). The wide applicability and efficacy of these models have led to rapidly growing research on understanding and improving the use of LLMs. In this work, we posit that LLMs also exhibit their own frames in thought, and it is not uncommon to observe a disparity between the frames used by humans and LLMs. It is widely acknowledged that the quality of the prompt generated by human critically influences the response quality of the LLMs, emphasizing the importance of effective queries that prioritize specificity, detail, and precision (OpenAI, 2022). However, because of an individual’s unique frame of thought, it can be challenging for humans to assess the clarity of their questions and to align their frames with those of LLMs. To illustrate this, we first present a motivating example by investigating a recent work (Allen-Zhu and Li, 2023) in detail.
In Allen-Zhu and Li (2023), the authors reported an important finding: LLMs such as GPT-4 may not efficiently reason with their internal knowledge even if they can retrieve information accurately. As shown in Figure 1, when posed with the query, “Was Mother Teresa born on an even month?” GPT-4 might mistakenly assert that August is an odd month. Based on this observation, Allen-Zhu and Li (2023) suggested that GPT-4 instead requires a Chain-of-Thought process—relying on user-led follow-up questions—to correct its previous wrong answers. When posed with the follow-up question “Do you know what even means?”, GPT-4 will correct itself. However, we take a step further to investigate the intrinsic reason for LLM’s inefficiency in answering such questions. As shown in the other three conversations in Figure 1, when GPT-4 explains its reasoning, it appears that the model has several ambiguities toward the questions. For example, it may consider February as odd due to its irregular number of days and sometimes consider an even/odd month to be months with an even/odd number of days.
Ambiguity in questions is a recognized concern in benchmark datasets. For instance, it has been observed that the NLI datasets such as MultiNLI (Williams et al., 2018) contains ambiguities, which are challenging even for human interpreters (Liu et al., 2023). Furthermore, our study uncovers that benchmark datasets commonly used for LLM evaluation (Wei et al., 2022; bench authors, 2023) possess ambiguities that are imperceptible to humans but challenging for language models. These ambiguities cause LLMs to provide mistaken responses to unintended queries. To address this issue, it is imperative to reduce ambiguity and contextualize information in a way that aligns with the existing frame of the LLMs.
In this paper, we highlight an often-overlooked aspect of studies in LLMs: the disparity between human and LLM thought frames. Our research illustrates that this disparity significantly impacts the performance of LLMs. To tackle this problem, we propose to let the LLM to rephrase the question and incorporate additional details for better answering. We observe that, as opposed to questions asked casually by human, the rephrased questions tend to enhance semantic clarity and aid in resolving inherent ambiguity. For example, the classification questions in Allen-Zhu and Li (2023) tend to be short. Upon rephrasing by the LLM itself, the newly generated question is more detailed and has a clearer question format, as presented in Figure 2. This self-rephrasing technique leads to significant improvement in accuracy compared with Allen-Zhu and Li (2023), as shown in the barplot of Figure 2. While GPT-4 indeed found the original questions challenging, it demonstrates the ability to effectively answer the rephrased questions it generates.
Building upon these insights, we introduce a method named Rephrase and Respond (RaR), which prompts the LLM to rearticulate the given question and respond in a single prompt. In addition to the simple RaR prompt, we also present a variation called Two-step RaR. Two-step RaR employs a rephrasing LLM to generate reworded questions that can be made available to any responding LLM. Our empirical results across diverse reasoning tasks show the effectiveness of both approaches. Notably, Two-step RaR facilitates the transfer of rephrased questions from more capable LLMs to clarify ambiguities for less advanced models. We also present both theoretical and empirical comparisons with the Chain-of-Thought (CoT) method (Kojima et al., 2022; Wei et al., 2022). On the one hand, like CoT, RaR is compatible with the black-box nature of the current powerful GPT-3.5/4 that operate through API services. On the other hand, while CoT focuses on augmentations either at the beginning or the end of a query, RaR directly modifies the query itself. Therefore, RaR is complimentary to CoT and can be easily combined for improvement, as confirmed by our experimental results. Furthermore, unlike methods that employ multiple LLMs for iterative prompt engineering based on accuracy scores (Zhou et al., 2022b; Pryzant et al., 2023), our method is both unsupervised and training-free, making it economical and applicable to all questions. Lastly, our work call forth the importance that the design of human-crafted tasks targeting specific LLM capabilities should be rigorously reviewed by both humans and LLMs to ensure clarity in intention.
The remainder of this paper is organized as follows. Section 2 introduces the RaR method in detail, including One-step RaR (Section 2.1) and Two-step RaR (Section 2.2). In Section 3, we present extensive empirical evaluations of the RaR method, including various benchmark tasks (Section 3.1), the performance on GPT-4 (Section 3.2) and other GPT models (Section 3.3). We also discuss the use of multiple rephrasing processes (Section 3.4) in this section. In Section 4, we compare RaR with CoT in detail both theoretically and experimentally. The related works are discussed in Section 5 and the conclusion is drawn in Section 6.
Rephrase and Respond
In this section, we introduce our proposed method in detail, encompassing two principled approaches, namely One-step RaR and Two-step RaR to facilitate better responses from LLMs by letting themselves rephrase the questions. In the following presentation, for the sake of simplicity, we will use RaR to refer to One-step RaR, unless a specific distinction is necessary.
In interpersonal communication, rephrasing is a commonly known technique. People rephrase another person’s question as a process of understanding, to ensure clarity and coherence in responding. Such a communication strategy can be similarly applied to an LLM, letting it generate a rephrased question first and provide an answer subsequently. Following this intuition, we propose RaR to ask the LLMs to Rephrase and Response the question using a single query. This approach can be viewed as a strategy to directly enhance the quality of the LLM’s response. In detail, we introduce the following prompt for the question-answering task:
As we will show in experiments, GPT-4 can achieve much better results using RaR prompt (2.1) across a wide range of tasks, and especially on human-crafted datasets that exhibit ambiguity to LLMs.
2 Two-step RaR: Rephrase the Question and Respond to the Rephrased Question
To further leverage the quality improvement of the questions rephrased by larger models, like GPT-4, we introduce a variation of RaR called Two-step RaR. Intuitively, even among humans, a more detailed and precise question elicits in more accurate and decisive responses. Two-step RaR follows this intuition by designing a two-step procedure to improve the quality of the questions: in the first step, given a query question, we generate a self-rephrased query rephrased_question by prompting a rephrasing LLM with the following prompt:
Then the original question and the rephrased question are combined to prompt a responding LLM with the following prompt:
Notably, the rephrasing LLM and the responding LLM can be either the same or different models. As we will show later in experiments, different LLMs exhibit distinct proficiency in question rephrasing. In particular, a question rephrased by GPT-4 can help a weaker LLM like Vicuna to produce more accurate responses.
The enhancement in the response quality can be leveraged to improve the benchmark datasets for a fairer evaluation of LLMs: Existing benchmark datasets, crafted by humans, are designed to assess the performance of LLMs across various reasoning skills. However, as demonstrated in our examples in Figure 2, these questions may lack the necessary clarity to fully showcase the specific abilities of LLMs. By the ‘Rephrase’ step in Two-step RaR, we can universally improve the question quality and enable a fairer comparison.
In addition, compared with the prompt of One-step RaR, in the Two-step version, we maintain the original context by including the user’s question, while adding the LLM-rephrased question to help better understanding. This prevents the possible divergence of LLMs from the original questions.
RaR Effectively Improves LLM Responses
In this section, we provide a comprehensive assessment of the applicability and efficacy of RaR. The results are presented in four primary dimensions: One-step RaR is a simple and effective prompt to improve LLM performances; Two-step RaR effectively enhances the response accuracy of GPT-4 across diverse tasks; LLMs, while all benefit from Two-Step RaR, have different proficiency in rephrasing questions; a weak LLM can benefit more from a question rephrased by a strong LLM.
We first introduce the benchmark tasks we use to evaluate our method.
We evaluate the capabilities of LLMs across multiple benchmark tasks in different categories.
Knowledge Classification (Allen-Zhu and Li, 2023). Sampling a pool of individuals with Wikipedia pages, this task challenges the LLM to decide if a renowned person was born on an even day, month, or year.
Knowledge Comparison (Allen-Zhu and Li, 2023). Using the same pool of individuals, this task instructs the LLM to compare the ages of two people and decide who was born earlier.
As GPT-4 responds poorly to many of these questionsAs the data are not open-sourced, we let GPT-4 generate famous individuals with their birth dates and Chinese idioms in its knowledge., it raises a concern that whether GPT-4, despite its proficiency in retrieving knowledge, falls behind in reasoning with its own knowledge. Furthermore, we consider the following widely-used datasets for a comprehensive evaluation, which are also considered in Wei et al. (2022).
CSQA (Talmor et al., 2019). The CommonSense QA data encompasses a range of questions that evaluate the ability of commonsense understanding of the world and involves intricate semantics.
Date Understanding (bench authors, 2023). Sourced from Big-bench (bench authors, 2023), the Date Understanding task emphasizes commonsense reasoning and deducing a date from a provided context. The task is also considered in Wei et al. (2022). We consider a more difficult version where we do not provide the choices of potential answers and let the LLM answer directly.
Last Letter Concatenation (Fortes, 2023). The task centers on symbolic reasoning, and asks the LLM to concatenate the final letters of a given list of names. We consider concatenation for two names as well as a more difficult task of concatenation for four names.
Coin Fliphttps://huggingface.co/datasets/skrishna/coin_flip. Sourced from Hugging Face, the task asks the LLM if the coin still heads up, given its initial condition and subsequent actions of people who either flipped or did not flip the coin. We add an additional “Flip means reverse.” to the questions.
Sports (bench authors, 2023). Sourced from Big-bench (bench authors, 2023), the Sports Understanding task primarily asks if a sentence is plausible or implausible, where a prominent sports figure is depicted performing specific sports-related actions.
The details of all evaluated tasks are summarized in Table 1.
We use the entire dataset for Dates Understanding, and randomly draw subsets of size for the rest tasks. We use accuracy to evaluate the performance of the LLM. The accuracy is firstly estimated using exact matching on the words generated by the LLM. Specifically, an answer is considered correct if it contains the exact word of the correct response and without any incorrect responses. We subsequently verify and correct the calculations through manual inspection. For certain tasks, to constrain the response format (e.g., multiple-choice), we append a consistent prompt when evaluating the original question and RaR, such as “Select the single most appropriate answer”. Details of the prompts are presented in Table 8 in Appendix A.
2 Performance on GPT-4
We conduct experiments on the aforementioned benchmark datasets using GPT-4We note that all our experiments accessed GPT-4 during 10/01-10/30. We also include results on GPT-4-0613. (OpenAI, 2023), both with One-step RaR and Two-step RaR. As shown in Figure 5, both One-step RaR and Two-step RaR enjoy superior performance compared with using original questions. We will discuss the findings in the sequel.
We investigate the performance of RaR, which allows the LLM to both rephrase and respond to the question in a single query. Such an approach can be considered as a simple black-box strategy to improve the LLM’s performance on any question.
In Figure 5 and detailed in Table 7 in Appendix A, we compare the accuracy of GPT-4 with One-step RaR (i.e., rephrasing and answering a question in a single prompt) and Two-step RaR (answering a pre-rephrased question in separate queries). Notably, One-step RaR improves GPT-4’s accuracy, and outperforms Two-step RaR on out of tasks. Indeed, similar to human communication, rephrasing and elaborating a question and then answering is an effective approach. The key takeaway from this experiment in highlighted below.
2.2 Two-step RaR: Rephrased Questions Improve Response Quality
We evaluate the quality improvement of the question using Two-step RaR. In detail, for each query, GPT-4 autonomously generates a rephrased question using the prompt (2.2) without any external intervention. The rephrased question is then combined with the original question using (2.3) to prompt GPT-4. We present the accuracy of GPT-4 using Two-step RaR and compare it with GPT-4 using the original questions as shown in Figure 5. Across a diverse span of tasks that emphasize different aspects of LLM’s capabilities, Two-step RaR consistently yields distinguishable improvements for GPT-4. Notably, for tasks that GPT-4 originally finds highly challenging (e.g., last letter concatenation), the Two-step RaR exhibits remarkable improvement even to almost accuracy. The numerical details of the accuracy are also presented in Table 7 in Appendix A. We conclude this experiment by the following takeaway.
3 Performance across Various LLMs
We further examine the performance of RaR on various LLMs, including GPT-3.5 and Vicuna (Chiang et al., 2023)size=,color=orange!20!white,]Quanquan: references. In particular, we employ Two-step RaR to investigate (1) if all these LLMs can provide consistent response improvement by rephrasing the questions; and (2) if the GPT-4-rephrased questions can improve the performance of other LLMs.
We investigate the rephrasing abilities of different LLMs by employing Two-step RaR to examine the quality of the rephrased questions. We evaluate the performance of several different LLMs, including GPT-4-0613, GPT-3.5-turbo-0613, and Vicuna-13b-v1.5, using Two-step RaR. We present the experiment results in Figure 6. Due to Vicuna-13b-v1.5’s near-zero performance on Last Letter Concatenation (4), we exclude this task from the evaluation of Vicuna-13b-v1.5. Remarkably, all examined LLMs demonstrate enhanced performance with Two-step RaR, resulting in a notable increase in accuracy across the majority of the tasks. More advanced models, such as GPT-4, benefit from the most significant gains across all tasks, while models of lesser complexity, like Vicuna, achieve modest improvements using our approach. On certain tasks such as CSQA and Sports, GPT-3.5 and Vicuna even exhibit slightly diminished performance. In Table 2, we closely examine specific examples of self-rephrased questions by different models. Initial observations suggest that Vicuna-13b-v1.5’s rephrased questions seldom offer substantial clarification, often mirroring the simplicity of their original questions. In the last instance of Table 2, Vicuna-13b-v1.5 perturbs the question’s intent by changing “yesterday” to “today”. While both GPT-3.5 and GPT-4 can elucidate questions, GPT-3.5 occasionally introduces extra details or misinterpretations. As shown in the second example of Table 2, GPT-3.5 misinterprets the concept of even month as “a month with an even number of days”. Similarly, in the third example, GPT-3.5 introduces a wrong constraint of “recent” game. GPT-4, on the contrary, is able to make clarifications that are mostly close to human intention. We also observe that GPT-3.5 tends to introduce the following phrase to rephrased questions in Sports ( out of ) and Dates ( out of ): “Please rephrase and provide additional details if necessary to enhance your response accuracy.”, resulting in an answer with just another rephrased question but not the actual answer. Therefore, we remove all sentences containing “rephrase” for GPT-3.5 on these two datasets.
We wrap up this experiment with the following key insight.
3.2 Are the Rephrased Questions Transferable?
Here, we examine if the rephrased questions generated by Two-step RaR are transferable across different LLMs. In particular, we would like to know if the rephrased questions generated by GPT-4 can benefit Vicuna’s performance. We detail Vicuna-13b-v1.5’s performance on questions rephrased by GPT-4, as compared to its own rephrased questions in Table 3. Consistent with our expectation that GPT-4 can better align with human intention and clarify the question, we observe that its rephrased questions remarkably enhance Vicuna-13b-v1.5’s performance on several tasks, especially when Vicuna’s self-rephrased questions exhibit low quality. Indeed, the questions can be clarified further for Vicuna, but more exploration needs to be made on its capability of self-rephrased questions. We conclude this experiment by the following key message.
4 Multiple Rephrasings: Will the Questions Converge?
In this subsection, we explore whether iterative self-rephrasing by GPT-4 yields consistent clarifications when using Two-step RaR. Specifically, we utilize prompt (2.2) in Two-step RaR to enable GPT-4 to rephrase a question, then feed its output back into the same prompt (2.2) for a second and third round of rephrasing. In Table 4, we consider “Was Abraham Lincoln born on an even day?” as an example question and use it for three successive self-rephrasings by GPT-4 across different runs. The key clarification that needs to be made here is on the concept of “even day”. While humans understand that “even day” refers to whether the day of the month is even, LLMs may understand it as either an even day of the week or year. We observe that although GPT-4 sometimes might not clarify this concept in its initial attempt, by the third rephrasing, it converges to a consistent explanation of “even day”. Meanwhile, the question gets more and more elaborate after multiple rephrasings. This conveys the following key message.
Comparison with Chain-of-Thought
In this section, we compare RaR with CoT. We first present the mathematical formualtions of RaR and CoT and compare them with each other. Then we present experimental results to show that (1) RaR offers improvements in scenarios where zero-shot CoT is ineffective; and (2) RaR addresses and corrects the shortcomings inherent in few-shot CoT.
In this subsection, we discuss the formulations of CoT and RaR, respectively. We denote the LLM model by . In detail, LLMs take the sequence as prompt, and generate the sentence following the distribution . Recently, there has been significant research effort emphasizing the use of instructions to enhance the quality of the generated text. Mathematically, instead of directly using the prompt to generate the response following , one can use a different prompt augmented by instruction to generate a different response following . We hypothesize that a successful instruction allows us to extract a better answer from LLM. In this subsection, we employ the symbol to represent the target answer we aim to generate. The notation is used to denote an extended text that encompasses the desired answer as well as additional details, such as the underlying reasoning. Often, is produced by prompting with instructions like a Chain of Thought (CoT).
The core concept behind the CoT is to generate a text such that includes intermediate CoT steps and the final answer . In particular,
where are intermediate CoT steps that progressively lead to the final answer . In essence, CoT consists of the following primary phases.
CoT • Find an instruction , and generate prompt for zero-shot CoT or for few-shot CoT. • Generate following the “step by step” format of (4.1). contains the intermediate CoT steps , and the desired question . • Extract the desired response from the sequence . For zero-shot CoT, the instruction is composed of some task-independent tokens, such as “Let’s think step by step”. For few-shot CoT, the instruction/context consists of some task-dependent tokens, which include several examples like , where is the number of in-context examples, i.e., . We give examples of zero-shot CoT and few-shot CoT as follows.
Consider , for zero-shot CoT, we have as the effective prompt and generate , where
.
.
.
Consider , for two-shot CoT, we have as the effective prompt, where
.
.
.
.
And we obtain as
.
1.2 One-step RaR
The foundation of our (one-step) RaR method is different from CoT. We generate a rephrased question that retains the same semantic content as , and the associated answer . Specifically, we define as
where is the rephrased question that induces the answer . In particular, RaR consists of two primary phases:
RaR • Find an instruction , generate following the “Rephrase and Respond” format of (4.2). contains the rephrased question and the desired question . • Extract the desired answer from the sequence . We provide an example of (One-step) RaR as follows.
Consider the question , for RaR, we have
as the instruction and generate , where
.
Unlike CoT outlined in (4.1), which generates numerous intermediate steps , RaR in (4.2) aims to come up with an improved question efficiently. In this sense, our method RaR is more cost-effective in terms of token usage than CoT.
1.3 Two-step RaR
Instead of generating an extended text that encompasses the answer , the Two-step RaR approach operates in a sequential way. Specifically, it first employs the rephrasing LLM, denoted by to generate the rephrased question . Then, we input both original and rephrased questions to the responding LLM, denoted by to generate the answer.
Two-step RaR • (Rephrase Step) Find an instruction , generate rephrased question . • (Respond Step) Generate the desired response following . We also give an example for Two-step RaR below.
Considering the question , for Two-step RaR, we have
Then we feed together into the LLM and get .
In our experiments, we find that Two-step RaR can consistently achieve better performance. The rephrased question can also be used by another LLM, which makes our RaR method more flexible.
1.4 Combining RaR and CoT
In addition, our method is complementary to CoT and can be naturally combined with CoT. For zero-shot CoT, we can simply concatenate the two instructions to obtain , such as “Given the above question, rephrase and expand it to help you do better answering. Lastly, let’s think step by step to answer”. For few-shot CoT, where the instruction/context is , we can use Two-step RaR to improve its few-shot examples by the following procedure.
RaR + CoT • Use RaR instruction , generate in the Rephrase step. • Apply to get . • Extract the response from the sequence by eliminating intermediate steps. Remark 4.1. Compared with Two-step RaR, which uses both the original question and the rephrased question to prompt the LLM to generate the response, RaR+CoT only uses the rephrased few-shot examples instead of combining with the original few-shot examples . This will save the token usage and prevent an increase in the number of in-context examples while maintaining a similar performance.
The following example showcases how to combine RaR with CoT.
Consider two-shot CoT instruction , where
Applying as in Figure 9, we get
Our experiment demonstrates that the integration of RaR with few-shot CoT significantly improves the performance of CoT. A comprehensive discussion of the results can be found in Section 4.3. Lastly, we present illustrations of our mathematical formulations for CoT and RaR in Figure 7.
2 Empirical Comparison with Zero-Shot CoT
It is widely known that zero-shot CoT, by appending the instruction “Let’s think step by step.” to queries, can effectively improve the performance of LLMs on reasoning tasks. However, we highlight some examples where zero-shot CoT fails to deliver improvements, sometimes even leading to diminished performance. In contrast, RaR consistently demonstrates effectiveness. We also emphasize the importance of question quality with an example, demonstrating that it should be prioritized before enhancing the model’s reasoning capabilities. Lastly, we note that our method is complementary to zero-shot CoT and can be combined together by simply adding “let’s think step by step” to (2.3) or (2.1).
We examine the Chinese Idiom task as introduced in Allen-Zhu and Li (2023), specifically the most difficult task of inferring the first letter. This task involves taking widely recognized four-character Chinese idioms and masking one character at each respective position. The task is to let LLM correctly infer the masked character. It has been discovered that GPT models suffer from inferring the masked character, particularly when it is located in the first position. Furthermore, we also use the StereoSet task (Nadeem et al., 2021), which assesses the stereotypical biases present in LLMs with respect to gender, race, profession, and religion. From the inter-sentence data, we sample examples, each comprising a context sentence and three choices: one stereotypical, one anti-stereotypical, and one unrelated. We adopt the prompt format used by Shaikh et al. (2022).
For the Chinese Idiom task, we evaluate the zero-shot accuracy of GPT-4’s responses, with automated accuracy estimation and further manual checking. For StereoSet, as suggested by Nadeem et al. (2021), two crucial evaluation metrics should be considered: the Language Modeling Score, which assesses whether the LLM selects related options over unrelated ones, and the Stereotype Score, which quantifies the percentage of data that a model favors stereotypical choices over anti-stereotypical ones. As identified by the authors, an ideal model would display no bias toward either stereotypical or anti-stereotypical associations, yielding an optimal score of for the Stereotype Score. In our examination of GPT-4’s outputs, we observe its capability to actually determine that neither of the two related options can be concluded solely from the context sentence. Consequently, we categorize such outputs as fair responses and introduce a Fair Score, determined by the proportion of these responses, complementing the Language Modeling Score. We provide an example of such a response below.
As illustrated in Table 5, even though RaR enhances LLM’s performance, accurately inferring the first character of the Chinese Idiom task remains a challenge. One might then ask: does zero-shot CoT provide consistent improvement to LLM in such tasks as it does on other reasoning tasks? Our discovery is, in fact, zero-shot CoT may result in worse performances () for such hard tasks, as the LLM tends to hallucinate during the intermediate steps—a phenomenon similar to hallucination snowballing (Zhang et al., 2023a). Furthermore, as Shaikh et al. (2022) discovered on other language models, zero-shot CoT may result in undesired reasoning towards bias and toxicity. Also in Table 5, we demonstrate the performance of GPT-4 on StereoSet. We can observe that, while zero-shot CoT fails to improve the Language Modeling Score, rephrased questions improve it significantly to . This implies that, with RaR, the LLM rarely opts for unrelated choices. Moreover, while zero-shot CoT improves the percentage of fair responses (choosing neither of them), RaR achieves the best performance.
With the following example, we emphasize that the attention to question quality is more important before considering to improve model reasoning. We examine the original Coin Flip questions. Specifically, an example question is
A coin is heads up. aluino flips the coin. arthor flips the coin. Is the coin still heads up?
This question, as originally crafted by Humans, appears clear to human interpreters that “flipping” the coin here means reversing the coin. However, LLMs like GPT-4 might perceive the flipping as a random toss. As shown in Figure 8, such misconception persists even when prompting the LLM to think step by step, which therefore results in an incorrect answer. Once we add a clarifying sentence stating that flip means reverse, GPT-4 can finally start answering the question as we desire. Since the clarification is also created by humans, LLM still exhibit an unsatisfactory performance of . With self-rephrased questions, the accuracy can finally be improved to . In light of such instances, we advocate for careful examination of human-crafted questions when evaluating LLMs, ensuring the removal of ambiguities for a fair assessment.
3 Empirical Improvement on Few-Shot CoT
Few-shot CoT (Wei et al., 2022) has been the most effective CoT technique. It employs a small set of human-crafted QA examples to facilitate LLMs in addressing similar questions with a congruent structure. LLMs, particularly advanced models such as GPT-4, are adept at extrapolating from the provided examples to improve their performance on new questions. Providing question-answer pairs effectively communicates the human-desired logical structure to the LLM. Instead of aligning the question to what the LLM best receives, few-shot CoT guides the LLM to reason using the supplied human logic. Nonetheless, a concern emerges: How do LLMs respond when the human-crafted examples are flawed or contain errors? As corroborated by a recent parallel study (Pawelczyk et al., 2023), we similarly observe that LLMs can be adversely influenced by bad few-shot examples.
We revisit the Last Letter Concatenation task and refer to the few-shot examples provided in Wei et al. (2022). As shown in Figure 9, the examples follow a specific logic: (1) obtain the last letter of the first word; (2) obtain the last letter of the second word; (3) concatenate these letters; resulting in (4) the answer. Such few-shot examples have been demonstrated to most effectively enhance the performance of a language model, achieving an accuracy of when concatenating the last letters of two words. Conversely, we explore an example that employs the following logic: (1) obtain the first letter of the first word; (2) obtain the first letter of the second word; (3) concatenate these letters; providing (4) the answer for last letter concatenation. Our aim is to investigate how this alternative few-shot prompt, despite bearing a logic similar to the original prompt and the correct answer, influences the performance of the GPT-4.
As illustrated in Figure 9, GPT-4 tends to stick to the logic of our modified prompt, resulting in an incorrect answer. It accurately concatenates all first letters, but concludes with a seemingly arbitrary final answer. In Table 6, we demonstrate the results of one-shot and four-shot CoT using such examples. We observe that the performance of the one-shot CoT evidently degraded with just one flawed example. As the number of these flawed examples increases, the performance of GPT-4 in a 4-shot setting for last letter concatenation of four words drops to only . This observation reveals a potential pitfall in employing few-shot CoT: given that these examples are user-crafted, their quality becomes vital. Meanwhile, we discovered that RaR enables GPT-4 to correct any pitfalls in the logic of the given examples.
Related Work
Since the advent of recent LLMs (OpenAI, 2023; Touvron et al., 2023; Chiang et al., 2023), a growing body of research has focused on prompt engineering for LLMs (Brown et al., 2020; Schick and Schütze, 2021; Zhou et al., 2022b; White et al., 2023; Wang et al., 2023). Manual guidelines have emerged to guide users in designing and revising their prompts (Reynolds and McDonell, 2021; Saravia, 2022). Notably, studies have shown that a well-crafted system message, such as ”You are a helpful assistant who always provides explanations,” preceding the main query can encourage an LLM to respond with greater subject expertise (Mukherjee et al., 2023; Ateia and Kruschwitz, 2023). OpenAI (2022) has also offered general recommendations for crafting queries, emphasizing specificity, detail, and precision. However, individuals often find it challenging to refine their own questions for clarity or to include necessary details for LLMs, as the questions are clear enough for humans themselves.
Subsequent research (Zhou et al., 2022b; Sorensen et al., 2022; Pryzant et al., 2023) has concentrated on the autonomous refinement of prompts. These methods often employ multiple LLMs to generate candidate prompts, evaluate and score these prompts, and iteratively refine them until a satisfactory prompt is produced. The evaluation of a prompt typically relies on either the accuracy of an LLM’s response (supervised, Zhou et al. (2022b); Pryzant et al. (2023)) or the mutual information of the question (unsupervised, Sorensen et al. (2022)). Given the nature of iterative computation and the necessity for qualitative evaluation, such methods are employed for refining single prompt templates; applying them universally to all questions would be expensive. Consequently, these techniques are less frequently adopted in daily user cases.
The method most frequently used by users and closely aligned with our approach is the Chain-of-Thought (CoT) prompting, which can be either zero-shot (Kojima et al., 2022) or few-shot (Wei et al., 2022). Given that these techniques do not require evaluation and iterative selection, they have gained widespread popularity and inspired a series of subsequent studies (Wang et al., 2022; Zhou et al., 2022a; Press et al., 2022; Yao et al., 2023; Zhang et al., 2023b; Shao et al., 2023). However, CoT methods are not without their limitations, as observed in our study. Recent investigations have also highlighted challenges with the reliability of both zero-shot CoT (Turpin et al., 2023) and few-shot CoT (Pawelczyk et al., 2023). Most recently, Zhou et al. (2023) propose Foresee and Reflect similarly as a zero-shot prompting method that targets the proposed task Thinking for Doing (T4D). Lastly, it is worth noting that our method is complementary to all the prompting techniques mentioned above and can be combined with them.
2 Self-correction Methods for LLMs
Another line of work aims at enhancing LLM performance (Madaan et al., 2023; Welleck et al., 2022; Kim et al., 2023; Pan et al., 2023; Shinn et al., 2023) by leveraging the LLM to refine its own responses, a concept known as post-hoc prompting. This encompasses terms such as ”self-correction,” ”self-refine,” and ”self-critique,” where LLMs revise their own responses drawing upon various feedback sources or critic models. As classified by Pan et al. (2023), automated critic models generally employ the LLM’s self-feedback (Madaan et al., 2023; Shinn et al., 2023; Yan et al., 2023), other trained LLMs (Yang et al., 2022; Lightman et al., 2023), or external references (Jung et al., 2022; Gao et al., 2023; Yu et al., 2023; Welleck et al., 2022). Yet, recent studies (Huang et al., 2023; Stechly et al., 2023) examine the self-correction capacities of LLMs and find potential limitations, suggesting that LLMs may not be able to self-correct their reasoning processes. Their findings reveal that self-correction is no better than self-consistency (Wang et al., 2022). Contrary to allowing the LLM to self-refine its responses, our methodology let the LLM instead rephrase questions originally crafted by humans.
Conclusion
In this paper, we have investigated the existing misunderstandings that occur between humans and LLMs and demonstrated that questions that appear clear to humans may still be misinterpreted by LLMs. Building on this insight, we introduced Rephrase and Respond (RaR), a novel approach that prompts an LLM to first rephrase and clarify the question before answering it. We also presented Two-step RaR, a variation of RaR that employs a rephrasing LLM to refine questions for subsequent use by any responding LLM. Our empirical evaluations, conducted across a range of benchmark datasets, confirm the effectiveness of our proposed methods. Further analysis reveals that while all models gain enhanced performance through question rephrasing, the more sophisticated models exhibit more substantial improvements. Crucially, we have found that the enhancement in question quality achieved through rephrasing is transferable across models. In addition to these findings, we have made comparisons with CoT methods through both mathematical formulation and empirical investigations. We also demonstrated that RaR is complementary to CoT, and can be leveraged to achieve additional performance gains.
Appendix A Experiment Details
Our experiments are done using the publicly available GPT-4 API, as well as the historical version of GPT-4-0613 and GPT-3.5-turbo-0613. We are also considering an open-source LLM model, Vicuna-13B-v1.5. Experiment results for Figure 5 are detailed in Table 7. Moreover, the prompts we used for formatting the answers are detailed in Table 8. The few-shot examples used in Section 4.3 are demonstrated in Tables 9 and 10. In the next appendix section, we provide comprehensive examples of the inputs and outputs of each task for the different methods.
Appendix B Input/Output Examples
In this section, we provide specific input and output examples of GPT-4 on each of task we considered, using either original questions or RaR.