REFINER: Reasoning Feedback on Intermediate Representations

Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, Boi Faltings

Introduction

Large language models (LLMs) have made significant strides in natural language processing (NLP) tasks (Brown et al., 2020). Recent work has shown that explicitly generating intermediate steps during reasoning tasks significantly improves a model’s performance and interpretability (Shwartz et al., 2020; Paul and Frank, 2021; Marasovic et al., 2022; Lampinen et al., 2022; Wei et al., 2022). Producing such intermediate representations provides insight into the model’s predictions and allows humans to inspect the model’s reasoning process. However, these intermediate representationsIn a reasoning task, the intermediate representations can be viewed as inference rules, explanations or reasoning steps. * Work done at EPFL can be unreliable (Ye and Durrett, 2022) and result in poor performance on downstream reasoning tasks. Most importantly, it is unclear how to meaningfully refine the intermediate representations to further improve the final performance.

The standard practice for correcting reasoning errors is to annotate new data and either retrain or finetune the model (Feng et al., 2021; Hedderich et al., 2021). However, fixing such errors by finetuning with more data is not only data- and resource-intensive but can also be insufficient to generalize well in complex reasoning tasks (Ward et al., 2022). Other works have explored improving models using feedback by providing a scalar reward (Ziegler et al., 2019; Martin et al., 2022) or directly revealing the correct missing answer (Mehta and Goldwasser, 2019; Elgohary et al., 2021; Tandon et al., 2022). However, in natural language reasoning tasks, defining a reward that captures different fine-grained reasoning error types (e.g., semantic consistency, logical, etc.) remains an open challenge (Golovneva et al., 2023). Additionally, such a reward provides a relatively sparse training signal.

In this work, we instead provide fine-grained and structured feedback on reasoning errors. We present REFINER, a novel interaction-based framework that allows a generator LM to iteratively use fine-grained feedback and refine its reasoning. The interaction happens between two models: a generator, which learns to solve the task by first generating the intermediate reasoning steps, and a critic, which provides structured feedback to the generator about errors in the intermediate steps.

To provide fine-grained feedback about reasoning errors, we develop a scheme to independently train the critic model on automatically constructed feedback data. More specifically, we create pairs of incorrect intermediate representations and structuredNote that we transform the structured feedback into semi-structured textual feedback using templates. feedback on their fine-grained reasoning errors. Then, we use this data to train the critic to provide fine-grained feedback on erroneous intermediate reasoning steps. Finally, the critic interacts with the generator LM, offering feedback both during the training of the generator and during inference.

Figure 1 illustrates an example of our REFINER framework where, given a math word problem, the generator generates an equation as an intermediate representation. The critic identifies the errors in the equation and provides semi-structured textual feedback (e.g., "the operator in #0\#0 is incorrect") to the generator. By interacting with the critic, REFINER enables the generator to reason over the semi-structured feedback and refine its generation.

Contributions. (i) We propose REFINER, a framework that refines LMs reasoning capabilities through feedback. Our work investigates how interacting with fine-grained reasoning feedback on intermediate reasoning steps impacts the performance of LMs on reasoning tasks. We evaluate REFINER on three natural language reasoning tasks: math word problems, synthetic natural language reasoning, and moral action generation. REFINER demonstrates significant performance gains across different LM architectures with different scales. Across different reasoning tasks, REFINER outperforms comparably-sized strong fine-tuned LM baselines (by +13.1, +3.2, +15 pts., respectively). (ii) We empirically demonstrate that for math word problems and synthetic natural language reasoning, our trained critic models alone are beneficial for improving intermediate representations as they help GPT-3.53.5 significantly increase its performance in a few-shot setting (by +3.5, +6.8 pts., respectively). We also demonstrate that providing structured feedback on fine-grained errors can benefit more than scalar value feedback for moral action generation and math word problem tasks. Our critic model acts as a ‘reasoning refinement tool’ for LLMs. (iii) We show that REFINER can substantially outperform other refinement methods that use feedback from large LMs, such as self-refine. (iv) Our analyses illustrate that (a) improving the intermediate representation generation improves the performance on the reasoning tasks, and (b) training a generator with an imperfect (noisy) critic is still beneficial. Our code is made publicly available https://github.com/debjitpaul/refiner.

Related Work

Intermediate Representations. While state-of-the-art LMs achieve incredible performances in a wide range of tasks, they have difficulty with many reasoning tasks Wang et al. (2022), especially ones with multiple constraints or sub-problems or requiring specialized knowledge Austin et al. (2021) – such as mathematical problem solving Ling et al. (2017); Andor et al. (2019); Ran et al. (2019); Geva et al. (2020); Piękos et al. (2021); Cobbe et al. (2021a); Kim et al. (2022).

For these tasks, both intermediate representations and rationales have been shown to be beneficial in learning mathematical skills Piękos et al. (2021), intermediate program execution computations Nye et al. (2021), or general reasoning outputs Wei et al. (2022); Golovneva et al. (2022).

Our work builds upon the observation that generating intermediate steps are valuable but distinguishes itself in several key aspects. Firstly, instead of prompting a large model, we finetune smaller models to learn to generate intermediate steps. Secondly, our framework can accommodate tasks that do not necessarily have unique closed-form correct answer, such as the Moral Norm task (see §3). Finally, our framework is trained with a critic providing feedback, improving the model’s reasoning process and teaching it how to leverage feedback.

Natural Language Feedback. Recent work has explored giving models richer and more complex feedback through the use of natural language Ziegler et al. (2019); Nguyen et al. (2021); Scheurer et al. (2022), used for aligning LLMs’ output with users’ preferences Christiano et al. (2017); Ziegler et al. (2019); Saunders et al. (2022); Scheurer et al. (2022); Bai et al. (2022), or to directly improve the model’s performance in its current task Weston (2016); Rupprecht et al. (2018); Elgohary et al. (2020); Austin et al. (2021); Madaan et al. (2023). This training depends on human-created feedback, generated in large quantities Bai et al. (2022), which takes up considerable resources. Though an external feedback provider can guide models to correct answers and reasoning Austin et al. (2021), demonstrably better than they can themselves Saunders et al. (2022), feedback has rarely been used in this way – and automated critics for reasoning tasks have proved to be difficult Scheurer et al. (2022); Wang et al. (2022); Huang et al. (2022).

Recently, Welleck et al. (2022) introduced a secondary model, the corrector, which improves the initial proposition of a generation model, by learning the kind of mistakes made by the generator and how to fix them. In this work, we also use a secondary model, a critic, but apply it quite differently as we integrate it into an interaction loop with the generator model during training. We further differ from previous works as we provide feedback at the intermediate reasoning steps of the model and not at the final output. The feedback is thus closer to the source of mistakes and guides the model’s reasoning toward the correct answer. Additionally, intermediate steps are often structured, allowing the critic to provide precise feedback.

REFINER

Problem Formulation. In this paper, we view natural language reasoning (NLR) as an autoregressive generation task where, given input context xx, a model needs to generate yy, such that yy satisfies the constraints of the task. Usually, to generate correct or plausible yy, the model needs to make the correct inference zz as intermediate steps. We use “inference steps/representations” and “hypothesis” interchangeably. We decompose NLR tasks as follows: p(y∣x)=p(y∣x,z)p(z∣x)p(y|x)=p(y|x,z)p(z|x). In practice, one can compute each conditional using an LM that includes its conditioning variables as a part of its input.

Before continuing with the model description, we describe three NLR tasks where we conduct our study and their respective intermediate representation zz. We deliberately chose these three tasks since they broadly cover two types of reasoning: (i) logical reasoning and (ii) normative reasoning. They are exemplified in Appx Fig. 6 and detailed below.

Math word problem (MWP), where given a word problem xx consisting of a context and question, the goal is to map xx to a valid mathematical expression zz (the intermediate representation) and then to a solution yy. This task requires the model to perform deduction using mathematical reasoning. Synthetic natural language reasoning (sNLR), where given a reasoning scenario xx consisting of 55 synthetic rules and a fact, the model needs to deduce a conclusion yy. This task requires the model to perform deductive reasoning and generate intermediate steps zz and the conclusion yy using closed-world rules and facts. Moral norm and action generation for moral stories (MS), where given a context xx consisting of a situation, an intention, and an immoral action, the model needs to generate the moral norm z{z} and the moral action y{y}. Moral actions are encouraged by the moral norm. This task requires the model to perform abductive reasoning to generate moral norms and deductive reasoning for moral action.

We propose to solve these tasks by forcing the model to generate intermediate hypotheses (zz) and improving them via structured feedback. We introduce an interactive framework, REFINER, made of two separate models: (a) a critic model (§3.1) trained to provide structured feedback on intermediate reasoning steps and (b) a generator model trained to solve the reasoning task by first generating intermediate reasoning steps (§3.2). The core idea of REFINER is to exploit the interaction between the generator model and the critic model, where the generator’s intermediate reasoning steps are improved via structured feedback from the critic.

REFINER presents several important properties. First, the generator is trained to incorporate and leverage feedback, which helps it converge towards better reasoning during training and makes it capable of integrating feedback at test time, whether from a trained critic or a human (see §5). Second, the trained critic can be useful on its own; we demonstrate that a generalist LLM like GPT-3.53.5 can significantly benefit from interacting with our trained critic on the reasoning tasks we consider (see §5). Finally, having two separate models allows us to easily measure the benefits of feedback during training and/or during inference (see §6).

The role of the critic is to provide feedback on the intermediate hypotheses produced by the generator model. One way to evaluate the quality of the hypothesis and produce feedback on the hypothesis zz, would be to compare it against a gold hypothesis z∗z^{*}. Previous works employed automatic metrics like BLEU, ROUGE, etc., as value functions (Wu et al., 2018; Ramamurthy et al., 2022). However, these scalar value functions are not suitable for natural language reasoning tasks because (i) it is unclear how to define a scalar value function that can encapsulate fine-grained reasoning errors (Golovneva et al., 2023) and (ii) during inference, these functions require access to the gold hypothesis (which is unavailable in practice). Therefore, we train a critic model and endow it with the ability to evaluate the hypothesis in a fine-grained manner and provide structured feedback.

Feedback Data Generation. To train the critic, we have to create example pairs of implausible hypotheses and their corresponding feedback with fine-grained reasoning errors. Inspired by Golovneva et al. (2023) and Talmor et al. (2020), we first define fine-grained reasoning error types for each reasoning task (see Table 1). For MWP, an equation can be incorrect due to: (i) the operands or operators in the equations being incorrect and/or (ii) one or more operators missing. For sNLR, an inference rule can be incorrect because it is (i) logically invalid and/or (ii) missing reasoning rules (failing to connect the correct facts with correct rules or missing implicit knowledge). For MS, a moral norm can be incorrect due to (i) contradiction and/or (ii) semantic misalignment.

Based on these error types, we propose two strategies to create the feedback data: (i) Rule-based perturbation strategy: we perturb the plausible hypotheses (zz) in the training data and collect a pool of data DD (xx: input, zz: plausible hypothesis, z′z^{\prime}: implausible hypothesis). We perturb by omitting, replacing or adding some tokens or some rules from the plausible hypothesis to create an implausible hypothesis automatically (details in Appendix F.1). (ii) Synthetic Generation strategy: we prompted OpenAI’s GPT-3.53.5 to generate implausible hypotheses based on the error types automatically. We used a few-shot setting where we varied the instruction, the number of demonstrations, and the formatting of the demonstrations (details in Appendix F.2).

Since our perturbations and automatic implausible hypotheses are based on logic and reasoning errors, we create structured feedback ff for every example (x,z,z′x,z,z^{\prime}) by stating the error type that occurs in z′z^{\prime} but not in zz (see Table 1). The basic structure of feedback ff for these tasks is ⟨\langleerror type, position (optional), hint (optional)⟩\rangle, where position denotes the error position in the implausible hypothesis (see Table 1). Despite the simplicity of the strategy we used for our tasks, this approach is easily generalisable to other reasoning tasks.

We also replace the correct judgment with random judgments to scale the number of implausible hypotheses per example. Finally, as feedback ff, we provide <<error type, hint>>. For non-monotonic reasoning tasks like norm and action generation, the critic should be able to provide hints that align the generator model’s objective to the reasoning task. Hence, as a hint, we provide verb phrases from the norms. Since the critic provides textual feedback to the generator, we convert the structured feedback into natural language feedback Further details about feedback are provided in Appx.F.. Formally, we create a data pool D={x,z,z′,f}D=\{x,z,z^{\prime},f\} to train a critic model.

Training the critic model. We train a supervised critic model (πβ\pi_{\beta}) with the context (xx) and (plausible or implausible) hypothesis (zz or z′z^{\prime}) as input and the textual feedback as output. We update the critic with the cross-entropy loss: L(β)=−log⁡pβ(f(u)∣x,u)L(\beta)=-\log p_{\beta}(f(u)|x,u) where u∈z,z′u\in z,z^{\prime}. The trained critic is only used during inference. The oracle critic is used while training the generator.

2 GENERATOR Model

This section presents a generator model that iteratively learns to interact with the critic model.

Warm-up. Given a context xx the generator model (πθ\pi_{\theta}) is trained to generate plausible hypotheses. The warm-up phase is critical to ensure that, when the critic comes in the loop, the generator does not produce random answers likely to be bad, given the size of the output space. As such, we use a small supervised dataset (10% training data) to fine-tune the model on the NLR task of interest. After the warm-up phase, we use the additional feedback ff from the critic model and learn πθ(z∣x,z′,f)\pi_{\theta}(z|x,z^{\prime},f).

Exploration. At each iteration (tt), the generator model generates multiple hypotheses (zkz^{k}) using nucleus sampling. The critic model randomly selects one hypothesis and provides feedback on that hypothesis. The exploration step aims at increasing the output variance such that the generator receives a wide range of feedback during training.

Learning. We update the generator model using the following cross-entropy loss: L(θ)=−∑t=1Tlog⁡pθ(zt∣x,zt′,ft(z′))L(\theta)=-\sum_{t=1}^{T}\log p_{\theta}(z_{t}|x,z_{t}^{\prime},f_{t}(z^{\prime})) where TT = total number of iterations. Since the feedback contains the error types and hints, which are (latent) fine-grained and logical, it should allow the model to learn and update its generation by addressing the reasoning errors mentioned in the feedback.

Inference. We use the trained critic along with the trained generator to generate a trajectory z0,z1,...,zTz_{0},z_{1},...,z_{T} and stop when either f(zt)f(z_{t}) is generated by the generator or “No hint” is generated by the critic. We also experimented with chain of thought prompting, where the generator generates a trajectory z0y0,z1y1,...,zTyTz_{0}y_{0},z_{1}y_{1},...,z_{T}y_{T} and stops when the critic generates “No hint”.

Experimental Setup

Datasets. We evaluate REFINER on three diverse tasks (examples in Fig. 6). We briefly describe the datasets used for each task below.Math Word Problem (MWP): We train our models on MAWPs (Koncel-Kedziorski et al., 2016) dataset and evaluated our models on a challenging dataset SVAMP (Patel et al., 2021). We evaluate our model on both the equation generation (zz) and answer prediction (yy) tasks. Similar to Ling et al. (2017); Amini et al. (2019) for equation generation, we replace the numeric values with variable names, for example, number0, number1, etc. Further, we also evaluated on GSM8K (Cobbe et al., 2021b) dataset which consists of 8.5K high-quality linguistically diverse grade school math word problems. For Synthetic Natural Language Reasoning (sNLR), we use the dataset from Liang et al. (2022) with the difficulty level as hard. We evaluate our model on both inference rule generation (zz) and consequent generation (yy). For Moral Story (MS), we use a dataset from (Emelin et al., 2021), where we evaluate our model on moral norm z{z} and the moral action y{y} generation.

Training Details. For each task, we train a UnifiedQa-T5-base model (UQA-base) (Khashabi et al., 2020) as a critic (§3.1). For exploration (§3.2), we use nucleus sampling with p=0.5p=0.5. We select the hyper-parameters by the validation loss: for both the generator and critic model, we use the Adam optimizer with a learning rate of 1e−41e^{-4}. Each model is trained for 2020 epochs with early stopping based on validation loss. We trained all models on one A100 GPU. We run our models with 33 random seeds and report the average results. For the human study, we selected outputs from the best models (baselines and our model) according to automatic metrics. We train models with T=3T=3 iterations.

At inference time, we use greedy decoding for the generator and critic model with T=1T=1 for the automatic critic and T=3T=3 for the oracle critic. On the MWP and sNLR tasks, we use the exact match (EM) metric for intermediate steps (equation generation and inference rules) and accuracy (Acc) for the final answers. For MS, we conduct a manual evaluation study to assess the relevance of norms and moral actionsSince the automatic scores such as BLUE, ROUGE, etc. only account for word level similarity between gold norms or actions and generate norms or actions.. Further evaluation details are provided in Appendix G. To train the critic model, we used the feedback data generated using the rule-based perturbation strategy (see §3.1).

Baselines. We compare our method with three different LMs as generator models: UQA-base, UQA-large (supervised setting), GPT-3.53.5-text-DaVinci-003 and ChatGPT (few-shot setting). We also compare REFINER to Proximal Policy Optimization (PPO) RL-based method (Schulman et al., 2017). We use the implementation of PPO from (Ramamurthy et al., 2022). For GPT-3.53.5, we provide 22 for demonstrations per class. We also experimented with chain of thought (COT) prompting (Wei et al., 2022) where the model is prompted first to generate the intermediate steps (zz) and then the final answer (yy). Note that the sNLR task is a synthetic task where the model needs to perform either one-hop or two-hop reasoning. Clark et al. (2021) showed that fine-tuning large language models (354354M parameter size) could achieve (99% accuracy) high performance. Hence, we only compare our REFINER model with the UQA-base model (220220M) (see Table 3). Since human annotation is expensive, we focus on comparing against the most meaningful baseline: UQA-large for MS task (see Table 4). It is important to highlight that our proposed framework is general, and one can use any other LMs as generator or critic.

Results

We evaluate our model on two aspects (i) performance on intermediate steps and (ii) performance on the final answer prediction. Tables 2, 3, and 4 show the performance comparisons.

Performance on Intermediate Steps. Table 2 reports the performance of the MWP task. We explored two different scenarios: (i) where the model only generates the equations (zz) with variable names replacing the numeric values, and (ii) where the model generates both the equations and the final answers together. We observe for both scenarios that REFINER significantly outperforms baseline models with comparable sizes. Notably, UQA-base benefits most (+13.1+13.1 EM) when adding a critic in the loop. We observe that GPT-3.53.5 significantly benefits from the REFINER trained critic. Since LLMs like GPT-3.53.5 (175175B parameters) are expensive to finetune, the improvement in equation generation of +3.2+3.2 EM without any modification is important. Interestingly, we observe that GPT-3.53.5 + COT manages to have significantly higher accuracy in answer yy than in equation zz (see Table 2). This result is similar to the observation made by Ye and Durrett (2022) and suggests that the intermediate equations can be unreliable. Finally, REFINER could even outperform PPO, which uses BLEU-score as a reward function. This suggests that semi-structured fine-grained textual feedback is more beneficial than value-based (where values are from automatic metrics) reward feedback. Note that this result may vary when these models are optimized directly with complex human values, as shown in Stiennon et al. (2020). Qualitatively, REFINER can correct incorrect equations through structured feedback, fixing the operators within a multistep solution (see Fig. 7).

For sNLR, similar to Liang et al. (2022), we observe that GPT-3.5 performs poorly (see Table 3). REFINER improves +2.9+2.9, and +6.8+6.8 EM scores over UQA-base, and GPT-3.53.5, respectively. Contrary to the MWP, the final answer yy is not a symbolic execution away from the intermediate step zz, but we still observe that REFINER focuses on improving the intermediate step zz, resulting in significant improvements in the answer yy prediction. Again, we observe that REFINER with a UQA-base can outperform few-shot prompted GPT-3.53.5. Thus, our critic can identify the fine-grained reasoning errors and help improve the performance on inference rules generation.

For MS, we assess the generation quality with three human judges who indicate whether the generated norms and moral actions are relevant to the given moral story. Table 4 summarises human evaluation results on 100100 moral story examples randomly sampled from the MS test dataset. More specifically, we report evaluation breakdown for both norm and moral action by the number of instances that are either Irrelevant, Unsure or Relevant along with Krippendorf’s α\alpha Krippendorff (2018) agreement scores. The results show an improvement of 2020 points, increasing the relevance over a strong UQA-large baseline. Hence, this suggests that a specialized critic model with 33 times fewer parameters than the generator can improve the performance on generating reasoning steps.

Performance on Final Answer Prediction. We observe that REFINER outperforms the strong LM baselines by +3.5,+3.2,+15+3.5,+3.2,+15 points for MWP, sNLR, and MS, respectively. These results support our hypothesis that generating better intermediate steps can result in better answer prediction. Notably, on the sNLR task, for GPT-3.53.5, we observe that by adding a critic, there is an improvement of +6.86.8 in inference step generation; however, only +1.5+1.5 in the consequent prediction. This result indicates that LLMs may either not use these intermediate steps to perform the deduction or fail to perform deduction.

Comparing REFINER with other refinement methods. In Table 5, we compare REFINER with two other recent refinement methods: Self-refine (Madaan et al., 2023) and Self-reflection (Shinn et al., 2023) method on the SVAMP and GSM8K datasets. Both these baseline methods use LLMs to generate automatic feedback. Similar to Madaan et al. (2023), we observe that self-refine has minor improvement for MWP tasks. On the contrary, we find that REFINER significantly improves the performance of GPT-3.5 and ChatGPT by +3.33.3 and +2.22.2 on SVAMP and GSM8K datasets, respectively. This highlights the benefit of training a specialised critic that is grounded to the task. It can make LLMs more accurate than feedback from a general-purpose model (GPT-3.53.5 or ChatGPT). In Appendix §6, we have provided more details about the quality of feedback generated using our trained critic and GPT-3.53.5 (see Table 8). Further, we assess the performance of REFINER in improving the CoT generated by two recent methods: Self-Consistency (Wang et al., 2023) and ReACT method (Yao et al., 2023). We observe that REFINER can improve self-consistency and ReACT by +2.022.02 and +2.92.9. This demonstrates that a trained critic can be used as a tool and can bring performance gains to different methods out-of-the-box (more details in Appendix §A.2).

Ablation. To obtain better insight into the contributions of the individual components of our models, we perform an ablation study (Table 6). We observe that there is a considerable drop in performance from 47.247.2 to 39.839.8 when we do not use the critic model during inference. Hence, this result indicates that our generator model can leverage the feedback from the critic at inference time. Further, we find that the exploration step improves the performance +3.3+3.3 over the baseline model. This result supports our hypothesis that the exploration step increases the output variance and gives the generator model the opportunity to learn over a wide range of feedback. We compared the performance with the critic model trained on two different training data (see §3.1). We find that the critic trained on small automatically generated data using GPT-3.5 works better than without the critic in the loop. This result motivates researchers to use this method to generate negative samples to train their critic or preference learning model. Finally, we also observe that if the critic was perfect (Oracle), then REFINER can significantly improve the performance by fixing the mistakes generated by the generator model. This result indicates that REFINER can be seen as a framework that allows AI-AI and human-AI interaction.

Analysis

Error Analysis. In order to get more insight into the performance of our method, we conduct a fine-grained error analysis on the MWP and MS datasets (Fig. 3). We note that the most frequent errors are Incorrect Numbers for MWP and Semantic Misalignment for MS. An intuitive reason can be that for the MWP task, the models are sensitive to the numbers order as argued in (Patel et al., 2021). For MS, generating norms grounded in the context is challenging. Our analyses show a clear trend that REFINER is able to considerably reduce the errors for both datasets. This indicates that our trained critic model could identify fine-grained reasoning errors during inference.

Noise Sensitivity. To further understand the behaviour of the REFINER framework, we run variations with noisy critics for the MWP task. We replace the oracle critic used during training with a noisy critic in (Fig. 4 (a)) to inspect how training with an imperfect critic impacts the generator. We also use a noisy critic at inference while keep the oracle critic during training (in Fig. 4 (b)). The noisy critics are generated by random perturbations of the oracle critic; for a noise-level ϵ\epsilon, the oracle feedback is replaced by random feedback with probability ϵ\epsilon.

Fig. 4 (a) shows that when training with a very noisy critic (>75%>75\% noise), the generator LM learns to ignore the critic, as there is no difference between using the trained critic or the oracle during inference. Interestingly, training with a bit of noise (<50%<50\%) does not seem to harm the model, as performances are not statistically different than training with the oracle (noise of 0%0\%). Fig. 4 (b) depicts the quality of the critic used at inference time has a huge impact. Having oracle provide feedback is by far the best scenario. Already with 25%25\% noise, the critic makes the generator perform worse than using our trained critic (REFINER). With more than 50%50\% noise, the critic significantly harms the generator. The generator, trained with an oracle critic, has learned to trust the critic and expects useful feedback.

Qualitative Analysis. To explain the findings in §6, we further manually analyze 100 instances for the MWP task. We observe two different scenarios when REFINER failed to fix the outputs generated by generator model: (a) when the critic model provides a correct feedback; however, the generator model still generates incorrect equation, and (b) the critic model provides an incomplete or partially correct feedback. The former case indicates that either the generator model makes mistakes in following the instruction from the critic or the feedback from the critic can be ambiguous. For example, in Appx Fig. 5, (b) we observe the case when the critic is correct, but the feedback could result in an incorrect equation. The latter case indicates that our trained critic model generates incorrect feedback, which can result in incorrect or partially correct equations. We also observe that our critic model failed to generate correct feedback when the generator model generates incorrect equations with multiple mistakes.

Quality of the feedback. To better understand the difference in the quality of the feedback, we compare our trained critic model with GPT-3.5. We assess the quality of the feedback on 500 instances per task and report the exact match scores in Table 8. Please note that we include instances where the critic feedback should say the solution is correct and hence generate ’No’. For GPT-3.5, we have provided (two) few-shot examples per type of error and two examples with ’No’ as feedback. Our results show that trained critic (UQA) can comprehensively outperform GPT-3.5. We observe that GPT-3.5 performs well in identifying when the answer is correct. However, it makes errors when asked to generate meaningful semi-structured feedback for incorrect reasoning steps.

Conclusion

In this paper, we propose REFINER, a framework to improve the reasoning abilities of LMs through an iterative feedback loop between two models, a generator and a critic. Our evaluation of this framework on three reasoning tasks showed structured and fine-grained feedback on intermediate reasoning errors results in significant performance gains, surpassing scalar value feedback. Our trained critic model alone, even when noisy, can improve intermediate representations of LMs, showing that REFINER can significantly boost LMs’ performance on reasoning tasks. Our REFINER framework is very general and, in principle, might be applied to steer language models in performing different reasoning tasks. More specifically, the critic model can be seen as a tool for LLMs to refine their generation quality.

Acknowledgment

We would like to thank Martin Josifoski, Syrielle Montariol, and Zeming Chen for their helpful feedback on a draft version of the paper. We acknowledge the support of the ICT-48 Network of AI Research Excellence Center “TAILOR” (EU Horizon 2020, GA No 952215). West’s lab is partly supported by grants from the Swiss National Science Foundation (200021_185043), Swiss Data Science Center (P22_08), H2020 (952215), Microsoft Swiss Joint Research Center, and Google, and by generous gifts from Facebook, Google, and Microsoft. Antoine Bosselut gratefully acknowledges the support of Innosuisse under PFFS-21-29, the EPFL Science Seed Fund, the EPFL Center for Imaging, Sony Group Corporation, and the Allen Institute for AI.

Limitations

Our REFINER framework could not be comprehensively evaluated on all applicable downstream reasoning tasks due to their sheer number. While deliberately distinct, we focused on only three different reasoning tasks in order to study how natural language reasoning feedback can impact downstream tasks. We believe this represents an initial but important step towards exploring automated natural language feedback on intermediate representations. In addition, the critic we presented here is specific for each task, while the ideal critic would be a general one, capable of providing feedback on a wide range of reasoning tasks. Similarly, we considered fine-grained reasoning errors specific to each reasoning task. Recent work has mentioned several other fine-grained reasoning errors (Golovneva et al., 2023), which can’t be fully covered by the reasoning tasks we considered. Generalizing both the critic and fine-grained error types emerges as both the main limitations of this paper and the directions of future work. Finally, with LLMs being deployed more and more for real-life applications (medical domain, making important decisions), we believe it is crucial to develop expert models and automatic feedback mechanisms to inspect model generations and improve them. LLMs are impressive and work well on several NLP tasks, but they are not expert systems. Our work aims to address this gap by showing that adding interventions/feedback from critics (specialised finetuned critics) can help the LLM model to be more accurate—additionally, making the whole process more transparent.

Ethical Considerations

In this paper, we experiment with existing datasets which are, to the best of our knowledge, adequately cited. Our proposed framework REFINER is designed to improve the reasoning abilities of LMs. These LMs have been shown to encode biases about race, gender, and many other demographic attributes Weidinger et al. (2021), Sheng et al. (2020). Since our framework does not offer a way to mitigate these biases, models improved using this framework could still reflect the same harmful behaviours normally exhibited by these models. We recommend anyone deploying our model off-the-shelf should first check whether the model is harmful towards any protected group, and appropriate mitigation should be taken. In addition, our MS task is based on a dataset of situations, intentions, and actions that heavily skew towards Western culture and social norms Emelin et al. (2021). Consequently, our human evaluation on the MS task was done with AMT workers based in the US who were paid adequately for the average time it took to solve the task.

References

Appendix A Additional Results

Please note we also include instances where the critic feedback should say the solution is correct and hence generate ’No’. Our exact match metric is not order-sensitive. We extract the sentences and match them individually to the oracle answers. Since we focused only on the semi-structured critic feedback, automatic evaluation can already capture (measure effectively) the quality of the feedback.

A.2 Details about ReACT and Self-consistency and Self-Correct

The ReACT method consists of the reason model (Reason-Only) LLM (GPT-3.5), which generates a single thought at each step, and the Action model LLM (another GPT-3.5) does the calculation and generates the intermediate outputs (observations). We propose to refine the intermediate steps generated by the above steps and report the results below. Please note ReAct is approx 3-4 times more expensive than GPT-3.5 + CoT. In our experiments, we assumed 33 reasoning steps for ReACT and a sample size of 55 for self-consistency to be more cost-effective. Interestingly, we observe that ReACT perform similarly to CoT for the SVAMP dataset. One intuitive reason is that the SVAMP dataset contains questions which require one or two-hop reasoning only. We find that REFINER performs (+2.2) better than Self-correct (Welleck et al., 2023) on the GSM8K dataset, indicating the importance of correcting the intermediate steps can lead to better performance. Please note that we have used GPT-Neo as the generator model and the Unified QA T5-base model as the critic model, consistent with the Self-correct paper by Welleck et al. (2022).

A.3 More results on SVAMP dataset

In the MWP, for the answer prediction task, we compare REFINER with the previously reported baselines from Jie et al. (2022) including Graph2Tree (Zhang et al., 2020) that uses quantity relations using GCN; GTS (Xie and Sun, 2019) which is a sequence-to-tree model that mainly uses a tree-based decoder with GRU; and DeductReasoner (Jie et al., 2022) which uses bottom-up DAG-structured decoding. Results of this comparison can be found in Table 9. For the sNLR task, we also experiment with a critic model trained on 50% of its original training data and we still observe a performance improvement over the baseline as can be seen in Table 14.

Appendix B REFINER Framework

Alg. 1 and Alg. 2 outline the training and inference algorithms for REFINER. We train a supervised critic model (πβ\pi_{\beta}) with the context (xx) and (plausible or implausible) hypothesis (zz or z′z^{\prime}) as input and the textual feedback as output. Given a context xx the generator model (πθ\pi_{\theta}) is trained to generate plausible hypotheses.

Appendix C Datasets and Models

In Table 10 and Table 12, we report the data statistics and dataset details. In Table 11, we report the details of the used models. Our research is conducted solely on datasets that are in the English language.

Appendix D Training Details

For each task, we train a UnifiedQa-T5-base model (UQA-base) (Khashabi et al., 2020) as a critic (§3.1). Further evaluation details are provided in Appendix G. For exploration (§3.2), we use nucleus sampling with p=0.5p=0.5. We select the hyper-parameters by the validation loss: for both the generator and critic model, we use the Adam optimizer with a learning rate of 1e−41e^{-4}. Each model is trained for 2020 epochs with early stopping based on validation loss. We trained all models on one A100 GPU. We run our models with 33 random seeds and report the average results. We perform a binomial sign test. We find that p-values are always <0.05 when we compare REFINER with all the baselines (GPT-3.5, Self-refine, Self-reflection), suggesting our results are not random and significant. For the human study, we selected outputs from the best models (baselines and our model) according to automatic metrics. We train models with T=3T=3 iterations. We trained the critic model for 8 hours and trained the generator model for 12 hours.

At inference time, we use greedy decoding for the generator and critic model with T=1T=1 for the automatic critic and T=3T=3 for the oracle critic. We evaluate our methods using the metrics presented in the original papers that proposed the tasks. On the MWP and sNLR tasks, we use the exact match (EM) metric for intermediate steps (equation generation and inference rules) and accuracy (Acc) for the final answers. For MS, we conduct a manual evaluation study to assess the relevance of norms and moral actions.Since the automatic scores such as BLUE, ROUGE, etc. only account for word level similarity between gold norms or actions and generate norms or actions.

Appendix E Qualitative Examples

Figure 7 and 20 depict a qualitative example of REFINER where REFINER could correct incorrect equations through structured feedback, fixing the operators within a multistep solution. Table 20 shows some qualitatively improved examples for MS.

Appendix F Feedback Data Generation

Based on these error types, we perturb the plausible hypotheses (zz) in the training data and collect a pool of data DD (xx: input, zz: plausible hypothesis, z′z^{\prime}: implausible hypothesis). We perturb by omitting, replacing or adding some tokens or some rules from the plausible hypothesis to automatically create an implausible hypothesis. For example, in Fig. 6, for sNLR we omit a few inference steps from the correct hypothesis "#0: viridian is green, #1: rose is green" and create an incorrect (incomplete) hypothesis (see Fig. 6). Since our perturbations are based on logic and reasoning errors, we create structured feedback ff for every example (x,z,z′x,z,z^{\prime}) by stating the error type that occurs in z′z^{\prime} but not in zz (see Table 1). The basic structure of feedback ff for these tasks is ⟨\langleerror type, position (optional), hint (optional)⟩\rangle, where position denotes the error position in the implausible hypothesis (see Appx Table 1). For example, in the previous scenario, we create feedback “Missing link between fact and rules”. Despite the simplicity of the strategy we used for our tasks, this approach is easily generalisable to other reasoning tasks.

For MWP and sNLR problems, the underlying reasoning requires symbolic systems with closed-world rules. Hence, we consider a simple rule-based method to automatically generate the pairs of errors and their corresponding structured feedback by considering the error types and position of the errors (see Fig. 6 and Table 1).

In the moral norm generation task, we consider two kinds of fine-grained errors: logical contradiction and semantic misalignment (incoherent, uninformative). Moral norms are people’s subjective judgments about the character and actions mentioned in the context. Each moral norm is a combination of two components (implicit structure): a moral judgment [You shouldn’t] and an action [criticize your family’s religion]. Firstly, to create logical contradictions, we use the concept of deontic logic from Kiehne et al. (2022) and derive new norms contrary to those of Moral Stories. Hence, we replace the correct moral judgments in the plausible hypothesis with inverse judgments. For example, replacing [You shouldn’t] from the plausible hypothesis to [It’s good], as depicted in Fig. 6. To scale such inverse norms (implausible hypothesis), we paraphrase them by substituting the adjectives with synonyms from WordNet. Secondly, to create semantic misalignments, we must collect implausible hypotheses that are either misaligned with the plausible hypothesis or incomplete in nature. To create them, we replace the correct action (verb phrase) from the plausible hypothesis with random verb phrases selected from the context of the plausible hypothesis.

F.2 Synthetic Feedback Generation

We used a few-shot setting where we varied the instruction, the number of demonstrations, and the formatting of the demonstrations. Since data generation with GPT-3.53.5 is expensive, we generated 3030K, 2020K, and 3030K implausible hypotheses for MWP, sNLR and MS tasks, respectively.

Appendix G Human Evaluation on Moral Stories

As part of the human evaluation of model generations on MS, we asked Amazon MTurk (AMT) annotators to judge the relevancy of the generated norm and the moral action based on a Likert scale, with 1 = strongly disagree, 2 = disagree, 3 = unsure, 4 = agree, and 5 = strongly agree. Ratings were subsequently aggregated, with scores ≥\geq 4 deemed to be Relevant and with scores, ≤\leq 2 deemed to be Irrelevant while ratings with score 3 (Unsure) left as is. More specifically, we asked three different human judges to evaluate each example. We performed majority voting over answers with the rating Unsure assigned to those examples with no clear majority winner. In Figures 8 and 9, we report a complete breakdown of evaluation results for both norm and moral action. We also report agreement scores computed according to Krippendorff’s α\alpha Krippendorff (2018) in Table 4. The low and moderate α\alpha values indicate that judging the plausibility of moral norms and actions is a challenging task. In Figures 10-18, we provide excerpts of HIT instructions given to AMT workers during moral norm and action evaluation. Each task was supplemented by an Acceptance and Privacy Policy (Figure 18) that explains participation and data collection terms. All workers were based in US and paid $0.10 per task which took around 5 minutes to complete on average.