Improving Code Generation by Training with Natural Language Feedback

Angelica Chen, Jérémy Scheurer, Tomasz Korbak, Jon Ander Campos, Jun Shern Chan, Samuel R. Bowman, Kyunghyun Cho, Ethan Perez

Introduction

An important task for the field of software engineering is program synthesis, the automatic generation of computer programs from an input specification (e.g. a natural language task description or a set of input-output examples) (Manna & Waldinger, 1971). Effective program synthesis can not only improve the efficiency of software developers (Ziegler et al., 2022), but also increase the accessibility of writing code in general. Recently, pre-trained large language models (LLMs) have demonstrated impressive success on program synthesis (Chen et al., 2021; Li et al., 2022; Austin et al., 2021; Nijkamp et al., 2022; Xu et al., 2022, inter alia) but still struggle to consistently generate correct code, even with large-scale pre-training (Chen et al., 2021).

We hypothesize that these failures can be largely attributed to modern LLM pre-training set-ups. For instance, code pre-training datasets consist mostly of unfiltered code scraped from the Internet, which contains a significant number of security vulnerabilities (Kang et al., 2022) and bugs (Chen et al., 2021). This training signal also consists exclusively of offline demonstrations, without any signal from trial-and-error or interactive guidance that penalizes the model’s buggy outputs. As such, we hypothesize that supervising LLMs with explicit human-written feedback on the model’s own outputs can be more effective at training models to produce functionally correct code.

In particular, an intuitive and rich form of feedback to provide to LLMs is natural language feedback. We argue that LLMs are naturally able to incorporate written feedback, which has been shown to significantly improve a code generation model’s pass rates when the feedback is provided at test time (Nijkamp et al., 2022; Austin et al., 2021). In our work, we build upon this observation by exploring the use of natural language feedback during the training process itself, rather than just during inference. We conjecture that such feedback provides expressive and targeted information about a code generation model’s current failings in a sample-efficient manner. More broadly, this approach also represents a weak version of scalable oversight (Bowman et al., 2022), in that model overseers can improve a model merely by evaluating its outputs, without manually generating new demonstrations, in a way that takes advantage of the capabilities that are being supervised.

To train LLMs with language feedback, we propose an algorithm called Imitation learning from Language Feedback (ILF; Algorithm 1), which extends the work of Scheurer et al. (2022), who study the impact of learning from language feedback on text summarization models. Scheurer et al. (2022) improves a summarization model by training the base model on improved summaries generated from the model’s original summaries and human-written feedback. Our work builds upon Scheurer et al. (2022) in a number of ways: (1) by formalizing the algorithm and generalizing it into a form that can be applied to any task (our ILF algorithm in Section 2.2), (2) by detailing how the reward function can be adapted for code generation, and (3) by demonstrating a proof-of-concept of ILF for code generation.We open-source our code and annotated data at https://github.com/nyu-mll/ILF-for-code-generation. ILF improves the correctness of programs generated by a baseline code generation model πθ\pi_{\theta} by training a separate model πRefine\pi_{\text{Refine}} to use language feedback to repair the incorrect πθ\pi_{\theta}-generated programs. (We refer to the repaired programs as refinements.) We then improve πθ\pi_{\theta} by fine-tuning it on the πRefine\pi_{\text{Refine}}-generated refinements that pass unit tests, yielding a final improved model πθ∗\pi_{\theta^{*}}. This procedure may be run iteratively to continue improving the model, which we show can be seen as minimizing the expected KL divergence from a target ground truth distribution (Section 2).

We demonstrate a proof of concept of ILF for code generation by showing that it improves a CodeGen-Mono 6.1B model’s pass@1 rate on the Mostly Basic Python Problems (MBPP) benchmark (Odena et al., 2021) by 38% relative (10% absolute) over its zero-shot performance. It also outperforms fine-tuning on the MBPP-provided code by 64% (14% absolute, see Section 3.2). We further find that the refinements generated during ILF do indeed leverage the human-written feedback (Section 3.1) – when the feedback is unhelpful or irrelevant, we observe steep drops in code correctness. The quality of the feedback is also crucial – LLM-generated feedback yields far lower final pass rates than human-written feedback (Section 3.3). Despite the success of our approach, we still observe concrete limitations – for instance, πRefine\pi_{\text{Refine}} is less effective at incorporating feedback when the feedback addresses multiple bugs (Section 3.5), which suggests headroom for future work or more capable LLMs to base πRefine\pi_{\text{Refine}} on. Overall, our results – as well as our additional results on text summarization, using a similar technique in Scheurer et al. (2023) – suggest that human-written feedback is a powerful, information-rich form of supervision for LLMs.

Method

Here, we formally describe the problem we aim to tackle, before introducing our algorithm. Suppose we start with vocabulary V\mathcal{V} and a pre-trained language model πθ\pi_{\theta} parameterized by θ\theta. πθ:V∗→\pi_{\theta}:\mathcal{V^{*}}\to is a probability distribution over sequences of tokens x∈V∗x\in\mathcal{V}^{*}, where V∗\mathcal{V}^{*} is the Kleene closure of V\mathcal{V}. We also have a dataset of tasks D={(t,u)}\mathcal{D}=\{(t,u)\}. A task (t,u)(t,u) consists of a task description t∈Tt\in\mathcal{T} (e.g. “Write a function that computes the prime factorization of an input integer.”) and a suite u=\textscUnitTests(t)∈Uu=\textsc{UnitTests}(t)\in\mathcal{U} of unit tests associated with task tt. Finally, let \textscEval:V∗×T→{0,1}\textsc{Eval}:\mathcal{V}^{*}\times\mathcal{T}\to\{0,1\} be a unit test verification function that indicates whether a program x∼πθ(⋅ ∣ t)x\sim\pi_{\theta}(\cdot\,|\,t) passes all the unit tests in \textscUnitTests(t)\textsc{UnitTests}(t):

We also define a fine-tuning function \textscFinetune(πθ,D)\textsc{Finetune}(\pi_{\theta},\mathcal{D}) that applies a gradient-based optimization algorithm to πθ\pi_{\theta} using the associated loss objective calculated over dataset D\mathcal{D}.

2 Imitation Learning From Language Feedback

Our goal is to sample a diverse set of high-quality programs x1∼πθ(⋅∣t)x_{1}\sim\pi_{\theta}(\cdot|t) for any given task tt sampled from the task distribution p(t)p(t). We do so by fitting an auto-regressive LLM πθ\pi_{\theta} to approximate a ground truth distribution πt∗(x1)\pi_{t}^{*}(x_{1}) that assigns a probability to x1x_{1} that is proportional to its quality, as measured by a reward function RR. Fitting πθ\pi_{\theta} to approximate πt∗\pi_{t}^{*} can be seen as minimizing the expected KL divergence from πt∗\pi_{t}^{*} to πθ\pi_{\theta} over the task distribution p(t)p(t):

In this work we use the unit test verification function Eval directly as our reward function RR, but RR can also be a function of any number of other signals, such as stack traces or compiler outputs.

Minimizing the objective in Equation 2 is equivalent to supervised learning, i.e. minimizing the cross-entropy loss:

Rather than computing this loss over the exponentially large space of all possible x1x_{1}’s, we instead use Monte-Carlo sampling over a small set of x1x_{1}’s drawn from πt∗\pi_{t}^{*}. However, this is still intractable because we cannot sample directly from πt∗\pi_{t}^{*}. Instead, we approximate πt∗\pi_{t}^{*} using importance sampling with a proposal distribution qt(x1)q_{t}(x_{1}):

which assigns higher weights to higher quality programs x1x_{1}.

3 Proposal Distribution q𝑞q

Intuitively, we aim to design qtq_{t} to be as close as possible to πt∗\pi_{t}^{*}, which we accomplish by incorporating pieces of natural language feedback ff that give information about how to transform a low-reward program x0x_{0} into a higher-reward program x1x_{1}. This can be achieved by (i) identifying a program x0∼πθ(⋅∣t)x_{0}\sim\pi_{\theta}(\cdot|t) that does not currently pass the test suite (i.e. \textscEval(x0,t)=0\textsc{Eval}(x_{0},t)=0), (ii) asking for natural language feedback ff about bugs in x0x_{0}, (iii) using ff to transform the original program x0x_{0} into a refinement x1x_{1} that incorporates the feedback and passes the test suite (i.e. \textscEval(x1,t)=1\textsc{Eval}(x_{1},t)=1), and (iv) assigning higher weight to x1x_{1}.

We can formalize this procedure as follows. Let πψ(x1∣t,x0,f)\pi_{\psi}(x_{1}|t,x_{0},f) be a distribution over programs x1x_{1} that improve x0x_{0} by incorporating the feedback ff and pF(f ∣ t,x0,\textscEval(x0,t)=0)p_{\mathcal{F}}(f\,|\,t,x_{0},\textsc{Eval}(x_{0},t)=0) be the distribution of pieces of feedback ff for incorrect program x0x_{0} and task tt. We can then define our proposal distribution as:

where δ0\delta_{0} and δ1\delta_{1} are the Dirac delta distributions centered at 0 and 1, respectively. Then this proposal distribution is guaranteed to place higher probability mass on higher-quality programs (in terms of unit test pass rate) than πθ\pi_{\theta} since the term δ1(\textscEval(x1,t) ∣ t,x1)\delta_{1}(\textsc{Eval}(x_{1},t)\,|\,t,x_{1}) equals 0 for incorrect programs x1x_{1}.

We approximate sampling from qq by considering each of the terms in Equation 7 in order:

We first sample from πθ(x0∣t)×δ0(\textscEval(x0,t) ∣ x0,t))\pi_{\theta}(x_{0}|t)\times\delta_{0}\left(\textsc{Eval}(x_{0},t)\,|\,x_{0},t)\right) by rejection sampling from πθ\pi_{\theta}. In other words, we sample programs x0x_{0} from πθ\pi_{\theta} for task tt and only keep those that fail the test suite (i.e. \textscEval(x0,t)=0\textsc{Eval}(x_{0},t)=0; step 2 of Algorithm 1).

We approximate sampling from pF(f∣t,x0,\textscEval(x0,t)=0)p_{\mathcal{F}}(f|t,x_{0},\textsc{Eval}(x_{0},t)=0) by having humans annotate programs x0x_{0} (paired with their corresponding task descriptions tt and test suites uu) with natural language feedback (step 3 of Algorithm 1).

We approximate sampling from πψ(x1∣t,x0,f)\pi_{\psi}(x_{1}|t,x_{0},f) by sampling from πRefine\pi_{\text{Refine}}, a model capable of generating refinements given the task description, original programs, and human-written feedback.

Finally, the term δ1(\textscEval(x1,t) ∣ t,x1)\delta_{1}(\textsc{Eval}(x_{1},t)\,|\,t,x_{1}) corresponds to another filter: we only keep refined programs x1x_{1} that pass the test suite.

Next, we consider more concrete details of how this sampling is accomplished.

ILF assumes the availability of feedback but not necessarily of the repaired code/refinements, for a variety of reasons. We assume that program synthesis may be a task for which writing high-level natural language feedback is often less laborious than performing program repair. Although writing feedback involves identifying at a high level what is wrong with the program and how it should be fixed, program repair may involve the additional steps of refactoring, looking through documentation, and testing. Moreover, past work (Austin et al., 2021; Nijkamp et al., 2022) has indicated that certain large LLMs can proficiently incorporate the feedback at inference time, assuming access to accurate and high-quality feedback. As such, ILF assumes access to some model πRefine\pi_{\text{Refine}} that is capable of producing a refinement given the original program and feedback.

πRefine\pi_{\text{Refine}} can take a variety of forms, but we fine-tune a pre-trained CodeGen-Mono 6.1B model as our πRefine\pi_{\text{Refine}}. We create a training dataset for πRefine\pi_{\text{Refine}} by further annotating a subset of CannotatedC_{\text{annotated}} with refinements x1x_{1} that repair incorrect programs x0x_{0} by incorporating feedback ff, such that \textscEval(x1,t)=1\textsc{Eval}(x_{1},t)=1 for (x0,f,t)∈Cannotated(x_{0},f,t)\in C_{\text{annotated}}. Further details of our dataset and annotation procedure are in Section 3.

Experiments and Results

Having described our high-level approach, we now explain the experimental setup we use to test ILF.

We train and evaluate our models on the Mostly Basic Python Problems (MBPP) dataset (Odena et al., 2021). MBPP contains 974 Python programming tasks designed to be solvable by entry-level coders. Each task contains a natural language task description tt (e.g., “Write a function to return the prime factorization of the input.”), a gold solution, and a suite uu of three unit tests. Since the task descriptions are sometimes ambiguous, we include one unit test in the task description. The addition of the unit test helps to specify the input and output format of each task. We hold out the remaining unit tests for the evaluation of our generated programs.

MBPP includes a designated prompt/training/validation/test split of the dataset, but we re-split the dataset into the following splits:

MBPPRefine: These are tasks with IDs in the range 111-310 for which CodeGen-Mono 6.1B did not generate any correct completions. This split is used to train πRefine\pi_{\text{Refine}}.

MBPPTrain: These are tasks with IDs in the range 311-974 for which Codegen-Mono 6.1B did not generate any correct completions. This split is first used to evaluate the correctness of refinements generated by πRefine\pi_{\text{Refine}}. Then, the correct refinements in this split are used to train πθ\pi_{\theta} to obtain πθ∗\pi_{\theta^{*}} (step 5 in Algorithm 1).

MBPPTest: These are tasks with IDs in the range 11-110 that we use to evaluate the final performance of πθ∗\pi_{\theta^{*}}. Unlike the previous two splits, we use all tasks in this split, rather than only the tasks for which CodeGen-Mono 6.1B did not originally generate correct programs for. This allows us to better compare the baseline performance of πθ\pi_{\theta} with that of πθ∗\pi_{\theta^{*}}.

We use this modified split so that a larger portion of the dataset can be used to train the final model πθ∗\pi_{\theta^{*}}, whereas smaller portions are allocated for training πRefine\pi_{\text{Refine}} and evaluating πθ∗\pi_{\theta^{*}}. We do not make use of the prompt split (IDs 1-10).

Models

Throughout this paper, we use a pre-trained CodeGen-Mono 6.1B model (Nijkamp et al., 2022) as our πθ\pi_{\theta}. It is pre-trained sequentially on ThePile (Gao et al., 2020), BigQuery (Nijkamp et al., 2022), and BigPython (Nijkamp et al., 2022). We selected this model because it is open-source, can be fine-tuned on a single 4×1004\times 100 A100 (80 GB) node, and demonstrated pass@k scores comparable to Codex-12B (Chen et al., 2021; Nijkamp et al., 2022).

To implement our algorithm, we independently fine-tune two separate instances of CodeGen-Mono 6.1B to create πRefine\pi_{\text{Refine}} and the final model πθ∗\pi_{\theta^{*}}. We train πRefine\pi_{\text{Refine}} using pairs of incorrect programs and human-written feedback as inputs, with human-written refinements as targets (using the format in Figure 2). In contrast, we train πθ∗\pi_{\theta^{*}} using natural language task descriptions from MBPP as the inputs and πRefine\pi_{\text{Refine}}-generated refinements as the targets. Further training details are in Appendix A.1.

Evaluation

We evaluate all code generations in this paper using the pass@k metric introduced in Kulal et al. (2019). It estimates the rate for which ≥\geq1 of kk model samples passes all the unit tests. We use the empirical estimate of this quantity from Chen et al. (2021), an unbiased estimator given by:

for nn total programs (where n≥kn\geq k) and cc correct programs for the given task.

Human Annotation

We hire annotators via Surge AIwww.surgehq.ai to write both natural language feedback and refinements for incorrect programs generated by CodeGen-Mono 6.1B. For each task that CodeGen-Mono 6.1B generated no correct programs for, we ask the workers to first select one of the incorrect programs to write feedback and refinement for. We specify that the workers should select a sample that seems relatively easy to correct (i.e. could be minimally corrected to pass the unit tests). Then, they are asked to write feedback that describes what is wrong with the current code and how to fix it. For the refinement, they are asked to copy over the original code and make the minimum number of edits necessary to incorporate the feedback and pass all the unit tests. The full set of worker instructions can be found in Appendix A.2.

Although the ILF algorithm only requires the collection of human-written feedback for the tasks in MBPPTrain (assuming access to some πRefine\pi_{\text{Refine}} that is already fine-tuned or can generate refinements via few-shot prompting), we collect both human-written feedback and refinement for all splits of the data so that we can conduct further analyses of our method. For instance, this allows us to compare fine-tuning on πRefine\pi_{\text{Refine}}-generated refinements with fine-tuning on human-written refinements. When scaled to other pairs of model and task, ILF requires new feedback annotations, but it is possible that using ILF on one dataset will improve the model’s abilities on another dataset for a similar task. We leave analyses of scaling ILF across different tasks and models to future work.

1 CodeGen-Mono 6.1B Incorporates Feedback

We first verify that our baseline model can use feedback to repair incorrect code, a pre-requisite for ILF to work. We evaluate CodeGen-Mono 6.1B’s ability to generate refinements given pairs of (incorrect code, natural language feedback), both in a few-shot manner and after fine-tuning. Feedback is only required for tasks for which πθ\pi_{\theta} is initially unable to produce a correct response, so we first evaluate CodeGen-Mono 6.1B zero-shot on all of MBPP, generating 30 programs per task with temperature 0.8. Table 1 shows the resulting pass rates. There were 321 tasks for which zero-shot CodeGen-Mono 6.1B yielded no correct samples (from Table 1: (100%−67%)×974 tasks≈321(100\%-67\%)\times 974\text{ tasks}\approx 321). We then annotate one incorrect program per task with both feedback and refinement, as described in Section 3.

We use the human feedback annotations to create few-shot feedback prompts, formatted as in Figure 2. We evaluate CodeGen-Mono 6.1B’s ability to produce refinements that incorporate the feedback and pass the unit tests. However, producing a refinement that passes the unit tests does not guarantee that the feedback has been incorporated; there can be multiple solutions to a programming task, including ones that are functional but completely different and not using the feedback to improve upon the original code. Alternatively, the model may already be able to repair programs without feedback. Thus, we also evaluate the pass rate after shuffling the feedback samples in the dataset, to evaluate if the model’s ability to repair code degrades when presented with unrelated feedback.

The results are shown in Table 2. CodeGen-Mono 6.1B’s ability to incorporate relevant feedback on this particular set of program is low, with pass@10 reaching only 13.8%. However, the gap in accuracy between CodeGen-Mono 6.1B-generated refinements on relevant versus irrelevant feedback is significant, with pass@10 decreasing by 71% (relative; 13.8% →\rightarrow 4.0%), indicating that the model is indeed using the feedback.

Next, we examine whether we can improve our ability to repair programs given feedback by fine-tuning a separate model specifically to perform this task. Our training examples consist of triples of incorrect program, human-written feedback, and human-written refinement. We train the model to maximize the likelihood of the refinement given the program and feedback. The incorrect programs were generated by CodeGen-Mono 6.1B zero-shot on MBPP tasks, and the feedback and refinements were written by human annotators, as discussed in Section 3. We only included tasks for which none of CodeGen-Mono 6.1B’s generated programs were correct, yielding 44 tasks in the training dataset (forming the split MBPPRefine) and 128 tasks in the evaluation dataset (forming the split MBPPTrain). We asked human annotators to write refinements of the original code that incorporated their own previously written feedback, passed the unit tests, and made only minimal edits to the code (see Section 3). The format of the training data also matched the few-shot prompt format (Figure 2) but without the in-context examples of refinements. We denote this model as πRefine\pi_{\text{Refine}}, as described in Section 2.3.

Table 3 shows the pass rates for πRefine\pi_{\text{Refine}} on the evaluation dataset, which were produced by sampling 30 refinements per task with temperature 0.8. Fine-tuning significantly improves CodeGen-Mono 6.1B’s ability to incorporate feedback compared to 1-shot refinement, increasing pass rates more than three-fold (2→\rightarrow19% pass@1, 13.8→\rightarrow47% pass@10, from Tables 2 and 3). Furthermore, 61% of tasks had at least one correct refinement. This is particularly significant when considering the fact that we selected only tasks for which a non-finetuned CodeGen-Mono 6.1B model did not originally output any correct programs for (the rightmost column in Table 3). For the 61% of validation tasks that πRefine\pi_{\text{Refine}} generated a correct refinement for, we randomly selected one such correct program for each task to form the training dataset for our final model πθ∗\pi_{\theta^{*}}, yielding a final training dataset of 78 examples.

2 ILF Yields Pass Rates Higher Than Fine-Tuning on Gold Data or Human-Written Programs Alone

Given that our refinements improve over the initial programs, we now fine-tune on the refinements to improve our code generation model. As discussed earlier, we use the correct refinements (as evaluated by the unit tests) that πRefine\pi_{\text{Refine}} generated for its evaluation dataset as the training dataset for πθ∗\pi_{\theta^{*}}. Since πθ∗\pi_{\theta^{*}} is meant to generate code from a natural language task description (rather than to incorporate feedback into a refinement), the inputs of our training dataset are the MBPP prompts and the targets are the 78 πRefine\pi_{\text{Refine}}-generated refinements described in the previous section. We also compare the performance of πθ∗\pi_{\theta}^{*} against that of CodeGen-Mono 6.1B evaluated in a zero-shot manner, CodeGen-Mono 6.1B fine-tuned on the gold programs from the MBPP dataset, and CodeGen-Mono 6.1B fine-tuned on our human-written refinements. For all fine-tuning experiments, we train on programs corresponding to the same set of task IDs as the ones used in πθ∗\pi_{\theta^{*}}’s training dataset.

Additionally, we evaluate the impact of ablating the human annotations in our algorithm by using an LLM in place of humans to generate the feedback and refinements (replacing steps 3 and 4 in Algorithm 1). For the LLM, we use GPT-3.5 fine-tuned with Feedback Made Easy (FeedME; text-davinci-002 on the OpenAI API)Details at beta.openai.com/docs/model-index-for-researchers. We refer to this model as InstructGPT, which is the series of OpenAI models that FeedME belongs to (OpenAI, 2022). We use InstructGPT to generate both the feedback and refinements on the original programs. We then fine-tune CodeGen-Mono 6.1B on the model-generated refinements.

The results of our ILF algorithm compared to the baselines and ablations are shown in Table 4. ILF yields the highest pass@1 and pass@10 rates, despite how few samples of feedback and refinements we use. The pass@1 rate in particular shows a significant increase in improvement over the zero-shot baseline, representing a 10% absolute increase (38% relative increase). Pass@1 improvements are especially helpful for assisting with software engineering, where it is more helpful to suggest a single correct completion rather than 10 possible completions for the user to select from.

Compared to the gold standards, ILF outperforms both fine-tuning on MBPP gold programs and human-written refinements on the pass@1 metric, yielding 14% absolute (64% relative) and 3% absolute (9% relative) increases in pass@1 rates, respectively. However, training on human-written refinements yielded comparable pass@10 rates as ILF, which is unsurprising since πRefine\pi_{\text{Refine}} was trained on human-written refinements. When human-written feedback and πRefine\pi_{\text{Refine}}-generated refinements are ablated (the “Ablations” section of Table 4), ILF also outperforms training on both 1-shot and 2-shot InstructGPT-generated refinements by 17% and 11% absolute (89% and 44% relative), respectively.

However, we also note the surprising fact that merely training on a small sample of the MBPP gold programs did not make a significant difference in accuracy over zero-shot inference. We speculate that the gold programs from the MBPP dataset may be somewhat out-of-distribution for CodeGen-Mono 6.1B. To test this hypothesis, we computed the perplexity of the MBPP gold programs, the πRefine\pi_{\text{Refine}}-generated refinements, and the human-written refinements using the pre-trained CodeGen-Mono 6.1B model. The results are shown in Figure 3. While the distributions of all three data sources look similar, the MBPP dataset contains more high-perplexity programs (i.e. programs with perplexity ≥102\geq 10^{2}) than either the πRefine\pi_{\text{Refine}}-generated refinements or the human-written refinements. As a result, it is likely easier for CodeGen-Mono 6.1B to learn from the latter two datasets, since they are closer to CodeGen-Mono 6.1B’s original distribution while still being functionally correct.

Furthermore, ILF is particularly useful for settings where large amounts of gold code are not available. In this setting, ILF can be thought of as a method of not only generating more training data, but training data that is closer to the model’s original outputs in data representation space and that specifically repairs the kinds of bugs that the original model generates. As a result, fine-tuning the model on πRefine\pi_{\text{Refine}}-generated refinements does not require adjusting the weights as much as fine-tuning the model on the MBPP gold programs would, even though both training datasets contain the same number of functionally correct programs.

3 Scaling Up Model Feedback Does Not Offer the Same Benefits As Human Feedback

Since high quality human feedback can be expensive to collect, we also evaluated how much model feedback might yield the same benefit as our sample of human-written feedback. To do so, we randomly select kk tasks from the set of MBPP tasks for which CodeGen-Mono 6.1B did not originally output a correct answer, and prompt InstructGPT to generate both the feedback and the refinement. We then evaluate the refinements for correctness and train CodeGen-Mono 6.1B on the correct refinements. We use k∈{50,100,200}k\in\{50,100,200\} and generate 30 output samples at temperature 0.8 for all stages of the experiment. We are limited to these kk values due to the small number of tasks we have in MBPPTrain, but future work may investigate scaling up these experiments by using larger datasets or automatically generating new tasks and unit tests for the training dataset. Further training details are listed in Appendix A.1.

The results are shown in Figure 4. Although increasing the quantity of InstructGPT-generated feedback offers modest improvements in pass rates, these improvements do not yield pass rates as high as those of πθ∗\pi_{\theta^{*}}, even though πθ∗\pi_{\theta^{*}} uses only a total of 122 pieces of feedback throughout its training process (44 for training πRefine\pi_{\text{Refine}} and 78 for generating refinements to train πθ∗\pi_{\theta^{*}} on). However, as pre-trained large language models continue to improve dramatically in quality, we expect that this gap between human- and model-written feedback will increasingly narrow.

4 Human Feedback Is More Informative Than InstructGPT Feedback

To better understand why human feedback produced greater improvements in pass rate than InstructGPT feedback, we randomly selected 50 samples of feedback for each source (i.e. human or InstructGPT) and annotated the number and types of bugs that each feedback sample addressed. The results are shown in Tables 5 and 6. We observed that InstructGPT often gave no feedback (e.g. “The code is correct” or “Great job!”), provided feedback that was irrelevant or incorrect, or restated the task description instead of addressing what should be repaired about the code. Despite this, InstructGPT’s refinements were often correct even if the feedback itself wasn’t. Human-written feedback addressed more bugs on average and never gave irrelevant feedback. We provide further examples of the differences between human and InstructGPT feedback in Appendix A.3.

Lastly, we explored whether the number of bugs addressed in the feedback affected πRefine\pi_{\text{Refine}}’s ability to repair the original code sample. The results are shown in Figure 5. The greater the number of bugs addressed, the lower the average pass rate of πRefine\pi_{\text{Refine}}’s refinements. This suggests that a promising direction for future work might consist of automatically decomposing the feedback into multiple steps and having πRefine\pi_{\text{Refine}} incorporate the feedback one step at a time. Indeed, Nijkamp et al. (2022) show that the CodeGen models are often more effective at following instructions when the instructions are given across multiple turns, and recent Chain-of-Thought work (Wei et al., 2022) illustrates a similar prompting technique.

Related Work

Our work builds on a large body of literature that explores the use of pre-trained LLMs for neural program synthesis. Many general purpose LLMs, although not pre-trained specifically for code generation, have demonstrated impressive proficiency at solving code challenges since they are pre-trained on large corpora of text such as The Pile (Gao et al., 2020) that contain a small percentage of code content (Austin et al., 2021; Wang & Komatsuzaki, 2021; Black et al., 2022; Nijkamp et al., 2022). Yet other recent LLMs for program synthesis are trained on solely source code files (Wang et al., 2021; Zan et al., 2022; Li et al., 2022; Xu et al., 2022), or on both text and source code documents – sometimes either in succession (Chen et al., 2021; Nijkamp et al., 2022; Bai et al., 2022a), in a mixed corpus (Workshop et al., 2022), or on mixed natural language-programming language documents (Feng et al., 2020).

Learning from Human Feedback

Our algorithm is inspired by a number of past works that have trained models to learn from feedback. A common technique is reinforcement learning from human feedback (RLHF Ziegler et al., 2019; Stiennon et al., 2020; Ouyang et al., 2022), which trains models to satisfy human preferences. However, our algorithm is closer to works that use natural language feedback, rather than comparisons between different choices. Elgohary et al. (2020); Austin et al. (2021); Nijkamp et al. (2022) all demonstrate that code LLM performance generally improves when prompted with natural language feedback, though Nijkamp et al. (2022) observes that the feedback is more effective when it is given one step at a time. Our work differs from these in that ILF learns from the feedback at training time, not at inference time.

Bai et al. (2022a) also uses natural language feedback during the training process, but as part of an RLHF algorithm instead where the feedback is used to solicit different responses from the digital assistant, the responses are ranked by crowdworkers, and the rankings are used to train the preference model. However, they note that this form of learning from natural language feedback does not measurably improve their code generation model more than simply prompting.

Outside of program synthesis, we show in our other work (Scheurer et al., 2023) that ILF is also effective for text summarization. In addition to re-formulating the reward function R(⋅)R(\cdot) for summarization, Scheurer et al. (2023) additionally demonstrates that an instruction-finetuned LLM can evaluate its own outputs and select the best one. Similar to our results on code generation, Scheurer et al. (2023) shows that ILF outperforms all supervised fine-tuning baselines on text summarization. This aligns with numerous other works that have explored supervision via natural language in other ways, such as via explanations (Camburu et al., 2018; Hase & Bansal, 2021; Pruthi et al., 2021; Lampinen et al., 2022, inter alia) and as part of RL systems (Fidler et al., 2017; Luketina et al., 2019; Lin et al., 2020, inter alia).

Conclusion

We have shown that ILF can significantly improve the quality of a code generation model, even with just a small sample of human-written feedback and refinements. This approach is theoretically justified as minimizing the expected KL divergence between πθ\pi_{\theta} and a target ground-truth distribution, where we acquire signal from the latter via human-written natural language feedback.

This approach is also appealing because it is not model-specific (in the sense that ILF can be used with any type of base model πθ\pi_{\theta}, assuming the existence of a sufficiently capable LLM to act as πRefine\pi_{\text{Refine}}), and can be conducted in multiple rounds to continuously improve the model. Furthermore, it is notable that our approach generates training data that is not only correct, but targets the specific kinds of bugs that the model is likely to output. In essence, it provides an online training signal that is missing from the offline pre-training set-up of modern LLMs. Our approach is also remarkably sample-efficient, yielding 38% and 64% relative increases in pass@1 rate over the zero-shot baseline and fine-tuning on MBPP data, despite fine-tuning on only 78 examples.

Our work opens up multiple avenues for promising future work. For instance, ILF can be applied iteratively over the course of multiple rounds whenever new information arrives (e.g. new Python syntax) or new bugs are discovered. As the pace of progress of modern LLM research continues to accelerate, it may soon be feasible to partially or fully automate the generation of natural language feedback (similar to ‘RL from AI feedback’ (RLAIF; Bai et al., 2022b) and our experiments in Section 3.3), greatly reducing both the time and cost necessary for collecting feedback. This direction of work is also particularly appealing because the learning signal is process-based rather than outcome-based, which has been shown to mitigate reward hacking and improve the correctness of intermediate reasoning steps (Uesato et al., 2022). Although further work is required to extend our method, ILF represents an exciting step forward in training LLMs with feedback that is rich, interactive, and sample-efficient.

Acknowledgements

We are grateful to Nitarshan Rajkumar, Jason Phang, Nat McAleese, Geoffrey Irving, Jeff Wu, Jan Leike, Cathy Yeh, William Saunders, Jonathan Ward, Daniel Ziegler, Seraphina Nix, Quintin Pope, Kay Kozaronek, Peter Hase, Talia Ringer, Asa Cooper Stickland, Jacob Pfau, David Lindner, Lennart Heim, Kath Lumpante, and Pablo Morena for helpful discussions and feedback about the design and implementation of this work. We are additionally thankful to Scott Heiner and Edwin Chen for extensive help with setting up our human annotation workflow and interface. EP thanks the National Science Foundation and Open Philanthropy for fellowship support. JAC is supported by a doctoral grant from the Spanish MECD. AC, SB, and KC are supported by National Science Foundation Awards 1922658 and 2046556. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the National Science Foundation. KC is additionally supported by 42dot, Hyundai Motor Company (under the project Uncertainty in Neural Sequence Modeling) and the Samsung Advanced Institute of Technology (under the project Next Generation Deep Learning: From Pattern Recognition to AI). This project has also benefited from financial support to SB by Eric and Wendy Schmidt (made by recommendation of the Schmidt Futures program), Open Philanthropy, and Apple. We also thank the NYU High-Performance Computing Center for in-kind support and OpenAI for providing access to and credits for their models via the API Academic Access Program.

References

Appendix A Appendix

For the experiments in Section 3.2, we run a hyperparameter sweep for all methods except for ILF. The hyperparameter value ranges that we sweep include learning rate ∈{1.0−6,5.0−6,1.0−5}\in\{1.0^{-6},5.0^{-6},1.0^{-5}\}, batch size ∈{32,64,128}\in\{32,64,128\}, and number of epochs ∈{1,2,5}\in\{1,2,5\}. The tasks for the training and validation datasets are from MBPPTrain and MBPPRefine, respectively, while the programs are sourced from the method (e.g. InstructGPT, MBPP, human-written, or zero-shot CodeGen-Mono 6.1B). For ILF, we use the best hyperparameters obtained for the sweep over MBPP programs instead of sweeping over ILF-generated programs, since the tasks in MBPPRefine are already used to train πRefine\pi_{\text{Refine}}. All pass rates reported in Table 4 are obtained by evaluating each method on MBPPTest using the best hyperparameters found during the sweep on MBPPRefine.

For the experiments in Section 3.3, we separately tune hyperparameters for each size of dataset. As in our other experiments, we train and validate using the tasks from MBPPTrain and MBPPRefine, respectively, coupled with the refinements generated by InstructGPT that pass the unit test suites. We sweep the same hyperparameter value ranges as the experiments in the previous section (i.e. learning rate ∈{1.0−6,5.0−6,1.0−5}\in\{1.0^{-6},5.0^{-6},1.0^{-5}\}, batch size ∈{32,64,128}\in\{32,64,128\}, and number of epochs ∈{1,2,5}\in\{1,2,5\}).

We implement all experimental pipelines with the HuggingFace transformers (v4.12.5) (Wolf et al., 2020), Huggingface datasets (v2.7.1) (Lhoest et al., 2021), and Pytorch (v1.11) (Paszke et al., 2019) libraries.

A.2 Annotator Instructions

A.3 Examples of Human Versus InstructGPT Feedback