Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Jacob Ginesin, Edward Berman, George Chakhnashvili, Anton Lozhkov, Carolyn Jane Anderson, Arjun Guha
Introduction
Large language models of code (Code LLMs) are starting to become an essential tool for software engineering practice and research. There has been significant research on synthesizing code from natural language instructions, but comparatively less attention has been given to code editing tasks. However, LLM users expect models to be capable of editing code. For example, the LMsys dataset of real-world conversations with chatbots (Zheng et al., 2023) has 4,188 conversations with code, and 831 (19%) of these involve edits, where the user prompts the model to update generated code based on natural language instructions (Appendix C). In general, code editing encompasses activities like feature addition or removal, bug fixing, and code refactoring (Zhang et al., 2023; Moon et al., 2023; Shinn et al., 2023; Chen et al., 2023; Olausson et al., 2023; Jin et al., 2023).
The ability to edit code is also essential for a model to be useful for an AI-focused code editor such as Cursor (Cursor, 2023), Copilot Chat (Copilot, 2023), or ChatGPT Advanced Data Analysis (ADA) (OpenAI, 2023b). Cursor and Copilot Chat facilitate edits with human-written instructions. In contrast, ADA uses both human-written instructions and model-generated reflections (Shinn et al., 2023) to extend and edit code. This approach represents a step towards fully AI-driven code assistance. In both scenarios, instructional code editing is employed, which we define as a function , where is the original code, is the instruction, and is the modified code. An example of this process can be seen in Figure 1, illustrating how the model edits a code segment from a given instruction.
Model-generated reflections and human-written instructions both describe desired code changes. However, they differ in the level of detail: reflections, usually more detailed, are generated by a model with access to the code, offering richer context and potentially a strategic plan for code modifications. In contrast, human-written instructions are typically shorter and less detailed but may express the true user’s intent more clearly. We refer to these as descriptive and lazy instructions, respectively.
In this work, we introduce CanItEdit, a novel dataset comprising 54 hand-crafted instructional code editing problems. These problems, featuring both descriptive and lazy instructions, are coupled with an extensive hidden test suite. Designed to assess a model’s proficiency in handling realistic code editing scenarios, CanItEdit serves as a platform for evaluating state-of-the-art Code LLMs in instructional code editing. Our evaluation focuses on measuring the accuracy of a given model’s ability to write correct code modifications without introducing superfluous code. We conduct comprehensive assessments of closed and open models, revealing significant performance disparities between the leading closed and open models in this domain (Section 5). To help address this gap, we propose a training dataset and methodology for code editing. Our findings demonstrate that fine-tuning open Code LLMs on this dataset can significantly enhance their performance (Section 4).
To summarize, we make the following contributions:
We introduce CanItEdit, an extensive and detailed collection of instructional code editing problems, designed to test a model’s ability to edit code under two levels of instruction detail (Section 3).
We propose a novel metric, ExcessCode, for assessing code editing models. This metric quantifies the volume of unnecessary code produced by a model when generating a correct solution (Section 5.1).
We perform a thorough evaluation of the latest Code LLMs in the context of code editing, providing insights into their current capabilities (Section 5).
We present a specially tailored dataset for code editing, along with an effective training methodology, demonstrating significantly enhanced code editing performance through fine-tuning a state-of-the-art Code LLM (Section 4).
The benchmark, models, training dataset, and the code to reproduce our work are available at:
Related Work
Correctly prompting an LLM is crucial for it to perform a desired task. There are multiple methods for instruction tuning LLMs to better adhere to natural language instructions. One method involves employing human annotators to create sample instructions and provide feedback on numerous model outputs (Ouyang et al., 2022; Köpf et al., 2023). However, this method is costly and demands substantial resources. An alternative, cost-effective method is to enable a proficient LLM to self-instruct, generating instructions from a smaller set of human-written seed instructions (Wang et al., 2023). These methods have been applied to generate datasets for instruction-tuning Code LLMs (Chaudhary, 2023; Muennighoff et al., 2023; Luo et al., 2023). Specific to code generation, another strategy to instruction tune an LLM is to use commit messages as instructions (Muennighoff et al., 2023). In this paper, we use commit messages as instructions for code editing. With regards to instruction-tuned models, our results demonstrate that while these models can edit code, they are not as effective as models that are explicitly trained for this task (Section 5).
Code Generation Benchmarks
Several benchmarks exist that test a model’s code generation ability. HumanEval and MBPP are two prominent benchmarks for evaluating Code LLMs in Python programming (Chen et al., 2021; Austin et al., 2021). MultiPL-E expands these benchmarks to 18+ additional programming languages (Cassano et al., 2023). These benchmarks assess model-generated candidate completions against a series of human-authored unit tests. EvalPlus (Liu et al., 2023) utilizes mutation testing to expand the test suites of the Python benchmarks. All of these benchmarks utilize the pass@k metric, which measures the likelihood of the model generating a completion that passes all of the tests in tries; we also adopt this metric in our evaluation (Section 5.1). However, these benchmarks are limited to the evaluation of a model’s ability to generate a single function from a natural language description and do not assess code editing capabilities. HumanEvalPack (Muennighoff et al., 2023) is a comprehensive benchmark designed for evaluating Code LLMs across various code generation tasks, such as synthesis, explanation for code understanding, and bug fixing. Specifically, HumanEvalFix, a bug-fixing variant of HumanEvalPack, is extensively used for assessing the models’ capabilities in code refinement (Moon et al., 2023; Muennighoff et al., 2023).
SWE-Bench (Jimenez et al., 2023) evaluates Code LLMs on a broad spectrum of tasks that are performed by real-world software engineers, and require planning, retrieval, code editing, and more for successful task completion. Our work is more narrowly focused on code editing, and we believe this focus will help guide model development. Another difference with SWE-Bench is that our benchmark is handcrafted, whereas SWE-Bench is based on PRs and issues from popular GitHub repositories. This increases the risk of contamination, particularly with models such as StarCoder, which is trained on several GBs of GitHub issues (Li et al., 2023a).
Code Editing Using Large Language Models
Previous studies on code editing with large language models (LLMs) have predominantly focused on bug fixing (Zhang et al., 2023; Moon et al., 2023; Shinn et al., 2023; Chen et al., 2023; Olausson et al., 2023; Jin et al., 2023; Joshi et al., 2023; Wei et al., 2023), a specific subset of code editing, fill-in-the-middle code completion (Bavarian et al., 2022; Fried et al., 2023; Yee and Guha, 2023; Rozière et al., 2023; AI, 2023), an inference strategy that requires specific insert locations, and intrinsic code editing (Li et al., 2023b; Gupta et al., 2023), which involves editing code without a specified instruction, exerting the model’s ability to intrinsically ascertain the desired code changes. Recently, LLMs have progressed in code editing guided by natural language without specific edit locations (Hu et al., 2023; Li et al., 2023a; Muennighoff et al., 2023). However, this advancement lacks benchmark evaluations to effectively measure the models’ code editing skills. Notably, StarCoder (Li et al., 2023a), the first LLM trained on an extensive dataset of commits using the format
The CanItEdit Dataset
This section presents CanItEdit, a dataset of Python code editing problems with natural language instructions, hand-written and cross-validated by experienced computer science experts for evaluating LLMs’ code editing capabilities.
CanItEdit features 54 Python code editing problems, each comprising a ‘before‘ and an ‘after‘ code segment, two types of natural language instructions (descriptive and lazy), and a hidden test suite. The task for models is to transform the ‘before‘ code segment into the ‘after‘ segment based on either instruction, aiming to pass the hidden tests. Inspired by HumanEval’s methodology (Austin et al., 2021), we hand-wrote the problems, avoiding public sources like GitHub to reduce pre-training exposure. We also verified that the instructions are unique to this dataset and not part of our fine-tuning data (section 4). Problems range from simple function edits to complex, multi-class challenges, covering data structures, algorithms, mathematics, language processing, and game programming. Some require popular external Python libraries like NumPy, Pandas, PyTorch, and Z3. Dataset statistics and example problems are detailed in Table 1 and Appendix B, respectively.
The ‘before‘ code segments in CanItEdit represent various starting states, ranging from functional programs needing additional features to those with bugs or incomplete implementations requiring fixes or optimizations. These segments are designed to mirror diverse real-world coding scenarios. Conversely, the ‘after‘ segments illustrate the correct solutions that fulfill the task requirements and clear the test suite.
We categorize code editing tasks into two types: evolve and revise. Evolve tasks involve adding or removing major features like new methods or classes. In contrast, Revise tasks focus on modifying existing functionalities, including bug fixing, logic changes, or refactoring, such as transitioning from imperative to object-oriented programming, as exemplified in Figure 7. The distinction between these categories is based on the edit’s primary objective, though some tasks may exhibit characteristics of both.
The dataset’s dual natural language instructions test model efficiency in two scenarios: 1) Descriptive: Detailed instructions replicate situations where users provide specific specifications or another model outlines a plan, similar to Reflexion prompting (Shinn et al., 2023; Fan et al., 2023; Phung et al., 2023). 2) Lazy: Informal instructions resemble typical user queries for LLMs in code generation (Babe et al., 2023).
In both, the model must generate code that meets the instruction and passes the hidden tests. Descriptive instructions offer detailed guidance, including function names and input-output examples, while lazy instructions provide minimal information, requiring the model to infer user intent and rely more on the ‘before‘ segment. Both instructions should lead to an equivalent ‘after‘ segment. For instance, in the hello_world problem (Figure 2), the descriptive instruction is comprehensive, whereas the lazy instruction is brief, pushing the model to deduce user intent. Further discussions on human-written and Reflexion-generated instructions are in Appendix C.
2. Test Suites
For our test suites, we ensure three essential properties:
Completeness: Each suite comprehensively covers all inputs and edge cases. This includes numerous test cases per problem, targeting edge and corner cases, with code coverage verified using Coverage.py (Batchelder and Contributors to Coverage.py, [n. d.]). It is worth noting that code coverage is not as robust as mutation testing, which is employed by EvalPlus (Liu et al., 2023).
Correctness: The suites are bug-free, passing all tests with the ‘after‘ code while failing at least one with the ‘before‘ code.
Concealment: Test suites are hidden from models during training and inference, achieved by excluding them from the training dataset and not presenting them during model evaluation.
We hand-crafted test suites for each problem, incorporating diverse testing methods ranging from simple unit tests to complex property-based testing, mocking, fuzzing, and integration tests. For instance, one of our benchmark problems involves implementing a strategy for a Tic-Tac-Toe game that outperforms a baseline strategy (Figure 9). The lazy instruction for this problem is: Create a strategy ‘GoodStrategy‘, that beats ‘CornerStrategy‘. Do not modify the ‘Game‘ class. To test this, the suite includes unit tests for both the ‘Game‘ and ‘CornerStrategy‘ classes, along with integration tests that evaluate the entire program, ensuring that ‘GoodStrategy‘ wins over ‘CornerStrategy‘. Additionally, we use Python’s inspect module to check if the ‘Game‘ class remains unmodified, adhering to the problem’s constraints.
Fine-tuning
This section outlines our methodology for fine-tuning a Code LLM specifically for code editing tasks. We fine-tune a model based on the DeepSeek-Coder-Base family of Code LLMs (AI, 2023), which are variants of Code Llama (Rozière et al., 2023) trained from scratch on 2T tokens comprised of 87% permissively licensed code from GitHub and 13% natural language, using the same filtering rules as StarCoder’s data collection (Li et al., 2023a). At the time of writing, these models are the top-performing, open-access foundational Code LLMs, excelling in various code generation benchmarks. Furthermore, they are distributed under a permissive open-source license, allowing free use and modification for research and commercial purposes. We selected these base models because they exhibit robust performance on CanItEdit, even without being specifically trained for this or any other instructional tasks, thus highlighting their exceptional generalization capabilities (Section 5). We hand-crafted a training dataset for code editing, which we describe in Section 4.1, and fine-tuned the DeepSeek-Coder-Base 6.7 billion parameter model on this dataset, which we refer to as EditCoder.
We experiment with a training dataset for code editing, which we refer to as EditPackFT. We created the EditPackFT dataset by further filtering the Python split of the CommitPackFT dataset (Muennighoff et al., 2023), which was used to train OctoCoder.
CommitPack is an extensive dataset comprising 4TB of permissively licensed commits from a 2016 GitHub snapshot across various programming languages. CommitPackFT is a subset of CommitPack, filtered for it to be amenable to instruction-tune Code LLMs. The primary criterion for CommitPackFT’s selection involved retaining commits whose messages begin with an imperative verb, mirroring the typical structure of natural language instructions. We apply a series of additional filtering steps, which make the dataset more suitable for code editing. We remove any item that passes any of the following predicates:
The presence of an empty ‘before‘ or ‘after‘ code segment, disregarding whitespace.
The inclusion of the words TODO, FIXME, or BUG in the ‘after‘ code segment, which signals an incomplete commit.
A Levenshtein distance below between the ‘before‘ and ‘after‘ code segments, indicative of trivial changes.
Incorrect parsing of the ‘after‘ code using the Python ast module.
Originally, the dataset contained 56,025 commits, and after applying the filtering steps, we are left with 22,602 commits. As shown by Figure 4 and Table 2, the mean number of lines in the ‘before‘ and ‘after‘ code segments is . We further analyzed the original CommitPackFT dataset, ensure that our filtering wasn’t the cause of the short code segments, and find that the mean number of lines is similar, with a mean of . We recognize that while this distribution may be suitable for small-scale code editing tasks, it is not representative of real-world scenarios, where the code segments are typically longer and more complex. We also analyze the distribution of the commit message lengths, and find that the mean token count is . Figure 3 illustrates a sunburst plot of the most frequent initial verbs in the commit messages of Commits2023FT, along with their corresponding root nouns. The set of initial verbs in EditPackFT is composed of unique verbs, and the most frequent verb is add, which appears in commits.
2. Training Tools and Configuration
For training EditCoder, we utilize the Finetuning-Harness (Cassano, 2023), a fine-tuning pipeline based on the HuggingFace Transformers library (Wolf et al., 2020). Additionally, we utilize DeepSpeed ZeRO 3 (Rajbhandari et al., 2020) to efficiently shard the model and dataset across multiple GPUs. We also use FlashAttention 2 (Dao, 2023) to speed up training on large context window sizes. All of our models are trained on a single machine equipped with 8 NVIDIA A100 (80GB) GPUs. The effective micro-batch size is set at ( gradient accumulation steps, with a single batch per GPU).
We employ a learning rate of with linear decay and warmup steps. All models underwent training for epochs, with a constant, unpadded context window of tokens. Prior to training, we shuffled the dataset randomly and deduplicatedDeduplication, achieved by concatenating the ‘before‘ and ‘after‘ code segments, helps mitigate overfitting to specific training examples (Lee et al., 2022). it following the method outlined by Li et al. (2023a). This process combines MinHash (Broder, 2000) and Locality Sensitive Hashing (LSH) (Leskovec et al., 2014). We format the training data as a prompt, with the ‘before‘ code segment followed by the ‘instruction‘ and the ‘after‘ code segment, as show in Figure 5.
Evaluation
In this section, we evaluate the performance of various open and closed-sourced models on the CanItEdit benchmark, as well our fine-tuned models.
We run the open-access models using HuggingFace Transformers (Wolf et al., 2020) and vLLM (Kwon et al., 2023). We use the following hyperparameters for all inference experiments: batch size , maximum new tokens, temperature , and top- sampling cutoff of . We run all tests in a Docker container to mitigate the risk of malicious code execution.
Models evaluated
We evaluate several state of the art models with varying sizes, and also fine-tune a model to build EditCoder. We group the models into three categories: open, closed-sourced, and distilled open models, where the latter are open models fine-tuned on data generated by proprietary models such as GPT-4. The full list of models and their sizes appears in Table 3.
Finally, we were careful in formatting each benchmark problem to use prompt formats that the models’ developers recommend. The specific formats appear in Appendix A.
1. Evaluation Metrics
We employed two metrics to assess the performance of different models: one for functional correctness and another for the conciseness of the code edits.
pass@1 calculates the average fraction of successful completions per problem in CanItEdit, where success is defined as a completion passing all unit tests. Following Cassano et al. (2023), we generated 20 completions per problem.
Besides functional correctness, we assess the conciseness of model-generated code edits using the ExcessCode metric. This metric evaluates the presence of unnecessary code in successful completions by calculating the percentage of superfluous code, as indicated by the percentage line coverage in the generated code. We calculate this metric by averaging the median line coverage for passing completions across all problems, omitting those with no successful completions.
2. Results with Existing Models
We draw several conclusions from the full results in Table 3.
Larger models are better at editing; small models generate more excess code. Generally, model size correlates positively with pass@1 and negatively with ExcessCode. This indicates that larger models are more adept at precise functionality addition.
Models pre-trained on commits are better at code editing. Of the open models, the StarCoder model family is unique because it is pre-trained on a sample of GitHub commits (Li et al., 2023a), and we use the StarCoder commit data format when we evaluate the StarCoder models. We find that StarCoder models are significantly better on our benchmark than the pre-trained DeepSeek and Code Llama models, despite the fact that the latter two models outperform StarCoder on code generation (AI, 2023).
Models are generally better at following descriptive instructions than lazy instructions. Models generally perform better with descriptive instructions, likely because these provide more specific code details. However, some smaller models like Deepseek-Coder-Base-1.3b and StarCoderBase-1b perform better with lazy instructions, possibly due to their limited capacity to process longer detailed instructions. Detailed statistics are available in Table 1.
Closed and distilled models outperform open models. The comparison between CodeLlama-Instruct, a generic instruction-following code generation model, and GPT-4, a broad instruction-following model, highlights the performance gap between open and closed-sourced models (Rozière et al., 2023; OpenAI, 2023a). In terms of pass@1, GPT-4 outperforms CodeLlama-Instruct-34b by and for descriptive and lazy instructions, respectively, confirming the significant gap in instructional code editing abilities between state-of-the-art open source and proprietary models.
3. Results after Fine-Tuning on Commits
In addition to evaluating existing open models, we also fine-tuned a pre-trained DeepSeek model (section 4) to build EditCoder, which we now evaluate.
Fine-tuning on open commits can significantly improve code editing performance. EditCoder surpasses all open models, showing an increase in pass@1 and a notable decrease in ExcessCode compared to StarCoderBase-7b for descriptive instructions. For lazy instructions, EditCoder performs better than StarCoder, the best performing open source model in this category, and has a perfect ExcessCode score of .
Conclusion
We present CanItEdit, a benchmark designed to assess the instructional code editing skills of Code LLMs. It includes hand-written code editing problems, each accompanied by dual natural language instructions: a “lazy” instruction that a human may write, and a “descriptive” instruction that may be generated by an agent revising code in a loop. Each problem has a comprehensive test suite. We evaluate contemporary state-of-the-art Code LLMs and reveal a significant gap between closed and open models. We also demonstrate that fine-tuning with a custom dataset and training methodology can significantly improve code editing capabilities across various model sizes. Our work provides a foundation for evaluating future enhancements in instructional code editing for Code LLMs, offering valuable tools and insights for AI-based software development research and practice.
We evaluated models in reproducing the entire ‘after‘ code segment, which may not be the most token-efficient method. A potentially more efficient strategy would involve generating a list of specific changes to be applied to the ‘before‘ code segment. Furthermore, our study does not explore varying prompt formats. Instead, we have adopted a format consistent with that used by other models (Li et al., 2023a). Another limitation is the size of our final training dataset, which is relatively modest. We have not investigated the potential benefits of utilizing larger datasets, which could notably enhance performance, particularly in larger models. We identify these areas as opportunities for future work.
References
Appendix A Prompts Used in Evaluation
We evaluate all of our models on CanItEdit using the same evaluation pipeline. However, for each model, we may utilize different prompts to generate the completions. These prompts are most aligned to how the model was trained, and are intended to maximize the model’s performance on the task, while keeping the prompts as similar as possible across models. Figure 6 shows the prompts used for each model.
Appendix B Example Benchmark Items
In this section, we showcase four examples from the CanItEdit benchmark, which are representative of the types of problems present.
Figure 7 details a task where the model refactors code using object-oriented programming (OOP) principles. Initially, the code is a function for formatting messages based on type. The refactoring involves creating TextMessage and ImageMessage as subclasses of an abstract Message class and implementing a MessageFactory for message construction.
This task provides an example of a revise edit (Section 3.1), focusing on reorganizing the code into an OOP style without adding new features. The transformation is quite significant, and the largest relative transformation in our dataset: from a single function to a multi-class OOP program.
The goal is to assess the model’s proficiency in converting functional code into well-structured OOP designs based on comprehensive instructions and for the model to restructure small programs into much larger ones. Our test suites verify both functional correctness and the proper hierarchical class structure.
B.2. group_theory
Figure 8 features a task to modify a class from representing group to group , including its operations like inverse and product. This task is an exemplary revise edit, focusing on significantly adapting an existing class rather than adding new features.
The problem also highlights domain-specific problems in CanItEdit, this one being set in the context of cyclic groups. Testing domain-specific edits is crucial, especially when comparing the capabilities of large proprietary models like GPT-4 with smaller open models. It requires the model to transform the C4 class (representing a 4-element cyclic group) into the C8 class (for an 8-element group), requiring extensive edits across various code sections. This complexity presents a significant test for other code editing approaches, such as fill-in-the-middle (Bavarian et al., 2022; Fried et al., 2023), which may struggle with multiple edit locations (Yee and Guha, 2023).
Key edits involve altering the size and elements methods. The necessary understanding for these modifications stems from group theory, which is not explicitly explained in the problem. This setup tests the model’s capability to execute domain-specific edits where contextual knowledge is implied rather than provided.
B.3. strategy
Figure 9 presents an open-ended problem where the model devises a game strategy to defeat the already implemented CornerStrategy in Tic Tac Toe. This task represents an evolve edit, focused on developing a new feature without altering existing classes.
The uniqueness in this problem lies in the lack of providing rules for the game, but rather requiring the model to infer them through understanding of the code. Additionally, it leaves the strategy design entirely to the model’s discretion. Our tests ensure that the Game class remain intact and that the model’s strategy consistently outperforms CornerStrategy in the game.
B.4. sudoku_solver
Figure 10 presents a sudoku solver problem leveraging the Z3 satisfiability modulo (SMT) solver. The problem starts with an incomplete solver that lacks checks for 3x3 subgrids, both in its solving logic and board validity function. In sudoku, each 3x3 grid must contain distinct numbers from 1 to 9. The task involves adding these checks to ensure the solver can correctly solve a sudoku board. This problem assesses the model’s capability to implement edits across different code sections. Although it uses Z3, in-depth knowledge of the library or SMT isn’t required; the necessary features needed to solve the problem can be inferred from the existing code, which already includes checks for row and column uniqueness.
Appendix C Using LLMs in Code Editing Tasks
In this section, we provide a brief overview of the use of LLMs in code editing tasks. We showcase two scenarios: (1) humans interacting with chat models to edit code, and (2) models automatically generating edits for code. For the former, we analyze a large dataset of LLM chatbot interactions, ”lmsys/lmsys-chat-1m” which can be found on HuggingFace’s hub, and for the latter, we analyze a sample reflection generated by GPT-4 using the Reflexion algorithm (Shinn et al., 2023).
We analyze a large dataset of human interactions with 25 different conversational LLMs, users to interact with a highly capable chatbot. The dataset, ”lmsys/lmsys-chat-1m”, contains 1-million real-world conversations from 25 conversational LLMs of varying sizes and capabilities. We analyze the dataset to understand how humans interact with LLMs to edit code. We find that 4188 of the 1-million conversations contain a code-related request, and that 831 of those conversations contain a code editing request. We found this number by searching for markdown-formatted code blocks in the conversations, therefore the actual number of code-related requests is likely higher. We analyzed a subset of code editing requests to understand the types of requests humans make to LLMs. We find that almost all of the requests are of the ”lazy” kind that we include in CanItEdit. We provide two examples of human editing requests in Figure 11. The first example is a request to refactor a Python code snippet, and the second example is a request to refactor a JavaScript code snippet. As shown, these requests are very informal and direct, and do not provide any information about the desired solution. Other instructions we found that we think exemplify the type of instructions humans give to LLMs include:
Please change use scrappy instead request.
Can you change above code to not use histogram but use two for loops to create the histogram?
Very cool. Now change it so that it compresses each file using lz4 and saves it to a file with the same name and extension, + ”.lz4”
C.2. Model-Generated Instructions for Editing Code
This section delves into an example of code editing guided by instructions generated by GPT-4 using the Reflexion algorithm. Reflexion is a versatile algorithm developed for enhancing model output through environmental feedback, as detailed in Shinn et al. (2023). While its application extends across various tasks, including reasoning and decision-making, its utility in program synthesis is particularly notable. The process starts with generating unit tests for a program given its natural language description, followed by the creation and evaluation of a candidate program against these tests. If the program fails, Reflexion induces the model to produce a reflection, identifying potential errors and suggesting corrections. This reflection serves as an instruction for modifying the failing program, which are both provided to the model to edit the failing program into a new candidate, iterating until it passes all tests or a predetermined stop condition is reached.
We provide an example of a model-generated instruction for code editing in Figure 12, where the model was tasked with addressing a problem from the LeetCode Hard problem set. The instruction, precise and detailed, pinpoints the specific issue in the function’s logic and suggests a clear approach for rectification. It emphasizes iterating through marble distributions to calculate the maximum and minimum scores, a method not implemented in the original code. This example showcases how Reflexion can guide models to not only identify errors in logic but also propose viable solutions. This kind of guided instruction is useful for enhancing the accuracy and efficiency of models in complex code editing tasks; however, it is important to note that the instruction is not a complete solution, and that these models may produce misleading or incorrect instructions. The instruction is quite verbose compared to the human examples shown in Figure 11, and it is unclear how humans would interact with such an instruction, as this amount of detail is not necessary for the task at hand.