Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Chunqiu Steven Xia, Yinlin Deng, Lingming Zhang

Introduction

Program synthesis is widely regarded as the holy-grail in the field of computer science. Recently, large language models (LLMs) have become the default choice for program synthesis due to its code reasoning capabilities acquired through training on large amounts of open-source code repositories. Popular LLMs like GPT-4 , Claude-3 , and Gemini have shown tremendous success in aiding developers on a wide-range of coding tasks such as code completion , repair , and test generation . Furthermore, researchers and industry practitioners have designed code LLMs (e.g., DeepSeeker Coder , CodeLlama , and StarCoder ) using a variety of training methods designed specifically for the code domain to improve LLM code understanding.

In order to evaluate the coding abilities of LLMs, benchmarks like HumanEval and MBPP have been handcrafted to evaluate the program synthesis task of turning natural language descriptions (e.g., docstrings) into code snippets. These code benchmarks measure functional correctness by evaluating LLM-generated solutions against a set of limited predefined tests. Recent work has further improved these benchmarks with augmented tests to rigorously evaluate the functional correctness of LLM generated code. However, apart from test inadequacy, existing popular code synthesis benchmarks have the following limitations:

Limited amount and variety of problems. Code benchmarks are mainly constructed by human annotators manually. Due to the high manual effort required, they only contain a limited amount of problems. For example, HumanEval only contains 164 handcrafted problems. Such a low amount of problems is not sufficient to fully measure the complete spectrum of program synthesis capability of state-of-the-art LLMs. Additionally, these code benchmarks include mostly self-contained coding problems that lack variety in both problem types and domains, where the final evaluation output only shows the percentage of problems solved. While they provide a baseline overview of the coding abilities, LLM builders and users cannot gain deeper insights to exactly what problem types or coding scenarios the particular LLM may excel or struggle in.

Prone to data leakage and training dataset composition. Popular benchmarks like HumanEval and MBPP were released almost 4 years ago, with example solutions available in third-party open-source repositories. While recent LLMs have been taking turns climbing the leaderboard by achieving higher pass@11 scores (often with less than 1 percent difference between the next best model), just how much of that is attributed to having leaked solutions as part of the training data? Furthermore, the problems within these benchmarks are often simple derivatives of common coding problems/concepts. In fact, recent work has shown that there are substantial overlap between benchmark solutions and open-source training corpuses. In addition, closed-source LLMs may even deliberately include benchmark groundtruths to artificially boost their leaderboard status . As such, it is unclear whether high scores achieved by LLMs are truly due to their learnt coding capability or instead obtained via memorizing benchmark solutions.

As more LLMs are being constructed, trained, and used especially for code, the insufficient evaluation benchmarks raise the question of validity: Is leaderboard performance on existing benchmarks reliable and comprehensive enough to measure the program synthesis ability of LLMs?

Our work. To address the limitation of existing benchmarks, we introduce \scalerel*○ EvoEval coincidentally similar pronunciation with EvilEval – a set of program synthesis benchmarks created by evolving existing problems. The key idea behind EvoEval is to use LLMs instead of humans to produce new code synthesis problems based on a variety of different instructions aimed at evolving or transforming the existing benchmark problems into targeted domains for more comprehensive evaluation. Different from prior benchmark constructions that either obtain problems from open-source repositories or databases – leading to data leakage or require manual construction of each problem – resulting in high manual effort and limited diversity, EvoEval directly uses LLMs with targeted transformation prompts to synthesis new coding problems. Specifically, we design 5 different targeted transformation prompts: Difficult, Creative, Subtle, Combine and Tool Use. We then prompt GPT-4 to independently transform any existing problem in previous benchmarks into a new problem in the targeted domain.

Figure 1 shows a concrete example of EvoEval in action starting with an initial problem in HumanEval– vowel_counts to count the number of vowels in the string. 1 We first observe the transformation to a more difficult problem by asking GPT-4 to add additional constraints or requirements. This new problem contains a separate custom vowel list that makes the overall program logic more complex. 2 We can also transform to a more creative problem of create_alias that still uses concepts like vowels and consonants but involves a much more creative and unusual problem description. 3 We can also make subtle changes to the problem where we only count the lowercase vowels to test if the LLM is simply memorizing the benchmark. 4 We can additionally combine concepts from multiple problems together. In the example, we use another problem bf to create a new problem that returns the vowels in each planet sorted based on the orbiting order. 5 Furthermore, we can test the ability for LLMs to utilize auxiliary helper functions (common place in real-world code repositories) to solve more complex problems. Again we reuse the concepts of vowels from the initial problem, where the frequency of each vowel should be computed. However instead of directly solving the problem, the LLM can directly use the provided check_vowel helper function to simplify the solution.

Together, each of these transformed benchmarks are designed to introduce more difficult and complex problems as well as test different aspects of the LLM code understanding and synthesis ability. In EvoEval, we additionally use GPT-4 to generate the groundtruth solution to each problem as well as rigorous test cases to ensure we can evaluate the functional correctness of LLM-synthesized code on EvoEval. Finally, we manually check each generated problem and corresponding groundtruth to ensure problem clarity and correctness. EvoEval serves as a way to further evolve existing benchmarks into more complex and well-suited problems for evaluation in order to keep up with the ever-growing LLM research.

Contribution. Our work proposes to evolve existing problems for benchmark creation:

Benchmark: We present EvoEval– a set of program synthesis benchmarks created by evolving existing popular HumanEval coding benchmark problems. EvoEval includes 828 problems across 5 semantic-altering and 2 semantic-preserving benchmarks. Furthermore, EvoEval also includes additional benchmarks to study program synthesis concepts like problem composition and decomposition. EvoEval is fully complete with groundtruth implementations and robust testcases to evaluate functional correctness.

Approach: We propose a complete pipeline to directly synthesize new coding problems for benchmarking by evolving existing problems through the use of targeted transformation prompts. Our pipeline aims to reduce manual checking effort using a self-consistency approach to automatically refine any problem inconsistencies and generate groundtruth as well as test cases. Our approach is general and can be used on other benchmark problems, adopted for transformation into additional domains or utilize different problem generation strategies .

Study: We conduct a comprehensive study on 51 different LLMs across all benchmarks in EvoEval. We found that compared to the high performance obtained on standard benchmarks like HumanEval, when evaluated on EvoEval, popular LLMs significantly drop in performance (on average 39.4%). Additionally, this drop is not uniform across all LLMs and can range from 19.6% to 47.7%, leading to drastic ranking changes amongst top performing models. We further demonstrate that certain LLMs cannot keep up their high performance obtained in HumanEval when evaluated on more challenging or problems in different domains, highlighting the possibilities of overfitting to existing benchmarks. Moreover, we observe that while instruction-following LLMs perform well in solving self-contained problems, they struggle with the tool using aspect of utilizing already provided auxiliary functions. Furthermore, they are particularly sensitive to the problem description where rephrasing or subtle changes to the problem docstring leads to degradation in output solutions compared to their base non-instruction-following counterparts. Additionally, we demonstrate that current state-of-the-art LLMs fail to effectively compose multiple general coding concepts to solve more complex variants, or address subproblems decomposed from previously solved difficult problem.

Approach

Figure 2 shows the overview of the benchmark creation pipeline for EvoEval. We start by taking the original problem and apply a chosen targeted transformation prompt aimed at prompting GPT-4 to produce a new code synthesis problem along the targeted domain. Using this initial transformed problem, we enter our refinement pipeline to fix any ambiguities or inconsistencies in the problem description, as well as generating the test cases and groundtruth solution for functional evaluation. Finally, to ensure correctness, we manually examine each produced problem along with the groundtruth and make corresponding changes to produce the final evolved benchmarks.

Targeted problem transformation. EvoEval uses zero-shot prompting to evolve an existing coding benchmark to produce new and diverse problems. Each transformation prompt, as shown in the examples in Figure 1, aims to transform the existing problem in a specific manner. In particular, we define two different types of transformation prompts: 1) semantic-altering – change the semantic meaning of the original problem and 2) semantic-preserving – modify the problem description while keeping the semantic meaning the same. While Figure 1 shows only semantic-altering transformation prompts to produce new problems, we can also produce semantic-preserving problems to test additional aspect of the LLM coding abilities.

Problem refinement & groundtruth Generation. The initial evolved problem produced by GPT-4 may include small inconsistencies such as contradicting sentences or incorrect I/O examples in the docstring. For coding benchmarks, such inconsistencies are especially damaging as it can detract from the problem specification, leading to inaccurate evaluation of LLM coding capabilities. As such, we introduce a refinement pipeline to iteratively rephrase and refine problem as needed. In addition, during this process, we also use GPT-4 to produce the necessary groundtruth implementation of the function as well as example test cases to be used for evaluation.

We first directly use GPT-4 to obtain a possible solution for the initial problem. Additionally, we also prompt GPT-4 to extract (if available in the initial problem docstring) or produce the test inputs for the transformed problem. We then evaluate the test inputs on the solution to derive the corresponding expected test outputs. Next, using these test inputs/outputs, we instruct GPT-4 to add or fix the example test cases in the docstring, providing further demonstrations of the task.

Using this refined problem, we again generate a solution. We then leverage self-consistency to check if the new solution on the test inputs produce the same outputs as the previous solution. The intuition is that since both solutions are generated by GPT-4 and the refined problem should only include minimal changes (e.g., adding new testcase examples), the solution output should then be the same in the absence of any potential inconsistencies or ambiguity in problem description. As such, if we observe differences between the two solution outputs, we ask GPT-4 to further rephrase and fix any inconsistencies in the original problem and repeat the process. On the other hand, if both solutions agree on outputs, we terminate the problem refinement stage and return the trio comprising of the new problem description, the solution as the groundtruth and the test cases for functional evaluation.

Manual examination & test augmentation. For each transformed problem, we carefully examine and adjust any final faults to ensure each problem and groundtruth is correctly specified and implemented. Additionally, using the initial set of test cases from the refinement stage, we further generate additional tests following the LLM-based test augmentation technique in EvalPlus . Finally, we produce EvoEval, a comprehensive code synthesis benchmark suite, which through the use of evolving transformations can generate diverse coding problems to evaluate LLM coding capability across various problem domains.

EvoEval Dataset Overview

We use the problems in HumanEval as seeds to produce EvoEval. Problems in EvoEval consist mainly of self-contained functions, except for Tool_Use that includes helper functions specifically designed to test the tool using capability of LLMs. Each problem uses a docstring to illustrate the problem specification, along with test cases and groundtruth to evaluate the functional correctness. Table 1 shows the statistics of the benchmarks in EvoEval. In total, EvoEval includes 828 problems across 7 different datasets (5 semantic-altering and 2 semantic-preserving):

Difficult: Introduce complexity by adding additional constraints and requirements, replace commonly used requirements to less common ones, or add additional reasoning steps to the original problem.

Creative: Generate a more creative problem compared to the original through the use of stories or uncommon narratives.

Subtle: Make a subtle and minor change to the original problem such as inverting or replacing a requirement.

Combine: Combine two different problems by integrating the concepts from both problems. In order to select problems that make sense to combine, we apply a simple heuristic to combine only problems of the same type together categorized based on the type of input arguments in the original problem.

Tool_Use: Produce a new problem containing a main problem and one or more helpers functions which can be used to solve it. Each helper function is fully implemented and provides hints or useful functionality for solving the main problem. The main problem does not explicitly reference individual helper functions, and we do not require the model to use the provided helpers.

Verbose: Reword the original docstring to be more verbose. These verbose docstrings can use more descriptive language to illustrate the problem, include detailed explanation of the example output, and provide additional hints.

Concise: Reword the original docstring to be more concise by removing unnecessary details and using concise language. Furthermore, simple examples that are not required to demonstrate edge cases may be removed.

For each of the semantic-altering benchmarks, we generate 100 problems each using different seed problems from HumanEval. For semantic-preserving benchmarks, we generate using all 164 problems in HumanEval as it requires less validation since we can reuse the original groundtruths. As shown in Table 1, compared to HumanEval, EvoEval contains longer coding questions with longer average problem length. Furthermore, EvoEval also uses more test cases to perform robust evaluation compared to base HumanEval.

Figure 3 shows the embedding visualization using t-SNE perplexity=50 and iter=1000 using text-embedding-3-large model from OpenAI by projecting high-dimension representation of the problems docstrings in both EvoEval and HumanEval into the 2D plane. First, we see that Creative and Tool_Use drastically change the embedding distribution compared to the original dataset. The arrow in Figure 3(a) shows one example of the shift in distribution from the original problem to a creative one. Next, we see that Subtle, Difficult and Combine largely retain the same distribution as the original problems. This is due to the high parity across these problem descriptions where Subtle only applies subtle changes and Difficult adds additional complex constraints while keeping the main problem descriptions largely the same. Specifically, for Combine, we can see from an example arrow in Figure 3(b), the new combined problem shifts the embedding for both of the original problems. Finally, we observe that for Verbose and Concise, the embeddings almost perfectly match the original problem, reflecting their semantic-preserving nature. In Appendix C, we present example problems for each benchmark in EvoEval.

Methodology

Setup. Each LLM generated sample is executed against the test cases in EvoEval and evaluated using differential testing – comparing against the groundtruth results to measure functional correctness. We report the functional correctness by using the popular pass@kk metric. We focus on greedy decoding (i.e., producing a deterministic sample per each problem with temperature = 0). We denote this as pass@11.

Models. We evaluate 51 popular state-of-the-art LLMs, including both proprietary and open-source models on EvoEval. We evaluate not only the popular general-purpose LLMs but also include recent code-based LLMs for comprehensive evaluation. Further, we classify the LLMs as either base or instruction-following and focus our analysis on discussing the effect of model variants have on EvoEval performance.

Input format. To produce the code solution using each LLM, we provide a specific input prompt: For base LLMs (i.e., not instruction-tuned variants), we simply use only the function header with the docstring and let the LLM autocomplete the solution. For instruction-following LLMs, we follow the model-makers' guide on the exact instruction and format to use and ask the LLM to generate a complete solution for the problem.

Evaluation

EvoEval produces more complex and challenging benchmarks for program synthesis. Table 2 shows the pass@11 performance along with the ranking of LLMs on each of the semantic-altering EvoEval benchmarks with the average pass@11 and ranking on all benchmarks in the last columns. First, compared to the success rate on HumanEval, when evaluated on EvoEval, all LLMs consistently perform worse. For example, the state-of-the-art GPT-4, GPT-4-Turbo and Claude-3 models solve close to 85% of all HumanEval problems but fall almost below 50% pass@11 when evaluated on the Difficult problems. On average, across all benchmarks, the performance of LLMs decreased by 39.4% (Difficult: 58.7%, Creative: 50.2%, Subtle: 5.0%, Combine: 78.1%, and Tool_Use: 4.9%). Additionally, this drop is not uniform across all LLMs and can range from 19.6% to 47.7%.

LLMs struggles on EvoEval benchmarks compare to high performance achieved on HumanEval. One surprising finding is that, on Subtle, where only small changes are made to original problem with the roughly the same level of difficulty, the average performance of LLMs drops by 24.0% across the same 100 problems. It is important to note that, as the pass@11 score is generally higher on the first 100 problems than the complete 164 HumanEval problems, this back-to-back performance drop is much higher than the performance drop from HumanEval to Subtle mentioned above (which is 5.0%). Furthermore, we can also identify LLMs which struggle heavily on specific types of problems compared to their relative performance on HumanEval. Figure 4 shows scatter plot of HumanEval+ and EvoEval scores of selected LLMs. As we saw before, the significant portions of the models tends to be worse on EvoEval than HumanEval (i.e., purple shaded region). However, there exists LLMs that have a much higher HumanEval score compared to their performance on EvoEval (i.e., blue shaded region). This highlights potential data leakage of popular benchmarks where LLM performances are artificially inflated but do not translate to more difficult or other program synthesis problems.

Significant ranking changes of LLMs across different EvoEval benchmarks. In Figure 5, compared to the existing parity – where top models all perform similarly on HumanEval, we observe drastic differences in ranking changes on EvoEval. We observe that while the relative difference between the top 5 models on HumanEval is less than 10%, the difference on EvoEval on average is over 20%. Due to such saturation in top model performance, existing benchmarks may not reliably rank the program synthesis ability of each model. Taking a closer look at specific models, while Claude-3 and GPT-4 are tied for the 2nd best HumanEval score, they both excel at different types of problems: GPT-4 performs best on difficult and creative problems while Claude-3 can better reason about helper functions in Tool_Use and are less affected by subtle changes from original HumanEval. Furthermore, while GPT-4-Turbo achieves the top HumanEval and HumanEval+ score, it falls off compare to the base GPT-4 variant where it is worse on Difficult, Creative and Combine problems. Such evaluation cannot be gained through naively reporting existing coding benchmark performance. Overall, by evolving the original benchmark into more difficult and diverse problems of different types, EvoEval can provide a more holistic evaluation and ranking of the coding ability of LLMs.

EvoEval can be used to comprehensively compare multiple models. Figure 6 shows two radar graphs of two sets of LLMs. In Figure 6(a), while both WizardCoder-1.1 and Phind-CodeLlama-2 are top performing LLMs and have similar HumanEval scores, they perform drastically differently across the benchmarks in EvoEval. WizardCoder-1.1 is better on Difficult and Creative and Phind-CodeLlama-2 are better on Combine problems. This can be partially explained through the training dataset used in each LLM where WizardCoder-1.1 uses an evolving dataset to generate more complex and difficult problems whereas Phind-CodeLlama-2 is fine-tuned on high quality programming problems that seems to boost the ability to solve programs which combines multiple smaller programming concepts. Similar phenomenon can also be observed in Figure 6(b). Different from just reporting a singular pass@kk score, EvoEval also allows detailed analysis across the different dimension of coding capability to identify particular domains or type of synthesis questions the LLM struggles or excels in.

Instruction-following LLMs are sensitive to subtle or rephrasing of problem docstring. Unlike the semantic-altering benchmarks in EvoEval, the semantic-preserving problems do not always lead to a decrease in performance. Figure 7 shows the HumanEval score (bar) and the relative performance drop or improvement (arrows) on Verbose and Concise separated into instruction-following and base LLMs. We observe that almost all instruction-following LLMs drops in performance (on average 3.4% and 4.0% decrease on Verbose and Concise respectively) when evaluated on the two semantic-preserving dataset compared to the original HumanEval. This is drastically different from the non-instruction-following variants where we even observe performance improvements (on average 0.5% and 2.1% increase on Verbose and Concise respectively). Verbose and Concise do not change the semantic meaning of the original problem except reword it in either a more verbose or concise manner. Prior work has shown that by smartly rephrasing the original problem description, one can further boost LLM performance and we observe the similar phenomenon here mostly only for non-instruction-following models. This further points to possibility of overfitting to the exact descriptions utilized in HumanEval especially for instruction-tuned LLMs.

Additionally, even on the semantic-altering benchmark of Subtle, where only subtle changes to the original problem are applied, on average, instruction-following LLMs drops by 7.6% whereas base models only decreases by less than 1% relative to their HumanEval performance. These findings across LLM types show that while instruction-tuning is expected to align better with detailed task instructions, it fails to distinguish between these subtle changes in docstring, indicating potential memorization or contamination of prior evaluation benchmarks.

2 Problem Composition

Composition problems. The ability to compose different known concepts to solve new problems is known as compositional generalization . This skill is essential for code synthesis, especially for complex problems in real-world programs. However, measuring compositional generalization in LLM presents a fundamental challenge since it requires controlling the relationship between training and test distributions . While it is not easy to control the pre-training data of LLMs, we have more control in the testing phase. Hence, we focus on program concepts that have been demonstrated to fall within the capabilities of an LLM, and explore whether this proficiency extends to the combination of program concepts. As such, we start by taking a deeper look at the Combine problems evolved from combining previous HumanEval problems.

First half of Table 3 shows the detailed breakdown of the Combine dataset results on the top 8 performing LLMs. We observe that almost all problems solved in Combine came from the pass both category, which is intuitive as we do not expect LLMs to solve a problem composed of subproblems that it cannot already solve. However, we see that overall, the composition percentage is quite low as only GPT-4 is able achieve greater than half.This demonstrates, for the first time, that while state-of-the-art LLMs can achieve a high pass rate on simple programming tasks in general-purpose languages like Python, they still struggle with generalizing and composing these known concepts to address more complex problems.

Naive combination problems. Since Combine problems are not guaranteed to not contain additional new logic or concepts, we build a simplified dataset for sequential composition. Let AA and BB be two separate problems with xx as input(s) for AA, we aim to create a new problem CC with same inputs where the solution can be written as B(A(x))B(A(x)). To accomplish this, the new problem includes a sequential docstring by attaching the docstring of problem AA followed by BB. Directly concatenating them will lead to unclear descriptions, as such, for each problem in HumanEval, we manually create two separate variants based on which order the problem may come in the new docstring. Figure 8 shows an example naive combination problem with the manual sequential instruction highlighted in red. Using these modified problem docstrings, we build a sequential combination dataset – Combine-naive, containing 1074 problems by randomly combining problems filtering for input output matching (i.e., type of A(x)A(x) should equal to type of yy in B(y)B(y))

The latter half of Table 3 shows the results on Combine-naive following the same setup as Combine. We observe that while the composition percentage on the naive dataset improves significantly compared to the evolved Combine dataset, it still fails to reach near perfection, with the best LLM being able to only solve 3/4 of prior pass both problems. While existing training or inference paradigms for LLMs for code focus on obtaining high quality datasets boosted with instruction-tuning, our result shows that existing LLMs still struggle with the concept of problem composition to tackle more complex problems. We hope future research can design novel training methods to tackle this limitation.

3 Problem Decomposition

Given our analysis and benchmark on combining different problems together, a nature follow-up would be to look at problem decomposition – decomposing larger problems into multiple subproblems. We start by selecting 50 HumanEval problems and then follow our approach in Section 2 to decompose each original problem into two smaller subproblems, creating 100 problems in our Decompose benchmark.

Table 4 shows the results of selected LLMs on Decompose (the same set of LLMs as Combine). We first observe that similar to the composition percentage in the Combine and Combine-naive problems, LLMs do not achieve a high decomposition percentage. One possible interpretation is that current LLMs are trained to memorize or recover seen outputs in their training data, and when used for program synthesis, they cannot generalize the concepts from training data. This is demonstrated by not being able to solve smaller subproblems obtained from solved more difficult parent problems. On the other hand, we show that LLMs can sometimes solve both smaller subproblems even when the original parent problem is not solved (i.e., recomposition percentage). Decompose is akin to breaking the harder problem down into easier subproblems, which is related to planning in prior work . We hope future work can again build on these insights to achieve the best of both worlds in being able to succcesfully generalize difficult concepts into subproblems and adopting decomposing/planning to solve additional challenging problems.

4 Tool Using

We further analyze the Tool_Use dataset, which contains pre-defined helper or auxiliary functions in addition to the main synthesis problem. Additionally, we construct Tool_Use-Main_Only dataset, which contains the same set of problem as Tool_Use, except that the input to the LLM consists only of the main problem description without including any helpers. Using both datasets together, we can evaluate the ability of LLMs to use helper functions to solve more complex problem. We observe that compared to scenarios without any helper functions (average pass@11 of 28.6%), LLMs on average improve by 81.3% when provided with the helper functions. This is to be expected as the helper functions provides additional utilities in aiding to solve the more complex problem. However, this improvement is not uniform, as we see that the average improvement when given the auxiliary functions for instruction-following models is only 60.4% compared to the non-instruction-following LLMs' improvement of 122.0%.

Figure 9 show the detailed comparison between 10 instruction-following and their base LLMs on both the Tool_Use-Main_Only and Tool_Use dataset. We observe that without the helpers, the instruction-following models significantly outperform their base LLMs. However, once the helpers are provided, this gap is drastically decreased, with cases even where the base models outperform their instruction-following counterparts. As real-world coding involves understanding, using, and then reusing existing functions across different places in the repository, being able to successfully leverage auxiliary methods is key. Current instruction-following LLMs are generally fine-tuned with data consisting of self-contained code snippets without the interaction and learning of function usages. This is further exacerbated by prior benchmarks, which mostly use self-contained functions, thus cannot expose the insufficient tool-using capability of such models. In EvoEval, with Tool_Use and Tool_Use-Main_Only, we demonstrate this gap in evaluation and hope to inspire future research on this important aspect of code LLMs.

Related Work

Large language models for code. Starting with the general development of LLMs for general purpose tasks, developers have applied LLMs to perform code-related tasks by further training LLMs using collected code snippets from open-source repositories. Such LLMs include Codex , PolyCoder , CodeT5 , CodeGen , InCoder , CodeLlama , StarCoder , StarCoder2 , DeepSeeker , etc. These LLMs can autoregressive complete code given the relevant prefix (e.g., docstrings for function completion). More recently, following the advancement in NLP, researchers have applied instruction-tuning methods to train code-specific LLMs that are well-versed in following instructions. Examples of such LLMs include CodeLlama-Inst and DeepSeeker-Inst . WizardCoder instruction-tunes the model using Evol-Instruct to create more complex instructions. Magicoder develops OSS-Instruct by synthesizing high quality instruction data from open-source code snippets. OpenCodeInterpreter additionally leverages execution feedback for instruction-tuning in order to better support multi-turn code generation and refinement.

Program synthesis benchmarking. HumanEval and MBPP are two of the most widely-used handcrafted code generation benchmarks complete with test cases to check for the correctness of LLM outputs. Building on these popular benchmarks, additional variants have been crafted including: HumanEval+ which improves the two benchmarks with more complete testcases; HumanEval-X which extends HumanEval to C++, Javascript and Go; MultiPL-E which further extends both HumanEval and MBPP to 18 coding languages. Similarly, other benchmarks have been developed for specific domains: DS-1000 and Arcade for data science APIs; ODEX for open-domain code generation covering a diverse range of libraries; CodeContests , APPS and LiveCodeBench for programming contests; ClassEval for class-level generations, and SWE-Bench for real-world software engineering tasks. Different from prior benchmarks which require handcraft problems from scratch – high manual effort or scrape open-source repositories or coding contest websites – leading to unavoidable data leakage, EvoEval directly uses LLMs to evolve existing benchmark problems to create new complex evaluation problems. Furthermore, contrasting with the narrow scope of prior benchmarks (often focusing on a single type or problem, i.e., coding contests), EvoEval utilizes targeted transformation to evolve problems into different domains, allowing for a more holistic evaluation of program synthesis using LLMs.

Conclusion

We present EvoEval– a set of program synthesis benchmarks created by evolving existing problems into different target domains. We build on top of the popular HumanEval benchmark to produce 828 problems across 7 different benchmarks for a holistic and comprehensive evaluation of LLM program synthesis ability. Our results on 51 LLMs show that compare to high performance on standard benchmarks, there is drastic drop in performance (on average 39.4%) when evaluated on EvoEval. Additionally, we observe significant ranking differences compared to previous leaderboards, indicating potential overfitting of popular LLMs on existing benchmarks. Throughout the paper, we provide additional insights, including the brittleness of instruction-following LLMs as well as problem composition and decomposition abilities. We hope EvoEval not only provides a valuable benchmarking suite for program synthesis but also inspires future code LLM builders to recognize the shown limitations of existing code LLMs and develop novel and targeted training approaches for code. We have open-sourced the EvoEval benchmarks, tools, and complete LLM generations available at https://github.com/evo-eval/evoeval

Acknowledgment

We thank Owen Colegrove for his help on starting this project and providing valuable feedback throughout, Jiawei Liu for providing helpful discussions and Yifeng Ding for his help in running experiments.

References

Appendix A Evaluation LLMs

Table 5 shows the overview of the 51 LLMs we evaluated in our work. For any LLMs which provide their open-source weights, we directly obtain them from huggingface model hub For certain LLMs, we may use the vLLM inference library for more efficient generation. For any close-sourced LLMs, we directly access their model endpoints using their providers. For more detail on the access of each LLM, please check our repository: https://github.com/evo-eval/evoeval

A.2 Detailed Evaluation Setup

LLM generation. As mentioned in Section 4, we report the pass@11 score for each LLM on our dataset generated using greedy decoding (i.e., sampling with temperature = 0). For each LLM, we provide a specific input prompt depending on the model type. For base LLMs (i.e., not instruction-following variants), we use only the function headers as input. For instruction-following, we make the best effort to follow examples provided by each model maker on the exact instruction and format to use at the time of writing. Specifically, for instruction-following LLMs, we ask the model to return the code snippet wrapped by code blocks (i.e., ```). Figure 10 shows an example input for GPT-4 on a Creative problem.

Furthermore, we also provide a custom sanitization script adopted from EvalPlus which parses the raw LLM outputs for code block parsing (e.g., removing ``` indicators for instruction-following models) and end-of-string identifiers (e.g., removing tokens like ). Each model generated output is passed into the sanitization script and the evaluation occurs on the sanitized outputs.

Oracle. To evaluate the functional correctness of each LLM synthesized solution, we use differential testing by comparing the model output with the groundtruth output on a set of testcase inputs. We build our evaluation framework on top of the EvalPlus evaluation script used for HumanEval and HumanEval+ benchmark which evaluates multiple problems and solutions in parallel for efficiency. For each testcase, we perform exact matching or check if the output is within an absolute difference threshold of 10−610^{-6} if the output is a floating point type. We additionally implement our evaluation script by recursively checking the type and performing the appropriate comparison (e.g., dictionary outputs are first length checked for equivalence and then matching is done for each value and key). Furthermore, we also implement custom oracles for specific problems where there could be multiple solutions or simple tolerance or exact matching cannot fully guarantee correctness. We refer the reader again to our repository https://github.com/evo-eval/evoeval which contains the full implementation of each of our custom oracle. Additionally, we also use timeout as another evaluation method. Our setting again follows EvalPlus default setup where the timeout per problem is defined as T=max(Tmax,f×tgt)T=max(T_{max},f\times t_{gt}) with default values of Tmax=1000msT_{max}=1000ms, f=4f=4 and tgtt_{gt} defined as the measured groundtruth solution time to produce the correct output. All timeout related factors can be adjusted to account for variance on different underlying machine and hardware.

Appendix B Transformation Prompts

Here we provide the exact targeted transformation prompts used to evolve existing benchmark problems into each of our transformation prompts. Figure 11, 12, 13, 14, 15, 16, 17 and 18 shows the prompt for Difficult, Creative, Subtle, Combine, Tool_Use, Verbose, Concise and Decompose respectively.

Figure 19 and 20 show the refinement and I/O extraction/fixing prompt used in EvoEval. The refinement prompt is used to refine the origin generated problem when inconsistency is detected (see Section 2). The extraction prompt is used to initially obtain a set of testcases from the problem docstring used for self-consistency evaluation. We further use an I/O fixing prompt (also in Figure 20) to fix any examples in the docstring which do not contain the right output (as computed by the groundtruth generated by GPT-4).

Appendix C Example Problems in EvoEval

Here we demonstrate a few example problems across the benchmarks in EvoEval and corresponding GPT-4 solution which cannot solve the problem. Figure 21, 22, 23, 24, 25, 26 and 27 show such example for the EvoEval Difficult, Creative, Subtle, Combine, Tool_Use, Verbose and Concise respectively.