OctoPack: Instruction Tuning Code Large Language Models
Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, Shayne Longpre
Introduction
Finetuning large language models (LLMs) on a variety of language tasks explained via instructions (instruction tuning) has been shown to improve model usability and general performance (Wei et al., 2022; Sanh et al., 2022; Min et al., 2022; Ouyang et al., 2022). The instruction tuning paradigm has also proven successful for models trained on visual (Liu et al., 2023a; Li et al., 2023a), audio (Zhang et al., 2023b) and multilingual (Muennighoff et al., 2022b; Wang et al., 2022b) data.
In this work, we instruction tune LLMs on the coding modality. While Code LLMs can already be indirectly instructed to generate desired code using code comments, this procedure is brittle and does not work when the desired output is natural language, such as explaining code. Explicit instructing tuning of Code LLMs may improve their steerability and enable their application to more tasks. Concurrently to our work, three instruction tuned Code LLMs have been proposed: PanGu-Coder2 (Shen et al., 2023), WizardCoder (Luo et al., 2023) and InstructCodeT5+ (Wang et al., 2023c). These models rely on more capable and closed models from the OpenAI APIhttps://openai.com/blog/openai-api to create their instruction training data. This approach is problematic as (1) closed-source APIs keep changing and have unpredictable availability (Pozzobon et al., 2023; Chen et al., 2023a), (2) it relies on the assumption that a more capable model exists (3) it can reinforce model hallucination (Gudibande et al., 2023) and (4), depending on legal interpretation, OpenAI’s terms of usehttps://openai.com/policies/terms-of-use forbid such models: “…You may not…use output from the Services to develop models that compete with OpenAI…”. Thus, we consider models trained on OpenAI outputs not usable for commercial purposes in practice and classify them as non-permissive in this work.
We focus on more permissively licensed data and avoid using a closed-source model to generate synthetic data. We benchmark four popular sources of code instruction data: (1) xP3x (Muennighoff et al., 2022b), which contains data from common code benchmarks, (2) Self-Instruct (Wang et al., 2023a) data we create using a permissive Code LLM, (3) OASST (Köpf et al., 2023), which contains mostly natural language data and few code examples and (4) CommitPack, our new 4TB dataset of Git commits. Instruction tuning’s primary purpose is to expand models’ generalization abilities to a wide variety of tasks and settings. Thus, we extend the code synthesis benchmark, HumanEval (Chen et al., 2021; Zheng et al., 2023), to create HumanEvalPack: A code benchmark covering code synthesis, code repair, and code explanation across six programming languages.
Instruction tuning StarCoder (Li et al., 2023b) on a filtered variant of CommitPack and OASST leads to our best model, OctoCoder, which surpasses all other openly licensed models (Figure 1), but falls short of the much larger GPT-4 (OpenAI, 2023). GPT-4 is close to maximum performance on the code synthesis variant, notably with a pass@1 score of 86.6% on Python HumanEval. However, it performs significantly worse on the code fixing and explanation variants of HumanEvalPack, which we introduce. This suggests that the original HumanEval benchmark may soon cease to be useful due to models reaching close to the maximum performance. Our more challenging evaluation variants provide room for future LLMs to improve on the performance of the current state-of-the-art.
CommitPack and CommitPackFT: 4TB of permissively licensed code commits across 350 programming languages for pretraining and a filtered variant containing high-quality code instructions for finetuning
HumanEvalPack: A benchmark for Code LLM generalization, spanning three scenarios (Code Repair, Code Explanation, Code Synthesis) and 6 programming languages (Python, JavaScript, Java, Go, C++, Rust)
OctoCoder and OctoGeeX: The best permissive Code LLMs
CommitPack: Code Instruction Data
Prior work has shown that models can generalize to languages included in pretraining, but absent during instruction tuning (Muennighoff et al., 2022b). However, they also show that including such languages during instruction tuning boosts their performance further. We hypothesize that code data exhibits the same behavior. To improve performance on code-related tasks, we thus construct a code instruction dataset leveraging the natural structure of Git commits.
To construct the dataset, we use commit metadata from the GitHub action dump on Google BigQuery.https://www.gharchive.org/ We apply several quality filters, filter for commercially-friendly licenses, and discard all commits that affect more than a single file to ensure commit messages are very specific and to avoid additional complexity from dealing with multiple files. We use the filtered metadata to scrape the affected code files prior to and after the commit from GitHub. This leads to close to 4 terabytes of data covering 350 programming languages (CommitPack). As instruction tuning does not necessarily require so much data (Zhou et al., 2023a; Touvron et al., 2023), we apply several strict filters to reduce the dataset to 2 gigabytes (CommitPackFT). These strict filters include filtering for samples where the commit message has specific words in uppercase imperative form at the start (e.g. "Verify …"), consists of multiple words and does not contain external references. All filters are detailed in Appendix D. Figure 2 depicts the distribution of both datasets and the tasks contained in CommitPackFT. For instruction tuning our models, we select 5,000 random samples from CommitPackFT across the 6 programming languages that we evaluate on.
We consider three additional datasets for instruction tuning presented in Table 1. xP3x: xP3x is a large-scale collection of multilingual instruction data with around 532 million samples (Muennighoff et al., 2022b). We focus only on the code subset of xP3x, excluding NeuralCodeSearch (Li et al., 2019) which is not licensed permissively, and select 5,000 samples. Self-Instruct: Using the Self-Instruct method (Wang et al., 2022a) and the StarCoder model (Li et al., 2023b), we create 5,003 synthetic instructions and corresponding answers. OASST: OASST is a diverse dataset of multi-turn chat dialogues (Köpf et al., 2023). While most dialogues center around natural language, some also contain code. We reuse a filtered variant of OASST from prior work (Dettmers et al., 2023) and additionally filter out moralizing assistant answers (Appendix D) leading to 8,587 samples.
HumanEvalPack: Evaluating Instruction Tuned Code Models
When instruction tuning LLMs using natural language (NL) data, the input is an NL instruction with optional NL context and the target output is the NL answer to the task (Wei et al., 2022). When instruction tuning with code (C) data, code may either appear only in the input alongside the NL instruction (NL+CNL, e.g. code explanation), only in the output (NLC, e.g. code synthesis), or in both input and output (NL+CC, e.g. code modifications like bug fixing). While prior benchmarks commonly only cover variants of code synthesis, users may want to use models in all three scenarios. Thus, we expand the code synthesis benchmark HumanEval (Chen et al., 2021; Zheng et al., 2023) to cover all three input-output combinations for six languages (Figure 3).
Given an incorrect code function with a subtle bug and accompanying unit tests, the model is tasked to fix the function. We manually add a bug to each of the 164 HumanEval solutions across all 6 languages (984 total bugs). For a given sample, the bugs are as similar as possible across the 6 languages enabling meaningful comparison of scores across languages. Bugs are written such that the code still runs but produces an incorrect result leading to at least one unit test failing. Bug statistics and examples are in LABEL:sec:bugs. We also evaluate an easier variant of this task where instead of unit tests, models are provided with the correct function docstring as the source of truth to fix bugs, see LABEL:sec:docstrings.
Given a correct code function, the model is tasked to generate an explanation of the code. Subsequently, the same model is tasked to regenerate the code given only its own explanation. The second step allows us to score this task via code execution and measure pass@k (Chen et al., 2021) instead of evaluating the explanation itself using heuristic-based metrics like BLEU (Papineni et al., 2002) or ROUGE (Lin, 2004) which have major limitations (Reiter, 2018; Schluter, 2017; Eghbali & Pradel, 2022; Zhou et al., 2023b). To prevent models from copying the solution into the description, we remove any solution overlap of at least 20 characters from the description. We further enforce a character length limit on the model-generated explanation equivalent to the length of the docstring describing the function. This limit is specified in the prompt for the model. Note that the function docstring itself is never provided to the model for this task.
Given a natural language docstring or comment describing the desired code, the model is tasked to synthesize the correct code. This task corresponds to the original HumanEval benchmark (Chen et al., 2021). For instruction tuned models, we add an explicit instruction to the input explaining what the model should do. For models that have only gone through language model pretraining, we follow Chen et al. (2021) and provide the model with the function header and docstring to evaluate its completion of the function.
For all tasks we execute the code generations to compute performance using the pass@ metric (Chen et al., 2021): a problem is considered solved if any of code generations passes every test case. We focus on the simplest version of pass@, which is pass@1: the likelihood that the model solves a problem in a single attempt. Like Chen et al. (2021), we use a sampling temperature of and to estimate pass@1. We generate samples, which is enough to get reliable pass@1 estimates (Li et al., 2023b). For GPT-4, we generate samples. Using instead of for GPT-4 only changes scores by around 2% while providing 20x cost savings.
Python HumanEval is the most commonly used code benchmark, thus many training datasets have already been decontaminated for HumanEval to enable fair evaluation. By reusing HumanEval and manually expanding it to more scenarios and languages, we ensure that existing decontamination remains valid. This enables a fair comparison across a large variety of models.
OctoCoder: Best Commercially Licensed Code LLM
We instruction tune the pretrained StarCoder model (Li et al., 2023b) on different combinations of our instruction datasets (§2). We evaluate all models on the Python subset of HumanEvalPack as depicted in Figure 4. Similar to prior work (Taori et al., 2023), we format all instructions into a consistent schema to distinguish question and answer (see Appendix K).
CommitPackFT is critical for the performance boost on code repair (HumanEvalFix), where instruction tuning on only OASST or other variants results in a significantly lower score. This is likely due to CommitPackFT including around 20% of bug fixes among other code-related tasks (Figure 2).
The pretrained StarCoder model, as well as the Self-Instruct variant, perform poorly on code explanation (HumanEvalExplain). This is because both models are only conditioned to write code instead of natural language. We find that to perform well at explaining code, it is necessary to include samples with natural language as the target output during instruction tuning. Only relying on data with code as the target, such as the Self-Instruct data, will lead to models always outputting code even if the question requires a natural language output. Thus, we mix all other ablations with OASST, which contains many natural language targets. While the xP3x subset also contains samples with natural language output, many of its target outputs are short, which leads to models with a bias for short answers. This is impractical for the explanation task leading to the comparatively low score of mixing xP3x with OASST.
All instruction datasets provide similar boosts for code synthesis (HumanEvalSynthesize), which has been the focus of all prior work on code instruction models (Wang et al., 2023c; Luo et al., 2023; Muennighoff et al., 2022b). We achieve the best average score by instruction tuning on CommitPackFT mixed with our filtered OASST data yielding an absolute 23% improvement over StarCoder. Thus, we select CommitPackFT+OASST for our final model dubbed OctoCoder. Using the same data, we also instruction tune the 6 billion parameter CodeGeeX2 (Zheng et al., 2023) to create OctoGeeX.
2 Comparing with other Models
We benchmark OctoCoder and OctoGeeX with state-of-the-art Code LLMs on HumanEvalPack in Table 2. For all models, we use the prompt put forward by the model creators if applicable or else a simple intuitive prompt, see Appendix K.
OctoCoder has the highest average score across all three evaluation scenarios among all permissive models. With just 6 billion parameters, OctoGeeX is the smallest model benchmarked, but still outperforms all prior permissive Code LLMs. GPT-4 (OpenAI, 2023) performs best among all models benchmarked with a significant margin. However, GPT-4 is closed-source and likely much larger than all other models evaluated.
Trained primarily on natural language, not code, BLOOMZ (Muennighoff et al., 2022b) performs worse than other models despite having 176 billion parameters. Go and Rust are not contained in BLOOMZ’s instruction data, yet it performs much better than the random baseline of 0.0 for these two languages across most tasks. This confirms our hypothesis that models are capable of generalizing instructions to programming languages only seen at pretraining, similar to crosslingual generalization for natural languages (Muennighoff et al., 2022b). To improve programming language generalization further, we tune OctoCoder and OctoGeeX on many languages from CommitPackFT, and this generalization improvement is reflected in the performance on HumanEvalPack’s new languages.
Prior work has shown that the performance on natural languages after instruction tuning is correlated with the weight of these languages during pretraining (Muennighoff et al., 2022b). The more weight during pretraining, the better the performance after instruction tuning. We find the same to be the case for programming languages. Python, Java, and JavaScript collectively make up around 30% of the pretraining data of StarCoder (Li et al., 2023b). After instruction tuning StarCoder to produce OctoCoder, we see the best performance among these three languages, especially for HumanEvalSynthesize. OctoCoder performs weakest on Rust, which is the lowest resource language of StarCoder among the languages we benchmark (1.2% of pretraining data).
HumanEvalFix is the most challenging task for most models. They commonly regenerate the buggy function without making any change (e.g. WizardCoder in Figure 31) or they introduce new bugs (e.g. GPT-4 in Figure 30). We analyze model performance by bug type in Appendix I and find bugs that require removing excess code are the most challenging. OctoCoder performs comparatively well across all languages. Instruction tuning on CommitPackFT has likely taught OctoCoder to make small, targeted changes to fix bugs.
Some models fail at HumanEvalExplain, as they do not generate natural language explanations. We manually inspect explanations for the first ten samples of the Python split and disqualify a model if none of them are explanations. This is the case for StarCoder and CodeGeeX2, which generate code instead of natural language explanations. BLOOMZ and InstructCodeT5+ also occasionally generate code. Other models exclusively generate natural language explanations, not containing any code for inspected samples.
HumanEvalExplain instructs models to fit their explanation within a given character limit (§3). Current models appear to have no understanding of how many characters they are generating. They commonly write very short and thus underspecified explanations (e.g. BLOOMZ in Figure 32) or excessively long explanations that end up being cut off (e.g. StarChat- in Figure 35). Future work could investigate how to enable models to be aware of their generated output length to improve HumanEvalExplain performance.
Pure code synthesis on HumanEvalSynthesize is the easiest task for all models. With a pass rate of 86.6% for a single solution, GPT-4 is close to fully saturating the Python subset. GPT-4 was originally found to score 67% on Python HumanEval (OpenAI, 2023) and 81% in later work (Bubeck et al., 2023). Our score for GPT-4 is significantly higher, possibly due to improvements made to the API by OpenAI, contamination of HumanEval in GPT-4 training, or slightly different prompting and evaluation. An example of our prompt is depicted in Figure 3 (right). We perform very careful evaluation to ensure every generation is correctly processed. We reproduce the HumanEval score of WizardCoder (Luo et al., 2023; Xu et al., 2023a) and find it to also perform well across other languages. For BLOOMZ and InstructCodeT5+ our evaluation leads to a higher Python score than they reported, likely because of our more careful processing of generations. OctoCoder has the highest performance for every language among permissively licensed models. With a pass@1 of 46.2% on the original Python split, OctoCoder improves by a relative 38% over its base model, StarCoder.
Related Work
There has been extensive work on code models tailored to a specific coding task, such as code summarization (Iyer et al., 2016; Ahmad et al., 2020; Zhang et al., 2022a; Shi et al., 2022) or code editing (Drain et al., 2021; Zhang et al., 2022c; He et al., 2022; Zhang et al., 2022b; Wei et al., 2023; Prenner & Robbes, 2023; Fakhoury et al., 2023; Skreta et al., 2023) (also see work on edit models more generally (Reid & Neubig, 2022; Schick et al., 2022; Dwivedi-Yu et al., 2022; Raheja et al., 2023)). These works use task-specific heuristics that limit the applicability of their methods to other tasks. In contrast, we aim to build models applicable to all kinds of tasks related to code and beyond.
Through large-scale pretraining more generally applicable code models have been developed (Nijkamp et al., 2022; 2023; Xu et al., 2022a; Christopoulou et al., 2022; Gunasekar et al., 2023; Li et al., 2023b; Bui et al., 2023; Scao et al., 2022a; b). However, these models only continue code making them hard to use for tasks such as explaining code with natural language (HumanEvalExplain). Teaching them to follow human instructions is critical to make them applicable to diverse tasks.
2 Instruction Models
Training models to follow instructions has led to new capabilities in text (Ouyang et al., 2022; Wang et al., 2022b; Chung et al., 2022) and visual modalities (Xu et al., 2023b; OpenAI, 2023). Prior work has shown its benefits for traditional language tasks (Sanh et al., 2022; Wei et al., 2022; Longpre et al., 2023a; Iyer et al., 2022), multilingual tasks (Muennighoff et al., 2022b; Yong et al., 2022), and helpfulness in dialog (Köpf et al., 2023; Bai et al., 2022; Ganguli et al., 2022). For coding applications, PanGu-Coder2 (Shen et al., 2023), WizardCoder (Luo et al., 2023) and InstructCodeT5+ (Wang et al., 2023c) are recent models trained with coding instructions. However, they all use the CodeAlpaca dataset (Chaudhary, 2023), which is synthetically generated from OpenAI models. Using data from powerful closed-source models provides a strong advantage, but limits the model use and has other limitations highlighted in §1. CoEditor (Wei et al., 2023) proposes an “auto-editing” task, trained on 1650 python commit history repositories. Our work expands this proposal to more general coding tasks (using instructions), more languages, and orders of magnitude more commit data.
3 Code Benchmarks
Many code synthesis benchmarks have been proposed (Wang et al., 2022d; c; Yu et al., 2023; Lai et al., 2023; Du et al., 2023). HumanEval (Chen et al., 2021; Liu et al., 2023b) has emerged as the standard for this task. Prior work has extended HumanEval to new programming languages via automatic translation mechanisms (Athiwaratkun et al., 2022; Cassano et al., 2023; Orlanski et al., 2023). These approaches are error-prone and only translate tests, not the actual solutions, which are needed for tasks like code explanation. Thus, we rely only on humans to create all parts of HumanEvalPack including test cases, correct solutions, buggy solutions, and other metadata across 6 languages.
Code repair is commonly evaluated on Quixbugs (Lin et al., 2017; Prenner & Robbes, 2021; Ye et al., 2021; Xia & Zhang, 2023; Jiang et al., 2023; Sobania et al., 2023) or Python bugs (He et al., 2022; Bradley et al., 2023). The latter does not support code execution, which limits its utility. While Quixbugs supports execution with unit tests, it only contains 40 samples in Python and Java. Further, the problems in Quixbugs are generic functions, such as bucket sort. This makes them easy to solve and hard to decontaminate training data for. Our benchmark, HumanEvalFix, contains 164 buggy functions for six languages with solutions and unit tests. Further, our coding problems, derived from HumanEval, are very specific, such as keeping track of a bank account balance (see Figure 12).
Prior work on evaluating code explanations (Lu et al., 2021; Cui et al., 2022) has relied on metrics such as METEOR (Banerjee & Lavie, 2005) or BLEU (Papineni et al., 2002). By chaining code explanation with code synthesis, we can evaluate this task using the execution-based pass@k metric overcoming the major limitations of BLEU and other heuristics-based metrics (Reiter, 2018).
Large-scale benchmarking has proven useful in many areas of natural language processing (Wang et al., 2019; Kiela et al., 2021; Srivastava et al., 2022; Muennighoff et al., 2022a). By producing 18 scores (6 languages across 3 tasks) for 9 models, we take a step towards large-scale benchmarking of code models. However, we lack many models capable of generating code (Black et al., 2021; Fried et al., 2022; Black et al., 2022; Wang & Komatsuzaki, 2021; Biderman et al., 2023b). Future work may consider more models or extending HumanEvalPack to new languages or tasks, such as code efficiency (Madaan et al., 2023a; Yetistiren et al., 2022) or code classification (Khan et al., 2023).
Conclusion
This work studies training and evaluation of Code LLMs that follow instructions. We introduce CommitPack, a 4TB dataset of Git commits covering 350 programming languages. We filter this large-scale dataset to create CommitPackFT, 2GB of high-quality code with commit messages that assimilate instructions. To enable a comprehensive evaluation of instruction code models, we construct HumanEvalPack, a human-written benchmark covering 3 different tasks for 6 programming languages. We ablate several instruction datasets and find that CommitPackFT combined with natural language data leads to the best performance. While our models, OctoCoder and OctoGeeX, are the best permissively licensed Code LLMs available, they are outperformed by closed-source models such as GPT-4. In addition to improving the instruction tuning paradigm, future work should consider training more capable base models.
Acknowledgements
We thank Hugging Face for providing compute instances. We are extremely grateful to Rodrigo Garcia for the Rust translations, Dimitry Ageev and Calum Bird for help with GPT-4 evaluation, Loubna Ben Allal for help on evaluation, Arjun Guha for insightful discussions on chaining evaluation tasks to avoid evaluating with BLEU, Lewis Tunstall for help on the OASST data, Victor Sanh and Nadav Timor for discussions, Jiaxi Yang for logo editing and domain classification prompting design, Ghosal et al. (2023); Zeng et al. (2023) for design inspiration, Harm de Vries for feedback and all members of BigCode for general support. Finally, we thank every programmer who takes the time to write informative commit messages.
References
Appendix A Contributions
Niklas Muennighoff created CommitPack and HumanEvalPack, wrote most of the paper and lead the project. Qian Liu devised many quality filters, ran SantaCoder ablations, investigated early training decisions and helped edit the paper. Armel Zebaze created the Self-Instruct data and ran numerous ablations. Niklas Muennighoff, Armel Zebaze and Qinkai Zheng created and evaluated OctoCoder and OctoGeeX. Binyuan Hui pretrained SantaCoder, made major contributions to the presentation and helped edit the paper. Terry Yue Zhuo ran GPT-4 evaluations and helped edit the paper. Xiangru Tang provided help on several experiments for evaluation and helped edit the paper. Leandro von Werra provided early guidance, suggested many quality filters and added the commit data to StarCoder pretraining. Niklas Muennighoff, Qian Liu, Binyuan Hui, Swayam Singh and Shayne Longpre conducted the data analysis. Shayne Longpre advised the project and made large contributions to the paper.
Appendix B Artifacts
Appendix C CommitPack and CommitPackFT Languages
Appendix D Dataset Creation
We use the GitHub archive available on GCP which contains metadata from GitHub commits up to 2016.https://www.gharchive.org/ It contains around 3TB of GitHub activity data for more than 2.8 million GitHub repositories including more than 145 million unique commits, over 2 billion different file paths and the contents of the latest revision for 163 million files.https://github.blog/2016-06-29-making-open-source-data-more-available/ We apply the filters in Table 5 to this dataset. The resulting dataset containing only metadata is uploaded at https://hf.co/datasets/bigcode/commitpackmeta. As the activity dataset only contains commit ids without the actual code changes, we scrape the code from GitHub. We use the metadata and the GitHub API to scrape the changed file prior and after the respective commit. Some repositories referenced in the activity data are no longer accessible, thus we discard them. This results in CommitPack with approximately 4 terabytes uploaded at https://hf.co/datasets/bigcode/commitpack.
Prior work has shown the importance of careful data filtering to maintain quality (Yin et al., 2018; Dhole et al., 2021; Laurençon et al., 2022; Longpre et al., 2023b). To create a smaller version focused on commits that resemble high-quality instructions, we further filter CommitPack to create CommitPackFT using the steps outlined in Table 6. We also checked for any contamination with HumanEval (Chen et al., 2021) but did not find any solution or docstring present in CommitPackFT. This is likely because our commit data only goes up to 2016, which is several years prior to the release of HumanEval. Our filters reduce the dataset by a factor of around 1000 resulting in close to 2 gigabytes uploaded at https://hf.co/datasets/bigcode/commitpackft. To gain a deeper understanding of the rich content within CommitPackFT, we analyze commits on its Python subset (56K samples). We first collect the most prevalent commit domain by prompting GPT-4 with: "I’d like to know the main types of commits on Github and aim to cover as comprehensively as possible.". Subsequently, we use GPT-4 to classify each sample using the prompt in Figure 5. The task distribution is visualized in Figure 2.
We use a subset of xP3x (Muennighoff et al., 2022b) focusing on code datasets consisting of APPS (Hendrycks et al., 2021), CodeContests (Li et al., 2022b), Jupyter Code Pairs,https://hf.co/datasets/codeparrot/github-jupyter-text-code-pairs MBPP (Austin et al., 2021), XLCoST (Zhu et al., 2022), Code Complex (Jeon et al., 2022), Docstring Corpus (Barone & Sennrich, 2017), Great Code (Hellendoorn et al., 2019) and State Changes.https://hf.co/datasets/Fraser/python-state-changes
We reuse a filtered variant of OASST (Köpf et al., 2023) from prior work (Dettmers et al., 2023) and apply additional filters to remove responses that refuse to comply with the user request. To compute the programming languages and code fraction for OASST depicted in Table 1, we count all responses containing e.g. ‘‘‘python or ‘‘‘py for the Python programming language. There are code samples that are not enclosed in backticks or do not specify the language, thus we are likely underestimating the actual fraction of code data for OASST in Table 1.
Appendix E Comparing Data Before and After Filtering
In Table 8 we compare word statistics prior to and after filtering CommitPack to create CommitPackFT. The mean commit subject and message length increases suggesting that messages are more informative in CommitPackFT. The code lengths decrease significantly as we limit the number of allowed tokens in the filters in Table 6. Notably, the percentage of code changed between pre- and post-commit is (a 31% increase) as opposed to (a 0.7% increase). Thus, the filtered data carries significantly more signal per token with fewer repetitions of the code prior to the commit.
Appendix F Comparing CommitPack and The Stack
In Table 9 we provide statistics on repositories and usernames of CommitPack and The Stack (Kocetkov et al., 2022). CommitPack contains a total of 1,934,255 repositories. Around half (49.3%) of them are also in The Stack. However, The Stack only provides the raw code files of these repositories from some fixed point in time. CommitPack contains the changes made to the code files in the form of commits. Thus, the same code file may appear multiple times in CommitPack for each change that was made to it. Therefore, The Stack only contains 3 terabytes of data, while CommitPack contains close to 4.
Appendix G Pretraining on CommitPack
Due to the scale of CommitPack, it is also adequate as a large-scale pretraining dataset. We have included parts of CommitPack during the pretraining of StarCoder (Li et al., 2023b) in the format of
In Table 10, we benchmark StarCoder and SantaCoderPack on HumanEvalFix using the above-detailed commit format. We find that the commit format leads to very strong performance for StarCoder often surpassing the instruction tuned OctoCoder from Table 2. However, this pretraining format is not suitable for HumanEvalExplain limiting its universality. For SantaCoderPack, we find performance comparable to SantaCoder, including checkpoints at 131B and 236B tokens. SantaCoderPack performs slightly worse on Python than SantaCoder. We hypothesize that this discrepancy is due to a multilingual tax, as SantaCoderPack needs to accommodate three additional coding languages (Go, C++ and Rust). SantaCoder has thus more capacity allocated to Python, JavaScript, and Java.
SantaCoderPack may also be bottlenecked by its small model size of 1.1B parameters. More research into what exactly happens during pretraining (Xia et al., 2022; Biderman et al., 2023a) and how to unify pretraining and instruction tuning are needed. Prior work has also found that including raw code data during pretraining benefits some natural language tasks (Muennighoff et al., 2023). Future work may consider the effects of including code commit data on natural language tasks.
Appendix H Line Diff Format for Fixing Code
We finetune SantaCoder to experiment with different formatting strategies for fixing bugs comparing full code generation and code diff generation. When fixing a code bug, usually only a small part of the code needs to change. Only generating the code diff corresponding to the necessary change can make inference significantly more efficient by avoiding repeated characters in the output generation. We finetune SantaCoder on the Python, Java and JavaScript subset of CommitPackFT. We exclude other languages as SantaCoder has only been pretrained on these three languages (Allal et al., 2023).
For full code generation, we reuse the format that we employed for commits in StarCoder pretraining from Appendix G:
For code diff generation, a simple solution is using the unified diff format,https://en.wikipedia.org/wiki/Diff#Unified_format which is a standard way to display changes between code files in a compact and readable format (Lehman et al., 2022; Jung, 2021; Xu et al., 2022b; Monperrus et al., 2021). We depict an example of this format in Figure 6. However, the unified diff format still requires the model to output several unchanged lines below and after the actual modification. Thus, its efficiency gains are limited and there is still unnecessary duplication of the input.
To address the inefficiencies of the unified diff format, we propose the line diff format for representing code differences. There are two requirements for our format: (1) The diff can be unambiguously applied to the code before the commit to generate the code after the commit, and (2) the code diff should be as short as possible to maximize efficiency by avoiding the inclusion of unchanged code. In Figure 7, we show how our format addresses these. The line diff format keeps track of each change sequentially line-by-line to ensure the code can be correctly modified. By focusing only on the lines that change, we reduce the number of characters in the diff by 70% compared to the unified diff representation in Figure 6.
Both the unified diff format and our line diff format require the model to predict line numbers. This is very challenging when training on raw code as models need to count and keep track of line numbers. To simplify line number prediction, we automatically add line numbers to the raw code in the finetuning dataset for the line diff format. This allows the model to simply copy the line number into the output simplifying the diff generation. However, it diminishes efficiency slightly by adding additional input tokens that the model needs to process.
As summarized in Table 11, finetuning SantaCoder using the line diff format significantly improves performance compared to prior finetuning on HumanEvalFix across all languages. It also outperforms finetuning using the commit format, which only provides gains on JavaScript and Java compared to no finetuning. However, finetuning on the diff format may converge slower than the commit format as the diff format significantly differs from the raw code seen during pretraining. Figures 8, 9, LABEL:fig:pyexample_linediff show line diff generations of our model. A limitation of our current line diff implementation is that it does not handle code insertion well. The inserted lines may change the line numbers of all following lines, which can result in problems when applying the diff. Further, the diff format is not useful for HumanEvalExplain and HumanEvalSynthesize. Future work could consider training models that can both be instructed to use the line diff format, such as for HumanEvalFix, but also explain or synthesize code without producing a diff.
Missing logic bug example. The buggy code (right) misses the ’abs’ statement.
Appendix I Performance Breakdown by HumanEvalFix Bug Type
All bugs in HumanEvalFix are categorized into bug types as described in LABEL:sec:bugs. In Table 12, we break down the HumanEvalFix performance of select models from Table 2 by bug type. We find that models struggle most with bugs that require removing excess logic (e.g. Figure 10). WizardCoder is only able to solve 11% of excess logic bugs while solving about four times more bugs that relate to value misuse. The performance of OctoGeeX and OctoCoder is more stable than WizardCoder across the different bug types, possibly due to the diversity of CommitPackFT as displayed in Figure 2. GPT-4 performs best across all bug types.
Appendix J Hyperparameters
For all experiments finetuning StarCoder, we use a learning rate of 5e-4 with a cosine schedule and linear warmup. We use a batch size of 32 and train for up to one epoch, as we did not observe benefits from more steps. OctoCoder was trained for 35 steps with a sequence length of 2048 and packing corresponding to 2.2 million total finetuning tokens.
To create OctoGeeX, we finetune CodeGeeX2 for 35 steps with a batch size of 48 and a learning rate of 5e-5 largely following the OctoCoder setup.
For all experiments finetuning SantaCoder, we use a learning rate of 5e-5 with a cosine schedule and linear warmup. We finetune SantaCoder using a batch size of 64 for up to 200,000 steps.
We follow the setup from Allal et al. (2023) to pretrain on CommitPack with the exception of using a sequence length of 8192 and using the tokenizer from StarCoder, which has special tokens for the commit format delimiters (see Appendix G). SantaCoderPack utilizes Multi Query Attention (MQA) (Shazeer, 2019) but removes Fill-in-the-Middle (FIM) (Bavarian et al., 2022). We conducted pretraining on 32 A100 GPUs, totaling 250k training steps, with a global batch size of 64. Other hyperparameter settings follow SantaCoder, including using Adam with , and a weight decay of 0.1. The learning rate is set to and follows a cosine decay after warming up for 2% of the training steps.
Appendix K Prompts
The prompting format can significantly impact performance. In the spirit of true few-shot learning (Perez et al., 2021) we do not optimize prompts and go with the format provided by the respective model authors or the most intuitive format if none is provided. For each task, we define an instruction, an optional context and an optional function start (Appendix K). The function start is provided to make sure the model directly completes the function without having to search for the function in the model output. These three parts are then combined in slightly different ways for each model (Figures K-K). We implement our evaluation using open-source frameworks (Ben Allal et al., 2022; Gao et al., 2021).
### Instruction: {instruction} {context} ### Response: {function_start}
### Instruction: {instruction} {context} ### Response:{function_start}
<|end|> <|user|> {instruction} {context}<|end|> <|assistant|> {function_start}
{instruction} Start your code with: {func_start}
L.2 GPT-4
However, if the function is supposed to calculate something specific related to a car race collision and it’s not doing that correctly, we would need more information about the expected behavior to fix it.
L.3 WizardCoder
A promising avenue for improving performance on HumanEvalFix is letting the model execute the given code or its own generated code and inspect its output (Chen et al., 2022; 2023c; Yasunaga & Liang, 2021; Li et al., 2022a; Gao et al., 2023; Dong et al., 2023; Zhang et al., 2023c; Madaan et al., 2023b; Ni et al., 2023; Gou et al., 2023; Hu et al., 2023; Taylor et al., 2022; Nye et al., 2021). This could allow the model to discover which unit tests are failing and for what reason. The model could then simply iterate on the function until all unit tests are passing. We leave explorations of this strategy to improve performance on HumanEvalPack to future work.
For the creation of CommitPack, we have filtered out any commits that affect multiple files to ensure commits are very specific and account for the fact that most current models are only capable of operating on a single file. Allowing models to take multiple files as input and modify multiple files given a single instruction is a promising direction for future work. There is active research on using repository-level context (Ding et al., 2022; Shrivastava et al., 2023a; b; Zhang et al., 2023a; Liu et al., 2023d) and the necessary long context windows (Dai et al., 2019; Press et al., 2021; Sun et al., 2021; Dao et al., 2022; Peng et al., 2023; Liu et al., 2023c; Chen et al., 2023b).
Current Code LLMs including OctoCoder struggle with awareness about the length of their generated output. For HumanEvalExplain, we instruct the models to limit their output to a given number of characters. While it is trivial for humans to count characters and adhere to the limit, all models tested frequently generate far too many characters. Prior work has shown that human raters are biased towards preferring longer texts (Wu & Aji, 2023) regardless of content. All models evaluated are instruction tuned on text that was at least indirectly assessed by human raters, hence they may be biased towards generating longer texts even if it means including literary bloat.
Evaluating code instruction models is challenging for several reasons: (1) Prompting: The prompt can significantly impact the performance of large language models (Brown et al., 2020; Zhou et al., 2022; Muennighoff, 2022; Babe et al., 2023). To ensure fair evaluation we use the prompting format put forth by the respective authors of the models and a simple intuitive prompt for models without a canonical prompt (see Appendix K). However, this may put models without a canonical prompt recommendation (e.g. BLOOMZ, GPT-4) at a slight disadvantage. OctoCoder and OctoGeeX perform best when prompted using the same format we use during training (Appendix K) and we recommend always using this format at inference. (2) Processing: Models may accidentally impair otherwise correct code by e.g. including a natural language explanation in their output. We largely circumvent this issue through the use of strict stopping criteria and careful postprocessing (e.g. for GPT-4 we check if it has enclosed the code in backticks, and if so, extract only the inner part of the backticks discarding its explanations). (3) Execution: When executing code to compute pass@k, it is important that the generated code matches the installed programming language version. Models may inadvertently use expressions from a different version (e.g. they may use the Python 2 syntax of print "hi", which would fail in a Python 3 environment). In our evaluation, we did not find this to be a problem, however, as models become more capable, it may make sense to specify the version. Future prompts may include the version (e.g. “use JDK 1.18.0”) or provide models with an execution environment that has the exact version installed that will be used for evaluation. (4) Comprehensiveness: Executing code can only reflect functional correctness lacking a comprehensive understanding of quality. Compared to execution-based evaluation, the human judgment of code quality can be considered more comprehensive as humans can consider factors beyond correctness. Directly hiring human annotators can be inefficient and expensive, and therefore researchers have explored approaches to automate human-aligned evaluation via LLMs (Fu et al., 2023; Liu et al., 2023e; Zhuo, 2023). However, recent work (Wang et al., 2023b) suggests LLM-based evaluation can be biased towards certain contexts. Future work on automating the human-aligned evaluation of instruction tuned Code LLMs while avoiding such bias is needed.
Our commit datasets, CommitPack and CommitPackFT, also lend themselves well for learning human preferences. The changed code after a commit generally represents a human-preferred version of the code (else the code would not have been modified). Thus, one could train a reward model that given the code before and after a commit, learns that the code afterward is better. Similar to prior work (Ouyang et al., 2022), this reward model could then be used to guide a language model to generate code that is preferred by humans.