Is ChatGPT the Ultimate Programming Assistant -- How far is it?
Haoye Tian, Weiqi Lu, Tsz On Li, Xunzhu Tang, Shing-Chi Cheung, Jacques Klein, Tegawendé F. Bissyandé
Introduction
Artificial intelligence (AI) has shown great potential in contributing to automate the execution of several challenging software engineering tasks, such as code generation (Svyatkovskiy et al., 2020; Li et al., 2022a; Siddiq et al., 2022; Bavishi et al., 2019), program repair (Gupta et al., 2017; Ye et al., 2022; Li et al., 2022b; Jiang et al., 2023b), and code summarization/explanation (Hu et al., 2018; Li et al., 2020; Wang et al., 2020; Stapleton et al., 2020). Indeed, recent state of the art approaches in code generation leverage AI models to automatically produce full or partial programs based on natural language descriptions or some code inputs. In the program repair realm, neural networks based model are becoming de facto standard for learning to generate fixes, thus reducing manual debugging efforts and enhancing software reliability and security. To address code summarization, researchers have employed generative AI models that output natural language descriptions given an input code, towards contributing to facilitating code comprehension, code maintainability, and code reuse.
The recent emergence of large-scale language models (LLMs) has received much attention in the society in general. In software engineering, it has brought AI-driven automation to new heights (Devlin et al., 2018; Chen et al., 2021; Zhang et al., 2020; Tian et al., 2020). LLMs, which are pre-trained on vast amounts of source code and natural language data, have demonstrated exceptional capabilities in understanding code structures and generating code or texts. Advances led by LLMs have further enhanced the effectiveness of automatic techniques for various software engineering challenges, such as code generation (Nijkamp et al., 2022; Fried et al., 2022; Tunstall et al., 2022; Dakhel et al., 2022; Chakraborty et al., 2022), program repair (Zhang et al., 2022a; Jain et al., 2022; Jiang et al., 2023a; Dakhel et al., 2022), and code summarization (Ahmed and Devanbu, 2022; Hu et al., 2020). A notable example of LLM is OpenAI’s Codex (Chen et al., 2021) (the model behind Github Copilot (Github, [n.d.])). Codex has been successfully applied to a range of software engineering tasks, demonstrating its ability to generate accurate and efficient code snippets (Yetistiren et al., 2022; Vaithilingam et al., 2022), identify and fix bugs or vulnerabilities in software (Pearce et al., 2021; Prenner and Robbes, 2021), and provide summarization for code segments (Sarsa et al., 2022; MacNeil et al., 2023). Literature experimental results showcase the enormous potential of LLMs in the software development community although the overall achieved performance remains relatively limited (Chen et al., 2022b; Fan et al., 2022).
ChatGPT (OpenAI, 2022a), a recently introduced LLM, has generated significant interest within the software engineering community due to its reported capabilities in the longstanding dream of software engineering: to repair software automatically with minimal human intervention (Sobania et al., 2023; Xia and Zhang, 2023; Surameery and Shakor, 2023). Reported results suggest that ChatGPT will be transformative for the field, and that LLM-driven software engineering has a bright future. We posit that further investigations are still needed to clearly scope the capabilities of LLMs. Indeed, existing program repair studies with ChatGPT tend to assess its performance using old publicly-available benchmark data (anterior to 2022), which may have been leaked into the training corpus of ChatGPT. This experimental bias potentially jeopardizes the generalizability of reported results to new and unseen problems. Furthermore, ChatGPT’s code generation and summarization/explanation capabilities have not been thoroughly investigated in the literature.
To fill these gaps, we propose to carry out an extensive empirical study to evaluate ChatGPT’s realistic potential as a fully automated programming assistant. Our study focuses on the key challenging tasks of code generation, program repair, and code summarization, as these tasks are not only repetitive but also fundamental to programmers’ daily work (Kazemitabaar et al., 2023; Becker et al., 2023; Xu et al., 2022; Le Goues et al., 2012; Yang et al., 2019; Rodeghero et al., 2014; Cordy, 2003). Our experimental setup targets common programming problems which are recurrently encountered by programmers (Capilla et al., 2019). We first evaluate the natural language-to-code (code generation) ability of ChatGPT on two LeetCode programming datasets. Second, we investigate ChatGPT as a program repair tool to repair a large set of diverse incorrect codes in a programming assignments benchmark. Since assignments are usually common coding problems and they are segments of code, the findings made on these assignments are likely to be generalizable (Hu et al., 2019; Le Goues et al., 2015). Finally, we explore whether ChatGPT can accurately explain the intention of the proposed correct and incorrect code samples in the student assignment benchmark. We will elaborate on how we mitigate the issue of data leakage in the experiments.
We conduct an empirical study of ChatGPT on the tasks of code generation, program repair, and code summarization on two benchmarks.
We investigate the realistic performance of ChatGPT on common programming problems, comparing it against state-of-the-art approaches.
We make interesting findings of ChatGPT performance and analyze their implications in Table 1. The findings contribute actionable insights to the adoption of LLMs for programming assistance and offer a better understanding of ChatGPT’s practical applications in the software engineering community.
The remainder of this paper is organised as follows: Section 2 presents background details as well as related work. The study design is presented in Section 3. Experimental results are described in Section 4 and threats to validity are enumerated in Section 5. We conclude the paper in Section 6.
Background and Related Work
Large Language Models (LLMs) are a class of deep learning models that have significantly impacted natural language processing (NLP) in recent years with a wide range of applications in various domains (Chen et al., 2021; Gu et al., 2021; Alayrac et al., 2022; Beltagy et al., 2019). Encoder model BERT (Bidirectional Encoder Representations from Transformers) (Devlin et al., 2018) is a widely-used LLM, built on a transformer-based architecture that uses a Masked Language Modeling (MLM) pre-training task, where it masks some of the input tokens and tries to predict them from the remaining tokens bidirectionally. The MLM is represented as:
where is the model’s prediction score for the masked word. This enables BERT to learn context-aware word representations and capture the relationships between words in a sentence. Therefore, Bert, for its ability on code/text understanding, has been applied as an embedding model (encoder) for downstream software engineering tasks (Mashhadi and Hemmati, 2021; Ciborowska and Damevski, 2022; Chen et al., 2022a; Tian et al., 2020). Decoder model GPT (Generative Pre-trained Transformer) (OpenAI, 2022a; Ouyang et al., 2022; Chen et al., 2021; Radford et al., 2019, 2018), on the other hand, was trained using the unsupervised learning technique of self-supervised pre-training. It uses a unidirectional, causal attention mechanism that allows each position in the sequence to attend only to preceding positions, preserving the input order. During pre-training, the model aims to maximize the likelihood of predicting a word based on its contextual information:
where is the likelihood, is the word at position , is the sequence up to , and are the model parameters. GPT-1 (Radford et al., 2018) attempts to predict the next word in a sequence given the previous words. This task is performed unidirectionally from left to right, allowing GPT-1 to learn the dependency relations between words and generate relevant and context-aware content. GPT-2 (Radford et al., 2019) is a larger model than GPT-1, with 1.5 billion parameters 10 times compared to GPT-1’s 117 million parameters. GPT-2 was trained with larger architecture on a larger and more diverse dataset than GPT-1, which includes web pages, books and scientific articles. GPT-2 demonstrated that training on a larger dataset and utilizing a greater number of parameters can enhance the capabilities of a language model even in zero-shot settings. Compared to GPT-1 and GPT-2, GPT-3 (Brown et al., 2020) represents a significant advancement in LLM development. It has a much larger 175 billion parameters. Codex (Chen et al., 2021) is initially created using GPT-3 weights derived from training on a natural language corpus, and then it is fine-tuned on an extensive dataset of code files. Afterwards, the InstructGPT (Ouyang et al., 2022) was secondly created to be fine-tuned on GPT-3 by using reinforcement learning from human feedback (RLHF).
ChatGPT (OpenAI, 2022a) is the most popular ChatBot (GPT-3.5 and GPT-4) developed by OpenAI. It is a sibling model of InstructGPT while they are trained with different training datasets because of the different purposes. The training dataset of InstructGPT mainly consists of instructional text such as explanations, and examples. Specifically, ChatGPT focuses on conversation, thus ChatGPT is trained with massive texts and codes, including a wide range of user queries, and responses which are coherently relevant to these queries. These training data make ChatGPT applicable to chatbots or virtual assistants. Furthermore, it also employs advanced supervised instruction fine-tuning techniques and RLHF to adapt more effectively to specific tasks or domains. A recent study (Sobania et al., 2023) shows that ChatGPT obtains significant improvement over other LLMs. Although OpenAI has not released official information regarding the technical details of ChatGPT, its enhanced performance is likely attributable to its training techniques and large corpus. Since ChatGPT is the most advanced and possibly the most powerful LLM, we conducted an extensive experiment to investigate ChatGPT’s capability in performing various software engineering tasks (§2.2).
2. Code-related Tasks
Code generation is the process of automatically creating entire program code or completing code snippets from a higher-level representation, such as a natural language description, a model or specification, to improve programming efficiency and reduce human error (Budinsky et al., 1996; Svyatkovskiy et al., 2020; Parvez et al., 2021). Recent advances in machine learning and natural language processing have led to the development of new techniques for code generation. Specifically, Liu et al. (Liu et al., 2020) proposed a self-attention neural architecture for code completion by utilizing hierarchical structural information, capturing long-term dependency, and employing a multi-task learning framework for knowledge sharing on real-world datasets. Wei et al. (Wei et al., 2019) presented a dual training framework that simultaneously trains code summarization and code generation tasks by exploiting their duality, resulting in improvements for both tasks on GitHub datasets. With the development of pre-trained LLMs, researchers attempt to use them for code generation.
Ahmad et al. (Ahmad et al., 2021) proposed PLBART, a sequence-to-sequence model, pre-training on a large dataset and effectively learning syntax, style, and logical flow. InCoder (Fried et al., 2022) is a unified generative model capable of program synthesis and editing through infilling, which leverages bidirectional context to improve performance. Nijkamp et al. (Nijkamp et al., 2022) introduced CodeGen, a family of large language models up to 16.1B parameters trained on natural language and programming data, demonstrating competitive performance on zero-shot Python code generation. More recently, Yetistiren et al. (Yetistiren et al., 2022) evaluated the quality of code generated by GitHub Copilot, producing a 91.5% success rate in generating valid code. However, the results suggest more comprehensive assessment is needed. In this paper, we explore LLM in generating code functions through the programming problem descriptions in natural language.
2.2. Program repair.
Program repair is designed to automatically identify and repair bugs or vulnerabilities in software code (Le Goues et al., 2021; Motwani et al., 2020; Shariffdeen et al., 2021; Zhang et al., 2022b). Specifically, Gulwani et al. (Gulwani et al., 2018) presented an automated program repair algorithm for introductory programming assignments by using correct student solutions to repair incorrect attempts. Hu et al. (Hu et al., 2019) proposed a search-based fully automated approach for generating real-time student program repairs by refactoring correct solutions and matching incorrect programs to refactored ones. The method requires fewer correct solutions and achieves better results with just a single correct reference solution. On the other hand, AI-based systems have recently achieved state-of-the-art performance in this task. Chakraborty et al. (Chakraborty et al., 2020) proposed a novel tree-based neural network system, CODIT, to learn and suggest code change patterns from large-scale open-source data. The evaluation demonstrates its effectiveness in suggesting patches and fixing bugs on the Defects4J dataset. Ye et al. (Ye et al., 2022) proposed a program repair model called RewardRepair which uses program-specific information during neural network weight optimization to produce patches that compile and are not prone to be overfitting.
With respect to the LLM techniques, Prenner et al. (Prenner et al., 2022) investigated pre-trained large language model Codex’s ability to localize and fix bugs in code, finding that it is surprisingly effective and competitive with state-of-the-art automated program repair techniques, especially in repairing Python bugs. Sobania et al. (Sobania et al., 2023) evaluated ChatGPT’s repair performance on the QuixBugs benchmark set. By providing hints through its conversational system, ChatGPT is able to repair 31 out of 40 bugs and outperform state-of-the-art approaches. Meanwhile, Xia et al. (Xia and Zhang, 2023) demonstrated the improved performance of conversational APR over prior LLM-based APR methods. In our study, we aim to investigate the automatic repair of diverse incorrect implementations submitted for common programming assignments.
2.3. Code summarization.
Code summarization refers to the automatic process of providing clear and concise summarizations or explanations for a piece of code or functionality (Hu et al., 2018; Parvez et al., 2021; Ahmad et al., 2020; Shi et al., 2022). Allamanis et al. (Allamanis et al., 2016) leveraged a simple Transformer model with self-attention mechanisms for source code summarization that significantly outperforms state-of-the-art techniques. Some researchers proposed to summarize code by leveraging the abstract syntax trees (AST) of code. For instance, Hu et al. (Hu et al., 2018) proposed a novel approach that uses a structure-based traversal (SBT) method to flatten the AST into a sequence to automatically generate code comments for Java methods. Shido et al. (Shido et al., 2019) proposed Multi-way Tree-LSTM, an extension of Tree-LSTM, to handle source code summarization by addressing the challenges of applying machine translation models to source code’s ASTs with nodes having arbitrary numbers of children and their order. Wan et al. (Wan et al., 2018) also incorporated an AST structure and sequential content into a deep reinforcement learning framework, using an actor-critic network to optimize word prediction and provide global guidance for exploration. LeClair et al. (LeClair et al., 2020) presented a graph-based neural architecture that combines both the source code sequence and structural information from ASTs to improve automatic source code summarization. In our paper, we focus on the summarization and explanation of the intention of the source code.
Study Design
In this section, we first present three research questions that we aim to investigate the ChatGPT’s ability in this study. Then, we introduce two benchmarks and evaluation metrics, describing how we elaborate them for the research questions. Finally, we present the application of ChatGPT.
RQ-1: To what extent can ChatGPT generate correct and efficient code for common programming problems? We propose to assess ChatGPT’s text-to-code generation ability by measuring the correctness and time complexity of its code generated for common programming problems on two datasets that are built from LeetCode 2016-2020 and LeetCode 2022 respectively.
RQ-2: Can ChatGPT repair diverse buggy code implemented for common programming problems effectively? To answer this research question, we leverage a benchmark of algorithmic programming assignments that contains 1783 diverse incorrect code submissions. We investigate whether ChatGPT can fix the incorrect code with or without the problem descriptions provided.
RQ-3: Can ChatGPT identify the intention of code? We propose to investigate the ability of ChatGPT to explain code. Specifically, we evaluate whether ChatGPT can infer the original intention of code, even in cases where the code contains errors.
2. Benchmark selection criteria
Python is a widely used high-level programming language for coding applications. In addition, Python is the programming language that descendant models of OpenAI GPT-3 (e.g. ChatGPT and Codex) are most capable in (OpenAI, 2022b, c). To investigate the ability of ChatGPT in code-related tasks, we thus adopt Python as the programming language for the benchmarks in our study. Furthermore, our study scope requires benchmarks that have a diversity of programming problems as well as implementations.
3. Datasets
We base our study on two benchmarks. The first benchmark is built on LeetCode (LeetCode, 2023), an online platform offering a diverse range of programming problems, such as algorithms, data structures, etc. LeetCode categorizes programming problems based on the types of problems (e.g. array, sorting, etc.) and difficulty level of problems (easy, medium, and hard). LeetCode also provides adequate test cases (avg. 135 per problem) to verify the correctness of submitted code and time-complexity-based rank percentile to evaluate the efficiency of the code. These informative features motivate us to select it as our first benchmark. In order to ensure the generalizability of our findings, we focus on the four most common (number) typesWe do not select Math (272) and Dynamic programming (267) as they are exclusively appearing in either the easy problem or hard problem respectively. of programming problems in the entire LeetCode repository: Array (862), String (398), Hash table (283), Sorting (203). We randomly sample 10 problems from each of the 4 problem types for the 3 levels of easy, medium, and hard. Finally, we obtain 120 (10*4*3) sampled programming problems from LeetCode in 2016-2020 and LeetCode in 2022 respectively. We describe how we leverage them at the end of this section.
The second benchmark is Refactory (Hu et al., 2019), a large Python repairing bug benchmark released in 2019. It contains 2442 correct and 1783 buggy programs submitted by undergraduate students at the National University of Singapore for five algorithmic common programming assignments. The various and diverse correct and incorrect Python code contribute to our study. Besides, Refactory also provides instructor-designed test cases (avg. 9 per problem), correct reference solutions and descriptionswe supplement with detailed descriptions provided in (Dakhel et al., 2022) for each assignment. The test suite is designed to provide students with evaluation results as feedback. We take this benchmark into our account as it contains plenty of implementations and well-designed test cases for common programming problems (e.g., searching/sorting/deduplication algorithms). Table 2 shows a summary of the benchmark.
We note that ChatGPT has been trained on a diverse and extensive dataset of text, including sources such as web pages, books, articles, and online forums. This actually presents a potential data leakage problem for the relevant research. Because when selecting experimental benchmarks on which ChatGPT will be evaluated, we are uncertain whether ChatGPT has already learned the solution references provided in these benchmarks in advance. In this case, the achieved performance and insights about ChatGPT may not be able to be generalized to new or unseen problems. To evaluate ChatGPT’s true capability and mitigate the possible cold start problemThe model cannot draw any inferences for items about which it has not yet gathered sufficient information. of ChatGPT, we elaborate on how we utilize the two benchmarks mentioned above for each research question. We do not aim to construct entirely new programming problems for ChatGPT; expecting so would be unrealistic. Similar information exists widely online, and this existing knowledge is essential for training language models like ChatGPT to effectively predict and address future issues.
RQ-1: Code generation. Refactory only contains five programming problems and all of them can be accomplished by ChatGPT and Codex (). Therefore, we mainly use LeetCode programming problems to evaluate the code generation ability of ChatGPT. Note that the knowledge cutoff of ChatGPT is September 2021. We first evaluate ChatGPT on 120 (10*4*3) programming problems from LeetCode in 2016-2020 that ChatGPT may have learned knowledge from. Then, to investigate whether ChatGPT can generalize its ability to new problems, we assess ChatGPT on another 120 (10*4*3) programming problems from LeetCode in 2022. We skip LeetCode in 2021 as it’s mixed of potentially trained problems and new problems for ChatGPT and we wish to evaluate ChatGPT separately on two kinds of datasets.
RQ-2: Program repair. To evaluate ChatGPT’s ability to program repair, we utilize the 1783 buggy programs of five programming problems in Refactory benchmark. The selection is based on: (1) The repair solutions of buggy programs are not publicly available, and therefore, ChatGPT is more likely not to have been trained on them in advance. (2) The buggy assignment code are diverse and implemented for common coding problems (searching/sorting algorithms), and they are segments of code, the findings made on these assignments are likely to be generalizable (Hu et al., 2019; Le Goues et al., 2015). Although ChatGPT can access the five teacher-written correct reference solutions, it is not an issue for such an assignment repair tool (ChatGPT’s role in this research question) that is allowed to learn from the correct reference solutions.
RQ-3: Code summarization. We assign ChatGPT the task of generating code summarization for the intention of various correct and incorrect codes. Therefore, we seek 2442 correct and 1783 buggy codes in Refactory. There are no natural language explanations provided for them on the Internet. We can be relatively safe to request ChatGPT to produce the natural language intention of code.
4. Application of ChatGPT
ChatGPT was a newly-released LLM-based conversational product by OpenAI in November 2022. There has not been a consensus on the standard protocol for the usage of ChatGPT. In this section, we describe how we apply ChatGPT in our experiments.
ChatGPT API. OpenAI has introduced the ChatGPT API for several of its developed models, including GPT-3 and GPT-4. Notably, the GPT-4 model—the latest in the series—comes with aggressive request rate limits https://platform.openai.com/docs/guides/rate-limits/overview, posing challenges for large-scale experimentation. Concurrently, a surge in recent reports and studies suggests that GPT-4 is exhibiting signs of instability and inaccuracy (Zakori, 2023; Madaan, 2023; Barr, 2023; Chen et al., 2023), which may stem from its radical redesign. In contrast, the GPT-3.5 family produce reliable results and has been widely used. Given this context, we have chosen the GPT-3.5 model for our investigation. Specifically, our research employs the gpt-3.5-turbo-0613 model with its default parameter configurations.
Randomness. Given the inherent randomness of GPT models, it is important to note that ChatGPT has the possibility to generate different responses for the same prompt input at different requests. In particular, when it comes to code generation and program repair, ChatGPT may attempt different data structures or algorithms to achieve these tasks. To deal with randomness and ensure reliable analysis and conclusion, there is a need to request multiple times to ChatGPT for the same prompt. We follow the prior works and run five independent conversational requests with ChatGPT for each prompt. We display the performance in two measurements: (1) TOP-5, the value is 1 if at least one of the five attempts of an approach solves the specific programming problem, otherwise 0. This metric indicates the robustness of an approach. Specifically, it assesses the presence or absence of a viable solution within the top five attempts, which, according to a survey (Noller et al., 2022), is the maximum number of generated code that most developers are willing to review. (2) AVG-5, the average success rate of five attempts of an approach to the programming problem. A score of 1 for AVG-5 means that all five attempts are successful. Given the inconsistency/randomness of the TOP-1 performance of LLMs, we therefore propose the AVG-5 to represent the average TOP-1 of an approach. A high AVG-5 score indicates that the approach consistently yields correct solutions, making it a reliable tool for programmers in need of precise outcomes.
Prompt. The prompt is a critical component in utilizing ChatGPT and other LLMs, as it provides the context and direction for the generation of texts or codes. We introduce the baseline prompt we apply for the research experiments. First, Figure 1 shows an example of a prompt that we request ChatGPT for the code generation task in RQ-1. The prompt is mainly composed of the programming task description given by LeetCode. The description comprises a programming problem introduction (lines 1: natural language description, lines 3-9: examples, lines 11-12: constraints) and a defined function to be completed (lines 14-17). To form the final prompt for Python code generation, we append a fixed request statement, ”Implement the above task in Python.” at the end. For the task of program repair (RQ-2), OpenAI set up a prompt template that we follow up on for our experiments, as shown in Figure 2. For the task of code summarization (RQ-3), we recall that our expectation is to obtain the explanation of ChatGPT about the semantic intention of code without describing what each line of code is doing. Therefore, we construct our prompt as to be the fixed request statement: “Can you explain the intention of the below function(s) within one sentence? Do not include any explanations of code details in your answer.” followed by the associated code.
5. Evaluation Metrics
Code correctness. To assess the correctness of generated code in RQ-1 and RQ-2, we rely on test cases that the benchmark (LeetCode) provides. Any code that passes all test cases is considered correct, while any code that fails any test case is considered incorrect. We then use TOP-5 and AVG-5 metrics as overall performance measures, as detailed in the last section (item (2) Randomness).
Code time complexity. For the RQ-1 task of code generation, we also pay attention to the efficiency of the code. We rely on the official time-complexity-based rank percentile of code submitted to LeetCode as a key metric to evaluate the performance of different code generation techniques and benchmark them against human-written code. More specifically, we compute the average time-complexity-based rank percentile. If the average rank is, e.g., 10%, this means that the generated code is among the top 10% of the most efficient solutions.
Repair rate. To evaluate the performance of automated program repair tools in RQ-2, we leverage the standard metric of repair rate, i.e. the ratio of the patched code generated by approaches that pass all instructor-designed test cases. We use TOP-5 and AVG-5 metrics as performance measures, as detailed in the last section (item (2) Randomness).
Experiments & Results
In this section, we present the experiments performed to answer our research questions. In Sections 4.1 and 4.2, we investigate the performance of ChatGPT in generating code and repairing buggy programs respectively. In Section 4.3, we examine whether ChatGPT can accurately explain the intention of code.
[Experiment Goal]: We aim to evaluate ChatGPT’s text-to-code generation ability for common programming problems by assessing its performance on two Leetcode datasets.
[Experiment Design]: As outlined in Section 3.3, we utilize the LeetCode datasets from 2016-2020 and 2022 to assess ChatGPT’s performance. The programming problem introduction consists of three parts: problem description, examples, and constraints. The problem description is a necessary block. We include examples because there may exist some edge cases without clear specifications explained in the descriptionhttps://leetcode.com/problems/determine-if-two-events-have-conflict/. The constraints specify the value range and format of the input data, which may not be missing in the descriptionhttps://leetcode.com/problems/minimum-total-cost-to-make-arrays-unequal/. We include these components in our prompt to ensure the completeness of requirement specifications. Regarding the metrics employed, we do not measure the ability of code completion since we have observed that ChatGPT and other approaches can generate code responses for all given problems. Our experiment mainly assesses ChatGPT’s performance with respect to the correctness of the generated solutions and the rank percentile of time complexity for the correct solutions, as determined by test suites and execution time calculations within the LeetCode grading system. We engage ChatGPT in code generation tasks with the prompt illustrated in Figure 1. To compare ChatGPT’s performance with two state-of-the-art code generation tools, we select: (1) Codex (Chen et al., 2021), an advanced GPT based LLM model from the OpenAI family that has been widely investigated and applied in software engineering tasks before (Yetistiren et al., 2022; Vaithilingam et al., 2022; Pearce et al., 2021; Prenner and Robbes, 2021; Sarsa et al., 2022; MacNeil et al., 2023);. When evaluating Codex, we employ the same prompt as ChatGPT and keep the default parameters, with the exception of setting max_tokens to 1024. This adjustment of max_tokens ensures that we receive the complete code output from Codex. (2) CodeGen (Nijkamp et al., 2022), an autoregressive transformer model outside the OpenAI family. We employ the same prompt and parameters as ChatGPT and Codex for a fair comparison. CodeGen serves as a baseline model outside the OpenAI family for evaluating the performance of ChatGPT in code generation. Finally, to explore the usage of LLMs for programmers, we further analyze the impact of the length of prompts (problem description) on the code generation performance of the approaches. We compute the boxplot distribution of the length of prompts for correct and incorrect predictions made by ChatGPT and Codex respectively on LeetCode 2022.
[Experimental Results]: Table 3 demonstrates that ChatGPT outperforms Codex and CodeGen in generating correct code for problems across all difficulty levels and problem types within the LeetCode 2016-2022 dataset. Remarkably, ChatGPT can generate the correct code for all selected easy problems and 65% (26/40) of the hard problems on the metric of TOP-5. However, for AVG-5, we note that ChatGPT can generate correct code for 43% of hard problems, which indicates requesting ChatGPT multiple times may improve the probability to discover correct solutions. On the other hand, regarding the time complexity of generated code, the overall rank percentile of ChatGPT-generated correct code surpasses that of Codex-generated correct code while both of them are basically lower 50% (i.e., always in the top 50%). These rank scores imply that the efficiency of code produced by ChatGPT and Codex is potentially superior to most human-written code (assuming that nearly all LeetCode submissions are from humans) in terms of time complexity.
Table 4 presents the performance of the prior three approaches on LeetCode 2022 dataset. It shows that both metrics of the correctness and time complexity drop across all difficulty levels and types, with greater reduction as the difficulty increases. These results suggest ChatGPT and other LLMs struggle to generalize their abilities to new and unseen problems, i.e., the cold start problem: The model cannot draw any inferences for problems about which it has not yet gathered sufficient information. Fortunately, when mitigating the impact of data leakage, the results demonstrate that ChatGPT indeed improves the performance and generalization over the prior state of the arts. It can solve most of the easy problems (33/40) and merely 2 hard problems. On the other hand, the efficiency rank of code generated by ChatGPT are still in the top 50% for easy and medium problems while not for hard problems.
Figure 3 presents the distributions of prompt lengths for the correct and incorrect prediction. We exclude the hard level of ChatGPT and the medium and hard levels of Codex and CodeGen from the analysis since the number of problems solved by them is not sufficient to achieve statistical significance. We observe that the length of prompts are smaller in the correct predictions than that in the incorrect predictions: This implies that the models make more correct predictions when the prompts have a relatively shorter length. Such a finding has already been demonstrated in the literature (Li and Liang, 2021). Our results suggest that programmers should provide a clear and concise prompt when using ChatGPT and Codex for code generation.
2. ChatGPT for Program Repair
[Experiment Goal]: We investigate ChatGPT’s ability to repair diverse incorrect code implemented for common programming problems.
[Experiment Design]: As discussed in Section 3.3, we attempt to repair a large-scale dataset of 1783 incorrect codes of five common problems in Refactory benchmark with the goal of alleviating data leakage. We design two experiments for the repair of ChatGPT. First, we provide plain buggy code to request ChatGPT to repair by following the prompt shown in Figure 2. We note that Sobania et al. (Sobania et al., 2023) investigated that ChatGPT can repair more bugs if more information about bugs are provided at multi-round dialogues. However, the bug related information is not often available in practice and the multi-round dialogue tends to be time costly. Therefore, to consider the scenarios practical, we in this experiment aim to explore the repair ability of ChatGPT by providing the often available problem (function) description (2) at a one-time dialogue request. We construct the second prompt by inserting the description of each problem into line 2 of the prior prompt (Figure 2). For comparison with state-of-the-art, we select two approaches: (1) Codex (Chen et al., 2021), an advanced GPT based LLM model from the OpenAI family that has been widely investigated and applied in software engineering tasks before (Yetistiren et al., 2022; Vaithilingam et al., 2022; Pearce et al., 2021; Prenner and Robbes, 2021; Sarsa et al., 2022; MacNeil et al., 2023); (2) Refactory (Hu et al., 2019) (same name as the Refactory benchmark), a semantic-based APR tool for repairing student programs in real-time by refactoring correct reference solutions. This tool retrieves the top-5 closest programs to construct patches. If any of them pass the test suites, they succeed in repairing the buggy code. We thus compare them with the metric of TOP-5. We use repair rate to measure approaches’ performance while not considering time cost as it is affected by OpenAI server stability.
[Experimental Results]: Table 5 shows the repair performance of the evaluated three approaches. The x% (y%) indicates the results of TOP-5 (AVG-5) (Section 3.4). We note that ChatGPT is a generic and multi-task model. It is not specifically developed for repairing incorrect assignments. Regarding the TOP-5 metric, APR tool Refactory achieves the highest repair rate of 90% while ChatGPT achieves competitive results, repairing 84% incorrect codes and Codex can repair 66% of them. Specifically, ChatGPT (or ChatGPT_D) can achieve better results on the second, third and fifth problems than Refactory, with a 100% repair rate on the problem of Top k Elements. Regarding AVG-5, ChatGPT can repair an average of 60% of incorrect codes which is higher than Codex_D. These results reveal that ChatGPT does not achieve high performance when it is just allowed to generate one patch (once-attempt repair). On the other hand, when provided with the problem descriptions, Codex_D can repair more incorrect codes than Codex. However, ChatGPT_D underperforms ChatGPT except for the second problem, i.e., providing a problem description reduces the repair performance of ChatGPT.
To investigate the cause, we first manually analyze the incorrect submissions that can be fixed by ChatGPT but not by ChatGPT_D from the remaining four problems. We find that such incorrect code commonly contains bugs that are not governed by the problem descriptions. For example, Figure 4 shows an incorrect student submission of problem Soring Tuples. The code is incorrect because it misses the break statement in the second ‘for’ loop which leads to mistakenly repeated elements in the returned list. This output fails to pass the test cases. Meanwhile, the provided problem description does not explicitly exhibit the deduplication requirement. ChatGPT_D, therefore, is prone to ignore this issue and thus cannot find a way to fix the bug. For the incorrect code in problem Unique Dates Months, we find that some of the functions among them only contain a return statement, e.g., unique_month and contains_unique_day as shown in Figure 5. For this scenario, providing more information (description) enhances ChatGPT’s comprehension and its ability to fix the code, resulting in the superior performance of ChatGPT_D on this problem. Additionally, ChatGPT_D struggles to address the exception incurred by empty input, as the provided problem description does not specify the input requirements, drawing ChatGPT’s attention to other aspects (non-bug related information). Overall, ChatGPT has limited attention when addressing a given prompt input. To improve repair performance, we suggest programmers can provide ChatGPT with informative problem descriptions or additional information related to the bug such as failing test cases. To validate our speculation, we manually analyze 35 sampled incorrect codes that cannot be fixed by ChatGPT_D. By providing problem descriptions and extra failing test cases, 30 (95% confidence, 5% error) of them can be fixed for the first time.
3. ChatGPT for Code Summarization
[Experiment Goal]: We aim to explore the ability of ChatGPT to explain code. To do so, we evaluate whether ChatGPT can explain the intention of correct and incorrect code against the description of programming problems.
[Experiment Design]: Given a programming problem, programmers aim to achieve the same goal by writing code to satisfy associated specifications (e.g., problem description). As a result, the correct and incorrect code they write both reflects the same intention as the problem description. To assess ChatGPT’s ability to summarize code, we consider investigating whether ChatGPT can explain the original intention of code even though they contain mistakes or errors. To that end, we first request ChatGPT to explain the intention (natural language) of 2443 correct and 1783 incorrect codes in Refactory benchmark respectively (see the prompt in Section 3.4). Second, we employ BERT (Devlin et al., 2018), a widely-used pre-trained large language Encoder model known for its capability of understanding and classifying natural language (Khurana et al., 2023; Min et al., 2021; Qiu et al., 2020; Koroteev, 2021). We use BERT to produce the semantic embedding representations of the above intentions of the code and the associated problem descriptions. Finally, we compute the distribution of similarities between intentions (i.e., semantic embeddings) generated for correct code and the ground-truth problem descriptions, which we compare against the distribution of similarities between intentions (i.e., semantic embeddings) generated for incorrect code and the ground-truth problem descriptions. Overall, our objective is to evaluate whether ChatGPT can consistently generate the intention for correct and incorrect code.
[Experimental Results]: Figure 6 shows the result of the distribution. The similarity distribution between the intent of incorrect code and the problem description is not significantly different from the similarity between correct code and the problem description. We calculate Mann-Whitney-Wilcoxon (MWW) (Wilcoxon, 1945; Mann and Whitney, 1947) and confirm the null hypothesis (no significant difference).
Based on the observation, we further explore the characteristics of correct and incorrect code in different problems that allow ChatGPT to successfully infer their intentions. As we see, the similarity distributions in the second problem (Unique Dates Months) are lower compared to others. This discrepancy arises due to the second problem comprising three functions that pose challenges for ChatGPT in summarizing the intention. To investigate the effect of the multi-functions, we separate the three single functions from the code in the second problem to reproduce the experiments. We find the intentions of (in)correct code can be identified by ChatGPT for functions unique_day while not for unique_month and contain_unique_days. After manual inspection, we observe that many of the incorrect implementations of functions unique_month and contain_unique_days only contain a single return statement (an example in Figure 5), increasing the difficulty of intention identification of incorrect code (lowest similarity). As for the other problems, ChatGPT generates highly similar intentions of incorrect code as correct code and problem descriptions. For instance, the ground truth intention for the third problem is described as “This function takes in a list and returns a new list with all repeated occurrences of any element removed.” Figure 7 shows a typical intention inference example for the third problem (Duplicate Elimination). In this example, the presented code is incorrect due to the lack of a return statement. When being asked about the intention of this code, ChatGPT explains that this function returns a new list containing only the unique elements, even though the return statement is mistakenly missing in this code. To avoid the effect of the function name, we replace the function name with “a”. We then still observe the same result. Overall, the results reveal that ChatGPT is able to identify the original intention of the given (in)correct code. This finding enables ChatGPT to address the oracle problem: in test generation, we do not always have a precise specification of what the output should be. Given a correctness-unknown code snippet without the test suites, ChatGPT can help reason the original intention of the code for programmers. With this intention, programmers can determine the expected output for a test input, and design the test case accordingly. We validate our speculation by analyzing 39 sampled incorrect codes that ChatGPT can identify the defects of 30 of them by reasoning the original intentions and automatically generating test cases. A recent study has demonstrated the potential of ChatGPT to find failure-inducing test cases by reasoning the intention of code (Li et al., 2023).
Threats to validity
Our empirical study may contain several threats to validity that we have attempted to relieve.
Threats to External Validity. Large language models usually generate varied responses for identical input across multiple requests due to their inherent randomness. One threat to the validity of our study is that conclusions drawn from random results may be misleading. To mitigate this threat, we have requested results from LLMs five times to obtain the TOP-5 and AVG-5 performance metrics and compare them against state-of-the-art approaches. Additionally, we conduct our experiment on a large dataset (e.g., 1783 code samples for program repair), which helps reduce the impact of randomness to some extent. Another potential threat to validity is the selection of benchmarks. We make the choice based on the available large dataset and the need to prevent data leakage in LLMs.
Threats to Internal Validity. A major threat to internal validity lies in the processing of ChatGPT’s responses. Although we request plain code, ChatGPT occasionally generates natural language explanations alongside code in its responses, which can lead to the responses not being able to pass test suites. To mitigate this threat, we employ scripts to remove natural language explanations from the received responses and manually check the code that raises exceptions. On the other hand, annotations in the original student submissions from Refactory benchmark may impact the evaluation of ChatGPT’s code summarization capability on submissions. We also use scripts to remove these annotations from the submissions. Finally, we release the ChatGPT-generated responses and submissions with annotations removed for the community to review.
Threats to Construct Validity. In our experiment, our evaluation is based on the GPT-3.5 model. At the time of our submission, the latest model is GPT-4. However, GPT-4 is subject to rigorous request rate restrictions and has garnered reports of instability and inaccuracy. It’s worth noting that OpenAI continually updates the behind large language model that supports ChatGPT. Therefore, the results that we have reported in this paper are likely an underestimation of the capability of the ChatGPT models. In future studies, we plan to investigate the GPT-4 version once it establishes itself as a stable research platform.
Conclusion
In this paper, we performed an empirical study on the ChatGPT generative large-scale language model (LLM), aiming to assess its true potential as an assistant bot for programmers. We investigate ChatGPT on three code-related tasks: code generation, program repair and code summarization. In the task of code generation, we evaluate ChatGPT and two prior LLMs on 120 common programming problems sampled from LeetCode 2016-2020 and LeetCode 2022 respectively. The results demonstrate the dominance of ChatGPT while revealing that ChatGPT struggles to generalize to new and unseen problems. We also validate the negative impact of long prompts on the inference capabilities of ChatGPT. For the program repair task, we investigate ChatGPT’s ability to repair diverse incorrect implementations submitted for common programming assignments from a large-scale benchmark. ChatGPT achieves competitive results compared against Refactory, a state of the art semantic-based assignments repair tool. Additionally, we find that providing prompts that are not related to bug information (e.g., generic problem description) makes ChatGPT perform even worse: ChatGPT has a limited attention span. For the code summarization task, we evaluate whether ChatGPT can produce consistent intention explanations for (in)correct codes against ground truth problem description. Surprisingly, the results imply that ChatGPT can reason about the original intention of (in)correct code. This insight provides the opportunity to leverage ChatGPT in the testing research field where the test oracle problem remains an open question.
Overall, we expect the findings and actionable insights of this study on ChatGPT to have a potent impact across the research on software engineering, and in particular for pushing further assistive technology for programmers.
Data Availability. All the artefacts of this study are available in the following public repository: https://github.com/HaoyeTianCoder/ChatGPT-Study