Binary Code Summarization: Benchmarking ChatGPT/GPT-4 and Other Large Language Models

Xin Jin, Jonathan Larson, Weiwei Yang, Zhiqiang Lin

Introduction

Binary code comprehension is pivotal for binary analysis and reverse engineering, which delve into the intricate details of software to understand its functionalities (votipka2020observational, ). While the scope and outcomes of reverse engineering can differ based on tasks such as malware analysis (or2019dynamic, ; or2019dynamic, ), vulnerability identification (gao2018vulseeker, ; luo2023vulhawk, ), and symbol recovery (jin2022symlm, ; pei2021stateformer, ); the underlying objective remains consistent: to reconstruct the program’s logic and garner insights tailored to task-specific goals (mantovani2022re, ). Mirroring this intersection of code comprehension and natural language, VirusTotal recently introduced “Code Insight”, which employs Google’s Sec-PaLM to produce natural language summaries of potentially malicious code segments, enhancing threat understanding and detection for security analysts (virustotal-blog-2023, ).

Although significant progress has been made in automating various binary analysis tasks (shoshitaishvili2016sok, ; baldoni2018survey, ), reverse engineering of binary code semantics largely remains a process where human expertise is indispensable (mantovani2022re, ). This task is challenging, demanding considerable skill and time investment, especially given the absence of high-level semantics such as symbol names and data types. For example, it took about 40 minutes on average for 31 reverse engineering professionals to understand malicious decompiled code with less than 150 code lines (yakdan2016helping, ). Efforts such as the DARPA Cyber Grand Challenge (CGC) (darpa-cyber-grand-challenge, ) have propelled advancements in computers’ capacity to autonomously reason about program binaries. However, these automated solutions are still distant from being able to rival the expertise of human reverse engineering experts.

Recent advances in machine learning, particularly the emergence of large language models (LLMs) such as ChatGPT and GPT-4 (openai2023gpt, ), offer a promising avenue for addressing these challenges. These models have unlocked the potential for “strong artificial intelligence”, allowing interactions that closely mimic human-like intelligence and comprehension. The capabilities of LLMs have already initiated groundbreaking transformations in binary analysis tasks. Notable examples include binary symbol recovery (chen2022augmenting, ; jin2022symlm, ) and memory alias analysis (pei2022neudep, ). Additionally, LLMs, particularly advanced ones like ChatGPT have demonstrated promising results in text/source code comprehension and summarization (tian2023chatgpt, ; yang2023exploring, ). The emergence and widespread adoption of LLMs across various domains pose a pressing question: Can large language models, including ChatGPT, GPT-4, and other LLMs, effectively summarize the semantics of binary code?

The task to assess the full potential of LLMs in binary code summarization is far from straightforward due to the following challenges.

First, unlike readily accessible textual and source code corpora (e.g., CommonCrawl (commoncrawl, ) and BigCode (kocetkov2022stack, )), the absence of a comprehensive and large-scale binary code summarization dataset with ground truth, encompassing various binaries that conform to real-world compilation settings, poses a significant obstacle. Additionally, while existing LLMs only accept text-like input, binary code can be represented in different formats, such as assembly and decompiled code, making it challenging to ascertain the most suitable input for LLMs.

Semantic gaps in black-box models

Models like ChatGPT and GPT-4 undergo pretraining on an extensive corpus of textual data with sparsely represented binary code. It necessities mitigating the semantic gap between pretraining and our task, i.e., binary code summarization. Also, LLMs are commonly adapted to downstream tasks via either fine-tuning or prompting strategies (liu2023pre, ). However, finetuning LLMs are resource-intensive thus computationally infeasible, especially for the closed-source ChatGPT and GPT-4. For prompting, it has been found that instructing LLMs with different prompts can yield distinct performance (zhou2022large, ). For our task, devising effective prompts to maximize LLMs’ performance remains challenging considering the black-box nature of LLMs, making it hard to interpret what features or words contribute the most to the summarization results.

Varied outputs for the same semantics

LLMs produce outputs in natural languages, and thus the generated binary code summarization is in fuzzy textual format with semantically similar but syntactically different expressions, e.g., acronyms and synonyms. Precisely evaluating such summarization results is necessary but presents its own set of difficulties. For instance, the same semantics can be expressed by different words, phrases, and sentences in ground truth and LLM-generated summaries.

2. Insights

Addressing the aforementioned challenges, we propose a systematic exploration of LLMs’ capacities in binary code comprehension. Specifically, our study builds upon the following key insights:

Developers typically annotate source code with comments to elucidate both the programmer’s intent and the code’s functionality, and consequently most existing research on source code summarization (shi2022evaluation, ) and binary code summarization (al2023extending, ) utilizes code comments as the ground truth. Moreover, the commercial off-the-shelf (COTS) binaries are compiled with various settings, i.e., cross computer architectures and optimization levels. With such diverse binaries, existing binary analysis efforts have used different binary code representations, e.g., raw bytes in XDA (pei2020xda, ), assembly code in SymLM (jin2022symlm, ), and decompiled code in DITRY (chen2022augmenting, ). Thus, we have constructed a large-scale dataset using code comments as ground truth summaries with binaries compiled into 4 architectures and 4 optimization levels, from which 4 different binary representations are further generated.

Exploiting in-context prompt synthesis and optimization

For our task, prompts crafted by experienced binary analysis or prompt engineering professionals may yield commendable performance of LLMs. Nevertheless, this approach is constrained by its subjectivity and may not fully optimize LLM performance. Meanwhile, LLMs have recently shown the great protential of automatic prompt systhensis and optimization. Specifically, LLMs have the capacity akin to that of human prompt engineers (zhou2022large, ) and they exhibit the ability to self-improve (huang2022large, ). Therefore, we propose to first search over a pool of prompt candidates that are generated by LLMs themselves in an in-context manner, then optimize the LLM-generated and human-crafted prompts, and finally select the optimal prompts that achieve the best performance evaluated on our task.

Using semantic-embeddings to calculate the similarity

The commonly used metrics (e.g., BLEU (papineni2002bleu, ), METEOR (banerjee2005meteor, ), and ROUGE (lin2004rouge, )) for evaluation code summarization predominantly focus on exact matching of n-grams. However, they fail to capture the same semantics in acronyms, synonyms, and syntactically different but semantically similar expressions. The assessment of LLM-generated binary code summaries should be performed at the semantic level. Therefore, we propose to evaluate LLMs’ performance by calculating similarity using semantic embeddings.

3. Contributions and Findings

These insights have guided our study design, resulting in 10 key measurement results (R1\mathcal{R}_{1}-R10\mathcal{R}_{10}) (and 6 additional findings (F1\mathcal{F}_{1}-F6\mathcal{F}_{6})). In the following, we outline the key contributions:

A large-scale, open source, and comprehensive binary code summarization dataset: We have curated 44 open source projects (with totally 11,475,734 lines of code), which are compiled into binaries across 4 architectures (i.e., x86, x64, ARM, and MIPS) and 4 optimization levels (O0-O3). With our comment extraction and function-comment matching, we obtain 557,664 binary functions with comment-based ground truth summaries. We further generated four binary code representations, including raw bytes, assembly code, intermediate representations (IR), and decompiled code (both with and without debugging symbols), along with their source code.

Novel techniques: To generate optimal prompts that maximize LLM performance, we have devised a four-step procedure, including in-context prompt synthesis, prompt variant generation, prompt optimization, and task-specific prompt selection. Moreover, we also introduce a novel semantic embedding-based evaluation metric by calculating the semantic similarity between ground truth and LLM-generated summaries for our task by semantic embeddings generated from a pretrained LLM.

Extensive measurement: We have focused on four state-of-the-art LLMs: GPT-4, ChatGPT (GPT-3.5), Llama 2 (touvron2023llama, ), and Code Llama (roziere2023code, ), that have shown significant impact and advanced performance in downstream tasks along with the existing binary code summarization model BinT5 (al2023extending, ). Our four critical sets of evaluations generate 4,058,297,977 tokens in total with a cost of 11,418 US dollars and 873 NVIDIA A100 GPU hours, resulting in the following findings. First, stripping debugging symbols from binaries can lead to significant semantic loss (55.0%) (R1\mathcal{R}_{1}), and decompiled code is the best representation for LLMs to understand binary code (R2\mathcal{R}_{2}). Second, among the LLMs, ChatGPT and Code Llama perform the best in understanding binary code when with and without debugging symbols, respectively (R3\mathcal{R}_{3}). Code Llama consistently outperforms Llama 2, the base model from which Code Llama is fine-tuned, but merely fine-tuning CodeT5 on decompiled code does not yield satisfactory performance for BinT5 (R4\mathcal{R}_{4}). Additionally, our efficiency study reveals that ChatGPT and Llama 2 models are up to 2.9×\times and 1.5×\times faster than GPT-4 and Code Llama models, respectively (R5\mathcal{R}_{5}). Third, among the four computer architectures, LLMs perform the best on the x64 and MIPS binaries (achieving up to 16.0% better scores) given decompiled code with and without symbols, respectively (R6\mathcal{R}_{6}). However, there are marginal performance gaps for LLMs at different optimization levels (R7\mathcal{R}_{7}). Evaluations of the outputs of different decompilers (Ghidra, Hex-Rays, and Angr) demonstrate that LLMs perform the best on Hex-Rays decompiled code with a 60.7% better score (R8\mathcal{R}_{8}). We further investigate which symbols (i.e., function names, variable names, and data types) contribute the most to binary code semantics and find function names contribute the most to this semantics (R9\mathcal{R}_{9}). Finally, we also study the different prompt engineering techniques, i.e., zero-shot, few-shot, and chain-of-thought prompting, for our task, and identify zero-shot prompts as the best jointly considering the performance and cost (R10\mathcal{R}_{10}). In addition to the measurement results, we have also obtained six findings (F1\mathcal{F}_{1}-F6\mathcal{F}_{6}) by case studies (Appendix §A.6, §A.7, and §A.10), regarding to what extent LLM can understand binary code, GPT-4 summaries, and LLM summarizaition vulnerabilities. Our code and dataset are released at https://github.com/xinjin95/BinSum.

Design

In this section, we present the detailed design of how we measure the binary code summarization capability of LLMs. Note that, similar to existing research (shi2022evaluation, ; wang2021codet5, ; feng2020codebert, ; al2023extending, ), we focus on summarizing binary code at the function level. For clarity, we formalize this task as: Given a binary function ff extracted from binary executables and represented as a textual representation xf∈Xx_{f}\in\mathcal{X}, which can manifest as raw bytes, assembly code, intermediate representation (IR) code, or decompiled code, we concatenate it with a predetermined prompt II, resulting in an augmented input xf′={I;xf}x_{f}^{\prime}=\{I;x_{f}\}. Using this input, a binary code summarization system, M:X→Y\mathcal{M}:\mathcal{X}\rightarrow\mathcal{Y}, generates a natural language summary y∈Yy\in\mathcal{Y}, encapsulating the semantic crux of ff.

Our study of the binary code summarization task follows three pivotal steps: (i) We first construct a comprehensive binary code summarization dataset by compiling open-source repositories, resulting in diverse binaries across different architectures and optimization levels. These binaries are then further processed to yield four distinct binary code representations. (ii) Next, we follow a four-step procedure to synthesize, optimize, and select the most suitable prompts, guided by task-specific evaluations. (iii) Finally, we evaluate four state-of-the-art generative LLMs and the existing binary code summarization model (BinT5 (al2023extending, )) on our dataset, using our proposed semantic summary evaluation metric. In the subsequent subsections, we provide the detailed design for these three steps.

Our study aims to gain a comprehensive understanding of LLMs, which requires an extensive binary code dataset with ground truth. Before constructing such a dataset, we explored existing ones that could be readily employed. To the best of our knowledge, BinT5 is the only work studying the same task as ours (al2023extending, ) but with a focus on decompiled code. Its released data https://github.com/AISE-TUDelft/Capybara-BinT5 only includes x86 binaries, with an imbalanced distribution in terms of optimization levels, i.e., a rough ratio of 1:2:6:2 for O0-O3 binaries. We conclude that such a dataset, while useful, falls short of meeting our objectives of obtaining profound insights. Consequently, we opt to construct a dataset from scratch to align with our research goals, as shown in Figure 1.

Our initial step involves curating public source projects from GNU Software (GNUSoftware, ), including well-known ones that are extensively used by prior works (jin2022symlm, ; pei2021stateformer, ; li2021palmtree, ), such as coreutils, binutils, and findutils. The comprehensive list of 44 GNU projects is presented in Table 3, encompassing a total of 11,475,734 lines of code. The reason for choosing these GNU software projects is that they have been well-developed and maintained. We then compile them into four architectures (x86, x64, ARM, and MIPS) and four optimization levels (O0-O3) using GCC-7.5, resulting in 8,760 unique binaries with debugging information. Next, we generate another 8,760 stripped binaries by removing symbols. It is worth noting that for ARM and MIPS binaries, we utilize cross-compiler and stripping tools as part of the process (details in §A.2).

Source Code Comment Extraction

As alluded in §1, we have decided to use developer-written comments as ground truth. However, our binaries do not contain any comments, as the comments have been removed by the compiler’s preprocessor. Therefore, we need to extract code comments from the source code to annotate binary functions, which is a completely new task, thus no publicly available tools have been designed for this. As such, we first build a source code parser to achieve this. Specifically, for each binary function in the DWARF entries of binaries with debugging information, we pinpoint and use the comment of its corresponding function in the source code as the ground truth.

However, this process requires the precise identification of both comments and source code functions, especially function names, which presents several challenges. First, developers typically compose comments not only for functions but also for other code elements, such as macro definitions, statements, and basic blocks, as shown in Figure 2. It becomes difficult to discern comments that are specifically intended for functions. For instance, only the comment at lines 11 and 12 in Figure 2 is the desired comment for the target function yylex_destroy. Moreover, rather than defining signatures in one line, this example shows a multiple-line signature definition (lines 13 to 15), which means that a simple text-parsing-based method may fail to identify the desired function name. Furthermore, in addition to commenting in the source (.c) files, developers also create comments in the header (.h) files, which further complex the problem.

Fortunately, we have developed an automated source code comment parser (step ❷ in Figure 1) to address all these challenges, as illustrated in Figure 3, upon srcML (collard2013srcml, ) and ANTLR (parr2013definitive, ). Specifically, a manual study of 450 GNU source functions revealed that: (i) Comments in header files usually recur in corresponding source file function definitions; (ii) C function comments are predominantly placed above their signatures, using ‘//’ and ‘/**/’ for single- and multi-line comments. Given these observations, our parser targets function and comment extraction from source files. The function identification module constructs abstract syntax trees (ASTs) via ANTLR, converts ASTs to XML documents with srcML, and locates XML-tagged functions, outputting function signatures (sigfsig_{f}) and line ranges (locfloc_{f}). In parallel, the comment parsing module uses regex to identify comment texts (textctext_{c}) and their line ranges (loccloc_{c}). To associate functions with comments, we pair comments and subsequent functions whose line numbers are adjacently sequential, yielding matched function-comment entries. It is worth noting that we exclude internal function comments (e.g., comments at lines 17 and 21) from our analysis. Because such comments do not fully encapsulate the semantics of the function. Additionally, mapping these comments to corresponding blocks in binary code is exceptionally challenging, due to the lack of indicative symbols.

Function-Comment Matching Between Source and Binaries

After obtaining source function-comment pairs, we must annotate binary functions with these comments. Our proposed matching approach is based on two insights: (i) Binaries with debugging information provide function names and addresses via DWARF entries. Given control over the source code, we ensure DWARF symbols are generated and we retain function addresses from non-stripped binaries; (ii) Internal binary functions possess unique, non-conflicting names, acting as identifiers and facilitating matching with source function-comment pairs. Based on these insights, Ghidra (ghidra, ) is employed to parse DWARF entries, obtaining binary function names and boundaries (start and end addresses). We then extract the function name from every source function-comment entry’s signature (sigfsig_{f}). The final matching process (❸ in Figure 1) pairs binary function addresses with their related source comments using function names. This method identifies comments for 557,664 binary functions across 8,760 binaries with debugging symbols. At step ❹ in Figure 1, four representations are generated for binary functions: raw bytes, assembly code, IR code, and decompiled code (detailed in §3).

2. Prompt Synthesis and Optimization

As discussed in §1.2, we have observed the great potential of using prompts to guide LLMs to generate summaries in a zero-shot manner. However, the ideal prompts for binary code summarization remain implicit but necessary because LLMs are highly sensitive to input prompts (zhou2022large, ). An intuitive solution is to seek prompts from experienced prompt engineers or binary analysis professionals, but this method has its limitations. First, human prompt creators might be limited by their knowledge and creativity, leading to suboptimal prompts. Furthermore, LLMs can exhibit behavior that deviates from human expectations (zhou2022large, ). Lastly, it is not clear whether such prompts can optimize LLMs’ performance as LLMs are black-box. In Figure 4, we propose a four-step solution, encompassing: (1) in-context prompt synthesis, (2) prompt variant generation, (3) prompt optimization, and (4) task-specific prompt selection.

In this approach, we harness LLMs to generate a pool of candidate prompts by themselves within the context of binary code summarization. This design choice is inspired by the way human users engage with LLMs when confronted with a new task. In this scenario, human users often need to experiment with various prompts to elicit the desired responses. Due to the virtually unlimited prompt search space, manually generating numerous prompts is very challenging. Meanwhile, LLMs themselves have shown very promising text generation capacities (openai2023gpt, ), which motivates us to propose an automated prompt synthesis approach.

Formally, we consider our binary code summarization task specified by the dataset D={xi,yi}i=1N\mathcal{D}=\{x_{i},y_{i}\}_{i=1}^{N} with a prompt targeting model M\mathcal{M}, where xix_{i} and yiy_{i} are the input binary code and the code summary. Our objective is to generate a suitable prompt, denoted as pp, in a manner that when M\mathcal{M} is given the concatenated input (p;x)(p;x), M\mathcal{M} generates the corresponding output yy. We cast this task as an optimization problem to find the optimal prompt p^\hat{p}. This optimization is centered around maximizing the expected score SM(p,x,y)S_{\mathcal{M}}(p,x,y), across all potential pairs of (x,y)(x,y):

where SM(⋅)S_{\mathcal{M}}(\cdot) could be computed in several ways, such as the aggregated token probability (zhou2022large, ) and the beam search probability (li2023guiding, ). However, we observe that API-based LLMs like ChatGPT are black-box generators. That is, the LLM output does not contain any probability products; thus, these probability-based score functions cannot apply to our task. Alternatively, we directly take the top-1 LLM-generated prompt as p^\hat{p} as it is produced by maximizing the autoregressive prediction probability (openai2023gpt, ), which can represent the maximum score of SM(p,x,y)S_{\mathcal{M}}(p,x,y). Apart from synthesized prompts, we also include human-crafted prompts to increase the size and diversity of our prompt candidate pool as demonstrated in Figure 4.

(2) Prompt Variant Generation

In our initial prompt candidate pool, there may be instances where this synthesis approach falls short, either due to a lack of diversity or an absence of candidates with sufficiently high scores. In response to such problems, we generate prompt variants that are semantically similar to the candidate prompts as shown in Figure 4. This approach essentially explores the search space in close proximity to the current best candidates. Therefore, it enables us to create new prompts that are more likely to yield favorable results.

(3) Prompt Optimization

To further enhance the performance of LLMs in the context of binary code summarization, we introduce a strategy to optimize the previously synthesized prompts and their variants using the LLMs themselves. Our design choice is motivated by prior research, showing that LLMs can optimize the problems of linear regression and traveling salesman (yang2023large, ), and LLMs can also self-improve (huang2022large, ). In our task, the advantage of employing LLMs for prompt optimization lies in their ability to comprehend natural language, i.e., enabling us to convey our prompt optimization objectives by high-level textual instructions without formal specifications. Given the candidate prompts, we instruct LLMs to serve as a prompt optimizer by a meta-instruction. This meta-instruction encapsulates the objective function, i.e., maximizing performance, and solution constraints, i.e., the context of the binary code summarization task, as shown in Figure 4.

(4) Task-Specific Prompt Selection

After generating a pool of prompt candidates by prompt synthesis and optimization, the subsequent selection step becomes crucial in determining the most promising candidates for our task. This process closely aligns with the well-established problem of best arm identification in Upper Confidence Bound (UCB) optimization (pryzant2023automatic, ). In this context, arms correspond to the prompt candidates, the hidden value of each arm represents its performance on our task, and the arm-pulling action corresponds to evaluating prompts with randomly selected data points. The ultimate objective is to identify the top-performing arms while minimizing the number of pull efforts required to achieve this outcome.

Evaluating each candidate prompt on our entire dataset is a resource-intensive process and prohibitively expensive. Inspired by previous work that efficiently estimates prompt performance (zhou2022large, ), we assess prompt candidates by evaluating them on a randomly selected subset of data samples from our dataset. Figure 4 presents how we proceed: We first generate the binary code summary by querying LLMs with concatenated every prompt candidate and the test binary sample. These generated summaries are then evaluated by calculating semantic similarity scores to ground truth (see the evaluator details in §2.3). Finally, the prompts with the best scores are selected as the final prompts for our subsequent evaluations.

3. Testing and Semantic Evaluation

Our goal is to assess existing LLMs’ performance on our task. For this, we follow a two-step process: (1) LLM testing on binary code summarization and (2) semantic summary similarity evaluation.

To generate summaries, we concatenate the selected prompt with the test binary code sample for model input preparation. This configuration essentially places us in a zero-shot setting, jointly considering both performance and cost considerations (see §4.5 for the evaluations of different prompting settings). As LLMs are generative models, they can generate summaries of varying lengths. Nevertheless, excessively verbose and lengthy summaries are generally regarded as low-quality (iyer2016summarizing, ). Therefore, we impose limits on response length for our target models. For instance, we explicitly specify these limits in our prompts with the instruction “summarize … in NN words”, where NN is the average number of words in the ground-truth summaries from our dataset. We observe that our test models exhibit an adaptive length behavior when generating summaries. That is, they tend to produce summaries with lengths that closely hover near the predetermined length limit, demonstrating a capacity to adjust the length intelligently without rigidly and dogmatically adhering to the predetermined length limit. This is beneficial considering the inherent diversity of binary code samples.

(2) Semantic Summary Similarity Evaluation

To assess LLM-generated binary code summaries, we aim to estimate how they align with the ground-truth summaries. For this, we notice that prior source and binary code summarization research (shi2022evaluation, ; al2023extending, ) both use exact matching-based metrics, e.g., BLEU (papineni2002bleu, ), METEOR (banerjee2005meteor, ), and ROUGE (lin2004rouge, ). We have found that these metrics fall short in capturing the essential semantics of summaries (see F4\mathcal{F}_{4} and Appendix §A.6). Alternatively, we propose solving this problem based on semantic measurement. For this, we observe that pre-trained LLMs, such as BERT and XLNet, have gained success in text semantic modeling tasks (min2021recent, ). Therefore, we can calculate the semantic similarity between LLM-generated and ground-truth summaries using these pre-trained models as the semantic measurement.

Here, we use the mean pooling as the aggregation function which has shown superior results in text embedding tasks (reimers2019sentence, ). Finally, we calculate the semantic similarity between generated summaries and ground-truth summary references by the cosine similarity function following the best practice (minaee2021deep, ):

The semantic similarity score sim(S,S^)sim(S,\hat{S}) quantifies the alignment between generated summaries and summary references. A higher similarity score indicates a better quality of the generated summary. Figure 5 illustrates the process of calculating this semantic similarity score.

Implementation

Dataset Construction. As depicted in §2.1, we have built a binary code summarization dataset by compiling 44 open-source projects with 11,475,734 total lines of code (Table 3 in Appendix has more details). Our dataset contains 557,664 binary functions with ground truth summaries. These functions are from various binaries, and their distributions across architectures (x64, x86, ARM, MIPS) and optimization levels (O0-O3) are presented in Table 4 in the appendix for readers of interest. For each function in our dataset, we generated four binary code representations along with source code:

Raw bytes are directly extracted from binaries based on the function start and end addresses (identified in §2.1).

Assembly code is generated by parsing ELF binaries using capstone (capstone-engine, ) and pyelftools (pyelftools, ) libraries.

IR code is the IDA microcode generated by our IDA plugin built upon the ida_hexrays.gen_microcode API (ida-hexrays-doc, ) where we have followed prior binary analysis research (yu2020order, ; yu2020codecmr, ) to use the widely adopted IDA microcode as IR code.

Decompiled code encompasses outputs generated by Ghidra, Hex-rays, and Angr, using the getDecompiledFunction (ghidra-decompiler-doc, ), idaapi.decompile (ida-idaapi-doc, ), and angr.analyses.Decompiler (angr-decompiler-doc, ) APIs, respectively. We have chosen Ghidra and Hex-rays as they are prominent decompilers. Angr, on the other hand, has been selected due to its rapid growth and comprehensive capabilities in binary analysis.

Source code is generated by parsing source files according to functions’ start and end line numbers. Note that we exclude the function comments to avoid information leakage to LLMs when testing their capability of summarization.

We have generated a set of 320 prompt candidates based on our prompt synthesis and optimization approach. As mentioned in the task-specific prompt selection section of §2.2, it is prohibitively expensive and thus impractical to evaluate all prompt candidates on the entire binary summarization dataset. Hence, we randomly select 1000 binary function samples from our dataset to conduct binary code summarization using each prompt candidate. Table 5 presents the top 40 prompts ranked by semantic similarity. We have selected the best prompt for our subsequent evaluations. For a fair comparison, we employ the same prompt for all models, except for BinT5, which only takes the decompiled code as input (see §A.5 for more details).

Evaluation Metrics

Our evaluation metrics include the semantic similarity metric that we proposed in §2.3. Specifically, we implemented our evaluation framework on top of Pytorch (paszke2019pytorch, ), Sentence-Transformers (reimers2019sentence, ), Scipy (virtanen2020scipy, ), and Sentencepiece (kudo2018sentencepiece, ). For the summary encoder model, we use the pre-trained all-mpnet-base-v2 model, which performs the best on 14 text embedding tasks (sbert-pretrained-models, ). To efficiently compute eSe_{S} and eS^e_{\hat{S}} in Equation 2, we store token embeddings in a Chroma vector database (trychroma, ). The calculated semantic similarity scores range from -1 to 1 where a score close to 1 indicates a high quality of LLM-generated code summaries. In addition to our proposed metric, we also calculate and report the scores of BLEU (papineni2002bleu, ), METEOR (banerjee2005meteor, ), and ROUGE-L (lin2004rouge, ) to facilitate future work, as these metrics have previously been employed in both source and binary code summarization research (al2023extending, ; wang2021codet5, ; shi2022evaluation, ). The calculation details of BLEU, METEOR, and ROUGE-L are in Appendix §A.4. BLEU, METEOR, and ROUGE-L scores are scaled between 0 and 1, where higher values signify superior summarization quality.

Targeted LLMs

The LLMs of our interest are GPT-4, ChatGPT, Llama 2 (7B), Llama 2 (13B), Code Llama (7B), Code Llama (13B), and BinT5 (al2023extending, ), as presented in Table 1, which has exhibited dominant performance across numerous NLP and code-related tasks (huggingface-llm-leaderboard, ; touvron2023llama, ; openai2023gpt, ). To test GPT-4 and ChatGPT, we use OpenAI’s chat completion API (openai.Completion.create) (openai-api-chat-complete, ), which have called gpt-4-0613 and gpt-3.5-turbo-16k-0613 backend models, corresponding to the 0613 snapshots, released on June 13th, 2023. To get consistent results, we have set the temperature parameter (openai-api-chat-complete, ), controlling models’ diversity and creative output, to 0.1. Furthermore, we set both top_p and n at 1 to obtain the top-1 results (openai-api-chat-complete, ). Regarding the parameters stop and max_tokens, we opt for the default values (openai-api-chat-complete, ), allowing the models to autonomously determine when to conclude their responses. To enhance inference speed, we have implemented multi-threading, utilizing five and six parallel threads for GPT-4 and ChatGPT, respectively, while ensuring compliance with the rate limits set by OpenAI (openai-rate-limits, ).

We have donwloaded Llama models from Huggingface, including meta-llama/Llama-2-7b-chat-hf, meta-llama/Llama-2-13b-chat-hf, codellama/CodeLlama-7b-Instruct-hf, and codellama/CodeLlama-13b-Instruct-hf (meta-llama, ; code-llama, ). To run these models, we implemented an inference framework based on Pytorch (paszke2019pytorch, ), transformers (wolf2019huggingface, ), Scikit-learn (pedregosa2011scikit, ) and DeepSpeed (rasley2020deepspeed, ). For efficient inference, we enable mixed precision in BF16 and batch size at 6. We directly use their tokenizers and pretrained model weights. We set temperature as 0.1 and top_k and top_p as 1 to get the best summaries. For BinT5, it only accepts decompiled code, thus we only use it for summarizing decompiled code. We set its inference batch size as 64, using its maximum context size, i.e., 512.

Evaluation

In this section, we present our evaluation results. While there are many research questions (RQ) centered around this measurement, we seek to answer the following ones:

RQ1: To what extent can LLMs comprehend binary code? What input of binary code impacts LLM’s output more?

RQ2: Which LLM performs the best on binary code comprehension? Which LLM is more efficient than others?

RQ3: How do the different computer architectures and optimization levels affect LLMs’ performance?

RQ4: What are additional factors of binary code input influencing LLMs’ comprehension capabilities?

Evaluation Environments. Our evaluation environments include a ThinkPad P15 desktop and a Dell cloud server. The ThinkPad desktop has an Intel i9-10885H vPro CPU (8 cores, 2.4 GHz), 128 GB RAM, 1TB storage, and an NVIDIA Quadro RTX 4000 Max-Q GPU, running Windows-11. The Dell cloud server has an AMD EPYC-7643 CPU (88 usable cores, 2.3 GHz), 921 GB RAM, 12.8 TB storage, and 4 NVIDIA A100 GPUs with 80 GB VRAM each, supercharged by NVLink with the RHEL-8.6 operating system.

We ran GPT-4 and ChatGPT on our binary code summarization datasets, completing the tasks in 20 and 6 days, respectively. GPT-4 tokenized all test samples into 309,401,126 tokens and generated responses totaling 19,678,448 tokens, summing up to 329,079,574 tokens, as reported in the first row of Table 1. For ChatGPT, the input samples were tokenized into 296,524,698 tokens, and it generated 16,338,012 tokens, resulting in a total of 312,862,710 tokens. In terms of the cost, we spent 10,462.74and10,462.74 and954.93 US dollars for querying GPT-4 and ChatGPT models, respectively. For Llama 2 and Code Llama models, they tokenize our input samples into 736,413,994 tokens and generate 33,195,771, 29,699,869, 40,016,196, and 33,323,576 tokens as responses. The execution of Llama models takes us 857 GPU hours to finish. Finally, BinT5 produced 326,007,985 tokens, including 326,007,985 input and 8,456,320 output tokens within 16 GPU hours. To sum up, the model execution costs 11,418 US dollars and 873 NVIDIA A100 GPU hours.

2. RQ1: Effectiveness

To answer RQ1, we first evaluate the performance of LLMs on different representations of binary code, including decompiled code, IR code, assembly code, and raw bytes along with source code, and present the results in Figure 6, in which BinT5’s performance is not included for a fair comparison, as it only accepts decompiled code. Among the different code representations, we find that LLMs perform the best in source code, achieving an average semantic similarity score of 0.474. For decompiled code, we observe a significant performance degradation when symbols are stripped from binaries, reducing the semantic similarity score from 0.449 to 0.202 (55.0% decrease).

Result 1 (R1\mathcal{R}_{1}) – Stripping debugging symbols significantly loses decompiled code semantics by 55.0%.

Even with such a significant information loss, decompiled code without symbols still outperforms other binary code representations. Specifically, the average semantic similarity scores of assembly code, IR code, and raw bytes are 0.188, 0.185, and 0.118, respectively. The same trend also appears in other evaluation metrics, e.g., the mean METEOR scores are 0.167, 0.141, 0.102, 0.052, 0.066, and 0.023 for source code, decompiled code (debug), decompiled code (stripped), IR code, assembly code, and raw bytes, respectively.

Result 2 (R2\mathcal{R}_{2}) – LLMs perform the best on decompiled code compared to IR code, assembly code, and raw bytes.

After obtaining the statistical results, we also conduct an in-depth study to understand to which extent LLMs can comprehend binary code, as well as the nuanced interpretation of specific score outcomes. The result of this study is presented in Appendix §A.6 for readers of interest. At a high level, our case studies of LLM-generated summaries provide additional confirmation of the deterioration of binary code semantics caused by symbol stripping (F1\mathcal{F}_{1}). Furthermore, rather than elucidating high-level functionalities, LLMs tend to generate operational summaries for low-level binary representations, i.e., IR and assembly code (F2\mathcal{F}_{2}). In the case of raw bytes, it is intriguing to note that the LLM-generated summaries closely resemble those of the assembly code, indicating the implicit code-lifting behavior (F3\mathcal{F}_{3}). And, compared to exact matching-based metrics, our proposed semantic evaluation metric has been shown to be better suited for our task, as it effectively captures summary semantics (F4\mathcal{F}_{4}).

3. RQ2: LLM Comparison

To answer RQ2, we study individual performance for each of our target LLMs. Having identified that decompiled code is the best for LLMs to understand binary code (R2\mathcal{R}_{2}), our report emphasizes the performance of LLMs on decompiled code, as shown in Figure 7. For decompiled code with symbols, we find that ChatGPT performs the best, achieving an average semantic similarity score of 0.543. Meanwhile, for stripped binaries, Code Llama models outperform all other LLMs achieving 0.284 and 0.283 average semantic similarity scores for 7B and 13B models. It is a surprise that GPT-4 is not the best-performing model. Therefore, we delve into a thorough examination of its results. Overall, we observe that GPT-4 can understand binary code. However, it places a greater emphasis on the intricate and superfluous details, compared to ChatGPT, resulting in noisy summaries (F5\mathcal{F}_{5}). A more detailed analysis and specific examples are provided in Appendix §A.7.

Result 3 (R3\mathcal{R}_{3}) – ChatGPT and Code Llama perform the best for decompiled code with and without symbols, respectively, compared to GPT-4, Llama 2, and BinT5.

Code Llama is fine-tuned on the code corpora from Llama 2 (roziere2023code, ). For decompiled code with symbols, Code Llama achieves 4.57%, and 11.3% better performance than Llama 2 on the 7B and 13B models. For decompiled code from stripped binaries, Code Llama models further outclass Llama 2 model by 12.3% and 22.0% on the 7B and 13B models. This outperformance of Code Llama demonstrates that Code Llama’s fine-tuning can improve its capacity for binary code comprehension. Meanwhile, BinT5 is the model fine-tuned from the CodeT5 model by retraining on decompiled code (al2023extending, ). Unfortunately, BinT5-generated summaries receive the lowest semantic similarity scores, i.e., 0.115 and 0.114, on binaries with and without symbols, compared to all other LLMs. It is worth noting that for a fair and unbiased comparison among LLMs, we have excluded test samples that exist in BinT5’s training dataset (See §A.8 for more discussions). The marginal performance difference in binaries with and without symbols for BinT5 suggests that it may not fully capture the intricate semantics within binary code, even with symbols. More importantly, we hypothesize that the different impact of fine-tuning in Code Llama and BinT5 may stem from disparities in their underlying base models, namely Llama 2 and CodeT5.

Result 4 (R4\mathcal{R}_{4}) – Code Llama, fine-tuned from Llama 2, consistently outperforms Llama 2 (up to 22.0% better scores). However, BinT5, fine-tuned from CodeT5, unfortunately, achieves the worst performance among all LLMs.

As listed in Table 1, different LLMs have different model sizes and parameters, which can potentially lead to different efficiency. Thus, we evaluate and compare the inference time cost per sample of our target models. Table 2 presents the average inference time per sample for different LLMs. Here, we have two different evaluation settings as stated in the model execution section of §4.1. For ChatGPT and GPT-4, their execution relies on OpenAI API calls, making it challenging to determine the computational infrastructure that underpins their operation. Their direct comparison shows that ChatGPT is 2.9×\times faster than GPT-4. For other models, we have run them on our own cloud server which gives us a fair comparison. Among these models, BinT5 is the fastest one among all models, e.g., 7.8×\times faster than Code Llama (13B). We also find that Code Llama takes more time for inference compared to Llama 2, e.g., 1.5×\times more time on the 13B models. In addition to the time cost, we have also observed the memory consumption difference among the open-source LLMs. Specifically, the 7B and 13B Llama models occupy 42.6GB and 61.9GB of GPU VRAM, respectively, while the BinT5 model only takes 3.4GB of VRAM.

Result 5 (R5\mathcal{R}_{5}) – Compared to GPT-4, ChatGPT is 2.9×\times faster. For the other LLMs, BinT5 is the fastest model, while Llama 2 is (up to 1.5×\times) faster than Code Llama.

4. RQ3: Various Binaries

To answer RQ3, we first study LLMs’ performance on binaries across different computer architectures. Figure 8 presents the performance on binaries across different architectures, i.e., x86, x64, ARM, and MIPS. For binaries with symbols, the average semantic similarity scores for x86, x64, ARM, and MIPS binaries are 0.483, 0.503, 0.497, and 0.485, respectively. The performance on x64 binaries outclasses these on x86 binaries by 4.14%. For stripped binaries, the similarity scores are 0.262, 0.299, 0.276, and 0.304 for the x86, x64, ARM, and MIPS binaries, respectively, in which LLMs’ performance on MIPS binaries is 16.0% better than these of x86 binaries.

Result 6 (R6\mathcal{R}_{6}) – LLMs perform the best on x64 binaries with symbols and MIPS stripped binaries. The cross-architecture performance gap can be up to 16.0%.

In Figure 9, we present the performance of LLMs when applied to binaries with O0, O1, O2, and O3 optimization levels. For binaries with symbols, LLMs attain average text similarity scores of 0.481, 0.477, 0.477, and 0.474 for O0, O1, O2, and O3 binaries, respectively. Conversely, for stripped binaries, the average text similarity scores for these binary optimization levels are 0.269, 0.268, 0.266, and 0.270. In both categories of binaries, we notice only a slight variance in performance across different optimization levels, with the performance gap between O0 and O3 binaries amounting to merely 1.47%.

Result 7 (R7\mathcal{R}_{7}) – LLMs show marginal performance gap across different optimization levels, e.g., 1.47% different text similarity scores between O0 and O3 binaries.

5. RQ4: Other Factors

Since we have identified that the decompiled code is the best for LLMs to understand (R2\mathcal{R}_{2}), next we undertake a further investigation of which decompiler generates the optimal output for code comprehension to answer RQ4. Figure 10 presents the LLMs’ performance on decompiled code generated from three different decompilers. Overall, we find that Hex-rays consistently outperforms the other two decompilers for both binaries with and without symbols. Specifically, for binaries with symbols, the output of Hex-rays achieves 5.41% and 21.9% higher semantic similarity scores than those of Ghidra and Angr. For stripped binaries, it further outclasses Ghidra and Angr by 60.7% and 19.8%, respectively. Among the decompilers, we observe that Angr is significantly slower than others. For instance, Angr requires about 15×\times more decompilation time than Ghira with setting the CFGFast analysis for Angr, while Ghidra and Hex-rays consume the comparable time for the same task.

Result 8 (R8\mathcal{R}_{8}) – The decompiled code output of Hex-rays obtains the best scores compared to those of Ghidra and Angr, achieving up to 60.7% outperformance.

Observing the performance decline from symbol stripping (R1\mathcal{R}_{1}), we delve deeper to discern which among the three symbol types—data types, variable names, and function names—most influence decompiled code semantics. Starting with the original decompiled code with all symbols as our baseline, we exclusively strip each symbol type by substituting them with non-informative placeholders (detailed in §A.9). This produces three sets of modified decompiled code, with each having only one symbol type replaced while the rest remains intact. Additionally, we use the decompiled code from stripped binaries as an all-symbol-removed benchmark. Figure 11 showcases the LLMs’ performance across these variations. Stripping just the function names led to the most substantial drop in performance—a 30.3% decrease in semantic similarity. Variable names and data types caused reductions of 30.0% and 19.5%, respectively. Moreover, we also find that the semantics of LLM-generated summaries can be manipulated by intentionally modified function names. For this, we perform a case study and identify a vulnerability of LLM-generated summaries by function name manipulation (F6\mathcal{F}_{6}) in §A.10.

Result 9 (R9\mathcal{R}_{9}) – Function name contributes the most to decompiled code semantics. Merely stripping function names can reduce LLM’s performance by 30.3%.

In our study, we predominantly employ zero-shot prompts. However, we have also observed the use of advanced prompts, e.g., few-shot and chain-of-thought prompts, in NLP tasks (liu2023pre, ); therefore, we investigate if these prompts can help LLMs in our task. Fundamentally, the few-shot prompts add demonstration examples (as shown in 15(b) in Appendix), e.g., pairs of binary code and corresponding summaries, to LLM input (brown2020language, ). For chain-of-thought prompting (wei2022chain, ), we send two queries to LLMs (as exemplified in 15(c)). In the first query, LLMs are first asked to explain the test code with the popular reasoning prompt “Let’s think step by step”. The resulting explanation is extracted and used to summarize the test code in the second query. We provide more details of how we construct these prompts in §A.11.

Figure 12 presents the performance of different prompts, where zero-shot, few-shot, and chain-of-thought prompts achieve average semantic similarity scores of 0.282, 0.312, and 0.271, respectively. Compared to zero-shot prompts, few-shot prompts improve LLM performance by 10.6%, while chain-of-thought reasoning does not help. Beyond disparities in performance, we have also noted substantial variations in cost. Figure 13 presents the length distribution of different prompts, where the median numbers of tokens for zero-shot, few-shot, and chain-of-thought prompts are 244, 1,763, and 944, respectively. The 7.23×\times more tokens of few-shot prompts mean significantly higher computational cost. For instance, since the use of ChatGPT and GPT-4 is changed based on token numbers, few-shot prompts would cost us 82,552 US dollars to finish the same amount of tokens listed in Table 1. Taking into account the scale of our evaluations, zero-shot learning can be a better solution based on the balance of performance and cost.

Result 10 (R10\mathcal{R}_{10}) – Zero-shot prompts are better than few-shot and chain-of-thought prompts when jointly considering both performance and cost for our large-scale task.

Lessons Learnt

Prioritizing Binary Code Comprehension is Paramount. The need for binary code summarization, albeit an area still in its infancy (al2023extending, ), has rapidly grown, with rising demand for natural language summaries of binary code (virustotal-next-step-2023, ; google-cloud-blog-ai-insights, ; microsoft-security-copilot, ). Notable entities like VirusTotal (virustotal-next-step-2023, ), Google Cloud (google-cloud-blog-ai-insights, ), and Microsoft Security Copilot (microsoft-security-copilot, ) are increasingly leveraging LLMs to understand malware/threats, aiming to produce ready-to-use summaries to security professionals. Yet, binary code comprehension remains a formidable challenge, especially for stripped binaries. In our studies, LLMs either fail to grasp the complete semantics (F1\mathcal{F}_{1}) or generate summaries missing the nuances of high-level code semantics (F2\mathcal{F}_{2}). The multifaceted nature of binaries, along with their diverse representations, compounds this challenge. For instance, LLM performance varies across binary representations (R2\mathcal{R}_{2}) and is sensitive to decompiler outputs (R8\mathcal{R}_{8}). Additionally, differing compilation settings, such as computer architectures, can influence LLM outcomes (R6\mathcal{R}_{6}). In navigating this intricate landscape, our findings offer pivotal insights that can guide future endeavours in comprehensively understanding binary code.

Harnessing LLMs Requires Further Advancements and Explorations. Machine learning models are extensively utilized in binary analysis, yet many face challenges in generalizing to unseen samples (jin2022symlm, ; chen2022augmenting, ). In contrast, LLMs like ChatGPT and GPT-4 boast superior generalizability—crucial for the vast array of real-world binaries. However, these promising models come with limitations. They are resource-intensive, with GPT and Llama models demanding expensive computational resources and specialized hardware accelerators (§4.3). Potential solutions may lie in model distillation and hardware acceleration techniques (deng2020model, ), especially as smaller models have proven to be more resource-friendly (R5\mathcal{R}_{5}). The black-box nature of LLMs presents another challenge, making it tricky to decipher their exact behaviors. Yet, our work suggests that refining and optimizing prompts might be a pathway forward, as highlighted by our prompt evaluations (Table 5 and R10\mathcal{R}_{10}). Additionally, while fine-tuning LLMs can enhance binary code comprehension, selecting the right base models is essential, as underscored by our findings in R4\mathcal{R}_{4}. In essence, while LLMs hold immense promise in binary code analysis, their effective deployment hinges on computational challenges, requiring selecting the correct models and inputs.

Augmenting Decompilers is Crucial for Improved Binary Comprehension. Our studies indicate that decompiled code proves most intuitive for LLMs (R2\mathcal{R}_{2}). Enhancing decompilers, particularly for stripped binaries, cannot be overstated. Our findings suggest that binary stripping can greatly diminish the semantics of decompiled code (R1\mathcal{R}_{1}). Furthermore, aside from BinT5 (al2023extending, ), all other LLMs exhibit a significant performance gap for binaries with and without symbols (Figure 7). This leads to LLMs’ inability to fully capture the semantics of stripped binary functions (F1\mathcal{F}_{1}). A promising approach to strengthen decompilers can be symbol recovery (chen2022augmenting, ; jin2022symlm, ; banerjee2021variable, ), as depicted in Figure 11. Our research pinpoints function name prediction (jin2022symlm, ; david2020neural, ; patrick2023xfl, ) as particularly influential, given its paramount contribution to the semantics of decompiled code (R9\mathcal{R}_{9}). The consistency check of binary function names is also important, particularly because LLM-generated summaries are susceptible to malicious manipulation through intentional alterations of function names (F6\mathcal{F}_{6}). Currently, Hex-Rays stands out as a leading decompiler, consistently producing output that LLMs perform the best on (R8\mathcal{R}_{8}). Moving forward, we believe that the insights from our study can serve as invaluable guidance for improving decompilers.

Threats to Validity

Scope of Study and Binary Diversity. In this study, we have performed an extensive study of LLM’s capabilities for binary code comprehension, assessing over 557K binary functions across different architectures and optimization levels. However, real-world binaries extend beyond the confines of our dataset, encompassing domain-centric binaries like IoT firmware and malware. Such binaries often possess unique semantics which might be further complicated by code obfuscations. Moreover, while our study has assessed specific binary code representations, it does not encompass all possible representations, such as those generated by other decompilers or IR code generators. Delving into these areas presents a promising avenue for future research.

Our study evaluated four leading LLMs recognized for their exceptional performance across numerous tasks (roziere2023code, ; openai2023gpt, ), as well as the binary code summarization model, BinT5 (al2023extending, ). Still, with academia and industry ceaselessly advancing, newer LLMs are continually being introduced, and these might be applicable for our task. Additionally, while our research employed a prompt synthesis and optimization method yielding 320 prompts (§4.1), there remains potential for discovering even more effective prompts and refining prompt engineering methodologies, like gradient-based prompt tuning (liu2023pre, ).

Code Summary and Semantic Representations

In our research, developer-written code comments served as our gold standard for code summaries. Alternative representations for these summaries exist as well, such as analysis logs or system calls (hao2023syzdescribe, ; pan2023automated, ). However, these representations are artificial, generated by third parties instead of the code creators, and might not always align semantically with the original code. We observe that developers’ comments present a more intuitive option. This aligns with conventional methodologies utilized in prior studies (feng2020codebert, ; husain2019codesearchnet, ; al2023extending, ). Similar to well-known code summary datasets, e.g., CodeSearchNet (husain2019codesearchnet, ), our dataset may also include semantic inconsistent noise (chen2021my, ) introduced by developers. While mitigating this noise is still an open question, this would be a worthwhile focus for subsequent research.

Related Work

LLM for Binary Reverse Engineering. Binary reverse engineering is labor-intensive and suffers from the absence of high-level semantics, especially for stripped binaries. Existing LLM-based solutions follow the pretraining-finetuning paradigm to learn binary semantics from large binary corpora and transfer the learned knowledge to downstream tasks, e.g., function similarity detection (li2021palmtree, ; pei2020trex, ; wang2022jtrans, ), variable name and type inference (pei2021stateformer, ; chen2022augmenting, ; banerjee2021variable, ), and function name prediction (jin2022symlm, ; david2020neural, ; patrick2023xfl, ). For example, Trex (pei2020trex, ) pretrains roBERTa on the assembly code and execution trace to detect function similarity. SymLM (jin2022symlm, ) predicts function names by learning calling context and execution behavior. Unlike existing works that focus on BERT-level LLMs, we focus on generative LLMs that exhibit significantly better generalizability and emergent abilities (wei2022emergent, ).

Automated code summarization aims at articulating code fragments, typically methods or functions, using natural language summaries (shi2022evaluation, ). Recently, the predominant focus of this research has been on source code (guo2020graphcodebert, ; wang2021codet5, ; feng2020codebert, ) As an example, CodeT5 (wang2021codet5, ) pretrains the T5 model with semantics-enriched code features, e.g., identifiers, to summarize code. GraphCodeBERT summarizes binary code semantics by modeling the data flow with the BERT model (guo2020graphcodebert, ). In addition to source code, there are also research efforts on binary code. For example, BinT5 finetunes CodeT5 on decompiled code to generate natural language descriptions (al2023extending, ). Different from existing research, we are the first to perform a large-scale and comprehensive study of different LLMs on the binary code summarization task.

Conclusion

We have presented a large-scale and comprehensive study of how LLMs can understand binary code semantics. We built BinSum, a comprehensive benchmark with an expansive dataset with over 557K binary functions, spanning various code representations, computer architectures, and optimization levels. We construct an extensive binary code summarization dataset, with over 557K binary functions in different binary code representations across different computer architectures and optimization levels. We designed the novel in-context prompt synthesis and optimization techniques for optimal prompt generation. We also devised a new semantic evaluation metric to measure code summaries. Our rigorous evaluations have resulted in 10 empirical results and 6 findings, which provide nuanced insights into both LLMs and binary code, serving as a reference for future research in binary code comprehension.

References

Appendix A Appendix

Table 3 presents the 44 open-source projects that we used to compile binaries.

A.2. Cross-compilation and Stripping

Our compilation and stripping processes are performed on a 64-bit Ubuntu machine. Compiling and stripping binaries with computer architectures different from the host machine, i.e., ARM and MIPS, requires cross-compilers and cross-stripping tools. For this, we have used cross-compilers, arm-linux-gnueabihf-gcc and mipsel-linux-gnu-gcc for ARM and MIPS binaries. For binary stripping, we use the strip command for x86 and x64 binaries. For ARM and MIPS binaries, we leverage arm-linux-gnueabihf-strip and mipsel-linux-gnu-strip.

A.3. Binary Functions with Ground-truth Summaries

Table 4 presents the distribution of binary functions across computer architectures and optimization levels with ground-truth summaries, extracted by methods defined in §2.1.

A.4. Evaluation Metrics for N-gram Matching

In this paper, we report three widely used n-gram matching-based evaluation metrics, i.e., BLEU, METEOR, and ROUGE-L.

Bilingual Evaluation Understudy (BLEU) (papineni2002bleu, ) is a commonly employed metric for assessing the quality of generated texts versus reference texts. It is a variant of precision metrics and quantifies similarity by computing the n-gram precision between a generated summary and a reference summary, with a penalty for excessively short lengths. The BLEU score is computed as follows:

where BP represents the brevity penalty, NN signifies the maximum order of n-grams taken into account, and pn\text{p}_{n} denotes the precision associated with n-grams. In this paper, we report the BLEU-1 score which is calculated on the unigram matching results. We have also calculated BLEU-2, BLEU-3, and BLEU-4, but their scores are 0 for most samples.

METEOR

Metric for Evaluation of Translation with Explicit ORdering (METEOR) (banerjee2005meteor, ) is proposed to improve the measurement of text ordering. In the context of comparing a pair of summaries, METEOR establishes a word alignment between them and subsequently computes similarity scores. Formally, METEOR is calculated by:

where PP and RR are the precision and recall of the mapped unigrams in generated summaries and summary references. ff is the fragmentation coefficient. The penalty parameters α\alpha, β\beta, and γ\gamma come with default values of 0.9, 3.0, and 0.5, respectively.

ROUGE-L

ROUGE-L is a variant of Recall-oriented Understudy for Gisting Evaluation (ROUGE), calculated based on the longest common subsequence (LCS) between the pair of texts. Specifically, ROGUE-L is calculated by:

where PlcsP_{lcs} and RlcsR_{lcs} are the precision and recall of the LCS between generated summaries and summary references. β\beta governs the relative significance of precision and recall, which is commonly set to 1.2.

A.5. Prompt Evaluation Results

Table 5 presents the top 40 prompts and their evaluation results, which are generated by our in-context prompt synthesis and optimization approach (proposed in §2.2). As mentioned in §3, we first generate 320 prompts by GPT-4, and then evaluate each of them on 1000 binary function samples that are randomly selected from our binary summarization dataset. To generate the 320 prompts, we chose to use GPT-4 because these tasks essentially involve test generation, and GPT-4 has shown advanced performance in this area according to OpenAI’s documentation (openai2023gpt, ). Based on the semantic similarity score, we select the best prompt in our subsequent evaluation. In addition to the prompt text, we also added instructions to limit the summary length as discussed in §2.3. To ease the process of retrieving summaries, we get structured summary output by adding the following text to LLM inputs, such as “Input decompiled code:\n¡CODE¿\n Function Summary:” for decompiled code summarization, in which ¡CODE¿ is filled with test decompiled code. For the GPT-4, ChatGPT, Llama 2, and Code Llama models, we use the same prompt for a fair comparison. For BinT5, it is fine-tuned on decompiled code without natural language instructions, therefore, we use the same input form, i.e., decompiled code, as its test input.

A.6. Case Study of LLM-generated Summaries

Following the acquisition of statistical results, we perform a more in-depth study into the extent to which LLMs can comprehend binary code, as well as the nuanced interpretation of specific score outcomes. For this, we perform a series of case studies on LLM-generated summaries. Our focus is on samples in which LLMs attain scores at the highest (100th percentile), median (50th percentile) and lowest (0th percentile) semantic similarity scores. This investigative approach aims to provide us with distributional insights across extreme and median scenarios, thereby shedding light on the robustness of LLM in binary code comprehension. For source code, our overarching observation is that LLM can capture most of its semantics (as presented in Table 6). For example, the summaries of source function 3 that obtain the median score show that the LLM has adeptly discerned the key semantics of “skip characters” and “return the index”. In addition, we observe a similarly good performance from summaries of decompiled code with debugging symbols (such as samples in Table 7). However, for decompiled code from stripped binaries, we find that LLMs cannot generate summaries explicitly matching the semantics of ground-truth summaries (as shown in Table 8). For example, LLMs only manage to capture partial or elementary semantics of functions 2 and 3, despite their scores falling within the median range. For function 2, the generated summary includes the semantics of “file name canonicalization” but it fails to include specific details of the ground truth, such as “remov2 dots”. For function 3, LLM only captures the ELF file format and loses all other semantic information.

Finding 1 (F1\mathcal{F}_{1}) – Although LLMs excel at source code and decompiled code with symbols, the absence of debugging symbols makes LLMs generate summaries with only partial or elementary semantics of decompiled code.

For lower-level code, e.g., IR code and assembly code, we observe that LLMs predominantly focus on describing the low-level operations, without providing comprehensive high-level semantic summaries. For instance, Table 9 presents the summaries and evaluation scores of the IR code. Across samples that obtained the highest to lowest scores, LLM-generated summaries mostly describe the operations of data movement, arithmetic calculations (e.g., addition and subtraction), memory loading/writing, and data flow transformation (e.g., jump operation). Similarly, we also observe such descriptions for assembly code as shown in Table 10. Divergent slightly from the IR code summaries, we find that in the case of the highest score for assembly code, LLM can capture essential information related to the string data structure and the copy operation. Finding 2 (F2\mathcal{F}_{2}) – LLM-generated summaries fail to encapsulate the high-level semantics of IR and assembly code, instead, focusing on elucidating the low-level operations, e.g., data movement and arithmetic calculations.

In the case of raw bytes, one would anticipate that LLMs generate summaries that describe a sequence of bytes. However, it is surprising to observe that LLM-generated summaries depict operations akin to those observed in IR code and assembly code (as shown in Table 11), particularly for OpenAI models, i.e., ChatGPT and GPT-4. For example, the summary of function 2 describes the stack frame operations, function calls, and arithmetic operations, which are not explicitly manifested in the raw byte input. It appears that LLMs generate summaries by implicitly lifting the raw bytes into higher-level representations, such as assembly code that can represent the above-mentioned operations.

Finding 3 (F3\mathcal{F}_{3}) – The LLM-generated summaries for raw bytes closely resemble those for assembly code, hinting at an implicit process of code lifting, performed by the LLMs.

In addition to understanding the performance of LLMs in individual samples, we have also confirmed the benefits of using our proposed semantic evaluation metric. Specificaly, as stated in §1.1, metrics based solely on exact matching fail to provide a precise evaluation outcome, as they cannot encapsulate the semantic nuances of summaries. For instance, for function 2 in Table 6, BLEU and METEOR scores are calculated as 0. However, function 2’s generated and ground truth summaries present close semantics, such as “routine” and “function”, as well as “a filename” and “the name of the file”. These words/phrases are syntactically different, but semantically the same. Moreover, its calculated scores of BLEU, METETOR, and ROUGE-L are lower than those of function 4, while function 2’s summaries have shown closer semantics. In contrast, our semantic evaluation metric computes semantic similarity at the semantic level, thus it generates fair scores that assign a higher score for closer semantics in function 2 compared to function 4. Therefore, we argue that our semantic evaluation metric is more suitable for evaluating binary code summaries.

Finding 4 (F4\mathcal{F}_{4}) – Our semantic evaluation metric can capture the essential semantics of binary code summaries, which is more suitable for our task than the exact matching-based metrics.

A.7. Case Study of GPT-4 Results

In RQ2 evaluations, it is surprising to observe that GPT-4 is not the best model, especially when ChatGPT outclasses GPT-4 on the binaries with symbols. For this, we have conducted a case study to manually investigate the GPT-4 results. Overall, we observe that GPT-4 occasionally focuses more on noisy details rather than capturing the global semantics of binary functions. Table 12 presents three types of such cases, showcasing GPT-4 generated summaries and ground truth along with those generated by ChatGPT for comparison. For function 1, we find that GPT-4 explains the implementation details of the input binary code, instead of providing a concise summary. Such an explanation can capture some semantics, such as its generated phrase “without changing its current position” has similar semantics as the ground truth “preserving the value”, but this explanation includes extra details that do not appear in the ground truth. Similarly, GPT-4 generates additional details for functions 2 and 3 as highlighted in the red text. Compared to GPT-4, ChatGPT focuses less on superfluous details and generates more concise summaries that represent the key semantics.

Finding 5 (F5\mathcal{F}_{5}) – For binaries with symbols, ChatGPT produces more concise summaries and focuses on essential semantics, showing improvement in brevity and clarity over GPT-4.

A.8. BinT5 Outliers and Sample Leakage

For RQ2 evaluation, we have noticed the presence of certain outliers in BinT5’s performance, which attain text similarity scores exceeding 0.8. These scores are notably higher than the median values. By comparing the function names of these outlier samples with BinT5’s training sets, we find that some of our Ghidra-generated decompiled code samples are in BinT5’s training set. We have removed these leaked samples from our reported results of BinT5.

A.9. Impact of Symbols and Exclusive Symbol Stripping

In RQ4 evaluations, we study which symbol type (data types, variable names, and function names) contributes the most to binary function semantics by exclusively stripping symbols of each type. We first treat the decompiled code with all symbols as the original code. Subsequently, to isolate the impact of each symbol type, we eliminate the respective symbols from the original code by substituting them with non-informative symbols. These non-informative symbols follows the same pattern as these in decompiler-generated symbols for stripped binaries. Specifically, for function names, we replace the original names with symbol Fun_addr in which addr is the address of functions. For variable names, we replace the original names with symbol Var_idx where idx is the index of variables in the function. For data types, we replace the original types with the symbol undefined.

A.10. Summary Manipulation by Modifying Function Names

While we have identified the significant semantic contribution of function names for binary code semantics in R9\mathcal{R}_{9}, we also note an intriguing vulnerability in LLM-generated summaries. Specifically, we’ve observed that these summaries can be manipulated by altering function names. To elaborate, Figure 14 showcases a decompiled function generated by Ghidra, with all symbols removed. When we apply our chosen prompt to this function, ChatGPT generates a summary, as demonstrated in row 1 of Table 13. However, the generated summary undergoes a substantial transformation when we make a simple adjustment: changing the original function name, FUN_0000ea01 into alternative names such as quick_sort, print_log, and DNS_flood. As depicted in rows 2 to 4 of Table 13, the resulting binary code summaries closely resemble the semantics of the manipulated function names, but they fail to accurately reflect the true semantics preserved within the function body.

This discovery underscores the susceptibility of LLM-generated summaries to manipulation through function name changes. In light of this vulnerability, adversaries could potentially modify the DWARF entries of malware binaries, thereby misleading LLMs and evading detection of malicious behavior.

Finding 6 (F6\mathcal{F}_{6}) – Summaries generated by ChatGPT are susceptible to manipulation via alterations in decompiled code function names.

A.11. Other Prompt Engineering Techniques

We also examine whether the other prompt engineering techniques can help LLMs understand binary code. Specifically, we focus on two popular techniques, including few-shot learning and chain-of-thought prompting. Few-shot prompting was initially introduced to enhance the generalizability of GPT-3 for tasks beyond its primary domain (brown2020language, ). This is achieved by imparting the model with contextual knowledge through in-context instructions and demonstration samples. To be more specific, few-shot learning entails the inclusion of a system instruction I\mathcal{I} along with either nn in-context demonstration pairs (Xd\mathcal{X}_{d}, Yd\mathcal{Y}_{d}), where Xd\mathcal{X}_{d} and Yd\mathcal{Y}_{d} represent sample inputs and outputs. This augmentation occurs prior to the introduction of the test input:

where the system instruction I\mathcal{I} and demonstration samples (xdi,ydi)(x_{d}^{i},y_{d}^{i}) are concatenated with the test input xtx_{t} to form the query input xx. 15(b) presents the example prompt that we used to perform few-shot prompting. Instead of using fixed demonstration examples, we randomly sample two pairs of example binary functions and summaries for each test binary code to avoid LLMs learning the unchanged pattern in the demonstration examples.

Chain-of-thought prompting is proposed to resolve complex reasoning problems by breaking them down into intermediate steps (wei2022chain, ). In the context of binary code understanding, we ask LLMs to perform an additional step of reasoning the input code semantics before generating the final code summaries. For this, we first send one request asking LLMs to reason about the code semantics, in which we use the popular reasoning prompt “Let’s think step by step”. Upon receiving the response, we parse the generated code reasoning results and concatenate them into the second request inquiring about the code summaries. 15(c) presents the sample of our chat-of-thought prompts. The baseline is the zero-shot prompt. As exemplified in 15(a), we directly concatenate the prompt with the test code to generate the request as the baseline approach.