Checking Smart Contracts with Structural Code Embedding
Zhipeng Gao, Lingxiao Jiang, Xin Xia, David Lo, John Grundy
Introduction
A Smart Contract, a term coined by Nick Szabo in 1994 , is a program that can be triggered to execute any task when specifically predefined conditions are satisfied. The conditions defined in smart contracts, and the execution of the contracts, are supposed to be trackable and irreversible in such a way that minimizes the need for trusted intermediaries. They are also supposed to minimize either malicious or accidental exceptions in order to ensure trustworthiness of any business transactions implied by the smart contracts.
In recent years, along with widely-deployed cryptocurrencies (e.g., Bitcoin, Ethereum, and many others) on distributed ledgers (a.k.a., blockchains), smart contracts have obtained much attention and have been applied to many business domains to enable more efficient and trustable transactions. The overall market capitalization of cryptocurrencies is more than 200 billions in USD as of August 2018 . Many crytocurrencies involve various kinds of smart contracts, and a smart contract in the blockchains often involves cryptocurrencies worthy of millions of USD (e.g., DAO , Parity and many more). This gives much incentive to hackers for discovering and exploiting potential problems in smart contracts, and there is a very significant need to check and ensure the robustness of smart contracts.
Even though there have been many studies on the characteristics of bugs in smart contracts and underlying blockchain systems (e.g., ) and detection of smart contract bugs (e.g., ), there are still increasing needs to detect and prevent more and more kinds of problems identified in smart contracts. A major disadvantage of these existing bug detection tools is that they require certain bug patterns or specification rules defined by human experts in order to construct bug detectors and/or code model checkers to check smart contracts against the defined rules. With the high stakes in smart contracts and race between attackers and defenders, it can be far too slow and costly to write new rules and construct new checkers in response to new bugs and exploits created by attackers.
In this paper, we propose a new approach that addresses the above issue. We aim to enable efficient checking of smart contracts and can evolve checking rules along with the evolution of code and/or bugs, based on our deep learning model for smart contracts. The main idea of our approach is two fold: (1) code and bug patterns, including their lexical, syntactical, and even some semantic information, can be automatically encoded into numerical vectors via techniques adapted from word embeddings (e.g., ) enhanced with basic program analyses and the availability of many smart contracts; (2) code checking can be essentially done through similarity checking among the numerical vectors representing various kinds of code elements of various levels of granularity in smart contracts. This idea, with suitable concrete code embedding and similarity checking techniques, can be general enough to be applied for various code debugging and maintenance tasks. These include repetitive (a.k.a. duplicate or cloned) contract detection, detection of specific kinds of bugs in a large contract corpus, or validation of a contract against a set of known bugs“Validation” in this paper is to check if a contract has no bug similar to the known bugs; it does not mean formal verification of the contract..
We have built a prototype based on the idea, named SmartEmbed, for smart contracts written in the Solidity programming language used in the Ethereum blockchain . We have collected 22,725 contracts in their Solidity source code that are labelled as “verified” in the Ethereum blockchain and 17 well-known buggy contracts from the Internet. Our tool can then automatically generate the vector embeddings from the contract code collected from the blockchain and provides a mechanism to compose vector embeddings for any code fragment, either buggy or correct. All of these vectors then go through similarity checking for different purposes. Our evaluation results against 22,725 contracts show that, for the tasks of clone detection, bug detection, and contract validation, our approach can achieve comparable results compared with specific tools such as Deckard, SmartCheck.
The main contributions of this paper are as follows:
We propose a new approach for Solidity code checking based on code embedding and similarity checking, which is applicable for various purposes, such as similar contract code detection, bug detection, and contract validation.
We built a prototype SmartEmbed based on the approach, and evaluated it on more than 22,000 Solidity contracts collected from the Ethereum blockchain.
Our clone detection results show that our tool can effectively identify many repetitive Solidity code where the clone ratio is around 90%, and we can detect more semantic clones accurately than the commonly used clone detection tool Deckard.
Our bug detection results show that SmartEmbed can identify more than 1,000 clone related bugs based on our bug databases efficiently and accurately, which can enable efficient checking of smart contracts with changing code and bug patterns. For contract validation, our approach can capture bugs similar to known ones with low false positive rates, the query for a clone or a bug is quite efficient which can be sufficient for practical uses.
This paper is organized as follows. Section 2 presents related work on smart contract security and relevant techniques. Section 3 presents our approach for smart contract code embedding. Section 4 evaluates our approach on actual contracts collected from the Ethereum blockchain. Section 6 discusses limitations of our approach and its evaluation. Section 7 concludes the paper.
Related Work
Despite the fact that Ethereum and smart contracts are relatively new, many studies have been performed on security aspects of smart contracts. Some studies focus on creating taxonomies of smart contract security vulnerabilities (e.g., ). Others focus on specific bug detection. For example, Loi et al. build a symbolic execution tool called OYENTE to detect four kinds of security bugs. Tikhomirov et al. build a static analysis tool called SmartCheck to automatically check for vulnerabilities and code smells. Brown et al. present a framework for analyzing runtime safety and functional correctness of smart contracts via formal verification; several types of vulnerability, such as reentrancy and exception disorders, can be identified by their tool. Chen et al. developed a security tool for identifying gas costly programming patterns in smart contracts.
Although the aforementioned research has proposed security analysis tools to find bugs in smart contracts, most of those tools are built to discover specific types of potential vulnerabilities, requiring manually constructed bug patterns or specifications. To the best of our knowledge, no one has yet considered how to make such tools more flexible and adaptive to arbitrary new bugs by using word embedding for smart contract code. Our work is the first to propose an approach for detecting smart contract bugs and validating contracts via similarity checking of contract code embeddings, especially the embeddings that take code structures into consideration.
2 Word Embedding and Code Similarity
Embedding (also known as distributed representation ) is a technique for learning vector representations of entities such as words, sentences and images. One of the typical embedding technique is word embedding, which represents each word as a fixed-size vector, so that similar words are close to one another in the vector space .
Recently, an interesting direction in software engineering is to use deep learning to compute and use vector representations of programs. For example, Mou et al. propose to learn vector representations of source code. They map the nodes of abstract syntax trees to vectors. Following their previous work, Mou et al. propose a tree-based convolutional neural network based on program abstract syntax trees to detect similar source code snippets. Ye et al. embed words into vector representations to score a pair of documents, and use StackOverflow questions and answers as document corpora to train word embeddings. White et al. propose an automatic program repair approach, DeepRepair, which leverages a deep learning model to identify similarity between code snippets.
Different from these existing tools, our code embedding methods are based on serialization of solidity parse tree for different level program elements. To the best of our knowledge, our work is the first to apply the code embeddings to the specific domain of Ethereum smart contracts as inspired by the promising results of employing deep learning to the many other software engineering tasks (e.g., ).
3 Clone Detection, Bug Detection, and Code Validation
A plethora of approaches have been investigated for different tasks such as code clone detection, bug detection, and code validation and/or program verification. All of the tasks can be viewed as variants of the problem of finding “similar” code, depending on the definition of similarity: code clone detection is to search for code in a code base “similar” to a given piece of code; bug detection is to search for code in a code base “similar” to a known bug; and code validation is to search for (non-existence of) code in a code base “similar” to any bug. As our approach based on code embedding and similarity checking is an instantiation of this general view, it is related to many such studies too.
For clone detection, many techniques in the literature generally begin by generating some intermediate representations for code before measuring similarity. According to source code representation, these techniques can be classified as text-based (e.g., ), token-based (e.g., ), tree-based (e.g., ), graph-based (e.g., ), semantic-based (e.g., ), deep-learning-based (e.g., ), or a mixture. Our approach complements those studies by applying word embedding to smart contract code and its syntax structures to search for smart contracts of various levels of granularity.
For bug detection, there also exists many conventional techniques tailored for smart contracts, such as those based on static analysis and model checking (e.g., SmartCheck , Securify ), symbolic execution and dynamic analysis (e.g., Oyente ), Manticore ), and a mix of techniques (e.g., Mythril ). “Conventional” here refers to the fact that they require human curated correctness and/or bug patterns or specifications in order to check whether the code complies with or violates the given patterns or specifications.
There are other bug detection techniques that do not require predefined bug patterns or specifications; instead, they often rely on statistically inconsistencies among multiple instances of code. For example, Juergens et al. report that inconsistencies among similar code are an important source of bugs in programs, and every second (possibly inconsistent) modification of a piece of similar code increases the chance of errors. This phenomenon has been explored in the literature to detect clone-related bugs (e.g., ), code porting errors (e.g., ), semantic bugs (e.g., ), etc.
Another category of bug detection techniques depending on historical known bugs is more similar to our approach. Those approaches learn patterns from known bugs using various techniques (e.g., graph pattern matching and heuristic rule matching ) and search for similar instances in a given code base. Recently, such techniques that require little or zero efforts in manually written specifications are often based on deep learning (e.g., ).
Our approach is relying on the existence of known bugs, as it automatically learns code and bug representations from known bugs based on code embedding. It is unsupervised; there is no need to handcraft features beforehand, which saves much manual effort in feature selection needed for many other techniques. Given a sufficiently comprehensive set of code and known bugs, our approach can potentially be applicable for both bug detection and contract code validation. On the downside, our “bug detection” and “contract validation” are both evaluated with respect to the known bugs: bug detection is to detect all instances of the known kinds of bugs in a large contract corpus; contract validation is to check if a contract is free of any instance of bugs similar to the known bugs. If no enough known bugs are available, our approach can utilize potential bugs reported by conventional techniques too, providing a complementary way to make bug detection and contract validation more comprehensive.
Approach
Fig.1 demonstrates the overall framework of SmartEmbed. Based on similarity checking and code embeddings, SmartEmbed is targeting three tasks: clone detection, bug detection, and contract validation. For clone detection and bug detection, we aim to identify code clones and clone-related bugs for smart contracts in the existing Ethereum blockchain. For contract validation, given a new smart contract, SmartEmbed will help to validate whether it contains vulnerable statements associated with our bug database.
To be more specific, the collected source code of smart contracts are loaded and parsed by our custom built parser, generating the abstract syntax trees (ASTs) for a smart contract. Then, we extracted a stream of tokens by serializing the ASTs. Following that, the normalizer reassembles the token stream to eliminate the differences (e.g., the stop word, values of constants or literals) between smart contracts. The result sequence that is output by the normalizer is then fed into our code representation learning sub-model. Through the model building and training, each code fragment would be embedded by a fixed-length dimension vector. All of the source code will be encoded into the code embedding matrix. In the meanwhile, all vulnerable source code would be embedded into the bug embedding matrix.
Next, clone detection, bug detection and contract validation are performed using similarity checking methods via vector space comparison. Similarity comparison is performed between the possible code snippet pairs, and a similarity threshold governs whether code fragments will be considered as code clones or clone-related bugs.
In following sub-sections, we elaborate our data collection, parsing, normalization, embedding learning, and similarity checking steps.
To prepare the smart contract code used for our approach and evaluation, firstly we collected Solidity smart contracts using EtherScanhttps://etherscan.io/, which is a block explorer and analytics platform for Ethereum. To be more specific, we built our own web scrapers to systematically search and download every HTML page on the entire site. After parsing HTML output from that page, needed information (e.g. contract address/source code/byte code/opcodes) were extracted from the HTML file for our further assessment.
By April 20, 2018 when we started our evaluation experiments, we had collected 22,725 verified smart contract. We counted the number of individual contracts (given the source code of a smart contract, there may include several individual contracts), functions, statements, and lines associated with these smart contracts. On average, each smart contract involves around 6 individual contracts, 27 functions, 85 statements, and 323 lines of code. Table I describes the statistics of our collected dataset.
2 Parsing
The abstract syntax tree (AST) is a structural representation of a program. In this step, for each smart contract, we used a custom-built Solidity parser to parse the smart contract into an AST. We built our code embeddings based on AST because its tree structured nature provides opportunities to capture structural information of programs.
More specifically, ANTLR and a custom Solidity grammar were used to generate the XML parse tree as an intermediate code representation. The source code was fully translated to this internal tree representation. After that, we built the code embeddings based on this abstract syntax tree. Listing 1 and Fig. 2 provides a simple example of a smart contract and its corresponding AST, defined in Solidity.
We serialized the parse tree of a smart contract differently for contract-level, function-level and statement-level program elements, depending on the types of the tree nodes that contain or are siblings of the relevant elements. The high level idea of such a processing is to capture the structural information (e.g., branch and loop conditions) in and around the focal elements. Further, non-trivial tokens and identifier names are processed and put into the code element sequences serialized from the trees, so that certain data flow information (via defining/using a same name) is added into the sequences too. We describe the details of the tokenization process below with the aforementioned sample Solidity code.
Contract Level Tokenization: We extracted all terminal tokens from the XML parse tree by performing an in-order traversal. Regarding the previous smart contract, the following tokens were extracted (1_10 stands for the line range of this contract).
Function Level Tokenization: Considering the function level tokenization, we appended the contract signature to the end of function tokens. For the previous smart contract, function level tokenization’s result was given as follows (6_9 represents this function starts at line 6 and ends at line 9).
Statement Level Tokenization: Different from the contract-level and function-level tokenization, for statement-level tokenization, based on the terminal tokens, we added more details of structural and semantic relations. For example, regarding the previous smart contract, structural information such as the chain of ancestors in ASTs as well as function signatures were retrieved from the XML parse tree. By adding the chain of ancestors in ASTs, our model can capture the structural relationship; by adding the diverse neighbourhood nodes, our model can capture the “context” information of a focal element.
Our parse tree based serialization of the code with respect to a focal element captures most structural (containment and neighbouring) and some semantic (data-flow) information, which serves the downstream applications.
3 Normalization
An important task during preprocessing is normalization. In this step, we normalized the token sequence to remove some semantic-irrelevant information. To be more specific, the following steps have been taken:
Stop words : For single-character variables, such as “i”, “j”, “a”, “b”, “k”, etc., we replaced them with “SimpleVar”. The below code snippet illustrates this step :
The main idea of our approach is based on code embedding and similarity checking for various similarity-based software engineering tasks. Herein, we evaluate how well our approach embeds code and checks similarity for the purposes of contract code clone detection, bug detection, and contract validation.
As we have introduced in previous sections, representation learning maps a symbol to a real-valued, distributed vector. the basic criterion of code embedding is that similar symbols should have similar representations. In particular, symbols that are similar in some aspects should have similar values in corresponding feature dimensions. To demonstrate the effectiveness of our code embedding, we pick top 100 frequent tokens, then draw the embeddings for the tokens on a 2D plot using T-SNE algorithm, which are shown in Fig. 3. Similar words that are close together in the vector space and are expected to be close in the 2D plot as well.
From the figure, we note that tokens sharing similar syntactic and lexical meaning are clustered together. For example, operators such as “”, “”, “”, “”, “”, “” are grouped together, and tokens such as “args”, “dynargs”, and “StringLiteral”, “decimals” are close to each other. This gives us confidence that high dimensional code representation can meaningfully capture co-occurrence statistics and distributed semantics for the tokens.
2 Similarity Checking Evaluation
To demonstrate the effectiveness of the similarity checking, we evaluate our approach with respect to three tasks: code clone detection, bug detection, and contract validation; and we compare the results with the following tools designed specifically for those tasks.
Deckard : a scalable, tree-based tool for source code clone detection. It has been widely used and extended to support the Solidity language, and we can compare with it on smart contract code clone detection.
SmartCheck : an extensive static analysis tool that can detect many kinds of vulnerabilities in smart contracts automatically. It works on Solidity source code, and has been shown to outperform many other tools in terms of bugs detected. Hence in our study, we choose SmartCheck to compare the performance of our approach in detecting bugs and validating contracts.
In the following sections, we aim to answer the following six key research questions:
RQ-1: How effective is our SmartEmbed for detecting code clones within smart contracts?
RQ-2: How effective is SmartEmbed for bug detection in smart contracts?
RQ-3: How effective is SmartEmbed for distinguishing the bug fixes from the bugs?
RQ-4: How effective is the structural and semantic information added to SmartEmbed?
RQ-5: How effective is SmartEmbed for smart contract validation?
3 RQ-1: Clone Detection Evaluation
Code clones are common in software and can be considered useful or harmful depending on different circumstances. They can appear more frequently in smart contracts than traditional software as smart contracts are irreversible and often intended to be self-contained, containing all the code implementing needed functionalities with little reference to other contracts. Maintaining smart contracts and managing duplications, redundancies, and inconsistencies are very important for contract quality assurance, and the detection of contract code clones is an important first step. The nature of the task is similarity based and very suitable for our approach.
Code clone detection is done through the vector space comparison via similarity checking, which is described in Section 3. A similarity threshold governs whether two code fragments are viewed as clones. We evaluate the code clone detection at the contract level, function level as well as the statement level by using our approach.
Contract-level clone detection: As mentioned in Section 3, each smart contract can be represented by a fixed dimensional vector. We construct a pairwise similarity matrix (in our case, M would be a 22718 22718 matrix, we removed 7 parsing error cases here), where each row and column corresponds to a smart contract, and each cell corresponds to the similarity score between smart contract and . Given a similarity threshold , if , the corresponding smart contract and would be considered as a clone pair.
Function-level clone detection: Theoretically we could also construct a pairwise similarity matrix the same as the above, for all functions. However, due to the large number of functions, which was 631261, the complexity of computing the pairwise similarity between every pair of functions directly is too expensive. Hence in this evaluation, we randomly sample 200 smart contracts from our repository and use the functions in the 200 contracts, which contain 5307 functions in total, as clone queries. Following that, a pairwise similarity matrix between the sampled 5307 functions and all of the functions in the whole contract set is generated (i.e., was a 5307 631261 matrix), where each cell represented the similarity score between the sampled function and the function . Same as the above, the associated functions and will be considered as a clone pair if .
Statement-level clone detection: Same with function-level clone detection, since it is too expensive to calculate the pairwise similarity between every pair of statements directly, we extract all the statements within the aforementioned 200 sampled contracts, which contain 16,350 statements in total. Following that, we construct a pairwise similarity matrix between the sampled 16,350 statements and all of the statements in the whole contract set (i.e., was a 16,350 1,944,513 matrix), where each cell represents the similarity score between the sampled statement and the statement . Same as the above, the associated statements and will be considered as a clone pair if .
3.2 Experimental Results
To justify our approach on the task of code clone detection, we compare our results with those of Deckard (with its default settings) by the numbers of lines of code that are detected as clones. We set the similarity threshold to 1.0 and 0.95 for Deckard and SmartEmbed respectively.The definitions of similarity used in SmartEmbed and Deckard are not exactly the same: SmartEmbed is based on the embedding vectors (cf. Section LABEL:sec:similarity); Deckard is based on tree structures. However, we simply assume the two are approximate of each other and treat them the same for easier comparison. The results are summarized in Table II. From the table, we can observe the following points.
There is a very high ratio of code clones among smart contracts. By using Deckard with its default settings with similarity threshold 1.0, the code clones may involve more than 6.6 million lines of code, while the total lines in 22725 contracts are just 7.3 million, which means more than 90% smart contracts on Ethereum are somehow cloned from others. The code clone ratio is even higher (more than 96%) if we set the similarity threshold to 0.95. Since SmartEmbed can detect code clones on contract-level, function-level and statement-level, we exclude the clone fragments in Deckard results that are smaller than a contract, function and statement respectively for a fair comparison. The clone ratios on both function-level and statement-level are consistent with the original clone ratio. We note that clone ratio drops at contract-level, this is because we just keep the results if the whole contract is a clone, removing all the non-whole contract clones.
SmartEmbed report less clones overall than Deckard on different levels of granularity and similarity thresholds. Regarding the SmartEmbed results, the clone ratio was 0.39 and 0.85 at the contract-level with respect to similarity threshold 1.0 and 0.95 respectively. At the function-level, as mentioned in the previous subsection, we randomly sample 200 contracts which include 5,307 functions, involving 27,945 lines of code int total. SmartEmbed detected 23,087 (85%) and 24,640 (91%) of them as clones with similarity threshold 1.0 and 0.95 respectively. Consistent with the function-level clone results, the clone ratio was 0.82 and 0.93 at statement-level with respect to the similarity threshold 1.0 and 0.95 respectively. We argue that the main reason for this phenomenon is that SmartEmbed is more precise than Deckard in detecting clones, this is because SmartEmbed encodes both structural and some contextual semantic information, while Deckard only considers structural information. So, SmartEmbed should have more constraints and detect less clones.
Most code clones detected by SmartEmbed are also detected by Deckard. To evaluate the quality of code clones reported by our approach, we count the numbers of lines of code in our results that overlap with clones reported by Deckard (assuming Deckard’s results are accurate), the results are summariized in Table III (for both 1.0 and 0.95 similarity) and the Venn diagrams in Fig. 4, Fig. 5 and Fig. 6 for the contract-level, function-level and statement-level respectively. We note that the overlap ratio is more stable at function-level, reflecting that SmartEmbed is better in finding functional clones while tolerating non-essential syntactic differences.
Regarding the relatively high clone ratio in smart contracts, we consider that the following reasons can be responsible for introducing clones:
One of the main reasons for introducing clones in smart contracts is the irreversibility of smart contracts stored in the Ethereum blockchain. Even when the same contract creator may want to evolve the contract code and create new versions of the smart contracts, the older versions are still kept visible in the blockchain. We consider such a scenario, and recount all the clones by creator addresses (i.e., if the detected clones are code belonging to a same creator, we do not report them), such clone results still report a considerable high clone ratio 51% for similarity threshold 0.95 on contract level, reflecting the fact that cloning contracts across different creators is more common than usual software.
ERC20 is the main technical standards for the implementation of tokens. The standardization allows contracts to operate on different tokens seamlessly, thus boosting interoperability between smart contracts. From the implementation perspective, ERC20 are interfaces defining a set of functions and events, such as totalSupply(), balanceOf(address owner), transfer(address to, uint value). For every contract in our database, if the contract has implemented all the interfaces required by ERC20, it will be considered as an ERC20 contract. Finally, we find that 15,514 out of 22,725 (68.3%) contracts contain the code blocks to support compliance to the ERC20 standard, reflecting that template contracts also plays an important role to cloning in Ethereum.
The experimental results reveals homogeneous of the Ethereum ecosystem. Our clone detection results can benefit the smart contract community as well as individual Solidity developers in the following ways:
The relatively high ratio of code clones in smart contracts may cause severe threats, such as security attacks, resource wastage, etc. Finding such clones can enable significant applications such as vulnerability discovery (clone-related bugs) and deployment optimization (reduce contract size and duplication), hence contribute to the overall health of the Ethereum ecosystem.
Our work in identifying clones can also help Solidity developers to check for plagiarism in smart contracts, which may cause a huge financial loss to the original contract creator.
3.3 Examples of clone detection
To compare the results of SmartEmbed and Deckard, we have manually checked the clones detected by SmartEmbed but not by Deckard. A sample code pair is shown in Fig. 7 and Fig. 8. The code pair has similar statements but some statements are added and modified, which can be considered as a type-III or even type-IV semantic clones and are hard for Deckard to detect as it was designed for syntactic clones.
We also manually checked the code clone pairs detected by Deckard but not by SmartEmbed. A sample code pair is shown in Fig. 9 and Fig. 10. Even though these two pieces of code are both functions about “addCompany”, since they use different data structures, they are not considered as syntactic clones. This is because Deckard ignores the different identifier names in the code, which results in detecting this clone by accident. Regarding SmartEmbed, it maintains these differences in identifier names, which increases the differences between associated code embedding vectors. This further justifies that SmartEmbed is more precise in clone detection than Deckard.
Answer to RQ-1: How effective is our SmartEmbed for detecting code clones within smart contracts? - we conclude that SmartEmbed is highly effective.
4 RQ-2: Bug Detection Evaluation
To quickly duplicate some functionality, programmers usually copy and paste code, which can introduce clone-related bugs into programs. It is also folklore that programmers often repeat similar bugs. Such intuitions give the basis for similarity-based bug detection using our approach. To pinpoint a bug accurately, we perform bug detection at the statement level of granularity. That is, for a given known buggy statement (simply called a bug), every statement in our code base whose similarity with respect to the bug exceeds a specific threshold is reported as a potential bug. As shown in the evaluation results later, compared with other analysis-based approach, our similarity-based approach can detect bugs similar to known ones across a large set of programs more efficiently and accurately, while analysis-based approach may detect more bugs in individual programs.
To detect bugs, we need to collect some known buggy statements to construct the bug database. Although there are many contracts in the wild reported to be vulnerable (e.g., ), there is a lack of a comprehensive list of references to pinpoint buggy statements in those contracts. We collected a list of 52 known buggy smart contracts belonging to 10 kinds of common vulnerabilities. These vulnerabilities are from real world events (e.g., Reentracy, Honeypot, Replay, Gas Limit) , previous research papers (e.g., Overflow/Underflow, Blockhash/Timestamp) and/or the CVE reported by some organizations (e.g., Transfer Flaw, Batch Overflow, Verify Reverse) .
We then tried our best to pinpoint buggy statements in those contracts by inspecting research papers, web articles, and community discussions. A list of vulnerable smart contracts and their vulnerabilities are summarized in Table IV. For each vulnerable smart contract in the table, one or more associated buggy lines are identified. We divide the 52 vulnerable smart contracts into two groups: 32 smart contracts marked with * are used for the bug detection evaluation, the other 20 are saved for the contract validation evaluation later. For the bug detection evaluation, 63 buggy statements are collected from the 32 vulnerable contracts. We create our bug database from the 63 buggy statements by using code embedding described in Section 3. That is, for each buggy statement, we compose a numerical vector by summing up the vectors for all relevant tokens in the statement. Each statement is thus mapped to a vector of 150 dimensions. Since we have 63 buggy statements, a bug embedding matrix is constructed and serves as our bug database.
The setting for bug detection herein is that, for each buggy statement embedding in our bug database (simply called a bug), we need to identify every possible statement that is in the set of all statements in the contracts we collect from the Ethereum blockchain and similar to the given bug. Given a similarity threshold , if the similarity score estimated between and is over , then will be reported as a potential bug similar to . We perform such bug detection to report bug candidates for every bug in our bug database. Following that, we validate each candidate bug to see whether it involves an actual bug or not by manually checking. To be more specific, we compare bug candidate lines reported by our approach with the real bug lines, the candidate bugs will be validated if one of the following conditions was satisfied:
The bug statements contain the exact identical code fragments same as the real bugs, which can be considered as type-I clone-related bugs.
The bug candidates involve syntactically equivalent fragments as real bugs, with some variations in identifiers, literals or types, which can be viewed as type-II clone-related bugs. A sample pair is shown in Fig. 11 and Fig. 12.
The candidate bug lines involve syntactically similar code with inserted, deleted or updated statements, which can be considered as type-III or type-IV clone-related bugs. A sample pair is shown in Fig. 13 and Fig. 14.
If the bug candidate is an actual clone-related bug, then it is counted as validated in Table V and Table VI. To demonstrate the advantages of SmartEmbed in clone-related bug detection, we also compare it with the detection results of SmartCheck.
4.2 Experimental Results
For different types of clones, the bug detection results of SmartEmbed are summarized in Table V. By setting the similarity threshold to 0.90, we count the number of reported bugs as well as validated bugs with respect to each clone type (i.e., type-I, type-II, type-III/type-IV). If the bug candidate does not belong to any of these clone types, it is identified as Not-Clones. From the table, we can observe the following points.
Most of the bug candidates reported by SmartEmbed are Type-II clones. This reflects that solidity developers do introduce the clone-related bugs by copying and pasting source code from somewhere else.
SmartEmbed can achieve 100% precision for detecting Type-I and Type-II clone-related bugs. This is because Type-I and Type-II clones do not involve structural changes and can be easily identified.
The performance of SmartEmbed drops for detecting the Type-III/IV clones. To identify the Type-III/IV clone-related bugs, we need to decrease the similarity threshold, which may also introduce more false positive cases at the same time.
The bug detection results of SmartEmbed with respect to different similarity threshold are summarized in Table VI. For each specific similarity threshold in the table, we show the number of reported bug candidates (i.e., the number of statements in our set of contracts that have a similarity higher than to some bug in our bug database), and the number of bugs validated by manual checking together with the precision. From Table VI, we can see that:
The precision of SmartEmbed increases as the similarity threshold increases. For thresholds higher than 0.96, SmartEmbed can have a 100% precision.
The lower the is, the more statements may be reported as potential bugs. When the similarity threshold is set to 0.91, SmartEmbed reports 1,052 statements as potential bugs, while maintaining a high precision of 95%.
When the similarity threshold is set to 0.90, SmartEmbed reports 1,311 potential bugs, 1,163 of them are validated as real bugs. The precision of SmartEmbed drops to 88.7%. This is reasonable because smaller similarity threshold will bring in more noises and hence incur more challenges for detecting clone related bugs. It also signals that setting the similarity threshold between 0.90 and 0.91 may be a good choice for the bug detection task.
Since it is too expensive to run SmartCheck on all the 20k+ contracts, we only run it on the manually validated contracts associated with the 1,163 statements. SmartCheck automatically checks a given contract for predefined vulnerability patterns and highlights the lines of code containing the vulnerabilities. For a fair comparison, we limit SmartCheck to the bug patterns we collected in Table IV. SmartCheck only reported 697 out of 1163 statements as bugs, which shows the advantage of our approach in detecting clone-related bugs.
4.3 Examples of bug detection
We manually checked some bugs reported by SmartEmbed but not by SmartCheck. Some types of bugs, such as “Honeypots” in Table IV can not be effectively checked by SmartCheck.
For example, the function multiplicate() above is the only function that does allow a call from anyone other than the owner. It looks like by sending a value higher than the current balance of the contract it is possible to withdraw the full balance from the contract. Both statements in line 7 and 9 try to reinforce the idea that this.balance is somehow credited after the function is finished. However, this is a trap since the this.balance is automatically updated before the multiplicate() function is called. So if(msg.value>=this.balance) is never true unless this.balance is initially zero.
Encoding such a bug type into tools like SmartCheck would require extra efforts in defining the bug specification, while our approach can just take the sample bug and automatically generate embeddings to recognize similar bugs. Of course, this advantage of our approach relies on good embeddding of all relevant structural and semantic information of code, which will be a continuing research direction in the future.
Answer to RQ-2: How effective is SmartEmbed for bug detection in smart contracts? - we conclude that SmartEmbed is very effective for clone-related bug detection in a large set of smart contracts.
5 RQ-3: Practical Analysis
Considering the cloning rate in Ethereum is remarkably higher than the traditional software, a key problem with code cloning is that the original piece of code should ideally be fixed in every copy of its later versions. Herein we perform a practical analysis to verify whether SmartEmbed can distinguish bug fixes from the original buggy statement.
Because the code file of deployed contracts is immutable, hence when a bug is identified in a smart contract, the developer should deploy a fixed version to the Ethereum blockchain. For each buggy smart contract in our bug database, we manually investigated the contract creation history of the contract creator to see if there is a fixed version contract for the specific buggy statement. Finally we found that 5 out of 52 buggy smart contracts include a fixed version. We pinpointed the fixed statement and estimated the similarity score between the buggy statement and its corresponding fixed statement.
5.2 Experimental Results
The practical analysis results of SmartEmbed are summarized in Table VII. A similarity score is calculated between the buggy statement and its corresponding fixed statement. From the table, we can see that:
By setting the similarity threshold to 0.90, all the fixed smart contracts can be correctly identified by SmartEmbed as not vulnerable. Even though the original version and fixed version are very similar, SmartEmbed can effectively identify the real clone-related bugs and neglect those fixed ones. This is because SmartEmbed focuses on statement-level for bug detection, any small fixes within the buggy statement will result in different code embedding vectors, which will also reduce the similarity scores.
There is a significant drop of similarity scores between the fixed version contracts and the original ones. This further justifies the ability of SmartEmbed to separate the real buggy statement and fixed statement.
5.3 Bug and Bug Fix Examples for Practical Analysis
We show a pair of original buggy statement and its corresponding fixed statement in Fig. 15 and Fig. 16. As illustrated in Fig. 15, the function batchTransfer() makes multiple transactions simultaneously. By passing several transferring addresses and amounts by the caller, the function would conduct some checks then transfer tokens by modifying balances. However, overflow might occur in line 193, uint256 amount = uint256(cnt) * _value, if _value is a huge number. It will make amount become a small value rather than cnt times of _value, then transfers out tokens exceeding balances[msg.sender]. For the fixed version of batchTransfer() function in Fig. 16, the buggy statement is updated to uint256 amount = _value. mul(uint256(cnt)), herein, the contract creator compute the multiplication by using secure mathematical operations such SafeMath. The change in the buggy statement as well as the function signatures reduce the similarity score between the buggy statement and the fixed statement.
Answer to RQ-3: How effective is SmartEmbed for distinguishing the bug fixes from the bugs? - we conclude that SmartEmbed is very effective for distinguishing the bug fixes from the clone-related bugs.
6 RQ-4: Ablation Analysis
When we perform the bug detection, one main novelty of SmartEmbed is adding details of structural (containment and neighbouring) and semantic (data-flow) information based on our serialization of parse trees. For example, we added the chain of ancestors in ASTs to capture sequence derivations and function signatures to capture the diverse neighbourhood relations of nodes. As shown in Section 4.4, this tree-based embedding technique is quite accurate and effective for bug detection in a large set of smart contracts. To verify the effectiveness of the structural and semantic information added to SmartEmbed, we perform an ablation analysis with respect to the bug detection task.
For the ablation analysis, we compare SmartEmbed with one of its incomplete variants, named BasicEmbed. Different from SmartEmbed, BasicEmbed removes all the structural and semantic relations from the statement tokenization results, and only keeps the simple statement token sequence. By going through the same steps of normalization, code embedding learning and embedding matrix building process, we can construct a new code embedding model for BasicEmbed. Following that, for each bug statement in Table IV, we apply BasicEmbed to the bug detection task via similarity checking.
6.2 Experimental Results
The bug detection results of BasicEmbed and SmartEmbed are summarized in Table VIII. Due to the very large number of bugs reported by BasicEmbed, which is more than 30k+, manually validating all these potential bugs is too expensive. Herein this evaluation, we randomly sampled 300 contracts and validated these contracts manually. From the table, we have the following observations.
The total number of bugs reported by BasicEmbed is very large, which is over 30k. At the same time, the overall precision of BasicEmbed is only around 5%, which means the majority of the bugs reported by BasicEmbed are false positives. This also reflects that by simply extracting the token sequence of the statement is not accurate enough for the bug detection task.
Regarding the precision of different similarity thresholds, SmartEmbed stably and substantially outperforms BasicEmbed, which reflects that the structural and semantic information have a major influence on the overall performance. This verifies the effectiveness and necessity of adding structural and context information based on parse trees.
87% of the bugs reported by BasicEmbed have a similarity threshold of 1.0, which means most of the bugs reported by BasicEmbed are type-I clone-related bugs. This is because without considering the context of the statement, code clones with respect to a single buggy statement can be easily identified in other smart contracts. It further supports our claims that the structural and semantic relations convey much valuable information.
6.3 Bug Detection Example for the Ablation Analysis
We manually checked some buggy statements that have a large number of clones reported by BasicEmbed. For example, BasicEmbed reported 10,679 potential bugs with respect to the following buggy smart contract.
The function above name DynamicPyramid should be Rubixi. The wrong name gives permissions to anyone to invoke the DynamicPyramid function to become the owner of the contract and withdraw fees from it. If the function had the same name as the contract Rubixi, then the Ethereum virtual machine would automatically block access from anyone except the contract creator. This bug happened at some point of time during the development of the contract: the contract name was changed from DynamicPyramid into Rubixi, but the programmers forgot to change the name of the constructor accordingly.
The buggy statement of this smart contract is pinpointed at line 5, which is owner = msg.sender. However, without considering context information, this simple statement can be easily identified in many other smart contracts with the exact identical code tokens, and most of these reported bugs are false positive cases. This is the reason for the extremely large number of bugs and very low precision by using BasicEmbed. For using SmartEmbed, we can encode the context of a statement, such as the function signatures function DynamicPyramid and contract ancestor node Rubixi into the code embedding vector, which can effectively reduce the false positive rate and identify the real bugs in other smart contracts.
Answer to RQ-4: How effective is the structural and semantic information added to SmartEmbed? - we conclude that the structural and semantic information added to SmartEmbed do have significant benefits for its overall performance.
7 RQ-5: Contract Validation Evaluation
Because a smart contract is immutable once it is deployed onto the blockchain, it would be better to ensure its correctness in its pre-deployment phase. The objective of the experiment here is to test the capability of SmartEmbed in catching all bugs in a smart contract that are similar to known bugs, so as to help validate the correctness of the contract. Although not a formal verification tool, our approach can grow its capability in validating a smart contract, as it is easily extensible to incorporate new known bugs into our bug database to check whether a smart contract contains similar bugs.
To help validate a given contract, for each statement in the contract, we generate a 150 dimensional vector for based on our model and query it against all the bugs in our bug database . If the similarity between and any bug in our bug database exceeds a threshold ( is set to 0.95, 0.90 & 0.85 for this task), can be reported as a potential bug.
To assess the effectiveness of our approach, we took the 20 smart contracts without * in Table IV for test. Also, a list of “bug-free” smart contracts can help to assess false positive and false negative rates. Therefore, we collected 20 audited smart contracts from Zeppelin, one of the most popular security audit firms. Each vulnerability discovered on them is automatically considered as a false positive. There are a total of 2857 statements associated with these 40 smart contracts (20 buggy and 20 bug-free); 45 statements from the 20 buggy contracts are labelled as bugs. We performed bug detection on these smart contracts by using both our SmartEmbed approach (SE) and SmartCheck (SC). The confusion matrix with respect to the bug reports generated by SE with three different similarity thresholds (0.95, 0.90 and 0.85) and SC are summarized in Table IX. We also calculated the Precision, Recall, F1 score, FPR (false positive rate), and FNR (false negative rate) based on the confusion matrix and show the metrics in Table X.
7.2 Experimental Results
From Table IX and Table X, it can be seen that:
The majority of the bugs can be checked with our approach, and our approach can identify clone-related bugs more accurately than SmartCheck, which is consistent with bug detection evaluation results.
By using our approach with the similarity threshold 0.90, the number of false positives was 8 and it decreased to 0 with the similarity threshold 0.95. SmartCheck reported far more false positives than ours. Since SmartCheck can check more kinds of bug patterns, it is worth noting that, for a fairer comparison, we only enabled the bug types listed in Table IV for SmartCheck. When other types of vulnerabilities were disabled, SmartCheck still had a 9.9% false positive rate; its FPR would be overwhelmingly higher if all bug types were enabled.
The number of clone-related bugs discovered by our approach increased from 27 to 36 with decreasing similarity thresholds from 0.95 to 0.90. A potential explanation is related to a common practice by developers who may do code cloning but make changes to the clones for various reasons. Such a practice may cause some cloned code to become dissimilar to each other, which would need lower thresholds to detect them.
The false negatives decreased to 0 when we set the similarity threshold to 0.85, which means all the bugs can be identified by our approach using this threshold. At the same time, the false positives reported by our approach increased to 116, but still far less than the results generated by SmartCheck. Looking at the F1 score of this similarity threshold, our approach is still much better than SmartCheck.
Answer to RQ-5: How effective is SmartEmbed for smart contract validation? - our results show that SmartEmbed is effective in capturing bugs similar to known ones with low false positive rates. Our future work will also continue to enrich the bug database with more real bugs and improve the embeddings.
8 RQ-6: Time Cost Analysis
The time cost of SmartEmbed is mostly for the training of code embeddings and the vector similarity checking, and is dependent on the sizes of contract codebase and bug database. To analyze the complexity of our proposed approach, we need to measure the time complexity in the computation of similarity as defined in Eqn.(2)(3). For our machine containing an Intel Xeon CPU E5-2640 v4 @ 2.40GHz, the training of code embedding took about a day for our dataset. The average time for a pairwise similarity calculation between two code snippets, as defined in Equation (2) and (3) (Sec. LABEL:sec:similarity) is around 250ns. We estimated the time by applying Deckard, SmartEmbed and SmartCheck service tool for clone detection, bug detection and contract validation tasks respectively. We use the same server described above for testing, it took on average 79.2ms and 416.3ms to check a single smart contract by using Deckard and SmartCheck respectively. Regarding SmartEmbed, for clone detection, computing the pairwise similarity matrix ( was a 2271822718 matrix) took on average 6.05s, checking each smart contract only cost 0.26ms. For bug detection, all statements in our contract codebase are queried against our bug embedding matrix, computing the similarity matrix ( was a 194451363 matrix) took on average 53.22s, checking each smart contract cost 2.3ms. For contract validation, a given contract is queried against our bug embedding matrix, which took on average 4.7ms.
Answer to RQ-6: How efficient is SmartEmbed? - The query for a clone or a bug using SmartEmbed is efficient for practical uses.
We selected several smart contract projects from Github, then contacted the Solidity developers by sending clone reports and bug reports generated by SmartEmbed for these projects. For clone detection, we reported the most similar smart contracts’ url on Etherscan associated with its similarity score. For bug detection, we reported the exact bug line and associated bug type. Some developers expressed interest in using our tool.
Clone Detection - Compared to Etherscan’s “find similar contract” function, which can only find “Exact Match” contracts, our tool is more flexible which can report code clone on contract level, function level or even statement level governed by a similarity threshold. One practitioner responded, “If the tool works with individual functions then that might be useful. I would give you a shout out on Twitter”. Another developer commented, “The clone detection isn’t useful to me, but I could believe it would be useful to authors of widely cloned contracts, such as cryptokitties or FOMO3D.”
Bug Detection - With the help of our techniques, developers could quickly check for vulnerabilities and improve confidence in the reliability of a contract. “It is nice to have such a tool to identify vulnerable bugs in smart contract, I probably will give it a try”. However, there are also some developers who mentioned that the bug report is not useful, “one intractable problem I found was that in smart contracts, everything is dangerous, and you can’t judge whether a contract is secure without understanding intent - any insecure pattern can be correct in the context of a contract designed to do that. ”
According to developers’ comments, we have implemented SmartEmbedhttp://www.smartembed.net as a standalone web application tool . Solidity developers can copy and paste their contract source code to the web application to find repetitive contract code and clone-related bugs in the given contract. The source code of SmartEmbed and contract data used in our experiments can be found in our Github repositoryhttps://github.com/beyondacm/SmartEmbed.
Some developers also suggested publishing the tool as an extension and enhancement to Etherscan so that developers who have already been familiar with Etherscan can easily utilize the tool, which can facilitate broader adoption of the tool and easier collections of new bugs. Since a lot of Solidity developers use the web IDE Remix to develop, deploy, and test a smart contract, developers also suggested integrating the tool as a plugin into an IDE (e.g., Remix and Visual Studio Code) to help detect clones and bugs early in development. The efficiency of SmartEmbed’s similarity checking step (excluding the embedding steps), as shown in Section 4.8, can be sufficient in supporting the uses in IDE in real-time when developers are writing their code. We will follow such suggestions to improve the tool in the near future.
Internal Validity. Code representations used for code embedding have significant effects on the embedding outcome and the downstream applications. The ways we calculated the code embedding for each code snippet is intuitive, which may bias our approach for detecting clones of different code sizes. There are a lof of related work have explored different ways to represent code and embed more semantic information into the code vectors, such as paths in control flow graphs, paths in ASTs, dynamic execution traces, API sequences and usage contexts , and many others. We will try to employ different code embedding techniques for the same tasks in the future.
Data Validity. We collected 22,725 solidity smart contracts with source code through Etherscan for our experiment. It is not complete as the number of smart contracts on Ethereum grows faster recently and the number of contracts on Etherscan is almost doubled, over 40,000 already, not to mention many other contracts that do not provide source code. In the future, we can retrain our model and gain a better code representation model with the enlarged Solidity source code data set, and may even extend the embedding techniques to Solidity bytecode. In addition, due to the lack of a comprehensive list of Ethereum contract vulnerabilities, the number of buggy contracts we collected is relatively small. Our bug database currently contains 52 buggy contracts covering 10 different bug types that are more relevant for Solidity smart contracts, ignoring bug types that may be common for other programming languages. The selected contracts may not be sufficiently diverse or representative of all contracts, and there can be a lot of false negatives if applying our approach to detect bug types not included in our bug database. We will keep expanding both our code base and bug database in the near future.
External Validity. We validated the clone-related bugs detected by SmartEmbed only from the SmartCheck benchmark. One of the threat is that SmartCheck can also have the false negative as well as false positive cases, hence the results may be biased and incomprehensive. There currently exists other security analysis tools to find bugs in smart contract, such as Oyente , Mythril , Gasper and Securify . We plan to do more large-scale evaluations with these tools in the near future. We also acknowledge that the sample size of the user study is not sufficient, we plan to get more feedback about our tool from practitioners in the future.
We have proposed a new approach, SmartEmbed, based on structural code embedding and similarity checking for clone detection, bug detection and contract validation tasks on smart contracts. We have evaluated our approach with more than 22,000 Solidity smart contracts from the Ethereum blockchain. For clone detection, SmartEmbed can effectively identify many instances of repetitive solidity code where the clone ratio is around 90%, and more semantic clones can be detected accurately by our tool than Deckard. For bug detection, SmartEmbed can identify more than 1000 clone-related bugs based on our bug databases efficiently and accurately, which can enable efficient checking of smart contracts with changing code and bug patterns. Such capabilities of SmartEmbed can be useful for facilitating contract validation in practice.