DeepAM: Migrate APIs with Multi-modal Sequence to Sequence Learning

Xiaodong Gu, Hongyu Zhang, Dongmei Zhang, Sunghun Kim

Introduction

Programming language migration is an important task in software development Mossienko (2003); Hassan and Holt (2005); Tonelli et al. (2010). A software product is often required to support a variety of devices and environments. This requires developing the software product in one language and manually porting it to other languages. This procedure is tedious and time-consuming. Building automatic code migration tools is desirable to reduce the effort in code migration.

However, current language migration tools, such as Java2CSharphttp://j2cstranslator.wiki.sourceforge.net/, require users to manually define the migration rules between the respective program constructs and the mappings between the corresponding Application Programming Interfaces (APIs) that are used by the software libraries of the two languages. For example, The API BufferedReader.read in Java should be mapped to StreamReader.read in C# . Such a manual procedure is tedious and error-prone. As a result, only a small number of API mappings are produced Zhong et al. (2010).

To reduce manual effort in API migration, several approaches have been proposed to automatically mine API mappings from a software repository Nguyen et al. (2014); Pandita et al. (2015); Zhong et al. (2010). For example, Nguyen et al. proposed StaMiner that applies statistical machine translation (SMT) Koehn et al. (2003) to bilingual projects, namely, projects that are released in multiple programming languages. It first aligns equivalent functions written in two languages that have similar names. Then, it extracts API mappings from the paired functions using the phrase-based SMT model Koehn et al. (2003).

However, existing approaches rely on the sparse availability of bilingual projects. The number of available bilingual projects is often limited due to the high cost of manual code migration. For example, we analyzed 11K Java projects on GitHub which were created between 2008 to 2014. Among them, only 15 projects have been manually ported to C# versions. Therefore, the number of API mappings produced by existing approaches is rather limited. In addition, given bilingual projects, they need aligning equivalent functions using name similarity heuristics. Only a portion of functions in a bilingual project have similar function names and can be aligned Zhong et al. (2010).

In this paper, we propose DeepAM (Deep API Migration), a novel, deep learning based system to API migration. Without the restriction of using bilingual projects, DeepAM can directly identify equivalent source and target API sequences from a large-scale commented code corpus. The key idea of DeepAM is to learn the semantic representations of both source and target API sequences and identify semantically related API sequences for the migration. DeepAM assigns to each API sequence a continuous vector in a high-dimensional semantic space in such a way that API sequences with similar vectors, or “embeddings”, tend to have similar natural language descriptions.

In our approach, DeepAM first extracts API sequences (i.e., sequences of API invocations) from each function in the code corpus. For each API sequence, it assigns a natural language description that is automatically extracted from corresponding code comments. With the ⟨\langleAPI sequence, description⟩\rangle pairs, DeepAM applies the sequence-to-sequence learning Cho et al. (2014) to embed each API sequence into a fixed-length vector that reflects the intent in the corresponding natural language description. By jointly embedding both source and target API sequences into the same space, DeepAM aligns the equivalent source and target API sequences that have the closest embeddings. Finally, the pairs of aligned API sequences are used to extract general API mappings using SMT.

To our knowledge, DeepAM is the first system that applies deep learning techniques to learn the semantic representations of API sequences from a large-scale code corpus. It has the following key characteristics that make it unique:

Big source code: DeepAM enables the construction of large-scale bilingual API sequences from big code corpus rather than limited bilingual projects. It learns API semantic representations from 10 million commented code snippets collected over seven years.

Deep model: The multi-modal sequence-to-sequence learning architecture ensures the system can learn deeper semantic features of API sequences than the traditional shallow ones.

Related Work

API migration has been investigated by many researchers Nguyen et al. (2014); Pandita et al. (2015); Zhong et al. (2010). Zhong et al. proposed MAM, a graph based approach to mine API mappings. MAM builds on projects that are released with multiple programming languages. It uses name similarity to align client code of both languages. Then, it detects API mappings between these functions by analyzing their API Transformation Graphs. Nguyen et al. proposed StaMiner that directly applies statistical machine translation to bilingual projects.

However, these techniques require the same client code to be available on both the source and the target platforms. Therefore, they rely on the availability of software packages that have been ported manually from the source to the target platform. Furthermore, they use name similarity as a heuristic in their API mapping algorithms. Therefore, they cannot align equivalent API sequences from client code which are similar but independently-developed.

Pandita et al. proposed TMAP, which applies the vector space model Manning et al. (2008), an information retrieval technique, to discover likely mappings between APIs. For each source API, it searches target APIs that have similar text descriptions in their API documentation. However, the vector space model they applied is based on the bag-of-words assumption; it cannot identify sentences with semantically related words and with different sequences of words.

Recently, deep learning technology Sutskever et al. (2014); Cho et al. (2014) has been shown to be highly effective in various domains (e.g., computer vision and natural language processing). Researchers have begun to apply this technology to tackle some software engineering problems. Huo et al. propose a neural model to learn unified features from natural and programming languages for locating buggy source code Huo et al. (2016). Gu et al. apply sequence-to-sequence learning to generate API sequences from natural language queries Gu et al. (2016). Hence, this study constitutes the first attempt to apply the deep learning approach to migrate APIs between two programming languages.

Method

Let A\mathcal{A}={a(i)}\{a^{(i)}\} denote a set of API sequences where a(i)a^{(i)}=[α1,...,αLa\alpha_{1},...,\alpha_{L_{a}}] denotes the sequence of API invocations in a function. Suppose we are given a set of source API sequences AS\mathcal{A}_{S}={aS(i)}\{a_{S}^{(i)}\} (i.e., API sequences in a source language) and a set of target API sequences AT\mathcal{A}_{T}={aT(i)}\{a_{T}^{(i)}\} (i.e., API sequences in a target language). Our goal is to find an alignment between AS\mathcal{A}_{S} and AT\mathcal{A}_{T}, namely,

so that each source API sequence aS(i)<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>∈</mo></mrow><annotationencoding="application/x−tex">∈</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5782em;vertical−align:−0.0391em;"></span><spanclass="mrel">∈</span></span></span></span></span>ASa_{S}^{(i)}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>∈</mo></mrow><annotation encoding="application/x-tex">\in</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em;"></span><span class="mrel">∈</span></span></span></span></span>\mathcal{A}_{S} is mapped to an equivalent target API sequence aT(j)<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>∈</mo></mrow><annotationencoding="application/x−tex">∈</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5782em;vertical−align:−0.0391em;"></span><spanclass="mrel">∈</span></span></span></span></span>ATa_{T}^{(j)}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>∈</mo></mrow><annotation encoding="application/x-tex">\in</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em;"></span><span class="mrel">∈</span></span></span></span></span>\mathcal{A}_{T}.

Since AS\mathcal{A}_{S} and AT\mathcal{A}_{T} are heterogeneous, it is difficult to discover the correlation ff directly. Our approach is based on the intuition of “third party translation”. That is, although AS\mathcal{A}_{S} and AT\mathcal{A}_{T} are heterogeneous, in the sense of vocabulary and usage patterns, they can all be mapped to high-level user intents described in natural language. Thus, we can bridge them through their natural language descriptions. For each a(i)<spanclass="katex−display"><spanclass="katex"><spanclass="katex−mathml"><mathxmlns="http://www.w3.org/1998/Math/MathML"display="block"><semantics><mrow><mo>∈</mo></mrow><annotationencoding="application/x−tex">∈</annotation></semantics></math></span><spanclass="katex−html"aria−hidden="true"><spanclass="base"><spanclass="strut"style="height:0.5782em;vertical−align:−0.0391em;"></span><spanclass="mrel">∈</span></span></span></span></span>Aa^{(i)}<span class="katex-display"><span class="katex"><span class="katex-mathml"><math xmlns="http://www.w3.org/1998/Math/MathML" display="block"><semantics><mrow><mo>∈</mo></mrow><annotation encoding="application/x-tex">\in</annotation></semantics></math></span><span class="katex-html" aria-hidden="true"><span class="base"><span class="strut" style="height:0.5782em;vertical-align:-0.0391em;"></span><span class="mrel">∈</span></span></span></span></span>\mathcal{A}, we assume that there is a corresponding natural language description d(i)d^{(i)}=[w1,...,wLdw_{1},...,w_{L_{d}}] represented as a sequence of words.

The idea can be formulated with Joint Embedding(a.k.a., multi-modal embedding) Xu et al. (2015), a technique to jointly embed/correlate heterogeneous data into a unified vector space so that semantically similar concepts across the two modalities occupy nearby regions of the space Andrej and Li (2015). In our approach, the joint embedding of AS\mathcal{A}_{S} and AT\mathcal{A}_{T} can be formulated as:

Through joint embedding, AS\mathcal{A}_{S} and AT\mathcal{A}_{T} can be easily correlated through their semantic vectors VASV_{\mathcal{A}_{S}} and VATV_{\mathcal{A}_{T}}. Figure 1 shows an illustration of joint semantic embedding between Java and C# API sequences. We are given a corpus of API sequences (in both Java and C# ) and the corresponding natural language descriptions. Each API sequence is embedded (through ϕ\phi or ψ\psi) and translated (through τ\tau) to its corresponding description. The yellow and blue points represent embeddings of Java and C# APIs respectively. Through traning, the Java API sequence BufferedWriter.new→\rightarrow BufferedWriter.write and the C# API sequence StreamWriter.new→\rightarrowStreamWriter.write are embedded into a nearby place in order to generate similar corresponding descriptions write text to file and save to a text file. Therefore, the two API sequences can be identified as semantically equivalent API sequences.

In our approach, the semantic embedding function (ϕ\phi or ψ\psi) and the translation function τ\tau are realized using the RNN-based sequence-to-sequence learning framework Cho et al. (2014). The sequence-to-sequence learning is a general framework where the input sequence is embedded to a vector that represents the semantic representation of the input, and the semantic vector is then used to generate the target sequence. The model that embeds the sequence to a vector (i.e., ϕ\phi or ψ\psi) is called “encoder”, and the model that generates the target sequence(i.e., τ\tau) is called “decoder”.

The framework of the sequence-to-sequence model applied to API semantic embedding is illustrated in Figure 2. Given a set of ⟨\langleAPI sequence, description⟩\rangle pairs {⟨a(i),d(i)⟩}\{\langle a^{(i)},d^{(i)}\rangle\}, The encoder (a bi-directional recurrent neural network Mikolov et al. (2010)) converts each API sequence aa=[α1,...,αLa\alpha_{1},...,\alpha_{L_{a}}], to a fixed-length vector a\bm{a} using the following equations iteratively from tt = 1 to LdL_{d}:

where ht(t\bm{h}_{t}(t=1,...,La)1,...,L_{a}) represents the hidden states of the RNN at each potion tt of the input; [aa;bb] represents the concatenation of two vectors, Wenc\bm{W}_{enc} and benc\bm{b}_{enc} are trainable parameters in the RNN, tanhtanh is the activation function.

The decoder then uses the encoded vector to generate the corresponding natural language description dd by sequentially predicting a word wtw_{t} conditioned on the vector a\bm{a} as well as previous words w1,...,wt−1w_{1},...,w_{t-1}.

where st(t\bm{s}_{t}(t=1,...,Ld)1,...,L_{d}) represents the hidden states of the RNN at each potion tt of the output; Wdeco\bm{W}_{dec}^{o}, bdeco\bm{b}_{dec}^{o}, Wdecs\bm{W}_{dec}^{s} and bdecs\bm{b}_{dec}^{s} are trainable parameters in the decoder RNN.

Both the encoder and decoder RNNs are implemented as a bidirectional gated recurrent neural network (GRU) Cho et al. (2014) which is a widely used implementation of RNN. Both GRUs have two hidden layers, each with 1000 hidden units.

2 Joint Semantic Embedding for Aligning Equivalent API Sequences

For joint embedding, we train the sequence-to-sequence model on both {⟨aS(i),dS(i)⟩}\{\langle a_{S}^{(i)},d_{S}^{(i)}\rangle\} and {⟨aT(i),dT(i)⟩}\{\langle a_{T}^{(i)},d_{T}^{(i)}\rangle\} to minimize the following objective function:

where NSN_{S} and NTN_{T} are the total number of source and target training instances, respectively. LdL_{d} is the length of each natural language sentence. θ\theta denotes model parameters, while pθ(w(it)∣a(i))p_{\theta}(w^{(it)}|a^{(i)}) (derived from Equation 3 to 7) denotes the likelihood of generating the tt-th target word given the API sequence a(i)a^{(i)} according to the model parameters θ\theta.

After training, each API sequence aa=[α1,...,αLa]\alpha_{1},...,\alpha_{L_{a}}] is embedded to a vector a\bm{a} that reflects developer’s high-level intent. We identify equivalent source and target API sequences as those having close semantic vectors.

Implementation

In this section, we describe the detailed implementation of DeepAM, a deep-learning based system we propose to migrate API usage sequences. Figure 3 shows the overall workflow of DeepAM. It includes four main steps. We first prepare a large-scale corpus of ⟨\langleAPI sequence, description⟩\rangle pairs for both Java and C# (Step 1). The pairs of both languages are jointly embedded by the sequence-to-sequence model as described in Section 3.2 (Step 2). Then, we identify related Java and C# API sequences according to their semantic vectors (Step 3). Finally, a statistical machine translation component is used to extract general API mappings from the aligned bilingual API sequences (Step 4).

In theory, our system could migrate APIs between any programming languages. In this paper we limit our scope to the Java-to-C# migration. The details of each step are explained in the following sections.

We first construct a large-scale database that contains ⟨\langleAPI sequence, description⟩\rangle pairs for training the model. We download Java and C# projects created from 2008 to 2014 from GitHubhttp://github.com. To remove toy or experimental programs, we only select the projects with at least one star. In total, we collected 442,928 Java projects and 182,313 C# projects from GitHub.

Having collected the code corpus, we extract API sequences and corresponding natural language descriptions: we parse source code files into ASTs (Abstract Syntax Trees) using Eclipse’s JDT compilerhttp://www.eclipse.org/jdt for Java projects, and Roslynhttps://roslyn.codeplex.com/ for C# projects. Then, we extract the API sequence from individual functions using the same approach in Gu et al. (2016).

To obtain natural language descriptions for the extracted API sequences, we extract function-level code summaries from code comments. In both Java and C# , it is the first sentence of a documentation commentA documentation comment in Java starts with ‘/**’ and ends with ‘*/’. A documentation comment in C# starts with a “<<summary>>” tag and ends with a “<</summary>>” tag. for a function. According to the Javadoc guidancehttp://www.oracle.com/technetwork/articles/java/index-137868.html, the first sentence of a documentation comment is used as a short summary of a function. Figure 4 shows an example of documentation comments for a C# function TextFile.ReadFilehttps://github.com/virtualmarc/gitlab-ci-runner-win/blob/master/gitlab-ci-runner/helper/TextFile.cs in the Gitlab CI project.

Finally, we obtain a database consisting of 9,880,169 ⟨\langleAPI sequence, description⟩\rangle pairs, including 5,271,526 Java pairs and 4,608,643 C# pairs.

2 Model Training

We train the sequence-to-sequence model on the collected ⟨\langleAPI sequence, description⟩\rangle pairs of both Java and C# . The model is trained using the mini-batch stochastic gradient descent algorithm (SGD) Bottou (2010) together with Adadelta Zeiler (2012). We set the batch size as 200. Each batch is constituted with 100 Java pairs and 100 C# pairs that are randomly selected from corresponding datasets. The vocabulary sizes of both APIs and natural language descriptions are set to 10,000. The maximum sequence lengths LaL_{a} and LdL_{d} are both set as 30. Sequences that exceed the maximum lengths will be excluded for training.

After training, we feed in the encoder with all API sequences and obtain corresponding semantic vectors from the last hidden layer of encoder.

3 API Sequence Alignment

After embedding all API sequences, we build pairs of equivalent Java and C# API sequences according to their semantic vectors. For each Java API sequence, we find the most related C# API sequence to align with by selecting the C# API sequence that has the most similar vector representation. We measure the similarity between the vectors of two API sequences using the cosine similarity, which is defined as:

where as\bm{a}_{s} and at\bm{a}_{t} are vectors of source and target API sequences. The higher the similarity, the more related the source and target API sequences are to each other.

Finally, we obtain a database consisting of aligned pairs of Java and C# API sequences.

4 Extracting General API Mappings

Experimental Results

We first evaluate how accurate DeepAM performs in mining API mappings. We focus on 1-to-1 API mappings that are currently used by many code migration tools such as Java2Csharp. We compare the 1-to-1 API mappings mined by DeepAM (Section 4) with a ground truth set of manually written API mappings provided by Java2CSharp. Metric We use the F-score to measure the accuracy. It is defined as: FF=2PR/(P2PR/(P+R)R) where PP=TPTP+FP\frac{TP}{TP+FP} and RR=TPTP+FN\frac{TP}{TP+FN}. TP is true positive, namely, the number of API mappings that are both in DeepAM results and in the ground truth set. FP is false positive which represents the number of resulting mappings that are not in the ground truth set. FN is false negative, which represents the number of mappings that are in the ground truth set but not in the results. Baselines We compare DeepAM with StaMiner Nguyen et al. (2014) and TMAP Pandita et al. (2015). StaMiner is a state-of-the-art API migration approach that directly utilizes statistical machine translation on bilingual projects. TMAP Pandita et al. (2015) is an API migration approach using information retrieval techniques. It aligns Java and C# APIs by searching similar descriptions in API documentation. For easy comparison, we use the same configuration as in TMAP Pandita et al. (2015). We manually examine the numbers of correctly mined API mappings on several Java SDK classes and make a direct comparison with the TMAP’s results presented in their paper. Results Table 1 shows the accuracy of both DeepAM and StaMiner. We evaluate the accuracy of mappings for both API classes and API methods. The results show that DeepAM is able to mine more correct API mappings. It achieves average recalls of 82.6% and 82.3% for class and method migrations respectively, which are significantly greater than StaMiner (60.2% and 57.6%). The average precisions of DeepAM are 82.7% and 71.9%, slightly less than but similar to StaMiner (77.9% and 81.1%). Overall, DeepAM performs better than StaMiner, with average F-measures of 81.9% and 76.3% compared to StaMiner ’s (66.2% and 65.0%).

Table 2 shows the number of correctly mined API mappings by TMAP and DeepAPI. The column # Methods lists the total numbers of API methods for each class. As shown in the results, DeepAM can mine many more correct API mappings than TMAP, which is based on text similarity matching.

The results indicate that without the restriction of a few bilingual projects, DeepAM yields many more correct API mappings.

2 The Scale of Mined API Mappings

Overall, the results indicate that DeepAM significantly increases the number of API mappings than StaMiner, with comparable quality. These results are expected because DeepAM does not rely upon bilingual projects, therefore significantly increasing the size of available training corpus.

Table 4 shows some concrete examples of API mappings. We selected 12 programming tasks that are commonly used in the literature Lv et al. (2015); Gu et al. (2016). The results show that DeepAM can successfully migrate API sequences for these tasks. DeepAM also performs well in longer API sequences such as copy file and play audio.

3 Effectiveness of Multi-modal API Sequence Embedding

As the most distinctive feature of our approach is the multi-modal semantic embedding of API sequences, we also evaluate DeepAM’s effectiveness in embedding API sequences, namely, whether the joint embedding is effective on API sequence alignment. As described in Section 4.3, we apply the semantic embedding and sequence alignment on raw API sequences, and obtain a database of semantically related Java and C# API sequences. We randomly select 500 aligned pairs of Java and C# API sequences from the database and manually examine whether each pair is indeed related. We calculate the ratio of related pairs of the 500 sampled pairs. Baseline We compare our results with an IR based approach. This approach aligns API sequences by directly matching corresponding descriptions using text similarities (e.g., the vector space model) Manning et al. (2008). We implement it using Lucenehttps://lucene.apache.org/. For each Java API sequence, we search the C# API sequence whose description is most similar to the description of the Java API sequence, and vice versa. We randomly select 500 aligned pairs from the results and manually examine the ratio of correctly aligned pairs. Results Table 5 shows the performance of sequence alignment. The column Java version shows the ratio of Java API sequences which are correctly aligned to C# API sequences. Likewise, the C# version column shows the ratio of C# API sequences that are correctly aligned to Java API sequences. The results show that the joint embedding is effective for the API sequence alignment. The ratio of successful alignments is 72.4%, which significantly outperforms the IR based approach (average accuracy is 40.8%). The results indicate that the deep learning model is more effective in learning semantics of API sequences than traditional shallow models such as the vector space model.

Conclusion

In this paper, we propose a deep learning based approach to the migration of APIs. Without the restriction of using bilingual projects, our approach can align equivalent API sequences from a large-scale commented code corpus through multi-modal sequence-to-sequence learning. Our experimental results have shown that the proposed approach significantly increases the accuracy and scale of API mappings the state-of-the-art approaches can achieve. Our work demonstrates the effectiveness of deep learning in API migration and is one step towards automatic code migration.

References