TranSmart: A Practical Interactive Machine Translation System

Guoping Huang, Lemao Liu, Xing Wang, Longyue Wang, Huayang Li, Zhaopeng Tu, Chengyan Huang, Shuming Shi

Introduction

Recent years have witnessed a breakthrough in automatic machine translation, thanks to the advances in neural machine translation (NMT) . The key idea in NMT is an encoder-decoder framework where a source sentence is represented by an encoder network and a target sentence is generate by an decoder network . By using a large-scale bilingual corpus, both encoder and decoder networks consisting of massive parameters can be trained sufficiently enough to yield excellent generalization ability on unseen sentences. As a result, NMT delivers state of the art performance in many machine translation benchmarks .

To the date, however, even a state-of-the-art NMT system is incapable of meeting the strict requirements in some translation scenarios, where human translation is still intensively involved even if it is inefficient and costly . Therefore, interactive machine translation (IMT) , which is believed to better trade off translation quality and efficiency, has drawn increasing attention : At each time a user corrects a translation, the machine automatically generates a new translation based on the human corrections until the translation process is finished . Compared to automatic machine translation, an interactive machine translation system includes more components and it is more complicated in essence. During the past decade, there were a few IMT systems available, for instance, TransType , CASMACAT , and LILT www.lilt.com. which are either on top of statistical machine translation or the advanced NMT. All these systems share a common characteristic: they perform translation in an incremental manner from left to right and naturally require human translators to follow this strict manner, as shown in Figure 1 (a & c), no matter what translation manner of translators is.

In this report, we describe a new human-machine interactive translation system, TranSmart, which was first made publicly accessible in 2018. Compared with other IMT systems, one advantage of TranSmart is its flexibility in the sense that users can employ any translation manners to interact with the machine rather than the incremental manner from left to right. Therefore, if some users prefer to translate difficult words first and then easy words, TranSmart is able to conduct interactive translation in their own manners, as shown in Figure 1 (b & d). Moreover, TranSmart includes an NMT engine augmented with translation memory, and it is very helpful to translate a document where similar translation errors may occur for an automatic machine translation engine. To the best of our knowledge, TranSmart is the first interactive neural machine translation (INMT) system which makes use of translation memory.

Specifically, TranSmart provides three key features as follows:

Word-level autocompletion: At word level, in order to input a correct word, users do not need to type all characters of this word from scratch but a few characters instead, and then the system can automatically complete this word. Unlike other INMT systems, the corrected word is not necessary to be adjacent to the translation prefix.

Sentence-level autocompletion: At sentence level, users do not need to translate all words from scratch but provide some of them (for instance, some difficult words), and then the system can automatically complete the translation based on user provided words. Different from other INMT systems, user provided words can be discontinuous.

Translation-memory-augmented NMT: The system is able to re-use translation results from users through translation memory and offers better translation. As there may be several similar sentences in a document to be translated, the system can efficiently avoid the occurrence of similar translation errors thanks to the knowledge from translation memory.

By repeatedly leveraging these key features, TranSmart is able to interactively collaborate human and a machine to generate a high-quality translation in an efficient way. In addition, TranSmart provides some extended features, including terminology translation, bilingual sentence examples, document translation, tag-preserving translation, and image translation. In the remaining part of this report, we first present our implementation of the key modules of TranSmart. Then the TranSmart API is briefly introduced. Finally, the effectiveness of some functions is demonstrated through empirical experiments.

System Features

At high level, TranSmart contains three key features as well as several extended features as shown in the left side of Figure 2. This section describes the basic ideas of these features, whose detailed implementations will be given in the next section.

Given a source sentence, a translation context consisting of translation pieces, and a human-typed character sequence, word level autocompletion aims to predict a target word which is compatible with the typed sequence. With the help of this feature, if a user is expected to type a target word to correct a translation, it is not necessary to manually type all the characters but only some of them. This feature is inspired from recent advances in input methods from monolingual scenario which are designed to improve the efficiency of human input.

Sentence Level Autocompletion

Given a source sentence and a translation context, sentence-level autocompletion aims to generate a complete translation for the source sentence on top of the context. The translation context can be human typed words (with or without word level autocompletion), or human-edited translation pieces from a translation generated by the system. With this feature, a user does not need to manually translate all words from scratch but some of them which are usually difficult to the system, and then the system tries to complete the translation by generating the remaining words automatically.

Memory-Aware Machine Translation

Memory-Aware Machine Translation aims to generate high-quality translation by making use of translation memory. Users can provide their own domain-specific translation data to our system as the memory or we can use our bilingual training corpus as the memory. In addition, after users finish translating a sentence, we have a mechanism to accumulate their translation history into the memory. In this way, our system has the ability to avoid the same translation errors occurring multiple times, and this feature is useful in translating a document where there are similar sentences.

The above three key features can be used as atomic operations, which can be applied multiple times to interact between the user and machine for a translation task. For example, by repeatedly leveraging word level autocompletion and sentence level autocompletion, TranSmart can generate a complete translation for a source sentence with high-quality, by using the memory-aware translation engine where translation memory is accumulated from translation history of users. It is worth noting that the translation pieces in the translation context may be discontinuous, and thus our word level and sentence level autocompletion is more general and flexible than those in existing INMT systems.

2 Extended Features

This feature is used to translate a formatted document in a source language to a corresponding formatted document in a target language. It supports many popular formats including TXT, HTML, XML, MARKDOWN, PDF, DOCX, PPTX and XLSX. To this end, it generally performs two steps as follows. First, it parses the input formatted document into a text document consisting of sentences with several tags. Each tag may indicate some structural information such as a paragraph or a font. For example, a sentence in the formatted document may be ” is a 1994 American drama film”, where ”” indicates the phrase Forrest Gump is with red color as its format. Second, it translates each tagged sentence by the automatic translation engine with a specially designed technique, which is named tag translation. Tag translation is crucial to document translation and we will present its challenges and our solution in the next section.

Image Translation

This feature aims to translate text contained in an image file in a source language to a text document in a target language. It supports many popular image formats such as JPG and PDF. Generally, it is implemented by a pipeline procedure consisting of two steps as follows. First, it employs an external OCR toolkit to extract a text document where any content beyond text is ignored. Second, it uses our translation engine to translate all sentences in the text document one by one. The first step, which is called text extraction from image, is critical to image translation and we will present its challenge and the technique to address in next section.

Terminology Translation

We collect more than 3 million of Chinese-English terminologies from websites. However, this corpus contains a large amount of noises, including non-terminology, unaligned, inconsistent format. We filter non-terminology words by considering the word frequency due to the long-tail property of terminology. More specifically, we built a phrase table from a large-scale parallel data and then extract high-frequency phrase pairs as a non-terminology list. Finally, we filter noises by comparing the stop list and collected data . Furthermore, we employ our in-house filtering scripts to filter unaligned terms according to various features such as length ratio and language identification. As a result, we obtained a clean version of terminology corpus that contains around 2 million terms.

Bilingual Examples

The input sentence is used to retrieve bilingual examples from the corresponding retrieval repository. We selected and used more than 200M bilingual sentences to build the retrieval repository. The three most similar bilingual examples are displayed to help users to translate the input sentence.

Implemented Techniques

We implemented the generic translation model on top of the Transformer architecture . To balance the translation performance and inference efficiency, we used a 24-layer encoder and a 6-layer decoder, whose hidden size is 1024. We trained the translation model on our in-house data after the following data manipulation methods, which consists of 200 millions of Chinese-English sentence pairs. We followed to train models with batches of approximately 460k tokens, using Adam with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98 and ϵ=10−8\epsilon=10^{-8}.

Data Rejuvenation

Large-scale parallel datasets lie at the core of the recent success of NMT models. However, the complex patterns and potential noises in the large-scale data make training NMT models difficult. We introduce data rejuvenation to improve the training of NMT models on large-scale datasets by exploiting inactive examples . The proposed framework consists of three phases, as shown in Figure 3. First, we train an identification model on the original training data to distinguish inactive examples and active examples by their sentence-level output probabilities. Then, we train a rejuvenation model on the active examples to re-label the inactive examples with forward-translation. Finally, we combined the rejuvenated examples and the active examples as the final bilingual data.

Data Augmentation

Although we have large-scale parallel data, there are limited amount parallel data for some specific domains. Data augmentation methods (e.g. self-training and back-translation) are a promising way to alleviate this problem by augmenting model training with synthetic parallel data. The common practice is to construct synthetic data based on a randomly sampled subset of large-scale monolingual data, which we empirically show is sub-optimal. In response to this problem, we improve the sampling procedure by selecting the most informative monolingual sentences to complement the parallel data. To this end, we compute the uncertainty of monolingual sentences using the bilingual dictionary extracted from the parallel data. Intuitively, monolingual sentences with lower uncertainty generally correspond to easy-to-translate patterns which may not provide additional gains. Accordingly, we design an uncertainty-based sampling strategy to efficiently exploit the monolingual data for self-training, in which monolingual sentences with higher uncertainty would be sampled with higher probability.

2 General Word-level Autocompletion

Word-level autocompletion aims to complete the target word based on human typed characters for a given a source sentence and translation context. Previous studies have explored word-level autocompletion task, but they either do not take into account translation context or they require the target word to be the next word of the translation prefix , which limit its applications in real-world scenarios such as post-editing . To this end, we propose a general word-level autocompletion task which can be applied to more general scenarios.

Suppose x\boldsymbol{x}=(x1,x2,…,xm)(x_{1},x_{2},\dots,x_{m}) is a source sequence, s\boldsymbol{s}=(s1,s2,…,sk)(s_{1},s_{2},\dots,s_{k}) is a sequence of human typed characters, and translation context is denoted by c\boldsymbol{c}=(cl,cr)(\boldsymbol{c}_{l},\boldsymbol{c}_{r}), where cl\boldsymbol{c}_{l}=(cl,1,cl,2,…,cl,i)(c_{l,1},c_{l,2},\dots,c_{l,i}) and cr\boldsymbol{c}_{r}=(cr,1,cr,2,…,cr,j)(c_{r,1},c_{r,2},\dots,c_{r,j}). The translation pieces cl\boldsymbol{c}_{l} and cr\boldsymbol{c}_{r} are on the left and right hand side of s\boldsymbol{s}, respectively. Formally, given a source sequence x\boldsymbol{x}, typed character sequence s\boldsymbol{s} and a context c\boldsymbol{c}, the general word-level autocompletion (GWLAN) task aims to predict a target word ww which is to be placed in the middle between cl\boldsymbol{c}_{l} and cr\boldsymbol{c}_{r} to constitute a partial translation. Note that cl\boldsymbol{c}_{l} or cr\boldsymbol{c}_{r} may be empty in some scenarios.

Methodology

Given a tuple (x,c,s)(\boldsymbol{x},\boldsymbol{c},\boldsymbol{s}), our approach decomposes the whole word autocompletion process into two parts: model the distribution of the target word ww based on the source sequence x\boldsymbol{x} and the translation context c\boldsymbol{c}, and find the most possible word ww based on the distribution and human typed sequence s\boldsymbol{s}.

In the first part, we propose a word prediction model (WPM) to define the distribution p(w∣x,c)p(w|\boldsymbol{x},\boldsymbol{c}) of the target word ww. We use a single placeholder [MASK] to represent the unknown target word ww, and use the representation of [MASK] learned from WPM to predict it. Formally, given the source sequence x\boldsymbol{x}, and the translation context c=(cl\boldsymbol{c}=(\boldsymbol{c}_{l}, cr)\boldsymbol{c}_{r}), the possibility of the target word ww is:

Suppose s\boldsymbol{s} denotes a human typed sequence of characters, in the second part, we predict the best word according to the constrained optimization:

where V(s)\mathcal{V}(\boldsymbol{s}) denotes a set of target words, whose element satisfies the sequences of s\boldsymbol{s}, for example, c\boldsymbol{c} is a prefix of ww. More details can be found in .

Data Generation

For training and evaluating GWLAN models above, firstly we should create a large scale dataset including tuples of (x,s,c,w)(\boldsymbol{x},\boldsymbol{s},\boldsymbol{c},w). Ideally, we may hire professional translators to manually annotate such a dataset, but it is too costly in practice. We instead propose to automatically construct the dataset from parallel datasets.

Assume we are given a parallel dataset {(xi,yi)}\{(\boldsymbol{x}^{i},\boldsymbol{y}^{i})\}, where yi\boldsymbol{y}^{i} is the reference translation of xi\boldsymbol{x}^{i}. We automatically construct the data ci\boldsymbol{c}^{i} and si\boldsymbol{s}^{i} by randomly sampling from yi\boldsymbol{y}^{i}. Specifically, we first sample a word w=ykiw=\boldsymbol{y}^{i}_{k}, and sample two spans [al,bl][a_{l},b_{l}] and [ar,br][a_{r},b_{r}] such that 0≤al≤bl≤k0\leq a_{l}\leq b_{l}\leq k and k+1≤ar≤br≤∣yi∣k+1\leq a_{r}\leq b_{r}\leq|\boldsymbol{y}^{i}|, leading to cl=⟨yali,⋯ ,ybl−1i⟩\boldsymbol{c}_{l}=\langle y^{i}_{a_{l}},\cdots,y^{i}_{b_{l}-1}\rangle and cr=⟨yari,⋯ ,ybr−1i⟩\boldsymbol{c}_{r}=\langle y^{i}_{a_{r}},\cdots,y^{i}_{b_{r}-1}\rangle. It is worth mentioning that cl\boldsymbol{c}_{l} or cr\boldsymbol{c}_{r} may not be adjacent to ykiy^{i}_{k}. In addition, we randomly sample a character sequence as ski\boldsymbol{s}^{i}_{k} for ykiy^{i}_{k} as follows: for languages like English and German, ski\boldsymbol{s}^{i}_{k} is a character prefix of ykiy^{i}_{k}; for languages like Chinese, ski\boldsymbol{s}^{i}_{k} consists of the first phonetic symbol of each character in ykiy^{i}_{k}. In this way, we can obtain a collection of {(xi,cl,cr,ski,yki)∣∀i,∀k}\{(\boldsymbol{x}^{i},\boldsymbol{c}_{l},\boldsymbol{c}_{r},\boldsymbol{s}^{i}_{k},y^{i}_{k})\mid\forall i,\forall k\}, which can be divided into training, validation and test datasets for the GWLAN task.

3 Sentence-level Autocompletion by Lexical Constraints

In our human-machine interactive translation scenario, human translators may pre-specify some constraint words (or lexical constraints) and our system requires to output a high-quality translation with knowledge from these pre-specified constraints. For example, the constraints can be typed by human translators through word level autocompletion in Section 3.2, and they can be a translation prefix corrected by translators, or a post-edited partial translation. It is worth noting that the constraints are not necessary to be continuous , unlike . In TranSmart, we implement two different approaches to incorporate these constraints into NMT. The first one relies on constrained decoding which requires the output translation to include constraints in a hard manner; whereas the second one makes use of them in a soft manner, i.e., the output translation may not include some of constraints.

Constrained decoding is essentially a constrained optimization problem, and it aims to search the translation which satisfies some constraints and is with the best model score. Formally, suppose P(y∣x)P(\boldsymbol{y}\mid\boldsymbol{x}) is a translation model, and c\boldsymbol{c} denotes a set of constraint words.

where Y(c)\mathcal{Y}(\boldsymbol{c}) denotes a set of translation hypotheses which include all constraint words in c\boldsymbol{c}. To address this constrained optimization problem, proposed a grid beam search algorithm (GBS). This algorithm maintains a beam along two dimensions, where one is the length of the hypotheses and the other is the number words in all constraints. Thus, its complexity is linear to the number of the constraint words, which is inefficient in our interactive scenario. presented an improved algorithm by dynamic beam allocation (DBA) which allows hypotheses in a beam to contain different number of constraint words. In our experiments, this improved algorithm is indeed more efficient in decoding speed but its translation quality is worse than grid beam search. To this end, we propose a variant of grid beam search algorithm to achieve a trade-off between quality and efficiency.

In our interactive scenario, we observe that constraint words in c\boldsymbol{c} exhibit two characteristics: some of constraints in c\boldsymbol{c} are consecutive to form a translation piece, especially when the size of c\boldsymbol{c} is very large; and c\boldsymbol{c} provided by human translators is placed in a fixed order such that the output translation should contain c\boldsymbol{c} in the same order. Based on the observation, we organize a beam along the length of hypotheses and the number of constraint pieces rather than constraint words. Thanks to the fixed order of pieces, any translations with the same number of pieces will naturally contain the same number of constraint words and thereby they can be fairly pruned in terms of model scores, similar to the grid beam search algorithm. Consequently, the complexity of this algorithm is linear to the number of pieces rather than the number of constraints, which leads to substantial speedup in practice because constraint pieces usually are long enough.

NMT with Soft Constraints

Unlike constrained decoding, we propose a new approach to make use of pre-specified constraints in a soft manner. The key is to treat the constraints as an external memory, integrate the memory into the standard Seq2Seq decoder and then train the memory-augmented NMT from a dataset such that it learns to constrain the output. Specifically, given x\boldsymbol{x}, c\boldsymbol{c} and a translation prefix y<t\boldsymbol{y}_{<t}, we generate a target word yt\boldsymbol{y}_{t} according to P(yt∣x,y<t,c;θ)P\left({y}_{t}\mid\boldsymbol{x},\boldsymbol{y}_{<t},\boldsymbol{c};\theta\right). Our model includes three components: the encoder of vanilla NMT model which encodes x\boldsymbol{x} into a sequence of hidden vectors; the Constraint Memory Encoder (CME) that encodes the constraints c\boldsymbol{c} into another sequence of hidden vectors E(c)E(\boldsymbol{c}), i.e., the constraint memory; and the Constraint Memory Integrator (CMI), which integrates the constraint memories into the decoder network to generate the next token yty_{t}. To train the memory-augmented model, we create a dataset from available bilingual corpus. We randomly sample some rare words from the reference r\mathbf{r} as constraints, since rare words are more difficult to translate than other words .

In inference, our decoding is an unconstrained optimization problem:

To approximately solve the above optimization problem, we employ the standard beam search algorithm which is similar to the decoding in our baseline model. Although we need additional overheads to handle the constraint memories, this is negligible to the time consuming of the decoding in NMT. Thus, its search is as efficient as the search algorithm of the standard NMT. In addition, unlike constrained decoding, the method does not force a feasible y\boldsymbol{y} to include all constraints in c\boldsymbol{c}. Hence, if the constraints in c\boldsymbol{c} include some noises (i.e., spelling mistakes), it has the potential to avoid copying these noises. More details about this function can be found in .

4 Graph based Translation Memory

The basic idea of translation memory based MT is to translate an input source sentence by using the translations of the source sentences which are similar to the input one, which is related to domain adaptation . Suppose we are given a translation memory (TM) for a source sentence which is a list of bilingual sentence pairs. Generally, there are two ways to improve translation models with translation memory: training model parameters with augmented data (i.e., memory) and summarizing knowledge from translation memory to augment MT decoder . For the latter idea, a typical solution to represent a TM is to encode each word in both the source and target sides by a neural memory. Unfortunately, since a word even a phrase may appear repeatedly in a TM, redundant words are encoded multiple times, leading to a large memory network as well as considerable computation. To address this issue, we propose an effective approach to representing a TM with a compact graph structure, which is further used to enhance the translation model. The model structure is illustrated in Figure 5 and its more details can be found in .

It is observed that most source words in the TM also appear in the input sentence and have already been represented by the encoder. In addition, we believe that those words in the TM yet beyond the source side may not be informative to translate the input sentence itself. Therefore, in our proposed model, we directly ignore the source sentences of the TM and only represent the target side. In addition, instead of sequentially encoding target sentences in a TM, we pack them into a compact graph such that some words in different sentences may correspond to the same node in the graph, which is inspired by the notion of lattice or hypergraph in statistical machine translation . To this end, we convert the target side in a TM into a confusion network by using the algorithm proposed by and .

NMT with Graph based TM

The graph based TM is further used to enhance the Transformer architecture. Generally, the enhanced Transformer shares the similar architecture as Transformer but with two major differences in encoding and decoding phases.

In the encoding phrase, besides encoding the input sequence, the proposed model also encodes the graph by using LL layers of networks in a similar fashion to the encoding of the input. Specifically, we firstly use a multi-head attention to encode each node in the TM graph where the query is the corresponding node, and key-value pairs are obtained from its first-order neighborhood nodes inspired by graph attention . Then we apply other sub-layers to the resulting vector obtained by the multi-head attention layer, which contain a residual layer, a feed-forward layer and a layer normalization. In this way, we can represent the graph into a list of vectors whose size is the same as the number of nodes in the graph.

In the decoding phase, similar to Transformer, the proposed model employs LL layers of networks but each layer includes sub-layers (i.e., multi-head attention, residual layer and layer normalization) to incorporate the list of vectors obtained from graph encoding, besides other six sub-layers. We place the extra three sub-layers nearest to the output of the decoder network in order to let the graph encoding fully influence the decoding process.

5 Others

Word alignment plays an important role in document translation for TranSmart. Since NMT is a blackbox model with massive parameters, previous work has made numerous efforts to induce word alignment from attention in NMT or other explanation methods . Other work improves alignment quality by building a word alignment model whose architecture is similar to NMT . Despite its success, statistical aligners are still respectful counterparts because of their training efficiency and alignment quality. Therefore, we employ statistical aligners to obtain word alignment. As TranSmart involves billion scale of bilingual sentences as its training data, the popular aligner GIZA++ can not train successfully due to memory consumption. Instead, we re-implement an aligner based on HMM with an adaptive strategy to prune the word translation table: for each high-frequency word we allow more words to be its translations whereas we allow less for each low-frequency word.

Tag Translation

Tag translation aims to translate a source language with tags into a target language with tags, and it serves as the key step for document translation. Since the standard translation engine are trained on top of bilingual sentences without tags, standard translation engines can not perform well on translating tagged sentences. In addition, there are no sufficient tagged bilingual sentences to train a customized tag translation engine. Therefore, we propose a simple post-processing approach based on word alignment as follows. First, we delete all tags from a tagged source sentence to obtain a tag-free sentence and then translate it into a target sentence by our default translation engine. Then we run our word aligner on both the tag-free sentence and its translation to obtain word level alignments. Second, we insert the corresponding tags from the source sentence into its translation. However, the second step is not straightforward because one tagged piece (or phrase) within the tagged source sentence may align to multiple pieces in the target side due to the essence of word alignment. We propose an algorithm based on dynamic programming to extract a piece-to-piece alignment: the piece-to-piece alignment is a one-to-one map between pieces in the source and target sides such that it makes the minimal violations according to the word alignment results. By using the piece-to-piece alignment, it is trivial to insert the tags from the source sentence into its target sentence.

Text Extraction from Image

Text Extraction from Image aims to extract the text content from an image file to form a text document in a source language while ignoring other content. Generally, this task is challenging due to two reasons. First, a sentence in an image file may not explicitly contain a special symbol to indicate its end and several such sentences may actually constitute one sentence. More importantly, OCR may recognize many blocks of text content from an image file to some extent, but it is difficult to organize these text blocks into a text document which preserves the same sequential structure of text blocks as in the original image file. To tackle the first challenge, we develop a language model to detect the end symbol for a sentence without an end symbol. In addition, we design another model to detect whether several sentences without an end symbol should be combined to one sentence. To address the second challenge, we employ the position information of each block from the OCR toolkit and uses it as a signal to decide the sequential order of all blocks, which is critical to form the final text document.

Discourse-Aware Translation

Existing translation models usually translate a text by considering isolated sentences based on a strict assumption that the sentences in a text are independent of one another. However, disregarding dependencies across sentences will harm translation quality especially in terms of coherence, cohesion, and consistency . To response this problem, we adapt document-level NMT to document-level training. Specifically, we add bilingual documents and paragraphs to our training data, which can helps models to learn discourse knowledge from larger contexts . This works well with the Document Translation Feature as stated in Section 2.2. Chinese is a pro-drop language, where pronouns are usually omitted when they can be inferable from the context. This leads to serious problems for Chinese-to-English translation models in terms of completeness and correctness. Thus, we recover missing pronouns in the informal domain of training data (e.g. conversations and movie subtitles) by leveraging our approaches .

System Usage

Two ways are available to use TranSmart. One way is to visit the website of TranSmarthttps://transmart.qq.com/index, which has a friendly interactive user-interface (UI). Another way is calling the TranSmart HTTP APIs. The HTTP APIs can be accessed by sending a JSON request to the servicehttps://transmart.qq.com/api/imt via POST method. The major fields of the JSON request of the translation APIs are shown in Table 1. In this section, we will focus more on the latter way and introduce the usage of some selected features.

One important API is the “dynamic_suggestion” function, aiming to save the effort of human translators when correcting the translation generated by AI. This API is capable of providing autocompletion at both word-level and sentence-level, which will be introduced one by one.

First, the word-level autocompletion is able to complete the unfinished word input by human translators. As shown in Figure 6, when the a character sequence “th” is tagged by “editing” in the “segment_list” field, TranSmart will return its autocompletion in the “ime_suggestion” field in response. Second, the sentence-level suggestion can complete the whole translation based on the user-specified spans. As shown in Figure 7, the span with the “prefix” type is promised to be the prefix of the re-generated translation in response, and spans with “std” are forced to be included in re-generated translation. The result of sentence-level autocompletion is in the “sentence_suggestion” field in response.

2 Memory-Aware Machine Translation

As shown in Figure 8, when users provide the information of translation memory in the “reference_list” field, TranSmart will be able to leverage those existing and relevant translations in history to improve the translation quality. The translation results are in the “auto_translation” field in response. Note that users can also disable this feature by removing the “reference_list” field, then TranSmart will translate only based on the source text, i.e., the setting of traditional automatic translation.

3 Extended Features

TranSmart also has two APIs that may be helpful for human translators. First, it provides a static_suggestion function to retrieve relevant information from our pre-constructed terminology and bilingual example databases. As shown in Figure 9, the “static_suggestion” function returns the most relevant terminology translations and bilingual examples in the “term_list” and “sentence_example_list” field in response, respectively. Second, our “selection_suggestion” is designed to translate specified spans in a source sentence. As shown in Figure 10, only spans with “selection” status in the “segment_list” field will be translated. The results of selection suggestions are placed in the “segment_suggestion” field in response.

System Evaluation

We conducted experiments on the widely used WMT14 English⇒\RightarrowGerman (En⇒\RightarrowDe) and English⇒\RightarrowFrench (En⇒\RightarrowFr) datasets, which consist of about 4.54.5M and 35.535.5M sentence pairs, respectively. Note that although the datasets used in the experiments are preprocessed by the standard toolkit in Moses (for a fair comparison between previous methods), in implementing the online system, we adopt TexSmart in data preprocessing (such as Chinese word segmentation) and postprocessing (e.g., restoring case information). We applied BPE with 32K merge operations for both language pairs. The experimental results were reported in case-sensitive BLEU score .

Systems

We validated our approach on a couple of representative NMT architectures:

Lstm that is implemented in the Transformer framework.

Transformer that is based solely on attention mechanisms.

DynamicConv that is implemented with lightweight and dynamic convolutions, which can perform competitively to the best reported Transformer results.

We adopted the open-source toolkit Fairseq to implement the above NMT models. We followed the settings in the original works to train the models. In brief, we trained the Lstm model for 100K steps with 32K (4096×84096\times 8) tokens per batch. For Transformer, we trained 100K and 300K steps with 32K tokens per batch for the Base and Big models respectively. We trained the DynamicConv model for 30K steps with 459K (3584×1283584\times 128) tokens per batch. We selected the model with the best perplexity on the validation set as the final model.

Results

Table 3 lists the results across model architectures and language pairs. Our Transformer models achieve better results than that reported in previous work , especially on the large-scale En⇒\RightarrowFr dataset (e.g., more than 1.0 BLEU points). showed that models of larger capacity benefit from training with large batches. Analogous to DynamicConv, we trained another Transformer-Big model with 459K tokens per batch (“+ Large Batch” in Table 3) as a strong baseline. We tested statistical significance with paired bootstrap resampling using compare-mthttps://github.com/neulab/compare-mt .

Clearly, our data rejuvenation consistently and significantly improves translation performance in all cases, demonstrating the effectiveness and universality of the proposed data rejuvenation approach. It’s worth noting that our approach achieves significant improvements without introducing any additional data and model modification. It makes the approach robustly applicable to most existing NMT systems.

2 Word Level Autocompletion

We carry out experiments on four GWLAN tasks including bidirectional Chinese–English tasks and German–English tasks. The training set for two directional Chinese–English tasks consists of 1.25M bilingual sentence pairs from LDC corpora. As discussed in §3.2, the training data for GWLAN is extracted from 1.25M sentence pairs. The validation data for GWLAN is extracted from NIST02 and the test datasets for GWLAN are constructed from NIST05 and NIST06. For two German–English tasks, we use the WMT14 dataset standard preprocessed by Stanford https://nlp.stanford.edu/projects/nmt/. The validation and test sets for our tasks are based on newstest13 and newstest14 respectively. All the sample operations in the data construction process are based on the uniform distribution. For each dataset, the models are tuned and selected based on the validation set.

Systems

In the experiments, we evaluate and compare the performance of our proposed approach (WPM) and a few baselines, which are illustrated below:

TransTable: We train an alignment model https://github.com/clab/fast_align on the training set and build a word-level translation table. While testing, we can find the translations of all source words based on this table, and select out valid translations based on the human input. The word with highest frequency among all candidates is regarded as the prediction. This baseline is inspired by .

Trans-PE: We train a vanilla NMT model using the Transformer-base model. During the inference process, we use the context on the left hand side of human input as the model input, and return the most possible words based on the probability of valid words selected out by the human input. This baseline is inspired by .

Trans-NPE: As another baseline, we also train an NMT model based on Transformer, but without position encoding on the target side. While testing, we use the averaged hidden vectors of all the target words outputted by the last decoder layer to predict the potential candidates.

Evaluation Metric

To evaluate the performance of the well-trained models, we choose accuracy as the evaluation metric:

where NmatchN_{match} is the number of words that are correctly predicted and NallN_{all} is the number of testing examples.

Results

Table 4 shows the main results of our method and three baselines on the test sets of Chinese-English and German-English datasets. The method Trans-PE, which assumes the human input is the next word of the given context, behaves poorly under the more general setting. As the results of Trans-NPE show, when we use the same model as Trans-PE and relax the constraint of position by removing the position encoding, the accuracy of the model improves. One interesting finding is that the TransTable method, which is only capable of leveraging the zero-context, achieves good results on the Chinese-English task when the target language is English. However, when the target language is Chinese, the performance of TransTable drops significantly. It is clear from the results that our method WPM significantly outperforms the three baseline methods.

3 Sentence Level Autocompletion

We conduct experiments on the Zh⇒\RightarrowEn, Fr⇒\RightarrowEn, De⇒\RightarrowEn, and translation tasks. The Zh⇒\RightarrowEn bilingual corpus includes news articles collected from several online news websites. After the standard preprocessing procedure as in , we obtain about 2 million bilingual sentences in total. Then we randomly select 2000 sentences as the development and test datasets, respectively, and leave other sentences as the training dataset. The Fr⇒\RightarrowEn bilingual corpus is from JRC-Acquis datasets and it is preprocessed following . This dataset is a collection of the parallel legislative text of European Union Law applicable in the EU member states and thus it is a highly related corpus focusing on a specific domain. We also evaluate our methods on the WMT 2018 English-German news translation task, which is composed of Europarl and news commentary data. We use the WMT newstest2013 and newstest2014 as the development and test set respectively. In addition, in our experiments, we use the subword technique to ensure that there are no unknown tokens in our input, even if a constraint word is not included in the training set. Therefore, if a human translator provides an unknown word as a constraint word, our system tokenizes it into several BPE tokens which are considered as several constraints accordingly. Since there are no real constraints available provided by human for all these datasets, we simulate this effect by randomly picking some words from the reference side as the constraints following .

Systems

As presented in section 3.3, we implement two methods: constrained decoding with ordered constraints denoted by O-Gbs and NMT with soft constraints denoted by SC. We compare both of our methods against three baselines:

Transformer: It is the standard Transformer, which is an in-house implementation using Pytorch following .

Gbs: It is the Grid Beam Search in built on top of the in-house Transformer.

Dba: It is the lexically constrained decoding with Dynamic Beam Allocation (an improved approach for ) in .

To ensure the translation quality, we set the default beam size for Gbs and Dba as suggested by and . The hyper-parameters for all the systems are following those of Transformer base model. We train all the models with Adam optimization algorithm and tune the number of iteration based on the performance of the development set.

Results

As shown in Table 5, although the Dba method would reduce the computational overhead compared to Gbs, its translation quality slightly sacrifices accordingly, which is similar to the finding in . Thanks to the ordered constraints, the O-Gbs is faster than Gbs in running speed and delivers the same performance as Gbs. The advantage of these three lexically decoding algorithm is that they do not need to retrain a translation model. Moreover, SC performs better than Gbs on the Fr⇒\RightarrowEn and Zh⇒\RightarrowEn tasks and outperforms Dba on all the three tasks, in terms of translation quality. This observation indicates that our model SC is able to learn how to integrate the lexical constraints correctly. Moreover, there is an additional advantage that the decoding speed of our model is comparable with that of the Transformer model and is much faster than lexically constrained decoding.

4 Translation Memory

Since there is no translation history provided by human translators, we simply use the training set as the memory for simulation. For each sentence, we retrieve 100 translation pairs from the training set by using Apache Lucene . We score the source side of each retrieved pair against the source sentence with fuzzy matching score and select top N=5N=5 translation sentence pairs as a translation memory for the sentence to be translated, following . Sentences from the target side in the translation memory are used to form a graph, with each word represented as a node and the connection between adjacent words in a sentence represented as an undirected edge.

Following the previous works incorporating TM into NMT models, we use the JRC-Acquis corpus for training and evaluating our proposed model. The JRC-Acquis corpus is a collection of parallel legislative text of European union Law applicable in the EU member states. The highly related text in the corpus is suitable for us to make evaluations. To fully explore the effectiveness of our proposed model, we conduct translation experiments on three language pair bidirectionally, namely, en-fr, en-es, and en-de. We manage to obtain preprocessed datasets from . For each language pair, we randomly select 3000 samples to form a development and a test set respectively. The rest of the pairs are used as the training set. Sentences longer than 80 and 100 are removed from the training and development/test set. The technique of Byte-pair Encoding is applied and the vocabulary size is set to be 20K for all the experiments.

Systems

The proposed graph based TM model is built on transformer , and it is denoted by G-TFM. We compare the proposed model against the following baselines:We notice that there are some recent advances on NMT with translation memory after the proposed G-TFM was implemented in TranSmart. We consider to update our TM model in future.

TFM: It is a natural baseline as the proposed model is directly built upon the Transformer architecture.

P-RNN: It is an in-house implementation of on top of RNN-search.

P-TFM: It similar to P-RNN but it is on top of Transformer rather than RNN-search as .

SEG-TFM: It implements the idea of on top of Transformer. Due to the architecture divergence between RNN-based NMT and Transformer, it only differs from the RNN-based counterpart in that two quantities ctc_{t} and ztz_{t} in are replaced by the hidden units obtained from the multi-head attention over the encoding units and the decoding hidden state units before the softmax operator.

SEQ-TFM: It sequentially encodes all target sentences in a TM as one of the baseline models. Specifically, each target sentence in TM goes through a multi-head mechanism and an immediate residual connection plus layer normalization in lthl_{th} layer. The derived representations for these sentences are then concatenated to form the representation of the translation memory, which can be utilized flexibly in lthl_{th} decoding layer.

For training all systems, we maintain the same hyper-parameters for fair comparison. Besides, we adopt the same training algorithm to learn the models as follows. We use a customized leaning rate decay paradigm following Tensor2Tensor package. The learning rate increases linearly on early stages for a certain number of steps, known as warm-up steps, and decay exponentially later on. We set the warm-up step to be 5 epochs and we early stop the model after training 20 epochs, typically the time when the development performance varies insignificantly. Furthermore, since there is a hyperparameter in the system P-TFM of which is sensitive to the specific translation task, we tune it carefully on the development set for all translation tasks. Its optimized value is 0.7 for es and de tasks while it is 0.8 for fr task. We run all 6 tasks with hyperparameters among [0.5, 1.5] with scale of 0.1, and manually pick the optimized value according to its performance on the development set.

Results

Table 6 shows the experiment results of all the systems on the es-en task in terms of BLEU. Several observations can be made from the results. First, the baseline TFM achieves substantial gains over RNN and even outperforms P-RNN by around 1 BLEU point on the test set. Compared with the strongest baseline P-TFM, the proposed SEQ-TFM and G-TFM are able to obtain some gains up to 1.9 BLEU points on the test set. This result verifies that our compact representation of TM is able to guide the decoding of the state-of-the-art model.

Second, it is observed that SEG-TFM is only comparable to TFM on this task, although its RNN based counterpart brought significant gains as reported in . This fact shows that the transformer architecture may need a sophisticated way to well define a key-value memory for TM encoding, which can be significantly different from that on RNN architecture. This is beyond the scope of this paper. Fortunately, this paper provides an easy yet effective approach to encode a TM, i.e. G-TFM, which does not rely on a context-based key-value memory.

Since the retrieval time can be neglected compared with the decoding time as found in , we thereby eliminate the retrieval time and directly compare running time for neural models as shown in Table 7. From this table, we observe that the proposed graph based model G-TFM saves significant running time compared with SEG-TFM and SEQ-TFM while achieving better translation performance.

Table 7 depicts the total number of source and target words encoded by the corresponding model for each test sentence on average. It’s observed that SEG-TFM needs to encode approximately 3 times and SEQ-TFM encodes approximately 2 times the number of words of our proposed model, G-TFM. There’s no wonder that TFM takes the fewest words to encode because no extra TM is included. These statistics indicate that under the scenario of incorporating TM in NMT, our m1odel requires the least memory.

We pick stronger baselines from the es-en task, i.e. TFM and P-TFM, and compare them with the proposed G-TFM model on other 5 translation tasks. Table 8 summarizes their results on both the development and test sets. From this table, we can see that on the test set, G-TFM steadily outperforms TFM by up to 3 BLEU points across all these 5 tasks, In addition, contrast to P-TFM, G-TFM demonstrates better performance by exceeding at least 1 BLEU point across all these tasks except the en-fr task. These results are consistent with the results on es-en task and further validates the effectiveness of integrating graph-based translation memory into the Transformer model.

Conclusion

In this technical report we have presented TranSmart, a practical interactive machine translation (IMT) system. Unlike the conventional IMT systems with the strict manner from left to right, TranSmart conducts interaction between a user and the machine in a flexible manner, and it particularly contains a translation memory technique to avoid similar mistakes occurring during translation process. We have introduced the main functions of TranSmart and key methods for implementing the features. Some instructions about how to use the TranSmart through online APIs have been described. We have also reported some evaluation results on major modules of TranSmart.

References