Translating Phrases in Neural Machine Translation
Xing Wang, Zhaopeng Tu, Deyi Xiong, Min Zhang
Introduction
Neural machine translation (NMT) has been receiving increasing attention due to its impressive translation performance Kalchbrenner and Blunsom (2013); Cho et al. (2014); Sutskever et al. (2014); Bahdanau et al. (2015); Wu et al. (2016). Significantly different from conventional statistical machine translation (SMT) Brown et al. (1993); Koehn et al. (2003); Chiang (2005), NMT adopts a big neural network to perform the entire translation process in one shot, for which an encoder-decoder architecture is widely used. Specifically, the encoder encodes a source sentence into a continuous vector representation, then the decoder uses the continuous vector representation to generate the corresponding target translation word by word.
The word-by-word generation philosophy in NMT makes it difficult to translate multi-word phrases. Phrases, especially multi-word expressions, are crucial for natural language understanding and machine translation Sag et al. (2002); Villavicencio et al. (2005) as the meaning of a phrase cannot be always deducible from the meanings of its individual words or parts. Unfortunately current NMT is essentially a word-based or character-based Chung et al. (2016); Costa-jussà and Fonollosa (2016); Luong and Manning (2016) translation system where phrases are not considered as translation units. In contrast, phrases are much better than words as translation units in SMT and have made a significant advance in translation quality. Therefore, a natural question arises: Can we translate phrases in NMT?
Recently, there have been some attempts on multi-word phrase generation in NMT Stahlberg et al. (2016b); Zhang and Zong (2016). However these efforts constrain NMT to generate either syntactic phrases or domain phrases in the word-by-word generation framework. To explore the phrase generation in NMT beyond the word-by-word generation framework, we propose a novel architecture that integrates a phrase-based SMT model into NMT. Specifically, we add an auxiliary phrase memory to store target phrases in symbolic form. At each decoding step, guided by the decoding information from the NMT decoder, the SMT model dynamically generates relevant target phrase translations and writes them to the memory. Then the NMT decoder scores phrases in the phrase memory and selects a proper phrase or word with the highest probability. If the phrase generation is carried out, the NMT decoder generates a multi-word phrase and updates its decoding state by consuming the words in the selected phrase.
Furthermore, in order to enhance the ability of the NMT decoder to effectively select appropriate target phrases, we modify the encoder of NMT to make it fit for exploring structural information of source sentences. Particularly, we integrate syntactic chunk information into the NMT encoder, to enrich the source-side representation. We validate our proposed model on the ChineseEnglish translation task. Experiment results show that the proposed model significantly outperforms the conventional attention-based NMT by 1.07 BLEU points on multiple NIST test sets.
The rest of this paper is organized as follows. Section 2 briefly introduces the attention-based NMT as background knowledge. Section 3 presents our proposed model which incorporates the phrase memory into the NMT encoder-decoder architecture, as well as the reading and writing procedures of the phrase memory. Section 4 presents our experiments on the ChineseEnglish translation task and reports the experiment results. Finally we discuss related work in Section 5 and conclude the paper in Section 6.
Background
Neural machine translation often adopts the encoder-decoder architecture with recurrent neural networks (RNN) to model the translation process. The bidirectional RNN encoder which consists of a forward RNN and a backward RNN reads a source sentence and transforms it into word annotations of the entire source sentence . The decoder uses the annotations to emit a target sentence in a word-by-word manner.
In the training phase, given a parallel sentence , NMT models the conditional probability as follows,
where is the target word emitted by the decoder at step and . The conditional probability is computed as
where is a non-linear function and is the hidden state of the decoder at step :
where is a non-linear function. Here we adopt Gated Recurrent Unit Cho et al. (2014) as the recurrent unit for the encoder and decoder. is the context vector, computed as a weighted sum of the annotations :
where is the annotation of source word and its weight is computed by the attention model.
We train the attention-based NMT model by maximizing the log-likelihood:
given the training data with bilingual sentences Cho (2015).
In the testing phase, given a source sentence , we use beam search strategy to search a target sentence that approximately maximizes the conditional probability
Approach
In this section, we introduce the proposed model which incorporates a phrase memory into the encoder-decoder architecture of NMT. Inspired by the recent work on attaching an external structure to the encoder-decoder architecture Gulcehre et al. (2016); Gu et al. (2016); Tang et al. (2016); Wang et al. (2017), we adopt a similar approach to incorporate the phrase memory into NMT.
Figure 1 shows an example. Given the generated words “President Bush emphasized that”, the model generates the next fragment either from a word generation mode or a phrase generation mode. If the model selects the word generation mode, it generates a word by the NMT decoder as in the standard NMT framework. Otherwise, it generates a multi-word phrase by enquiring a phrase memory, which is written by an SMT decoder based on the dynamic decoding information from the NMT model for each step. The trade-off between word generation mode and phrase generation mode is balanced by a weight , which is produced by a neural network based balancer.
Formally, a generated translation consists of two sets of fragments: words generated by NMT decoder and phrases generated from the phrase memory . The probability of generating is calculated by
where is the probability of generating the word (see Equation 2), is that of generating the phrase which will be described in Section 3.2, and is the decoding step to generate the corresponding fragment.
The balancing weight is produced by the balancer – a multi-layer network. The balancer network takes as input the decoding information, including the context vector , the previous decoding state and the previous generated word :
where is a sigmoid function and is the activation function. Intuitively, the weight can be treated as the estimated importance of the phrase to be generated. We expect to be high if the phrase is appropriate at the current decoding step.
We employ a source-side chunker to chunk the source sentence, and only phrases that corresponds to a source chunk are used in our model. We restrict ourselves to the well-formed chunk phrases based on the following considerations: (1) In order to take advantage of dynamic programming, we restrict ourselves to non-overlap phrases.Overlapped phrases may result in a high dimensionality in translation hypothesis representation and make it hard to employ shared fragments for efficient dynamic programming. (2) We explicitly utilize the boundary information of the source-side chunk phrases, to better guide the proposed model to adopt a target phrase at an appropriate decoding step. (3) We enable the model to exploit the syntactic categories of chunk phrases to enhance the proposed model with its selection preference for special target phrases. With these information, we enrich the context vector to enable the proposed model to make better decisions, as described below.
Following the commonly-used strategy in sequence tagging tasks Xue and Shen (2003), we allow the words in a phrase to share the same chunk tag and introduce a special tag for the beginning word. For example, the phrase “ 信息 安全 (information security)” is tagged as a noun phrase “NP”, and the tag sequence should be “NP_B NP”. Partially motivated by the work on integrating linguistic features into NMT Sennrich and Haddow (2016), we represent the encoder input as the combination of word embeddings and chunking tag embeddings, instead of word embeddings alone in the conventional NMT. The new input is formulated as follows:
2 Phrase Memory
The phrase memory stores relevant target phrases provided by an SMT model, which is trained on the same bilingual corpora. At each decoding step, the memory is firstly erased and re-written by the SMT model, the decoding of which is based on the translation information provided by the NMT model. Then, the proposed model enquires phrases along with their probabilities from the memory.
Given a partial translation generated from NMT, the SMT model picks potential phrases extracted from the translation table. The phrases are scored with multiple SMT features, including the language model score, the translation probabilities, the reordering score, and so on. Specially, the reordering score depends on alignment information between source and target words, which is derived from attention distribution produced by the NMT model Wang et al. (2017). SMT coverage vector in Wang et al. (2017) is also introduced to avoid repeat phrasal recommendations. In our work, the potential phrase is phrase with high SMT score which is defined as following:
where is a target phrase and is its corresponding source span. is a SMT feature function and is its weight. The feature weights can be tuned by the minimum error rate training (MERT) algorithm Och (2003).
This leads to a better interaction between SMT and NMT models. It should be emphasized that our memory is dynamically updated at each decoding step based on the decoding history from both SMT and NMT models.
The proposed model is very flexible, where the phrase memory can be either fully dynamically generated by an SMT model or directly extracted from a bilingual dictionary, or any other bilingual resources storing idiomatic translations or bilingual multi-word expressions, which may lead to a further improvement. Bilingual resources can be utilized in two ways: First, we can store the bilingual resources in a static memory and keep all items available to NMT in the whole decoding period. Second, we can integrate the bilingual resources into SMT and then dynamically feed them into the phrase memory.
Reading Phrase Memory
When phrases are read from the memory, they are rescored by a neural network based score function. The score function takes as input the phrase itself and decoding information from NMT ( denotes the current decoding step):
where is either an identity or a non-linear function. is the representation of phrase , which is modeled by a recurrent neural networks. Again, is the decoder state, is the lastly generated word, and is the context vector. The scores are normalized for all phrases in the phrase memory, and the probability for phrase is calculated as
The probability calculation is controlled with parameters, which are trained together with the parameters from the NMT model.
3 Training
Formally, we train both the default parameters of standard NMT and the new parameters associated with phrase generation on a set of training examples :
where is defined in Equation 7. Ideally, the trained model is expected to produce a higher balance weight and phrase probability when a phrase is selected from the memory, and lower scores in other cases.
4 Decoding
During testing, the NMT decoder generates a target sentence which consists of a mixture of words and phrases. Due to the different granularities of words and phrases, we design a variant of beam search strategy: At decoding step , we first compute for all phrases in the phrase memory and for all words in NMT vocabulary. Then the balancer outputs a balancing weight , which is used to scale the phrase and word probabilities : and . Now outputs are normalized probabilities on the concatenation of phrase memory and the general NMT vocabulary. At last, the NMT decoder generates a proper phrase or word of the highest probability.
If a target phrase in the phrase memory has the highest probability, the decoder generates the target phrase to complete the multi-word phrase generation process, and updates its decoding state by consuming the words in the selected phrase as described in Equation 3. All translation hypotheses are placed in the corresponding beams according to the number of generated target words.
Experiments
In this section, we evaluated the effectiveness of our model on the ChineseEnglish machine translation task. The training corpora consisted of about 1.25 million sentence pairsThe corpus includes LDC2002E18, LDC2003E07, LDC2003E14, Hansards portion of LDC2004T07, LDC2004T08 and LDC2005T06. with 27.9 million Chinese words and 34.5 million English words respectively. We used NIST 2006 (NIST06) dataset as development set, and NIST 2004 (NIST04), 2005 (NIST05) and 2008 (NIST08) datasets as test sets. We report experiment results with case-insensitive BLEU scoreftp://jaguar.ncsl.nist.gov/mt/resources/mteval-v11b.pl.
We compared our proposed model with two state-of-the-art systems:
Moses: a state-of-the-art phrase-based SMT system Koehn et al. (2007) with its default settings, where feature function weights are tuned by the minimum error rate training (MERT) algorithm Och (2003).
RNNSearch: an in-house implementation of the attention-based NMT system Bahdanau et al. (2015) with its default settings.
For Moses, we used the full bilingual training data to train the phrase-based SMT model and the target portion of the bilingual training data to train a 4-gram language model using KenLMhttps://kheafield.com/code/kenlm/. We ran Giza++ on the training data in both Chinese-to-English and English-to-Chinese directions and applied the “grow-diag-final” refinement rule Koehn et al. (2003) to obtain word alignments. The maximum phrase length is set to 7.
For RNNSearch, we generally followed settings in the previous work Bahdanau et al. (2015); Tu et al. (2017a, b). We only kept a shortlist of the most frequent 30,000 words in Chinese and English, covering approximately 97.7% and 99.3% of the data in the two languages respectively. We constrained our source and target sequences to have a maximum length of 50 words in the training data. The size of embedding layer of both sides was set to 620 and the size of hidden layer was set to 1000. We used a minibatch stochastic gradient descent (SGD) algorithm of size 80 together with Adadelta Zeiler (2012) to train the NMT models. The decay rates and were set as and . We clipped the gradient norm to 1.0 Pascanu et al. (2013). We also adopted the dropout technique. Dropout was applied only on the output layer and the dropout rate was set to 0.5. We used a simple beam search decoder with beam size 10 to find the most likely translation.
For the proposed model, we used a Chinese chunkerhttp://www.niuparser.com/ Zhu et al. (2015) to chunk the source-side Chinese sentences. 13 chunking tags appeared in our chunked sentences and the size of chunking tag embedding was set to 10. We used the trained phrase-based SMT to translate the source-side chunks. The top 5 translations according to their translation scores (Equation 10) were kept and among them multi-word phrases were used as phrasal recommendations for each source chunk phrase. For a source-side chunk phrase, if there exists phrasal recommendations from SMT, the output chunk tag was used as its chunking tag feature as described in Section 3.1. Otherwise, the words in the chunk were treated as general words by being tagged with the default tag. In the phrase memory, we only keep the top 7 target translations with highest SMT scores at each decoding step. We used a forward neural network with two hidden layers for both the balancer (Equation 8) and the scoring function (Equation 11). The numbers of units in the hidden layers were set to 2000 and 500 respectively. We used a backward RNN encoder to learn the phrase representations of target phrases in the phrase memory.
Table 1 reports main results of different models measured in terms of BLEU score. We observe that our implementation of RNNSearch outperforms Moses by 2.34 BLEU points. (+memory) which is the proposed model with the phrase memory obtains an improvement of 0.47 BLEU points over the baseline RNNSearch. With the source-side chunking tag feature, (+memory+chunking tag) outperforms the baseline RNNSearch by 1.07 BLEU points, showing the effectiveness of chunking syntactic categories on the selection of appropriate target phrases. From here on, we use “+memory+chunking tag” as the default setting in the following experiments if not otherwise stated.
We also check the number of translations that contain phrases generated by the proposed model, as shown in Table 2. As seen, a large portion of translations take the recommended phrases, and the number increases when the chunking tag feature is used.The numbers on NIST08 are relatively lower since part of the test set contains sentences from Web forums, which contain less multi-word expressions. Considering BLEU scores reported in Table 1, we believe that the chunking tag feature benefits the proposed model on its phrase generation.
2 Analysis on Generated Phrases
We first investigate which category of phrases is more likely to be selected by the proposed approach. There are some phrases, such as noun phrases (NPs, e.g., “national laboratory” and “vietnam airlines”) and quantifier phrases (QPs, e.g., “15 seconds” and “two weeks”) , that we expect to be favored by our approach. Statistics shown in Table 3 confirm our hypothesis. Let’s first concern all generated phrases (i.e., column “All”): most selected phrases are noun phrases (81.0%) and quantifier phrases (10.8%). Among them, 44.5% percent of them are fully correctFully correct means that the generated phrases can be retrieved in corresponding references as a whole unit.. Specifically, NPs have relative higher generation accuracy (i.e., ) while VPs have lower accuracy (i.e., ). By looking into the wrong cases, we found most errors are related to verb tense, which is the drawback of SMT models.
Concerning the newly introduced phrases that cannot be found in baseline translations (i.e., column “New”), 13.2% of generated phrases are both new and fully correct, which contribute most to the performance improvement. We can also find that most newly introduced verb phrases and quantifier phrases are not correct, the patterns of which can be well learned by word-based NMT models.
Number of Words in Generated Phrases
Table 4 lists the distribution of generated phrases based on the number of inside words. As seen, most generated phrases are short phrases (e.g., 2-gram and 3-gram phrases), which also contribute most to the new and fully correct phrases (i.e., ). Focusing on long phrases (e.g., order), most of them are newly introduced ( out of ). Unfortunately, only a few portion of these phrases are fully correct, since long phrases have higher chance to contain one or two unmatched words.
Effect of Generated Phrases on Translation Performance
Note that the proposed model benefits not only from fully matched phrases, but also from partially matched phrases. For example, the baseline system translates “ 国家 航空 暨 太空 总署” in a word-by-word manner and outputs “state aviation and space department”. The generated phrase provided by SMT is “national aviation and space administration”, but the only correct reference is “national aeronautics and space administration”. The generated phrase is not fully correct but still useful.
To directly measure the improvement obtained by the phrase generation, we replace the generated target phrases with a special symbol “NULL” in test sets. As shown in Table 5, when deleting the generated target phrases, (“+memory+chunking tag”) and (“+memory”) translation performances decrease by 2.74 BLEU points and 1.32 BLEU points respectively. Moreover, translation performances on NIST08 decrease less than those on NIST04 and NIST05 in both settings. The reason is that NIST08 which contains sentences from web data has little influence on generating target phrases which are provided from a different domain The parallel training data are mainly from news domain.. The overall results demonstrate that neural machine translation benefits from phrase translation.
3 Effect of Balancer
The balancer which is used to coordinate the phrase generation and word generation is very crucial for the proposed model. We conducted an additional experiment to validate the effectiveness of the neural network based balancer. We use the setting “+memory +chunking tag” as baseline system to conduct the experiments. In this experiment, we fixed the balancing weight (Equation 8) to 0.1 during training and testing and report the results. As shown in Table 6, we find that using the fixed value for the balancing weight (Constant () ) decreases the translation performance sharply. This demonstrates that the neural network based balancer is an essential component for the proposed model.
4 Comparison to Word-Level Recommendations and Discussions
Our approach is related to our previous work Wang et al. (2017) which integrates the SMT word-level knowledge into NMT. To make a comparison, we conducted experiments followed settings in Wang et al. (2017). The comparison results are reported in Table 7. We find that our approach is marginally better than the word-level model proposed in Wang et al. (2017) by 0.28 BLEU points.
In our approach, the SMT model translates source-side chunk phrases using the NMT decoding information. Although we use high-quality target phrases as phrasal recommendations, our approach still suffers from the errors in segmentation and chunking. For example, the target phrase “laptop computers” cannot be recommended by the SMT model if the Chinese phrase “手 提 电脑” is not chunked as a phrase unit. This is the reason why some sentences do not have corresponding phrasal recommendations (Table 2). Therefore, our approach can be further enhanced if we can reduce the error propagations from the segmenter or chunker, for example, by using n-best chunk sequences instead of the single best chunk sequence.
Additionally, we also observe that some target phrasal recommendations have been also generated by the baseline system in a word-by-word manner. These phrases, even taken as parts of final translations by the proposed model, do not lead to improvements in terms of BLEU as they have already occurred in translations from the baseline system. For example, the proposed model successfully carries out the phrase generation mode to generate a target phrase “guangdong province” (the translation of Chinese phrase “广东省”) which has appeared in the baseline system.
As external resources, e.g., bilingual dictionary, which are complementary to the SMT phrasal recommendations, are compatible with the proposed model, we believe that the proposed model will get further improvement by using external resources.
Related work
Our work is related to the following research topics on NMT:
In these studies, the generated NMT multi-word phrases are either from an SMT model or a bilingual dictionary. In syntactically guided neural machine translation (SGNMT), the NMT decoder uses phrase translations produced by the hierarchical phrase-based SMT system Hiero, as hard decoding constraints. In this way, syntactic phrases are generated by the NMT decoder Stahlberg et al. (2016b). Zhang and Zong (2016) use an SMT translation system, which is integrated an additional bilingual dictionary, to synthesize pseudo-parallel sentences and feed the sentences into the training of NMT in order to translate low-frequency words or phrases. Tang et al. (2016) propose an external phrase memory that stores phrase pairs in symbolic forms for NMT. During decoding, the NMT decoder enquires the phrase memory and properly generates phrase translations. The significant differences between these efforts and ours are 1) that we dynamically generate phrase translations via an SMT model, and 2) that at the same time we modify the encoder to incorporate structural information to enhance the capability of NMT in phrase translation.
Incorporating linguistic information into NMT
NMT is essentially a sequence to sequence mapping network that treats the input/output units, eg., words, subwords Sennrich et al. (2016), characters Chung et al. (2016); Costa-jussà and Fonollosa (2016), as non-linguistic symbols. However, linguistic information can be viewed as the task-specific knowledge, which may be a useful supplementary to the sequence to sequence mapping network. To this end, various kinds of linguistic annotations have been introduced into NMT to improve its translation performance. Sennrich and Haddow (2016) enrich the input units of NMT with various linguistic features, including lemmas, part-of-speech tags, syntactic dependency labels and morphological features. García-Martínez et al. (2016) propose factored NMT using the morphological and grammatical decomposition of the words (factors) in output units. Eriguchi et al. (2016) explore the phrase structures of input sentences and propose a tree-to-sequence attention model for the vanilla NMT model. Li et al. (2017) propose to linearize source-side parse trees to obtain structural label sequences and explicitly incorporated the structural sequences into NMT, while Aharoni and Goldberg (2017) propose to incorporate target-side syntactic information into NMT by serializing the target sequences into linearized, lexicalized constituency trees. Zhang et al. (2016) integrate topic knowledge into NMT for domain/topic adaptation.
Combining NMT and SMT
A variety of approaches have been explored for leveraging the advantages of both NMT and conventional SMT. He et al. (2016) integrate SMT features with the NMT model under the log-linear framework in order to help NMT alleviate the limited vocabulary problem Luong et al. (2015); Jean et al. (2015) and coverage problem Tu et al. (2016). Arthur et al. (2016) observe that NMT is prone to making mistakes in translating low-frequency content words and therefore attempt at incorporating discrete translation lexicons into the NMT model, to alliterate the imprecise translation problem Wang et al. (2017). Motivated by the complementary strengths of syntactical SMT and NMT, different combination schemes of Hiero and NMT have been exploited to form SGNMT Stahlberg et al. (2016a, b). Wang et al. (2017) propose an approach to incorporate the SMT model into attention-based NMT. They combine NMT posteriors with SMT word recommendations through linear interpolation implemented by a gating function which dynamically assigns the weights. Niehues et al. (2016) propose to use SMT to pre-translate the inputs into target translations and employ the target pre-translations as input sequences in NMT. Zhou et al. (2017) propose a neural system combination framework to directly combine NMT and SMT outputs. The combination of NMT and SMT has been also introduced in interactive machine translation to improve the system’s suggestion quality Wuebker et al. (2016). In addition, word alignments from the traditional SMT pipeline are also used to improve the attention mechanism in NMT Cohn et al. (2016); Mi et al. (2016); Liu et al. (2016).
Conclusion
In this paper, we have presented a novel model to translate source phrases and generate target phrase translations in NMT by integrating the phrase memory into the encoder-decoder architecture. At decoding, the SMT model dynamically generates relevant target phrases with contextual information provided by the NMT model and writes them to the phrase memory. Then the proposed model reads the phrase memory and uses the balancer to make probability estimations for the phrases in the phrase memory. Finally the NMT decoder selects a phrase from the phrase memory or a word from the vocabulary of the highest probability to generate. Experiment results on ChineseEnglish translation have demonstrated that the proposed model can significantly improve the translation performance.
Acknowledgments
We would like to thank three anonymous reviewers for their insightful comments, and also acknowledge Zhengdong Lu, Lili Mou for useful discussions. This work was supported by the National Natural Science Foundation of China (Grants No.61525205, 61373095 and 61622209).