Modeling Past and Future for Neural Machine Translation

Zaixiang Zheng, Hao Zhou, Shujian Huang, Lili Mou, Xinyu Dai, Jiajun Chen, Zhaopeng Tu

Introduction

Neural machine translation (NMT) generally adopts an encoder-decoder framework [Kalchbrenner and Blunsom (2013, Cho et al. (2014, Sutskever et al. (2014], where the encoder summarizes the source sentence into a source context vector, and the decoder generates the target sentence word-by-word based on the given source. During translation, the decoder implicitly serves several functionalities at the same time:

Building a language model over the target sentence for translation fluency (Lm).

Acquiring the most relevant source-side information to generate the current target word (Present).

Maintaining what parts in the source have been translated (Past) and what parts have not (Future).

However, it may be difficult for a single recurrent neural network (RNN) decoder to accomplish these functionalities simultaneously. A recent successful extension of NMT models is the attention mechanism [Bahdanau et al. (2015, Luong et al. (2015], which makes a soft selection over source words and yields an attentive vector to represent the most relevant source parts for the current decoding state. In this sense, the attention mechanism separates the Present functionality from the decoder RNN, achieving significant performance improvement.

In addition to Present, we address the importance of modeling Past and Future contents in machine translation. The Past contents indicate translated information, whereas the Future contents indicate untranslated information, both being crucial to NMT models, especially to avoid under-translation and over-translation [Tu et al. (2016]. Ideally, Past grows and Future declines during the translation process. However, it may be difficult for a single RNN to explicitly model the above processes.

In this paper, we propose a novel neural machine translation system that explicitly models Past and Future contents with two additional RNN layers. The RNN modeling the Past contents (called Past layer) starts from scratch and accumulates the information that is being translated at each decoding step (i.e., the Present information yielded by attention). The RNN modeling the Future contents (called Future layer) begins with holistic source summarization, and subtracts the Present information at each step. The two processes are guided by proposed auxiliary objectives. Intuitively, the RNN state of the Past layer corresponds to source contents that have been translated at a particular step, and the RNN state of the Future layer corresponds to source contents of untranslated words. At each decoding step, Past and Future together provide a full summarization of the source information. We then feed the Past and Future information to both the attention model and decoder states. In this way, our proposed mechanism not only provides coverage information for the attention model, but also gives a holistic view of the source information at each time.

We conducted experiments on Chinese-English, German-English, and English-German benchmarks. Experiments show that the proposed mechanism yields 2.7, 1.7, and 1.1 improvements of BLEU scores in three tasks, respectively. In addition, it obtains an alignment error rate of 35.90%, significantly lower than the baseline (39.73%) and the coverage model (38.73%) by ?). We observe that in traditional attention-based NMT, most errors occur due to over- and under-translation, which is probably because the decoder RNN fails to keep track of what has been translated and what has not. Our model can alleviate such problems by explicitly modeling Past and Future contents.

Motivation

In this section, we first introduce the standard attention-based NMT, and then motivate our model by several empirical findings.

The attention mechanism, proposed in ?), yields a dynamic source context vector for the translation at a particular decoding step, modeling Present information as described in Section 1. This process is illustrated in Figure 1.

Formally, let x={x1,…,xI}{\mathbf{x}}=\{x_{1},\dots,x_{I}\} be a given input sentence. The encoder RNN—generally implemented as a bi-directional RNN [Schuster and Paliwal (1997]—transforms the sentence to a sequence of annotations with \mathbf{h}_{i}=\big{[}\overrightarrow{\bf h}_{i};\overleftarrow{\bf h}_{i}\big{]} being the annotation of xix_{i}. (h→i\overrightarrow{\bf h}_{i} and h←i\overleftarrow{\bf h}_{i} refer to RNN’s hidden states in both directions.)

Based on the source annotations, another decoder RNN generates the translation by predicting a target word yty_{t} at each time step tt:

where g(⋅)g(\cdot) is a non-linear activation, and st\mathbf{s}_{t} is the decoding state for time step tt, computed by

Here f(⋅)f(\cdot) is an RNN activation function, e.g., the Gated Recurrent Unit [Cho et al. (2014, GRU] and Long Short-Term Memory [Hochreiter and Schmidhuber (1997, LSTM]. ct\mathbf{c}_{t} is a vector summarizing relevant source information. It is computed as a weighted sum of the source annotations

where the weights (αt,i\alpha_{t,i} for i=1⋯ ,Ii=1\cdots,I) are given by the attention mechanism:

Here, a(⋅)a(\cdot) is a scoring function measuring the degree to which the decoding state and source information match to each other.

Intuitively, the attention-based decoder selects source annotations that are most relevant to the decoder state, based on which the current target word is predicted. In other words, ct\mathbf{c}_{t} is some source information for the Present translation.

The decoder RNN is initialized with the summarization of the entire source sentence \big{[}\overrightarrow{\bf h}_{I};\overleftarrow{\bf h}_{1}\big{]}, given by

After we analyze existing attention-based NMT in detail, our intuition arises as follows. Ideally, with the source summarization in mind, after generating each target word yty_{t} from the source contents ct\mathbf{c}_{t}, the decoder should keep track of (1) translated source contents by accumulating ct\mathbf{c}_{t}, and (2) untranslated source contents by subtracting ct\mathbf{c}_{t} from the source summarization. However, such information is not well learned in practice, as there lacks explicit mechanisms to maintain translated and untranslated contents. Evidence show that attention-based NMT still suffers from serious over- and under-translation problems [Tu et al. (2016, Tu et al. (2017b]. Examples of under-translation are shown in Table 1a.

Another piece of evidence also shows the decoder may lack a holistic view of the source information, explained as below. We conduct a pilot experiment by removing the initialization of the RNN decoder. If the “holistic” context is well exploited by the decoder, translation performance would significantly decrease without the initialization. As shown in Table 1b, however, translation performance only decreases slightly after we remove the initialization. This indicates NMT decoders do not make full use of source summarization ; that the initialization only helps the prediction at the beginning of the sentence. We attribute the vanishing of such signals to the overloaded use of decoder states (e.g., Lm, Past, and Future functionalities), and hence we propose to explicitly model the holistic source summarization by Past and Future contents at each decoding step.

Related Work

Our research is built upon an attention-based sequence-to-sequence model [Bahdanau et al. (2015], but is also related to coverage modeling, future modeling, and functionality separation. We discuss these topics in the following.

?) and ?) maintain a coverage vector to indicate which source words have been translated and which source words have not. These vectors are updated by accumulating attention probabilities at each decoding step, which provides an opportunity for the attention model to distinguish translated source words from untranslated ones. Viewing coverage vectors as a (soft) indicator of translated source contents, we take one step further following this idea. We model translated and untranslated source contents by directly manipulating the attention vector (i.e., the source contents that are being translated) instead of attention probability (i.e., the probability of a source word being translated).

In addition, we explicitly model both translated (with Past-RNN) and untranslated (with Future-RNN) instead of using a single coverage vector to indicate translated source words. Another difference with ?) is that the Past and Future contents in our model are fed to not only the attention mechanism but also the decoder’s states.

In the context of semantic-level coverage, ?) propose a memory-enhanced decoder and ?) propose a memory-enhanced attention model. Both implement the memory with a Neural Turing Machine [Graves et al. (2014], in which the reading and writing operations are expected to erase translated contents and highlight untranslated contents. However, their models lack an explicit objective to guide such intuition, which is one of the key ingredients for the success in this work. In addition, we use two separate layers to explicitly model translated and untranslated contents, which is another distinguishing feature of the proposed approach.

Standard neural sequence decoders generate target sentences from left to right, thus failing to estimate some desired properties in the future (e.g., the length of target sentence). To address this problem, actor-critic algorithms are employed to predict future properties [Li et al. (2017, Bahdanau et al. (2017]; in their models, an interpolation of the actor (the standard generation policy) and the critic (a value function that estimates the future values) is used for decision making. Concerning the future generation at each decoding step, ?) guide the decoder’s hidden states to not only generate the current target word, but also predict the target words that remain untranslated. Along the direction of future modeling, we introduce a Future layer to maintain the untranslated source contents, which is updated at each decoding step by subtracting the source content being translated (i.e., attention vector) from the last state (i.e., the untranslated source content so far).

Recent work has revealed that the overloaded use of representations makes model training difficult, and such problem can be alleviated by explicitly separating these functions [Reed and Freitas (2015, Ba et al. (2016, Miller et al. (2016, Gulcehre et al. (2016, Rocktäschel et al. (2017]. For example, ?) separate the functionality of look-up keys and memory contents in memory networks [Sukhbaatar et al. (2015]. ?) propose a key-value-predict attention model, which outputs three vectors at each step: the first is used to predict the next-word distribution; the second serves as the key for decoding; and the third is used for the attention mechanism. In this work, we further separate Past and Future functionalities from the decoder’s hidden representations.

Modeling Past and Future for Neural Machine Translation

In this section, we describe how to separate Past and Future functions from decoding states. We introduce two additional RNN layers (Figure 2):

Future Layer (Section 4.1) encodes source contents to be translated.

Past Layer (Section 4.2) encodes translated source contents.

Let us take y={y1,y2,y3,y4}\mathbf{y}=\{y_{1},y_{2},y_{3},y_{4}\} as an example of the target sentence. The initial state of Future layer is a summarization of the whole source sentence, indicating that all source contents need to be translated. The initial state of Past layer is a all-zero vector, indicating no source content is yet translated.

After c1\mathbf{c}_{1} is obtained by the attention mechanism, we (1) update the Future layer by “subtracting” c1\mathbf{c}_{1} from the previous state, and (2) update the Past layer state by “adding” c1\mathbf{c}_{1} to the previous state. The two RNN states are updated as described above at every step of generating y1y_{1}, y2y_{2}, y3y_{3}, and y4y_{4}. In this way, at each time step, the Future layer encodes source contents to be translated in the future steps, while the Past layer encodes translated source contents up to the current step.

The advantages of Past and Future layers are two-fold. First, they provide coverage information, which is fed to the attention model and guides NMT systems to pay more attention to untranslated source contents. Second, they provide a holistic view of the source information, since we would anticipate “Past + Future = Holistic.” We describe them in detail in the rest of this section.

Formally, the Future layer is a recurrent neural network (the first gray layer in Figure 2) , and its state at time step tt is computed by

When calculating attention context at time step tt, we feed the attention model with the Future state from the last time step, which encodes source contents to be translated. We rewrite Equation 4 as

After obtaining attention context ct\mathbf{c}_{t}, we update Future states via Equation 6, and feed both of them to decoder states:

where ct\mathbf{c}_{t} encodes the source context of the present translation, and stF\mathbf{s}_{t}^{F} encodes source context on future translation.

We design several variants of RNN activation functions to better model the subtractive operation (Figure 3):

A natural choice is standard GRU,Our work focuses on GRU, but can be applied to any RNN architectures such as LSTM. which learns subtraction directly from the data:

where rt\mathbf{r}_{t} is a reset gate determining the combination of the input with the previous state, and ut\mathbf{u}_{t} is an update gate defining how much of the previous state to keep around. The standard GRU uses a feed-forward neural network (Equation 10) to model the subtraction without any explicit operation, which may lead to the difficulty of the training.

Compared with Equation 10, the differences between GRU-i and standard GRU are

The reset gate rt\mathbf{r}_{t} is used to control the amount of information flowing from inputs instead of from the previous state st−1F\mathbf{s}_{t-1}^{F}.

Note that for both GRU-oo and GRU-ii, we leave enough freedom for GRU to decide the extent of integrating with subtraction operations. In other words, the information subtraction is “soft.”

2 Modeling Past

Formally, the Past layer is another recurrent neural network (the second gray layer in Figure 2), and its state at time step tt is calculated by

We feed the Past state from last time step to both attention model and decoder state:

3 Modeling Past and Future

We integrate Past and Future layers together in our final model (Figure 2):

In this way, both of the attention model and decoder state are aware of what has been translated, and what has not yet.

4 Learning

We introduce additional loss functions to estimate the semantic subtraction and addition, which guide the training of the Future layer and Past layer, respectively.

In other words, we explicitly guide the Future layer by this subtractive loss, expecting ΔtF\mathbf{\Delta}_{t}^{F} to be discriminative of the current word yty_{t}.

Likewise, we introduce another loss function to measure the information incrementation of the Past layer. Notice that ΔtP=stP−st−1P≈ct\mathbf{\Delta}_{t}^{P}=\mathbf{s}_{t}^{P}-\mathbf{s}_{t-1}^{P}\approx\mathbf{c}_{t}, which is defined similar to ΔtF\mathbf{\Delta}_{t}^{F} except a minus sign. In this way, we can reasonably assume the Future and Past layers are indeed doing subtraction and addition, respectively.

We train the proposed model θ^\hat{\theta} on a set of training examples {[xn,yn]}n=1N\{\left[{\bf x}^{n},{\bf y}^{n}\right]\}_{n=1}^{N}, and the training objective is

Experiments

We conduct experiments on Chinese-English (Zh-En), German-English (De-En), and English-German (En-De) translation tasks.

For Zh-En, the training set consists of 1.6m sentence pairs, which are extracted from the LDC corporaThe corpora includes LDC2002E18, LDC2003E07, LDC2003E14, Hansards portion of LDC2004T07, LDC2004T08 and LDC2005T06. The NIST 2003 (MT03) dataset is our development set; the NIST 2002 (MT02), 2004 (MT04), 2005 (MT05), 2006 (MT06) datasets are test sets. We also evaluate the alignment performance on the standard benchmark of ?), which contains 900 manually aligned sentence pairs. We measure the alignment quality with the alignment error rate [Och and Ney (2003].

For De-En and En-De, we conduct experiments on the WMT17 [Bojar et al. (2017] corpus. The dataset consists of 5.6M sentence pairs. We use newstest2016 as our development set, and newstest2017 as our testset. We follow ?) to segment both German and English words into subwords using byte-pair encoding [Sennrich et al. (2016, BPE].

We measure the translation quality with BLEU scores [Papineni et al. (2002]. We use the multi-bleu script for Zh-En https://github.com/moses-smt/mosesdecoder/blob/master/scripts/generic/multi-bleu.perl, and the multi-bleu-detok script for De-En and En-De https://github.com/EdinburghNLP/nematus/blob/master/data/multi-bleu-detok.perl.

We use the Nematus https://github.com/EdinburghNLP/nematus [Sennrich et al. (2017b], implementing a baseline translation system, RNNSearch. For Zh-En, we limit the vocabulary size to 30K. For De-En and En-De, the number of joint BPE operations is 90,000. We use the total BPE vocabulary for each side.

We tie the weights of the target-side embeddings and the output weight matrix [Press and Wolf (2017] for De-En. All out-of-vocabulary words are mapped to a special token UNK.

We train each model with sentences of length up to 50 words in the training data. The dimension of word embeddings is 512, and all hidden sizes are 1024. In training, we set the batch size as 80 for Zh-En, and 64 for De-En and En-De. We set the beam size as 12 in testing. We shuffle the training corpus after each epoch.

We use Adam [Kingma and Ba (2014] with annealing [Denkowski and Neubig (2017] as our optimization algorithm. We set the initial learning rate as 0.0005, which halves when the validation cross-entropy does not decrease.

For the proposed model, we use the same setting with the baseline model. The Future and Past layer sizes are 1024. We employ a two-pass strategy for training the proposed model, which has proven useful to ease training difficulty when the model is relatively complicated [Shen et al. (2016, Wang et al. (2017, Wang et al. (2018]. Model parameters shared with the baseline are initialized by the baseline model.

1 Results on Chinese-English

We first evaluate the proposed model on the Chinese-English translation and alignment tasks.

Table 2 shows the translation performances on Chinese-English. Clearly the proposed approach significantly improves the translation quality in all cases, although there are still considerable differences among different variants.

(Rows 1-4). All the activation functions for the Future layer obtain BLEU socre improvements: GRU +0.52, GRU-oo +1.03, and GRU-ii +1.12. Specifically, GRU-oo is better than a regular GRU for its minus operation, and GRU-ii is the best, which shows that our elaborately designed architecture is more proper for modeling the decreasing phenomenon of the future semantics.

Adding subtractive loss gives an extra 0.68 BLEU score improvement, which indicates that adding g is beneficial guided objective for Frnn to learn the minus operation.

(Rows 5-6). We observe the same trend on introducing Past layer: using it alone achieves a significant improvement ( +1.19), and with the additional objective it further improves the translation performance ( +0.57).

(Rows 7-8). The model’s final architecture outperforms our intermediate models (1-6) by combining Frnn and Prnn. By further separating the functionaries of past contents modeling and language modeling into different neural components, the final model is more flexible, obtaining a 0.91 BLEU improvement over the best intermediate model (Row 4) and an improvement of 2.71 BLEU points over the RNNSearch baseline.

(Rows 9-11). We also conduct experiments with multi-layer decoders [Wu et al. (2016] to see whether NMT system can automatically model the translated and untranslated contents with additional decoder layers (Rows 9-10). However, we find that the performance is not improved using a two-layer decoder (Row 9), until a deeper version (three-layer decoder, Row 10) is used. This indicates that enhancing performance is non-trivial by simply adding more RNN layers into the decoder without any explicit instruction, which is consistent with the observation of ?)

Our model also outperforms the word-level Coverage [Tu et al. (2016], which considers the coverage information of the source words independently. Our proposed model can be regarded as a high-level coverage model, which captures higher level coverage information, and gives more specific signals for the decision of attention and target prediction. Our model is more deeply involved in generating target words, by being fed not only to the attention model as in ?), but also to the decoder state.

1.2 Subjective Evaluation

Following ?), we conduct subjective evaluations to validate the benefit of modeling Past and Future (Table 3). Four human evaluators are asked to evaluate the translations of 100 source sentences, which are randomly sampled from the testsets without knowing from which system the translation is selected. For the Base system, 1.7% of the source words are over-translated and 8.8% are under-translated. Our proposed model alleviates these problems by explicitly modeling the dynamic source contents by Past and Future layers, reducing 11.8% and 35.2% of over-translation and under-translation errors, respectively. The proposed model is especially effective for alleviating the under-translation problem, which is a more serious translation problem for NMT systems, and is mainly caused by lacking necessary coverage information [Tu et al. (2016].

1.3 Alignment Quality

Table 4 lists the alignment performances of our proposed model. We find that the Coverage model do improve attention model. But our model can produce much better alignments compared to the word level coverage [Tu et al. (2016]. Our model distinguishes the Past and Future directly, which is a higher level coverage mechanism than the word coverage model.

2 Results on German-English

We also evaluate our model on the WMT17 benchmarks for both De-En and En-De. As shown in Table 5, our baseline gives comparable BLEU scores to the state-of-the-art NMT systems of WMT17. Our proposed model improves the strong baseline on both De-En and En-De. This shows that our proposed model work well across different language pairs. ?) and ?) obtain higher BLEU scores than our model, because they use additional large scaled synthetic data (about 10M) for training. It maybe unfair to compare our model to theirs directly.

3 Analysis

We conduct analyses on Zh-En, to better understand our model from different perspectives .

As shown in Table 6, the baseline model (Base) has 80M parameters. A single Future or Past layer introduces 15M to 17M parameters, and the corresponding objective introduces 18M parameters. In this work, the most complex model introduces 65M parameters, which leads to a relatively slower training speed. However, our proposed model does not significant slow down the decoding speed. The most time consuming part is the calculation of the subtraction and addition losses. As we show in the next paragraph, our system works well by only using the losses in training, which further improve the decoding speed of our model.

Adding subtraction and addition loss functions helps in twofold: (1) guiding the training of the proposed subtraction and addition operation, and (2) enabling better reranking of generated candidates in testing. Table 7 lists the improvements from the two perspectives. When applied only in training, the two loss functions lead to an improvement of 0.48 BLEU points by better modeling subtraction and addition operations. On top of that, reranking with Future and Past loss scores in testing further improves the performance by +0.99 BLEU points.

The baseline model does not obtain abundant accuracy improvement by feeding the source summarization into the decoder (Table 1). We also experiment to not feed the source summarization into the decoder of the proposed model, which leads to a significant BLEU score drop on Zh-En. This shows that our proposed model better use the source summarization with explicitly modeling the Future compared to the conventional encoder-decoder baseline.

We also compare the translation cases for the baseline, word level coverage and our proposed models. As shown in Table 9, our baseline system suffers from the over-translation problems (case 1), which is consistent with the results of human evaluation (Section 3). The Base system also incorrectly translates “the royal family” into “the people of hong kong”, which is totally irrelevant here. We attribute the former case to the lack of untranslated future modeling, and the latter one to the overloaded use of the decoder state where the language modeling of the decoder leads to the fluent but wrong predictions. In contrast, the proposed approach almost address the errors in these cases.

Conclusion

Modeling source contents well is crucial for encoder-decoder based NMT systems. However, current NMT models suffer from distinguishing translated and untranslated translation contents, due to the lack of explicitly modeling past and future translations. In this paper, we separate Past and Future functionalities from decoder states, which can maintain a dynamical yet holistic view of the source content at each decoding step. Experimental results show that the proposed approach significantly improves translation performances across different language pairs. With better modeling of past and future translations, our approach performs much better than the standard attention-based NMT, reducing the errors of under and over translations.

Acknowledgement

We would like to thank the anonymous reviewers for their insightful comments. Shujian Huang is the corresponding author. This work is supported by the National Science Foundation of China (No. 61672277, 61772261), the Jiangsu Provincial Research Foundation for Basic Research (No. BK20170074).

References