Supervised Attentions for Neural Machine Translation

Haitao Mi, Zhiguo Wang, Abe Ittycheriah

Introduction

Neural machine translation (NMT) has gained popularity in recent two years [Bahdanau et al. (2014, Jean et al. (2015, Luong et al. (2015], especially for the attention-based models of ?).

The attention model plays a crucial role in NMT, as it shows which source word(s) the model should focus on in order to predict the next target word. However, the attention or alignment quality of NMT is still very low [Mi et al. (2016a, Tu et al. (2016].

In this paper, we alleviate the above issue by utilizing the alignments (human annotated data or machine alignments) of the training set. Given the alignments of all the training sentence pairs, we add an alignment distance cost to the objective function. Thus, we not only maximize the log translation probabilities, but also minimize the alignment distance cost. Large-scale experiments over Chinese-to-English on various test sets show that our best method for a single system improves the translation quality significantly over the large vocabulary NMT system (Section 5) and beats the state-of-the-art syntax-based system.

Neural Machine Translation

As shown in Figure 1, attention-based NMT [Bahdanau et al. (2014] is an encoder-decoder network. the encoder employs a bi-directional recurrent neural network to encode the source sentence x=(x1,...,xl){\bf{x}}=({x_{1},...,x_{l}}), where ll is the sentence length (including the end-of-sentence ⟨eos⟩\langle\text{eos}\rangle), into a sequence of hidden states h=(h1,...,hl){\bf{h}}=({h_{1},...,h_{l}}), each hih_{i} is a concatenation of a left-to-right hi→\overrightarrow{h_{i}} and a right-to-left hi←\overleftarrow{h_{i}}.

Given h{\bf h}, the decoder predicts the target translation by maximizing the conditional log-probability of the correct translation y∗=(y1∗,...ym∗){\bf y^{*}}=(y^{*}_{1},...y^{*}_{m}), where mm is the sentence length (including the end-of-sentence). At each time tt, the probability of each word yty_{t} from a target vocabulary VyV_{y} is:

where gg is a two layer feed-forward neural network over the embedding of the previous word yt−1∗y^{*}_{t-1}, and the hidden state sts_{t}. The sts_{t} is computed as:

where qq is a gated recurrent units, HtH_{t} is a weighted sum of h{\bf h}; the weights, α\alpha, are computed with a two layer feed-forward neural network rr:

We put all αt,i\alpha_{t,i} (t=1...mt=1...m, i=1...li=1...l) into a matrix A′\mathcal{A^{\prime}}, we have a matrix (alignment) like (c) in Figure 2, where each row (for each target word) is a probability distribution over the source sentence x{\bf x}.

The training objective is to maximize the conditional log-probability of the correct translation y∗y^{*} given xx with respect to the parameters θ\theta

where nn is the nn-th sentence pair (xn,y∗n)({\bf x}^{n},{\bf y^{*}}^{n}) in the training set, NN is the total number of pairs.

Alignment Component

The attentions, αt,1...αt,l\alpha_{t,1}...\alpha_{t,l}, in each step tt play an important role in NMT. However, the accuracy is still far behind the traditional MaxEnt alignment model in terms of alignment F1 score [Mi et al. (2016b, Tu et al. (2016]. Thus, in this section, we explicitly add an alignment distance to the objective function in Eq. 5. The “truth” alignments for each sentence pair can be from human annotated data, unsupervised or supervised alignments (e.g. GIZA++ [Och and Ney (2000] or MaxEnt [Ittycheriah and Roukos (2005]).

Given an alignment matrix A\mathcal{A} for a sentence pair (x,y{\bf x,y}) in Figure 2 (a), where we have an end-of-source-sentence token ⟨eos⟩=xl\langle\text{eos}\rangle=x_{l}, and we align all the unaligned target words (y3∗y^{*}_{3} in this example) to ⟨eos⟩\langle\text{eos}\rangle, also we force ym∗y^{*}_{m} (end-of-target-sentence) to be aligned to xlx_{l} with probability one. Then we conduct two transformations to get the probability distribution matrices ((b) and (c) in Figure 2).

The first transformation simply normalizes each row. Figure 2 (b) shows the result matrix A∗\mathcal{A^{*}}. The last column in red dashed lines shows the alignments of the special end-of-sentence token ⟨eos⟩\langle\text{eos}\rangle.

2 Smoothed Transformation

Given the original alignment matrix A\mathcal{A}, we create a matrix A∗\mathcal{A^{*}} with all points initialized with zero. Then, for each alignment point At,i=1\mathcal{A}_{t,i}=1, we update A∗\mathcal{A^{*}} by adding a Gaussian distribution, g(μ,σ)g(\mu,\sigma), with a window size ww (tt-ww, … tt … tt+ww). Take the A1,1=1\mathcal{A}_{1,1}=1 for example, we have A∗1,1\mathcal{A^{*}}_{1,1} += 11, A∗1,2\mathcal{A^{*}}_{1,2} += 0.610.61, and A∗1,3\mathcal{A^{*}}_{1,3} += 0.140.14 with ww=2, g(μ,σ)g(\mu,\sigma)=g(0,1)g(0,1). Then we normalize each row and get (c). In our experiments, we use a shape distribution, where σ\sigma = 0.50.5.

3 Objectives

Alignment Objective: Given the “true” alignment A∗\mathcal{A^{*}}, and the machine attentions A′\mathcal{A^{\prime}} produced by NMT model, we compute the Euclidean distance bewteen A∗\mathcal{A^{*}} and A′\mathcal{A^{\prime}}.

NMT Objective: We plug Eq. 6 to Eq. 5, we have

There are two parts: translation and alignment, so we can optimize them jointly, or separately (e.g. we first optimize alignment only, then optimize translation). Thus, we divide the network in Figure 1 into alignment A and translation T parts:

A: all networks before the hidden state sts_{t},

If we only optimize A, we keep the parameters in T unchanged. We can also optimize them jointly J. In our experiments, we test different optimization strategies.

Related Work

In order to improve the attention or alignment accuracy, ?) adapted the agreement-based learning [Liang et al. (2006, Liang et al. (2008], and introduced a combined objective that takes into account both translation directions (source-to-target and target-to-source) and an agreement term between the two alignment directions. By contrast, our approach directly uses and optimizes NMT parameters using the “supervised” alignments.

Experiments

We run our experiments on Chinese to English task. The training corpus consists of approximately 5 million sentences available within the DARPA BOLT Chinese-English task. The corpus includes a mix of newswire, broadcast news, and webblog. We do not include HK Law, HK Hansard and UN data. The Chinese text is segmented with a segmenter trained on CTB data using conditional random fields (CRF). Our development set is the concatenation of several tuning sets (GALE Dev, P1R6 Dev, and Dev 12) initially released under the DARPA GALE program. The development set is 4491 sentences in total. Our test sets are NIST MT06 (1664 sentences) , MT08 news (691 sentences), and MT08 web (666 sentences).

For all NMT systems, the full vocabulary size of the training set is 300k300k. In the training procedure, we use AdaDelta [Zeiler (2012] to update model parameters with a mini-batch size 80. Following ?), the output vocabulary for each mini-batch or sentence is a sub-set of the full vocabulary. For each source sentence, the sentence-level target vocabularies are union of top 2k2k most frequent target words and the top 10 candidates of the word-to-word/phrase translation tables learned from ‘fast_align’ [Dyer et al. (2013]. The maximum length of a source phrase is 4. In the training time, we add the reference in order to make the translation reachable.

The Cov. LVNMT system is a re-implementation of the enhanced NMT system of ?), which employs a coverage embedding model and achieves better performance over large vocabulary NMT ?). The coverage embedding dimension of each source word is 100.

Following ?), we dump the alignments, attentions, for each sentence, and replace UNKs with the word-to-word translation model or the aligned source word.

Our traditional SMT system is a hybrid syntax-based tree-to-string model [Zhao and Al-onaizan (2008], a simplified version of the joint decoding [Liu et al. (2009, Cmejrek et al. (2013]. We parse the Chinese side with Berkeley parser, and align the bilingual sentences with GIZA++ and MaxEnt. and extract Hiero and tree-to-string rules on the training set. Our language models are trained on the English side of the parallel corpus, and on monolingual corpora (around 10 billion words from Gigaword (LDC2011T07).We tune our system with PRO [Hopkins and May (2011] to minimize (Ter- Bleu)/2 The metric used for optimization in this work is (Ter-Bleu)/2 to prevent the system from using sentence length alone to impact Bleu or Ter. Typical SMT systems use target word count as a feature and it has been observed that Bleu can be optimized by tweaking the weighting of the target word count with no improvement in human assessments of translation quality. Conversely, in order to optimize Ter shorter sentences can be produced. Optimizing the combination of metrics alleviates this effect [Arne Mauser and Ney (2008]. on the development set.

2 Translation Results

Table 1 shows the translation results of all systems. The syntax-based statistical machine translation model achieves an average (Ter-Bleu)/2 of 13.36 on three test sets. The Cov. LVNMT system achieves an average (Ter-Bleu)/2 of 14.24, which is about 0.9 points worse than Tree-to-string SMT system. Please note that all systems are single systems. It is highly possible that ensemble of NMT systems with different random seeds can lead to better results over SMT.

Zh →\rightarrow En (one direction of GIZA++),

GDFA (the “grow-diag-final-and” heuristic merge of both directions of GIZA++),

MaxEnt (trained on 67k67k hand-aligned sentences).

The alignment quality improves from Zh →\rightarrow En to MaxEnt. We also test different optimization strategies: J (jointly), A (alignment only), and T (translation model only). A combination, A →\rightarrow T, shows that we optimize A only first, then we fix A and only update T part. Gau. denotes the smoothed transformation (Section 3.2). Only the last row uses the smoothed transformation, all others use the simple transformation.

Experimental results in Table 1 show some interesting results. First, with the same alignment, J joint optimization works best than other optimization strategies (lines 3 to 6). Unfortunately, breaking down the network into two separate parts (A and T) and optimizing them separately do not help (lines 3 to 5). We have to conduct joint optimization J in order to get a comparable or better result (lines 3, 5 and 6) over the baseline system.

Second, when we change the training alignment seeds (Zh →\rightarrow En, GDFA, and MaxEnt) NMT model does not yield significant different results (lines 6 to 8).

Third, the smoothed transformation (J + Gau.) gives some improvements over the simple transformation (the last two lines), and achieves the best result (1.2 better than LVNMT, and 0.3 better than Tree-to-string). In terms of Bleu scores, we conduct the statistical significance tests with the sign-test of ?), the results show that the improvements of our J + Gau. over LVNMT are significant on three test sets (p<0.01p<0.01).

At last, the brevity penalty (BP) consistently gets better after we add the alignment cost to NMT objective. Our alignment objective adjusts the translation length to be more in line with the human references accordingly.

3 Alignment Results

Table 2 shows the alignment F1 scores on the alignment test set (447 hand aligned sentences). The MaxEnt model is trained on 67k67k hand-aligned sentences, and achieves an F1 score of 75.96. For NMT systems, we dump the alignment matrixes and convert them into alignments with following steps. For each target word, we sort the alphas and add the max probability link if it is higher than 0.2. If we only tune the alignment component (A in line 3), we improve the alignment F1 score from 45.76 to 47.87. And we further boost the score to 50.97 by tuning alignment and translation jointly (J in line 7). Interestingly, the system using MaxEnt produces more alignments in the output, and results in a higher recall. This suggests that using MaxEnt can lead to a sharper attention distribution, as we pick the alignment links based on the probabilities of attentions, the sharper the distribution is, more links we can pick. We believe that a sharp attention distribution is a great property of NMT.

Again, the best result is J + Gau. in the last row, which significantly improves the F1 by 5 points over the baseline Cov. LVNMT system. When we use MaxEnt alignments, J + Gau. smoothing gives us about 1.7 points gain over J system. So it looks interesting to run another J + Gau. over GDFA alignment.

Together with the results in Table 1, we conclude that adding the alignment cost to the training objective helps both translation and alignment significantly.

Conclusion

In this paper, we utilize the “supervised” alignments, and put the alignment cost to the NMT objective function. In this way, we directly optimize the attention model in a supervised way. Experiments show significant improvements in both translation and alignment tasks over a very strong LVNMT system.

Acknowledgment

We thank the anonymous reviewers for useful comments.

References