Accurate Word Alignment Induction from Neural Machine Translation
Yun Chen, Yang Liu, Guanhua Chen, Xin Jiang, Qun Liu
Introduction
The task of word alignment is to find lexicon translation equivalents from parallel corpus Brown et al. (1993). It is one of the fundamental tasks in natural language processing (NLP) and is widely studied by the community Dyer et al. (2013); Brown et al. (1993); Liu and Sun (2015). Word alignments are useful in many scenarios, such as error analysis Ding et al. (2017); Li et al. (2019), the introduction of coverage and fertility models Tu et al. (2016), inserting external constraints in interactive machine translation Hasler et al. (2018); Chen et al. (2020) and providing guidance for human translators in computer-aided translation Dagan et al. (1993).
Word alignment is part of the pipeline in statistical machine translation (Koehn et al., 2003, SMT), but is not necessarily needed for neural machine translation (Bahdanau et al., 2015, NMT). The attention mechanism in NMT does not functionally play the role of word alignments between the source and the target, at least not in the same way as its analog in SMT. It is hard to interpret the attention activations and extract meaningful word alignments especially from Transformer Garg et al. (2019). As a result, the most widely used word alignment tools are still external statistical models such as Fast-Align Dyer et al. (2013) and GIZA++ Brown et al. (1993); Och and Ney (2003).
Recently, there is a resurgence of interest in the community to study word alignments for the Transformer Ding et al. (2019); Li et al. (2019). One simple solution is Naive-Att, which induces word alignments from the attention weights between the encoder and decoder. The next target word is aligned with the source word that has the maximum attention weight, as shown in Fig. 1. However, such schedule only captures noisy word alignments Ding et al. (2019); Garg et al. (2019). One of the major problems is that it induces alignment before observing the to-be-aligned target token Peter et al. (2017); Ding et al. (2019). Suppose for the same source sentence, there are two alternative translations that diverge at decoding step , generating and which respectively correspond to different source words. Presumably, the source word that is aligned to and should change correspondingly. However, this is not possible under the above method, because the alignment scores are computed before prediction of or .
To alleviate this problem, some researchers modify the transformer architecture by adding alignment modules that predict the to-be-aligned target token Zenkel et al. (2019, 2020) or modify the training loss by designing an alignment loss computed with full target sentence Garg et al. (2019); Zenkel et al. (2020). Others argue that using only attention weights is insufficient for generating clean word alignment and propose to induce alignments with feature importance measures, such as leave-one-out measures Li et al. (2019) and gradient-based measures Ding et al. (2019). However, all previous work induces alignment for target word at step , when is the decoder output.
In this work, we propose to induce alignment for target word at step rather than at step as in previous work. The motivation behind this is that the hidden states in step are computed taking word as the input, thus they can incorporate the information of the to-be-aligned target token easily. Following this idea, we present Shift-Att and Shift-AET, two simple yet effective methods for word alignment induction. Our contributions are threefold:
We introduce Shift-Att (see Fig. 1), a pure interpretation method to induce alignments from attention weights of vanilla Transformer. Shift-Att is able to reduce the Alignment Error Rate (AER) by 7.0-10.2 points over Naive-Att and 5.5-7.9 points over Fast-Align on three publicly available datasets, demonstrating that if the correct decoding step and layer are chosen, attention weights in vanilla Transformer are sufficient for generating accurate word alignment interpretation.
We further propose Shift-AET , which extracts alignments from an additional alignment module. The module is tightly integrated into vanilla Transformer and trained with supervision from symmetrized Shift-Att alignments. Shift-AET does not affect the translation accuracy and significantly outperforms GIZA++ by 1.4-4.8 AER points in our experiments.
We compare our methods with Naive-Att on dictionary-guided decoding Alkhouli et al. (2018), an alignment-related downstream task. Both methods consistently outperform Naive-Att, demonstrating the effectiveness of our methods in such alignment-related NLP tasks.
Background
Let and be source and target sentences. Neural machine translation models the target sentence given the source sentence as :
In this paper, we use Transformer Vaswani et al. (2017) to implement the NMT model. Transformer is an encoder-decoder model that only relies on attention. Each decoder layer attends to the encoder output with multi-head attention. We refer to the original paper Vaswani et al. (2017) for more model details.
2 Alignment by Attention
and extract word alignments with maximum a posterior strategy following Garg et al. (2019):
where indicates is aligned to . We call this approach Naive-Att. Garg et al. (2019) show that attention weights from the penultimate layer, i.e., , can induce the best alignments.
Although simple to implement, this method fails to obtain satisfactory word alignments Ding et al. (2019); Garg et al. (2019). First of all, instead of the relevance between and , measures the relevance between decoder hidden state and encoder output . Considering that the decoder input is and the output is at step , may better represent instead of , especially for bottom layers. Second, since is computed before observing , it becomes difficult for it to induce the aligned source token for the target token , as discussed in Section 1.
As a result, it is necessary to develop novel methods for alignment induction. This method should be able to (i) take into account the relationship of , and , and (ii) adapt the alignment induction with the to-be-aligned target token.
Method
In this section, we propose two novel alignment induction methods Shift-Att and Shift-AET. Both methods adapt the alignment induction with the to-be-aligned target token by computing alignment scores at the step when the target token is the decoder input.
Naive-Att Garg et al. (2019) induces alignment for target token at step when is the decoder output and defines the alignment score matrix with Eq. 2. They find the best layer to extract alignments by evaluating the AER of all layers on the test set.
We instead propose to induce alignment for target token at step when is the decoder input. We define the alignment score matrix as:
This is because measures the relevance between and , and we use and to represent and respectively. With the alignment score matrix , we can extract word alignments using Eq. 5. We call this method Shift-Att. Fig. 1 shows an alignment induction example to compare Naive-Att and Shift-Att.
Shift-Att uses to represent the to-be-aligned target token while Naive-Att uses . We argue using is better. First, at bottom layers, we hypothesize that could better represent the decoder input than output . Therefore we can use with small to represent . Second, is computed after observing , indicating that Shift-Att is able to adapt the alignment induction with the to-be-aligned target token.
Our proposed method involves inducing alignments from source-to-target and target-to-source vanilla Transformer models. Following Zenkel et al. (2019), we merge bidirectional alignments using the grow diagonal heuristic Koehn et al. (2005).
Layer Selection Criterion
To select the best layer to induce alignments, we propose a surrogate layer selection criterion without manually labelled word alignments. Experiments show that this criterion correlates well with the AER metric.
Given parallel sentence pairs , we train a source-to-target model and a target-to-source model . We assume that the word alignments extracted from these two models should agree with each other Cheng et al. (2016). Therefore, we evaluate the quality of the alignments by computing the AER score on the validation set with the source-to-target alignments as the hypothesis and the target-to-source alignments as the reference. For each model, we can obtain word alignments from different layers. In total, we obtain AER scores. We select the one with the lowest AER score, and its corresponding layers of the source-to-target and target-to-source models are the layers we will use to extract alignments at test time:
2 Shift-AET: Alignment from Alignment-Enhanced Transformer
To further improve the alignment accuracy, we propose Shift-AET, a word alignment induction method that extracts alignments from Alignment-Enhanced Transformer (AET). AET extends the Transformer architecture with a separate alignment module, which observes the hidden states of the underlying Transformer at each step and predicts the alignment scores for the current decoder input. Note that this module is a plug and play component and it neither makes any change to the underlying NMT model nor influences the translation quality.
To train the alignment module, we use the symmetrized Shift-Att alignments extracted from vanilla Transformer models as labels. Specifically, while the underlying Transformer is pretrained and fixed (Fig. 2), we train the alignment module with the loss function following Garg et al. (2019):
where is the alignment score matrix predicted by the alignment module, and denotes the normalized reference symmetrized Shift-Att alignments.We simply normalize rows corresponding to target tokens that are aligned to at least one source token of . In this way, we transfer the alignment knowledge implicitly learned in two vanilla Transformer models and into the alignment module of a single AET model.
Once the alignment module is trained, we extract alignment scores from it given a parallel sentence pair and induce alignments using Eq. 5.
Experiments
We follow previous work Zenkel et al. (2019, 2020) in data setup and conduct experiments on publicly available datasets for German-English (de-en)https://www-i6.informatik.rwth-aachen.de/goldAlignment/, Romanian-English (ro-en) and French-English (fr-en)http://web.eecs.umich.edu/˜mihalcea/wpt/index.html. Since no validation set is provided, we follow Ding et al. (2019) to set the last 1,000 sentences of the training data before preprocessing as validation set. We learn a joint source and target Byte-Pair-Encoding Sennrich et al. (2016) with 10k merge operations. Table 1 shows the detailed data statistics.
NMT Systems
We implement the Transformer with fairseq-pyhttps://github.com/pytorch/fairseq and use the transformer_iwslt_de_en model configuration following Ding et al. (2019). We train the models with a batch size of 36K tokens and set the maximum updates as 50K and 10K for Transformer and AET respectively. The last checkpoint of AET is used for evaluation. All models are trained in both translation directions and symmetrized with grow-diag Koehn et al. (2005) using the script from Zenkel et al. (2019).https://github.com/lilt/alignment-scripts
Evaluation
We evaluate the alignment quality of our methods with Alignment Error Rate (Och and Ney, 2000, AER). Since word alignments are useful for many downstream tasks as discussed in Section 1, we also evaluate our methods on dictionary-guided decoding, a downstream task of alignment induction, with the metric BLEU Papineni et al. (2002). More details are in Section 4.3.
Baselines
We compare our methods with two statistical baselines Fast-Align and GIZA++ and nine other baselines:
Naive-Att Garg et al. (2019): the approach we discuss in Section 2.2, which induces alignments from the attention weights of the penultimate layer of the Transformer.
Naive-Att-LA Garg et al. (2019): the Naive-Att method without layer selection. It induces alignments from attention weights averaged across all layers.
Shift-Att-LA: Shift-Att method without layer selection. It induces alignments from attention weights averaged across all layers.
SmoothGrad Li et al. (2016): the method that induces alignments from word saliency, which is computed by averaging the gradient-based saliency scores with multiple noisy sentence pairs as input.
SD-SmoothGrad Ding et al. (2019): an improved version of SmoothGrad, which defines saliency on one-hot input vector instead of word embedding.
PD Li et al. (2019): the method that computes the alignment scores from Transformer by iteratively masking each source token and measuring the prediction difference.
AddSGD Zenkel et al. (2019): the method that explicitly adds an extra attention layer on top of Transformer and directly optimizes its activations towards predicting the to-be-aligned target token.
Mtl-Fullc Garg et al. (2019): the method that trains a single model in a multi-task learning framework to both predict the target sentence and the alignment. When predicting the alignment, the model observes full target sentence and uses symmetrized Naive-Att alignments as labels.
Mtl-Fullc-GZ Garg et al. (2019): the same method as Mtl-Fullc except using symmetrized GIZA++ alignments as labels. It is a statistical and neural method as it relies on GIZA++ alignments.
Among these nine baselines and our proposed methods, SmoothGrad, SD-SmoothGrad and PD induce alignments using feature importance measures, while the others from some form of attention weights. Note that the computation cost of methods with feature importance measures is much higher than those with attention weights.For each sentence pair, PD forwards once with masked sentence pairs as the input, while SmoothGrad and SD-SmoothGrad forward and backward once with ( in Ding et al. (2019)) noisy sentence pairs as the input. In contrast, attention weights based methods forward once with one sentence pair as the input.
2 Alignment Results
Table 2 compares our methods with all the baselines. First, Shift-Att, a pure interpretation method for the vanilla Transformer, significantly outperforms Fast-Align and all neural baselines, and performs comparable with GIZA++. For example, it outperforms SD-SmoothGrad, the state-of-the-art method with feature importance measures to extract alignments from vanilla Transformer, by - AER points across different language pairs. The success of Shift-Att demonstrates that vanilla Transformer has captured alignment information in an implicit way, which could be revealed from the attention weights if the correct decoding step and layer are chosen to induce alignments.
Second, the method Shift-AET achieves new state-of-the-art, significantly outperforming all baselines. It improves over GIZA++ by - AER across different language pairs, demonstrating that it is possible to build a neural aligner better than GIZA++ without using any alignments generated from statistical aligners to bootstrap training. We also find Shift-AET performs either marginally better (de-en and ro-en) or on-par (fr-en) when comparing with MTL-Fullc-GZ, a method that uses GIZA++ alignments to bootstrap training. We evaluate the model sizes: the number of parameters in vanilla Transformer and AET are 36.8M and 37.3M respectively, and find that AET only introduces 1.4% additional parameters to the vanilla Transformer. In summary, by supervising the alignment module with symmetrized Shift-Att alignments, Shift-AET improves over Shift-Att and GIZA++ with negligible parameter increase and without influencing the translation quality.
Comparison with Zenkel et al. (2020)
Concurrent with our work, Zenkel et al. (2020) propose a neural aligner that can outperform GIZA++. Table 3 compares the performance of Shift-AET and the best method BAO-Guided (Birdir. Att. Opt. + Guided) in Zenkel et al. (2020). We observe that Shift-AET performs better than BAO-Guided in terms of alignment accuracy.
Shift-AET is also much simpler than BAO-Guided. The training of BAO-Guided includes three stages: (i) train vanilla Transformer in source-to-target and target-to-source directions; (ii) train the alignment layer and extract alignments on the training set with bidirectional attention optimization. This alignment extraction process is computational costly since bidirectional attention optimization fine-tunes the model parameters separately for each sentence pair in the training set; (iii) re-train the alignment layer with the extracted alignments as the guidance. In contrast, Shift-AET can be trained much faster in two stages and does not involve bidirectional attention optimization.
Similar with Mtl-Fullc Garg et al. (2019), BAO-Guided adapts the alignment induction with the to-be-aligned target token by requiring full target sentence as the input. Therefore, BAO-Guided is not applicable in cases where alignments are incrementally computed during the decoding process, e.g., dictionary-guided decoding Alkhouli et al. (2018). In contrast, Shift-AET performs quite well on such cases (Section 4.3). Therefore, considering the alignment performance, computation cost and applicable scope, we believe Shift-AET is more appropriate than BAO-Guided for the task of alignment induction.
Performance on Distant Language Pair
To further demonstrate the superiority of our methods on distant language pairs, we also evaluate our methods on Chinese-English (zh-en). We use NIST corporaThe corpora include LDC2002E18, LDC2003E07, LDC2003E14, LDC2004T07, LDC2004T08 and LDC2005T06 as the training set and v1-tstset released by TsinghuaAligner Liu and Sun (2015) as the test set. The test set includes 450 parallel sentence pairs with manually labelled word alignments.TsinghuaAligner labels the word alignments based on segmented Chinese sentences and does not provide the segmentation model. Therefore, we convert the manually labelled word alignments to our segmented Chinese sentences for evaluation. We use jiebahttps://github.com/fxsjy/jieba for Chinese text segmentation and follow the settings in Section 4.1 for data pre-processing and model training. The results are shown in Table 4. It presents that both Shift-Att and Shift-AET outperform Naive-Att to a large margin. When comparing the symmetrized alignment performance with GIZA++, Shift-AET performs better, while Shift-Att is worse. The experimental results are roughly consistent with the observations on other language pairs, demonstrating the effectiveness of our methods even for distant language pairs.
3 Downstream Task Results
In addition to AER, we compare the performance of Naive-Att, Shift-Att and Shift-AET on dictionary-guided machine translation Song et al. (2020), which is an alignment-based downstream task. Given source and target constraint pairs from dictionary, the NMT model is encouraged to translate with provided constraints via word alignments Alkhouli et al. (2018); Hasler et al. (2018); Hokamp and Liu (2017); Song et al. (2020). More specifically, at each decoding step, the last token of the candidate translation will be revised with target constraint if it is aligned to the corresponding source constraint according to the alignment induction method. To simulate the process of looking up dictionary, we follow Hasler et al. (2018) and extract the pre-specified constraints from the test set and its reference according to the golden word alignments. We exclude stop words, and sample up to 3 dictionary constraints per sentence. Each dictionary constraint includes up to 3 source tokens.
Table 5 presents the performance with different alignment methods. Both Shift-Att and Shift-AET outperform Naive-Att. Shift-AET obtains the best translation quality, improving over Naive-Att by 1.1 and 1.5 BLEU scores on deen and ende translations, respectively. The results suggest the effectiveness of our methods in application to alignment-related NLP tasks.
4 Analysis
To test whether the layer selection criterion can select the right layer to extract alignments, we first determine the best layer and based on the layer selection criterion. Then we evaluate the AER scores of alignments induced from different layers on the test set, and check whether the layers with the lowest AER score are consistent with and . The experiment results shown in Table 6 verify that the layer selection criterion is able to select the best layer to induce alignments. We also find that the best layer is always layer 3 under our setting, consistent across different language pairs.
Relevance Measure Verification
To investigate the relationship between and , we design an experiment to probe whether contain the identity information of and , following Brunner et al. (2019). Formally, for decoder hidden state , the input token is identifiable if there exists a function such that . We cannot prove the existence of analytically. Instead, for each layer we learn a projection function to project from the hidden state space to the input token embedding space and then search for the nearest neighbour within the same sentence. We say that can identify if . Similarly, we follow the same process to identify the output token . We report the identifiability rate defined as the percentage of correctly identified tokens.
Fig. 3 presents the results on the validation set of deen translation. We try three projection functions: a naive baseline , a linear perceptron and a non-linear multi-layer perceptron . We observe the following points: (i) With trainable projection functions and , all layers can identify the input tokens, although more hidden states cannot be mapped back to their input tokens anymore in higher layers. (ii) Overall it is easier to identify the input token than the output token. For example, when projecting with mlp, all layers can identify more than 98% of the input tokens. However, for the output tokens, we can only identify 83.5% even from the best layer. Since even may not be able to identify , this observation partially verifies that it is better to represent using than . (iii) At bottom layers, the input tokens remain identifiable and the output tokens are hard to identify, regardless of the projection function we use. This confirms our hypothesis that for small , is more relevant to than .
AER v.s. BLEU
During training, vanilla Transformer gradually learns to align and translate. To analyze how the alignment behavior changes at different layers with checkpoints of different translation quality, we plot AER on the test set v.s. BLEU on the validation set for deen translation. We compare Naive-Att and Shift-Att, which align the decoder output token (align output) and decoder input token (align input) to the source tokens based on current decoder hidden state, respectively.
The experiment results are shown in Fig. 4. We observe that at the beginning of training, layers and learn to align the input token, while layers and the output token. However, with the increasing of BLEU score, layer tends to change from aligning input token to aligning output token, and layer and begin to align input token. This suggests that vanilla Transformer gradually learns to align the input token from middle layers to bottom layers. We also see that at the end of training, layer ’s ability to align output token decreases. We hypothesize that layer already has the ability to attend to the source tokens which are aligned to the output token, therefore attention weights in layer may capture other information needed for translation. Finally, for checkpoints with the highest BLEU score, layer aligns the output token best and layer aligns the input token best.
Alignment Example
In Fig. 5, we present a symmetrized alignment example from de-en test set. Manual inspection of this example as well as others finds that our methods Shift-Att and Shift-AET tend to extract more alignment pairs than GIZA++, and extract better alignments especially for sentence beginning compared to Naive-Att.
Related Work
Alignment induction from RNNSearch Bahdanau et al. (2015) has been explored by a number of works. Bahdanau et al. (2015) are the first to show word alignment example using attention in RNNSearch. Ghader and Monz (2017) further demonstrate that the RNN-based NMT system achieves comparable alignment performance to that of GIZA++. Alignment has also been used to improve NMT performance, especially in low resource settings, by supervising the attention mechanisms of RNNSearch Chen et al. (2016); Liu et al. (2016); Alkhouli and Ney (2017).
There is also a number of other studies that induce word alignment from Transformer. Li et al. (2019); Ding et al. (2019) claim that attention may not capture word alignment in Transformer, and propose to induce word alignment with prediction difference Li et al. (2019) or gradient-based measures Ding et al. (2019). Zenkel et al. (2019) modify the Transformer architecture for better alignment induction by adding an extra alignment module that is restricted to attend solely on the encoder information to predict the next word. Garg et al. (2019) propose a multi-task learning framework to improve word alignment induction without decreasing translation quality, by supervising one attention head at the penultimate layer with GIZA++ alignments. Although these methods are reported to improve over head average baseline, they ignore that better alignments can be induced by computing alignment scores at the decoding step when the to-be-aligned target token is the decoder input.
Conclusion
In this paper, we have presented two novel methods Shift-Att and Shift-AET for word alignment induction. Both methods induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work. Experiments on three public alignment datasets and a downstream task prove the effectiveness of these two methods. Shift-AET further extends Transformer with an additional alignment module, which consistently outperforms prior neural aligners and GIZA++, without influencing the translation quality. To the best of our knowledge, it reaches the new state-of-the-art performance among all neural alignment induction methods. We leave it for future work to extend our study to more downstream tasks and systems.
Acknowledgments
This work was supported by the National Key R&D Program of China (No. 2018YFB1005103), National Natural Science Foundation of China (No. 61925601), the Fundamental Research Funds for the Central Universities and the funds of Beijing Advanced Innovation Center for Language Resources (No. TYZ19005). We thank the anonymous reviewers for their insightful feedback on this work.