Context-Aware Cross-Attention for Non-Autoregressive Translation
Liang Ding, Longyue Wang, Di Wu, Dacheng Tao, Zhaopeng Tu
Introduction
Different from autoregressive translation [Bahdanau et al. (2015, Vaswani et al. (2017, AT] models that generate each target word conditioned on previously generated ones, non-autoregressive translation [Gu et al. (2018, NAT] models break the autoregressive factorization and produce the target words in parallel. Given a source sentence , the probability of generating its target sentence with length is defined by NAT as: , where is a separate conditional distribution to predict the length of target sequence. As NAT models can predict all tokens independently and simultaneously, recent works have fully investigated their superiority on decoding efficiency [Lee et al. (2018, Ghazvininejad et al. (2019, Gu et al. (2019, Kasai et al. (2020, Sun et al. (2019, Shu et al. (2020, Ran et al. (2019]. However, there still exists a gap between AT and NAT models in terms of effectiveness.
In encoder-decoder frameworks, the cross-attention module dynamically selects relevant source-side information (key) given a target-side token (query) [Yang et al. (2020, Wang and Tu (2020]. Through qualitative and quantitative analyses, we found that it is difficult for the NAT decoder to adequately capture the source context due to the lack of autoregressive factorization. As shown in Table 1, when translating the Chinese word “交往”, the source context word “女孩” should play a significant role in predicting the candidate word “dating”. However, the NAT model inappropriately generates “socializing with”, resulting in lexical choice errors. As seen, the AT model gives relatively higher attention weights to local contexts on the source side while the NAT model pays less attention on them (0.15 vs 0.04). We make further statistical analysis in Section 2 to prove the universality of this localness perception problem. Similar to our findings, ?) showed that distributions of cross-attention in NAT models are more ambiguous than those in AT ones.
To alleviate this localness perception problem in NAT, we propose a context-aware cross-attention to model both local and global contexts simultaneously. For local attention, we limit the scope of cross-attention to adjacent tokens surrounding the source word with the maximum alignment probability. We then combine the local attention weights with the original global ones by a gating mechanism (in Section 3).
Experiments are conducted on four commonly-cited datasets on translation task (i.e. WMT16 RomanianEnglish, WAT17 JapaneseEnglish, WMT14 EnglishGerman and WMT17 ChineseEnglish) and show that our approach can consistently improve translation quality by around 0.5 BLEU point over advanced NAT models (in Section 4). Further analyses reveal that our method can enhance abilities of NAT to learn syntactic and semantic information as well as phrase patterns (in Section 5).
Localness Perception Problem
We compare the locality entropy of NAT and AT models on En-De and Zh-En. As shown in Table 2, the locality entropy “LE” of NAT model is higher than that of AT, showing that the localness perception problem in NAT is more severe. With the help of our method (in Section 3), this problem can be alleviated (LE), leading to better translation quality (BLEU ). This observation confirms the universality and side effect of localness perception problem in NAT, validating our hypothesis in Section 1.
Context-Aware Cross-Attention for NAT
In this section, we introduce the detail of our proposed context-aware cross-attention networks (CCAN), which perceives the original and local cross-attention simultaneously.
For the target-side query , source-side key and value . The -th original cross-attention can be calculated with dot-product: . The original attention of the -th element is the weighted sum of values (in Figure 1(a)).
Our Approach
For the -th position in target side, we propose a locally-sensitive cross-attention component for NAT to capture the neighbor signals. For simplicity, we adopt a straightforward but has been proven effective way [Luong et al. (2015, Xu et al. (2019, You et al. (2020]: constricting the attention scope to a nearby window around the aligned -th element. In practice, we choose the source element with the highest attention weight as the aligned element, and the local range can be modeled as follows:
where denotes the attention correlation between the - and - elements in encoder and decoder parts, respectively. The is the hard-coded localness modeling window. Furthermore, we design an interpolation gating mechanism to wisely combine the original and local cross-attention:
where is the interpolation weight conditioned on the decoder side query and denotes the sigmoid function. Note that is the only additional parameter to estimate the importance of original cross-attention operation, and we share it for different cross-attention heads.
Experiments
Experiments are conducted on four widely-used translation datasets, including the small-scale WMT16 Romanian-English (Ro-En, ?)), the medium-scale WMT14 English-German (En-De, ?)), the large-scale WMT17 Chinese-English (Zh-En, ?)), and word-order-divergent WAT17 Japanese-English (Ja-En, ?)), which consist of 0.6M, 4.5M, 20M and 2M sentence pairs, respectively. We preprocessed data via BPE [Sennrich et al. (2016] with 32K merge operations. We used BLEU [Papineni et al. (2002] as metric with statistical significance test [Collins et al. (2005].
Models
We follow ?) to apply sequence-level knowledge distillation [Kim and Rush (2016] to simplify the training data. About AT Teachers, we train both Base and Big Transformer [Vaswani et al. (2017] models with corresponding training data. In Big model, we adopt large batch strategy (458K tokens per batch) to optimize the performance. The main results employ Transformer-Big for all directions except Ro-En, which is distilled by Base. Our approach can be applied to different NAT architectures. In this paper, we mainly implement it on conditional masked language models [Ghazvininejad et al. (2019, CMLMs] and leave further investigation to future work. The model contains 6-layer encoder and 6-layer decoder, where the decoder trained with conditional mask language model fashion. The model dimension is 512 on 8 heads, with 2048 feed forward dimensions. We follow the common practices [Ghazvininejad et al. (2019, Kasai et al. (2020] to average the top three checkpoints to avoid stochasticity.
2 Ablation Study
In order to make best use of our proposed component for NAT, we conducted extensive ablation studies. All models are trained and validated on WMT14 En-De training and validation sets.
We investigate the localness window size within and report the translation performance in Table 3 (left). As seen, our context-aware cross-attention with the window size of 9 achieves the best BLEU, which is therefore used as the default setting.
Effects of Decoder Layers
As shown in Table 3 (right), deploying CCAN on the top-layer slightly outperforms deploying on the bottom-layer (“”>“”). In NAT, multiple decoding layers can be cast as the refiner, and the source central word chosen by the bottom-layer cross-attention is not as accurate as of the top-layer one. Our method, highly conditioned on the predicted central words, thus can gain a better effect on the top-layer compared to the bottom layer. In the end, modelling all layers (“”) achieves the best performance and we thus use this setting in the following experiments.
3 Main Results
Table 4 lists main results and comparison with previous NAT models on WMT16 Ro-En, WMT14 En-De, WMT17 Zh-En and WAT17 Ja-En datasets. We mainly implemented our approach on top of the advanced CMLMs model. As seen, our approach (Row 9) consistently improves translation performance (BLEU) over CMLMs on four language pairs. Note that our approaches only modify the cross-attention module and introduce fewer extra parameters, leading to negligible loss on latency. Encouragingly, our approach even slightly outperforms its AT teachers (Transformer-Base) on three tasks.
Analysis
In this section, we conduct extensive analyses on WMT14 En-De to better understand how our method contribute to performance gains.
The importance of localness should be different over layers. We explore it through gating (in Equation 2) analyzing. Specifically, we cast the weighting scalar of local cross-attention as its importance degree and calculate the importance of localness for each decoder layer. As shown in Figure 2(a), during information flow evolving from bottom to top layers, the importance of localness continues to decline till the penultimate layer, and then increases. The possible reason for the increase in the last two layers is that the top layer followed by softmax, requiring more source-side context to choose lexicons.
Phrasal Patterns
Our approach is expected to pay more attention to the most relevant source token and its neighbours, such that the phrasal translation can be improved. To evaluate the accuracy of phrase translations, we calculate the improvement on n-gram tokens in Figure 2(b), where the golden dashed line indicates that the window size is 9. As seen, CCAN consistently outperforms the baseline (Accuracy>0), indicating that our method can enhance the ability of NAT model on capturing the phrasal information, which is similar with ?)’s findings.
Linguistic Properties
Intuitively, our proposed cross-attention component brings context-aware representation, may affecting the linguistic properties learned by the encoder. We quantitatively investigate it from linguistic perspectives with probing tasks [Conneau et al. (2018]. These tasks can be categorized into three types: “Surface” focuses on the simple surface properties learned from the sentence embedding; “Syntactic” quantifies the syntactic reservation ability; and “Semantic” assesses the deeper semantic representation ability. To evaluate the representation ability of CCAN equipped NAT model, we compare the pre-trained vanilla NAT and CCAN equipped NAT encoders, followed by a MLP classifier. Specifically, the mean of the top encoding layer, as sentence representation, will be passed to the classifier. We can see from Table 5, the CCAN equipped NAT encoder preserves rich syntactic and semantic information.
Conclusion and Future Work
We reveal a localness perception problem in NAT. To alleviate it, we propose the context-aware approach to make the cross-attention pay more attention to source-side local words, which in turn improves the translation performance over several benchmarks. In future work, we will investigate selectively choosing the context [Geng et al. (2020, Yang et al. (2019] rather than the fixed window size. Besides, it is interesting to enhance NAT model with extra signals, such as cross-lingual position embedding [Ding et al. (2020], larger context [Wang et al. (2017] and pre-trained initialization [Liu et al. (2020].
Acknowledgements
This work was supported by Australian Research Council Projects under grants FL-170100117, DP-180103424, and IC-190100031. We are grateful to the anonymous reviewers and the area chair for their insightful comments and suggestions.