Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information

Zehui Lin, Xiao Pan, Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou, Lei Li

Introduction

Pre-trained language models such as BERT have been highly effective for NLP tasks Peters et al. (2018); Devlin et al. (2019); Radford et al. (2019); Conneau and Lample (2019); Liu et al. (2019); Yang et al. (2019). Pre-training and fine-tuning has been a successful paradigm. It is intriguing to discover a “BERT” equivalent – a pre-trained model – for machine translation. In this paper, we study the following question: can we develop a single universal MT model and derive specialized models by fine-tuning on an arbitrary pair of languages?

While pre-training techniques are working very well for NLP task, there are still several limitations for machine translation tasks. First, pre-trained language models such as BERT are not easy to directly fine-tune unless using some sophisticated techniques (Yang et al., 2020). Second, there is a discrepancy between existing pre-training objective and down-stream ones in MT. Existing pre-training approaches such as MASS Song et al. (2019) and mBART Liu et al. (2020) rely on auto-encoding objectives to pre-train the models, which are different from translation. Therefore, their fine-tuned MT models still do not achieve adequate improvement. Third, existing MT pre-training approaches focus on using multilingual models to improve MT for low resource or medium resource languages. There has not been one pre-trained MT model that can improve for any pairs of languages, even for rich resource settings such as English-French.

In this paper, we propose multilingual Random Aligned Substitution Pre-training (mRASP), a method to pre-train a MT model for many languages, which can be used as a common initial model to fine-tune on arbitrary language pairs. mRASP will then improve the translation performance, comparing to the MT models directly trained on downstream parallel data. In our method, we ensure that the pre-training on many languages and the down-stream fine-tuning share the same model architecture and training objective. Therefore, this approach lead to large translation performance gain. Consider that many languages differ lexically but are closely related at the semantic level, we start by training a large-scale multilingual NMT model across different translation directions, then fine-tuning the model in a specific direction. Further, to close the representation gap across different languages and make full use of multilingual knowledge, we explicitly introduce additional loss based on random aligned substitution of the words in the source and target sentences. Substituted sentences are trained jointly with the same translation loss as the original multilingual parallel corpus. In this way, the model is able to bridge closer the representation space across different languages.

We carry out extensive experiments in different scenarios, including translation tasks with different dataset scales, as well as exotic translation tasks. For extremely low resource (<<100k), mRASP obtains gains up to +22 BLEU points compared to directly trained models on the downstream language pairs. mRASP obtains consistent performance gains as the size of datasets increases. Remarkably, even for rich resource (>>10M, e.g. English-French), mRASP still achieves big improvements.

We divide ”exotic translation” into four categories with respect to the source and target side.

Exotic Pair Both source and target languages are individually pre-trained while they have not been seen as bilingual pairs.

Exotic Source Only target language is pre-trained, but source language is not.

Exotic Target Only source language is pre-trained, but the target language is not.

Exotic Full Neither source nor target language is pre-trained.

Surprisingly, even when mRASP is fine-tuned on ”exotic full” language pair, the resulting MT model is still much better than the directly trained ones (+3.3 to +14.1 BLEU). We finally conduct extensive analytic experiments to examine the contributing factors inside the mRASP method for the performance gains.

We highlight our contributions as follows:

We propose mRASP, an effective pre-training method that can be utilized to fine-tune on any language pairs in NMT. It is very efficient in the use of parallel data in multiple languages. While other pre-trained language models are obtained through hundreds of billions of monolingual or cross-lingual sentences, mRASP only introduces several hundred million bilingual pairs. We suggest that the consistent objectives of pre-training and fine-tuning lead to better model performance.

We explicitly introduce a random aligned substitution technique into the pre-training strategy, and find that such a technique can bridge the semantic space between different languages and thus improve the final translation performance.

We conduct extensive experiments 42 translation directions across different scenarios, demonstrating that mRASP can significantly boost the performance on various translation tasks. mRASP achieves 14.1 BLEU with only 12k pairs of Dutch and Portuguese sentences even though neither appears in the pre-training data. mRASP also achieves 44.3 BLEU on WMT14 English-French translation. Note that our pre-trained model only use parallel corpus in 32 languages, unlike other methods that also use much more monolingual raw corpus.

Methodology

In this section, we introduce our proposed mRASP and the training details.

We adopt a standard Transformer-large architecture Vaswani et al. (2017) with 6-layer encoder and 6-layer decoder. The model dimension is 1,024 on 16 heads. We replace ReLU with GeLU Hendrycks and Gimpel (2016) as activation function on feed forward network. We also use learned positional embeddings.

Methodology

A multilingual neural machine translation model learns a many-to-many mapping function ff to translate from one language to another. More formally, define L={L1,…,LM}L=\left\{L_{1},\ldots,L_{M}\right\} where LL is a collection of languages involving in the pre-training phase. Di,j\mathcal{D}_{i,j} denotes a parallel dataset of (Li,Lj)(L_{i},L_{j}), and E\mathcal{E} denotes the set of parallel datasets {D}i=1i=N\{\mathcal{D}\}_{i=1}^{i=N}, where NN the numbers of the bilingual pair. The training loss is then defined as:

where xi\mathbf{x}^{i} represents a sentence in language LiL_{i}, and θ\theta is the parameter of mRASP, and C(xi)C(\mathbf{x}^{i}) is our proposed alignment function, which randomly replaces the words in xi\mathbf{x}^{i} with a different language. In the pre-training phase, the model jointly learns all the translation pairs.

Language Indicator

Inspired by Johnson et al. (2017); Ha et al. (2016), to distinguish from different translation pairs, we simply add two artificial language tokens to indicate languages at the source and target side. For instance, the following En→\rightarrowFr sentence “How are you? -> Comment vas tu? ” is transformed to “ How are you? -> Comment vas tu?”

Multilingual Pre-training via RAS

Recent work proves that cross-lingual language model pre-training could be a more effective way to representation learning Conneau and Lample (2019); Huang et al. (2019). However, the cross-lingual information is mostly obtained from shared subword vocabulary during pre-training, which is limited in several aspects:

The vocabulary sharing space is sparse in most cases. Especially for dissimilar language pairs, such as English and Hindi, they share a fully different morphology.

The same subword across different languages may not share the same semantic meanings.

The parameter sharing approach lacks explicit supervision to guild the word with the same meaning from different languages shares the same semantic space.

Inspired by constructive learning, we propose to bridge the semantic gap among different languages through Random Aligned Substitution (RAS). Given a parallel sentence (xi,xj)(\mathbf{x}^{i},\mathbf{x}^{j}), we randomly replace a source word in xti\mathbf{x}^{i}_{t} to a different random language LkL_{k}, where tt is the word index. We adopt an unsupervised word alignment method MUSELample et al. (2018b), which can translate xti\mathbf{x}^{i}_{t} to di,k(xti)d_{i,k}(\mathbf{x}^{i}_{t}) in language LkL_{k}, where di,k(⋅)d_{i,k}(\cdot) is the dictionary translating function. With the dictionary replacement, the original bilingual pair will construct a code-switched sentence pair (C(xi),xj)(C(\mathbf{x}^{i}),\mathbf{x}^{j}). As the benefits of random sampling, the translation set {di,k(xti)}k=1k=M\{d_{i,k}(\mathbf{x}^{i}_{t})\}_{k=1}^{k=M} potentially appears in the same context. Since the word representation depends on the context, the word with similar meaning across different languages can share a similar representation. Figure 1 shows our alignment methodology.

2 Pre-training Data

We collect 32 English-centric language pairs, resulting in 64 directed translation pairs in total. English is served as an anchor language bridging all other languages. The parallel corpus are from various sources: tedCompiled by Qi et al. (2018). For simplicity, we deleted zh-tw and zh (which is actually Cantonese), and merged fr-ca with fr, pt-br with pt., wmthttp://www.statmt.org, europarlhttp://opus.nlpl.eu/Europarl-v8.php, paracrawlhttps://paracrawl.eu/, open-subtitleshttp://opus.nlpl.eu/OpenSubtitles-v2018.php, qedhttp://opus.nlpl.eu/QED-v2.0a.php. We refer to our pre-training data as PC32(Parallel Corpus 32). PC32 contains a total size of 197M pairs of sentences. Detailed descriptions and summary for the datasets can be found in Appendix.

For RAS, we utilize ground-truth En-X bilingual dictionarieshttps://github.com/facebookresearch/MUSE, where X denotes languages involved in PC32. Since not all languages in PC32 have ground-truth dictionaries, we only use available dictionaries.

3 Pre-training Details

We use learned joint vocabulary. We learn shared BPE Sennrich et al. (2016b) merge operations (with 32k merge ops) across all the training data and added monolingual data as a supplement (limit to 1M sentences). We do over-sampling in learning BPE to balance the vocabulary size of languages, whose resources are drastically different in size. We over-sampled the corpus of each language based on the volume of the largest language corpus. We keep tokens occurring more than 20, which results in a subword vocabulary of 64,808 tokens.

In pre-training phase, we train our model with the full pairs of the parallel corpus. Following the training setting in Transformer, we use Adam optimizer with ϵ=1e−8,β2=0.98\epsilon=1e-8,\beta_{2}=0.98. A warm-up and linear decay scheduling with a warm-up step of 4000 is used. We pre-train the model for a total of 150000 steps.

For RAS, we use the top 1000 words in dictionaries and only substitute words in source sentences. Each word is replaced with a probability of 30%\% according to the En-X bilingual dictionaries. To address polysemy, we randomly select one substitution from all candidates.

Experiments

This section shows that mRASP obtains consistent performance gains in different scenarios. We also compare our method with existing pre-training methods and outperforms the baselines on En→\rightarrowRo dataset. The performance further boosts by combining back-translationSennrich et al. (2016a) technique. Otherwise stated, for all experiments, we use the pre-trained model as initialization and fine-tune with the downstream target parallel corpus.

We collect 14 pairs of parallel corpus to simulate different scenarios. Most of the En-X parallel datasets are from the pre-training phase to avoid introducing new information. Most pairs for fine-tuning are from previous years of WMT and IWSLT. Specifically, we use WMT14 for En-De and En-Fr, WMT16 for En-Ro. For pairs like Nl(Dutch)-Pt(Portuguese) that are not available in WMT or IWSLT, we use news-commentary instead. For a detailed description, please refer to the Appendix.

Based on the volume of parallel bi-texts, we divide the datasets into four categories: extremely low resource (<<100K), low resource(>>100k and <<1M), medium resource (>>1M and <<10M), and rich resource (>>10M).

For back translation, we include 2014-2018 newscrawl for the target side, En. The total size of the monolingual data is 3M.

Baseline

To better quantify the effectiveness of the proposed pre-training models, we also build two baselines.

mRASP w/o RAS. To measure the effect of alignment information, we also pre-train a model on the same PC32. We do not include alignment information on this pre-training model.

Direct. We also train randomly initialized models directly on downstream bilingual parallel corpus as a comparison with pre-training models.

Fine-tuning

We fine-tune our obtained mRASP model on the target language pairs. We apply a dropout rate of 0.3 for all pairs except for rich resource such as En-Zh and En-Fr with 0.1. We carefully tune the model, setting different learning rates and learning scheduler warm-up steps for different data scale. For inference, we use beam-search with beam size 5 for all directions. For most cases, We measure case-sensitive tokenized BLEU. We also report de-tokenized BLEU with SacreBLEU Post (2018) for a fair comparison with previous works.

2 Main Results

We first conduct experiments on the (extremely) low-resource and medium-resource datasets, where multilingual translation usually obtains significant improvements. As illustrated in Table 1, we obtain significant gains in all datasets. For extremely low resources setting such as En-Be (Belarusian) where the amount of datasets cannot train an NMT model properly, utilizing the pre-training model boosts performance.

We also obtain consistent improvements in low and medium resource datasets. Not surprisingly, We observe that with the scale of the dataset increasing, the gap between the randomly initialized baseline and pre-training model is becoming closer. It is worth noting that, for En→\rightarrowDe benchmark, we obtain 1.0 BLEU points gainsWe report results of En→\rightarrowDe on newstest14. The baseline result is reported in Ott et al. (2018). Extra experiment results on public testsets are provided in Table 9.

To verify mRASP can further boost performance on rich resource datasets, we also conduct experiments on En→\rightarrowZh and En→\rightarrowFr. We compare our results with two strong baselines reported by Ott et al. (2018); Li et al. (2019). As shown in Table 2, surprisingly, when large parallel datasets are provided, it still benefits from pre-training models. In En→\rightarrowFr, we obtain 1.1 BLEU points gains.

We compare our mRASP to recently proposed multilingual pre-training models. Following Liu et al. (2020), we conduct experiments on En-Ro, the only pairs with established results. To make a fair comparison, we report de-tokenized BLEU.

As illustrated in Table 4 , Our model reaches comparable performance on both En→\rightarrowRo and Ro→\rightarrowEn. We also combine Back Translation Sennrich et al. (2016a) with mRASP, observing performance boost up to 2 BLEU points, suggesting mRASP is complementary to BT. It should be noted that the competitors introduce much more pre-training data.

mBART contucted experiments on extensive language pairs. To illustrate the superiority of mRASP, we also compare our results with mBART. We use the same test sets as mBART. As illustrated in Table 5, mRASP outperforms mBART for most of language pairs by a large margin. Note that while mBART underperforms baseline for benchmarks En-De and En-Fr, mRASP obtains 4.3 and 2.9 BLEU gains compared to baseline.

3 Generalization to Exotic Translation

To illustrate the generalization of mRASP, we also conduct experiments on exotic translation directions, which is not included in our pre-training phase. For each category, we select language pairs of different scales.

The results are shown in Table 3. As is shown, mRASP obtains significant gains for each category for different scales of datasets, indicating that even trained with exotic languages, with pre-training initialization, the model still works reasonably well.

Note that in the most challenging case, Exotic Full, where the model does not have any knowledge of both sides, with only 11K parallel pairs for Nl(Dutch)-Pt(Portuguese), the pre-training model still reaches reasonable performance, while the baseline fails to train appropriately. It suggests the pre-train model does learn language-universal knowledge and can transfer to exotic languages easily.

Analysis

In this section, we conduct a set of analytical experiments to better understand what contributes to performance gains. Three aspects are studied. First, we study whether the main contribution comes from pre-training or fine-tuning by comparing the performance of fine-tuning and no-fine-tuning. The results suggest that the performance mainly comes from pre-training, while fine-tuning further boosts the performance. Second, we thoroughly analyze the difference between incorporating RAS at the pre-training phase and pre-training without RAS. The finding shows that incorporating alignment information helps bridge different languages and obtains additional gains. Lastly, we study the effect of data volume in the fine-tuning phase.

In the pre-training phase, the model jointly learns from different language pairs. To verify whether the gains come from pre-training or fine-tuning, we directly measure the performance without any fine-tuning, which is, in essence, zero-shot translation task.

We select datasets covering different scales. Specifically, En-Af (41k) from extremely low resource, En-Ro (600k) from low resource, En-De (4.5M) from medium resource, and En-Fr (40M) from rich resource are selected.

As shown in Table 6 , we find that model without fine-tuning works surprisingly well on all datasets, especially in low resource where we observe model without fine-tuning outperforms randomly initialized baseline model. It suggests that the model already learns well on the pre-training phase, and fine-tuning further obtains additional gains. We suspect that the model mainly tunes the embedding of specific language at the fine-tuning phase while keeping the other model parameters mostly unchanged. Further analytical experiments can be conducted to verify our hypothesis.

Note that we also report pre-trained model without RAS (NA-mRASP). For comparison, we do not apply fine-tuning on NA-mRASP. mRASP consistently obtains better performance that NA-mRASP, implying that injecting information at the pre-training phase do improve the performance.

The effectiveness of RAS technique

In the pre-training phase, we explicitly incorporate RAS. To verify the effectiveness of RAS, we first compare the performance of mRASP and mRASP without RAS.

As illustrated in Table 7, We find that utilizing RAS in the pre-training phase consistently helps improve the performance in datasets with different scales, obtaining gains up to 2.5+ BLEU points.

To verify whether the semantic space of different languages draws closer after adding alignment information quantitatively, we calculate the average cosine similarity of words with the same meaning in different languages. We choose the top frequent 1000 words according to MUSE dictionary. Since words are split into subwords through BPE, we simply add all subwords constituting the word. As illustrated in Figure 3, we find that for all pairs in the Figure, the average cosine similarity increases by a large margin after adding RAS, suggesting the efficacy of alignment information in bridging different languages. It is worth mentioning that the increase does not only happen on similar pairs like En-De, but also on dissimilar pairs like En-Zh.

To further illustrate the effect of RAS on semantic space more clearly, we use PCA (Principal Component Analysis) to visualize the word embedding space. We plot En-Zh as the representative for dissimilar pairs and En-Af for similar pairs. More figures can be found in the Appendix.

As illustrated in Figure 2, we find that for both similar pair and dissimilar pair, the overall word embedding distribution becomes closer after RAS. For En-Zh, as the dashed lines illustrate, the angle of the two word embedding spaces becomes smaller after RAS. And for En-Af, we observe that the overlap between two space becomes larger. We also randomly plot the position of three pairs of words, with each pair has the same meaning in different languages.

Fine-tuning Volume

To study the effect of data volume in the fine-tuning phase, we randomly sample 1K, 5K, 10K, 50K, 100K, 500K, 1M datasets from the full En-De corpus (4.5M). We fine-tune the model with the sampled datasets, respectively. Figure 4 illustrates the trend of BLEU with the increase of data volume. With only 1K parallel pairs, the pre-trained model works surprisingly well, reaching 24.46. As a comparison, the model with random initialization fails on this extremely low resource. With only 1M pairs, mRASP reaches comparable results with baseline trained on 4.5M pairs.

With the size of dataset increases, the performance of the pre-training model consistently increases. While the baseline does not see any improvement until the volume of the dataset reaches 50K. The results confirm the remarkable boosting of mRASP on low resource dataset.

Related Works

aims at taking advantage of multilingual data to improve NMT for all languages involved, which has been extensively studied in a number of papers such as Dong et al. (2015); Johnson et al. (2017); Lu et al. (2018); Rahimi et al. (2019); Tan et al. (2019). The most related work to mRASP is Rahimi et al. (2019), which performs extensive experiments in training massively multilingual NMT models. They show that multilingual many-to-many models are effective in low resource settings. Inspired by their work, we believe that the translation quality of low-resource language pairs may improve when trained together with rich-resource ones. However, we are different in at least two aspects: a) Our goal is to find the best practice of a single language pair with multilingual pre-training. Multilingual NMT usually achieves inferior accuracy compared with its counterpart, which trains an individual model for each language pair when there are dozens of language pairs. b) Different from multilingual NMT, mRASP can obtain improvements with rich-resource language pairs, such as English-Frence.

Unsupervised Pretraining

has significantly improved the state of the art in natural language understanding from word embedding Mikolov et al. (2013b); Pennington et al. (2014), pretrained contextualized representations Peters et al. (2018); Radford et al. (2019); Devlin et al. (2019) and sequence to sequence pretraining Song et al. (2019). It is widely accepted that one of the most important factors for the success of unsupervised pre-training is the scale of the data. The most successful efforts, such as RoBERTa, GPT, and BERT, highlight the importance of scaling the amount of data. Following their spirit, we show that with massively multilingual pre-training, more than 110 million sentence pairs, mRASP can significantly boost the performance of the downstream NMT tasks.

On parallel, there is a bulk of work on unsupervised cross-lingual representation. Most traditional studies show that cross-lingual representations can be used to improve the quality of monolingual representations. Mikolov et al. (2013a) first introduces dictionaries to align word representations from different languages. A series of follow-up studies focus on aligning the word representation across languages Xing et al. (2015); Ammar et al. (2016); Smith et al. (2017); Lample et al. (2018b). Inspired by the success of BERT, Conneau and Lample (2019) introduced XLM - masked language models trained on multiple languages, as a way to leverage parallel data and obtain impressive empirical results on the cross-lingual natural language inference (XNLI) benchmark and unsupervised NMTSennrich et al. (2016a); Lample et al. (2018a); Garcia et al. (2020). Huang et al. (2019) extended XLM with multi-task learning and proposed a universal language encoder.

Different from these works, a) mRASP is actually a multilingual sequence to sequence model which is more desirable for NMT pre-training; b) mRASP introduces alignment regularization to bridge the sentence representations across languages.

Conclusion

In this paper, we propose a multilingual neural machine translation pre-training model (mRASP). To bridge the semantic space between different languages, we incorporate word alignment into the pre-training model. Extensive experiments are conducted on different scenarios, including low/medium/rich resource and exotic corpus, demonstrating the efficacy of mRASP. We also conduct a set of analytical experiments to quantify the model, showing that the alignment information does bridge the gap between languages as well as boost the performance. We leave different alignment approaches to be explored in the future. In future work, we will pre-train on larger corpus to further boost the performance.

Acknowledgments

We would like to thank the anonymous reviewers for their valuable comments. We would also like to thank Liwei Wu, Huadong Chen, Qianqian Dong, Zewei Sun, and Weiying Ma for their useful suggestion and help with experiments.

References

Appendix A Appendices

In addition to visualization of En-Zh and En-Af presented in main body of paper, we also plot visualization of En-Ro, En-Ar, En-Tr and En-De. As shown in Figure 5,6,7,8, the overall word embedding distribution becomes closer after RAS.

A.2 Case Study

A.3 Results on public testsets

A.4 Data Description

As listed in Table 10, we collect 32 English-centric language pairs, resulting in a total pairs of 110M. The parallel corpus are from various source, ted, wmt, europarl, paracrawl, opensubtitles and qed.