Contrastive Learning for Many-to-many Multilingual Neural Machine Translation
Xiao Pan, Mingxuan Wang, Liwei Wu, Lei Li
Introduction
Transformer Vaswani et al. 2017 has achieved decent performance for machine translation with rich bilingual parallel corpora. Recent work on multilingual machine translation aims to create a single unified model to translate many languages Johnson et al. 2017; Aharoni et al. 2019; Zhang et al. 2020; Fan et al. 2020; Siddhant et al. 2020. Multilingual translation models are appealing for two reasons. First, they are model efficient, enabling easier deployment Johnson et al. 2017. Further, parameter sharing across different languages encourages knowledge transfer, which benefits low-resource translation directions and potentially enables zero-shot translation (i.e. direct translation between a language pair not seen during training) Ha et al. 2017; Gu et al. 2019; Ji et al. 2020.
Despite these benefits, challenges still remain in multilingual NMT. First, previous work on multilingual NMT does not always perform well as their corresponding bilingual baseline especially on rich resource language pairs Tan et al. 2019; Zhang et al. 2020; Fan et al. 2020. Such performance gap becomes larger with the increasing number of accommodated languages for multilingual NMT, as model capacity necessarily must be split between many languages Arivazhagan et al. 2019. In addition, an optimal setting for multilingual NMT should be effective for any language pairs, while most previous work focus on improving English-centric “English-centric” means that having English as the source or target language directions Johnson et al. 2017; Aharoni et al. 2019; Zhang et al. 2020. A few recent exceptions are Zhang et al. 2020 and Fan et al. 2020, who trained many-to-many systems with introducing more non-English corpora, through data mining or back translation.
In this work, we take a step towards a unified many-to-many multilingual NMT with only English-centric parallel corpora and additional monolingual corpora. Our key insight is to close the representation gap between different languages to encourage transfer learning as much as possible.
As such, many-to-many translations can make the most of the knowledge from all supervised directions and the model can perform well for both English-centric and non-English settings. In this paper, we propose a multilingual COntrastive Learning framework for Translation (mCOLT or mRASP2) to reduce the representation gap of different languages, as shown in Figure 1.
The objective of mRASP2 ensures the model to represent similar sentences across languages in a shared space by training the encoder to minimize the representation distance of similar sentences. In addition, we also boost mRASP2 by leveraging monolingual data to further improve multilingual translation quality. We introduce an effective aligned augmentation technique by extending RAS Lin et al. 2020 – on both parallel and monolingual corpora to create pseudo-pairs. These pseudo-pairs are combined with multilingual parallel corpora in a unified training framework.
Simple yet effective, mRASP2 achieves consistent translation performance improvements for both English-centric and non-English directions on a wide range of benchmarks. For English-centric directions, mRASP2 outperforms a strong multilingual baseline in 20 translation directions on WMT testsets. On 10 WMT translation benchmarks, mRASP2 even obtains better results than the strong bilingual mBART model. For zero-shot and unsupervised directions, mRASP2 obtains surprisingly strong results on 36 translation directions 6 unsupervised directions 30 zero-shot directions, with 10+ BLEU improvements on average.
Methodology
mRASP2 unifies both parallel corpora and monolingual corpora with contrastive learning. This section will explain our proposed mRASP2. The overall framework is illustrated in Figure 1
A multilingual neural machine translation model learns a many-to-many mapping function to translate from one language to another. To distinguish different languages, we add an additional language identification token preceding each sentence, for both source side and target side. The base architecture of mRASP2 is the state-of-the-art Transformer Vaswani et al. 2017. A little different from previous work, we choose a larger setting with a 12-layer encoder and a 12-layer decoder to increase the model capacity. The model dimension is 1024 on 16 heads. To ease the training of the deep model, we apply Layer Normalization for word embedding and pre-norm residual connection following Wang et al. 2019a for both encoder and decoder. Therefore, our multilingual NMT baseline is much stronger than that of Transformer big model.
More formally, we define where is a collection of M languages involving in the training phase. denotes a parallel dataset of , and denotes all parallel datasets. The training loss is cross entropy defined as:
where represents a sentence in language , and is the parameter of multilingual Transformer model.
2 Multilingual Contrastive Learning
Multilingual Transformer enables implicitly learning shared representation of different languages. mRASP2 introduces contrastive loss to explicitly bring different languages to map a shared semantic space.
The key idea of contrastive learning is to minimize the representation gap of similar sentences and maximize that of irrelevant sentences. Formally, given a bilingual translation pairs , is the positive example and we randomly choose a sentence from language to form a negative example It is possible that = . The objective of contrastive learning is to minimize the following loss:
where calculates the similarity of different sentences. and denotes positive and negative respectively. denotes the average-pooled encoded output of an arbitrary sentence . is the temperature, which controls the difficulty of distinguishing between positive and negative examples Higher temperature increases the difficulty to distinguish positive sample from negative ones.. In our experiments, it is set to . The similarity of two sentences is calculated with the cosine similarity of the average-pooled encoded output. To simplify implementation, the negative samples are sampled from the same training batch. Intuitively, by maximizing the softmax term , the contrastive loss forces their semantic representations projected close to each other. In the meantime, the softmax function also minimizes the non-matched pairs .
During the training of mRASP2, the model can be optimized by jointly minimizing the contrastive training loss and translation loss:
where is the coefficient to balance the two training losses. Since is calculated on the sentence-level and is calculated on the token-level, therefore should be multiplied by the averaged sequence length .
3 Aligned Augmentation
We then will introduce how to improve mRASP2 with data augmentation methods, including the introduction of noised bilingual and noised monolingual data for multilingual NMT. The above two types of training samples are illustrated in Figure 2.
Lin et al. 2020 propose Random Aligned Substitution technique (or RAS They apply RAS only on parallel data) that builds code-switched sentence pairs () for multilingual pre-training. In this paper, we extend it to Aligned Augmentation (AA), which can also be applied to monolingual data.
For a bilingual or monolingual sentence pair (, ) is in language and is in language , where , AA creates a perturbed sentence by replacing aligned words from a synonym dictionary We will release our synonym dictionary. For every word contained in the synonym dictionary, we randomly replace it to one of its synonym with a probability of 90%.
For a bilingual sentence pair , AA creates a pseudo-parallel training example . For monolingual data, AA takes a sentence and generates its perturbed to form a pseudo self-parallel example . and is then used in the training by calculating both the translation loss and contrastive loss. For a pseudo self-parallel example , the translation loss is basically the reconstruction loss from the perturbed sentence to the original one.
Experiments
This section shows that mRASP2 can achieve substantial improvements over previous many-to-many multilingual translation on a wide range of benchmarks. Especially, it obtains substantial gains on zero-shot directions.
We use the parallel dataset PC32 provided by Lin et al. 2020. It contains a large public parallel corpora of 32 English-centric language pairs. The total number of sentence pairs is 97.6 million.
We apply AA on PC32 by randomly replacing words in the source side sentences with synonyms from an arbitrary bilingual dictionary provided by Lample et al. 2018 https://github.com/facebookresearch/MUSE. For words in the dictionaries, we replace them into one of the synonyms with a probability of 90% and keep them unchanged otherwise. We apply this augmentation in the pre-processing step before training.
Monolingual Dataset MC24
We apply AA on MC24 by randomly replacing words in the source side sentences with synonyms from a multilingual dictionary. Therefore the source side might contain multiple language tokens (preserving the semantics of the original sentence), and the target is just the original sentence. The replace probability is also set to 90%. We apply this augmentation in the pre-processing step before training. We will release the multilingual dictionary and the script for producing the noised monolingual dataset.
Evaluation Datasets
For supervised directions, most of our evaluation datasets are from WMT and IWSLT benchmarks, for pairs that are not available in WMT or IWSLT, we use OPUS-100 instead.
For zero-shot directions, we follow Zhang et al. 2020 and use their proposed OPUS-100 zero-shot testset. The testset is comprised of 6 languages (Ru, De, Fr, Nl, Ar, Zh), resulting in 15 language pairs and 30 translation directions.
We report de-tokenized BLEU with SacreBLEU Post 2018. For tokenized BLEU, we tokenize both reference and hypothesis using Sacremoses https://github.com/alvations/sacremoses toolkit then report BLEU using the multi-bleu.pl script https://github.com/moses-smt/mosesdecoder. For Chinese (Zh), BLEU score is calculated on character-level.
Experiment Details
We use the Transformer model in our experiments, with 12 encoder layers and 12 decoder layers. The embedding size and FFN dimension are set to 1024. We use dropout = 0.1, as well as a learning rate of 3e-4 with polynomial decay scheduling and a warm-up step of 10000. For optimization, we use Adam optimizer Kingma and Ba 2015 with = 1e-6 and = 0.98. To stabilize training, we set the threshold of gradient norm to be 5.0 and clip all gradients with a larger norm. We set the hyper-parameter in Eq.3 during training. For multilingual vocabulary, we follow the shared BPE Sennrich et al. 2016 vocabulary of Lin et al. 2020, which includes 59 languages. The vocabulary contains 64808 tokens. After adding 59 language tokens, the total size of vocabulary is 64867.
Experiment Results
This section shows that mRASP2 provides consistent performance gains for supervised and unsupervised English-centric translation directions as well as for non-English directions.
As shown in Table 1, mRASP2 clearly improves multilingual baselines by a large margin in 10 translation directions. Previously, multilingual machine translation underperforms bilingual translation in rich-resource scenarios. It is worth noting that our multilingual machine translation baseline is very competitive. It is even on par with the strong mBART bilingual model, which is fine-tuned on a large scale unlabeled monolingual dataset. mRASP2 further improves the performance.
We summarize the key factors for the success training of our baseline many-to-many Transformer trained on PC32 as in Johnson et al. 2017 except that we apply language indicator the same way as Fan et al. 2020 m-Transformer:
Unsupervised Directions
In Table 2, we observe that mRASP2 achieves reasonable results on unsupervised translation directions. The language pairs of En-Nl, En-Pt, and En-Pl are never observed by m-Transformer. m-Transformer sometimes achieves reasonable BLEU for XEn, e.g. for PtEn, since there are many similar languages in PC32, such as Es and Fr. Not surprisingly, it totally fails on EnX directions. By contrast, mRASP2 obtains +14.13 BLEU score on an average without explicitly introducing supervision signals for these directions.
Furthermore, mRASP2 achieves reasonable BLEU scores on NlPt directions even though it has only been trained on monolingual data of both sides. This indicates that by simply incorporating monolingual data with parallel data in the unified framework, mRASP2 successfully enables unsupervised translation through its unified multilingual representation.
2 Zero-shot Translation for non-English Directions
Zero-shot Translation has been an intriguing topic in multilingual neural machine translation. Previous work shows that the multilingual NMT model can do zero-shot translation directly. However, the translation quality is quite poor compared with pivot-based model.
We evaluate mRASP2 on the OPUS-100 Zhang et al. 2020 zero-shot test set, which contains 6 languages Arabic, Chinese, Dutch, French, German, Russian and 30 translation directions in total. To make the comparison clear, we also report the results of several different baselines. mRASP2 w/o AA only adopt contrastive learning on the basis of m-Transformer. mRASP2 w/o MC24 excludes monolingual data from mRASP2.
The evaluation results are listed in Appendix and we summarize them in Table 3. We find that our mRASP2 significantly outperforms m-Transformer and substantially narrows the gap with pivot-based model. This is in line with our intuition that bridging the representation gap of different languages can improve the zero-shot translation.
The main reason is that contrastive loss, aligned augmentation and additional monolingual data enable a better language-agnostic sentence representation. It is worth noting that, Zhang et al. 2020 achieves BLEU score improvements on zero-shot translations at sacrifice of about 0.5 BLEU score loss on English-centric directions. By contrast, mRASP2 improves zero-shot translation by a large margin without losing performance on English-Centric directions. Therefore, mRASP2 has a great potential to serve many-to-many translations, including both English-centric and non-English directions.
Analysis
To understand what contributes to the performance gain, we conduct analytical experiments in this section. First we summarize and analyze the performance of mRASP2 in different scenarios. Second we adopt the sentence representation of mRASP2 to retrieve similar sentences across languages. This is to verify our argument that the improvements come from the universal language representation learned by mRASP2. Finally we visualize the sentence representations, mRASP2 indeed draws the representations closer.
To make a better understanding of the effectiveness of mRASP2, we evaluate models of different settings. We summarize the experiment results in Table 4:
1 v.s.3: 3 performs comparably with m-Transformer in supervised and unsupervised scenarios, whereas achieves a substantial BLEU improvement for zero-shot translation. This indicates that by introducing contrastive loss, we can improve zero-shot translation quality without harming other directions.
2 v.s.4: 2 performs poorly for zero-shot directions. This means contrastive loss is crucial for the performance in zero-shot directions.
5 : mRASP2 further improves BLEU in all of the three scenarios, especially in unsupervised directions. Therefore it is safe to conjecture that by accomplishing with monolingual data, mRASP2 learns a better representation space.
2 Similarity Search
In order to verify whether mRASP2 learns a better representation space, we conduct a set of similarity search experiments. Similarity search is a task to find the nearest neighbor of each sentence in another language according to cosine similarity. We argue that mRASP2 benefits this task in the sense that it bridges the representation gap across languages. Therefore we use the accuracy of similarity search tasks as a quantitative indicator of cross-lingual representation alignment.
We conducted comprehensive experiments to support our argument and experiment on mRASP2 and mRASP2 w/o AA .We divide the experiments into two scenarios: First we evaluate our method on Tatoeba dataset Artetxe and Schwenk 2019, which is English-centric. Then we conduct similar similarity search task on non-English language pairs. Following Tran et al. 2020, we construct a multi-way parallel testset (Ted-M) of 2284 samples by filtering the test split of ted http://phontron.com/data/ted_talks.tar.gz that have translations for all 15 languages Arabic, Czech, German, English, Spanish, French, Italian, Japanese, Korean, Dutch, Romanian, Russian, Turkish, Vietnamese, Chinese.
Under both settings, we follow the same strategy: We use the average-pooled encoded output as the sentence representation. For each sentence from the source language, we search the closest sentence in the target set according to cosine similarity.
We display the evaluation results in Table 5. We detect two trends: (i) The overall accuracy follows the rule: m-Transformer mRASP2 w/o AA mRASP2. (ii) mRASP2 brings more significant improvements for languages with less data volume in PC32. The two trends mean that mRASP2 increases translation BLEU score in a sense that it bridges the representation gap across languages.
Non-English: Ted-M
It will be more convincing to argue that mRASP2 indeed bridges the representation gap if similarity search accuracy increases on zero-shot directions. We list the averaged top-1 accuracy of 210 non-English directions 15 languages, resulting in 210 directions in Table 6. The results show that mRASP2 increases the similarity search accuracy in zero-shot scenario. The results support our argument that our method generally narrows the representation gap across languages.
To better understanding the specifics beyond the averaged accuracy, we plot the accuracy improvements in the heat map in Figure 3. mRASP2 w/o AA brings general improvements over m-Transformer. mRASP2 especially improves on Dutch(Nl). This is because mRASP2 introduces monolingual data of Dutch while mRASP2 w/o AA includes no Dutch data.
3 Visualization
In order to visualize the sentence representations across languages, we retrieve the sentence representation for each sentence in Ted-M, resulting in 34260 samples in the high-dimensional space.
To facilitate visualization, we apply T-SNE dimension reduction to reduce the 1024-dim representations to 2-dim. Then we select 3 representative languages: English, German, Japanese and depict the bivariate kernel density estimation based on the 2-dim representations. It is clear in Figure 4 that m-Transformer cannot align the 3 languages. By contrast, mRASP2 draws the representations across 3 languages much closer.
Related Work
While initial research on NMT starts with building translation systems between two languages, Dong et al. 2015 extends the bilingual NMT to one-to-many translation with sharing encoders across 4 language pairs. Hence, there has been a massive increase in work on MT systems that involve more than two languages Chen et al. 2018; Choi et al. 2018; Chu and Dabre 2019; Dabre et al. 2017. Recent efforts mainly focuses on designing language specific components for multilingual NMT to enhance the model performance on rich-resource languages Bapna and Firat 2019; Kim et al. 2019; Wang et al. 2019b; Escolano et al. 2020. Another promising thread line is to enlarge the model size with extensive training data to improve the model capability Arivazhagan et al. 2019; Aharoni et al. 2019; Fan et al. 2020. Different from these approaches, mRASP2 proposes to explicitly close the semantic representation of different languages and make the most of cross lingual transfer.
Zero-shot Machine Translation
Typical zero-shot machine translation models rely on a pivot language (e.g. English) to combine the source-pivot and pivot-target translation models Chen et al. 2017; Ha et al. 2017; Gu et al. 2019; Currey and Heafield 2019. Johnson et al. 2017 shows that a multilingual NMT system enables zero-shot translation without explicitly introducing pivot methods. Promising, but the performance still lags behind the pivot competitors. Most following up studies focused on data augmentation methods. Zhang et al. 2020 improved the zero-shot translation with online back translation. Ji et al. 2020; Liu et al. 2020 shows that large scale monolingual data can improve the zero-shot translation with unsupervised pre-training. Fan et al. 2020 proposes a simple and effective data mining method to enlarge the training corpus of zero-shot directions. Some work also attempted to explicitly learn shared semantic representation of different languages to improve the zero-shot translation. Lu et al. 2018 suggests that by learning an explicit “interlingual” across languages, multilingual NMT model can significantly improve zero-shot translation quality. Al-Shedivat and Parikh 2019 introduces a consistent agreement-based training method that encourages the model to produce equivalent translations of parallel sentences in auxiliary languages. Different from these efforts, mRASP2 attempts to learn a universal many-to-many model, and bridge the cross-lingual representation with contrastive learning and m-RAS. The performance is very competitive both on zero-shot and supervised directions on large scale experiments.
Contrastive Learning
Contrastive Learning has become a rising domain and achieved significant success in various computer vision tasks Zhuang et al. 2019; Tian et al. 2020; He et al. 2020; Chen et al. 2020; Misra and van der Maaten 2020. Researchers in the NLP domain have also explored contrastive Learning for sentence representation. Wu et al. 2020 employed multiple sentence-level augmentation strategies to learn a noise-invariant sentence representation. Fang and Xie 2020 applies the back-translation to create augmentations of original sentences. Inspired by these studies, we apply contrastive learning for multilingual NMT.
Cross-lingual Representation
Cross-lingual representation learning has been intensively studied in order to improve cross-lingual understanding (XLU) tasks. Multilingual masked language models (MLM), such as mBERTDevlin et al. 2019 and XLMConneau and Lample 2019, train large Transformer models on multiple languages jointly and have built strong benchmarks on XLU tasks. Most of the previous works on cross-lingual representation learning focus on unsupervised training. For supervised learning, Conneau and Lample 2019 proposes TLM objective that simply concatenates parallel sentences as input. By contrast, mRASP2 leverages the supervision signal by pulling closer the representations of parallel sentences.
Conclusion
We demonstrate that contrastive learning can significantly improve zero-shot machine translation directions. Combined with additional unsupervised monolingual data, we achieve substantial improvements on all translation directions of multilingual NMT. We analyze and visualize our method, and find that contrastive learning tends to close the representation gap of different languages. Our results also show the possibilities of training a true many-to-many Multilingual NMT that works well on any translation direction. In future work, we will scale-up the current training to more languages, e.g. PC150. As such, a single model can handle more than 100 languages and outperforms the corresponding bilingual baseline.
References
Appendix A Case Study
We plot the location of multi-way parallel sentences in the representation space of mRASP2 in Figure 5 and list sentences number 1 and 100 in Table 7
Appendix B Details of Evaluation Results
We list detailed results of evaluation on a wide range of test sets.
Detailed results on OPUS-100 zero-shot evaluation set are listed in Table 8
B.2 Results on WMT
Detailed results on WMT evaluation set are listed in Table 9
Appendix C Example of AA
We show two results of sentences after AA in Figure 6
Appendix D Details of MC24
We describe the detail of MC24 in Table 10