Multi-task Learning for Multilingual Neural Machine Translation

Yiren Wang, ChengXiang Zhai, Hany Hassan Awadalla

Introduction

Multilingual Neural Machine Translation (MNMT), which leverages a single NMT model to handle the translation of multiple languages, has drawn research attention in recent years Dong et al. (2015); Firat et al. (2016a); Ha et al. (2016); Johnson et al. (2017); Arivazhagan et al. (2019). MNMT is appealing since it greatly reduces the cost of training and serving separate models for different language pairs Johnson et al. (2017). It has shown great potential in knowledge transfer among languages, improving the translation quality for low-resource and zero-shot language pairs Zoph et al. (2016); Firat et al. (2016b); Arivazhagan et al. (2019).

Previous works on MNMT has mostly focused on model architecture design with different strategies of parameter sharing Firat et al. (2016a); Blackwood et al. (2018); Sen et al. (2019) or representation sharing Gu et al. (2018). Existing MNMT systems mainly rely on bitext training data, which is limited and costly to collect. Therefore, effective utilization of monolingual data for different languages is an important research question yet is less studied for MNMT.

Utilizing monolingual data (more generally, the unlabeled data) has been widely explored in various NMT and natural language processing (NLP) applications. Back translation (BT) Sennrich et al. (2016), which leverages a target-to-source model to translate the target-side monolingual data into source language and generate pseudo bitext, has been one of the most effective approaches in NMT. However, well trained NMT models are required to generate back translations for each language pair, it is computationally expensive to scale in the multilingual setup. Moreover, it is less applicable to low-resource language pairs without adequate bitext data. Self-supervised pre-training approaches Radford et al. (2018); Devlin et al. (2019); Conneau and Lample (2019); Lewis et al. (2019); Liu et al. (2020), which train the model with denoising learning objectives on the large-scale monolingual data, have achieved remarkable performances in many NLP applications. However, catastrophic forgetting effect Thompson et al. (2019), where finetuning on a task leads to degradation on the main task, limits the success of continuing training NMT on models pre-trained with monolingual data. Furthermore, the separated pre-training and finetuning stages make the framework less flexible to introducing additional monolingual data or new languages into the MNMT system.

In this paper, we propose a multi-task learning (MTL) framework to effectively utilize monolingual data for MNMT. Specifically, the model is jointly trained with translation task on multilingual parallel data and two auxiliary tasks: masked language modeling (MLM) and denoising auto-encoding (DAE) on the source-side and target-side monolingual data respectively. We further present two simple yet effective scheduling strategies for the multilingual and multi-task framework. In particular, we introduce a dynamic temperature-based sampling strategy for the multilingual data. To encourage the model to keep learning from the large-scale monolingual data, we adopt dynamic noising ratio for the denoising objectives to gradually increase the difficulty level of the tasks.

We evaluate the proposed approach on a large-scale multilingual setup with 1010 language pairs from the WMT datasets. We study three English-centric multilingual systems, including many-to-English, English-to-many, and many-to-many. We show that the proposed MTL approach significantly boosts the translation quality for both high-resource and low-resource languages. Furthermore, we demonstrate that MTL can effectively improve the translation quality on zero-shot language pairs with no bitext training data. In particular, MTL achieves even better performance than the pivoting approach for multiple low-resource language pairs. We further show that MTL outperforms pre-training approaches on both NMT tasks as well as cross-lingual transfer learning for NLU tasks, despite being trained on very small amount of data in comparison to pre-training approaches.

The contributions of this paper are as follows. First, we propose a new MTL approach to effectively utilize monolingual data for MNMT. Second, we introduce two simple yet effective scheduling strategies, namely the dynamic temperature-based sampling and dynamic noising ratio strategy. Third, we present detailed ablation studies to analyze various aspects of the proposed approach. Finally, we demonstrate for the first time that MNMT with MTL models can be effectively used for cross-lingual transfer learning for NLU tasks with similar or better performance than the state-of-the-art massive scale pre-trained models using single task.

Background

NMT adopts the sequence-to-sequence framework, which consists of an encoder and a decoder network built upon deep neural networks Sutskever et al. (2014); Bahdanau et al. (2014); Gehring et al. (2017); Vaswani et al. (2017). The input source sentence is mapped into context representations in a continuous representation space by the encoder, which are then fed into the decoder to generate the output sentence. Given a language pair (x,y)(x,y), the objective of the NMT model training is to maximize the conditional probability P(y∣x;θ)P(y|x;\theta) of the target sentence given the source sentence.

NMT heavily relies on high-quality and large-scale bitext data. Various strategies have been proposed to augment the limited bitext by leveraging the monolingual data. Back translation Sennrich et al. (2016) utilizes the target-side monolingual data. Self learning Zhang and Zong (2016) leverages the source-side monolingual data. Dual learning paradigms utilize monolingual data in both source and target language He et al. (2016); Wang et al. (2019); Wu et al. (2019). While these approaches can effectively improve the NMT performance, they have two limitations. First, they introduce additional cost in model training and translation generation, and therefore are less efficient when scaling to the multilingual setting. Second, back translation requires a good baseline model with adequate bitext data to start from, which limits its efficiency on low-resource settings.

Multilingual NMT

MNMT aims to train a single translation model that translates between multiple language pairs Firat et al. (2016a); Johnson et al. (2017). Previous works explored the model architecture design with different parameter sharing strategies, such as partial sharing with shared encoder Dong et al. (2015); Sen et al. (2019), shared attention Firat et al. (2016a), task-specific attention Blackwood et al. (2018), and full model sharing with language identifier Johnson et al. (2017); Ha et al. (2016); Arivazhagan et al. (2019). There are also extensive studies on representation sharing that shares lexical, syntactic, or sentence level representations across different languages Zoph et al. (2016); Nguyen and Chiang (2017); Gu et al. (2018). The models in these works rely on bitext for training, and the largely available monolingual data has not been effectively leveraged.

Self-supervised Learning

This work is motivated by the recent success of self-supervised learning for NLP applications Radford et al. (2018); Devlin et al. (2019); Lample et al. (2018a, b); Conneau and Lample (2019); Lewis et al. (2019); Liu et al. (2020). Different denoising objectives have been designed to train the neural networks on large-scale unlabeled text. In contrast to previous work in pre-training with separated self-supervised pre-training and supervised finetuning stages, we focus on a multi-task setting to jointly train the MNMT model on both bitext and monolingual data.

Multi-task Learning

Multi-task learning (MTL) Caruana (1997), which trains the model on several related tasks to encourage representation sharing and improve generalization performance, has been successfully used in many different machine learning applications Collobert and Weston (2008); Deng et al. (2013); Ruder (2017). In the context of NMT, MTL has been explored mainly to inject linguistic knowledge Luong et al. (2015); Niehues and Cho (2017); Eriguchi et al. (2017); Zaremoodi and Haffari (2018); Kiperwasser and Ballesteros (2018) with tasks such as part-of-speech tagging, dependency parsing, semantic parsing, etc. In this work, we instead focus on auxiliary self-supervised learning tasks to leverage the monolingual data.

Approach

The main task in the MTL framework is the translation task trained on bitext corpora DBD_{B} of sentence pairs (x,y)(x,y) with the cross-entropy loss:

With the large amount of monolingual data in different languages, we can train language models on both source-side For the English-to-Many translation model, the source-side language is English; for Many-to-English and Many-to-Many, it refers the set of all other languages. Similarly for the target-side language. and target-side languages. We introduce two denoising language modeling tasks to help improve the quality of the translation model: the masked language model (MLM) task and the denoising auto-encoding (DAE) task.

In the masked language model (MLM) task Devlin et al. (2019), sentences with tokens randomly masked are fed into the model and the model attempts to predict the masked tokens based on their context. MLM is beneficial for learning deep bidirectional representations. We introduce MLM as an auxiliary task to improve the quality of the encoder representations especially for the low-resource languages. As is illustrated in Figure 1(a), we add an additional output layer to the encoder of the translation model and train the encoder with MLM on source-side monolingual data. The output layer is dropped during inference. The cross entropy loss for predicting the masked tokens is denoted as LMLM\mathcal{L}_{MLM}.

Following BERT Devlin et al. (2019), we randomly sample RM%R_{M}\% units in the input sentences and replace them with a special [MASK] token. A unit can either be a subword token, or a word consists of one or multiple subword tokens. We refer to them as token-level and word-level MLM.

Denoising Auto-Encoding (DAE)

Denoising auto-encoding (DAE) Vincent et al. (2008) has been demonstrated to be an effective strategy for unsupervised NMT Lample et al. (2018a, b). Given a monolingual corpus DMD_{M} and a stochastic noising model CC, DAE minimizes the reconstruction loss as shown in Eqn 2:

As is illustrated in Figure 1(b), we train all model parameters with DAE on the target-side monolingual data. Specifically, we feed the target-side sentence to the noising model CC and append the corresponding language ID symbol; the model then attempts to reconstruct the original sentence.

We introduce three types of noises for the noising model CC. 1) Text Infilling Lewis et al. (2019): Following Liu et al. (2020), we randomly sample RD%R_{D}\% text spans with span lengths drawn from a Poisson distribution (λ=3.5\lambda=3.5). We replace all words in each span with a single blanking token. 2) Word Drop & Word Blank: we randomly sample words from each input sentence, which are either removed or replaced with blanking tokens for each token position. 3) Word Swapping: we slightly shuffle the order of words in the input sentence. Following Lample et al. (2018a), we apply a random permutation σ\sigma with condition ∣σ(i)−i∣≤k,∀i∈{1,n}|\sigma(i)-i|\leq k,\forall i\in\{1,n\}, where nn is the length of the input sentence, and k=3k=3 is the maximum swapping distance.

Joint Training

In the training process, the two self-learning objectives are combined with the cross-entropy loss for the translation task:

In particular, we use bitext data for the translation objective, source-side monolingual data for MLM, and target-side monolingual data for the DAE objective. A language ID symbol [LID] of the target language is appended to the input sentence in the translation and DAE tasks.

2 Task Scheduling

The scheduling of tasks and data associated with the task is important for multi-task learning. We further introduce two simple yet effective scheduling strategies in the MTL framework.

One serious yet common problem for MNMT is data imbalance across different languages. Training the model with the true data distribution would starve the low-resource language pairs. Temperature-based batch balancing Arivazhagan et al. (2019) is demonstrated to be an effective heuristic to ease the problem. For language pair ll with bitext corpus DlD_{l}, we sample instances with probability proportional to (∣Dl∣∑k∣Dk∣)1T(\frac{|D_{l}|}{\sum_{k}|D_{k}|})^{\frac{1}{T}}, where TT is the sampling temperature.

While MNMT greatly improves translation quality for low-resource languages, performance deterioration is generally observed for high resource languages. One hypothesized reason is that the model might converge before well trained on high-resource data Bapna and Firat (2019). To alleviate this problem, we introduce a simple heuristic to feed more high-resource language pairs in the early stage of training and gradually shift more attention to the low-resource languages. To achieve this, we modify the sampling strategy by introducing dynamic sampling temperature T(k)T(k) as a function of the number of training epochs kk. We use a simple linear functional form for T(k)T(k):

Where T0T_{0} and TmT_{m} are the initial and maximum value for sampling temperature respectively. NN is the number of warm-up epochs. The sampling temperature starts from a smaller value T0T_{0}, resulting in sampling leaning towards true data distribution. T(k)T(k) gradually increases in the training process to encourage over-sampling low-resource languages more to avoid them getting starved.

Dynamic Noising Ratio

We further schedule the difficulty level of MLM and DAE from easier to more difficult. The main motivation is that training algorithms perform better when starting with easier tasks and gradually move to harder ones as promoted in curriculum learning Elman (1993). Furthermore, increasing the learning difficulty can potentially help avoid saturation and encourage the model to keep learning from abundant data.

Given the monolingual data, the difficulty level of MLM and DAE tasks mainly depends on the noising ratio. Therefore, we introduce dynamic noising ratio R(k)R(k) as a function of training steps:

Where R0R_{0} and RmR_{m} are the lower and upper bound for noising ratio respectively and MM is the number of warm-up epochs. Noising ratio RR refers to the masking ratio RMR_{M} in MLM and the blanking ratio RDR_{D} of the Text Infilling task for DAE.

Experimental Setup

We evaluate MTL on a multilingual setting with 10 languages to and from English (En), including French (Fr), Czech (Cs), German (De), Finnish (Fi), Latvian (Lv), Estonian (Et), Romanian (Ro), Hindi (Hi), Turkish (Tr) and Gujarati (Gu).

The bitext training data comes from the WMT corpus. Detailed desciption and statistics can be found in Appendix A.

Monolingual Data

The monolingual data we use is mainly from NewsCrawlhttp://data.statmt.org/news-crawl/. We apply a series of filtration rules to remove the low-quality sentences, including duplicated sentences, sentences with too many punctuation marks or invalid characters, sentences with too many or too few words, etc. We randomly select 55M filtered sentences for each language. For low-resource languages without enough sentences from NewsCrawl, we leverage data from CCNet Wenzek et al. (2019).

Back Translation

We use the target-to-source bilingual models to back translate the target-side monolingual sentences into the source domain for each language pair. The synthetic parallel data from back translation is mixed and shuffled with bitext and used together for the translation objective in training. We use the same monolingual data for back translation as the multi-task learning in all our experiments for fair comparison.

2 Model Configuration

We use Transformer for all our experiments using the PyTorch implementationhttps://github.com/pytorch/fairseq Ott et al. (2019). We adopt the transformer_big setting Vaswani et al. (2017) with a 66-layer encoder and decoder. The dimensions of word embeddings, hidden states, and non-linear layer are set as 10241024, 10241024 and 40964096 respectively, the number of heads for multi-head attention is set as 1616. We use a smaller model setting for the bilingual models on low-resource languages Tr, Hi and Gu (with 33 encoder and decoder layers, 256256 embedding and hidden dimension) to avoid overfitting and acquire better performance.

We study three multilingual translation scenarios including many-to-English (X→\toEn), English-to-many (En→\toX) and many-to-many (X→\toX). For the multilingual model, we adopt the same Transformer architecture as the bilingual setting, with parameters fully shared across different language pairs. A target language ID token is appended to each input sentence.

3 Training and Evaluation

All models are optimized with Adam Kingma and Ba (2015) with β1=0.9\beta_{1}=0.9, β2=0.98\beta_{2}=0.98. We set the learning rate schedule following Vaswani et al. (2017) with initial learning rate 5×10−45\times 10^{-4}. Label smoothing Szegedy et al. (2016) is adopted with 0.10.1. The models are trained on 88 V100 GPUs with a batch size of 40964096 and the parameters are updated every 1616 batches. During inference, we use beam search with a beam size of 55 and length penalty 1.01.0. The BLEU score is measured by the de-tokenized case-sensitive SacreBLEUSacreBLEU signatures: BLEU+case.mixed+lang.l1−l1-l2 numrefs.1+smooth.exp+test.SET+tok.13a+version.1.4.3,whereSET+tok.13a+version.1.4.3, wherel1, l2arethelanguagecode(Table9),l2 are the language code (Table 9),SET is the corresponding test set for the language pair. Post (2018).

Results

We compare the performance of the bilingual models (Bilingual), multilingual models trained on bitext only, trained on both bitext and back translation (+BT) and trained with the proposed multi-task learning (+MTL). Translation results of the 1010 languages translated to and from English are presented in Table 1 and 2 respectively. We can see that:

1. Bilingual vs. Multilingual: The multilingual baselines perform better on lower-resource languages, but perform worse than individual bilingual models on high-resource languages like Fr, Cs and De. This is in concordance with the previous observations Arivazhagan et al. (2019) and is consistent across the three multilingual systems (i.e., X→\toEn, En→\toX and X→\toX).

2. Multi-task learning: Models trained with multi-task learning (+MTL) significantly outperform the multilingual baselines for all the languages pairs in all three multilingual systems, demonstrating the effectiveness of the proposed framework.

3. Back Translation: With the same monolingual corpus, MTL achieves better performance on some language pairs (e.g. Fr→\toEn, Gu→\toEn), while getting outperformed on some others, especially on the En→\toX direction. However, back translation is computationally expensive as it involves the additional procedure of training 1010 bilingual models (2020 for the X→\toX system) and generating translations for each monolingual sentence. Combining MTL with BT (+BT+MTL) introduces further improvements for most language pairs without using any additional monolingual data. This suggests that when there is enough computation budget for BT, MTL can still be leveraged to provide good complementary improvement.

2 Zero-shot Translation

We further evaluate the proposed approach on zero-shot translation of non English-centric language pairs. We compare the performances of the pivoting method, the X→\toX baseline system, X→\toX with BT, and with MTL. For the pivoting method, the source language is translated into English first, and then translated into the target language De Gispert and Marino (2006); Utiyama and Isahara (2007). We evaluate on a group of high-resource languages with a multi-way parallel test set for De, Cs, Fr and En, constructed by newstest2009 with 30273027 sentences and that of a group of low-resource languages Et, Hi, Tr and Hi (995995 sentences). The results are shown in Table 3 and 4 respectively.

Utilizing monolingual data with MTL significantly improves the zero-shot translation quality of the X→\toX system, further demonstrating the effectiveness of the proposed approach. In particular, MTL achieves significantly better results than the pivoting approach on the high-resource pair Cs→\toDe and almost all low-resource pairs. Furthermore, leveraging monolingual data through BT does not perform well for many low-resource language pairs, resulting in comparable and even downgraded performances. We conjecture that this is related to the quality of the back translations. MTL helps overcome such limitations with the auxiliary self-supervised learning tasks.

3 MTL vs. Pre-training

We also compare MTL with mBART Liu et al. (2020), the state-of-the-art multilingual pre-training method for NMT. We adopt the officially released mBART model pre-trained on CC25 corpushttps://dl.fbaipublicfiles.com/fairseq/models/mbart/mbart.CC25.tar.gz and finetune the model on the same bitext training data used in MTL for each language pair. As shown in Figure 2, MTL outperforms mBART on all language pairs. This suggests that in the scenario of NMT, jointly training the model with MT task and self-supervised learning tasks could be a better task design than the separated pre-training and finetuning stages. It is worth noting that mBart is utilizing much more monolingual data; for example, it uses 5555B English tokens and 1010B French tokens, while our approach is using just 100100M tokens each. This indicates that MTL is more data efficient.

4 Multi-task Objectives

We present ablation study on the learning objectives of the multi-task learning framework. We compare performance of multilingual baseline model with translation objective only, jointly learning translation with MLM, jointly learning translation with DAE, and the combination of all objectives. Table 5 shows the results on a high-resource pair De↔\leftrightarrowEn and low-resource pair Tr↔\leftrightarrowEn. We can see that introducing MLM or DAE can both effectively improve the performance of multilingual systems, and the combination of both yields the best performance. We also observe that MLM is more beneficial for ‘→\toEn’ compared with ‘En→\to’ direction, especially for the low-resource languages. This is in concordance with our intuition that the MLM objective contributes to improving the encoder quality and source-side language modeling for low-resource languages.

5 Dynamic Sampling Temperature

We study the effectiveness of the proposed dynamic sampling strategy. We compare multilingual systems using a fixed sampling temperature T=5T=5 with systems using dynamic temperature T(k)T(k) defined in Equation 4. We set T0=1,Tm=5,N=5T_{0}=1,T_{m}=5,N=5, which corresponds to gradually increasing the temperature from 11 to 55 with 55 training epochs and saturate to T=5T=5 afterwards. The results for X→\toEn and En→\toX systems are presented in Figure 3 and 4 respectively, where we report Δ\DeltaBLEU relative to their corresponding bilingual baseline model that was evaluated on the individual validation sets for each language pairs. The dynamic temperature strategy improves the quality for high-resource language pairs (e.g. Fr→\toEn, De→\toEn, En→\toFr), while introducing minimum effect for mid-resource languages (Lv). Surprisingly, the proposed strategy also greatly boosts performance for low-resource languages Tr and Gu, with over +1+1 BLEU gain for both to and from English direction.

6 Noising Scheme

We study the effect of different noising schemes in the MLM and DAE objectives. As introduced in Section 3.1, we have token-level and word-level masking scheme for MLM depending on the unit of masking. We also have two noising schemes for DAE, where the Text Infilling task blanks a span of words (span-level), and the Word Blank task blanks the input sentences at word-level. We compare performance of these different noising schemes on X→\toEn system as shown in Figure 5.

We report Δ\DeltaBLEU relative to the multilingual X→\toEn baseline on the corresponding language pairs for each noising scheme. As we can see, the model benefits most from the word-level MLM and the span-level Text Infilling task for DAE. This is in concordance with the intuition that the Text Infilling task teaches the model to predict the length of masked span and the exact tokens at the same time, making it a harder task to learn. We use the word-level MLM and span-level DAE as the best recipe for our MTL framework.

7 Noising Ratio Scheduling

In our initial experiments, we found that the dynamic noising ratio strategy does not effectively improve the performance. We suspect that it is due to the limitation of data scale. We experiment with a larger scale setting by increasing the amount of monolingual data from 55M sentences for each language to 2020M. For low-resource languages without enough data, we take the full available amount (1818M for Lv, 1111M for Et, 5.25.2M for Gu).

Table 6 shows results on X→\toEn MNMT model with large-scale monolingual data setting. We compare the performance of multilingual with back translation baseline, a model with MTL and a model with both MTL and dynamic noising ratio. For the dynamic noising ratio, we set the masking ratio for MLM to increase from 10%10\% to 20%20\% and blanking ratio for DAE to increase from 20%20\% to 40%40\%. As we can see, the dynamic noising strategy helps boost performance for mid-resource languages like Lv and Et, while introducing no negative effect to other languages. For future study, we would like to cast the dynamic noising ratio over different subsets of monolingual datasets to prevent the model from learning to copy and memorize.

8 MTL for Cross-Lingual Transfer Learning for NLU

Large scale pre-trained cross-lingual language models such as mBERT Devlin et al. (2019) and XLM-Roberta Conneau et al. (2020) are the state-of-the-art for cross-lingual transfer learning on natural language understanding (NLU) tasks, such as XNLI Conneau et al. (2018) and XGLUE Liang et al. (2020). Such models are trained on massive amount of monolingual data from all language as a masked language model. It has been shown that massive MNMT models are not able to match the performance of pre-trained language models such as XLM-Roberta on NLU downstream tasks Siddhant et al. (2020). In Siddhant et al. (2020), the MNMT models are massive scale models trained only on the NMT task. They are not able to outperform XLM-Roberta, which is trained with MLM task without any parallel data. In this work, we evaluate the effectiveness of our proposed MTL approach for cross-lingual transfer leaning on NLU application. Intuitively, MTL can bridge this gap since it utilizes NMT, MLM and DAE objectives.

In the experiment, we train a system on 66 languages using both bitext and monolingual data. For the bitext training data, we use 3030M parallel sentences per language pair from in-house data crawled from the web. For the monolingual data, we use 4040M sentences per language from CCNet Wenzek et al. (2019). Though this is a relatively large-scale setup, it only leverages a fraction of the data used to train XLM-Roberta for those languages. We train the model with 1212 layers encoders and 66 layers decoder. The hidden dimension is 768768 and the number of heads is 88. We tokenize all data with the SentencePiece model Kudo and Richardson (2018) with the vocabulary size of 6464K. We train a many-to-many MNMT system with three tasks described in Section 3.1: NMT, MLM, and DAE. Once the model is trained, we use the encoder only and discard the decoder. We add a feedforward layer for the downstream tasks.

As shown in Table 7, MTL outperform both XLM-Roberta and MMTE Siddhant et al. (2020) which are trained on massive amount of data in comparison to our system. XLM-Roberta is trained only on MLM task and MMTE is trained only on NMT task. Our MTL system is trained on three tasks. The results clearly highlight the effectiveness of multi-task learning, and demonstrate that it can outperform single-task systems trained on massive amount of data. We observe the same pattern in Table 8 with XGLUE NER task, which outperforms SOTA XLM-Roberta model.

Conclusion

In this work, we propose a multi-task learning framework that jointly trains the model with the translation task on bitext data, the masked language modeling task on the source-side monolingual data and the denoising auto-encoding task on the target-side monolingual data. We explore data and noising scheduling approaches and demonstrate their efficacy for the proposed approach. We show that the proposed MTL approach can effectively improve the performance of MNMT on both high-resource and low-resource languages with large margin, and can also significantly improve the translation quality for zero-shot language pairs without bitext training data. We showed that the proposed approach is more effective than pre-training followed by finetuning for NMT. Furthermore, we showed the effectiveness of multitask learning for cross-lingual downstream tasks outperforming SOTA larger models trained on single task.

For future work, we are interested in investigating the proposed approach in a scaled setting with more languages and a larger amount of monolingual data. Scheduling the different tasks and different types of data would be an interesting problem. Furthermore, we would also like to explore the most sample efficient strategy to add a new language to a trained MNMT system.

Acknowledgment

We would like to thank Alex Muzio for helping with zero-shot scoring and useful discussions. We also would like to thank Dongdong Zhang for helping with mBART comparison and Felipe Cruz Salinas for helping with XNLI and XGLUE scoring.

References

Appendices

Appendix A Bitext Training Data

We concatenate all resources except WikiTitles provided by WMT of the latest available year and filter out duplicated pairs and pairs with the same source and target sentence. For Fr and Cs, we randomly sample 1010M sentence pairs from the full corpus. The detailed statistics of bitext data can be found in Table 9.

We randomly sample 1,0001,000 sentence pairs from each individual validation set and concatenate them to construct a multilingual validation set. We tokenize all data with the SentencePiece model Kudo and Richardson (2018), forming a vocabulary shared by all the source and target languages with 3232k tokens for bilingual models (1616k for Hi and Gu) and 6464k tokens for multilingual models.