Larger-Scale Transformers for Multilingual Masked Language Modeling
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, Alexis Conneau
Introduction
The goal of this paper is to present a study of the impact of larger capacity models on cross-lingual language understanding (XLU). We scale the capacity of XLM-R by almost two orders of magnitude while training on the same CC100 dataset Wenzek et al. (2019). Our two new multilingual masked language model dubbed XLM-RXL and XLM-RXXL, with 3.5 and 10.7 billion parameters respectively, significantly outperform the previous XLM-R model (trained in a similar setting) on cross-lingual understanding benchmarks and obtain competitive performance with the multilingual T5 models Raffel et al. (2019); Xue et al. (2020). We show that they can even outperform RoBERTa-Large Liu et al. (2019) on the GLUE benchmark Wang et al. (2018).
Recent multilingual masked language models (MLM) like mBERT Devlin et al. (2018) or XLM Lample and Conneau (2019) improved cross-lingual language understanding by pretraining large Transformer models Vaswani et al. (2017) on multiple languages at once. The XLM-R model Conneau et al. (2019) extended that approach by scaling the amount of data by two orders of magnitude, from Wikipedia to Common-Crawl and training longer, similar to RoBERTa Liu et al. (2019). These models are particularly effective for low-resource languages, where both labeled and unlabeled data is scarce. They enable supervised cross-lingual transfer, where labeled data in one language can be used to solve the same task in other languages, and unsupervised cross-lingual transfer, where low-resource language self-supervised representations are improved using additional unlabeled data from higher-resource languages. Furthermore, they reduce the need for training one model per language, and allows the use of a single - potentially much larger - pretrained model that is then fine-tuned on annotated data from many languages.
The better performance of self-supervised cross-lingual models on low-resource languages comes however at the cost of lower performance on higher-resource languages Arivazhagan et al. (2019). When the number of languages becomes large, Conneau et al. (2019) even observed an overall decrease of performance on all languages. It was hypothesized that when multilingual models get more capacity, they may showcase strong performance on both high-resource languages and low-resource languages. With only 550M parameters, the XLM-R model is now relatively small compared to new standards. Recent work scaled language models to hundreds of billions Brown et al. (2020) or even multiple trillion parameters Fedus et al. (2021), showing consistent gains in doing so. Recently, multilingual T5 showed impressive increase in performance by scaling the model capacity to tens of billions of parameters. Our study complements these findings by showing the impact of larger capacity models on the important pretraining task of multilingual masked language modeling. We show promising results for cross-lingual understanding: XLM-RXXL can both obtain a new state of the art on some cross-lingual understanding benchmarks and outperform the RoBERTa-Large model on the English GLUE benchmark Wang et al. (2018). This suggests that very large-scale multilingual models may be able to benefit from the best of both worlds: obtaining strong performance on high-resource languages while still allowing for zero-shot transfer and low-resource language understanding.
Pretraining and evaluation
In this section, we describe the model we use and how we scale it, as well as the data and tasks we use for pretraining and evaluation.
We use a Transformer model Vaswani et al. (2017) trained with the multilingual MLM objective Devlin et al. (2018); Lample and Conneau (2019) using only monolingual data. We sample streams of text from each language and train the model to predict the masked tokens in the input. We use the same learning procedure as XLM-R. We apply subword tokenization directly on raw text data using Sentence Piece Kudo and Richardson (2018) with a unigram language model Kudo (2018) just like in XLM-R. We sample batches from different languages using the same sampling distribution as Conneau et al. (2019), with , and without language embeddings. We use a large vocabulary size of 250K with a full softmax and train two different models: XLM-RXL (L = 36, H = 2560, A = 32, 3.5B params) and XLM-RXXL (L = 48, H = 4096, A = 32, 10.7B params). We pretrain the models on the CC100 dataset, which corresponds to 167B tokens in 100 languages. We compare our approach to previous results as well as the mT5 baselines, which were pretrained on the larger mC4 corpus of 6.4T tokens.
2 Evaluation
To evaluate our models, we use cross-lingual natural language inference and question answering for cross-lingual understanding, and the GLUE benchmark for monolingual English evaluation.
The XNLI dataset Conneau et al. (2018) comes with ground-truth dev and test sets in 15 languages, and a ground-truth English training set. The training set has been machine-translated to the remaining 14 languages, providing synthetic training data for these languages as well. We evaluate our model on cross-lingual transfer from English to other languages. We also consider two machine translation baselines: (i) translate-test: dev and test sets are machine-translated to English and a single English model is used (ii) translate-train-all: the English training set is machine-translated to each language and we fine-tune a multilingual model on all training sets. For the translations, we use the original data provided by the XNLI project for consistency.
Cross-lingual Question Answering.
We use MLQA and XQuad benchmarks from Lewis et al. (2019) and Artetxe et al. (2019), which extend SQuAD Rajpurkar et al. (2016) to more languages. We report F1 score and exact match (EM) score for cross-lingual transfer from English.
The English GLUE Benchmark.
We evaluate English performance on the GLUE benchmark Wang et al. (2018) which gathers multiple classification tasks, such as MNLI Williams et al. (2017), SST-2 Socher et al. (2013) or QNLI Rajpurkar et al. (2018).
3 Training details
We use model parallelism based on tensor parallel Shoeybi et al. (2019) for scaling models. XLM-RXL uses model parallel size of 2 and XLM-RXXL used 8. Compared to previous XLM-R models, we reduce the batch size and number of updates significantly to keep the compute of the new models similar (see Table 5). For both models, we use batch size of 2048 and train for 500,000 updates. We use pre-LayerNorm setting for both the models which was more stable during training.
For all the tasks in finetuning, we use batch size of 32 and train for 10 epochs. We do early stopping based on the average valid metrics across all languages and report test results.
Analysis and Results
In this section, we present our results and compare XLM-RXL and XLM-RXXL performance to other methods from previous work.
On XNLI, we observe in Table 1 that scaling the capacity from XLM-RLarge to XLM-RXL leads to an average accuracy improvement of 1.4 on zero-shot cross-lingual transfer and 1.8 on multilingual fine-tuning. When scaling even further to XLM-RXXL, we observe a total improvement of 2.2 on zero-shot and 2.4 on translate-train-all compared to XLM-RXL, with a new state of the art on French, Vietnamese and Hindi. On MLQA, in Table 4, we observe even larger gains for cross-lingual zero-shot transfer, where scaling from XLM-RLarge to XLM-RXXL leads to improvements of 4.1 F1 and 3.9 EM scores on average. Similarly, on XQuad we observe improvements of 4.4 F1 and 5.5 scores, with new state-of-the-art results on Arabic, German, Greek and Russian (see Table 3).
Comparison to monolingual English model.
For smaller-capacity models like the Base and Large version of XLM-R, it was shown that the more languages are considered the lower the performance Conneau et al. (2019), in particular on high-resource languages. For instance, XLM-RLarge was outperformed by RoBERTaLarge by 1% accuracy on average on several downstream tasks from the GLUE benchmark, as illustrated in Table2. With larger capacity, we now observe that XLM-RXXL is able to outperform RoBERTaLarge by 0.3 dev points, going from 92.9 to 93.2 average accuracy, while handling 99 more languages. While a RoBERTaXXL model may outperform XLM-RXXL, we believe it interesting to notice that with more capacity, a multilingual model can get strong high-resource performance while not losing its cross-lingual transfer ability for lower-resource languages. Given the compute needed for training such large-scale models, the possibility of training a single very large model on hundreds of languages with state-of-the-art performance on high-resource languages is an encouraging result.
Discussion and comparison to mT5.
Both mT5 and XLM-R models obtain strong performance on cross-lingual understanding benchmarks, as well as high performance on English benchmarks (see the score of 91.6 of mT5XXL on English XNLI). Many hyperparameters are however different between mT5 and XLM-R models which makes difficult an apple-to-apple comparison. First, as shown in Table 5, the mT5 models are pretrained on the much larger mC4 dataset which contains around 6.4T tokens, which is 38 times bigger than CC100 (167B tokens). While XLM-RLarge was pretrained with more updates (6T tokens), the XLM-RXL and XLM-RXXL models have seen less tokens (0.5T) during pretraining than their mT5 counterparts, although it also uses a bigger batch size (2048 over 1024 for mT5). Another difference is the context sequence length of 512 for XLM-R and 1024 for mT5. The mT5-XXL model also has slightly more parameters (13B over 10.7B). The larger number of updates combined with the larger dataset size may explain the larger improvement from the XL model to the XXL model in the case of mT5 (+3 average accuracy on XNLI), in which the additional capacity can exploit the large quantity of unlabeled mC4 data. We note however that the mT5XL is outperformed by XLM-RXL on XNLI by 0.6% on average, on XQuad by 1.3% and on MLQA by 0.9% when considering average EM score. In comparison, gains of XLM-R from the XL to the XXL architecture are only of 0.6 on average. Another explanation may be that generative models scale better than masked language models. The difference in the nature of the pretraining dataset is particularly striking when looking at the variance of performance across languages. For example the mT5XXL outperforms XLM-RXXL by 8.4 points on Swahili on XNLI zero-shot, while it only outperforms XLM-RXXL by 1.4 average accuracy. These results may suggest that the CC100 dataset gets saturated with current larger-capacity models.
Conclusion
In this study, we scaled the model capacity of the XLM-R model up to 10.7B parameters and obtained stronger performance than previous XLM-R models on cross-lingual understanding benchmarks. We show that the additional capacity allows a multilingual model to outperform a the RoBERTaLarge baseline on English benchmarks. Our technical study suggests that larger capacity multilingual model can obtain state-of-the-art cross-lingual understanding results while maintaining strong performance on high-resource languages. Our work provides an alternative to mT5 models, with new state-of-the-art performance on some languages, and publicly released code and models.