XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei
Introduction
It has become a de facto trend to use a pretrained language model Devlin et al. (2019); Dong et al. (2019); Yang et al. (2019b); Bao et al. (2020) for downstream NLP tasks. These models are typically pretrained with masked language modeling objectives, which learn to generate the masked tokens of an input sentence. In addition to monolingual representations, the masked language modeling task is effective for learning cross-lingual representations. By only using multilingual corpora, such pretrained models perform well on zero-shot cross-lingual transfer Devlin et al. (2019); Conneau et al. (2020), i.e., fine-tuning with English training data while directly applying the model to other target languages. The cross-lingual transferability can be further improved by introducing external pre-training tasks using parallel corpus, such as translation language modeling Conneau and Lample (2019), and cross-lingual contrast Chi et al. (2021b). However, previous cross-lingual pre-training based on masked language modeling usually requires massive computation resources, rendering such models quite expensive. As shown in Figure 1, our proposed XLM-E achieves a huge speedup compared with well-tuned pretrained models.
In this paper, we introduce ELECTRA-style tasks Clark et al. (2020b) to cross-lingual language model pre-training. Specifically, we present two discriminative pre-training tasks, namely multilingual replaced token detection, and translation replaced token detection. Rather than recovering masked tokens, the model learns to distinguish the replaced tokens in the corrupted input sequences. The two tasks build input sequences by replacing tokens in multilingual sentences, and translation pairs, respectively. We also describe the pre-training algorithm of our model, XLM-E, which is pretrained with the above two discriminative tasks. It provides a more compute-efficient and sample-efficient way for cross-lingual language model pre-training.
We conduct extensive experiments on the XTREME cross-lingual understanding benchmark to evaluate and analyze XLM-E. Over seven datasets, our model achieves competitive results with the baseline models, while only using 1% of the computation cost comparing to XLM-R. In addition to the high computational efficiency, our model also shows the cross-lingual transferability that achieves a reasonably low transfer gap. We also show that the discriminative pre-training encourages universal representations, making the text representations better aligned across different languages.
Our contributions are summarized as follows:
We explore ELECTRA-style tasks for cross-lingual language model pre-training, and pretrain XLM-E with both multilingual corpus and parallel data.
We demonstrate that XLM-E greatly reduces the computation cost of cross-lingual pre-training.
We show that discriminative pre-training tends to encourage better cross-lingual transferability.
Background: ELECTRA
ELECTRA Clark et al. (2020b) introduces the replaced token detection task for language model pre-training, with the goal of distinguishing real input tokens from corrupted tokens. That means the text encoders are pretrained as discriminators rather than generators, which is different from the previous pretrained language models, such as BERT Devlin et al. (2019), that learn to predict the masked tokens. The ELECTRA pre-training task has shown good performance on various data, such as language Hao et al. (2021), and vision Fang et al. (2022).
ELECTRA trains two Transformer Vaswani et al. (2017) encoders, serving as generator and discriminator, respectively. The generator is typically a small BERT model trained with the masked language modeling (MLM; Devlin et al. 2019) task. Consider an input sentence containing tokens. MLM first randomly selects a subset as the positions to be masked, and construct the masked sentence by replacing tokens in with [MASK]. Then, the generator predicts the probability distributions of the masked tokens . The loss function of the generator is:
The discriminator is trained with the replaced token detection task. Specifically, the discriminator takes the corrupted sentences as input, which is constructed by replacing the tokens in with the tokens sampled from the generator :
Then, the discriminator predicts whether is original or sampled from the generator. The loss function of the discriminator is
where represents the label of whether is the original token or the replaced one. The final loss function of ELECTRA is the combined loss of the generator and discriminator losses, .
Compared to generative pre-training, ELECTRA uses more model parameters and training FLOPs per step, because it contains a generator and a discriminator during pre-training. However, only the discriminator is used for fine-tuning on downstream tasks, so the size of the final checkpoint is similar to BERT-like models in practice.
Methods
Figure 2 shows an overview of the two discriminative tasks used for pre-training XLM-E. Similar to ELECTRA described in Section 2, XLM-E has two Transformer components, i.e., generator and discriminator. The generator predicts the masked tokens given the masked sentence or translation pair, and the discriminator distinguishes whether the tokens are replaced by the generator.
The pre-training tasks of XLM-E are multilingual replaced token detection (MRTD), and translation replaced token detection (TRTD).
The multilingual replaced token detection task requires the model to distinguish real input tokens from corrupted multilingual sentences. Both the generator and the discriminator are shared across languages. The vocabulary is also shared for different languages. The task is the same as in monolingual ELECTRA pre-training (Section 2). The only difference is that the input texts can be in various languages.
We use uniform masking to produce the corrupted positions. We also tried span masking Joshi et al. (2019); Bao et al. (2020) in our preliminary experiments. The results indicate that span masking significantly weakens the generator’s prediction accuracy, which in turn harms pre-training.
Translation Replaced Token Detection
Parallel corpora are easily accessible and proved to be effective for learning cross-lingual language models Conneau and Lample (2019); Chi et al. (2021b), while it is under-studied how to improve discriminative pre-training with parallel corpora. We introduce the translation replaced token detection task that aims to distinguish real input tokens from translation pairs. Given an input translation pair, the generator predicts the masked tokens in both languages. Consider an input translation pair . We construct the input sequence by concatenating the translation pair as a single sentence. The loss function of the generator is:
where is the operator of concatenation, and stand for the randomly selected masked positions for and , respectively. This loss function is identical to the translation language modeling loss (TLM; Conneau and Lample 2019). The discriminator learns to distinguish real input tokens from the corrupted translation pair. The corrupted translation pair is constructed by replacing tokens with the tokens sampled from with the concatenated translation pair as input. Formally, is constructed by
The same operation is also used to construct . Then, the loss function of the discriminator can be written as
where represents the label of whether the -th input token is the original one or the replaced one. The final loss function of the translation replaced token detection task is .
2 Pre-training XLM-E
The XLM-E model is jointly pretrained with the masked language modeling, translation language modeling, multilingual replaced token detection and the translation replaced token detection tasks. The overall training objective is to minimize
over large scale multilingual corpus and parallel corpus . We jointly pretrain the generator and the discriminator from scratch. Following Clark et al. (2020b), we make the generator smaller to improve the pre-training efficiency.
3 Gated Relative Position Bias
Inspired by the gating mechanism of Gated Recurrent Unit (GRU; Cho et al. 2014), we compute gated relative position bias via:
Compared with relative position bias Parikh et al. (2016); Raffel et al. (2020); Bao et al. (2020), the proposed gates take the content into consideration, which adaptively adjusts the relative position bias by conditioning on input tokens. Intuitively, the same distance between two tokens tends to play different roles in different languages.
4 Initialization of Transformer Parameters
Properly initializing Transformer parameters is critical to stabilize large-scale training. First, all the parameters are randomly initialized by uniformly sampling from a small range, such as . Second, for the -th Transformer blockEach block contains a self-attention layer and a feed-forward network layer., we rescale the attention output weight and the feed-forward network output matrix by . Notice that the Transformer block after the embedding layer is regarded the first one.
Experiments
We use the CC-100 Conneau et al. (2020) dataset for the replaced token detection task. CC-100 contains texts in 100 languages collected from the CommonCrawl dump. We use parallel corpora for the translation replaced token detection task, including translation pairs in 100 languages collected from MultiUN Ziemski et al. (2016), IIT Bombay Kunchukuttan et al. (2018), OPUS Tiedemann (2012), WikiMatrix Schwenk et al. (2019), and CCAligned El-Kishky et al. (2020).
Following XLM Conneau and Lample (2019), we sample multilingual sentences to balance the language distribution. Formally, consider the pre-training corpora in languages with examples for the -th language. The probability of using an example in the -th language is
The exponent controls the distribution such that a lower increases the probability of sampling examples from a low-resource language. In this paper, we set .
Model
We use a Base-size -layer Transformer Vaswani et al. (2017) as the discriminator, with hidden size of , and FFN hidden size of . The generator is a -layer Transformer using the same hidden size as the discriminator Meng et al. (2021). See Appendix A for more details of model hyperparameters.
Training
We jointly pretrain the generator and the discriminator of XLM-E from scratch, using the Adam Kingma and Ba (2015) optimizer for 125K training steps. We use dynamic batching of approximately 1M tokens for each pre-training task. We set , the weight for the discriminator objective to 50. The whole pre-training procedure takes about 1.7 days on 64 Nvidia A100 GPU cards. See Appendix B for more details of pre-training hyperparameters.
2 Cross-lingual Understanding
We evaluate XLM-E on the XTREME Hu et al. (2020b) benchmark, which is a multilingual multi-task benchmark for evaluating cross-lingual understanding. The XTREME benchmark contains seven cross-lingual understanding tasks, namely part-of-speech tagging on the Universal Dependencies v2.5 Zeman et al. (2019), NER named entity recognition on the Wikiann Pan et al. (2017); Rahimi et al. (2019) dataset, cross-lingual natural language inference on XNLI Conneau et al. (2018), cross-lingual paraphrase adversaries from word scrambling (PAWS-X; Yang et al. 2019a), and cross-lingual question answering on MLQA Lewis et al. (2020), XQuAD Artetxe et al. (2020), and TyDiQA-GoldP Clark et al. (2020a).
We compare our XLM-E model with the cross-lingual language models pretrained with multilingual text, i.e., Multilingual BERT (mBert; Devlin et al. 2019), mT5 Xue et al. (2021), and XLM-R Conneau et al. (2020), or pretrained with both multilingual text and parallel corpora, i.e., XLM Conneau and Lample (2019), InfoXLM Chi et al. (2021b), and XLM-Align Chi et al. (2021c). The compared models are all in Base size. In what follows, models are considered as in Base size by default.
Results
We use the cross-lingual transfer setting for the evaluation on XTREME Hu et al. (2020b), where the models are first fine-tuned with the English training data and then evaluated on the target languages. In Table 1, we report the accuracy, F1, or Exact-Match (EM) scores on the XTREME cross-lingual understanding tasks. The results are averaged over all target languages and five runs with different random seeds. We divide the pretrained models into two categories, i.e., the models pretrained on multilingual corpora, and the models pretrained on both multilingual corpora and parallel corpora. For the first setting, we pretrain XLM-E with only the multilingual replaced token detection task. From the results, it can be observed that XLM-E outperforms previous models on both settings, achieving the averaged scores of 67.6 and 69.3, respectively. Compared to XLM-R, XLM-E (w/o TRTD) produces an absolute 1.2 improvement on average over the seven tasks. For the second setting, compared to XLM-Align, XLM-E produces an absolute 0.4 improvement on average. XLM-E performs better on the question answering tasks and sentence classification tasks while preserving reasonable high F1 scores on structured prediction tasks. Despite the effectiveness of XLM-E, our model requires substantially lower computation cost than XLM-R and XLM-Align. A detailed efficiency analysis in presented in Section 4.5.
3 Ablation Studies
For a deeper insight to XLM-E, we conduct ablation experiments where we first remove the TRTD task and then remove the gated relative position bias. Besides, we reimplement XLM that is pretrained with the same pre-training setup with XLM-E, i.e., using the same training steps, learning rate, etc. Table 2 shows the ablation results on XNLI and MLQA. Removing TRTD weakens the performance of XLM-E on both downstream tasks. On this basis, the results on MLQA further decline when removing the gated relative position bias. This demonstrates that XLM-E benefits from both TRTD and the gated relative position bias during pre-training. Besides, XLM-E substantially outperform XLM on both tasks. Notice that when removing the two components from XLM-E, our model only requires a multilingual corpus, but still achieves better performance than XLM, which uses an additional parallel corpus.
4 Scaling-up Results
Scaling-up model size has shown to improve performance on cross-lingual downstream tasks Xue et al. (2021); Goyal et al. (2021). We study the scalability of XLM-E by pre-training XLM-E models using larger model sizes. We consider two larger model sizes in our experiments, namely Large and XL. Detailed model hyperparameters can be found in Appendix A. As present in Table 3, XLM-E achieves the best performance while using significantly fewer parameters than its counterparts. Besides, scaling-up the XLM-E model size consistently improves the results, demonstrating the effectiveness of XLM-E for large-scale pre-training.
5 Training Efficiency
We present a comparison of the pre-training resources, to explore whether XLM-E provides a more compute-efficient and sample-efficient way for pre-training cross-lingual language models. Table 4 compares the XTREME average score, the number of parameters, and the pre-training computation cost. Notice that InfoXLM and XLM-Align are continue-trained from XLM-R, so the total training FLOPs are accumulated over XLM-R.
Table 4 shows that XLM-E substantially reduces the computation cost for cross-lingual language model pre-training. Compared to XLM-R and XLM-Align that use at least 9.6e21 training FLOPs, XLM-E only uses 9.5e19 training FLOPs in total while even achieving better XTREME performance than the two baseline models. For the setting of pre-training with only multilingual corpora, XLM-E (w/o TRTD) also outperforms XLM-R using 6.3e19 FLOPs in total. This demonstrates the compute-effectiveness of XLM-E, i.e., XLM-E as a stronger cross-lingual language model requires substantially less computation resource.
6 Cross-lingual Alignment
To explore whether discriminative pre-training improves the resulting cross-lingual representations, we evaluate our model on the sentence-level and word-level alignment tasks, i.e., cross-lingual sentence retrieval and word alignment.
We use the Tatoeba Artetxe and Schwenk (2019) dataset for the cross-lingual sentence retrieval task, the goal of which is to find translation pairs from the corpora in different languages. Tatoeba consists of English-centric parallel corpora covering 122 languages. Following Chi et al. (2021b) and Hu et al. (2020b), we consider two settings where we use 14 and 36 of the parallel corpora for evaluation, respectively. The sentence representations are obtained by average pooling over hidden vectors from a middle layer. Specifically, we use layer-7 for XLM-R and layer-9 for XLM-E. Then, the translation pairs are induced by the nearest neighbor search using the cosine similarity. Table 5 shows the average accuracy@1 scores under the two settings of Tatoeba for both the xx en and en xx directions. XLM-E achieves 74.4 and 72.3 accuracy scores for Tatoeba-14, and 65.0 and 62.3 accuracy scores for Tatoeba-36, providing notable improvement over XLM-R. XLM-E performs slightly worse than InfoXLM. We believe the cross-lingual contrast Chi et al. (2021b) task explicitly learns the sentence representations, which makes InfoXLM more effective for the cross-lingual sentence retrieval task.
For the word-level alignment, we use the word alignment datasets from EuroParlwww-i6.informatik.rwth-aachen.de/goldAlignment/, WPT2003web.eecs.umich.edu/~mihalcea/wpt/, and WPT2005web.eecs.umich.edu/~mihalcea/wpt05/, containing 1,244 translation pairs annotated with golden alignments. The predicted alignments are evaluated by alignment error rate (AER; Och and Ney 2003):
where and stand for the predicted alignments, the annotated sure alignments, and the annotated possible alignments, respectively. In Table 6 we compare XLM-E with baseline models, i.e., fast_align Dyer et al. (2013), XLM-R, and XLM-Align. The resulting word alignments are obtained by the optimal transport method Chi et al. (2021c), where the sentence representations are from the -th layer of XLM-E. Over the four language pairs, XLM-E achieves lower AER scores than the baseline models, reducing the average AER from to 19.32. It is worth mentioning that our model requires substantial lower computation costs than the other cross-lingual pretrained language models to achieve such low AER scores. See the detailed training efficiency analysis in Section 4.5. It is worth mentioning that XLM-E shows notable improvements over XLM-E (w/o TRTD) on both tasks, demonstrating that the translation replaced token detection task is effective for cross-lingual alignment.
7 Universal Layer Across Languages
We evaluate the word-level and sentence-level representations over different layers to explore whether the XLM-E tasks encourage universal representations.
As shown in Figure 3, we illustrate the accuracy@1 scores of XLM-E and XLM-R on Tatoeba cross-lingual sentence retrieval, using sentence representations from different layers. For each layer, the final accuracy score is averaged over all the 36 language pairs in both the xx en and en xx directions. From the figure, it can be observed that XLM-E achieves notably higher averaged accuracy scores than XLM-R for the top layers. The results of XLM-E also show a parabolic trend across layers, i.e., the accuracy continuously increases before a specific layer and then continuously drops. This trend is also found in other cross-lingual language models such as XLM-R and XLM-Align Jalili Sabet et al. (2020); Chi et al. (2021c). Different from XLM-R that achieves the highest accuracy of 54.42 at layer-7, XLM-E pushes it to layer-9, achieving an accuracy of 63.66. At layer-10, XLM-R only obtains an accuracy of 43.34 while XLM-E holds the accuracy score as high as 57.14.
Figure 4 shows the averaged alignment error rate (AER) scores of XLM-E and XLM-R on the word alignment task. We use the hidden vectors from different layers to perform word alignment, where layer-0 stands for the embedding layer. The final AER scores are averaged over the four test sets in different languages. Figure 4 shows a similar trend to that in Figure 3, where XLM-E not only provides substantial performance improvements over XLM-R, but also pushes the best-performance layer to a higher layer, i.e., the model obtains the best performance at layer-9 rather than a lower layer such as layer-7.
On both tasks, XLM-E shows good performance for the top layers, even though both XLM-E and XLM-R use the Transformer Vaswani et al. (2017) architecture. Compared to the masked language modeling task that encourages the top layers to be language-specific, discriminative pre-training makes XLM-E producing better-aligned text representations at the top layers. It indicates that the cross-lingual discriminative pre-training encourages universal representations inside the model.
8 Cross-lingual Transfer Gap
We analyze the cross-lingual transfer gap Hu et al. (2020b) of the pretrained cross-lingual language models. The transfer gap score is the difference between performance on the English test set and the average performance on the test set in other languages. This score suggests how much end task knowledge has not been transferred to other languages after fine-tuning. A lower gap score indicates better cross-lingual transferability. Table 7 compares the cross-lingual transfer gap scores on five of the XTREME tasks. We notice that XLM-E obtains the lowest gap score only on PAWS-X. Nonetheless, it still achieves reasonably low gap scores on the other tasks with such low computation cost, demonstrating the cross-lingual transferability of XLM-E. We believe that it is more difficult to achieve the same low gap scores when the model obtains better performance.
Related Work
Learning self-supervised tasks on large-scale multilingual texts has proven to be effective for pre-training cross-lingual language models. Masked language modeling (MLM; Devlin et al. 2019) is typically used to learn cross-lingual encoders such as multilingual BERT (mBERT; Devlin et al. 2019) and XLM-R Conneau et al. (2020). The cross-lingual language models can be further improved by introducing external pre-training tasks using parallel corpora. XLM Conneau and Lample (2019) introduces the translation language modeling (TLM) task that predicts masked tokens from concatenated translation pairs. ALM Yang et al. (2020) utilizes translation pairs to construct code-switched sequences as input. InfoXLM Chi et al. (2021b) considers an input translation pair as cross-lingual views of the same meaning, and proposes a cross-lingual contrastive learning task. Several pre-training tasks utilize the token-level alignments in parallel data to improve cross-lingual language models Cao et al. (2020); Zhao et al. (2021); Hu et al. (2020a); Chi et al. (2021c).
In addition, parallel data are also employed for cross-lingual sequence-to-sequence pre-training. XNLG Chi et al. (2020) presents cross-lingual masked language modeling and cross-lingual auto-encoding for cross-lingual natural language generation, and achieves the cross-lingual transfer for NLG tasks. VECO Luo et al. (2020) utilizes cross-attention MLM to pretrain a variable cross-lingual language model for both NLU and NLG. mT6 Chi et al. (2021a) improves mT5 Xue et al. (2021) by learning the translation span corruption task on parallel data. LM Ma et al. (2021) proposes to align pretrained multilingual encoders to improve cross-lingual sequence-to-sequence pre-training.
Conclusion
We introduce XLM-E, a cross-lingual language model pretrained by ELECTRA-style tasks. Specifically, we present two pre-training tasks, i.e., multilingual replaced token detection, and translation replaced token detection. XLM-E outperforms baseline models on cross-lingual understanding tasks although using much less computation cost. In addition to improved performance and computational efficiency, we also show that XLM-E obtains the cross-lingual transferability with a reasonably low transfer gap.
Ethical Considerations
Our work introduces ELECTRA-style tasks for cross-lingual language model pre-training, which requires much less computation cost than previous models and substantially reduces the energy cost.
References
Appendix
Appendix A Model Hyperparameters
Table 8 and Table 9 shows the model hyperparameters of XLM-E in the sizes of Base, Large, and XL. For the Base-size model, we use the same vocabulary with XLM-R Conneau et al. (2020) that consists of 250K subwords tokenized by SentencePiece Kudo and Richardson (2018). For the models in Large size and XL size, we use VoCap Zheng et al. (2021) to allocate a 500K vocabulary for models in Large size and XL size.
Appendix B Hyperparameters for Pre-Training
As shown in Table 10, we present the hyperparameters for pre-training XLM-E. We use the batch size of 1M tokens for each pre-training task. In multilingual replaced token detection, a batch is constructed by 2,048 length-512 input sequences, while the input length is dynamically set as the length of the original translation pairs in translation replaced token detection.
Appendix C Hyperparameters for Fine-Tuning
In Table 11, we report the hyperparameters for fine-tuning XLM-E on the XTREME end tasks.