When Being Unseen from mBERT is just the Beginning: Handling New Languages With Multilingual Language Models
Benjamin Muller, Antonis Anastasopoulos, Benoît Sagot, Djamé Seddah
Introduction
Language models are now a new standard to build state-of-the-art Natural Language Processing (NLP) systems. In the past year, monolingual language models have been released for more than 20 languages including Arabic, French, German, and Italian (Antoun et al., 2020; Martin et al., 2020; de Vries et al., 2019; Cañete et al., 2020; Kuratov and Arkhipov, 2019; Schweter, 2020, inter alia). Additionally, large-scale multilingual models covering more than 100 languages are now available (XLM-R by Conneau et al. (2020) and mBERT by Devlin et al. (2019)). Still, most of the 6500+ spoken languages in the world Hammarström (2016) are not covered—remaining unseen—by those models. Even languages with millions of native speakers like Sorani Kurdish (about 7 million speakers in the Middle East) or Bambara (spoken by around 5 million people in Mali and neighboring countries) are not covered by any available language models at the time of writing.
Even if training multilingual models that cover more languages and language varieties is tempting, the curse of multilinguality (Conneau et al., 2020) makes it an impractical solution, as it would require to train ever larger models. Furthermore, as shown by Wu and Dredze (2020), large-scale multilingual language models are sub-optimal for languages that are under-sampled during pretraining.
In this paper, we analyze task and language adaptation experiments to get usable language model-based representations for under-studied low resource languages. We run experiments on 15 typologically diverse languages on three NLP tasks: part-of-speech (POS) tagging, dependency parsing (DEP) and named-entity recognition (NER).
Our results bring forth a diverse set of behaviors that we classify in three categories reflecting the abilities of pretrained multilingual language models to be used for low-resource languages. We dub those categories Easy, Intermediate and Hard.
Hard languages include both stable and endangered languages, but they predominantly are languages of communities that are majorly under-served by modern NLP. Hence, we direct our attention to these Hard languages. For those languages, we show that the script they are written in can be a critical element in the transfer abilities of pretrained multilingual language models. Transliterating them leads to large gains in performance outperforming non-contextual strong baselines. To sum up, our contributions are the following:
We propose a new categorization of the low-resource languages that are unseen by available language models: the Hard, the Intermediate and the Easy languages.
We show that Hard languages can be better addressed by transliterating them into a better-handled script (typically Latin), providing a promising direction towards making multilingual language models useful for a new set of unseen languages.
Background and Motivation
As Joshi et al. (2020) vividly illustrate, there is a large divergence in the coverage of languages by NLP technologies. The majority of the 6500+ of the world’s languages are not studied by the NLP community, since most have few or no annotated datasets, making systems’ development challenging.
The development of such models is a matter of high importance for the inclusion of communities, the preservation of endangered languages and more generally to support the rise of tailored NLP ecosystems for such languages Schmidt and Wiegand (2017); Stecklow (2018); Seddah et al. (2020). In that regard, the advent of the Universal Dependencies project (Nivre et al., 2016) and the WikiAnn dataset (Pan et al., 2017) have greatly increased the number of covered languages by providing annotated datasets for more than 90 languages for dependency parsing and 282 languages for NER.
Regarding modeling approaches, the emergence of multilingual representation models, first with static word embeddings (Ammar et al., 2016) and then with language model-based contextual representations Devlin et al. (2019); Conneau et al. (2020) enabled transfer from high to low-resource languages, leading to significant improvements in downstream task performance (Rahimi et al., 2019; Kondratyuk and Straka, 2019). Furthermore, in their most recent forms, these multilingual models process tokens at the sub-word level (Kudo and Richardson, 2018). As such, they work in an open vocabulary setting, only constrained by the pretraining character set. This flexibility enables such models to process any language, even those that are not part of their pretraining data.
When it comes to low-resource languages, one direction is to simply train contextualized embedding models on whatever data is available. Another option is to adapt/fine-tune a multilingual pretrained model to the language of interest. We briefly discuss these two options.
Even though the amount of pretraining data seems to correlate with downstream task performance (e.g. compare BERT and RoBERTa (Liu et al., 2020)), several attempts have shown that training a model from scratch can be efficient even if the amount of data in that language is limited. Indeed, Ortiz Suárez et al. (2020) showed that pretraining ELMo models (Peters et al., 2018) on less than 1GB of text leads to state-of-the-art performance while Martin et al. (2020) showed that pretraining a BERT model on as few as 4GB of diverse enough data results in state-of-the-art performance. Micheli et al. (2020) further demonstrated that decent performance was achievable with only 100MB of raw text data.
Adapting large-scale models for low-resource languages
Multilingual language models can be used directly on unseen languages, or they can also be adapted using unsupervised methods. For example, Han and Eisenstein (2019) successfully used unsupervised model adaptation of the English BERT model to Early Modern English for sequence labeling. Instead of fine-tuning the whole model, Pfeiffer et al. (2020) recently showed that adapter layers (Houlsby et al., 2019) can be injected into multilingual language models to provide parameter efficient task and language transfer.
Still, as of today, the availability of monolingual or multilingual language models is limited to approximately 120 languages, leaving many languages without access to valuable NLP technology, although some are spoken by millions of people, including Bambara and Sorani Kurdish, or are an official language of the European Union, like Maltese.
What can be done for unseen languages?
Unseen languages strongly vary in the amount of available data, in their script (many languages use non-Latin scripts such as Sorani Kurdish and Mingrelian), and in their morphological or syntactical properties (most largely differ from high-resource Indo-European languages). This makes the design of a methodology to build contextualized models for such languages challenging at best. In this work, by experimenting with 15 typologically diverse unseen languages, (i) we show that there is a diversity of behavior depending on the script, the amount of available data, and the relation to the pretraining languages; (ii) Focusing on the unseen languages that lag in performance compared to their easier-to-handle counterparts, we show that the script plays a critical role in the transfer abilities of multilingual language models. Transliterating such languages to a script which is used by a related language seen during pretraining can lead to significant improvement in downstream performance.
Experimental Setting
We will refer to any languages that are not covered by pretrained language models as “unseen.” We select a small portion of those languages within a large scope of language families and scripts. Our selection is constrained to 15 typologically diverse languages for which we have evaluation data for at least one of our three downstream tasks. Our selection includes low-resource Indo-European and Uralic languages, as well as members of the Bantu, Semitic, and Turkic families. None of these 15 languages are included in the pretraining corpora of mBERT. Information about their scripts, language families, and amount of available raw data can be found in the Appendix in Table 12.
To perform pretraining and fine-tuning on monolingual data, we use the deduplicated datasets from the OSCAR project Ortiz Suárez et al. (2019). OSCAR is a corpus extracted from a Common Crawl Web snapshot.http://commoncrawl.org/ It provides a significant amount of data for all the unseen languages we work with, except for Buryat, Meadow Mari, Erzya and Livvi for which we use Wikipedia dumps and for Narabizi, Naija and Faroese, for which we use data collected by Seddah et al. (2020), Caron et al. (2019) and Biemann et al. (2007) respectively.
2 Non-contextual Baselines
For parsing and POS tagging, we use the UDPipe future system (Straka, 2018) as our baseline. This model is a LSTM-based (Hochreiter and Schmidhuber, 1997) recurrent architecture trained with pretrained static word embedding (Mikolov et al., 2013) (hence our non-contextual characterization) along with character-level embeddings. This system was ranked in the very first positions for parsing and tagging in the CoNLL shared task 2018 (Zeman and Hajič, 2018). For NER we use the LSTM-CRF model with character and word level embedding using Qi et al. (2020) implementation.
3 Language Models
In all our study, we train our language models using the Transformers library (Wolf et al., 2020).
The first approach we evaluate is to train a dedicated language model from scratch on the available raw data we have. To do so, we train a language-specific SentencePiece tokenizer (Kudo and Richardson, 2018) before training a Masked-Language Model (MLM) using the RoBERTa (base) architecture and objective functions Liu et al. (2019). As we work with significantly smaller pretraining sets than in the original setting, we reduce the number of layers to 6 layers in place of the original 12 layers.
Multilingual Language Models
We want to assess how large-scale multilingual language models can be used and adapted to languages that are not in their pretraining corpora. We work with the multilingual version of BERT (mBERT) trained on the concatenation of Wikipedia corpora in 104 languages (Devlin et al., 2019). We also ran experiments with the XLM-R base version (Conneau et al., 2020) trained on 100 languages using data extracted from the Web. As the observed behaviors are very similar between both models, we only report results using mBERT. Note that mBERT is highly biased toward Indo-Europeans languages written in the Latin script. More than 77% of the subword vocabulary are in the Latin script while only 1% are in the Georgian script Ács (2019).
Adapting Multilingual Language Models to unseen languages with MLM-tuning
Following previous work Han and Eisenstein (2019); Karthikeyan et al. (2019); Pfeiffer et al. (2020), we adapt large-scale multilingual models by fine-tuning them with their Mask-Language-Model objective directly on the available raw data in the unseen target language. We refer to this process as MLM-tuning. We will refer to a MLM-tuned mBERT model as mBERT+MLM.
4 Downstream Tasks
We perform experiments on POS tagging, Dependency Parsing (DEP), and Name Entity Recognition (NER). We use annotated data from the Universal Dependency project (Nivre et al., 2016) for POS tagging and parsing, and the WikiAnn dataset (Pan et al., 2017) for NER. For POS tagging and NER, we append a linear classifier layer on top of the language model. For parsing, following Kondratyuk and Straka (2019), we append a Bi-Affine Graph prediction layer (Dozat and Manning, 2017). We refer to the process of fine-tuning a language model in a task-specific way as Task-tuning.Details about optimization can be found in Appendix B
5 Dataset Splits
For each task and language, we use the provided training, validation and test dataset split except for the ones that have less than 500 training sentences. In this case, we concatenate the training and test set and perform 8-folds cross-Validation and use the validation set for early stopping. If no validation set is available, we isolate one of the folds for validation and report the test scores as the average of the other folds. This enables us to train on at least 500 sentences in all our experiments (except for Swiss German for which we only have 100 training examples) and reduce the impact of the annotated dataset size on our analysis. Since cross-validation results in training on very limited number of examples, we refer to training in this cross-validation setting as few-shot learning.
The Three Categories of Unseen Languages
For each unseen language and each task, we experiment with our three modeling approaches: (a) Training a language model from scratch on the available raw data and then fine-tuning it on any available annotated data in the target language. (b) Fine-tuning mBERT with Task-tuning directly on the target language. (c) Finally, adapting mBERT to the unseen language using MLM-tuning before fine-tuning it in a supervised way on the target language. We then compare all these experiments to our non-contextual strong baselines. By doing so, we can assess if language models are a practical solution to handle each of these unseen languages.
Interestingly we find a large diversity of behaviors across languages regarding those language model training techniques. As summarized in Figure 1, we observe three clear clusters of languages.
The first cluster, which we dub “Easy", corresponds to the languages that do not require extra MLM-tuning for mBERT to achieve good performance. mBERT has the modeling abilities to process such languages without relying on raw data and can outperform strong non-contextual baselines as such. In the second cluster, the “Intermediate" languages require MLM-tuning. mBERT is not able to beat strong non-contextual baselines using only Task-tuning, but MLM-tuning enables it to do so. Finally, Hard languages are those on which mBERT fails to deliver any decent performance even after MLM- and Task- fine-tuning. mBERT simply does not have the capacity to learn and process such languages.
We emphasize that our categorization of unseen languages is only based on the relative performance of mBERT after fine-tuning compared to strong non-contextual baseline models. We leave for future work the analysis of the absolute performance of the model on such languages (e.g. analysing the impact of the fine-tuning data set size on mBERT’s downstream performance).
In this section, we present our results in detail in each of these language clusters and provide insights into their linguistic properties.
Easy languages are the ones on which mBERT delivers high performance out-of-the-box, compared to strong baselines. We classify Faroese, Swiss German, Naija and Mingrelian as easy languages and report performance in Table 1.
We find that those languages match two conditions:
They are closely related to languages used during MLM pretraining
These languages use the same script as such closely related languages.
Such languages benefit from multilingual models, as cross-lingual transfer is easy to achieve and hence quite effective.
More details about those languages can be found in Appendix C.
2 Intermediate
The second type of languages (which we dub “Intermediate”) are generally harder to process for pretrained MLMs out-of-the-box. In particular, pretrained multilingual language models are typically outperformed by a non-contextual strong baselines. Still, MLM-tuning has an important impact and leads to usable state-of-the-art models.
A good example of such an intermediate language is Maltese, a member of the Semitic language but using the Latin script. Maltese has not been seen by mBERT during pretraining. Other Semitic languages though, namely Arabic and Hebrew, have been included in the pretraining languages. As seen in Table 2, the non-contextual baseline outperforms mBERT. Additionally, a monolingual MLM trained on only 50K sentences matches mBERT performance for both NER and POS tagging. However, the best results are reached with MLM-tuning: the proper use of monolingual data and the advantage of similarity to other pretraining languages render Maltese a tackle-able language as shown by the performance gain over our strong non-contextual baselines.
Our Maltese dependency parsing results are in line with those of Chau et al. (2020), who also showed that MLM-tuning leads to significant improvements. They also additionally showed that a small vocabulary transformation allowed fine-tuning to be even more effective and gain 0.8 LAS points more. We further discuss the vocabulary adaptation technique of Chau et al. (2020) in section 6.
We consider Narabizi (Seddah et al., 2020), an Arabic dialect spoken in North-Africa written in the Latin script and code-mixed with French, to fall in the same Intermediate category, because it follows the same pattern. For both POS tagging and parsing, the multilingual models outperform the monolingual NarabiziBERT. In addition, MLM-tuning leads to significant improvements over the non-language-tuned mBERT baseline, also outperforming the non-contextual dependency parsing baseline.
We also categorize Bambara, a Niger-Congo Bantu language spoken in Mali and surrounding countries, as Intermediate, relying mostly on the POS tagging results which follow similar patterns as Maltese and Narabizi. We note that the BambaraBERT that we trained achieves notably poor performance compared to the non-contextual baseline, a fact we attribute to the extremely low amount of available data (1000 sentences only). We also note that the non-contextual baseline is the best performing model for dependency parsing, which could also potentially classify Bambara as a “Hard" language instead.
Our results in Wolof follow the same pattern. The non-contextual baseline achieves a 77.0 in LAS outperforming mBERT. However, MLM-tuning achieves the highest score of 77.9.
We now turn our focus to Uralic languages. Finnish, Estonian, and Hungarian are high-resource representatives of this language family that are typically included in multilingual LMs, also having task-tuning data available in large quantities. However, for several smaller Uralic languages, task-tuning data are generally very scarce.
We report in Table 2 the performance for two low-resource Uralic languages, namely Livvi and Erzya using 8-fold cross-validation, with each run only using around 700 training instances. Note the striking difference between the parsing performance (LAS) of mBERT on Livvi, written with the Latin script, and on Erzya that uses the Cyrillic script. This suggests that the script could be playing a critical role when transferring to those languages. We explore this hypothesis in detail in section 5.2.
3 Hard
The last category of the hard unseen language is perhaps the most interesting one, as these languages are very hard to process. mBERT is outperformed by non-contextual baselines as well as by monolingual language models trained from scratch on the available raw data. At the same time, MLM-tuning on the available raw data has a minimal impact on performance.
Uyghur, a Turkic language with about 10-15 million speakers in central Asia, is a prime example of a hard language for current models. In our experiments, outlined in Table 3, the non-contextual baseline outperforms all contextual variants, both monolingual and multilingual, in all the tasks with up to 20 points difference compared to mBERT for parsing. Additionally, the monolingual UyghurBERT trained on only 105K sentences outperforms mBERT even after MLM-tuning.
We attribute this discrepancy to script differences: Uyghur uses the Perso-Arabic script, when the other Turkic languages that were part of mBERT pretraining use either the Latin (e.g. Turkish) or the Cyrillic script (e.g. Kazakh).
Sorani Kurdish (also known as Central Kurdish) is a similarly hard language, mainly spoken in Iraqi Kurdistan by around 8 million speakers, which uses the Sorani alphabet, a variant of the Arabic script. We can solely evaluate on the NER task, where the non-contextual baseline and the monolingual SoraniBERT perform similarly around 80.5 F1-score outperforming significantly mBERT which only reaches 70.4 in F1-score. MLM-tuning on 380K sentences of Sorani texts improves mBERT performance to 75.6 F1-score, but it is still lagging behind the baseline. Our results in Sindhi follow the same pattern. The non-contextual baseline achieves a 51.4 F1-score outperforming with a large margin our language models (a monolingual SindhiBERT achieves an F1-score of 45.2, and mBERT is worse at 42.3).
Tackling Hard Languages with Multilingual Language Models
Our intermediate Uralic language results provide initial supporting evidence for our argument on the importance of having pretrained LMs on languages with similar scripts, even for generally high-resource language families. Our hypothesis is that the script is a key element for language models to correctly process unseen languages.
To test this hypothesis, we assess the ability of mBERT to process an unseen language after transliterating it to another script present in the pretraining data. We experiment on six languages belonging to four language families: Erzya, Bruyat and Meadow Mari (Uralic), Sorani Kurdish (Iranian, Indo-European), Uyghur (Turkic) and Mingrelian (Kartvelian). We apply the following transliterations:
Erzya/Buryat/Mari: Cyrillic Latin Script
Uyghur: Arabic Script Latin Script
Sorani: Arabic Script Latin Script
Mingrelian: Georgian Script Latin Script
The strategy we used to transliterate the above-listed language is specific to the purpose of our experiments. Indeed, our goal is for the model to take advantage of the information it has learned during training on a related language written in the Latin script. The goal of our transliteration is therefore to transcribe each character in the source script, which we assume corresponds to a phoneme, into the most frequent (sometimes only) way this phoneme is rendered in the closest related language written in the Latin script, hereafter the target language. This process is not a transliteration strictly speaking, and it needs not be reversible. It is not a phonetization either, but rather a way to render the source language in a way that maximizes the similarity between the transliterated source language and the target language.
We have manually developed transliteration scripts for Uyghur and Sorani Kurdish, using respectively Turkish and Kurmanji Kurdish as target languages, only Turkish being one of the languages used to train mBERT. Note however that Turkish and Kurmanji Kurdish share a number of conventions for rendering phonemes in the Latin script (for instance, /\textesh/, rendered in English by “sh”, is rendered in both languages by “ş”; as a result, the Arabic letter “”, used in both languages, is rendered as “ş” by both our transliteration scripts). As for Erzya, Buryat and Mari, we used the readily available transliteration package transliterate,https://pypi.org/project/transliterate/ which performs a standard transliteration.In future work, we intend to develop dedicated transliteration scripts using the strategy described above, and to compare the results obtained with it with those described here. We used the Russian transliteration module, as it covers the Cyrillic script. Finally, for our control experiments on Mingrelian, we used the Georgian transliteration module from the same package.
2 Transfer via Transliteration
We train mBERT with MLM-tuning and Task-tuning as well as monolingual BERT model trained from scratch on the transliterated data. We also run controlled experiments on high-resource languages written in the Latin script on which mBERT was pretrained on, namely Arabic, Japanese and Russian (reported in Table 5).
Our results with and without transliteration are listed in Table 4. Transliteration for Sorani and Uyghur has a noticeable positive impact. For instance, transliterating Uyghur to Latin leads to an improvement of 16 points in parsing and 20 points in NER. For one of the low-resource Uralic languages, Meadow Mari, we observe an 8 F1-score points improvement on NER, while for other Uralic languages like Erzya the effect of transliteration is very minor. The only case where transliteration to the Latin script leads to a drop in performance for mBERT and mBERT+MLM is Mingrelian.
We interpret our results as follows. When running MLM-tuning and Task-tuning, mBERT associates the target unseen language to a set of similar languages seen during pretraining based on the script. In consequence, mBERT is not able to associate a language to its related language if they are not written in the same script. For instance, transliterating Uyghur enables mBERT to match it to Turkish, a language which accounts for a sizable portion of mBERT pretraining. In the case of Mingrelian, transliteration has the opposite effect: transliterating Mingrelian in the Latin script is harming the performance as mBERT is not able to associate it to Georgian which is seen during pretraining and uses the Georgian script.
This is further supported by our experiments on high resource languages (cf. Table 5). When transliterating pretrained languages such as Arabic, Russian or Japanese, mBERT is not able to compete with the performance reached when using the script seen during pretraining. Transliterating the Arabic script and the Cyrillic script to Latin does not automatically improve mBERT performance as it does for Sorani, Uyghur and Meadow Mari. For instance, transliterating Arabic to the Latin script leads to a drop in performance of 1.5, 4.1 and 6.9 points for POS tagging, parsing and NER respectively.Details and complete results on these controlled experiments can be found in Appendix E.
Our findings are generally in line with previous work. Transliteration to English specifically Lin et al. (2016); Durrani et al. (2014) and named entity transliteration Kundu et al. (2018); Grundkiewicz and Heafield (2018) has been proven useful for cross-lingual transfer in tasks like NER, entity linking Rijhwani et al. (2019), morphological inflection Murikinati et al. (2020), and Machine Translation Amrhein and Sennrich (2020).
The transliteration approach provides a viable path for rendering large pretrained models like mBERT useful for all languages of the world. Indeed, as reported in Table 4, transliterating both Uyghur and Sorani leads to matching or outperforming the performance of non-contextual strong baselines and deliver usable models (e.g. +12.5 POS accuracy in Uyghur).
Discussion and Conclusion
Pretraining ever larger language models is a research direction that is currently receiving a lot of attention and resources from the NLP research community Raffel et al. (2019); Brown et al. (2020). Still, a large majority of human languages are under-resourced making the development of monolingual language models very challenging in those settings. Another path is to build large scale multilingual language models.Even though we explore a different research direction, recent advances in small scale and domain specific language models suggest such models could also have an important impact for those languages Micheli et al. (2020). However, such an approach faces the inherent zipfian structure of human languages, making the training of a single model to cover all languages an unfeasible solution (Conneau et al., 2020). Reusing large scale pretrained language models for new unseen languages seems to be a more promising and reasonable solution from a cost-efficiency and environmental perspective (Strubell et al., 2019).
Recently, Pfeiffer et al. (2020) proposed to use adapter layers (Houlsby et al., 2019) to build parameter efficient multilingual language models for unseen languages. However, this solution brings no significant improvement in the supervised setting, compared to a more simple Masked-Language Model finetuning. Furthermore, developing a language agnostic adaptation method is an unreasonable wish with regard to the large typological diversity of human languages.
On the other hand, the promising vocabulary adaptation technique of Chau et al. (2020) which leads to good dependency parsing results on unseen languages when combined with task-tuning has so far been tested only on Latin script languages (Singlish and Maltese). We expect that it will be orthogonal to our transliteration approach, but we leave for future work the study of its applicability and efficacy on more languages and tasks.
In this context, we bring empirical evidence to assess the efficiency of language models pretraining and adaptation methods on 15 low-resource and typologically diverse unseen languages. Our results show that the “Hard" languages are currently out-of-the-scope of any currently available language models and are therefore left outside of the current NLP progress. By focusing on those, we find that this challenge is mostly due to the script. Transliterating them to a script that is used by a related higher resource language on which the language model has been pretrained on leads to large improvements in downstream performance. Our results shed some new light on the importance of the script in multilingual pretrained models. While previous work suggests that multilingual language models could transfer efficiently across scripts in zero-shot settings (Pires et al., 2019; Karthikeyan et al., 2019), our results show that such cross-script transfer is possible only if the model has seen related languages in the same script during pretraining.
Our work paves the way for a better understanding of the mechanics at play in cross-language transfer learning in low-resource scenarios. We strongly believe that our method can contribute to bootstrapping NLP resources and tools for low-resource languages, thereby favoring the emergence of NLP ecosystems for languages currently under-served by the NLP community.
Acknowledgments
The Inria authors were partly funded by two French Research National agency projects, namely projects PARSITI (ANR-16-CE33-0021) and SoSweet (ANR-15-CE38-0011), as well as by Benoit Sagot’s chair in the PRAIRIE institute as part of the “Investissements d’avenir” programme under the reference ANR-19-P3IA-0001. Antonios Anastasopoulos is generously supported by NSF Award 2040926 and is also thankful to Graham Neubig for very insightful initial discussions on this research direction.
References
Appendix A Languages
We list the 15 typologically diverse unseen languages we experiment with in Table 12 with information on their language family, script, origin and number of sentences available along with the categories we classified them in.
We base our experiments on data originated from two sources. The Universal Dependency project Nivre et al. (2016) downloadable here https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-2988 and the WikiNER dataset Pan et al. (2017). We also use of the CoNLL-2003 shared task NER English dataset https://www.clips.uantwerpen.be/conll2003/
Appendix B Reproducibility
Our experiments were ran on a shared cluster on the equivalent of 15 Nvidia Tesla T4 GPUs.https://www.nvidia.com/en-sg/data-center/tesla-t4/
Optimization
For all pretraining and fine-tuning runs, we use the Adam optimizer (Kingma and Ba, 2015). For fine-tuning, following Devlin et al. (2019), we only back-propagate through the first token of each word. We select the hyperparameters that minimize the loss on the validation set. The reported results are the average score of 5 runs with different random seeds computed on the test splits. We report the hyperparameters in Table 6-7.
Appendix C Easy Languages
We describe here in more details languages that we classify as Easy in section 4.1.
In practice, one can obtain very high performance even in zero-shot settings for such languages, by performing task-tuning on related languages.
Perhaps the best example of such an “easy” setting is Faroese. mBERT has been trained on several languages of the north Germanic genus of the Indo-European language family, all of which use the Latin script. As a result, the multilingual mBERT model performs much better than the monolingual FaroeseBERT model that we trained on the available Faroese text (cf rows 1–2 and 5–6 in Table 8). Fine-tuning mBERT on the Faroese text is even more effective (rows 3 and 6 in Table 8), leading to further improvements, reaching more than 96.5% POS-tagging accuracy, 86% LAS for dependency parsing, and 58% NER F1 in the few-shot setting, surpassing the non-contextual baseline. In fact, even in zero-shot conditions, where we task-tune only on related languages (Danish, Norwegian, and Swedish), the model achieves remarkable performance of over 83% POS-tagging accuracy and 67.8% LAS dependency parsing.
Swiss German is another example of a language for which one can easily adapt a multilingual model and obtain good performance even in zero-shot settings. As in Faroese, simple MLM fine-tuning of the mBERT model with 200K sentences leads to an improvement of more than 25 points in both POS tagging and dependency parsing (Table 9) in zero-shot settings, with similar improvement trends in the few-shot setting.
The potential of similar-language pretraining along with script similarity is also showcased in the case of Naija (also known as Nigerian English or Nigerian Pidgin), an English creole spoken by millions in Nigeria. As Table 10 shows, with results after language- and task-tuning on 6K training examples, the multilingual approach surpasses the monolingual baseline.
On a side note, we can rely on Han and Eisenstein (2019) to also classify Early Modern English as an easy language. Similarly, the work of Chau et al. (2020) allows us to also classify Singlish (Singaporean English) as an easy language. In both cases, these languages are technically unseen by mBERT, but the fact that they are variants of English allows them to be easily handled by mBERT.
Appendix D Additional Uralic languages experiments
Following a similar procedure as in the Appendix C, we start with mBERT, perform task-tuning on Finnish and Estonian (both of which use the Latin script) and then do zero-shot experiments on Livvi, and Komi, all low-resource Uralic languages (results on the top part of Table 11). We also report results on the Finnish treebanks after task-tuning, for better comparison. The difference in performance on Livvi (which uses the Latin script) and the other languages that use the Cyrillic script is striking.
Although they are not easy enough to be tackled in a zero-shot setting, we show that the low-resource Uralic languages fall in the “Intermediate” category, since mBERT has been trained on similar languages: a small amount of annotated data are enough to improve over mBERT using task-tuning.
For both Livvi and Erzya, the multilingual model along with MLM-tuning achieves the best performance, outperforming the non-contextual baseline by more than 1.5 point for parsing and POS tagging.
Appendix E Controlled experiment: Transliterating High-Resource Languages
To have a broader view on the effect of transliteration when using mBERT (section 5.2), we study the impact of transliteration to the Latin script on high resource languages seen during mBERT pretraining such as Arabic, Japanese and Russian. We compare the performance of mBERT fine-tuned and evaluated on the original script with mBERT fine-tuned and evaluated on the transliterated text. As reported in Table 5, transliterating those languages to the Latin script leads to large drop in performance for all the three tasks.