Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT

Shijie Wu, Mark Dredze

Introduction

Pretrained language representations with self-supervised objectives have become standard in a variety of NLP tasks Peters et al. (2018); Howard and Ruder (2018); Radford et al. (2018); Devlin et al. (2019), including sentence-level classification Wang et al. (2018), sequence tagging (e.g. NER) Tjong Kim Sang and De Meulder (2003) and SQuAD question answering Rajpurkar et al. (2016). Self-supervised objectives include language modeling, the cloze task Taylor (1953) and next sentence classification. These objectives continue key ideas in word embedding objectives like CBOW and skip-gram Mikolov et al. (2013a).

At the same time, cross-lingual embedding models have reduced the amount of cross-lingual supervision required to produce reasonable models; Conneau et al. (2017); Artetxe et al. (2018) use identical strings between languages as a pseudo bilingual dictionary to learn a mapping between monolingual-trained embeddings. Can jointly training contextual embedding models over multiple languages without explicit mappings produce an effective cross-lingual representation? Surprisingly, the answer is (partially) yes. BERT, a recently introduced pretrained model Devlin et al. (2019), offers a multilingual model (mBERT) pretrained on concatenated Wikipedia data for 104 languages without any cross-lingual alignment Devlin (2018). mBERT does surprisingly well compared to cross-lingual word embeddings on zero-shot cross-lingual transfer in XNLI Conneau et al. (2018), a natural language inference dataset. Zero-shot cross-lingual transfer, also known as single-source transfer, refers trains and selects a model in a source language, often a high resource language, then transfers directly to a target language.

While XNLI results are promising, the question remains: does mBERT learn a cross-lingual space that supports zero-shot transfer? We evaluate mBERT as a zero-shot cross-lingual transfer model on five different NLP tasks: natural language inference, document classification, named entity recognition, part-of-speech tagging, and dependency parsing. We show that it achieves competitive or even state-of-the-art performance with the recommended fine-tune all parameters scheme Devlin et al. (2019). Additionally, we explore different fine-tuning and feature extraction schemes and demonstrate that with parameter freezing, we further outperform the suggested fine-tune all approach. Furthermore, we explore the extent to which mBERT generalizes away from a specific language by measuring accuracy on language ID using each layer of mBERT. Finally, we show how subword tokenization influences transfer by measuring subword overlap between languages.

Background

Cross-lingual transfer learning is a type of transductive transfer learning with different source and target domain (Pan and Yang, 2010). A cross-lingual representation space is assumed to perform the cross-lingual transfer. Before the widespread use of cross-lingual word embeddings, task-specific models assumed coarse-grain representation like part-of-speech tags, in support of a delexicalized parser Zeman and Resnik (2008). More recently cross-lingual word embeddings have been used in conjunction with task-specific neural architectures for tasks like named entity recognition Xie et al. (2018), part-of-speech tagging Kim et al. (2017) and dependency parsing Ahmad et al. (2019).

Cross-lingual Word Embeddings.

The quality of the cross-lingual space is essential for zero-shot cross-lingual transfer. Ruder et al. (2017) surveys methods for learning cross-lingual word embeddings by either joint training or post-training mappings of monolingual embeddings. Conneau et al. (2017) and Artetxe et al. (2018) first show two monolingual embeddings can be aligned by learning an orthogonal mapping with only identical strings as an initial heuristic bilingual dictionary.

Contextual Word Embeddings

ELMo Peters et al. (2018), a deep LSTM Hochreiter and Schmidhuber (1997) pretrained with a language modeling objective, learns contextual word embeddings. This contextualized representation outperforms stand-alone word embeddings, e.g. Word2Vec Mikolov et al. (2013b) and Glove Pennington et al. (2014), with the same task-specific architecture in various downstream tasks. Instead of taking the representation from a pretrained model, GPT Radford et al. (2018) and Howard and Ruder (2018) also fine-tune all the parameters of the pretrained model for a specific task. Also, GPT uses a transformer encoder Vaswani et al. (2017) instead of an LSTM and jointly fine-tunes with the language modeling objective. Howard and Ruder (2018) propose another fine-tuning strategy by using a different learning rate for each layer with learning rate warmup and gradual unfreezing.

Concurrent work by Lample and Conneau (2019) incorporates bitext into BERT by training on pairs of parallel sentences. Schuster et al. (2019) aligns pretrained ELMo of different languages by learning an orthogonal mapping and shows strong zero-shot and few-shot cross-lingual transfer performance on dependency parsing with 5 Indo-European languages. Similar to multilingual BERT, Mulcaire et al. (2019) trains a single ELMo on distantly related languages and shows mixed results as to the benefit of pretaining.

Parallel to our work, Pires et al. (2019) shows mBERT has good zero-shot cross-lingual transfer performance on NER and POS tagging. They show how subword overlap and word ordering effect mBERT transfer performance. Additionally, they show mBERT can find translation pairs and works on code-switched POS tagging. In comparison, our work looks at a larger set of NLP tasks including dependency parsing and ground the mBERT performance against previous state-of-the-art on zero-shot cross-lingual transfer. We also probe mBERT in different ways and show a more complete picture of the cross-lingual effectiveness of mBERT.

Multilingual BERT

Devlin et al. (2019) is a deep contextual representation based on a series of transformers trained by a self-supervised objective. One of the main differences between BERT and related work like ELMo and GPT is that BERT is trained by the Cloze task Taylor (1953), also referred to as masked language modeling, instead of right-to-left or left-to-right language modeling. This allows the model to freely encode information from both directions in each layer. Additionally, BERT also optimizes a next sentence classification objective. At training time, 50% of the paired sentences are consecutive sentences while the rest of the sentences are paired randomly. Instead of operating on words, BERT uses a subword vocabulary with WordPiece Wu et al. (2016), a data-driven approach to break up a word into subwords.

Fine-tuning BERT

BERT shows strong performance by fine-tuning the transformer encoder followed by a softmax classification layer on various sentence classification tasks. A sequence of shared softmax classifications produces sequence tagging models for tasks like NER. Fine-tuning usually takes 3 to 4 epochs with a relatively small learning rate, for example, 3e-5.

Multilingual BERT

mBERT Devlin (2018) follows the same model architecture and training procedure as BERT, except with data from Wikipedia in 104 languages. Training makes no use of explicit cross-lingual signal, e.g. pairs of words, sentences or documents linked across languages. In mBERT, the WordPiece modeling strategy allows the model to share embeddings across languages. For example, “DNA” has a similar meaning even in distantly related languages like English and Chinese “DNA” indeed appears in the vocabulary of mBERT as a stand-alone lexicon.. To account for varying sizes of Wikipedia training data in different languages, training uses a heuristic to subsample or oversample words when running WordPiece as well as sampling a training batch, random words for cloze and random sentences for next sentence classification.

Transformer

For completeness, we describe the Transformer used by BERT. Let x\mathbf{x}, y\mathbf{y} be a sequence of subwords from a sentence pair. A special token [CLS] is prepended to x\mathbf{x} and [SEP] is appended to both x\mathbf{x} and y\mathbf{y}. The embedding is obtained by

where EE is the embedding function and LN is layer normalization Ba et al. (2016). MM transformer blocks are followed by the embeddings. In each transformer block,

In each attention, referred to as attention head,

Tasks

Does mBERT learn a cross-lingual representation, or does it produce a representation for each language in its own embedding space? We consider five tasks in the zero-shot transfer setting. We assume labeled training data for each task in English, and transfer the trained model to a target language. We select a range of different tasks: document classification, natural language inference, named entity recognition, part-of-speech tagging, and dependency parsing. We cover zero-shot transfer from English to 38 languages in the 5 different tasks as shown in Tab. 1. In this section, we describe the tasks as well as task-specific layers.

2 Natural Language Inference

We use XNLI Conneau et al. (2018) which cover 15 languages for natural language inference. The 3-way classification includes entailment, neutral, and contradiction given a pair of sentences. We feed a pair of sentences directly into mBERT and the task-specific classification layer is the same as § 4.1. We evaluate by classification accuracy.

3 Named Entity Recognition

We use the CoNLL 2002 and 2003 NER shared tasks Tjong Kim Sang (2002); Tjong Kim Sang and De Meulder (2003) (4 languages) and a Chinese NER dataset Levow (2006). The labeling scheme is BIO with 4 types of named entities. We add a linear classification layer with softmax to obtain word-level predictions. Since mBERT operates at the subword-level while the labeling is word-level, if a word is broken into multiple subwords, we mask the prediction of non-first subwords. NER is evaluated by F1 of predicted entity (F1). Note we use a simple post-processing heuristic to obtain a valid span.

4 Part-of-Speech Tagging

We use a subset of Universal Dependencies (UD) Treebanks (v1.4) Nivre et al. (2016), which cover 15 languages, following the setup of Kim et al. (2017). The task-specific labeling layer is the same as § 4.3. POS tagging is evaluated by the accuracy of predicted POS tags (ACC).

5 Dependency parsing

Following the setup of Ahmad et al. (2019), we use a subset of Universal Dependencies (UD) Treebanks (v2.2) Nivre et al. (2018), which includes 31 languages. Dependency parsing is evaluated by unlabelled attachment score (UAS) and labeled attachment score (LAS) Punctuations (PUNCT) and symbols (SYM) are excluded.. We only predict the coarse-grain dependency label following Ahmad et al. We use the model of Dozat and Manning (2016), a graph-based parser as a task-specific layer. Their LSTM encoder is replaced by mBERT. Similar to § 4.3, we only take the representation of the first subword of each word. We use masking to prevent the parser from operating on non-first subwords.

Experiments

We use the base cased multilingual BERT, which has N=12N=12 attention heads and M=12M=12 transformer blocks. The dropout probability is 0.1 and dhd_{h} is 768. The model has 179M parameters with about 120k vocabulary.

For each task, no preprocessing is performed except tokenization of words into subwords with WordPiece. We use Adam Kingma and Ba (2014) for fine-tuning with β1\beta_{1} of 0.9, β2\beta_{2} of 0.999 and L2 weight decay of 0.01. We warm up the learning rate over the first 10% of batches and linearly decay the learning rate.

Maximum Subwords Sequence Length

At training time, we limit the length of subwords sequence to 128 to fit in a single GPU for all tasks. For NER and POS tagging, we additionally use the sliding window approach. After the first window, we keep the last 64 subwords from the previous window as context. In other words, for a non-first window, only (up to) 64 new subwords are added for prediction. At evaluation time, we follow the same approach as training time except for parsing. We threshold the sentence length to 140 words, including words and punctuation, following Ahmad et al. (2019). In practice, the maximum subwords sequence length is the number of subwords of the first 140 words or 512, whichever is smaller.

Hyperparameter Search and Model Selection

We select the best hyperparameters by searching a combination of batch size, learning rate and the number of fine-tuning epochs with the following range: learning rate {2×10−5,3×10−5,5×10−5}\{2\times 10^{-5},3\times 10^{-5},5\times 10^{-5}\}; batch size {16,32}\{16,32\}; number of epochs: {3,4}\{3,4\}. Note the best hyperparameters and model are selected by development performance in English.

1 Question #1: Is mBERT Multilingual?

We include two strong baselines. Schwenk and Li (2018) use MultiCCA, multilingual word embeddings trained with a bilingual dictionary Ammar et al. (2016), and convolution neural networks. Concurrent to our work, Artetxe and Schwenk (2018) use bitext between English/Spanish and the rest of languages to pretrain a multilingual sentence representation with a sequence-to-sequence model where the decoder only has access to a max-pooling of the encoder hidden states.

mBERT outperforms (Tab. 2) multilingual word embeddings and performs comparably with a multilingual sentence representation, even though mBERT does not have access to bitext. Interestingly, mBERT outperforms Artetxe and Schwenk (2018) in distantly related languages like Chinese and Russian and under-performs in closely related Indo-European languages.

XNLI

We include three strong baselines, Artetxe and Schwenk (2018) and Lample and Conneau (2019) are concurrent to our work. Lample and Conneau (2019) with MLM is similar to mBERT; the main difference is that it only trains with the 15 languages of XNLI, has 249M parameters (around 40% more than mBERT), and MLM+TLM also uses bitext as training data They also use language embeddings as input and exclude the next sentence classification objective. Conneau et al. (2018) use supervised multilingual word embeddings with an LSTM encoder and max-pooling. After an English encoder and classifier are trained, the target encoder is trained to mimic the English encoder with ranking loss and bitext.

In Tab. 3, mBERT outperforms one model with bitext training but (as expected) falls short of models with more cross-lingual training information. Interestingly, mBERT and MLM are mostly the same except for the training languages, yet we observe that mBERT under-performs MLM by a large margin. We hypothesize that limiting pretraining to only those languages needed for the downstream task is beneficial. The gap between Artetxe and Schwenk (2018) and mBERT in XNLI is larger than MLDoc, likely because XNLI is harder.

NER

We use Xie et al. (2018) as a zero-shot cross-lingual transfer baseline, which is state-of-the-art on CoNLL 2002 and 2003. It uses unsupervised bilingual word embeddings Conneau et al. (2017) with a hybrid of a character-level/word-level LSTM, self-attention, and a CRF. Pseudo training data is built by word-to-word translation with an induced dictionary from bilingual word embeddings.

mBERT outperforms a strong baseline by an average of 6.9 points absolute F1 and an 11.8 point absolute improvement in German with a simple one layer 0th{}^{\text{th}}-order CRF as a prediction function (Tab. 4). A large gap remains when transferring to distantly related languages (e.g. Chinese) compared to a supervised baseline. Further effort should focus on transferring between distantly related languages. In § 5.4 we show that sharing subwords across languages helps transfer.

POS

We use Kim et al. (2017) as a reference. They utilized a small amount of supervision in the target language as well as English supervision so the results are not directly comparable. Tab. 5 shows a large (average) gap between mBERT and Kim et al. Interestingly, mBERT still outperforms Kim et al. (2017) with 320 sentences in German (de), Polish (pl), Slovak (sk) and Swedish (sv).

Dependency Parsing

We use the best performing model on average in Ahmad et al. (2019) as a zero-shot transfer baseline, i.e. transformer encoder with graph-based parser Dozat and Manning (2016), and dictionary supervised cross-lingual embeddings Smith et al. (2017). Dependency parsers, including Ahmad et al., assume access to gold POS tags: a cross-lingual representation. We consider two versions of mBERT: with and without gold POS tags. When tags are available, a tag embedding is concatenated with the final output of mBERT.

Tab. 6 shows that mBERT outperforms the baseline on average by 7.3 point UAS and 0.4 point LAS absolute improvement even without gold POS tags. Note in practice, gold POS tags are not always available, especially for low resource languages. Interestingly, the LAS of mBERT tends to weaker than the baseline in languages with less word order distance, in other words, more closely related to English. With the help of gold POS tags, we further observe 1.6 points UAS and 4.7 point LAS absolute improvement on average. It appears that adding gold POS tags, which provide clearer cross-lingual representations, benefit mBERT.

Summary

Across all five tasks, mBERT demonstrate strong (sometimes state-of-the-art) zero-shot cross-lingual performance without any cross-lingual signal. It outperforms cross-lingual embeddings in four tasks. With a small amount of target language supervision and cross-lingual signal, mBERT may improve further; we leave this as future work. In short, mBERT is a surprisingly effective cross-lingual model for many NLP tasks.

2 Question #2: Does mBERT vary layer-wise?

The goal of a deep neural network is to abstract to higher-order representations as you progress up the hierarchy Yosinski et al. (2014). Peters et al. (2018) empirically show that for ELMo in English the lower layer is better at syntax while the upper layer is better at semantics. However, it is unclear how different layers affect the quality of cross-lingual representation. For mBERT, we hypothesize a similar generalization across the 13 layers, as well as an abstraction away from a specific language with higher layers. Does the zero-shot transfer performance vary with different layers?

We consider two schemes. First, we follow the feature-based approach of ELMo by taking a learned weighted combination of all 13 layers of mBERT with a two-layer bidirectional LSTM with dhd_{h} hidden size (Feat). Note the LSTM is trained from scratch and mBERT is fixed. For sentence and document classification, an additional max-pooling is used to extract a fixed-dimension vector. We train the feature-based approach with Adam and learning rate 1e-3. The batch size is 32. The learning rate is halved whenever the development evaluation does not improve. The training is stopped early when learning rate drop below 1e-5. Second, when fine-tuning mBERT, we fix the bottom nn layers (nn included) of mBERT, where layer 0 is the input embedding. We consider n∈{0,3,6,9}n\in\{0,3,6,9\}.

Freezing the bottom layers of mBERT, in general, improves the performance of mBERT in all five tasks (Fig. 1). For sentence-level tasks like document classification and natural language inference, we observe the largest improvement with n=6n=6. For word-level tasks like NER, POS tagging, and parsing, we observe the largest improvement with n=3n=3. More improvement in under-performing languages is observed.

In each task, the feature-based approach with LSTM under-performs fine-tuning approach. We hypothesize that initialization from pretraining with lots of languages provides a very good starting point that is hard to beat. Additionally, the LSTM could also be part of the problem. In Ahmad et al. (2019) for dependency parsing, an LSTM encoder was worse than a transformer when transferring to languages with high word ordering distance to English.

3 Question #3: Does mBERT retain language specific information?

mBERT may learn a cross-lingual representation by abstracting away from language-specific information, thus losing the ability to distinguish between languages. We test this by considering language identification: does mBERT retain language-specific information? We use WiLI-2018 Thoma (2018), which includes over 200 languages from Wikipedia. We keep only those languages included in mBERT, leaving 99 languages Hungarian, Western-Punjabi, Norwegian-Bokmal, and Piedmontese are not covered by WiLI.. We take various layers of bag-of-words mBERT representation of the first two sentences of the test paragraph and add a linear classifier with softmax. We fix mBERT and train only the classifier the same as the feature-based approach in § 5.2.

All tested layers achieved around 96% accuracy (Fig. 2), with no clear difference between layers. This suggests each layer contains language-specific information; surprising given the zero-shot cross-lingual abilities. As mBERT generalizes its representations and creates cross-lingual representations, it maintains language-specific details. This may be encouraged during pretraining since mBERT needs to retain enough language-specific information to perform the cloze task.

4 Question #4: Does mBERT benefit by sharing subwords across languages?

As discussed in § 3, mBERT shares subwords in closely related languages or perhaps in distantly related languages. At training time, the representation of a shared subword is explicitly trained to contain enough information for the cloze task in all languages in which it appears. During fine-tuning for zero-shot cross-lingual transfer, if a subword in the target language test set also appears in the source language training data, the supervision could be leaked to the target language explicitly. However, all subwords interact in a non-interpretable way inside a deep network, and subword representations could overfit to the source language and potentially hurt transfer performance. In these experiments, we investigate how sharing subwords across languages effects cross-lingual transfer.

Discussion

We show mBERT does well in a cross-lingual zero-shot transfer setting on five different tasks covering a large number of languages. It outperforms cross-lingual embeddings, which typically have more cross-lingual supervision. By fixing the bottom layers of mBERT during fine-tuning, we observe further performance gains. Language-specific information is preserved in all layers. Sharing subwords helps cross-lingual transfer; a strong correlation is observed between the percentage of overlapping subwords and transfer performance.

mBERT effectively learns a good multilingual representation with strong cross-lingual zero-shot transfer performance in various tasks. We recommend building future multi-lingual NLP models on top of mBERT or other models pretrained similarly. Even without explicit cross-lingual supervision, these models do very well. As we show with XNLI in § 5.1, while bitext is hard to obtain in low resource settings, a variant of mBERT pretrained with bitext Lample and Conneau (2019) shows even stronger performance. Future work could investigate how to use weak supervision to produce a better cross-lingual mBERT, or adapt an already trained model for cross-lingual use. With POS tagging in § 5.1, we show mBERT, in general, under-performs models with a small amount of supervision while Devlin et al. (2019) show that in English NLP tasks, fine-tuning BERT only needs a small amount of data. Future work could investigate when cross-lingual transfer is helpful in NLP tasks of low resource languages. With such strong cross-lingual NLP performance, it would be interesting to prob mBERT from a linguistic perspective in the future.

References