Improving Multilingual Models with Language-Clustered Vocabularies

Hyung Won Chung, Dan Garrette, Kiat Chuan Tan, Jason Riesa

Introduction

Multilingual models such as mBERT Devlin et al. (2019), XLM Lample and Conneau (2019), and XLM-R Conneau et al. (2020) have built on the advances of deep contextualized language modeling by pretraining on texts from many languages at once. One trait common to all of these models is the use of a single vocabulary containing subwords from all languages, used to segment the input text before transforming it into a sequence of embeddings. Conneau et al. (2020) showed that increasing the vocabulary size can produce quality gains, but unlike similar monolingual models, the vocabulary embedding matrix in each of these multilingual models constitutes a significant fraction of its total parameters; for example, 47% of XLM-R’s parameters are in its embedding matrix. Therefore, scaling up a model’s vocabulary size requires the construction of inductive biases that will guide the training procedure to effectively and efficiently learn these parameters.

The multilingual subword vocabularies used by the state-of-the-art models are generated by algorithms such as WordPiece Schuster and Nakajima (2012); Wu et al. (2016), SentencePiece Kudo and Richardson (2018),SentencePiece uses BPE or unigram language model. or Byte Pair Encoding (BPE) Sennrich et al. (2016). Given a desired vocabulary size, these algorithms select an inventory of subwords that compactly represents the training corpora, which means preferring subwords that occur frequently, and, by extension for multilingual models, occur frequently across languages.

Because these algorithms look at overall subword frequencies in the combined multilingual corpus, they may learn suboptimal decompositions for low-resource languages that happen to have character combinations that resemble subwords of high-resources languages.For example, the fact that la is a good subword in French and Spanish does not mean that the algorithm should split the prefix la away from words in all languages. Additionally, subwords in common scripts like Latin and Cyrillic have a higher chance of selection since their counts are combined across a large number of languages Wu and Dredze (2019).

By attempting to optimize for the best overall vocabulary across all languages without regard for the differences among those languages, the joint procedure over-emphasizes cross-lingual subword sharing—even across languages with little lexical overlap—which is at odds with the finding of K et al. (2020) that subword sharing is not the principal reason for the effectiveness of multilingual models. It also under-emphasizes the need for all languages—particularly low-resource languages or those written in scripts used by few languages—to contribute subwords that are most effective for their own representations.

In this paper, we propose a novel approach to multilingual vocabulary generation that seeks to balance the trade-off between optimizing for cross-lingual subword sharing and the need for robust representation of individual languages. At a high level, we: 1) automatically group languages into clusters based on the distributional similarities of their individual subword vocabularies, 2) apply the SentencePiece algorithm separately to the data for each individual cluster, and finally 3) combine all cluster-vocabularies together to form a single unified multilingual model vocabulary.

We evaluate our approach on three distinct downstream tasks: TyDi QA, XNLI, and WikiAnn NER. Our experimental results show that our method improves model performance on all three tasks, and achieves a state-of-the-art result for TyDi QA and zero-shot cross-lingual NER. Crucially, our method improves performance without any changes to the model, and since it does not depend on the model architecture, it can be applied to any model that uses a vocabulary.

Clustered Vocabularies

We then generate a vocabulary VcjV^{c_{j}} for each cluster cj∈Cc_{j}\in C by pooling all of the pretraining data for all languages l∈cjl\in c_{j} and running the SentencePiece algorithm. We set the target vocabulary size to be proportional to the size of the union of the individual vocabularies VlV^{l} of the languages belonging to the cluster i.e. ∣Vcj∣∝∣⋃l∈cjVl∣|V^{c_{j}}|\propto|\bigcup_{l\in c_{j}}V^{l}|, ensuring that the proportion of the overall vocabulary allocated to a particular cluster is guided by two factors: it will increase if the cluster has more languages, and decrease if the cluster’s languages have more vocabulary overlap. Finally, the multilingual model’s overall vocabulary is the union of all of the cluster vocabularies: VC=⋃cj∈CVcjV^{C}=\bigcup_{c_{j}\in C}V^{c_{j}}.

Note that our method is a generalization of the conventional approach, which has ∣C∣=1|C|=1.

Intrinsic Analysis

In this section, we directly examine the vocabularies produced by the standard recipe (Joint) and our clustering-based method (Cluster) in order to assess the degree to which each approach is able to capture and balance the inductive biases introduced in §1: a multilingual vocabulary should encourage subword sharing across languages when appropriate, but each language should also have the freedom to contribute the subwords that are most effective for its own representation.

For all of our experiments, we generate vocabularies from the Wikipedia articles of 104 languages with 906M sentences. Each multilingual vocabulary is 488k subwords. The cluster assignments and sizes generated by Cluster are found in Table 2.

In order to quantify the extent of subword sharing between languages, we look at each language as a distribution over the vocabulary. In particular, we apply the multilingual vocabulary VV’s segmentation to each language ll’s monolingual corpus in order to count the frequency of each subword. The empirical distribution of a language ll over the vocabulary VV is defined by normalizing the frequencies of the monolingual corpus. Given these distributional representations of languages, we can use the Wasserstein-1 distance, W1W_{1}, between two languages to quantify the extent of subword sharing (Table 3). Unlike Joint for which W1W_{1} is relatively small even for languages with distinct scripts (e.g., English and Urdu), Cluster manifests much larger values for empirically different languages (and hence in different clusters) while having even smaller values for languages in the same cluster (e.g., Japanese and Chinese). This suggests that the clustering based approach not only minimizes the subword sharing when languages are dissimilar but it also puts more emphasis on subword sharing when necessary.

We quantify the degree of language freedom granted by each approach by examining the fraction of the vocabulary’s subwords that contain rare scripts (Table 4). Cluster has a higher percentage of Chinese, Japanese and Korean (grouped together as CJK) subwords because CJK languages are in the same cluster and by themselves, so CJK subwords can be selected independent of other languages. A similar pattern exists for Arabic script.

Experiments

The principal goal of this work is to investigate the effect of improved vocabulary composition on multilingual models. We make a good-faith effort to control all other variables, e.g., hyperparameters, training/evaluation procedures. In particular, we keep the number of languages constant, since per-language model capacity is known to affect the performance of multilingual models as shown in Conneau et al. (2020). In addition, we keep the number of parameters constantThe only exception is the full scale model in §4.2., including the vocabulary size, since the performance of Transformer-based Vaswani et al. (2017) models is strongly correlated with number of parameters (Lepikhin et al., 2020; Kaplan et al., 2020; Raffel et al., 2020; Conneau et al., 2020; Brown et al., 2020).

In order to demonstrate the effectiveness of our approach across languages and downstream tasks, we evaluate our method on three distinct datasets:

TyDi QA Clark et al. (2020): question answering in 10 languages. Results in Table 5.

XNLI Conneau et al. (2018): natural language inference in 15 languages. Results in Table 6.

WikiAnn NER Pan et al. (2017): named entity recognition in 40 languages. Results in Table 7.

We pretrain using the masked language modeling (MLM) task on the raw text without applying any language-specific pre-processing.See Appendix A for additional training details.

At a high level, the experimental results demonstrate that our Cluster approach to multilingual vocabulary generation improves over the standard Joint recipe, increasing the macro average performance across languages for all datasets and tasks. For TyDi QA, we see improvements on all languages. For XNLI, Cluster perform better or equally well in all languages except French. The results on individual languages are more mixed for NER, with Cluster providing large gains on low-resource or rare-script languages, and small losses on high-resource languages.

In the remainder of this section, we provide more in-depth analyses of these results (§4.1), and show that our approach continues to perform well when used to train a large scale model (§4.2).

We use the minimum description length principle (MDL) (Rissanen, 1989) to aid in our analysis. Consider an example where we want to encode data with a codebook. For example, each word in the data can be encoded by a unique sequence of bits, which is referred to as a codeword. The MDL principle favors the codebook with the minimal description length. Following Goldwater (2007), we define the description length (DL) as the length of combined codebook and encoded data; that is, the sum of the lengths of all codewords in our codebook plus the length of the encoded data.

We apply this to the setting of learning a subword vocabulary for neural network models. Our codebook is a mapping from a subword to a unique integer index of fixed length, typically 32 bits. Therefore, the description length is equivalent to the sum of the number of unique subwords, i.e. the vocabulary size, and the number of integers of the encoded or tokenized input data. Comparing the description length of two vocabularies is therefore equivalent to comparing the total number of subwords after the input corpus is tokenized by each vocabulary. Without loss of generality, we use the average number of tokens per sentence as an equivalent measure of the description length.

As shown in Table 8, Cluster does indeed reduce the DL of the training corpus, which correlates with the performance improvements we see on downstream tasks. With smaller DL, the input text is encoded with longer subwords, which we might expect to lead to a higher out-of-vocabulary (OOV) rate Arivazhagan et al. (2019). However, Cluster has an 8 times smaller OOV rate than Joint (Table 8). We believe that this is related to the larger extent of language freedom as evidenced by the particularly large OOV rate reductions observed in rare-script languages. For example, the OOV rate is reduced by a factor of 26 (Korean), 18 (Japanese) and 17 (Chinese).

The longer DL of Joint means the average subword length is shorter. As a result, the model has learn to map from finer-grained input sequences to semantic meaning Arivazhagan et al. (2019). As an extreme example, a character-based model would have to learn how to reconstruct each word, while a word-based model is exempt from this task. Though the difference in DL is smaller than this extreme case, the same logic applies and Joint must learn a more complex function than Cluster.

Finally, we note that this correlation between DL and downstream performance means we can use DL as a proxy metric to compare vocabularies, allowing for faster iteration over various vocabulary generation approaches without having to run the expensive model training.

Languages like Swahili and Yoruba use Latin script but have small amounts of data and Cluster outperforms Joint for all three tasks on these languages, which we believe is attributable to better segmentation. In §1, we highlighted one example where over-segmentation can occur,While Joint segments the common Swahili adverb lazima into la and zima, Cluster keeps it intact. but we can quantify this analysis more generally with DL: Cluster has 9.4% (Swahili) and 11.1% (Yoruba) shorter DL compared to Joint, and this matches well with 7.5 and 8.0 higher NER F1, respectively.

Rare-script languages.

For languages with rare scripts (e.g., CJK and Thai), Cluster strongly outperforms Joint in all tasks. For NER, particularly large gains are achieved for Arabic-script languages in cluster c2c_{2} (e.g., 28.9 F1 improvement in Urdu) and Indian languages in c8c_{8}. We hypothesize that the gain for this group of languages is due to the higher coverage of subwords in rare scripts, and consequently lower OOV rates (Table 4).

On clusters with languages in different scripts.

Our vocabularies were trained from the Wikipedia articles, which frequently contain translations/transliterations. For example, the first sentence of both the Hindi- and Tamil-language versions of the article on “India” contain the exact English phrase “Republic of India”. Since many of the same people/places will be topics in articles across Indic-script languages, it is not too surprising that the clustering algorithm groups some of these languages. We note that the NER performance on the Indic languages in cluster c8c_{8} are especially strong (Table 7) possibly due to such shared representation of named entities.

A similar phenomenon can be seen with Korean, since Chinese characters are often appended in parentheses to a Korean word to disambiguate polysemy, which is particularly common for formal contents like Wikipedia. This is why Chinese, Japanese (having large lexical overlap with Chinese) and Korean are in the same cluster c6c_{6}. We note that these languages show especially strong improvement in all three tasks we considered as well as drastically large reduction in OOV rate.

Therefore, we see these behaviors as a strength of our data-driven approach, flexibly capturing the characteristics of the data which may not be obvious from a purely linguistic perspective.

2 Full-scale model

To evaluate the effectiveness of our method on a large scale model, we train a 24-layer Transformer model with Cluster and report the macro-averaged results in Table 9. We drastically outperform the baseline for TyDi QA with about 10.7 F1 absolute improvement on MinSpan and 13.3 F1 on SelectP, and NER with 8.2 F1 absolute improvement over XLM-R. There is a small loss on XNLI, though this may be due to training on less data, and only Wikipedia domain.

Conclusion

We describe a novel clustering-based multilingual vocabulary generation algorithm. We showed that this empirically motivated clustering-based method consistently outperforms the standard vocabulary generation recipe used by most multilingual pretrained language modeling work without increasing the model size, compute or data.

Acknowledgements

We would like to thank Vera Axelrod, Tim Dozat, Melvin Johnson, Thibault Févry, Xavier Garcia, and Karthik Raman for their feedback on this work.

References

Appendix A Experiment details

We did not use any hyperparameter search for pretraining. We use LAMB optimizer You et al. (2020) with batch size of 4096. You et al. (2020) also recommended learning rate of 0.0018, warm-up proportion of 2.5% of the total number of steps. We use linear warm-up and linear learning rate decay down to 0 in the last step. We use gradient clipping with a norm of 1.0.

For all the experiments except for the full-scale model, we trained using 64 Google Cloud TPUs. We pretrained models with two different sequence lengths: 128 and 512 with 500k and 125k steps, respectively. The model with sequence length of 512 runs at about 2.3 steps/second, which takes about 16 hours to finish; this model was used to run WikiAnn NER and TyDi QA. The model with sequence length of 128 runs at 9.8 steps/second and finishes in about 15 hours; this model was used to run XNLI. A step refers to one gradient update of the LAMB optimizer.

During pretraining, we sampled each language’s data with the following strategy. First we compute the empirical distribution for each language

where nln^{l} is the number of sentences in language ll’s corpus. Then we use the exponential smoothing value of 0.7 following Devlin et al. (2019), i.e., we exponentiate plp^{l} and renormalize to compute the sampling probabilities of each language.

We used whole-word masking during pretraining. However, since some languages do not typically use whitespace between words (e.g., Thai), we used the heuristic of SentencePiece meta symbol U+2581 to designate the beginning of the word. Therefore, a word is defined as the token span between two successive U+2581 symbols.

For the full-scale model, we pretrained with 256 Google Cloud TPUs for 1.5 million steps with a batch size of 4096 and learning rate of 0.0018, which took 8 days. We only trained one model, and used a sequence length of 512.

A.2 SentencePiece configurations

We used the following configuration to train a SentencePiece model: unigram language model, character coverage of 0.9995, and 1M seed sentences.

A.3 Model architecture

Except for the full-scale model, all models were 12 layers of Transformers, with hidden size of 768 and 12 attention heads. For faster experimentation, we used an embedding size of 128, similar to Lan et al. (2020). The total number of parameters is 150M. The full-scale model has 24 Transformer layers, hidden size of 1024, and 16 attention heads. We used an embedding size of 512, totaling 550M parameters. We chose this number of parameters to mimic XLM-R.

A.4 Fine-tuning

We ran experiments with two seed values and chose the best model based on the average of the two runs. For fine-tuning, we used Adam optimizer Kingma and Ba (2015).

For WikiAnn NER, we used a learning rate of 4×10−54\times 10^{-5}, a batch size of 32, and 2 training epochs. We found that the performance is robust with respect to the set of hyperparameters, so we did not change this setting. The training was run with 4 TPUs which took about one hour to finish. The evaluation metric is span-level F1 score. Our evaluation code was tested against the seqeval library: https://github.com/chakki-works/seqeval, and produces the same scores.

For XNLI, we used a batch size of 32, performed grid search over the learning rate of [1×10−5, 2×10−5, 3×10−5][1\times 10^{-5},\ 2\times 10^{-5},\ 3\times 10^{-5}], and trained for 3 epochs. We chose the best model on the development set based on the macro-averaged accuracy and then used that model to report on the test set. The training was run with 8 TPUs which took about 2-3 hours to finish. The evaluation metric is classification accuracy.

For TyDi QA, we found that larger batch sizes improved the training stability, so we used a batch size of 512. With the larger batch size, longer epochs were helpful, so we used a grid search over the learning rate of [3×10−5, 4×10−5, 5×10−5][3\times 10^{-5},\ 4\times 10^{-5},\ 5\times 10^{-5}] and training epochs of [7, 8, 9][7,\ 8,\ 9]. We chose the best model based on the macro-averaged F1 score on 10 languages, excluding English following Clark et al. (2020). For the hyperparameters that are specific to TyDi QA, we used the same settings as the baseline model from Clark et al. (2020): 45 maximum passages, 0.1 include unknown rates, sequence length of 512, and window stride of 128. The evaluation metric is F1 score, which we computed with the official evaluation script from https://github.com/google-research-datasets/tydiqa.

Appendix B Datasets

For pretraining, we use 906M sentences of Wikipedia data covering 104 languages. The number of examples of the three datasets used for evaluation is summarized in Table 15. The XNLI data has a training set in English and development and test sets in 15 languages, which can be downloaded from https://cims.nyu.edu/ sbowman/multinli/.

The TyDi QA datasets are in 11 languages including English, which is excluded from the official evaluation. The link to download the dataset is https://github.com/google-research-datasets/tydiqa.

For Wikiann NER data, we follow XTREME Hu et al. and used the balanced train, development, test set splits of Rahimi et al. (2019), for 40 languages. The dataset can be downloaded from https://github.com/afshinrahimi/mmner.

We did not exclude any examples from the three datasets. For Wikiann NER, Greek, Thai, Japanese, Korean and Chinese data have incorrect IOB2 encoding, e.g., I-PER following O. We fixed those encoding with a simple rule such that a tag with I prefix starting after O is corrected to have B prefix.

We did not use any preprocessing for XNLI or TyDi QA.

Appendix C Performance on development set

The main body of the paper contains test set results for XNLI and WikiAnn NER. In this section, we report the development set results so that researchers can try to reproduce the results without consulting to the test set.

Table 10 shows the results for XNLI and the results for WikiAnn NER are summarized in Table 11. Both tables contain the best performing hyperparameters in the caption. Since the TyDi QA test set is private, we performed all experiments on the development set for TyDi QA except for the full-scale model for which we submitted to the official leaderboard.

For the full-scale model, Table 14 shows the development set results on WikiAnn NER, Table 13 for XNLI and, Table 12 for TyDi QA. We found that this large Transformer model requires different sets of hyperparameters to be effective. We used the LAMB optimizer to match the pretraining since it made the training more stable. For XNLI we did a grid search on over learning rates [9×10−5, 1×10−4][9\times 10^{-5},\ 1\times 10^{-4}], training epochs [9, 10, 11][9,\ 10,\ 11], and a fixed batch size of 512. For TyDi QA we did a grid search over learning rates [8×10−5, 9×10−5][8\times 10^{-5},\ 9\times 10^{-5}] and training epochs [13, 14, 15][13,\ 14,\ 15], and a fixed batch size of 512. For WikiAnn NER, we chose the learning rate from [2×10−5, 3×10−5][2\times 10^{-5},\ 3\times 10^{-5}], and used a fixed batch size of 32 and 2 training epochs.

C.2 WikiAnn NER test set results on all 40 languages

In this section, we expand the results in and Table 7 and Table 9, which only show subset of languages (or only average for the latter.

For these test set results, Table 16 expands on Table 7 and show results on all 40 languages considered in XTREME. Table 17 expands on Table 9 in a similar manner.

C.3 Language information

Table 18 lists all 104 languages considered in this paper and their script information.