XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, Michael Auli
Introduction
Self-supervised learning of generic neural representations has gathered much recent interest with a large body of work in natural language processing (NLP; Radford et al. 2018; Baevski et al. 2019; Devlin et al. 2019; Raffel et al. 2019), computer vision (Chen et al., 2020; He et al., 2020; Caron et al., 2021) as well as speech processing (van den Oord et al., 2018; Schneider et al., 2019; Baevski et al., 2020b; Hsu et al., 2021b; Chung et al., 2021). Self-supervised learning provides general representations that can be used across domains and languages.
Multilingually pretrained NLP models such as mBERT (Devlin et al., 2019), XLM-R (Conneau et al., 2020) or mT5 (Xue et al., 2020) brought significant improvements in multilingual language understanding (Conneau et al., 2018; Hu et al., 2020; Ruder et al., 2021). These models offer a promising path towards more ubiquitous NLP technology by improving performance for low-resource languages through leveraging data from high-resource languages. Furthermore, it is only necessary to maintain a single multilingual model instead of a myriad of monolingual models.
For speech processing, self-supervised approaches such as wav2vec 2.0 (Baevski et al., 2020b; Xu et al., 2021) have also been extended to the multilingual setting (Kawakami et al., 2020; Conneau et al., 2021). The recent XLSR (Conneau et al., 2021) leverages cross-lingual transfer from high-resource languages to build better representations for languages with little unlabeled data. The largest model, XLSR-53, was trained on about 50K hours of public training data in 53 languages and comprises about 300M parameters (Conneau et al., 2021). But such models only scratch the surface of self-supervised cross-lingual speech representation learning.
In natural language processing, language models are trained on very large datasets, spanning billions of documents such as CC100 (Wenzek et al., 2020) or mC4 (Xue et al., 2020) to fit models with tens of billions and even trillions of parameters (Brown et al., 2020; Goyal et al., 2021; Fedus et al., 2021) with strong results on established benchmarks. In contrast, scaling efforts in speech have focused either on supervised multilingual models (Li et al., 2021a) or monolingual self-supervised models, counting a billion or more parameters (Zhang et al., 2020; 2021), while cross-lingually pretrained speech models are much smaller in scale.
To this end, we present XLS-R, a large-scale cross-lingually pretrained wav2vec 2.0 model (see illustration in Figure 1) whose name is inspired by XLM-R in NLP. It leverages new publicly available VoxPopuli data, comprising 372K hours of unannotated speech (Wang et al., 2021a), the MLS corpus (Pratap et al., 2020), CommonVoice (Ardila et al., 2020), BABEL (Gales et al., 2014) and VoxLingua107 (Valk & Alumäe, 2020) to cover 128 different languages from various regions of the world. To our knowledge, this is the largest effort to date, in making speech technology accessible for many more languages using publicly available data.
Background
Our work builds on Conneau et al. (2021) who pretrain wav2vec 2.0 models on data from multiple languages. wav2vec 2.0 contains a convolutional feature encoder to map raw audio to latent speech representations which are input to a Transformer to output context representations (Baevski et al., 2020a). Each represents 25ms of audio strided by 20ms and the Transformer architecture follows BERT (Vaswani et al., 2017; Devlin et al., 2019).
During training, feature encoder representations are discretized to with a quantization module to represent the targets in the objective. The quantization module uses a Gumbel softmax to choose entries form the codebooks and the chosen entries are concatenated to obtain (Jegou et al., 2011; Jang et al., 2016; Baevski et al., 2020a).
The model is trained on multiple languages to obtain cross-lingual representations. Specifically, training batches contain samples from multiple languages (Devlin et al., 2019; Lample & Conneau, 2019; Conneau et al., 2021) by sampling from a distribution where , while is the amount of unlabeled data for each language, and is the upsampling factor which controls the trade-off between high- and low-resource languages during pretraining.
Data and Evaluation
In this section, we outline the datasets on which XLS-R is pretrained as well as the language coverage. We also describe the downstream tasks on which we evaluate our models.
We pretrain our models on a total of 436K hours of publicly available data from the following sources:
VoxPopuli (VP-400K) comprises a total of 372K hours of dataThere are 400K hours before removing leading and trailing silences. in 23 European languages of parliamentary speech from the European parliament (Wang et al., 2021a). This makes it the largest publicly available speech corpus for semi-supervised learning.
Multilingual Librispeech (MLS) contains data in eight European languages totaling around 50K hours of data (Pratap et al., 2020). The majority of the data is English (44K hours).
CommonVoice (CV) is a corpus of read speech. We use the December 2020 release (v6.1; Ardila et al. 2020) which covers 60 languages and over 7K hours of speech audio, ranging from over 1.6K hours for English to less than one hour for languages such as Hindi.
VoxLingua107 (VL) is a dataset of 6.6K hours of data in 107 languages based on YouTube content (Valk & Alumäe, 2020) with an average of 62 hours of data per language.
BABEL (BBL) is a multilingual corpus of conversational telephone speech of about 1K hours of data in 17 African and Asian languages (Gales et al., 2014).
To the best of our knowledge, this is the largest dataset used for training a publicly available self-supervised speech model to date. Figure 2 shows the data distribution across the 128 different languages in our training dataset. There are about 24 high-resource languages with more than 1K hours of data each, almost all of which are European, except for Kinyarwanda which is African. Then there is a small number of 17 mid-resource languages which have more than 100 hours of data (but less than 1K hours) which includes Catalan, Persian, Turkish, Russian, and Basque. Finally, the remaining 88 languages are low-resource and have less than 100 hours of data each. Table 1 lists all the languages, together with their ISO code, language family, sub-grouping and the amount of training data.
2 Downstream Evaluation
We evaluate on a broad and diverse set of downstream tasks to showcase the generalization ability of our pretrained models across tasks, data regimes, domains and languages.
For speech translation evaluation we adopt CoVoST-2 (Wang et al., 2020), a multilingual speech translation benchmark based on CommonVoice (Ardila et al., 2020).https://github.com/facebookresearch/covost It provides data for translating from English into 15 languages (En X) and from 21 languages into English (X En). The En X languages are: Arabic (ar), Catalan (ca), Welsh (cy), German (de), Estonian (et), Persian (fa), Indonesian (id), Japanese (ja), Latvian (lv), Mongolian (mn), Slovenian (sl), Swedish (sv), Tamil (ta), Turkish (tr), Chinese (zh) where each direction comprises about 430 hours of training data. The X En languages include all target languages of En X as well as Spanish (es), French (fr), Italian (it), Dutch (nl), Portuguese (pt). We group the latter into high-resource (136-264h of train data; fr, de, es, ca), mid-resource (10-49h of train data; fa, it, ru, pt, zh), and low-resource (2-7h of train data; tr, ar, et , mn, nl, sv, lv, sl, ta, ja, id, cy) for ease of presentation.
2.2 Automatic Speech Recognition (ASR)
BABEL is a challenging speech recognition benchmark from IARPA consisting of noisy telephone conversational data.https://catalog.ldc.upenn.edu/byyear LDC2016S06, LDC2016S13, LDC2017S05, LDC2017S08, LDC2016S12 We evaluate on five languages: Assamese (as), Tagalog (tl), Swahili (sw), Lao (lo), and Georgian (ka). Training sets comprise between 30 and 76 hours of annotated data. Following Conneau et al. (2021), we use 10% of the training set for validation, and report test results on the BABEL dev set. We report word error rate (WER) and use n-gram language models trained on CommonCrawl data.
Multilingual LibriSpeech (MLS; Pratap et al. 2020) is a large corpus derived from read audiobooks of Librivox and consists of eight European languages: Dutch (nl), English (en), French (fr), German (de), Italian (it), Polish (pl), Portuguese (pt), Spanish (es).https://github.com/flashlight/wav2letter/blob/main/recipes/mls Training sets comprise between 104 hours for Polish and 44.7K hours for English. We use the 10 hour training splits of Conneau et al. (2021) and report word error rate with the official n-gram models provided by the MLS dataset.
Following Rivière et al. (2020), we use ten languages of CommonVoice for ASR evaluation: Spanish (es), French (fr), Italian (it), Kyrgyz (ky), Dutch (nl), Russian (ru), Swedish (sv), Turkish (tr), Tatar (tt) and Chinese-Hong Kong (zh-HK).https://dl.fbaipublicfiles.com/cpc_audio/common_voices_splits.tar.gz CommonVoice contains read speech primarily from Wikipedia sentences. Following prior work (Rivière et al., 2020; Conneau et al., 2021), we fine-tune models on just one hour of labeled data per language, a few-shot scenario. Results are reported in terms of phoneme error rate (PER) without a language model.
Following Wang et al. (2021a), we evaluate on languages which have at least ten hours of labeled data which is a total of 14 languages: English (en), German (de), Italian (it), French (fr), Spanish (es), Polish (pl), Romanian (ro), Hungarian (hu), Dutch (nl), Czech (cs), Slovenian (sl), Finnish (fi), Croatian (hr), Slovakian (sk).https://github.com/facebookresearch/voxpopuli Models are fine-tuned on the full train set, which ranges from 543 hours (English) to ten hours (Slovenian). We report word error rate without language models.
LibriSpeech is a widely-used evaluation benchmark for speech recognition research (Panayotov et al., 2015). It consists of 960 hours of English annotated data. Following Baevski et al. (2020b), we use the 10 minute, 1 hour and 10 hour training splits. We compare the performance of XLS-R models against English-only wav2vec 2.0 models. We report word error rate without language models.
2.3 Speech classification (LID and Speaker ID)
For spoken language identification we consider VoxLingua107 (Valk & Alumäe, 2020) which spans 107 languages.http://bark.phon.ioc.ee/voxlingua107/ It consists of short speech segments automatically extracted from YouTube videos. The total amount of speech data in the training set is 6,628 hours and the average per language is 62 hours. We report accuracy on the official test set which covers 33 different languages.
We use VoxCeleb1 for speaker identification (Nagrani et al., 2017).https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html VoxCeleb1 is an audio-visual dataset consisting of short clips of human speech, extracted from interview videos uploaded to YouTube. It consists of 1,251 unique speakers and 153K utterances. We use the official dataset splits.
Experimental Setup
In this section, we give more details on architectures and hyperparameters used during pretraining and finetuning.
We use the wav2vec 2.0 implementation available in fairseq (Ott et al., 2019) and evaluate several model architectures detailed in Table 2. We consider models with between 0.3B parameters to 2B parameters. To optimize GPU memory usage, we use a fully sharded backend (Rajbhandari et al., 2021) as well as activation checkpointing (Chen et al., 2016) as implemented in FairScale (Baines et al., 2021).
Models are optimized with Adam (Kingma & Ba, 2015) and the learning rate is warmed up for the first 32K steps followed by polynomial decay to zero for the remainder of training. Training audio sequences are cropped to a maximum of 320K samples, or 20 seconds and all models were pretrained for a total of one million updates. XLS-R (0.3B) was trained on 128 GPUs with nearly 2M samples on each GPU, totaling about 4.3h of data in a batch. Larger models were trained on 200 GPUs with 800K to 1M samples on each GPU giving an effective batch size of about 2.8-3.6 hours.
Our training data covers 128 languages and five training corpora with different characteristics. To balance data from the different languages and corpora we upsample both training corpora and languages. We first upsample the languages within a particular corpus using the strategy outlined in §2 and then balance the different corpora using the same strategy by treating each corpus as a different language. We use in all cases.
2 Speech translation
To build speech translation models, we multilingually fine-tune XLS-R models by training on the combined labeled data of all En X or X En language directions without upsampling any direction. We stack a decoder network on top of XLS-R which is a Transformer network with 12 layers, embedding size 1024, 16 attention heads and feed forward network dimension 4096. The decoder network is initialized with weights from multilingually fine-tuned mBART (Liu et al., 2020; Li et al., 2021b; Tang et al., 2021) and uses the same vocabulary with 250K types. The total size of the decoder network is 459M parameters.
In our ablations, we also consider bilingually fine-tuned speech translation models which use a much smaller decoder of 16M parameters which has seven layers, embedding size 256, 4 attention heads and feed forward network dimension 2048. For this, a 10K byte-pair encoding (BPE; Sennrich et al. 2016) vocabulary is built on the CoVoST 2 target text for each target language.
We fine-tune with Adam (Kingma & Ba, 2015), a learning rate of 3e-4, label smoothing with probability 0.1, an effective batch size of 66M samples, or nearly 68 minutes, layer drop 0.05, a masking strategy similar to wav2vec 2.0 with mask length 5 and mask probability 0.15. During fine-tuning, the wav2vec 2.0 encoder is not updated for the first 10K updates. Models are fine-tuned for 250K updates in total and the best checkpoint is selected based on validation BLEU. We choose the learning rate by searching in the interval . Translations are generated with a beam size of 5.
3 Automatic Speech Recognition
For fine-tuning, we follow the settings of Baevski et al. (2020b) by adding a linear layer on top of the pretrained model to predict the output vocabulary and train using Connectionist Temporal Classification (CTC; Graves et al. 2006). The output vocabulary is characters for all benchmarks, except for CommonVoice where we use phonemes. We fine-tune using Adam and the learning rate is warmed up for the first 10% of total updates, kept constant for the next 40% and then decayed to zero in the remaining 50% of updates. Since the amount of labeled data differs widely for each dataset we found the following number of training updates to be effective: 20K updates for BABEL, 13K updates for CommonVoice, 20K updates for MLS and 50K updates for VoxPopuli.
We found large batch sizes to be very effective. For XLS-R (0.3B) and XLS-R (1B), we use an effective batch size of 0.44 hours and for XLS-R (2B) we used 1.06 hours. Learning rate as well as batch size was tuned based on dev error rate and we searched the range for XLS-R (0.3B) and XLS-R (1B) as well as for XLS-R (2B). To reduce overfitting for XLS-R (2B), we increase stochastic depth to 15% compared to 10% for all other setups (Huang et al., 2016; Fan et al., 2019). We use a masking probability of between 30-75%, depending on the setup and which is determined on the development set.
For MLS and BABEL, we use a language model for decoding. We tune the language model weight within the interval using Bayesian optimization.https://github.com/facebook/Ax We run 128 trials with beam 500 and choose the best set of parameters according to the dev error rate. For Common Voice and Vox Populi we do not use a language model.
Results
Next, we analyze the results of our pretrained models on all downstream tasks.
We conduct an extensive study on the CoVoST-2 speech translation benchmark. The task entails translating speech audio in one language into another language with text as output. Performance is evaluated in terms of BLEU (Papineni et al., 2002). Models are simultaneously fine-tuned either on all 21 translation directions with English as target language (X En) or on all the 15 directions where English is the input language (En X; see §3.2.1), resulting in only two models instead of 36.
For X English directions we group languages into high-resource, mid-resource and low-resource directions (§3.2.1) and compare to several baselines: in order to directly compare to XLSR-53 (Conneau et al., 2021), and VP-100K (Wang et al., 2021a), we fine-tune these publicly available models following the same protocol as XLS-R. We also compare to Li et al. (2021b), the best known results from the literature who either use an English-pretrained wav2vec 2.0 model (XMEF-En) for En X directions or the multilingually pretrained XLSR-53 (XMEF-X) for X En directions.
Table 3 shows a new state of the art with XLS-R (2B), improving over the previous best result (Li et al., 2021b), by 7.4 BLEU on average over all 21 directions (14.7 BLEU vs. 22.1 BLEU). This is largely due to improvements on mid-resource (+7.5 BLEU) and low-resource (+9.2 BLEU) language directions. Model capacity has a large impact: XLS-R (1B) improves over XLS-R (0.3B) by an average of 6.1 BLEU and XLS-R (2B) improves by an average of 2.8 BLEU compared to XLS-R (1B). Appendix A shows detailed results for all translation directions.
There is a trend of larger capacity in pretrained models enabling few-shot learning for speech translation, similar to wav2vec 2.0 enabling few-shot speech recognition (Baevski et al., 2020b; Xu et al., 2020). For example, on language pairs with only two hours of labeled speech translation data, XLS-R (2B) improves over XLS-R (0.3B) as follows: from 10.3 BLEU to 29.6 BLEU on Swedish-English, from 1.4 BLEU to 16.5 BLEU on Indonesian-English and from 3.0 BLEU to 17.1 BLEU on Arabic-English (see Appendix A).
1.2 English →→\rightarrow X
For English X directions we compare to previous cross-lingually pretrained models (XLSR-53, VP-100K) as well as baselines with English-only pretraining: XMEF JT, the best performing setup of Li et al. (2021b) for En X directions as well as wav2vec 2.0 pre-trained on 60K hours of English Libri-light data and fine-tuned following the same protocol as XLS-R (Kahn et al., 2020; Baevski et al., 2020b). The latter has the advantage of being pre-trained on exactly the same language as the input data for all translation directions while cross-lingually pretrained models need to be able to represent many different languages which puts them at a disadvantage.
Table 4 shows that XLSR-53 now performs similarly to XLS-R (0.3B) while for X English XLS-R (0.3B) performed much better (see §5.1.1). This is likely because English data dominates the training corpus of XLSR-53 which is not the case for XLS-R (§3.1). Both XLS-R (1B) and XLS-R (2B) outperform XMEF JT showing that larger capacity results in better performance.
We also compare to prior work using the English-only pretrained wav2vec 2.0 LV-60K model (Wang et al., 2021c) which additionally uses self-training and a language model for decoding. We do not use these techniques. Their results represent the state of the art on these four directions. Wang et al. (2021c) achieves an average BLEU of 25.6 on the four directions while as XLS-R (2B) rivals this at an average BLEU of 25.5. We note that self-training and LM decoding methods are equally applicable to our approach.
XLS-R (2B) also performs well compared to English-only pretraining at 27.8 average BLEU compared to 26.6 BLEU for a wav2vec 2.0 model pretrained on 60K hours of Libri-light data and 720M parameters. This confirms that with sufficient capacity, cross-lingual pretraining can perform as well as strong monolingual models (Conneau et al., 2021).
1.3 Ablations
We build speech translation models by adopting two design decision from Li et al. (2021b): multilingual fine-tuning of pretrained models on labeled speech translation data in multiple translation directions and initializing the decoder network with mBART (Li et al., 2021b; Tang et al., 2021). In the following we ablate these two choices to better understand their impact.
We first compare multilingual fine-tuning to bilingual fine-tuning. For faster experimental turn-around we consider a reduced setup of four English X language directions (en-ca, en-ar, en-de, en-tr) as well as all high-resource and mid-resource X English directions.We also did not use mBART initialization for this ablation. We compare bilingual fine-tuning to models fine-tuned on all 15 English X or all 21 X English directions.
Table 5 shows that multilingual fine-tuning is particularly effective for X English directions where the average improvement is 3.3 BLEU (20.9 BLEU to 24.2 BLEU). The amount of labeled data ranges from 264 hours for French English to 10 hours for Chinese English and multilingual fine-tuning leverages supervision from high-resource languages to improve performance for languages with less labeled data. Languages with less data benefit both from cross-lingual transfer during pretraining, through training on unlabeled data in other languages, and fine-tuning, through labeled data from other languages (Arivazhagan et al., 2019; Conneau et al., 2021). For English X, multilingual fine-tuning performs roughly on par to bilingual fine-tuning which is a desirable outcome given that transfer between language directions is diminished. This is supported by the larger size of the decoder network in multilingual fine-tuning (§3.2.1).
Next, we analyze the impact of initializing the decoder network with mBART which was pretrained on additional labeled text-to-text machine translation data (Liu et al., 2020; Li et al., 2021b; Tang et al., 2021). Specifically, we use MBART-ML50N1 (49 languages to English) for X English directions and MBART-ML501N (English to 49 languages) for English X directions.https://github.com/pytorch/fairseq/tree/main/examples/multilingual#mbart50-models We observe that mBART initialization has little impact on English X but it leads to large improvements for X English, especially on mid- and low-resource language directions.
Initializing the decoder network with mBART resulted in some low-resource languages moving from 1-3 BLEU to 10+ BLEU. The labeled translation data used to train mBART helps speech translation to adapt faster to the low supervision in mid/low-resource language pairs of the CoVoST-2 benchmarks where many language directions have only a few hours of labeled data. This shows that pretraining both the encoder and decoder, multilingual fine-tuning, as well as the use of extra machine translation data through mBART, enables few-shot learning for some speech translation directions which have only a few hours of labeled data.
2 Speech Recognition
Our experiments cover four common speech recognition benchmarks, 26 different languages, three different domains and both low-data and high-data regimes. The BABEL dataset evaluates models on noisy speech (§5.2.1), CommonVoice presents a few-shot setup with just one hour of labeled data per language (§5.2.2), MLS contains read speech in multiple European languages (§5.2.3), and VoxPopuli contains parliamentary speech with varying amounts of labeled data (§5.2.4).
BABEL consists of the hardest speech recognition setting among our four benchmarks which results in higher word error rates. Languages are low-resource, the speech is very noisy and corresponds to natural telephone conversation. Many competitions have tackled this dataset (Alumäe et al., 2017; Ragni et al., 2018; Inaguma et al., 2019) and baselines are thus well tuned. We compare to the best numbers we have found in the literature, as well as our own best baselines.
Table 7 shows that XLS-R (0.3B) outperforms the equally sized XLSR-53, which was the previous state of the art on all languages by an average of 1.4 WER, e.g., on Assamese (as), WER decreases from 44.1 to 42.9, on Swahili (sw) WER decreases from 26.5 to 24.3 and on Georgian (ka) WER drops from 31.1 to 28.0 WER. XLSR-53 and XLS-R were both pretrained on the same BABEL data, and the better performance of XLS-R (0.3B) shows that pretraining on additional out-of-domain datasets such as VoxPopuli does help performance on BABEL. This is similar to findings for monolingual pretraining (Hsu et al., 2021a).
Using additional capacity, XLS-R (1B) outperforms XLS-R (0.3B) by 2.5 WER on average. On Georgian (ka), this corresponds to improvements of 6 WER and 7.1 WER compared to Conneau et al. (2021) and Alumäe et al. (2017), respectively. Compared to results from only three years ago from Ragni et al. (2018) and Inaguma et al. (2019), XLS-R (1B) reduces word error rate by more than 10 WER, from 40.6 to 30.6 on Tagalog and from 35.5 to 21.2 on Lao. XLS-R (2B) improves over XLS-R (1B) by 0.8 BLEU on average showing that additional capacity can further improve performance.
2.2 CommonVoice
CommonVoice is an easier task than BABEL because it is read-speech but we use a reduced labeled data setup which introduces a different challenge. Specifically, we use the train/dev/test splits of Rivière et al. (2020) which corresponds to a few-shot scenario where only one hour of training data is available per language.
On English speech recognition, pretraining has been shown to be particularly beneficial for low labeled data settings (Baevski et al., 2020b). This is similar to cross-lingual pretraining (Conneau et al., 2021) where pretraining on the large MLS corpus significantly improved performance over pretraining only on CommonVoice data, e.g., on Dutch accuracy improved from 14 PER to 5.8 PER.
Table 8 shows that the additional training data of XLS-R compared to XLSR-53 results in better performance of 1.1 PER on average for XLS-R (0.3B). XLS-R uses the same training data as XLSR-53 plus the very large VP-400K corpus of parliamentary speech as well as the much smaller VoxLingua-107 which consists of YouTube data, both of which are out of domain with respect to the read audiobook domain of CommonVoice. This confirms that pretraining on more out of domain data can still improve performance (Hsu et al., 2021a).
Furthermore, accuracy improves even on languages for which XLS-R does not add any pretraining data compared to XLSR-53, e.g., Kyrgyz (ky) improves from 6.1 PER to 5.1 PER for XLS-R (0.3B) and 4.1 PER for XLS-R (1B) and both models are pretrained on only about 11 hours of Kyrgyz data - 0.003% of the total pretraining data. This shows that there is cross-lingual transfer that benefits low-resource languages and that additional capacity is important to realize this effect.
Chinese improves the least and gains are particularly large for languages for which the training corpus of XLS-R contains more data due to VoxPopuli, e.g., for Swedish VP-400K adds more than 16K hours of unannontated speech and performance improves from 12.2 PER to 5.5 PER when comparing XLSR-53 to XLS-R (1B). Finally, XLS-R (2B) performs slightly better than XLS-R (1B) on average with some languages improving while as others are performing slightly worse. The modest average improvement is likely because error rates are already low on this benchmark.
2.3 Multilingual LibriSpeech
Multilingual LibriSpeech is a common benchmark for evaluating multilingual speech recognition on eight European languages. We consider a setup where we use ten hours of labeled data for each language (Conneau et al., 2021) and compare to the prior work of Pratap et al. (2020) which uses all available labeled data as well as XLSR-53 (Conneau et al., 2021) which uses the same ten hour reduced labeled data setup.
Table 9 shows that XLS-R can outperform XLSR-53 on average by 1 WER at equal capacity and that additional model capacity results in an improvement of 2.9 WER on average for XLS-R (1B). This result rivals the performance of the supervised models of Pratap et al. (2020) which is based on significantly more labeled data compared to the ten hour setup of XLS-R. Finally, on average XLS-R (2B) does not show improvements over XLS-R (1B), which is similar to CommonVoice.
2.4 VoxPopuli
The VoxPopuli corpus provides about 1.8K hours of labeled speech data in 14 languages, ranging from 543 hours for English to 35 hours for Slovakian, as well as about 400K hours of unlabeled speech. This dataset is representative of a setting where a lot of unannotated data in the same domain as the labeled data is available. We compare to the work of Wang et al. (2021a) which used a cross-lingually pretrained wav2vec 2.0 Base model on an earlier version of VoxPopuli that contained about 10K hours of unlabeled speech.
Table 10 shows that cross-lingual pretraining (VP-10K) reduces WER from an average of 37.5 for supervised-only training (No pretraining) to 15.3 WER. XLS-R uses a lot more unlabeled VoxPopuli data and this results in improved performance: XLS-R (0.3B) improves over VP-10K by an average of 2.5 WER. The largest gains are on English, where WER improves from 16.2 to 10.2, likely due to the use of more English data from MLS during pretraining (over 44K hours). Increasing model capacity to 1B parameters results in even better performance, reducing WER from an average of 15.3 for VP-10K to 10.6.
2.5 LibriSpeech
On LibriSpeech English ASR, we compare XLS-R (0.3B) and XLS-R (1B) to the wav2vec 2.0 English baseline. We see in Table 11 that with the same capacity and same fine-tuning procedure, the English wav2vec 2.0 significantly outperforms the XLS-R (0.3B) in all data regimes, showing the capacity dilution and interference problem of our multilingual model. However, when increasing the capacity, the model is able to catch up with the monolingual results. In particular, XLS-R (1B) outperforms wav2vec 2.0 LV-60k on the 10 minute setting, but is at a disadvantage in the 10 hour setting, where the English-focused monolingual model outperforms it by 0.7 WER on average. This suggests that higher-capacity models can circumvent the interference problem and can get strong results on high-resource languages, while still leveraging their cross-lingual transfer ability for lower-resource languages Conneau et al. (2021).
3 Speech Classification
Finally, we evaluate our approach on two speech utterance classification tasks, language identification and speaker identification. For these tasks we use our smallest model as these tasks require less capacity given the lower complexity of the tasks compared to the structured prediction problems of speech recognition and speech translation.
For language identification we adopt the setup of VoxLingua107 (Valk & Alumäe, 2020) which provides data for 107 different languages. We train our model on the official train set, and report results on the development set, comprising 33 languages.
Table 12 shows that our best model outperforms previous work, improving the best known prior work of Ravanelli et al. (2020) by 1% absolute, a relative error reduction of 15%. For comparison, we also fine-tune the English-only wav2vec 2.0 pretrained on Libri-Light which performs surprisingly well on this multilingual task but is outperformed by the XLS-R model by 1.5% error rate on average.
3.2 Speaker Identification
Finally, we consider speaker identification on VoxCeleb1 where we fine-tune our model to distinguish between a fixed set of speakers given an utterance. We compare to prior work including results published as part of the recently introduced SUPERB benchmark (Yang et al., 2021) but note that their results are not strictly comparable because they do not fine-tune the underlying pre-trained model. All parameters of XLS-R are fine-tuned, similar to the evaluation of all other tasks. The results (Table 13) show that our cross-lingual model also performs very well for speaker identification, even though utterances are mostly in English.
4 Discussion
Cross-lingual training results in a single model for multiple languages compared to a separate model for each language. Training a cross-lingual model requires more effort than a single monolingual model but the resulting model can be used for many different languages. Advances in architectures and training can also be deployed more easily since we only need to retrain a single model rather than many different ones.
In terms of accuracy, prior work in self-supervised learning for speech established that cross-lingually pretrained models are very competitive to monolingually pretrained models for speech recognition (Conneau et al., 2021). Our experiments show a similar trend for speech-translation: XLS-R can perform very competitively to English-only pretrained models for English X speech translation where the encoder only needs to encode English speech - a setting which favors monolingually pretrained models.
Overall, XLS-R performs best for low-resource and mid-resource languages. For speech translation, we observe strong improvements for low- and mid-resource X English directions and comparatively smaller gains on high-resource directions. Many low-resource directions which previously had performance in the 1-5 BLEU range improve to over 10-20 BLEU due to the better cross-lingual speech representations. For English X directions, large enough cross-lingual models can even surpass the performance of English-only pretrained models.
Similarly, for speech recognition, we see strong improvements on BABEL, CommonVoice and VoxPopuli, benchmarks which include low- and mid-resource tasks.MLS is a notable exception and we attribute the different performance pattern to prior work having pretrained on large amounts of in-domain data. We find that models trained on more data from more languages can perform as well or better than comparable models of the same size and we we observe this trend across all speech recognition benchmarks. Keeping everything else equal, larger capacity models often further improve performance.
Conclusion
XLS-R is a new self-supervised cross-lingual speech representation model which scales the number of languages, the amount of training data as well as model size. The training corpus is an order of magnitude larger than prior work and covers 128 languages in 436K hours of recorded speech audio. The resulting model enables state of the art results for X English speech translation on CoVoST-2, outperforming prior art by a sizeable margin with the largest improvements on mid- and low-resource directions. It also performs competitively to the best English X work, without the use of equally applicable techniques such as self-training and language model decoding.
On speech recognition, XLS-R sets a new state of the art on CommonVoice, VoxPopuli, several languages of BABEL, while performing competitively on MLS with much less labeled data. These datasets cover a wide range of languages, data regimes and domains, demonstrating the generalization ability of XLS-R. Our model also sets a new state of the art on the VoxLingua107 language identification benchmark. The largest XLS-R model comprises 2B parameters which enables it to outperform a strong English-only pretrained model on English X speech translation, a setting which favors monolingually pretrained models. This shows that cross-lingually trained models with sufficient capacity can perform as well as specialized monolingually pretrained models. We hope XLS-R will help catalyze research in speech technology for many more languages of the world. Models and code are publicly available on several platforms.
We thank Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Emmanuel Dupoux for early access to VP-400K, and Christophe Ropers for his advise on language categorization. We also thank Min Xu, Jacob Kahn, Shruti Bhosale, Anjali Sridhar, and Tatiana Likhomanenko for help with infrastructure.