Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, Michael Auli
Introduction
Self-supervised learning of speech representations has received a lot attention and has been demonstrated to work well both in low- and high-resource labeled data settings for automatic speech recognition (ASR) . However, the majority of studies focus on settings where there is little domain mismatch between the unlabeled data for pre-training, the labeled data for fine-tuning and the domain of the test data, or the target domain. It is well known that the performance of ASR systems trained from scratch with conventional supervised objectives can degrade significantly when tested on domains mismatched from training data . However, the impact of domain mismatch in self-supervised speech representation learning has been much less studied.
In this paper we present a series of experiments to better understand the impact of the various domain mismatches that can occur in the self-supervised data pipeline. We study the effect of increasing the amount of both in-domain and out-of-domain unlabeled data when fine-tuning the resulting model on both in-domain and out-of-domain labeled data. We also investigate the robustness of models on domains not seen during pre-training or fine-tuning. Our findings include that adding unlabeled data whose domain matches the test data always improves performance, even if the labeled data for fine-tuning is out-of-domain. This has immediate practical applications since it is much easier to obtain unlabeled data for a particular target domain than labeled data. We also find that pre-training on multiple domains increases robustness to completely unseen domains.
Related Work
This paper is related to a large body of work on robust ASR and domain adaptation. There are two popular lines of approaches. The first one is feature-based, which focuses on creating robust features . Both signal processing-based and learned features have been explored. The other line is model-based, which exposes a model to diverse data while minimally pre-processing speech input in order to exploit the model capacity. This includes data augmentation , self-training on target domain , domain adversarial training , and joint training . The self-supervised approach explored in this paper can be categorized as model-based; however, unlike the aforementioned model-based methods, it does not require any labeled data during pre-training by using a self-supervised objective.
The most related work to this paper is , which investigated domain-shift for self-supervised learning, but did not dissect the domains of data used during pre-training. We extend this work by also examining the effect of pre-training data domain. Furthermore, the pre-trained feature extractor in is fixed during supervised fine-tuning, making it more similar to feature-based approaches. Other related work includes pre-training on multiple languages and investigating how representations transfer between languages .
Experimental Setup
We experiment with wav2vec 2.0 which consists of a convolutional feature encoder to map raw audio to latent speech representations input to a Transformer to output context representations . Each represents about 25ms of audio strided by 20ms and the Transformer architecture follows BERT . During training, latent representations are discretized to with a quantization module to represent the targets in the objective. The quantization module uses a Gumbel softmax to choose entries from codebooks with entries each and the chosen entries are concatenated to obtain . The model is trained to identify the true quantized latent using for each masked time-step within a set of distractors sampled from other masked time steps.
Table 1 summarizes the setups explored in each subsequent section. Let in-domain and out-of-domain (OOD) be defined relative to the test data. We first study in Section 4.1 whether adding in-domain data for pre-training is beneficial when OOD data is used for fine-tuning. In Section 4.2, we study if adding data that are OOD to pre-training can improve the performance, since previous work has only verified the effectiveness of increasing in-domain pre-training data . In Section 4.3, robustness of a model is evaluated through testing on domains that are not seen during pre-training or fine-tuning. All of the experiments mentioned above are fine-tuned on labeled data of 10 hours for low-resource setups. In Section 4.4, a 100-hour labeled set is used for fine-tuning to validate if the conclusions from earlier sections still hold with more labeled data.
Careful ablation studies are conducted in the following two sections. Section 4.5 keeps the total amount of data used in pre-training constant and varies the domain similarity by mixing OOD and in-domain data with different ratios. Section 4.6 dives deeper into the question raised in Section 4.1 by showing how much test performance is improved with respect to the amount of in-domain pre-training data. In the last section, we scale up the experiments by using even more data for pre-training/fine-tuning and a larger wav2vec 2.0 model to compare with previous work.
We consider English datasets from six domains, where datasets are regarded as from same domain if they are collected with the same process. The six domains are: (1) LibriSpeech(LS) and Libri-light (LL), which contain about 960 and 60K hours of 16kHz crowd-sourced audiobook recordings, respectively, derived from the LibriVox project; (2) TED-LIUM v3(TD), which contains 452 hours of TED conference audio in 16kHz; (3) Switchboard(SB) and Fisher (collectively denoted as SF), which consist of about 300 and 2K hours telephone conversational speech recorded in 8KHz, respectively; (4) Common Voice(CV) with almost 700 hours crowd-sourced recordings of Wikipedia sentences in 48kHz; (5) Wall Street Journal(WS) with over 80 hours of read news text recordings, and (6) VoxPopuli(VP) , which contains 552 hours 16kHz speech recordings sourced from European Parliament plenary sessions.
Standard train/dev/test splits are considered for these datasets. Since SB/SF do not have a standard split, we follow the setup of that uses RT-03S for validation, and Hub05 Eval2000 SwitchBoard (H-SB) and CallHome (H-CH) subsets for testing. The chosen datasets cover a wide variety of linguistic and acoustic domains, enabling comprehensive studies for various domain shift scenarios. We re-sample all datasets to 16kHz for consistency. Transcripts are pre-processed following , which upper-cases letters and removes punctuation except for apostrophes, resulting in 27+1(space) symbols.
2 Pre-training
For experiments from Section 4.1 to Section 4.6, we consider pre-training on various combinations of LS, TD, and SF. The three largest datasets (LL, SF, CV) are combined for the scaling experiments in the last section. The wav2vec 2.0 Base architecture is used for all but the last section, which uses the Large model. Base/Large contain 12/24 transformer blocks, each of which with 768/1024 input and output dimensions, 3,072/4,096 inner (FFN) dimension and 12/16 attention heads. For Base, 10% dropout is applied to the quantize/Transformer input, and output after attention, activation, and FFN layer within each transformer block. Additionally, LayerDrop is applied with . LayerNorm is used at each convolution layer in the feature encoder for both Base and Large models.
The same pre-training hyperparameters are used for all Base models following regardless what combinations of datasets they are trained on, except for the number of updates. Models are trained for 400K steps for the ablation studies in Section 4.5 and 4.6 for efficiency, and 800k steps elsewhere. In our preliminary studies, we found that the improvement is marginal beyond 800k.
3 Fine-tuning
We consider 10 hour subsets from LS, SF, and TD for supervised fine-tuning as the low-resource setup in most sections, 100 hour subset of TD as the mid-resource in Section 4.4, and full LS and SB in Section 4.7 for the high-resource setup. LS-10h is taken from the official Libri-light split, and other subsets are sampled from the corresponding training set with genders balanced.
Pre-trained models are fine-tuned with connetionist temporal classification . The same fine-tuning hyperparameters are used for all pre-trained models when fine-tuned on the same labeled set. These parameters are tuned using the in-domain pre-training and fine-tuning setup. For example, TD-10h parameters are selected based on the TD-dev greedy decoding word error rate (WER) of a TD pre-trained model fine-tuned on TD-10h.
4 Language model and decoding
We decode each fine-tuned model using the word-level -gram language model (LM) from the target domain built with KenLM and report its WER. Wav2letter++ beam search decoder are used with a beam size 50, beam threshold 100, and Bayesian optimizationhttps://github.com/facebook/Ax is used to find decoding hyper-parameters over 100 trials: LM weight (), and silence score ($$). We use the LMs provided in for LS, SB, TD, CV, WS, and that from for VP.
Results
Often times, it is much easier to obtain unlabeled speech for a particular domain than labeled data which requires annotation. Motivated by this, we examine the benefit of adding unlabeled in-domain data to pre-training. To answer this question, we consider the following setup: We pre-train models on all 7 possible combinations of LS, TD, SB, and fine-tune each of them on the ten hour subsets of the labeled version of each corpus, LS-10h, TD-10h, SB-10h, respectively. WERs on the validation sets of these three domains are presented in Table 2. Unshaded columns are setups where the fine-tuning data are out-of-domain, and red numbers denotes models pre-trained with in-domain data.
For this section, we compare each black number with the red number on its right, where the red number is pre-trained additionally with the in-domain data. Both results are based on fine-tuning and testing on the same data. The benefit of adding in-domain data is clearly shown as all the red numbers are lower than the black numbers in Table 2.
2 Does adding pre-training data help if out-of-domain?
Here we still pay attention to Table 2 but compare numbers vertically. We split the question to be answered into two scenarios: (1) the original pre-training data does not contain in-domain data, and (2) otherwise. For the former, we compare black numbers in the second and the third row (pre-trained on one OOD dataset) with those in the last row (on two OOD ones). The black numbers in the last row are consistently better than those above, confirming the benefit of adding OOD data in this case.
For the other scenario, we compare red numbers within each column, where we found the hypothesis holds most of the time when increasing pre-training data from one domain (row 1) to two domains (row 2 and 3), with the only exception being the SB RT03 WERs when fine-tuned on SB-10h. However, when further increasing from two to three domains (row 4), the results are mixed, where about half of the cases improve.
3 Does pre-training on diverse data improve robustness?
We test the 21 fine-tuned models (7 pre-training dataset combinations 3 fine-tuning datasets) from earlier sections on three domains not seen during pre-training or fine-tuning: Wall Street Journal (WS), Common Voice (CV), and VoxPopuli (VP), and report the results in Table 3. In general, by comparing numbers within each column, one can observe that a model pre-trained on more domains tends to perform better than those pre-trained on fewer. To derive a summary statistic, we report the average WER over the six domains (LS, SB, TD, CV, WS, VP) for each fine-tuned model in the last three columns of Table 3. It shows that pre-training on three domains achieves better performance than on two domains, which in turn is better than on one domain, regardless of what labeled data they are fine-tuned on.
4 Is it still effective and robust with more labeled data?
We fine-tune four models on a larger TD-100h labeled set and test on LS dev-other to verify if the conclusions from Section 4.1 and 4.2 hold when more labeled data are available. Results in Table 4 confirm that both adding in-domain pre-training data (black vs. red) and adding out-of-domain data (row 1 vs row 2 vs row 3) are still effective.
5 Effect of pre-training data similarity to target domain
Section 4.1 shows that adding in-domain unlabeled data helps (red vs. black), but the improvement may be not just be due to domain similarity but it may also be due to simply increasing pre-training data size. To better understand the effect of domain similarity alone, we fix the amount of pre-training data to 450 hours and vary the ratio of TD/LS data to control domain similarity with respect to the test data, LS dev-other.
Table 5 shows performance improvements when increasing the amount of in-domain unlabeled data up to 50% of all pre-training data. If there is perfect domain match for the unlabeled data, labeled data and the target domain, then more unlabeled data leads consistently to better performance (grey shaded results). However, if the labeled data domain differs, then performance saturates at either 50% of in-domain unlabeled data for fine-tuning with TD-10h and 75% with SB-10h. The effect for this is particularly strong for TD-10h and we believe that it is beneficial to have some of the unlabeled data match the domain of the labeled data for fine-tuning in order for fine-tuning to be most beneficial.
6 Effect of in-domain pre-training data size
In this section, we study the relation between the amount of in-domain data used during pre-training and performance. Two strategies are considered: (1) joint training, which pre-trains on TD + LS of size for 400k steps, and (2) continual training, which first pre-trains on TD for 400k steps, and then pre-trains only on unlabeled LS of size for the numbers of steps specified in Table 6. In practice, it is convenient to pre-train a model on a large dataset, and then adapt it to a new domain of interest by running additional pre-training steps on that domain.
Result in Table 6 show that adding more in-domain unlabeled data continually improves performance for both joint training and continual training, and both strategies achieve similar performance. Compared to Table 5, using all in-domain unlabeled data still performs well when fine-tuned on TD-10h, because pre-training always includes all TD data which matches the domain of labeled data for fine-tuning.
7 Larger model, more pre-training and fine-tuning data
Finally, we pre-train a single large wav2vec 2.0 model with 300M parameters on three domains (LL, SF and CV) for 800K steps and fine-tune it on 10 hours as well as all data of the LS and SB datasets. We evaluate each model on the validation and test splits for LS, SB, CV and TD, showing in-domain and OOD performance. We compare our model to supervised models trained on single or multiple domains reported in (Table 7).
On in-domain data (LS/SB), our model achieves superior performance to all single- and multi-domain models in except on the CallHome (H-CH) set. Pre-training on multiple domains is particularly effective when testing on domains different from fine-tuning: when fine-tuned on LS, WER is reduced by a relative 35% to 50% on SB, CV, TD compared to the single-dataset baseline trained on full LS in . Moreover, we even achieve better OOD performance compared to that baseline when fine-tuning on just 10 hours of LS data (LS-10h). The same trend holds when comparing fine-tuning on SB-10h and for the baseline in trained on all of SF (200x more labeled data). Our model is not pre-trained or fine-tuned on any TD data, and when fine-tuning it with LS, it achieves better performance on TD than the supervised SOTA trained on TD as well as the joint RASR model trained on five labeled datasets including TD .
8 Effectiveness of pre-training
In the supervised learning paradigm, practitioners who would like to build a system for a new domain, can either train on existing OOD labeled data or build a corpus of labeled data in the new domain. With pre-training, we have a third option: collect unlabeled data in the new domain and fine-tune on existing labeled OOD data. This has the clear advantage of unlabeled in-domain being often much easier to obtain than transcribed in-domain data.
To get a better sense of how effective this third choice is, we re-examine our earlier experiments (Table 2). We measure the performance gap between the ideal setting where we have access to the labels of the in-domain data (Topline; PT-on-all, FT-on-InD) and a setting where we have only access to labeled out-of-domain data (Baseline; (PT,FT)-on-OOD). Table 8 shows how much of this gap is closed by the third option which in this case is pre-training on a variety of domains, including in-domain unlabeled data, while fine-tuning on labeled OOD data (Proposed; PT-on-all, FT-on-OOD). Pre-training on in-domain data closes at least 73% of the gap across a variety of settings.
We repeat the same analysis for the much larger and more competitive systems presented in § 4.7. Performance is measured on the evaluation sets of Switchboard (RT03/H-SB/H-CH). As baseline, we consider the competitive supervised model of trained on labeled LS only (Baseline; No PT, FT-on-OOD). Table 9 shows that pre-training on unlabeled in-domain data and fine-tuning on OOD closes between 66%-73% of the performance gap between the Topline and the Baseline performance. This bodes very well for practitioners who would like to build a model for a new domain since it is generally much easier to obtain unlabeled data for a new domain compared to transcribed data.
Conclusion
We present the first controlled study to better understand domain-shift in self-supervised learning for ASR. Results show that adding unlabeled in-domain data improves performance, even when the fine-tuning data does not match the test domain. With no access to in-domain labeled data, pre-training on unlabeled in-domain data closes 66-73% of the performance gap between the ideal setting of in-domain labeled data and a competitive supervised out-of-domain model. Moreover, self-supervised representations trained on a variety of domains are robust and lead to better generalization performance on domains completely unseen during pre-training and fine-tuning. Retaining some unlabeled data from the same domain as the fine-tuning data is beneficial though.