Self-training and Pre-training are Complementary for Speech Recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, Michael Auli
Introduction
Speech recognition models trained on labeled speech data has progressed substantially in the recent past . A drawback of these models is that they require a lot of labeled data to perform well which is usually only available for English and a few other languages. Therefore, purely supervised training is impractical for the vast majority of the 7,000 languages spoken around the world which is why there has been a lot of interest in how to better use unlabeled speech data .
This includes classical self-training which demonstrated strong results by pseudo-labeling unannotated audio data and then retraining the final system with the additional labeled data. Another line of work is pre-training representations on unlabeled speech followed by fine-tuning on labeled data .
In this paper we combine self-training and unsupervised pre-training which are different approaches to leveraging unlabeled data. Both achieved excellent results on competitive benchmarks and the central question we explore is whether the two methods are complementary to each other. Specifically, we build on the recently introduced wav2vec 2.0 model and the self-training approach of Kahn et al. (2020; ) and Xu et al. (2020; ). We explore training models on the pseudo-labeled data from scratch or by fine-tuning the pre-trained model. To better understand how complementary the two methods are, we use the same unlabeled data for both.
Experiments on the full Librispeech corpus as well as the low-resource labeled data setups of Libri-light show that self-training and unsupervised pre-training are indeed complementary, a finding that is inline with recent work in natural language understanding . In a very low resource setup with just 10 minutes of labeled data and LibriVox as unlabeled data, the combination of wav2vec 2.0 and self-training achieves a WER of 3.0%/5.2% on the clean and other test sets of Librispeech, a relative WER reduction of 25% and 40% over recent work in pre-training alone . Using just the acoustic model without a language model achieves WER 3.7%/6.5% - supporting the hypothesis that self-training distills the language model used for pseudo-labeling into the final model. When all 960 hours of labeled training data are used we achieve 1.5%/3.1% WER on Librispeech.
Background
We experiment with the recently introduced wav2vec 2.0 model of Baevski et al. (2020; ). This model contains a convolutional feature encoder to map raw audio to latent speech representations which are input to a Transformer to output context representations . Each represents about 25ms of audio strided by 20ms and the Transformer architecture follows BERT . During training, feature encoder representations are discretized to with a quantization module to represent the targets in the objective. The quantization module uses a Gumbel softmax to choose entries from codebooks with entries each and the chosen entries are concatenated to obtain .
2 Self-training Approach
We adopt the pseudo-labeling strategy of Kahn et al. (2020; ) and Synnaeve et al. (2020; ). This first trains an initial acoustic model on the available labeled data and then labels the unlabeled data with the initial model as well as a language model in a step we call pseudo-labeling. Finally, a new acoustic model is trained on the pseudo-labeled data as well as the original labeled data.
Previous work considered multiple rounds of pseudo-labeling where the labeling step is repeated with each new model to train another model . While iterative pseudo-labeling is more accurate, we opt for a single iteration which is computationally less demanding while still enabling us to reason about whether unsupervised pre-training and pseudo-labeling are complementary. Another line of work investigated filtering the resulting pseudo-labeled data to match the distribution of the original labeled data . Both methods may improve results and we leave them to future work.
3 Combining the two Approaches
To combine the approaches, we replace the initial model for pseudo-labeling with a pre-trained model. The resulting training pipeline is as follows: we first pre-train a wav2vec 2.0 model on the unlabeled data, fine-tune it on the available labeled data, use the model to label the unlabeled data, and finally use the pseudo-labeled data to train the final model. In our experiments, we also consider a variant where we fine-tune the original wav2vec 2.0 model on the pseudo-labeled data.
Experimental Setup
As unlabeled data for pre-training and self-training we consider the speech audio of the Librispeech corpus (LS-960; ) without transcriptions containing 960h of audio as well as the audio data of LibriVox (LV-60k). For the latter we follow the pre-processing of Kahn et al. (2020; ) resulting in 53.2k hours of audio. We consider five labeled data setups: all 960h of transcribed Librispeech, the train-clean-100 subset comprising 100h, as well as the Libri-light limited resource training subsets of train-10h (10h), train-1h (1h), and train-10min (10min). We evaluate on the standard Librispeech dev-other/clean and test-clean/other sets.
2 Pre-trained models
Pre-trained models are implemented in fairseq and we obtain them from the public fairseq github repository.https://github.com/pytorch/fairseq/tree/master/examples/wav2vec The repository provides fine-tuned models for the five labeled data setups we consider (§ 3.1). We experiment with the Large configuration comprising 24 transformer blocks with model dimension 1,024, inner dimension 4,096 and 16 attention heads, comprising a total of about 300M parameters. The feature encoder contains seven blocks and the temporal convolutions in each block have 512 channels with strides (5,2,2,2,2,2,2) and kernel widths (10,3,3,3,3,2,2), resulting in a receptive field of about 25ms and a stride of about 20ms. After pre-training on the unlabeled data, this model is fine-tuned on the labeled data using Connectionist Temporal Classification (CTC; ) and a letter-based output vocabulary.
3 Self-training
We pseudo-label the audio data of either LS-960 or LV-60k using wav2vec 2.0 Large fine-tuned on different labeled data splits. For labeling, we follow the two-pass rescoring procedure of Synnaeve et al. (2020; ): first, we generate a list of candidate transcriptions by combining wav2vec 2.0 and the standard Librispeech 4-gram language model during beam-search with beam 800. Next, the n-best list is pruned to the 50 highest scoring entries and then rescored with a Transformer LM trained on the Librispeech language corpus . The Transformer LM has 20 blocks with model dimension 1,280, inner dimension 6,144 and 16 attention heads. The n-gram model obtains perplexity 150.3 on the development set and the Transformer language model 49.2. We found this to be more efficient than directly integrating the Transformer LM into beam search at little loss in accuracy. Decoding and rescoring hyper-parameters are tuned on dev-other of Librispeech for each experiment using a random parameter search. The LM weight and the word insertion penalty is tuned by randomly sampling values in the range of and over 128 trials.
4 Final Model
We follow Synnaeve et al. (2020; ) and train a Transformer-based sequence to sequence model with log-Mel filterbank inputs after pseudo-labeling using wav2letter++ . The encoder uses a convolutional frontend containing 4 layers of temporal convolutions with kernel width 3, followed by 36 Transformer blocks with model dimension , 4 attention heads and feed-forward network (FFN) dimension . The model contains about 300M parameters.
We use a 10k word piece output vocabulary computed from the training transcriptions if the whole Librispeech training set is used as labeled data . Otherwise, we switch to the 5k WP estimated on the train-clean-100 transcriptions . Language models are incorporated similar to § 3.3. We use a 4-gram language model and then rescore with a Transformer LM. The beam size used in both decoding and rescoring is 50.
Results
Pre-training has been shown to be very effective in both high- and low-resource labeled training data setups whereas self-training has been most effective when at least a moderate amount of labeled data is available ( 100h; ). To get a sense of whether the combination of both methods can be even more effective, we start with experiments on the Libri-light setups with 10min, 1h and 10h of labeled data. For pre-training and pseudo-labeling we use either the 960h of Librispeech without transcriptions or the 53.2k hours of LibriVox (§ 3.1). As baseline we consider wav2vec 2.0 pre-trained on Librispeech and fine-tuned on one of the labeled data splits.
We use the publicly available wav2vec 2.0 models to pseudo-label (ST) the unlabeled data and then evaluate two options to train the final model on the resulting labels: one is to train a new sequence to sequence model from random initialization with a word-piece vocabulary (s2s scratch) following Synnaeve et al. (2019; ; § 3.4). Another option is to fine-tune wav2vec 2.0 on the pseudo-labeled data with CTC and a letter-based vocabulary (ctc ft).
Table 1 shows that the combination of pre-training and self-training (wav2vec 2.0 + ST) outperforms pre-training alone (wav2vec 2.0) across all low-resource setups. It also achieves a very large improvement over iterative pseudo-labeling in the 10h labeled setup. This is because the initial model is much stronger due to pre-training and it is very difficult to train a good supervised-only model on just 10h of labeled data.
With just 10 minutes of labeled data, the combination of pre-training and pseudo-labeling with LibriVox achieves WER 5.2% on test-other. Using Librispeech (LS-960) as unlabeled data and 10 minutes of labeled data, wav2vec 2.0 + ST achieves 4.0%/7.2% WER on test-clean/other compared to 4.2%/8.6% for the best known pseudo-labeling approach which uses 100 hours of labeled data. More unlabeled data leads to large improvements, reducing WER from 4.0%/7.2% for LS-960 to 3.0%/5.2% for LV-60k, a relative WER reduction of 25-28%. But increasing the amount of labeled data without more unlabeled data leads to diminishing returns - an issue we return to in § 5. Fine-tuning (ctc ft) generally outperforms from scratch training of a sequence to sequence model with a WP vocabulary (s2s scratch). This is likely because the model can leverage the pre-trained representations.
2 High-Resource Labeled Data
Next, we evaluate performance with more labeled data. We consider the 100h clean subset of Librispeech as well as all 960h of labeled data in Librispeech. Table 2 shows that LS-960 as unlabeled data is not enough to outperform the baseline when 100h of labeled data is available. However, performance improves when using the much larger LV-60k, achieving a 10% relative WER reduction on test-other over wav2vec 2.0.
When using the full Librispeech benchmark as labeled data, combining wav2vec 2.0 and pseudo-labeling achieves WER 1.5%/3.1%. This result was achieved with a strong sequence to sequence model trained from scratch. While less effective than fine-tuning with CTC on smaller setups, a powerful sequence to sequence model excels in this larger setting since the decoder part of the model, which acts in part like a language model, does not overfit. The lower performance of fine tuning is likely due to CTC not being as competitive as more elaborate sequence to sequence models when a lot of pseudo-labeled data is available .
3 Results without a Language Model at Inference Time
Table 3 shows that combined training models have very good performance even without a language model. This is because the language model used during pseudo-labeling was partly distilled into the pseudo-labeled data . This effect is particularly striking for the 10 min labeled setup without LM where wav2vec 2.0 + ST (s2s scratch) reduces the WER of the baseline (wav2vec 2.0 - LM) by 83% relative on test-other. As more labeled data becomes available, the performance of the acoustic model without a language model improves but there is still a clear effect of self-trained models having distilled the language model. Generally, the sequence to sequence model in the (s2s scratch) setting is better able to distill the language model used at pseudo-labeling time compared to the CTC model used in fine-tuning.
Analysis
We previously saw that improvements decreased with more labeled data (§ 4.1). To better understand this, we perform an experiment on Librispeech where we consider data setups with a fixed ratio between the unlabeled and labeled data. Table 4 shows that relative improvements are a function of the amount of unlabeled data relative to the labeled data, rather than the amount of labeled data alone. Table 1 showed much larger improvements for the 10 min labeled split but with a fixed ratio of labeled and unlabeled data, the relative improvement is comparable to the 100h labeled setup (Table 2).
Conclusion
Unsupervised pre-training and pseudo-labeling are complementary for speech recognition. This enables building speech recognition systems with as little as 10 minutes of transcribed speech with word error rates that only a year ago were reserved to the best systems trained on 960 hours of labeled data.