wav2vec: Unsupervised Pre-training for Speech Recognition

Steffen Schneider, Alexei Baevski, Ronan Collobert, Michael Auli

Introduction

Current state of the art models for speech recognition require large amounts of transcribed audio data to attain good performance (Amodei et al., 2016). Recently, pre-training of neural networks has emerged as an effective technique for settings where labeled data is scarce. The key idea is to learn general representations in a setup where substantial amounts of labeled or unlabeled data is available and to leverage the learned representations to improve performance on a downstream task for which the amount of data is limited. This is particularly interesting for tasks where substantial effort is required to obtain labeled data, such as speech recognition.

In computer vision, representations for ImageNet (Deng et al., 2009) and COCO (Lin et al., 2014) have proven to be useful to initialize models for tasks such as image captioning (Vinyals et al., 2016) or pose estimation (Pavllo et al., 2019). Unsupervised pre-training for computer vision has also shown promise (Doersch et al., 2015; Hénaff et al., 2019). In natural language processing (NLP), unsupervised pre-training of language models (Devlin et al., 2018; Radford et al., 2018; Baevski et al., 2019) improved many tasks such as text classification, phrase structure parsing and machine translation (Edunov et al., 2019; Lample & Conneau, 2019). In speech processing, pre-training has focused on emotion recogniton (Lian et al., 2018), speaker identification (Ravanelli & Bengio, 2018), phoneme discrimination (Synnaeve & Dupoux, 2016a; van den Oord et al., 2018) as well as transferring ASR representations from one language to another (Kunze et al., 2017). There has been work on unsupervised learning for speech but the resulting representations have not been applied to improve supervised speech recognition (Synnaeve & Dupoux, 2016b; Kamper et al., 2017; Chung et al., 2018; Chen et al., 2018; Chorowski et al., 2019).

In this paper, we apply unsupervised pre-training to improve supervised speech recognition. This enables exploiting unlabeled audio data which is much easier to collect than labeled data. Our model, wav2vec, is a convolutional neural network that takes raw audio as input and computes a general representation that can be input to a speech recognition system. The objective is a contrastive loss that requires distinguishing a true future audio sample from negatives (Collobert et al., 2011; Mikolov et al., 2013; van den Oord et al., 2018). Different to previous work (van den Oord et al., 2018), we move beyond frame-wise phoneme classification and apply the learned representations to improve strong supervised ASR systems. wav2vec relies on a fully convolutional architecture which can be easily parallelized over time on modern hardware compared to recurrent models used in previous work (§2).

Pre-training Approach

Given an audio signal as input, we optimize our model (§2.1) to predict future samples from a given signal context. A common problem with these approaches is the requirement to accurately model the data distribution p(x)p(\mathbf{x}), which is challenging. We avoid this problem by first encoding raw speech samples x\mathbf{x} into a feature representation z\mathbf{z} at a lower temporal frequency and then implicitly model a density ratio p(zi+k∣zi…zi−r)/p(zi+k)p(\mathbf{z}_{i+k}|\mathbf{z}_{i}\dots\mathbf{z}_{i-r})/p(\mathbf{z}_{i+k}) similar to van den Oord et al. (2018).

Our model takes raw audio signal as input and then applies two networks. The encoder network embeds the audio signal in a latent space and the context network combines multiple time-steps of the encoder to obtain contextualized representations (Figure 1). Both networks are then used to compute the objective function (§2.2).

The layers in both the encoder and context networks consist of a causal convolution with 512 channels, a group normalization layer and a ReLU nonlinearity. We normalize both across the feature and temporal dimension for each sample which is equivalent to group normalization with a single normalization group (Wu & He, 2018). We found it important to choose a normalization scheme that is invariant to the scaling and the offset of the input. This choice resulted in representations that generalize well across datasets.

2 Objective

where we denote the sigmoid σ(x)=1/(1+exp⁡(−x))\sigma(x)=1/(1+\exp(-x)), and where σ(zi+k⊤hk(ci))\sigma(\mathbf{z}_{i+k}^{\top}h_{k}(\mathbf{c}_{i})) is the probability of zi+k\mathbf{z}_{i+k} being the true sample. We consider a step-specific affine transformation hk(ci)=Wkci+bkh_{k}(\mathbf{c}_{i})=W_{k}\mathbf{c}_{i}+\mathbf{b}_{k} for each step kk, that is applied to ci\mathbf{c}_{i} (van den Oord et al., 2018). We optimize the loss L=∑k=1KLk\mathcal{L}=\sum_{k=1}^{K}\mathcal{L}_{k}, summing (1) over different step sizes. In practice, we approximate the expectation by sampling ten negatives examples by uniformly choosing distractors from each audio sequence, i.e., pn(z)=1Tp_{n}(\mathbf{z})=\frac{1}{T}, where TT is the sequence length and we set λ\lambda to the number of negatives.Similar to van den Oord et al. (2018), we found that sampling negatives from different sequences and speakers yields inferior results.

After training, we input the representations ci\mathbf{c}_{i} produced by the context network to the acoustic model instead of log-mel filterbank features.

Experimental Setup

We consider the following corpora: For phoneme recognition on TIMIT (Garofolo et al., 1993b) we use the standard train, dev and test split where the training data contains just over three hours of audio data. Wall Street Journal (WSJ; Garofolo et al. (1993a); Woodland et al. (1994)) comprises about 81 hours of transcribed audio data. We train on si284, validate on nov93dev and test on nov92. Librispeech (Panayotov et al., 2015) contains a total of 960 hours of clean and noisy speech for training. For pre-training, we use either the full 81 hours of the WSJ corpus, an 80 hour subset of clean Librispeech, the full 960 hour Librispeech training set or a combination of all of them.

2 Acoustic Models

We use the wav2letter+​+ toolkit for training and evaluation of acoustic models (Pratap et al., 2018). For the TIMIT task, we follow the character-based wav2letter+​+ setup of Zeghidour et al. (2018a) which uses seven consecutive blocks of convolutions (kernel size 55 with 10001000 channels), followed by a PReLU nonlinearity and a dropout rate of 0.70.7. The final representation is projected to a 39-dimensional phoneme probability. The model is trained using the Auto Segmentation Criterion (ASG; Collobert et al., 2016)) using SGD with momentum.

Our baseline for the WSJ benchmark is the wav2letter+​+ setup described by Collobert et al. (2019) which is a 17 layer model with gated convolutions (Dauphin et al., 2017). The model predicts probabilities for 31 graphemes, including the standard English alphabet, the apostrophe and period, two repetition characters (e.g. the word ann is transcribed as an1), and a silence token (|) used as word boundary.

All acoustic models are trained on 8 Nvidia V100 GPUs using the distributed training implementations of fairseq and wav2letter+​+. When training acoustic models on WSJ, we use plain SGD with learning rate 5.6 as well as gradient clipping (Collobert et al., 2019) and train for 1,000 epochs with a total batch size of 64 audio sequences. We use early stopping and choose models based on validation WER after evaluating checkpoints with a 4-gram language model. For TIMIT we use learning rate 0.12, momentum 0.9 and we train for 10001000 epochs on 8 GPUs with a batch size of 16 audio sequences.

3 Decoding

For decoding the emissions from the acoustic model we use a lexicon as well as a separate language model trained on the WSJ language modeling data only. We consider a 4-gram KenLM language model (Heafield et al., 2013), a word-based convolutional language model (Collobert et al., 2019), and a character based convolutional language model (Likhomanenko et al., 2019). We decode the word sequence y\mathbf{y} from the output of the context network c\mathbf{c} or log-mel filterbanks using the beam search decoder of Collobert et al. (2019) by maximizing

where fAMf_{\text{AM}} is the acoustic model, pLMp_{\text{LM}} is the language model, π=π1,...,πL\pi=\pi_{1},...,\pi_{L} are the characters of y\mathbf{y}. Hyper-parameters α\alpha, β\beta and γ\gamma are weights for the language model, the word penalty, and the silence penalty.

For decoding WSJ, we tune the hyperparameters α\alpha, β\beta and γ\gamma using a random search. Finally, we decode the emissions from the acoustic model with the best parameter setting for α\alpha, β\beta and γ\gamma. We use a beam size of 40004000 and beam score threshold of 250250 for the word based language models, and a beam size of 15001500 with beam score threshold 4040 for the character based language model.

4 Pre-training Models

Results

Different to van den Oord et al. (2018), we evaluate the pre-trained representations directly on downstream speech recognition tasks. We measure speech recognition performance on the WSJ benchmark and simulate various low resource setups (§4.1). We also evaluate on the TIMIT phoneme recognition task (§4.2) and ablate various modeling choices (§4.3).

Table 1 shows that pre-training on more data leads to better accuracy on the WSJ benchmark. Pre-trained representations can substantially improve performance over our character-based baseline which is trained on log-mel filterbank features. This shows that pre-training on unlabeled audio data can improve over the best character-based approach, Deep Speech 2 (Amodei et al., 2016), by 0.67 WER on nov92. In comparison to Hadian et al. (2018), wav2vec performs as well as their phoneme-based model and wav2vec large outperforms it by 0.37 WER. The phoneme-based approach of Ghahremani et al. (2017) pre-trains on the labeled version of Librispeech and then fine-tunes on WSJ. wav2vec large still outperforms Ghahremani et al. (2017) despite a weaker baseline model and not using Librispeech transcriptions.

2 Pre-training for TIMIT

On the TIMIT task we use a 7-layer wav2letter+​+ model with high dropout (§3; Synnaeve & Dupoux (2016b)). Table 2 shows that wav2vec pre-training on Librispeech and WSJ audio data can lead to results matching the state of the art. Accuracy steadily increases with more data for pre-training and the best accuracy is achieved with the largest amount of data for pre-training.

3 Ablations

In this section we analyze some of the design choices we made for wav2vec. We pre-train on the 80 hour subset of clean Librispeech and evaluate on TIMIT. Table 3 shows that increasing the number of negative samples only helps up to ten samples. Thereafter, performance plateaus while training time increases. We suspect that this is because the training signal from the positive samples decreases as the number of negative samples increases. In this experiment, everything is kept equal except for the number of negative samples.

Table 5 also shows that predicting more than 12 steps ahead in the future does not result in better performance and increasing the number of steps increases training time.

Conclusions

Acknowledgements

We thank the Speech team at FAIR, especially Jacob Kahn, Vineel Pratap and Qiantong Xu for help with wav2letter+​+ experiments, and Tatiana Likhomanenko for providing convolutional language models for our experiments.

References