Joint Masked CPC and CTC Training for ASR

Chaitanya Talnikar, Tatiana Likhomanenko, Ronan Collobert, Gabriel Synnaeve

Introduction

Deep learning has been impactful in building state-of-the-art end-to-end speech recognition systems . But, they typically require large amounts of annotated speech data in the form of transcripts. Whereas, humans are able to learn language and speech with little supervision.

Recently, self-supervised learning (SSL) has been proposed as a method for training automatic speech recognition (ASR) models by pre-training on large amount of unlabeled data and then fine-tuning the speech recognition model on labeled data, for example contrastive predictive coding (CPC) . While these methods have achieved impressive results on low-resource speech datasets, their goal is to learn speech representations that are useful for multiple speech-related tasks. Training an ASR model using SSL methods is a two-stage process as it requires running separate pre-training and fine-tuning experiments and jointly tuning hyperparameters for both stages. It is unclear how much pre-training is required to achieve reasonable performance on the downstream task of speech recognition.

In this paper, we propose a training method for ASR models that combines SSL and supervised learning in a single stage. The model is trained by jointly minimizing a loss on labeled data, and a loss on unlabeled data. The supervised loss is the Connectionist Temporal Classification (CTC) loss , while the unsupervised loss is based on a masked variant of CPC. As both losses are optimized jointly, our method allows early stopping by measuring the performance of the model for the downstream task on the validation dataset.

We show that a model trained using our method (with no quantization) achieves equivalent word error rate (WER) when trained on 960-hours of unlabeled data and 100-hours of labeled data to a model that is trained using the two-stage process of wav2vec 2.0 (with quantization), which is a method based on masked CPC. Additionally, we verify that our method provides a regularization to the supervised loss when only using labeled data.

Joint Training

We propose to train our speech recognition model in a single stage, by jointly minimizing a supervised and an unsupervised loss. Our training procedure alternates between minimizing the unsupervised loss on unlabeled data and minimizing the supervised loss on labeled data.

Our model is a neural network architecture which gets as input raw audio (x{\bm{x}}) and outputs token (y{\bm{y}}) probabilities pθ(yt∣x)p_{\bm{\theta}}({\bm{y}}_{t}|{\bm{x}}) at time tt with the following functions:

2 Unsupervised and supervised losses

3 Alternate minimization

The model is trained by alternately minimizing the two losses. Using a minibatch from the unlabeled data, the gradient of the unsupervised loss is used to update the model parameters for NN steps, followed by the gradient of the supervised loss (using a minibatch from labeled data) for 11 step. This process is repeated until convergence of the word error rate on the validation dataset. A brief description is shown in Algorithm 1.

Separate adaptive momentum optimizers are used for each of the two losses with different learning rates: ηu\eta_{u} for the unsupervised loss and ηs\eta_{s} for the supervised loss. The two optimizers maintain their state independently, while sharing the parameters of the model. This ensures that the momentum averaging for one loss is not affected by the gradient updates from the other loss, leading to faster convergence. Experiments with a single optimizer show worse performance on the downstream task compared to the usage of two optimizers.

The ratio of unsupervised to supervised loss updates, NN:1, is chosen to be 1:1. This results in equal opportunity for the unsupervised and supervised tasks to affect the weights of the network as a function of the total number of updates. Choosing an update ratio that favors the unsupervised task results in a more computationally expensive training. While, an update ratio that is biased towards the supervised task produces an ASR model that does not improve over a supervised baseline.

The learning rate ratio is biased towards the unsupervised task as compared to the supervised task. Using a learning rate ratio of 1:1 or one that favors the supervised task results in an ASR model that does not improve over a supervised baseline.

Experimental Setup

The experiments use the Librispeech 960-hours dataset as the unsupervised dataset. The supervised dataset is a subset of Librispeech: either 100-hours or 960-hours (full). During training, samples in the dataset that are smaller than 2 seconds or longer than 33 seconds are filtered out. The performance of the trained model is validated on the dev-clean/other datasets of Librispeech and tested on the test-clean/other datasets.

2 Architecture details

Similar to wav2vec 2.0 , the convolutional encoder network consists of a stack of 7 convolutions with kernel size (10,3,3,3,3,2,2)(10,3,3,3,3,2,2) and strides (5,2,2,2,2,2,2)(5,2,2,2,2,2,2) respectively. The number of input and output channels in the convolution is 512512. Additionally, the input audio is normalized in the time dimension before it is passed into the convolutional encoder.

We use two versions of the model, Base and Large. The transformer context network for the Base model is composed of a convolutional relative positional embedding layer with kernel size 128 and group size 16, followed by a stack of 12 transformer layers with 8 heads. The hidden dimension is 768 and the feed-forward network dimension is 3072. Each transformer layer uses layer dropout with probability 0.05 and dropout with probability 0.1. The transformer context network for the Large model uses a stack of 24 transformer layers with 16 heads. The hidden dimension is 1024 and the feed-forward network dimension is 4096. Each transformer layer uses layer dropout with probability 0.2 and dropout with probability 0.1. The linear classifier is trained to output letter-based tokens, which consist of 26 English alphabet letters, augmented with the apostrophe and a word boundary token. The total number of parameters for the Base model is 94.3M and the Large model is 315M. The masking probability is 0.0750.075 for the Base model and 0.0650.065 for the Large model. The number of masked tokens per sample is 1010. The number of negative samples used in the contrastive loss is 100 and the temperature is 0.1. A variation of SpecAugment that uses the same masking procedure as the contrastive loss is used for data augmentation in the ASR task.

3 Training

The model is trained using the Adam optimizer () for both losses with β1=0.9, β2=0.98, ϵ=10−6\beta_{1}=0.9,\,\beta_{2}=0.98,\,\epsilon=10^{-6} and weight decay 0.010.01. The gradient for the convolutional encoder is scaled by 0.10.1 for each of the two losses. The ratio of unsupervised to supervised loss updates is set to 1:1. The learning rate (LR) for the unsupervised loss is 5×10−45\times 10^{-4} and for the supervised loss is 2.5×10−52.5\times 10^{-5} for the Base model, whereas the LR for the unsupervised loss is 3×10−43\times 10^{-4} and for the supervised loss is 2×10−52\times 10^{-5} for Large model when using the 100-hours dataset as the labeled data. The LR for the unsupervised loss is 5×10−45\times 10^{-4} and for the supervised loss is 1×10−41\times 10^{-4} for the Base model when using the 960-hours dataset as the labeled data.

The total number of updates is 500K. The LR for the both losses is warmed up from to their respective values in 20K updates. After the warmup period, the LR of the unsupervised loss ηu\eta_{u} is decayed to 0.1ηu0.1\eta_{u} at the end of training, whereas the LR of the supervised loss is kept constant. SpecAugment in the supervised loss update is activated after the warmup period.

Training is performed on 64 V100 GPUs with a batch size per GPU equal to 87.587.5s of audio for the Base model and on 256 V100 GPUs with a batch size per GPU equal to 4040s of audio for the Large model. The audio samples are batched together such that the total length of the samples does not exceed the batch size. The model is trained using the wav2letter++ toolkit for approximately 4 days.

4 Beam-search decoding and rescoring

Besides reporting word error rate (WER) without a language model (LM), we also perform a one-pass beam-search decoder with a 4-gram word-level LM and further the beam rescoring with a strong word-level Transformer LM . We rely on the beam-search decoder from the wav2letter++ toolkit and follow the procedure from .

Results and Discussion

The single-stage training pipeline is evaluated in a setting where there is a large amount of unlabeled data compared to labeled data.

Table 1 shows word error rates (with and without an LM, see Section 3.4) for the Base model trained on Librispeech 960-hours unlabeled data and 100-hours labeled data. The joint training procedure generates an ASR model that matches the WER of the wav2vec 2.0 Base model on both the test-clean and test-other datasets. Unlike the wav2vec 2.0 model, this model does not include quantization, operates in the continuous space and does not use any unsupervised loss penalty terms during training. Using the two-stage pipeline of wav2vec 2.0 (reproduced in wav2letter++) to train the continuous Base model results in slightly worse ASR performance compared to the quantized wav2vec 2.0 Base model.

Table 2 shows word error rates (with and without an LM, see Section 3.4) for the Large model trained on Librispeech 960-hours unlabeled data and 100-hours labeled data. The joint training procedure generates an ASR model that matches the WER of the wav2vec 2.0 Large model on both the test-clean and test-other datasets.

2 Effect of hyperparameters on downstream task

Table 3 shows the effect of different hyperparameters on the ASR performance of the model trained using the single-stage training method. All models are trained for 500K updates using the Librispeech 960-hours dataset as the unsupervised dataset and the 100-hours dataset as the supervised dataset. The baseline model uses a Lu\mathcal{L}_{u} to Ls\mathcal{L}_{s} update ratio equal to 1:1, Lu\mathcal{L}_{u} to Ls\mathcal{L}_{s} learning rate ratio equal to 20:1 and separate optimizers for each of the two losses. Using a lower Lu\mathcal{L}_{u} to Ls\mathcal{L}_{s} learning rate ratio or using a single optimizer results in a higher WER on the dev-other dataset compared to the baseline. The training pipeline is not sensitive to the update ratio as can be seen by the negligible difference in WER between the models with a Lu\mathcal{L}_{u} to Ls\mathcal{L}_{s} loss update ratio 1:1 and 5:1.

3 Regularization effect on supervised loss

Figure 1 shows a plot of the unsupervised loss Lu\mathcal{L}_{u} and the supervised loss Ls\mathcal{L}_{s} on the train (Librispeech 960-hours) and validation (Librispeech dev-other) datasets as a function of total number of updates for the Base model trained using either joint training or supervised only training. Both models are trained for the same total number of updates, 500K. The supervised loss attains a lower value on the validation dataset and a higher value on the train dataset with joint training in comparison to supervised only training. Furthermore, Table 4 shows that a model trained using joint training achieves lower WER (with and without an LM) compared to a model trained using supervised loss only, even though it has a lower number of updates from this loss. This suggests that our method provides a regularizing effect to the supervised loss.

Related Work

This paper draws upon recent advances in self-supervised contrastive learning . It uses the principle of contrastive learning: similarity between an anchor and positive samples is compared against similarity with negative samples. But, the goal of self-supervised learning is to learn representations that are useful for multiple downstream tasks. Whereas, our method is designed to maximize performance on a single downstream task.

More broadly, our single-stage training method can be linked to semi-supervised learning or self-training methods for ASR. These methods bootstrap an acoustic model (AM) from transcriptions (labeled data), transcribe unlabeled audio with the trained AM (optionally with the help of an LM) and then retrain the AM on the generated pseudo-labels. Self-training methods are complementary to our method and there is potential to combine the two methods.

As our approach addresses both, a contrastive learning task and speech recognition task, this paper is related to the field of multi-task learning . Recent approaches to multi-task learning solve the tasks by minimizing a loss, containing multiple terms, on the same supervised datasets. Whereas, in our method, the unsupervised and supervised losses are minimized on their respective datasets.

Conclusion

Our single-stage training method simplifies the process for learning speech recognition models jointly from labeled and unlabeled data and allows directly optimizing the model on the downstream task. Furthermore, the trained models match the performance of state of the art self-supervised models for speech that use a two-stage pipeline. Finally, we demonstrate that solving the contrastive task provides a regularizing effect on the supervised loss when only using a labeled dataset.

Finally, we would like to thank Alexei Baevski and Michael Auli for helpful discussions regarding wav2vec 2.0.

References