TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech

Andy T. Liu, Shang-Wen Li, Hung-yi Lee

I Introduction

Unlike humans, capable of self-learning through experiences and interactions, current real-world speech applications like automatic speech recognition (ASR) rely heavily on large amounts of human annotations. For the next generation of speech processing systems to exhibit similar cognitive intelligence levels as humans, machines should be designed to learn from unlabeled data as humans do. In the era of big data, self-supervised learning has emerged as an attractive approach to leverage knowledge from many unlabeled data, and are shown effective for improving downstream systems .

In self-supervised learning, an auxiliary task (or pre-training task) is formulated, and models are trained to solve it. While solving the auxiliary task, the network is learning a function that maps input to desired representations. Hence the fundamental tenet of self-supervised learning is the design of an auxiliary task, which allows the model to leverage knowledge from unlabeled data. As such, the formulation of the auxiliary task should be carefully chosen. The task should be challenging enough for the model to learn high-level semantic properties and not be too amiable to exploit low-level shortcuts.

After self-supervised pre-training, learned models could be applied to downstream Speech and Language Processing (SLP) tasks through feature-based speech representation extraction, or fine-tuning as part of the downstream model. Speech representations are compact vectors which aim to capture high-level semantic information from raw speech . Thus, the goal of speech representation learning is to find a transform that maps the input acoustic features into such vectors. When the pre-trained networks are re-used as features, it provides a useful speech representation to reduce classifier complexity, makes high-level information more accessible, and ultimately improves downstream SLP tasks. Besides, speech representations also help transfer learning and adaptation across different data distributions . On the other hand, the fine-tuning approach uses the pre-trained model to initialize a downstream model for supervised training. The parameters of self-supervised learned models are found to be good initialization for ASR .

In this work, we propose TERA: Transformer Encoder Representations from Alteration, where we use alteration on data to pre-train Transformer Encoders . We introduce a total of three types of alteration to form the self-supervised pre-training scheme: 1) time alteration: reconstructing from corrupted blocks of time steps. 2) frequency alteration: reconstructing from missing blocks of frequency bins. 3) magnitude alteration: reconstructing from altered feature magnitudes. These alterations can be applied together or separately in the pre-training process. We apply alteration on data by dynamically sampling through a probabilistic policy to create random alterations. The model acquires information about the content around the corrupted or altered portions, and by reconstructing them, the model learns a more contextualized representation. We illustrated the framework in Fig. 1.

We use the following downstream tasks to evaluate TERA: phoneme classification, keyword spotting, speaker recognition, and automatic speech recognition (ASR). Also, we compare the effectiveness of each alteration method separately and in combination. As a result, we confirm that each of the proposed alteration methods guides the model to learn a distinct aspect of speech: 1) The time alteration effectively enforces a more accurate phoneme prediction, keyword detection, and speech recognition, as it leads the model to learn richer phonetic content. 2) The frequency alteration effectively improves speaker prediction accuracy, as it leads the model to encode speaker identity. 3) The magnitude alteration effectively improves performance for all tasks, as it potentially increases data diversity for pre-training.

Different self-supervised frameworks have been widely studied, in Section II we provide a thorough review. Previous work explored mostly for reconstruction on the temporal axis, for example unidirectional (or autoregressive) reconstruction of magnitude or phase from past frames , or bidirectional reconstruction of a temporal frame from both past and future slices . Our work contrasts with prior work in several ways. Firstly, unlike previous work that only employs reconstruction on the temporal axis, we use reconstruction loss and apply alteration on data along three orthogonal axes, including time, frequency, and magnitude axis. Secondly, most works evaluated their approach with classification tasks only . In contrast, we moved beyond classification and applied our model to ASR. For a comprehensive investigation, we evaluate our method with four downstream tasks. Thirdly, we explore knowledge transfer between pre-trained models and downstream tasks, an under-investigated problem in speech compared to NLP . We leverage two ways to incorporate the pre-trained model with downstream tasks, where most of the previous work only explored one way of transferring their pre-trained models. Fourthly, we study how self-supervised models behave when pre-trained on a different amount of unlabeled data. Surprisingly, we find that methods that learn from time-only masked reconstruction methods can sometimes not benefit from more unlabeled data. This is because of the reconstruction nature, memorizing all the details, including the unnecessary noise. Additionally, we study the effect of pre-training on various features. We explore four different acoustic features in this work, including log Mel, fMLLR, MFCC, and FBANK. We report results primarily on log Mel and fMLLR. The usage of fMLLR, is not explored before for reconstruction-based methods. Also, none of the previous work explores more than one acoustic feature for their method. Our study finds that using different acoustic features in reconstruction-based learning has a significant effect on pre-trained models and is a parameter choice for researchers. Furthermore, we find smaller pre-trained models perform well for feature extraction over larger models, and larger models tend to be more effective for fine-tuning. Finally, we show explicitly that our approach continues to work well in the face of domain mismatch between pre-training and downstream datasets. For reproducibility of our results, we provide our implementation with pre-trained models and evaluation scripts in the S3PRL ToolkitThe S3PRL Toolkit: https://github.com/s3prl/s3prl .

II Related Work

There are two major branches of speech pre-training methods: Contrastive Predictive Coding (CPC) and Reconstruction.

The CPC paper describes a form of unidirectional modeling in the feature space, where the model learns to predict the near future frames in an acoustic sequence while contrasting with frames from other sequences or frames from a more distant time. In wav2vec , the CPC loss is used to pre-train speech representations for speech recognition, and experiment results show self-supervised pre-training improves supervised speech recognition.

II-A2 CPC with Quantization

In the work of vq-wav2vec , the wav2vec approach is incorporated with the well-performing Natural Language Processing (NLP) algorithm – Bidirectional Encoder Representations from Transformers (BERT) . The vq-wav2vec approach learns BERT-like speech representations through a two-stage training pipeline. In a follow-up work of BERT + vq-wav2vec , the pre-trained vq-wav2vec model is directly fine-tuned on transcribed speech using a Connectionist Temporal Classification (CTC) loss instead of feeding the representations into a task-specific model. In wav2vec 2.0 , the vq-wav2vec method is improved to a single-stage training scheme through time masking in the latent space.

II-A3 Improving CPC

In recent research, the CPC loss has also been extended and applied to bidirectional context networks. Bidirectional CPC (Bidir-CPC) representations are learned from bidirectional predictive models. The multi-layer CNN encoder network is shared for both directions, but two autoregressive GRU context networks read encoded observations from the forward and backward contexts. The Modified CPC focuses on improving the CPC approach by introducing several modifications to the original. These modifications include changing the batch normalization to channel-wise normalization, replacing the linear prediction layer with a Transformer layer , and replacing the GRU context network with LSTM cells.

In this work, TERA is compared with several contrastive learning methods, including CPC , wav2vec , vq-wav2vec , BERT + vq-wav2vec , wav2vec 2.0 , Bidir-CPC , and Modified CPC .

II-B Reconstruction Losses

Another recently emerged branch of speech pre-training approach devotes its attention on reconstruction losses.

Primarily inspired by language models (LM) for text, the Autoregressive Predictive Coding (APC) model can be seen as a speech version of LM. The APC approach uses an autoregressive model to encode temporal information of past acoustic sequence; the model then predicts future frames like a recurrent-based LM while conditioning on past frames. In , the APC objective is extended to multi-target training. The new objective predicts not only the future frame conditioning on previous context but also past memories. In VQ-APC , a vector quantization (VQ) layer is used with the APC objective, which imposes a bottleneck and forces the model to learn better representations. In DeCoAR , combining the bidirectionality of ELMo and the reconstruction objective of APC , models were able to learn deep contextualized acoustic representations.

II-B2 Time-only Masked Reconstruction

Largely inspired by the Masked Language Model (MLM) task from BERT and Permutation Language Modeling (PLM) from XLNet , recent work have explored using BERT-style tasks to pre-train speech encoders. These approaches adopt the NLP pre-training technique to continuous speech. In Mockingjay , input frames of speech are masked to zero to pre-train Transformer Encoders. In Audio ALBERT , Mockingjay is modified to have shared parameters across Transformer layers. In , Mockingjay is shown to be effective in defending adversarial black-box attacks. And in , the self-attention of Mockingjay is shown to be meaningful and explainable. On the other hand, TERA can be seen as an improved version of Mockingjay . Using the time alteration alone as pre-training objective reduces TERA to Mockingjay.

In , time-only masked reconstruction following the standard BERT masking policy is employed to pre-train ASR encoders. In , a simpler masking policy is employed, where input features are divided into chunks of four frames, and masking on chunks are applied with a probability of 15%. In Speech-XLNet , models learn by reconstructing from shuffled input speech frame orders rather than masked frames. In , SpecAugment is applied on input frames to pre-train ASR encoders (bi-GRUs). In wav2vec 2.0 , time masking is applied in the latent space. In Non-autoregressive Predictive Coding (NPC) , time masking is introduced through Masked Convolution Blocks, rather than on the input data. In , phoneme posterior vectors are used to train a standard BERT model. The phoneme posterior vectors are output from a supervised acoustic model, which requires CTC loss training over the ground-truth phonemes. Also, in , CTC loss is used along with time-only masked reconstruction training to learn phonetic representations. As and both use phoneme labels for CTC training, they diverge from other works that are fully self-supervised.

II-B3 Learning from Other Reconstruction Losses

Other than autoregressive and time-only masked reconstruction losses, previous works have also explored the reconstruction of different targets or frameworks, including temporal slice estimation, gap estimation, autoencoders, phase prediction, and Markov Models. In Audio2Vec , the model learns through reconstructing a spectrogram slice from past and future slices; this can be seen as a speech version of the NLP Word2Vec variants CBoW (continuous bag-of-words) and skip-gram. The TemporalGap approach learns through estimating the temporal gap between two short audio segments extracted at random from the same audio clip. In , speech representations are learned by applying autoencoding neural networks to speech waveform. Apart from reconstructing spectrograms, in , representations are learned through reconstructing the phase of the short-time Fourier transform from its magnitude. In PASE , a single neural encoder learns to solve multiple self-supervised tasks at once, including reconstruction of waveform, Log power spectrum, MFCC, prosody, and other binary discrimination tasks. The ConvDMM approach learns speech representations with convolutional neural networks and Markov Models. Although The design of the auxiliary task fundamentally decides what the model learns through its reconstruction.

In this work, TERA is compared with several reconstruction learning methods, including APC , VQ-APC , DeCoAR , Mockingjay , Audio ALBERT , SpecAugment , and NPC .

III Proposed Methodology

As illustrated in Fig. 1A, the input acoustic frames (outlined in the red box) and target predicted frames (outlined in the green box) could be any acoustic features, such as MFCC, FBANK, fMLLR, or log Mel. We show a sample of 80-dimensional log Mel feature sequence from the LibriSpeech train-clean-100 subset in Fig. 2A. We denote the entire speech corpus as X\mathcal{X} and the acoustic features of the utterance sampled from X\mathcal{X} as x→\overrightarrow{x}. The length (the number of frames) and the height (the number of frequency bins) of x→\overrightarrow{x} is denoted as LxL_{x} and HxH_{x}, respectively. Below, we introduce how we use different methods to alter the input x→\overrightarrow{x}.

Our model learns bidirectional representations from past and future contexts by altering contiguous segments along the time axis. In time alteration, a certain percentage of input frames are altered during training, and the model attempts to reconstruct the corrupted span from neighboring frames. We randomly select TnumT_{num} amount of starting indexes ITI_{T} without replacement to alter the input utterance. The amount TnumT_{num} is given as the maximum time alteration percentage PTP_{T} normalized by the time alteration width WTW_{T}:

Note that if time alteration width WT=1W_{T}=1, then Tnum=PT×LxT_{num}=P_{T}\times L_{x}. For each starting index location iti_{t} in ITI_{T}, we alter WTW_{T} consecutive frames from iti_{t} according to the following stochastic alteration policy: 1) 80% of the time, we mask all the selected frames to zero. 2) 10% of the time, we replace all with random segments of frames. 3) For the rest 10% of the time, we do nothing and leave the frames in x→\overrightarrow{x} unchanged. The design of Case 3) is to allow the model to receive real inputs during training and addresses the train-test inconsistency problem. This inconsistency problem comes from the absence of alteration during inference time, and the model will only receive acoustic features without alteration.

We illustrate the masking and replacing of frames in Fig. 2B and 2C, respectively. Our time alteration policy is more sophisticated than other time-only masked reconstruction approaches , where they simply mask a percentage with zeroed-out spans, unlike ours that have random and real frames. We set the time alteration width WTW_{T} to 7 frames, which corresponds to 85ms of speech. The time alteration width WTW_{T} then lies in the average phoneme duration range (average phone duration is around 50 ms to 100 ms at usual rates of 10 to 20 phones per second). We set the PTP_{T} percentage of total altered frames to 15%, as suggested in . We allow time alteration blocks to overlap each other, hence resulting in the larger highlighted yellow box in the left of Fig. 2B and 2C. With overlapping, we generate a longer altered span (>WT>W_{T}) and force the model to infer on more global structure rather than a fixed local span (WTW_{T}). The idea behind time alteration is that a model that can predict the partial loss of small speech segments should provide a contextualized understanding of previous and future content. Our experiments show that the proposed time alteration is the essential element that drives models to learn bidirectional understanding, resulting in a substantial improvement compared to models that are not using the time alteration method.

III-A2 Frequency Alteration

Our second pre-training method is frequency alteration. It is largely inspired by SpecAugment proposed for ASR augmentation , and the ASR pre-training scheme proposed in . We randomly mask the values of a block of consecutive frequency bins to zero for all time steps across the input sequence for this alteration. The block of masked frequency is selected by first sampling the width of block WCW_{C} from {0,1,...,WC}\{0,1,...,W_{C}\} uniformly. Then, we sample a frequency index ICI_{C} from {0,1,...,Hx−Wc−1}\{0,1,...,H_{x}-W_{c}-1\}, where HxH_{x} is the number of frequency bins in input sequence x→\overrightarrow{x}. The frequency bins from ICI_{C} to IC+Wc−1I_{C}+W_{c}-1 are those to be masked. Note that the policy will mask none of the frequencies for 1/(WC+1)1/(W_{C}+1) of the time. Thus, from time to time, the model will receive inputs with all of the frequency information. By allowing the model to receive real inputs again addresses the inconsistency between training and inference time.

We illustrate this alteration’s effect on input sequence in Fig. 2D. Unlike the time alteration case, where we sample many blocks for alteration as visualized in Fig. 2B and 2C, we only sample a single block for frequency alteration in each utterance. The reason is that acoustic sequences can be arbitrarily long and temporally smooth , while there are only a limited and fixed number of frequencies HxH_{x}. Hence, we select multiple blocks for time alteration, but only one block for the frequency axis. Following the work of , we set the maximum frequency alteration width WCW_{C} to 16 frequency bins (20% of the 80-dimension feature). The intuition behind frequency alteration is that a model that can predict the partial loss of frequency information should learn a high-level understanding along the frequency axis.

As we will show in our experiments, we find that using frequency alteration provides a more linearly spreadable speaker representation and a stronger speaker recognizer. Surprisingly, encoding speaker information through this objective does not compromise phoneme classification or ASR performance but instead increases performance. This makes TERA not only suitable for tasks that only require speaker information (e.g. speaker recognition) and tasks that only require phonetic information (e.g. speech recognition), but also beneficial for tasks that requires both speaker and phonetic information at the same time (e.g. voice conversion ).

III-A3 Magnitude Alteration

We introduce the third method, magnitude alteration, by applying sampled Gaussian noise to augment input sequences’ magnitude with a probability PNP_{N}. For PNP_{N} of the time, we sample a random magnitude matrix z→\overrightarrow{z} of dimensions LxL_{x} and HxH_{x}, which has the same shape as x→\overrightarrow{x}. Each element in z→\overrightarrow{z} is sampled from the normal distribution N\mathcal{N} with zero mean and 0.2 variance. We then add z→\overrightarrow{z} on top of the real frames of x→\overrightarrow{x}. We show the effect of magnitude alteration in Fig. 2E, where we apply magnitude alteration to the original utterances x→\overrightarrow{x}. By altering input magnitude, we potentially increase the amount of pre-training data (which is similar to the idea of data augmentation ). Additionally, magnitude alteration offers another variation to all the ‘mask to zero‘ cases described in Section III-A1 (time alteration) and Section III-A2 (frequency alteration). We illustrate this new ‘mask to noise‘ variation in Fig. 2F, where the selected blocks of time and frequency are now with random magnitudes instead of zeros. Empirically, altering magnitude provides a performance benefit for all downstream tasks, thanks to the increased input data variety. Also, the gain from magnitude alteration is additive to that of other alteration techniques.

III-B Pre-training TERA

We use the three alteration techniques to formulate the TERA self-supervised pre-training task, where the model is required to minimize the reconstruction error of acoustic features given altered frames as input. The proposed three alteration methods can be used separately or used together as a mixture, as shown in Fig. 2. The complete pre-training process of TERA is described in Algorithm 1, where we show how three alteration on data are deployed through the proposed stochastic policy. The stochastic policy dynamically samples random pattern based on the applied alterations every time we feed an input sequence to the model. We denote the altered input as x^\hat{x}. In the case where multiple alterations are applied on the same input, x→\overrightarrow{x} is used again to represent carried-on alterations. After input alteration, we feed x^\hat{x} into the Transformer Encoders TencT_{enc} and the prediction network PnetP_{net}. The architecture of PnetP_{net} is consist of a 2-layer feed-forward network with a hidden size of 768. For TencT_{enc}, we use Transformer Encoders with a hidden size of 768, number of self-attention heads as 12, dropout rate of 0.1, and the intermediate feed-forward layer hidden size as 3072. We primarily report results on three model sizes: base (Layer=3, Parameters=21.3M), medium (Layer=6, Parameters=42.6M), and large (Layer=12, Parameters=85.1M). The Transformer Encoders TencT_{enc} and the prediction network PnetP_{net} are connected to reconstruct x→\overrightarrow{x} from x^\hat{x}.

L1 reconstruction loss is then computed between input x→\overrightarrow{x} and network output from PnetP_{net} to update the network parameters θtenc\theta_{tenc} and θpnet\theta_{pnet}. We use gradient descent training with mini-batches of size 3232 to find model parameters that minimize the L1 loss under the self-supervised pre-training task. We employ the AdamW optimizer for updating model parameters, where the learning rate is warmed up over the first 7%7\% of total training steps TstepsT_{steps} to a peak value of 2e−42e^{-4} and then linearly decayed. Our pre-training setup can be accommodated in a single 1080Ti GPU with 11GB of memory. Computational efficiency allows interested parties to easily train our model with their data without massive computational resources. The models are trained with fixed total training steps TstepsT_{steps} (details in Section IV-A). After pre-training, the parameters θtenc\theta_{tenc} of the Transformer Encoders TencT_{enc} are retained for downstream tasks, while the prediction network PnetP_{net} is discarded, as illustrated in Figure 1.

In this work, our models’ input is an 80-dimensional log Mel spectrogram if not specified otherwise. We also explore pre-training with other acoustic features, including 39-dimensional MFCC, 80-dimensional FBANK, and 40-dimensional fMLLR. We extract all of the features with a window of 25 ms and a stride of 10 ms. We apply per-utterance CMVN (cepstral mean and variance normalization) to the features. We set the total pre-training steps TstepsT_{steps} of TERA as 200k and 1M for 100 hours and 960 hours of pre-training data, respectively.

III-C Incorporating with Downstream Tasks

We investigate different ways to incorporate the learned TERA model to downstream tasks.

We first use the weighted sum technique to investigate which layer is best for representation extraction. We use a learnable weighted sum to integrate hidden states from all layers of the self-supervised model when training with downstream tasks, similar to the ELMO approach. We then analyze the learned weights and find that for TERA, APC , and vq-wav2vec , the weighted sum always favors the last layer the most (we investigate with phoneme and speaker classification tasks). Hence in this work, we extract representations from TERA’s deepest layer, which is essentially the hidden states of the last Transformer Encoder layer. The extracted representation is fed to downstream classifiers as input and replacing surface features. Parameters of TERA are frozen when training downstream tasks in this approach. In later experiments, we use this approach if not specified otherwise.

III-C2 Fine-tuning

The second approach is to fine-tune the TERA model with downstream models. Here the output of TERA is connected to a downstream model of any kind, as illustrated in Fig. 1. We then update the pre-trained TERA together with random initialized downstream models. We denote this approach as fine-tune to distinguish from representation extraction in later experimental tables and figures.

IV Experimental Setup

We use four downstream tasks for evaluation: phoneme classification, speaker recognition, keyword spotting, and speech recognition. We use publicly known and available settings for our downstream tasks to link our works with previous ones and allow easier comparison. For all experiments, we train with fixed random seeds for consistency.

We use a total of three datasets. We consider three subsets of LibriSpeech for the pre-training of TERA: the train-clean-100, the train-clean-360, and the train-other-500 subset. The three subsets add up to a total of 960 hours of data. We also use the LibriSpeech dataset for downstream tasks of phone classification, speaker recognition, and speech recognition. The train-clean-100 subset is used in the phoneme classification and speaker recognition task. For ASR, we use the train-clean-100 as labeled data, and the dev-clean and test-clean subset are used for evaluation. We use the TIMIT dataset to evaluate the transferability of pre-trained models. We do not use the TIMIT dataset for pre-training, but we use it for downstream phoneme classification and ASR tasks. We consider two subsets of TIMIT: the training set and the complete test set. As is usually done, we derive the development set and the core test set from the complete test set, according to the Kaldi TIMIT recipe and . We report results by training downstream tasks on the training set and evaluating with the development and core test sets. For keyword spotting, we use the Speech Commands dataset. We consider two subsets of Speech Commands: the training set and the test set. We derive the development set from the validation list provided in Speech Commands . We report results by training keyword spotting models on the training set, and evaluating on the development set and test set. The Speech Commands dataset is also not used for pre-training.

IV-B Phoneme Classification Setup

We measure the frame-wise phoneme prediction performance with classifiers trained on top of representations. Following previous work , we adapt the common setup using 41 possible phoneme classes and the train-clean-100 subset of LibriSpeech . To make the comparison as fair as possible, we use aligned phoneme labels and train/test split provided in the CPC and Modified CPC literature. We derive 10% of the data from the training split and use it as the development set. Following the previous work , we utilize linear classifiers to measure the linear separability of phonemes. We denote these classifiers as linear. Additionally, following the work of Modified CPC , we also report results from performing linear classification with a concatenation of 8 windows, which matches the average length of a phoneme. We denote this type of classifier setting as linear concat. As not all the information encoded is linearly accessible, in addition to measuring linear separability, we also use classifiers with one single hidden layer, following the same settings as in . We denote such setting as 1-hidden.

IV-B2 Phoneme Classification on TIMIT

For TIMIT, we obtain a phoneme label for each frame from the manual phoneme transcriptions provided in TIMIT. We measure accuracy after mapping the 48 phonemes to the smaller set of 39 phoneme classes, as is customary for TIMIT since (the same applies to ASR on TIMIT). Following the classifier settings for LibriSpeech, we also use the linear, linear concat, and 1-hidden classifiers on TIMIT. For both LibriSpeech and TIMIT, we report phoneme classification accuracy (%) on the test set. We investigate the domain shift issue with this setup, where models are pre-trained on LibriSpeech then used on TIMIT for the downstream task.

IV-C Keyword Spotting Setup

We evaluate representations on the keyword spotting task. We follow the standard setup , where the keyword spotting task is a 12-class classification problem. The test set provides equal numbers of instances for each class; hence class accounts for approximately 8.3%. To probe representations, we did not use complex models for this task, as performance may saturate. Instead, we use a two hidden layer feed-forward classifier, with mean pooling overtime applied just before the output layer. We report keyword classification accuracy (%) on the test set. This setup allows us to study the domain shift issue, where models are pre-trained on LibriSpeech then used on Speech Commands for the detection task.

IV-D Speaker Classification Setup

We evaluate representations on speaker prediction tasks. Following the common experiment setting in , we adopt the LibriSpeech train-clean-100 subset, which consists of 251 speakers. We use the same train/test split as provided in . We evaluate the pre-trained models with two types of speaker classification tasks, frame-wise linear classification, and utterance-wise linear classification. For frame-wise speaker classification, the classifier predicts the speaker for each input frame. We denote this experiment setting as speaker frame. As for utterance-wise classification, we average the representation of each utterance over time. Then the classifier predicts speaker identity conditioning on the averaged vector. We denote this setting as speaker utter. For both settings, we employ a simple linear classification model. In , they also investigate these two speaker classification settings. In general, the speaker frame classification task is more difficult than speaker utter, however speaker utter is a more common scenario for speaker classification. Also, good results in speaker frame always imply good results in speaker utter, but not the other way around. Hence we report both speaker frame and speaker utter for completeness. For both tasks, we report speaker classification accuracy (%) on the test set.

IV-E Hybrid DNN/HMM ASR Setup

We evaluate the performance of ASR models on top of representations. We employ the Hybrid DNN/HMM ASR modeling implemented with the PyTorch-Kaldi toolkit. We investigate two types of DNN settings, MLP and liGRU. MLP is a simple single-layer multilayer perceptron model. liGRU is a 5-layer light-gated recurrent unit followed by 2-layers of fully-connected network. In first-pass decoding, we use a 4-gram language model with a beam-search algorithm. The WER score is computed with the NIST SCTK scoring toolkit . For lattice rescoring, we use the Kaldi ConstArpaLm format language model. The decoding and rescoring scripts are inherited from the Kaldi s5 recipe .

When we utilize TERA as a representation extractor, we feed the output of TERA to liGRU or MLP and freeze the parameters of TERA during training. As for fine-tuning TERA, we update the TERA model with liGRU or MLP as part of the DNN component in the hybrid ASR framework. For our ASR experiments, we pre-train TERA on 40-dimensional fMLLR features instead of 80-dimensional log Mel features since the PyTorch-Kaldi ASR toolkit works best with fMLLR inputs. To highlight the effect of pre-training, we use a limited amount of labeled data, i.e., the train-clean-100 subset, for supervised ASR training. Hyperparameters are tuned on the dev-clean subset, and testing results measured from the test-clean subset are reported. We report ASR modeling results of LibriSpeech in terms of WER. In addition to evaluating with LibriSpeech, we also benchmark ASR results with TIMIT, where we measure PER. With this setup, we investigate the domain shift issue, where models are pre-trained using data in the domains different than the downstream tasks.

IV-F Training Downstream Tasks

For both representation extraction and fine-tuning, we use the AdamW optimizer to update models when training with downstream tasks. We use a learning rate of 2e−42e^{-4} with a batch size of 1616. When applying representation extraction for ASR tasks, we use the RMSPROP optimizer with a learning rate of 2e−42e^{-4}. We half the learning rate every epoch if the development set error does not drop more than a threshold of 0.0010.001. We use a batch size of 1616, and update for 24 epochs. We use the same setting as above for fine-tuning on ASR, except we use a different learning rate for each TERA model. Learning rate is set to 2e−42e^{-4} when we fine-tune base, 1e−41e^{-4} for medium, and 5e−55e^{-5} for large. Empirically, we find that larger models require a lower learning rate during fine-tuning. Hence, we half the learning rate every time the model depth is doubled; this helps stabilize the entire fine-tuning process.

During downstream fine-tuning, we also apply the SpecAugment LD policy. During SpecAugment, the chosen (and possibly overlapping) time and frequency blocks are zeroed out. We omit the use of time warping as masking provides enough regularization. We find SpecAugment is additive to the proposed pre-training approach, as it delays overfitting and improves the final accuracy numbers in the downstream tasks.

V Results

In Section V-A, we study the effect of applying different combinations of time, frequency, and magnitude alteration. In Section V-B, we compare TERA with recent self-supervised representation methods on downstream tasks. In Section V-C, we study multiple aspects of the TERA pre-training process, including fine-tuning with downstream models, the amount of unlabeled data, the effect of pre-training on various features, various model sizes, and different masking policies. In Section V-D, we compare TERA with approaches where we train ASR models on frozen speech representations. In Section V-E, TERA is compared with approaches that fine-tune their pre-trained models. In Section V-F, we demonstrate transferring TERA representations from one dataset to another by presenting ASR PER on TIMIT .

With three types of alterations, there are a total of seven combinations, namely ”time”, ”freq”, ”mag”, ”time+freq”, ”time+mag”, ”freq+mag”, and ”time+freq+mag”. We pre-train TERA with all seven combinations of alterations on 100 hours of LibriSpeech, then measure the learned representations with downstream tasks. The results are presented in Figure 3. We use TERA for representation extraction (from the last layer), i.e., TERA parameters were frozen when adapting downstream models. We split our methods into two categories: one with the time alteration that allows the model to encode bidirectional information (denoted in a solid color). The second category is the models that don’t include the time alteration method (denoted in dotted color). The conclusions made in the following sections are statistically significant, where we measured p-values with statistical significance tests (paired samples t-test or Fisher’s exact test). Overall, applying all alterations ”time+freq+mag” achieves the best performance on average (with an averaged p-value of 0.0492 across nine tasks when compared to ”time”). Now, we discuss the effect of each alteration in detail.

We discuss the effect of time alteration for each downstream task. For phoneme classification, the models with time alteration (in solid color) perform reasonably well, while the models without time alteration (in dotted color) perform relatively poorly. The results of seen data (LibriSpeech) and unseen data (TIMIT) agree with each other and yield similar trends. For keyword spotting, the ”time” model achieves the best performance. Adding other alterations compromise the performance. The models with time alteration perform better than the models that do not have time alteration. This observation suggests that bidirectional understanding is an essential aspect of detection tasks. For speaker recognition, the existence of time alteration does not compromise the performance. With time alteration alone, the ”time” model achieves slightly lower but satisfactory performance. From the above discussion, we thus deem the time alteration to be indispensable. We speculate that although Transformer Encoders are bidirectional due to their multi-head self-attention, without the time alteration, the network fails to encode proper context and yield sub-optimal performance. However, through alteration on the time axis, the model establishes a bidirectional understanding of the audio and gives better results.

V-A2 The Effect of Frequency Alteration

For the phoneme classification tasks, adding the frequency alteration during pre-training is effective in boosting performance. The ”time+freq” and ”time+freq+mag” model improves over the ”time” model for both seen and unseen data. For keyword spotting, adding the frequency alteration does not help. For speaker recognition, we find that the model can achieve strong speaker classification performance by employing the frequency alteration. The ”freq” model, with frequency alteration alone, is enough to achieve high performance. The ”mag” and ”time” models, without frequency alteration, have lower performance. By employing the frequency alteration, models benefit in the phoneme classification and speaker recognition performance, which is surprising and counter-intuitive. It is well known that a robust phonetic system should be speaker invariant . We surmise this phenomenon results from that TERA representations provide more accessibility to phonetic and speaker information and maintain the separability between the two types of information. Downstream models can efficiently learn to extract task-specific information. To sum up this section, the frequency alteration effectively encodes speaker identity while also fine grinding the phonetic quality.

V-A3 The Effect of Magnitude Alteration

For the phoneme classification tasks, adding the magnitude alteration is effective in boosting performance. The ”time+freq+mag” model improves over the ”time+freq” model for both seen and unseen data. For keyword spotting, adding the magnitude alteration improves most cases but does not improve over time alteration alone. For speaker recognition, adding the magnitude alteration improves performance for all cases. We observe that by applying the magnitude alteration during pre-training, the model learns to be robust to various inputs. As a result, we improve downstream tasks performance by adding the magnitude alteration.

V-B Comparison of Recent Speech Representation Approaches

In this section, we compare TERA with recent speech representation learning methods through evaluating with downstream tasks. We select eight different methods, which cover a wide range of representation learning techniques. We list these methods and their details in Table I. We organize Table I by first grouping methods with similar learning styles, then by their publication date in the order of first to last. We use publicly available pre-trained weights or pre-training scripts provided by their original authors for all of the listed models. Note that all models are pre-trained on LibriSpeech except for CPC , which is pre-trained on LibriLight . All of the models and baseline features have a stride of 10 ms, with only one exception where wav2vec 2.0 has a downsample rate of 320, which generates a representation every 20 ms. Interestingly, all of the methods that employ reconstruction loss in their learning objective use an 80-dim log Mel input feature. We also use 80-dim log Mel for consistency and fair comparison. On the other hand, all methods that employ contrastive loss in their learning objective use raw waveform as an input feature. In Table I, we also list the details of baseline features. We present the performance of different representations on downstream tasks in Table II. The pre-trained models are all used for representation extraction (from the last layer), i.e., we freeze parameters when adopting the downstream model and feed representations to the downstream model as input. In the last column, we average the accuracy of all tasks across each row. The above results can be reproduced with the S3PRL toolkit1, where all the self-supervised models and downstream tasks are available with easy-to-use setups.

For phoneme classification on seen data (LibriSpeech), TERA outperforms other methods. For phoneme classification on unseen data (TIMIT), TERA outperforms other methods except for CPC on the linear concat classifier. Some representations encounter a more significant degradation in performance for unseen data. On the other hand, TERA is not affected by the domain change. In particular, we find that adding the TERA magnitude alteration helps performance for unseen data. As a result, by adding magnitude alteration, we increase accuracy for phoneme classification on TIMIT. We observe that all representations benefit from concatenating neighboring frames, as their performance on linear concat is better than linear. The linear concat classifier provided neighboring information through the concatenated frames. Several representations benefit more from linear concat than others, depending on how much temporal information the single feature frame already encodes. We also observe that all representations benefit from a deeper downstream model, as their performance on 1-hidden is better than linear. The 1-hidden classifier is able to extract more information than linear classifiers. The wav2vec 2.0 method achieves the lowest accuracy for phoneme classification, which contrasts its high performance when used as an initialization for ASR encoder fine-tuning . The under-performing of wav2vec 2.0 here suggests that although large and deep models are suitable for downstream initialization and fine-tuning, they may not be a good choice for feature extraction. Our later experiment shows that TERA-large under-perform TERA-base for feature extraction, but TERA-large outperforms TERA-base for ASR fine-tuning.

V-B2 Keyword Spotting Results

For keyword spotting, all representations achieve a good score. However, wav2vec 2.0 and NPC are not able to surpass the performance of baseline features. The wav2vec 2.0 method achieves the lowest score, which again suggests that pre-trained models may work well for fine-tuning, but it does not always imply that they can generate meaningful representations. The NPC method and Mockingjay achieve comparable performance, where Mockingjay barely exceeds the performance of MFCC. By adding frequency alteration and magnitude alteration to Mockingjay, TERA can boost the performance by 4.51% over Mockingjay. The best performing representation on keyword spotting is vq-wav2vec, which is surprising as it does not perform as well as the others on phoneme classification and speaker recognition. CPC achieves comparable performance with vq-wav2vec, suggesting that contrastive learning over time may be the key for detection tasks. Other than CPC and vq-wav2vec, VQ-APC and TERA achieve high scores on this task too.

V-B3 Speaker Recognition Results

For the speaker recognition task, TERA achieves the highest speaker classification accuracy for both the frame and utter setting. We can observe a discriminative comparison under the frame setting. Methods that use autoregressive, contrastive, vector-quantization have a lower frame-wise speaker recognition accuracy. The autoregressive prediction method focuses more on the local dependencies between time steps. Also, contrastive learning discriminates input from a distant time, and vector-quantization is a bottleneck that limits information flow over the network. The above techniques seem to have encouraged the model not to encode speaker information in each frame, resulting in poor frame classification performance. Although the NPC technique uses time-masked reconstruction like Mockingjay and Audio ALBERT, it performs poorly for the frame setting. We speculate that this is because Masked Convolution Blocks’ masking is fixed and designed to focus on local dependencies. As a result, each representation doesn’t preserve speaker information. On the other hand, the time mask of Mockingjay, Audio ALBERT, and TERA is applied through dynamic masking , where masks may overlap each other, creating various time mask lengths. The overlapped masks encourage the model to also focus on global information. As speaker characteristics tend to persist across time, this thus preserves the speaker information. For the utter setting of speaker recognition, all of the representations yield satisfactory results except for vq-wav2vec. However, remember that vq-wav2vec achieves the highest score on keyword spotting. We conclude that there is a trade-off for some representation learning methods, vq-wav2vec is good on keyword spotting but lacks speaker information.

V-C Analysis on TERA

To better understand how TERA derives representation from speech, we study various aspects of TERA’s pre-training. We first investigate how TERA can benefit from fine-tuning, where we present results in Table III. Next, we investigate the effect of increasing the amount of pre-training data. Also, we study the effect of learning with various acoustic features. Furthermore, we investigate how the network depth (number of layers) affects TERA’s performance. Finally, we study the difference between directly applying SpecAugment and the TERA time and frequency alteration. We intentionally did not apply the magnitude alteration here to rule out the variation. We visualize results in Figure 4. We denote models trained on 100 hours of LibriSpeech in red and denote models trained on 960 hours of LibriSpeech in blue. All TERA models are pre-trained with time and frequency alteration. We freeze all models for representation extraction from the last layer.

In Table III, we show the effect of fine-tuning the pre-trained TERA with downstream tasks. In the last column, we average the accuracy of all tasks across each row. We fine-tune TERA ”time+freq+mag”, the best performing TERA-base model, which is pre-trained on 960 hours of LibriSpeech. In the first row of Table III, we include the results from the last row of Table II of the frozen TERA ”time+freq+mag” model to link the two tables. With fine-tuning, we see considerable performance increases for all tasks. By adding SpecAugment during fine-tune, we improve performance for some tasks. However, the improvement of adding SpecAugment during fine-tuning is limited because the pre-trained weights are good initialization for the tasks. Empirically, we find that when compared to training downstream classifiers on frozen TERA, fine-tuning TERA with downstream tasks dramatically reduces the number of training steps for the downstream models to converge. We also observe that fine-tuning for phoneme classification on LibriSpeech brings a more significant improvement than on TIMIT. The reason is that there are more labeled data for LibriSpeech. It is worth noticing that fine-tuning with limited label data (TIMIT) still benefits performance, where we observe no sign of overfitting.

We train the same TERA network (with randomly initialized parameters) on downstream tasks, where we distort the input features with SpecAugment to prevent overfitting. The above setting serves as a baseline for fine-tuning pre-trained TERA, and we denote it as random init + SpecAug in Table III. Overall, the downstream models trained with randomly initialized TERA and SpecAugment either underperform or perform poorly. The baseline random init + SpecAug results of phoneme classification on LibriSpeech are close to the pre-trained models because the amount of label is sufficient. However, there is a performance gap between baseline results and pre-trained models on phoneme classification for TIMIT. The reason is the limited amount of label data. The baseline random init + SpecAug overfits for the keyword spotting and speaker recognition task. Note that the reported accuracy is from the test set; therefore, a low testing score means severe overfitting on the training set. It does not indicate that the network is untrainable. We conclude that without pre-training, model architecture alone provides no benefit.

V-C2 Pre-training on More Data

Unsurprisingly, pre-training TERA on more data increases the performance of all downstream tasks. In Figure 4, base 960hr improves significantly over base 100hr. Presumably, all pre-trained models should benefit from more unlabeled data. However, we find that this is not the case for all methods. In particular, we find that models trained from time-only masked reconstruction (See Table I) may not improve by training on more data, especially when the added data is unclean. To be more specific, we find that Mockingjay and NPC got worse performance when we add train-other-500 to the pre-training data. The Mockingjay method uses dynamic masking on data along the time-axis. The NPC method uses a mask in its convolution module to achieve masking on the time-axis. These methods use the idea of reconstructing a masked time span from surrounding context, which is unlike other pre-training methods such as the line of work of APC that uses autoregressive prediction and the line of work of CPC that uses contrastive learning. The time-only masked reconstruction process forces the model to remember all the information of speech, including speaker characteristics and other noise. However, with the regularization of frequency masking, the TERA representations do not exhibit this weakness.

This observation is presented in Figure 5, where NPC , Mockingjay , and TERA ”time+freq” are pre-trained with increasing amounts of unlabeled data. We evaluate the learned representations on downstream tasks. For all tasks and all types of classifiers, both NPC and Mockingjay got worse performance when we add the train-other-500 subset for pre-train. There is only one exception with NPC, where it improves on the keyword spotting task by pre-training on more data. On the other hand, TERA can benefit from training on noisy data. As a result, TERA improves for all tasks by pre-training on more data.

We want to point out that pre-training on more data has not been explored in NPC and Mockingjay, where they only report results with models pre-trained on 360 hours of LibriSpeech. We believe this is a unique phenomenon for methods that learn from the time-only masked reconstruction. To the best of our knowledge, only NPC and Mockingjay examine this kind of behavior. Note that we pre-train Mockingjay and TERA with the same implementation and setting. The only difference is by adding the frequency masking.

V-C3 Learning with Different Acoustic Features

In Figure 4, we pre-trained TERA with MFCC and FBANK as both input and output targets, instead of log Mel. The setting of these acoustic features is identical to the ones listed in Table I. For phoneme classification on both LibriSpeech and TIMIT, pre-training on log Mel outperforms the other two, while pre-training on FBANK is better than MFCC. By using a more primitive feature, the model can preserve more phoneme information. However, the results are opposite for keyword spotting, where models pre-trained on MFCC slightly outperforms FBANK, and models pre-trained on FBANK slightly exceeds log Mel. Learning from all features successfully preserves the speaker information, with the model pre-trained on log Mel achieving the highest score for speaker recognition. In general, pre-training with log Mel features results in the best performing model. This aligns with previous work (See Table I) that also uses 80-dim log Mel.

V-C4 Effect of Different Network Depth

We present classification accuracy of three TERA models, including base (3-layer), medium (6-layer), large (12-layer) in Figure 4. We observe that performance decay for all tasks as model depth increases. We conclude that a small model is sufficient to solve the proposed pre-training task, which is reasonable as a lot of work also uses 3-layer models (See Table I). Note that here TERA is used for representation extraction and not fine-tune. In our later ASR experiments, we find that large models outperform small models during fine-tuning. The above discovery aligns with our previous finding with wav2vec 2.0, that large models are excellent for fine-tuning but not representation extraction. On the other hand, small models are adequate for feature extraction than large models.

V-C5 Comparing TERA with SpecAugment

SpecAugment is a regularization technique for improving ASR training . It shares a similar ideology with our time and frequency alteration. We consider the LD Policy of SpecAugment, which is the best performing policy described in the paper. Following the notations in SpecAugment , the LD Policy has the following parameters: time mask parameter T=100 (maximum length of the consecutive time mask), frequency mask parameter F=9 (maximum length of the consecutive frequency mask), number of time masks mT=2 (amount of consecutive mask blocks in time), and number of frequency masks mF=2 (amount of consecutive mask blocks in frequency). Here we first point out some key differences between the proposed method and SpecAugment , then we present the experiment results of comparing TERA with SpecAugment.

The difference in time mask length. SpecAugment uses a longer time mask length of up to 100 compared to TERA’s length of 7. In SpecAugment, the time mask length applied is chosen from a uniform distribution from 0 to the maximum consecutive length (T) of 100. The considerable variation of time mask length introduces some problems during pre-training. Empirically, we find that pre-training loss of reconstruction from SpecAugment is volatile and not stable. Also, in our experiments, we find that the large variation of time mask length makes it hard for the pre-trained model to encode phonetic information.

The difference in number of time masks. SpecAugment uses two consecutive mask blocks (mT=2) for each input. A fixed amount of consecutive mask blocks is employed. Whereas TERA determines the number of mask blocks by the maximum time alteration percentage PT=15%P_{T}=15\%. The amount of mask blocks TnumT_{num} is then given as Tnum=⌊PT×Lx÷WT⌉T_{num}=\lfloor P_{T}\times L_{x}\div W_{T}\rceil, where LxL_{x} is the length of input. Simply put, the number of masking blocks will change according to different input lengths. The number of time masks will vastly affect the model’s training for long utterances. By fixing the number of consecutive mask blocks, SpecAugment does not utilize the advantage of longer training samples.

The difference in masking policy. The SpecAugment masks selected time blocks to zero. In TERA, following the idea of BERT , a more sophisticated policy is applied. There are three cases for the selected time blocks, 1) mask to zero, 2) replace with random time blocks, and 3) do nothing. Case 2) helps TERA learn the order of speech and not always reconstruct it from zero, case 3) helps TERA ease the training and testing inconsistency problem. Overall, TERA uses a more advanced time alteration policy, and SpecAugment uses a simple mask-to-zero method.

Comparing experimentally. To compare with SpecAugment experimentally, we apply the SpecAugment LD Policy to self-supervised speech training. We use the same pre-training settings (960 hours) and network architecture for both TERA and SpecAugment (3-layer Transformer Encoders TencT_{enc} and the 2-layer prediction network PnetP_{net}), with only the difference between masking policies discussed above. We show the results in dark blue and blue dotted color in Figure 4. TERA base 960hr largely outperforms SpecAug on phoneme classification for both LibriSpeech and TIMIT, and on the keyword spotting task. Also, TERA wins over SpecAugment on the speaker recognition task. We conclude that SpecAugment is suitable for regularizing ASR training but not self-supervised learning. On the other hand, TERA is more effective for self-supervised representation learning.

V-D Speech Representations for ASR

We further apply TERA to speech recognition tasks. In Table IV, we list results of TERA and recent work in terms of WER, where all ASR models are trained on top of frozen representations without fine-tuning. All TERA models use a combination of ”time+freq+mag” alteration as the auxiliary objective. All of the works report WER and LM rescored WER (denoted as Rescore) on the test-clean subset of LibriSpeech . All methods are pre-trained with 960 hours of LibriSpeech, except for one experiment setup in where it uses 8000 hours of data for pre-training. All of the work use only the train-clean-100 for downstream adaption, except for the setup in and that use 96 hours and 960 hours of label, respectively. We use the decoding and rescoring setup described in Section IV-E for all liGRU + TERA and its variation, as well as liGRU with baseline features. Beam-search decoding with a 4-gram language model is used for Bidir-CPC . For wav2vec-large and DeCoAR , a 4-gram LM is used in the first-pass decoding.

First, we observe that the model sizes of TERA (i.e., base, medium and large) have little influence on the ASR performance when TERA is used as an extractor for speech representation. This observation aligns with our previous discovery on other downstream tasks. The representations from the base model are sufficient to improve supervised ASR. Since all the cited models use different LM setups, it is hard to conclude the WER comparison in Table IV. However, we cite other work’s performance to show that the WER achieved in this work is well within the expected range. We also investigate three baseline features of MFCC, FBANK, and fMLLR. We use an identical ASR framework and setting of TERA representations for the three features. Our results suggest that TERA yields constant improvement over surface features in the same ASR framework.

V-E Speech Pre-training for ASR Comparison

In this section, we compare the results of fine-tuning various pre-trained models for ASR. All TERA models use a combination of ”time+freq+mag” alteration as the auxiliary objective. We summarize the results from previous literature as well as fine-tuning TERA with liGRU or MLP framework in Table V. We also list results from recent literature, where all results are from fine-tuning the pre-trained model as an ASR encoder. Similar to the previous section, we report WER and LM rescored WER on the test-clean subset of LibriSpeech . The first-pass decoding and LM rescoring setting are described in Section IV-E. All the methods investigated here were pre-trained with 960 hours and use 100 hours of labels, except for Masked Pre-trained Encoders trained with 360 hours of labels. The wav2vec 2.0 uses a Transformer language model with beam search size of 500 for decoding. The vq-wav2vec uses a 4-gram LM during first-pass decoding, and Masked Pre-trained Encoders adopt beam search and RNN LM with CTC decoding. When fine-tuning TERA with liGRU models, performance roughly correlates with the depth of TERA, and the large TERA achieved the best WER. Remember that large models (wav2vec 2.0, TERA-large, etc) did not perform well for feature extraction. However, they are effective for ASR fine-tuning. By comparing the liGRU results in Table IV and Table V, we see that fine-tuning TERA consistently outperforms the case when TERA is simply used for extraction of speech representation. The ASR model adopting base TERA improves from 6.01% to 5.84%, the medium TERA improves from 6.05% to 5.90%, and the large TERA from 6.01% to 5.80%.

Here we also cite other work’s performance to show that the WER achieved for the proposed method is within the expected range. The wav2vec 2.0 large model achieves a high score of 2.3%. However, we argue that it is consists of seven convolution blocks plus 24 transformer blocks. In contrast, our base model contains only a 3-layer Transformer Encoder layers, which brings substantial low-footprint benefits. The massive wav2vec 2.0 model needs to be trained on 128 V100 GPUs, where the TERA model can be trained on a single GPU. The discrete BERT + vq-wav2vec achieves a high score of 4.5%. However, we argue that the two-step pre-training is computation-intensive during model training. First, a discrete vocabulary of the data is learned from vq-wav2vec , and then in the second step a standard BERT is trained on these discrete units. Also, the discrete BERT + vq-wav2vec is built by stacking a standard BERT model of 12 Transformer Encoder layers on top of vq-wav2vec , which consists of an 8-layer encoder network and a 12-layer aggregator network (or context network, as described in ). Our small encoder architecture (3-layer) benefits from less computational cost and can run on edge devices during inference for downstream tasks.

We also fine-tune TERA with MLP models, and we find a similar trend but sometimes higher WER compared to TERA with liGRU. Using a deeper model with MLP gives performance benefit, and large achieves the best WER among the MLP models. The reason is that the simple architecture of MLP can benefit from a deeper TERA model. Comparing MLP with liGRU, MLP achieved superior performance than liGRU on the medium model, and similar performance for the rest of the model size. Although in general MLP outperforms liGRU, however MLP has the advantage of a fast training and inference time, thanks to the absence of recurrent units. Additionally, the parameters of the 1-layer MLP is significantly less than the 5-layer liGRU models. To conclude this section for our proposed method, using a deeper model increases ASR performance during fine-tuning.

V-F Transferring to TIMIT

We then explore how the mismatch of domains between pre-training and downstream tasks affects performance. For the exploration, we pre-train TERA with LibriSpeech , and apply the resulting networks to the supervised TIMIT ASR task. The same Hybrid ASR setting and framework described above for LibriSpeech ASR are used, except that we use a learning rate of 4e−44e^{-4} and a batch size of 88. Testing results of TERA and another self-supervised learning technique, wav2vec , are summarized in Table VI in terms of PER. We also list the results of strong supervised systems . All of the TERA models use a combination of ”time+freq+mag” alteration as the auxiliary objective, and are pre-trained with various amounts of data. As expected, pre-training on a larger amount of data gives performance benefit, and we achieved the best WER (14.5%) with 960 hours of pre-training data. We find that for TIMIT ASR as the downstream task, fine-tuning is not helpful, and extracting speech representations from the last layer provides the best performance. The reason is likely because there is not enough labeled data in TIMIT, which aligns with our discovery in Section V-C1 for other downstream tasks. Also, there is no significant gain when extracting features from a larger model medium, which aligns with our previous discussion that smaller models are better for feature extraction.

VI Conclusion

We propose a novel self-supervised training scheme called TERA, where the model learns from the reconstruction of altered input. We pre-train TERA using a large amount of unlabeled data, and adapt TERA to downstream SLP tasks using a limited amount of labeled data. We demonstrate strong results in phone classification, keyword spotting, speaker recognition, and speech recognition. We conduct a complete ablation study and a thorough comparison of recent representation learning and pre-training approaches. We show that TERA pre-trained on one dataset can be easily transferred to another downstream dataset. We study how self-supervised models behave on more pre-training data and find that time-only masked reconstruction methods cannot benefit from extensive data. We also study the choice of acoustic features for pre-training. We show that it plays a crucial role in reconstruction-based self-supervised learning, as various surface features will lead to significantly different downstream performance. We investigate networks with different depths and find that small models are more suitable for feature extraction than large models. On the other hand, large models are more effective for fine-tuning than small models.

Acknowledgment

The authors are grateful to the National Center for High-performance Computing for computer time and facilities. They thank Shu-wen Yang for implementing a significant part of the S3PRL toolkit and pre-training the APC, VQ-APC, and NPC models; and Yist Y. Lin for implementing the keyword spotting task.

References