UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
Chengyi Wang, Yu Wu, Yao Qian, Kenichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, Xuedong Huang
Introduction
In the past several decades, enormous progress has been made by the speech recognition (SR) community. SR systems have achieved remarkable quality and even reach human parity in many domains (Watanabe et al., 2018; Li et al., 2020; Wang et al., 2020; Xiong et al., 2016; Chiu et al., 2018). Unfortunately, the successful techniques require many thousands of hours of human-annotated speech recordings for training, which is not available for the vast majority of the nearly 7000 languages spoken worldwide (Katzner & Miller, 2002). This poses a real challenge for building an accurate and robust SR system for low-resource languages. Even for the rich-resource languages, lack of training data is also a serious problem for specific domains, especially when the background noise and distortion conditions vary greatly from the general domain (Li et al., 2014).
Current works tackle the low-resource speech recognition in either supervised or unsupervised manners. In the supervised case, transfer learning methods learn features on large, high-resource datasets and directly use them in similar but data-poor tasks (Ghoshal et al., 2013). Though effective, it requires massive supervised corpora and neglects the large scale unlabeled data. In contrast, the unsupervised method attempts to learn powerful contextual representations from audio alone and then fine-tune the model on labeled data. For instance, wav2vec 2.0 (Baevski et al., 2020b) demonstrates a remarkable performance with 60k hours of unpaired data and 1 hour labeled data on Librispeech. Conneau et al. (2020) further extend the model to a multilingual setting and reduce the phoneme error rate (PER) significantly. The model jointly learns contextual speech representations and a discrete codebook of latent representations, which serves to train the model with a contrastive loss. However, the self-supervised paradigm needs to be carefully designed and such representations may be difficult to interpret. There is no guarantee that the model learns “good” speech representations in terms of the most valuable information for recognition.
In the most of cases, it is less challenging to obtain labeled high-resource data and unlabeled low-resource data, while the labeled low-resource data is hard to collect. Our goal is to leverage all accessible data to learn robust representation across different languages or domains, which is capable of capturing SR-specific content, e.g. phoneme identities, while being invariant to confounding details like the background noise. With such representation, limited amounts of labeled data is sufficient to achieve acceptable performance.
In this work, we propose a unified approach, named UniSpeech, to learn phonetically-aware contextual representations. We follow the model structure of wav2vec2.0 which consists of a feature extractor to extract latent speech representations, a Transformer context network to learn contextual representations and a quantizer to discrete latent representations. We first pre-train the model on the labeled high-resource data and unlabeled low-resource data. Then we freeze the feature extractor and fine-tune the Transformer part on a small amount of labeled low-resource data. For pre-training, we use a multitask learning manner. For labeled data, we train the model towards two objectives: the first is a sequence-level CTC loss (Graves et al., 2006) applied to phoneme labels for phonetic representation learning; the second is a contrastive task defined over the masked contextual representations and the discrete latent representations as in wav2vec2.0. The CTC loss aligns each contextual representation with a phoneme label. Meanwhile, the contrastive loss implicitly closes the distance between discrete representations and contextual representations, with the hope that each codeword from the codebook can also be aligned with a meaningful phoneme unit. However, this simple loss combination method leads to limited improvements. Besides, in contrastive learning, the quantizer is prone to collapse problom where only a small portion of codewords are used. And sometimes it results in locally-optimal codebooks to enable a good contrastive loss, like Voice Activity Decection coodbook or temporally invariant coodbook (Sadhu et al., 2021). Thus, we go further to explicitly guide the quantizer to learn SR-specific information. Specifically, we randomly replace a proportion of the contextual representations with quantized latent representations in the corresponding time steps and calculate the CTC loss upon the mixed representations. In our experiment, we find this method can activate more codewords and helps for learning a phonetic-aware codebook. For those unlabeled data from low-resource setting, we only conduct contrastive learning. As the codebook is already located in the phonetic level, the model is easily adapted to the target domain.
We evaluate the proposed method on cross-lingual SR and domain transfer tasks. On the CommonVoice dataset (Ardila et al., 2019), our UniSpeech outperforms both supervised transfer learning and unsupervised contrastive learning by a large margin in three settings: English to single low-resource language setting (one-to-one), multi-lingual high resource languages to single low-resource language setting (many-to-one), and multi-lingual high resource languages to multi-lingual low-resource language setting (many-to-many). In addition, we test the domain transferability from the Librispeech (Panayotov et al., 2015) to the Tedlium3 dataset (Hernandez et al., 2018), where the source domain and target domain are audiobook reading and live presentation, respectively. UniSpeech achieves a relative 6 word error rate reduction against the baseline.
The main contributions of this paper can be summarized as three-folds: First, we provide a paradigm to use both labeled and unlabeled data to improve the SR performance in low-resource scenario. To the best of our knowledge, it is the first attempt to pre-train speech encoder with both supervised method and self-supervised method in a multitask learning manner; Second, we propose a new learning strategy to explicitly align the discrete latent representation to linguistic units, resulting in a meaningful speech codebook; Third, our approach significantly outperforms both self-supervised learning and supervised learning, and achieves state-of-the-art performance on CommonVoice dataset without the help of extra data.
Related Work
A common way to improve the performance of low-resource ASR models is to leverage data from other high-resource settings. Transfer learning and multitask learning are commonly used methods. Transfer learning (Kunze et al., 2017; Huang et al., 2014; Joshi et al., 2020) first trains the model on the high-resource setting and then fine-tune it on the target data-scarce setting. The parameters learned from the first setting serve as a starting point, also known as supervised pre-training. In multitask learning (Huang et al., 2013; Knill et al., 2014; Chen & Mak, 2015), the model is simultaneously trained on multiple languages with shared components. Both methods depend on labeled data from multiple languages to yield consistent improvements while large amount of unlabeled data cannot be used.
Recently, self-supervised learning has received great attention as it does not require any labeled data. Based on the training objectives, self-supervised methods can be categorized into reconstructive learning (recreating audio frames) and contrastive learning (discriminating true sample from set of negative samples). Chen et al. (2019) use an autoencoder to perform full reconstruction. (Chorowski et al., 2019) use a high capacity WaveNet autoencoders to learn meaningful speech representations. They compare three different variants of constraints: a simple bottleneck, a Gaussian Variational Autoencoder (VAE) and a Vector Quantized VQE (VQ-VAE). Autoregressive predictive coding (APC) (Chung et al., 2019) reconstructs the future frame with an unidirectional encoder for phone classification and SR task. Masked reconstruction (Liu et al., 2020; Ling et al., 2020a, b) has also been widely investigated which masks part of the input and learn to reconstruct it. In the research line of contrastive learning, CPC (Oord et al., 2018) uses an autoregressive model to classify future frames from negative examples. Wav2vec (Schneider et al., 2019) evaluates the effectiveness of contrastive learning on speech recognition task. Kawakami et al. (2020) and Rivière et al. (2020) show bi-directional and modified CPC transfers well across domains and languages. Vq-wav2vec (Baevski et al., 2020a) uses a vector quantization module to learn discrete representations. They further introduce Wav2vec2.0 (Baevski et al., 2020b), which masks the speech input in the latent space and solves a contrastive task defined over contextual representations in the masked region and a quantization of the latent representations. They show the discrete latent speech representations learnt by quantizer are related to phonemes. Conneau et al. (2020) try the idea on multilingual settings, named XLSR.
Some other recent work employ multitask learning strategy for speech representation learning. (Pascual et al., 2019) and (Ravanelli et al., 2020) propose to learn a problem-agnostic speech encoder which jointly solves different self-supervised tasks, including reconstructive loss and contrastive loss. Compared to our work, they use only unlabeled data and the models are evaluated on speaker identification, emotion classification and ASR tasks. While we focus on improving ASR performance in low-resource scenarios with all availabel data. (Talnikar et al., 2020) alternatively minimize the unsupervised masked CPC loss and the supervised CTC loss. They focus on simplify the training pipeline and reports equivalent word error rate as in wav2vec2.0. In contrast, our method can improve the recognition performance significantly.
UniSpeech
2 Model Structure
Our model architecture is shown in Figure 1. Following the design choices in wav2vec2.0 (Baevski et al., 2020b), the model contains a convolutional feature encoder, a Transformer context encoder and a vector quantizer. The convolutional feature encoder maps raw audio to a latent space . It is composed of seven blocks of temporal convolution followed by layer normalization and GELU activation layer. The temporal convolutions in each block have 512 channels with strides (5,2,2,2,2,2,2) and kernel widths (10,3,3,3,3,2,2), resulting in each represents about 25ms of audio strided by 20ms. Then the representations are fed into the Transformer (Vaswani et al., 2017; Devlin et al., 2019) network to output context representations . The Transformer network is equipped with a convolutional layer with kernal size 128 and 16 groups to replace absolute positional embedding. We evaluate on two different settings with the same feature encoder: Base with 12 Transformer blocks, model dimension 768, inner dimension 3072 and 8 attention heads; and Large with 24 Transformer blocks, model dimension 1024, inner dimension 4096 and 16 attention heads Acting as an information bottleneck on the latent representation, the quantizer module discretizes to a finite set of speech representations . The quantizer has codebooks with entries each. For each frame, we select two entries from two codebooks independently to obtain quantized representation , resulting in over 100K codewords in total ().
3 Multitask Learning with Unified Representation
Specifically, given a data pair , the model learns its context representations . We add a linear layer with softmax to predict a distribution over observed labels, including phoneme tokens and a blank token, denoted as . A legal CTC path is a variation of the transcription by allowing occurrences of blank tokens and repetitions. The CTC objective trains the model to maximize the sum of conditional probability of all possible legal paths:
where is the set of all valid alignments.
where is a non-negative temperature, and are uniform samples from . In the forward pass, the quantizer finds a nearest prototype to the input from each codebook, denoted as , where . We concatenate the resulting vectors and apply a linear transformation to obtain . has the same dimension as Transformer encoder. In the backward pass, the gradient of the loss with respect to the pre-quantized vector is approximated using the straight-through estimator, that is .
For each centered over masked time step , the model needs to identify the true quantized latent speech representation in a set of quantized candidates . The distractors are uniformly sampled from the other timesteps from the same utterance. This frees up the model from using its capacity to represent speaker-dependent information and instead focuses on phonetic features. The loss is defined as
The objective is augmented by a codebook diversity loss with a loss weight 0.1. It encourages the equal use of all entries by maximizing the entropy of the averaged softmax distribution over the codebook entries:
The final pre-training loss can be defined as:
The quantization module was qualitatively shown to learn a representation which separates phonetic content within an utterance from the speaker identity (Baevski et al., 2020b). Moreover they discover the tokens learnt in an unsupervised manner can be mapped to phonemes in a limited setting. However, it cannot guarantee that the discrete representations are as useful as those learnt by the supervised learning for the ASR tasks. And we find the above multitask method leads to limited gain (as shown in the Table 1 of the experiment part). To address it, when calculating the CTC loss, we replace the continuous representation with its quantized versions with probability . Mathematically, the conditional probability of Eq. 1 is changed as
where is either or . Since is a phoneme sequence, predicting with can explicitly guide the quantizer to cluster phonemes and learn SR specific knowledge into codebooks.
With Eq. 7, representations in supervised learning and unsupervised learning are forced to project into the same space, and it avoids the two objective functions optimize themselves individually. Although the method is simple, it is effective according to our experiment results. Since the representations are unified in two tasks and different languages, we call our model UniSpeech.
Experiments
We evaluate our methods on Multilingual ASR task and domain transfer ASR task.
Regarding multilingual ASR, we first train the UniSpeech model on high-resource languages, and then transfer it to low-resource languages. We employ the CommonVoice (CV) dataset (Ardila et al., 2019) https://commonvoice.mozilla.org/en/datasets. We use the June 2020 release version for training our models., which is a multilingual corpus of read speech comprising more than 5k hours of speech data in 60 languages. To be comparable with XLSR (Conneau et al., 2020), we consider the following eight languages for evaluationFiles in the test set of Turkish(tr) and Chinese(zh) are missing in CommonVoice June2020 release, so we exclude these languages.: Spanish (es), French (fr), Italian (it), Kyrgyz (ky), Dutch (nl), Russian (ru), Swedish (sv) and Tatar (tt). English (en) is always regarded as a high-resource language. Our pre-training data is not exactly same as Conneau et al. (2020) as we use different dataset version, but we use the same data size for each language without selection. We define three settings based on the number of pre-training languages and fine-tuning langauges: one-to-one, many-to-one and many-to-many. The pre-training details will be illustrated in the corresponding experiment settings. For fine-tuning, we use the evaluation splits from Rivière et al. (2020), which contains 1 hour paired data for training, 20 minutes for validation and 1 hour for testing. We retrieve phoneme transcriptions by running open-source phonemizerhttps://github.com/bootphon/phonemizer and report PER following prior work (Conneau et al., 2020).
Models are implemented in fairseq (Ott et al., 2019). To train the UniSpeech model, we use mask probability , loss weight and replace probability unless otherwise stated. During pre-training, we crop each utterance to 250k samples for Base model and 320k samples for Large model. Each batch on one GPU contains max up to 1.4m samples for Base and 1.2m samples for Large. The models are trained on 64 GPUs. We use Adam optimizer where the learning rate is warmed up for the first 10% of updates to a peak of 5e-4(Base) or 1e-3(Large) and then linearly decayed over a total of 250k updates. The model is fine-tuned with 2 GPUs. We still use Adam optimizer and the learning rate is warmed up for 2k updates to 2e-5, keep constant for 8k updates and then linearly decay for 10k updates. Dropout 0.1 is always used for both pre-training and finetuning.
1.2 Results
From Table 1, we can see that CTC-Transfer is better than XLSR as it uses the label information in English data. However, when the target language unpaired data is available, the performance of outperforms CTC-Transfer significantly, indicating in-language data is the key to the success of unsupervised method. Compared with CTC-Transfer and XLSR, our UniSpeech obtain PER reductions of 9.6 and 13.4 respectively, which is mainly because our method combines transfer learning and self-supervised learning, and the two methods are complementary. However, when we set the replace probability as 0, our method degrades to multitask learning and it is worse than UniSpeech by 7.7 relatively. Furthermore, when target unlabeled data is available, achieves 8.8 PER score, outperforming by relative 12.9%. This indicates our model can well utilize both supervised and unsupervised data and it learns robust speech representation which is easily transferred across different languages. The conclusion is the same for Large setting. The best result for the one-to-one setting is obtained by UniSpeech-L+, which gets 7.7 PER on average, a relative PER reduction of 19.8% compared to XLSR+ Large model.
In Table 2, the observation is consistent with one-to-one setting. The UniSpeech reaches relative 17.8% and 9.8 PER reduction compared to CTC-Transfer and XLSR respectively, indicating that our method is effective on the cross-lingual transfer task. An interesting phenomenon is the performances of CTC-Transfer and even the UniSpeech with are worse than XLSR, which is different from the conclusion in the previous setting. A possible explanation is that it is hard to deal with multiple vocabularies in CTC pre-training. As there are many overlapped phonemes appear in multiple vocabularies, some of them share similar pronunciation while the others have totally different pronunciations. Instead, the unsupervised method does not suffer from this issue. This indicates the codebook in unsupervised learning has the ability to adapt to multiple languages. When the target unlabeled data is available, outperforms by 21.1% relative PER. It also significantly outperforms which uses unpaired data from 10 languages. For Large model, our UniSpeech+ even improves baseline by 26.9% relative. It shows UniSpeech is more robust and transferable than either supervised learning or unsupervised learning alone. Furthermore, many-to-one results are better than one-to-one results on the 5 unseen low-resource languages. It suggests that large and diverse training data is beneficial to our model.
We also evaluate the model for multilingual fine-tuning. In the pre-training stage, we use the same pre-trained model as in the many-to-one setting for UniSpeech experiment. For , instead of pre-training one model for each low-resource scenario, we merge the 5 unlabeled low-resource datasets and pre-train the model on the joint set. In the fine-tuning stage, the 5 labeled datasets are always merged. As there are overlapped phonemes across languages, we can either regard each as a single label for shared vocabulary or as different labels with different language ids for separate vocabulary.
As Table 3 shows, UniSpeech outperforms CTC-Transfer and XLSR by 10.4 and 5.1 respectively. When the unlabeled datasets from target languages are used, the gain for UniSpeech+ against XLSR+ is 6 PERR. The overall PER is higher than the many-to-one setting, because the multi-lingual outputs make the task harder. The performances of shared vocabulary and separated vocabulary are similar. After checking the shared phoneme vocabulary, we find that the same phoneme unit generated by the phonemizer may have different pronunciation and thus represent different sounds in different languages. A better universal phonemizer is worth trying in the future.
2 Domain Transfer
We also conduct an experiment for domain transfer task. We use our UniSpeech model to pre-train phoneme representations in reading English domain, namely on Librispeech, and transfer them to spoken English domain on Tedlium3 (Hernandez et al., 2018) dataset. We only use the 960 hours of labeled speech as pre-training dataset in this experiment. To train the UniSpeech model, we also extract the phoneme sequences by phonemizer to calculate the CTC loss. During fine-tuning, we discard the phoneme CTC layer and use character based CTC loss. We reports WER on dev and test sets. The model is pre-trained on 64 GPUs to 400k steps and fine-tuned on 8 GPUs to 320k steps. Other training parameters are the same in the multilingual experiments.
We compare our model with supervised CTC-Transfer and unsupervised wav2vec2.0 pre-training. As we use character-level fine-tuning loss, we report results using character-level pre-training for CTC-Transfer baseline. This leads to better performance than phonetic-level pre-training. We also compare with results from Kawakami et al. (2020). They first train a bidirectional CPC model on an unlabeled audio dataset, either Librispeech or 8k mixed audio dataset spanning a range of recording conditions, noise levels, speaking styles and languages. Next they freeze the model’s parameter and use its output representations as input to train a TDNN based CTC model. Their methods are denoted as CPC-Librispeech and CPC-8k respectively. They also report results when using LogFilterbank as input feature.
From Table 4, we can draw the following conclusions. There is no doubt that wav2vec 2.0 outperforms CPC-based models largely because of a better network structure (Transformer v.s. TDNN) and a better pretraining method. The wav2vec2.0 baseline obtains similar performance compared with CTC transfer learning, indicating that self-supervised learning can learn powerful speech representation. UniSpeech is better than wav2vec 2.0 and CTC transfer, which shows the advantages of our model on domain transfer task.
3 Discussion and Analysis
In this section, we investigate whether the discrete latent speech representations learnt by the quantizer can be mapped to the meaningful phonetic units. Following Baevski et al. (2020b), we compute the discrete latents on the training data of TIMIT, which contains 5 hours of audio recordings with human annotated phonemes. We use the multilingual UniSpeech model without any fine-tuning. We then compute the conditional probability based on the co-occurrence between phonemes and the latents. The alignments are built by choosing the phoneme which is most represented in the receptive field of each . Figure 2 shows that many discrete latents appear to specialize in specific phonetic sounds, indicating our methods can obtain a good alignment between latents and labeled phonemes. The silence phoneme (bcl) is aligned with the most latents, since there are many blanks tokens in CTC training and TIMIT data has slience in every utterance.
We define two metrics to quantitatively evaluate our method against the baselines. 1) The number of active discrete codewords calculated with TIMIT training data. There are over 100k entries by combining two codebooks and most of them are not active (not triggered by the gumbel softmax layer for any frames of TIMIT training data). The more active latents, the more diverse the codebook. 2) The average entropy of alignments, which is computed by . High entropy indicates is closed to a uniformed distribution, thus a low entropy is preferred. We compare our model to multilingual XLSR. For UniSpeech, the number of active codewords is 25743 and the entropy is 0.83. While for XLSR, the two numbers are 3645 and 1.34 respectively. This indicates our model learns a more diverse codebook and it’s better at phoneme clustering. It is a possible interpretation of why our model outperforms XLSR on CommonVoice dataset with the same diversity loss weight.
Table 5 shows the impact of various hyperparameter choices. The experiments are conducted in the one-to-one setting. First, we show the impact of different replacement probability . When we set = 0, the UniSpeech model degrades to a simple multitask model and the performance drops about 7.8% relatively. The model performs similarly when ranges from 0.3 to 0.7. We set as 0.5 in all experiments as it performs best according to our hyperparameter search. Setting the loss weight too low leads to lower performance. While increasing it from 0.5 to 0.7 leads to no improvements. Finally, we find that increasing the mask probability has little impact on the model’s performance.
Conclusion
In this paper, we propose the UniSpeech to learn speech representations with unlabeled and labeled data, in which supervised CTC labeling and phonetically-aware contrastive learning are unified with a multitask learning framework. Unispeech consists of a convolutional feature extractor, a Transformer encoder and a quantizer, and the quantizer is explicitly guided to learn SR-specific information. The results show that UniSpeech outperforms both self-supervised and supervised pre-training alone by a large margin on multilingual SR and domain transfer tasks. In the future, we will scale up our model with over one million hours of unlabeled data, explore more complex SR architecture as well as the usage of pre-trained language model in speech recognition.