Exploring wav2vec 2.0 on speaker verification and language identification

Zhiyun Fan, Meng Li, Shiyu Zhou, Bo Xu

Introduction

Recently, neural networks trained with a large amount of labeled data can meet most industrial needs in the field of speech processing . However, purely supervised learning seems to be inconsistent with the mechanism of human learning. Early on in their lives, human infants learn language by watching and listening to adults around them, which resembles an unsupervised learning process. Later, they learn reading and writing, which seems to be a supervised learning process. To simulate the two-stage learning process, a lot of self-supervised frameworks are proposed .

In the field of speech processing, most self-supervised methods can be divided into two categories. One kind of method is conducted by the reconstruction loss, such as autoregressive predictive coding (APC) , masked predictive coding (MPC) and so on. The other kind of method is conducted by contrastive predictive loss. The most representative work is the contrastive predictive coding (CPC) and wav2vec . The wav2vec 2.0 used in this paper belongs to the latter category. Most of these self-supervised pre-training methods are applied to speech recognition. However, there is almost no work on whether pre-training methods could work on the speaker verification (SV) or the language identification (LID). In this paper, we use the framework of wav2vec 2.0 to explore this feasibility.

We denote the model structure used in wav2vec 2.0 as w2v-encoder in this paper. It is illustrated in the dashed box of Fig. 1. It mainly consists of a convolutional neural network (CNN) encoder and a Transformer . The CNN transfers raw waveform input to latent speech representations. They are fed to the Transformer after being masked and converted to context representations. A quantization module converts the latent speech representations to a discrete version which is used as the target. The whole model is trained to solve a contrastive task, which requires identifying the true quantized latent speech representations for a masked time step within a set of distractors . After pre-training, Baevski et al. applied it to ultra-low resource speech recognition. Using only ten minutes labeled data, their approach achieved word error rate (WER) of 5.7/10.1% on the clean/noisy test sets of Librispeech. The results demonstrate that the phoneme-related information is preserved during the pre-training of w2v-encoder and the drownstream task such as speech recognition can benefit a lot from it. Audio is a complex signal that contains not only phoneme-related information but also factors about speaker, language, environment, noise, etc. However, there is very little work on pre-training for the SV and the LID.

In this paper, we explore the effectiveness of self-supervised pre-training on the SV and the LID tasks. We utilize the pre-trained w2v-encoder to extract context representations, and use t-SNE tools to visualize them. We find that they have distinguishability among different speakers and languages even if pre-training of wav2vec 2.0 is problem-agnostic. Moreover, we find the lower layer has the stronger distinguishability. This distinguishability is exactly what the SV and the LID tasks need. It also verifies the feasibility of applying the self-supervised pre-training to the two tasks. Thus, we attempt to fine-tune the pre-trained model on these two downstream tasks respectively. For the SV task, we obtain an EER of 3.61% on the test set of the VoxCeleb1 dataset. For the LID task, we obtain an EER of 12.02% on the 1 second condition and 3.47% on the full-length condition of the AP17-OLR dataset. Furthermore, in order to simplify the fine-tuning process and reduce model parameters, we utilize the multi-task learning to conduct the fine-tuning on the two tasks simultaneously.

Method

In this section, we first review the pre-training of the wav2vec 2.0 . Then we introduce how to apply the pre-trained model to downstream tasks. Fig. 1 illustrates the pre-training and fine-tuning.

The left side of Fig. 1 gives an illustration of the w2v-encoder and its pre-training. The main body of the model consists of a CNN-based feature encoder, a Transformer-based context network and a quantization module. The CNN encoder stacks seven blocks, and in each block the temporal convolutions followed by a GELU activation function have 512 channels with strides (5,2,2,2,2,2,2)(5,2,2,2,2,2,2) and kernel widths (10,3,3,3,3,2,2)(10,3,3,3,3,2,2). The CNN encoder maps the raw audio XX into latent speech representations ZZ.

The context network stacks 12 Transformer blocks with model dimension 768768, inner dimension 3,0723,072, and 88 attention heads. Before sending ZZ into the context network, all time steps of ZZ are randomly sampled as starting indices with p=0.065p=0.065, and M=10M=10 consecutive time steps from every sampled index are masked. Then the relative positional embedding is added to the masked representations. The Transformer contextualizes the masked representations and generates context representations CC.

The pre-training process is optimized with Adam . During the first 8% of the updates, the learning rate warms up to a peak of 5×10−35\times 10^{-3}, and then it decays linearly. For more details about the pre-training of wav2vec 2.0, we refer readers to .

2 Fine-tuning

Before the post-training, we add an average pooling layer and a fully connected layer on the top of w2v-encoder. The average pooling layer converts the frame-level context representations given by w2v-encoder into sentence-level representations, and the fully connected layer classifies each sentence into some speaker or some language.

The newly added fully connected layer is randomly initialized, and w2v-encoder is initialized with the base model released by Baevski et al https://github.com/pytorch/fairseq/blob/master/examples/wav2vec/. The cross-entropy criteria is employed as the loss function for the classification of speakers or languages. Specially, for the training of speaker classification, AM-softmax is used to increase the discrimination of the learned embedding to the speaker.

In the multi-task fine-tuning, we add a pooling layer and two parallel fully connected layers to predict the speaker and language respectively. The training loss is obtained by the weighted sum of the losses of these two tasks. The LsvL_{sv} and LlidL_{lid} in Eq. 5 represent the CE loss of the SV and the LID tasks respectively.

Due to the problem of unbalanced data volume in the datasets of the speakers and languages, the batch is generated by sampling from two datasets with equal probability to ensure the data used in the training process is balanced. In addition, the two tasks also have the problem of inconsistent convergence speed. We mitigate this issue by adjusting the weight of the loss of the two tasks through the development set.

Experiments

Various informative factors are mixed in speech signals, including semantics, speaker, emotion, channel, background noise, etc. Baevski et al. have shown that the representations underlying pre-trained w2v-encoder can capture the linguistic factors. It remains unclear whether the problem-agnostic pre-training of wav2vec 2.0 can learn about any other factors. In the experiment part, we take speaker and language factors as examples to explore this question, and try to apply wav2vec 2.0 to the SV and the LID tasks.

VoxCeleb1 and AP17-OLR datasets are used in our experiments for the SV and the LID respectively.

Speaker verification dataset: VoxCeleb1 contains over 100,000 utterances from 1,251 celebrities. It can be used for both speaker identification and verification. We use the VoxCeleb1 to conduct the SV task. And the consine distance is used to calculate the similarity score. The data split of the VoxCeleb1 dataset for verification is listed in Table 1.

Language identification dataset: AP17-OLR consists of 10 different languages (Mandarin, Cantonese, Indonesian, Japanese, Russian, Korean, Vietnamese, Kazakh, Tibetan and Uyghur). The duration of training data for each language is about 10 hours with the speech sampled at 16 kHz. The test set contains three subsets with different durations (1 second, 3 second, and full length). These subsets respectively contain 17964, 16404 and 17964 utterances.

2 Model description

In the experiments, we utilize the base model released by Baevski et al. and three models fine-tuned by us. For simplicity, we use some symbols to represent them, and the explanations are as follows:

M-nofinetune: the base model pre-trained on the Librispeech corpus .

M-sv: We fine-tune M-nofinetune on the VoxCeleb1 dataset for speaker verification.

M-lid: We fine-tune M-nofinetune on the AP17-OLR dataset for language identification.

M-multi: We fine-tune M-nofinetune on the AP17-OLR and VoxCeleb1 dataset simultaneously in a multi-task form.

3 Feasibility analysis

In this section, we explore whether the speaker and language factors are retained during the pre-training of wav2vec 2.0. It determines whether the pre-training method can be used for these two tasks.

We directly extract context representations from the test set of AP17-OLR and VoxCeleb1 with the M-nofinetune model. Then we visualize the context representations by t-SNE , a nonlinear dimensionality reduction algorithm for visualizing high-dimensional data. The results are shown in Fig. 2. The left three images are the visualization results of the features from the three layers of the Transformer. Different colors represent different speakers. It is not difficult to find that all the three layers have certain speaker distinguishability, and this distinguishability is more obvious at the bottom of the Transformer. In the three images on the right, different colors represent different languages. It can also be found that these features are distinguished by languages, and the lower the layer, the stronger the distinction. The phenomena presented in Fig. 2 show that the model pre-trained by wav2vec 2.0 can effectively extract the characteristics of the speaker and language of the speech.

We further quantify this claim by performing the SV and the LID with a simple fully connected layer. The pre-trained w2v-encoder (M-nofinetune) acts as a feature extractor. The fully connected layer is optimized to distinguish the 10 languages or 1211 speakers for the two tasks respectively. The test results are listed in Table 2.

The randomrandom results are evaluated on the randomly initialized w2v-encoder. The comparison of these two results in Table 2 further illustrates that the model pre-trained by wav2vec 2.0 can extract speaker and language-related characteristics, which provides a basis for the application of wav2vec 2.0 to the SV and the LID.

4 Speaker verification

From the experiments in the previous section, we can see that the pre-trained w2v-encoder, M-nofinetune, can extract features that contain a certain speaker distinguishability. This kind of speaker distinguishing learning is exactly required in the SV task. Then we attempt to fine-tune the pre-trained w2v-encoder, M-nofinetune, to finish the SV task. We initialize the w2v-encoder with M-nofinetune, and add a randomly initialized fully connected layer on the top of it to predict speakers. The fine-tuning is conducted on the VoxCeleb1 dataset . All parameters are adjustable during fine-tuning. However, at the first 1000010000 steps, the w2v-encoder is frozen. We optimize the model with Adam, the learning rate warms up to 5×10−35\times 10^{-3} during the first 60006000 steps, and then it decays linearly during the remaining 70007000 steps.

The no preno\ pre-trainingtraining in Table 3 represents that using VoxCeleb1 to train the w2v-encoder added a fully connected layer without pre-training. Our fine-tuning model, M-sv, outperforms the no preno\ pre-trainingtraining result by a significant margin (EER of 3.61%3.61\% vs 24.28%24.28\%). The gap between them illustrates the benefits of pre-training. Moreover, our model outperforms all baselines in Table 3, and obtains new state-of-the-art results on the VoxCeleb1 dataset. It means the pre-training of wav2vec 2.0 is useful to the SV task and can work well without any task-specific adjustment of model structure.

5 Language identification

Although Baevski et al. only used English data during the pre-training of M-nofinetune, it can be seen from the visualization results in section 3.3 that the features extracted by M-nofinetune still retain the distinction of language. It means that the model obtained by this pre-training method may be useful to the language identification system. Similarly, we add a fully connected layer on top of the w2v-encoder to predict language. We initialize w2v-encoder with M-nofinetune and randomly initialize the extra fully connected layer. Then the whole model is fine-tuned on the AP17-OLR dataset . We optimize the model with Adam, the learning rate warms up to 5×10−35\times 10^{-3} during the first 50005000 steps, and then it decays linearly during the remaining 80008000 steps. The parameters of the w2v-encoder part are frozen at the first 50005000 steps. After training, we test on the model, which obtains the best performance on the development set.

The first two rows in Table 4 are the two baselines released by the organizer of the AP17-OLR challenge. The TSMTSM-DNNDNN-BNBN-LSTMLSTM is one of the best models on this benchmark. The no preno\ pre-trainingtraining in Table 4 means that the fine-tuning starts from scratch on the AP17-OLR dataset. The M-lid, which is fine-tuned from the pre-trained M-nofinetune, outperforms the no preno\ pre-trainingtraining result by a large margin on both the 1 second condition and the full-length condition. The gap between them illustrates the benefits brought by pre-training to the LID task. Compared with baselines released by the organizer, M-lid shows a clear performance advantage on the two test conditions. However, it is far from the best results. It means that wav2vec 2.0 is useful to the LID task. However, its effectiveness on the LID task is not good as the SV task. We consider that the use of multiple languages during pre-training (not just English) can mitigate this issue. In addition, we find that the performance of the no preno\ pre-trainingtraining is influenced by overfitting seriously. This problem is obviously alleviated during the fine-tuning of M-lid, which benefits from pre-training.

6 Multi-task system

The parameters of the w2v-encoder have reached 9494M. Fine-tuning two models for the SV and the LID tasks independently will take up a lot of resources. Hence, we try to use one model to finish these two tasks simultaneously. On the top of the w2v-encoder we connect two fully connected layers in parallel to predict the speaker and language respectively. We follow the experiment settings described in section 2.2. The λ\lambda in Eq. 5 is set to 0.70.7.

Results in Table 5 show that compared with single-task training, although the performance of multi-task form is a bit reduced, it achieves good results with fewer parameters on the SV and the LID tasks. It shows that the pre-training of wav2vec 2.0 can be combined with multi-task learning to achieve unified modeling of the two tasks. This greatly simplifies the use of pre-trained model and can save a lot of time spent on fine-tuning to each task. In addition, it can reduce the demand for storage.

Conclusion

In this paper, we explore the application of wav2vec 2.0 on speaker verification and language identification. First of all, through some preliminary experiments and visualization methods, we find that the features extracted by the pre-trained w2v-encoder have the distinction between speakers and languages, and this distinction is more obvious in lower layers. This illustrates the feasibility of using the pre-trained model for the SV and the LID tasks. Then we verify the effectiveness of the pre-trained model on the two tasks and obtain competitive results on the VoxCeleb1 and the AP17-OLR datasets. Finally, in order to simplify the fine-tuning process on multiple tasks and reduce parameters, we use a multi-task learning mechanism, so as to realize the unified modeling for the SV and the LID. In future work, we are planning to extend wav2vec 2.0 to more speech processing tasks with the multi-task learning.

References