VQVAE Unsupervised Unit Discovery and Multi-scale Code2Spec Inverter for Zerospeech Challenge 2019
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, Satoshi Nakamura
Introduction
Current spoken language technologies only cover about two percent of the world’s languages. This is because most groundworks require a large amount of paired data resources, including a sizeable collection of spoken audio data and corresponding text transcription. On the other hand, most of the world’s languages are severely under-resourced, some of which even lack a written form. Zero resource speech research is an extreme case from low-resourced approaches that learn the elements of a language solely from untranscribed raw audio data. This completely unsupervised technique attempts to mimic the early language acquisition of humans. The zero resource speech challenge (ZeroSpeech) is directly addressing this issue and offers participants the opportunity to advance the state-of-the-art in the core tasks of zero resource speech technology.
In ZeroSpeech 2015 and 2017, the goal was to discover an appropriate speech representation of the underlying language of a dataset . The ZeroSpeech 2019 challenge confronts the problem of constructing a speech synthesizer without any text or phonetic labels: TTS without T. The task requires the full system not only to discover subword units in an unsupervised way but also to re-synthesize the speech with a same content to a different target speaker. It includes both ASR and TTS components. In this paper, we describe our submitted system for the ZeroSpeech Challenge 2019 and focus on constructing end-to-end systems.
The top performances in discovering speech representation in ZeroSpeech 2015 and 2017 are dominated by a Bayesian non-parametric approach with unsupervised cluster speech features using a Dirichlet process Gaussian mixture model (DPGMM) . However, the DPGMM model is too sensitive to acoustic variations and often produces too many subword units and a relatively high-dimensional posteriogram, which implies high computational cost for learning and inference as well as more tendencies for overfitting . Therefore it is difficult to synthesize speech waveform from the resulting DPGMM-based acoustic units.
To tackle these problems and achieve the best trade-off, an optimization method is required to balance and improve both components. Recently, Tjandra et al. proposed a machine speech chain (see Figure 1) that enables ASR and TTS to assist each other when they receive unpaired data by allowing them to infer the missing pair and optimize both models with reconstruction loss. However, since the architecture is based on an attention-based sequence-to-sequence framework that transforms from a dynamic-length input into a dynamic-length output without decoding at the frame-level (one symbol per frame), it is less suitable for this challenge.
Inspired by a similar idea, we propose to utilize a frame-based vector quantized variational autoencoder (VQ-VAE) and a multi-scale codebook-to-spectrogram (Code2Spec) inverter trained by mean square error (MSE) and adversarial loss. VQ-VAE extracts the speech to a latent space and forces itself to map onto the nearest codebook, leading to compressed representation. Next, the inverter generates a magnitude spectrogram to the target voice, given the codebook vector from VQ-VAE. In our experiments, we also investigate other clustering algorithms such as K-Means and GMM and compare them with the VQ-VAE result on ABX scores and bit rate.
Vector Quantized Variational Autoencoder (VQ-VAE)
A vector quantized variational autoencoder (VQ-VAE) is a variant of variational autoencoder architecture. It has several differences compared to a standard autoencoder or a variational autoencoder (VAE). First, the encoder generates discrete latent variables instead of continuous latent variables to represent the input data. Second, instead of one-to-one mapping between the input data and the latent variables, VQ-VAE forces the latent variables to be represented by the closest codebook vector.
In the discretization process, we choose closest codebook vector based on the index of the closest distance (e.g., L2-norm distance) from continuous representation . To decode the data, we use codebook and speaker embedding and feed both into decoder to reconstruct original data .
In VQ-VAE, we formulate the training objective:
where function stops the gradient, defined as:
There are three terms in loss . The first is a negative log-likelihood that resembles a reconstruction loss and optimizes the encoder and decoder parameters. The second optimizes codebook vectors , named codebook loss. The third forces the encoder to generate a representation near the codebook, called commitment loss. Coefficient is used to scale the commitment loss.
Codebook-to-Spectrogram Inverter
After we define the multiple objectives for training, we update each module parameter and with the following equation:
where Optim() is a gradient optimization function (e.g., SGD, Adam ), and is the coefficient to balance the loss between the MSE and the adversarial loss. In the inference stage, given the predicted linear magnitude spectrogram , we reconstruct the missing phase spectrogram with the Griffin-Lim algorithm and applied the inverse short-term Fourier transform (STFT) to generate the waveform.
Experiment
In this section, we describe the feature extraction, the preliminary models, and our proposed models for this challenge. All of the results were evaluated using evaluate.sh from the English test set.
There are two datasets for two languages, English data for the development dataset, and a surprise Austronesian language for the test dataset. Each language dataset contains subset datasets: (1) a Voice Dataset for speech synthesis, (2) a Unit Discovery Dataset, (3) an Optional Parallel Dataset from the target voice to another speaker voice, and (4) a Test Dataset. The source corpora of the surprise language are describe here , and further details can be found here . In this work, we only use (1)-(2) for training and (4) for testing.
For the speech input, we experimented with several feature types, such as Mel-spectrogram (80 dimensions, 25-ms window size, 10-ms time-steps) and MFCC (13 dimensions (total=39 dimensions), 25-ms window size, 10-ms time-steps). Both MFCC and Mel-spectrogram are generated by the Librosa package .
2 Official baseline and topline model
ZeroSpeech 2019 provides official baselines and toplines. The baseline consists of a pipeline with a simple acoustic unit discovery system based on DPGMM and a speech synthesizer based on Merlin, and the topline uses gold phoneme transcription to train a phoneme-based ASR system with Kaldi and a phoneme-based TTS with Merlin. The performance is shown in Table 1.
3 Preliminary model
We started to explore this challenge using a simpler method and gradually increased our model’s complexity.
We directly evaluated the ABX and the bit rate of Mel-spectrogram and MFCC as speech representations. In Table 2, we report each feature extraction method with respect to their ABX and bit rates. In our preliminary experiments, MFCC produced better performances on the ABX metric than the Mel-spectrogram. Therefore, for the rest of our discussion, we only focus on utilizing MFCC features. However, even the MFCC has better ABX score, the bit rate still remains too high.
3.2 K-Means
We trained Minibatch K-Means (with scikit-learn toolkit ) on the MFCC feature and varied the cluster size: 64, 128, 256. We represent a data point (a speech frame) K-Means by using the closest centroid vector to the data frame and calculate the ABX with the DTW cosine. Table 3 reports all the models and their configurations with respect to their ABX and bit rate.
3.3 Gaussian Mixture Model (GMM)
We trained GMM with diagonal covariance matrices (with scikit-learn toolkit ) on the MFCC features. We varied the number of mixtures: 64, 128, and 256. We represent a data point (a speech frame) with the posterior probability from each component with a Bayes rule and calculate the ABX with DTW KL-divergence. In Table 4, we report all of the models and their configurations with respect to their ABX and bit rate.
4 Proposed model
Next we describe our encoder and decoder architecture in Fig. 4 with four times the sequence length reduction. For the input and output targets, we use the MFCC features and explore different stride sizes to reduce the time length from 1, 2, 4, 8. We use speaker embedding with 32 dimensions and codebook embedding with 64 dimensions. We varied the number of codebooks: 64, 128, 256, 512. Batch normalization and LeakyReLU activation were applied to every layer, except the last encoder and decoder layer. The decoder input is a concatenation between codebook and speaker embedding in the channel axis. We set commitment loss coefficient .
4.2 Multi-scale Code2Spec inverter
In Fig. 4, we describe our inverter architecture. Our input is a codebook sequence with 64 dimensions and our target output is a sequence of linear magnitude spectrogram with 1025 dimensions. The first four layers have multiple kernels with different sizes across the time-axis. All convolution layers have stride = 1 and the “same” padding. Batch normalization and LeakyReLU activation are applied to every layer, except the last one before the output prediction. For the adversarial loss, we found LSGAN is stabler, thus LSGAN with is used in every model. We independently trained the inverter to generate a voice target speaker with a train/voice set. We have two inverters for the English set and one for the surprise set.
4.3 Model training
We used Adam as our first-order optimizer for both VQ-VAE and the Code2Spec inverter. All of our models are implemented with PyTorch framework.
4.4 Results and Discussion
Table 5 reports all models and their configurations with respect to their ABX and bit rate. Considering the balance between the discrimination score ABX and the bit-rate compression rate, we submitted two proposed systems: (1) 256 codebooks and 4 stride size to reduce the time length and (2) 256 codebooks and 2 stride size to reduce the time length.
We also attempted further enhancement of the synthesized voice using several techniques, such as WaveNet and GAN-based voice conversion . WaveNet decoder is conditioned by frame-wise linguistic features or acoustic features with a 5ms timeshift (80 times smaller than the speech samples). As the sample rate of the codebook embeddings of our system was 320 times smaller than the speech samples, the Wavenet couldn’t produced satisfying result. GANs are known to be effective for achieving high-quality voice conversion with clean input data . However, our task is more challenging due to the fact that our generated voice will always have some distortion. Therefore, GAN-based voice conversion approach failed to improve our performance. As a future work, we will investigate the use of GAN-based speech enhancement approaches to further improve our results.
Conclusions
We described our approach for the ZeroSpeech Challenge 2019 for unsupervised unit discovery. We explored many different possibilities: feature extraction, clustering algorithm, and embedding representation. For our final submission, we utilized VQ-VAE to extract a sequence of codebook vectors. The codebook generated by VQ-VAE has a better trade-off between ABX and the bit rate compared to the other models such as K-Means, GMM, or direct feature representation. To reconstruct speech from the codebook, we trained a Code2Spec inverter to generate a corresponding linear magnitude spectrogram. The combination between VQ-VAE and Code2Spec significantly improved the intelligibility (in CER), the MOS, and the discrimination ABX scores compared to the official ZeroSpeech 2019 baseline or even the topline.
Acknowledgements
Part of this work was supported by JSPS KAKENHI Grant Numbers JP17H06101 and JP17K00237.