Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
Andy T. Liu, Po-chun Hsu, Hung-yi Lee
Introduction
Despite that human speech inherently carries linguistic features that represent textual information, modern text-to-speech training pipelines still require parallel speech and text transcription pairs . Parallel speech and transcripts may not always be available, and is costly to acquire, however human speech alone can be easily gathered. In this work, we focus on utilizing the advantage of unlabeled speech to discover discrete linguistic units, where machine learns to uncover the linguistic features hidden in human utterance without any supervision. We use these discovered linguistic units for voice conversion (VC) and achieved outstanding results.
Embedding audio signals into latent representations has been a well studied practice . Former studies have also attempted to encode speech content into various representations for voice conversion, including the use of disentangled autoencoders , VAEs or GANs . continuous vectors are the most common approach, however when it comes to encoding speech content, they may not be the best choice. As they are unlike the discrete phonemes that we often used to represent human language. Previous works also attempt to learn discrete representations from audio, however they did not apply the learned units for VC.
In VC, the state-of-the-art approach requires additional loss in the framework of GAN to guarantee the disentanglement of learned encodings, whereas the proposed approach naturally possesses the direct ability to separate speaker style from speech content. In , VC is achieved through multiple losses together with a K-way quantized embedding space and autoregressive WaveNet, on the other hand, the proposed encoding space does not require additional training constraints.
In this work, we use an ASR-TTS autoencoder to discovers discrete linguistic units from speech, without any alignment, text label, or parallel data provided. The ASR-Encoder learns to encode speech from different speakers to a common set of small linguistic symbols. These finite set of linguistic symbols are represented by Multilabel-Binary Vectors (MBV), vectors consist of arbitrary number of zeros and ones. The proposed MBV method is differentiable, hence allowing backpropagation of gradients and end-to-end training of an autoencoder reconstruction setting. While the ASR-Encoder learns a many-to-one mapping from speech to discrete subword units, a TTS-Decoder learns a one-to-many mapping from discrete subword units back to speech. The discrete nature of MBV allows it to innately separate linguistic content from speech, removing speaker characteristics. Given an utterance of a source speaker, we were able to encode its speech content using the ASR-Encoder, and perform voice conversion to generate speech with the same linguistic content but style of a target speaker using the TTS-Decoder.
Proposed method
We present an unsupervised end-to-end learning framework for distinct linguistic units discovery. In this stage we learn a common set of discrete encodings for all the speakers. Let be an acoustic feature sequence where is a set of all such sequences from all source and target speakers, and is a fixed length- segment randomly sampled from . Let be a speaker where is the group of all source and target speakers who produce the sequence collection . In a training pair , the sequence is produced by speaker .
In Figure 1, an ASR-TTS autoencoder framework is used to learn discrete linguistic units. The ASR-Encoder is trained to map input acoustic feature sequence to a latent discrete encoding representation:
where we propose to use Multilabel-Binary Vectors (MBV) to represent generated in (1), the discrete encoding is designed to represents the linguistic content of input speech . We define MBV as:
1.2 Voice reconstruction and conversion
Given ASR-Encoder’s output, the TTS-Decoder is trained to generate an output acoustic feature sequence defined as:
where is a reconstruction of from given the speaker identity . Mean Absolute Error (MAE) is employed to evaluate the ASR-TTS autoencoder reconstruction loss, as MAE is reported to be able to generate sharper outputs than Mean Square Error . The reconstruction loss is given by:
where and are the parameters of the ASR-Encoder and TTS-Decoder, respectively. We uniformly sample for training in (4). Because the speaker identity is provided to the TTS-Decoder, the proposed MBV encodings is able to learn an abstract space that is invariant to speaker identity and only encodes the content of speech, without using any form of linguistic supervision.
At inference time, given a source speech , and target speaker , the TTS-Decoder can generate the voice of the target speaker using the linguistic content from :
is the output of TTS-Decoder, which has the linguistic content of but style of speaker .
2 Target guided adversarial learning
With the learned speaker invariant ASR-Encoder mapping in (1), we successfully represent speech content with discrete binary vectors MBVs. In this section, we describe how adversarial training is used to boost VC quality based on the discrete linguistic units learned in Section 2.1. We train a TTS-Patcher in an unsupervised manner, in which the TTS-Patcher generates a spectrum mask that residually augments the output of equation (5), resulting in a more precise voice conversion result.
We define two sets of speakers, source speaker set and target speaker set, where we aim to convert the speech of a source speaker into a target speaker’s style while preserving its content. Let and be sequences where and are the set of sequences from source speakers and target speakers, respectively. Let be a speaker from the set of all source speakers who produce , and let be a speaker from the set of all target speakers who produce .
In Figure 2, a TTS-Patcher is trained as a generator under an adversarial learning setting. The TTS-Patcher takes and as input, and generates a spectrum mask that ranges from zero to one, and modifies through a residual augmentation:
The and symbols indicate element-wise addition and multiplication, respectively. In (6), is a randomly sampled input speech segment from source speaker , where is a randomly sampled target speaker. is the voice conversion utterance obtained in (5), is the output of TTS-Patcher, and finally the augmented spectrum.
A discriminator is trained to distinguish whether an input acoustic feature sequence is real or reconstructed by machine. Since naive GAN is notoriously hard to train, we minimize the Wasserstein-1 distance between real and fake distributions instead, as proposed in the WGAN formulation:
The discriminator computes the Wasserstein-1 distance of two distributions: real data sampled from the target speaker set , and augmented voice conversion outputs from (6). We use the alternative WGAN-GP to enforce the 1-Lipschitz constraint required by , where weight clipping is replaced with gradient penalty. On the last layer of the discriminator, we stretch an additional layer that learns a classifier to predict speaker from a given speech. This allows the discriminator to consider input spectrum’s fidelity and speaker identity at the same time .
2.2 Target guided training step
The decoupled learning of ASR-TTS autoencoder and TTS-Patcher stabilizes the GAN training process. However we found that under the adversarial learning scheme, the TTS-Patcher can easily learn to deceive the discriminator by over-adding style, this greatly compromises the original speech content. This is caused by the discriminator’s inability to discriminate utterances with incorrect or ambiguous content, the discriminator only learns to focus on speaker style. As a solution, we propose to add an additional target guided training step, we apply additional reconstruction loss after every adversarial step, as shown in Figure 2. Instead of converting , the ASR-Encoder now takes a segment of target speech as input, equation (6) then becomes:
and we minimize MAE between and :
where is the parameter of the TTS-Patcher. This loss effectively guides the TTS-Patcher’s update under adversarial settings, as the added style is regularized to preserve intelligibility.
Implementation
The ASR-Encoder is inspired by the CBHG module , where the linear output of ASR-Encoder is fed to the MBV encoding module. We add noise in training by adding dropout layers in the ASR-Encoder as suggested in . The TTS-Decoder and TTS-Patcher have identical model architectures, where we use pixel shuffle layers to generate high resolution spectrum . We add speaker embedding on the feature map of all layers, where a distinct embedding is learned for all different layers as different information may be needed for each layer. The discriminator is consist of 2D-convolution blocks for temporal texture capturing, and convolutions projection layers followed by fully-connected output layers. We trained the network using Adam optimizer and a batch size of 16. In the discrete linguistic units discovery stage, we train the ASR-TTS autoencoder for 200k mini-batches. In the target guided adversarial learning stage we train the model for 50k mini-batches, in one batch we train a step of adversarial learning including 5 iterations of discriminator update and 1 iteration of generator update, followed by a target guided reconstruction step.
We train and evaluate our model using the ZeroSpeech 2019 English dataset . In particular, we use the “Train Unit Dataset” as our source speaker set , the “Train Voice Dataset” as our target speaker set , and we evaluate models with the “Test Dataset”. We used log-magnitude spectrograms as acoustic features, the detailed settings are in Table 1. During training, our model is trained to process 128 consecutive overlapping frames of spectrogram, where we uniformly sample from the dataset. At inference time, for a given input with more than 128 frames, the model process them as segments and concatenate the outputs on the time-axis. Source code are publicly availablehttps://github.com/andi611/ZeroSpeech-TTS-without-T.
Experiments
We compare the proposed MBV encodings with one-hot encodings, continuous encodings, and continuous encodings with additional loss , all of which under the same autoencoder training setting as described in Section 3.
2 Subjective and objective evaluation
We perform subjective human evaluation on the converted voices. We use 20 subjects to grade each method on a 1 to 5 scale under two measures: the naturalness of speech and the similarity in speaker characteristics to the target speaker. In Table 3 we show the result of our evaluation, the proposed method results in significant increase of target similarity with a slight degrade of naturalness. We easily achieve comparable speech intelligibility as ordinary continuous methods, while achieving better voice conversion quality with more disentanglement (Table 2). Subjective and objective evaluations suggest that the proposed MBV method eliminates speaker identity while reserving content within speech, and is suitable for voice conversion.
3 Encoding dimension analysis
We use several objective measures to determine the quality of an encoding, these measurements are shown in Table 4. The column is the output Character Error Rate (CER) from a pre-trained ASR, where we use the ASR results of real input voice as ground truth, which measures intelligibility of the converted speech. The column measures the bitrate (amount of information) that encodings carry in average with respect to the testing set, as suggested in . The column measures the machine ABX score, which indicates the goodness of encoding quality . The column indicates the number of unique symbols used to encode speech in the test set. Lower values suggest a better performance for all the measures described above. In Table 4, we compare the proposed method with other approaches along side with the baseline model demonstrated in . When compared to other approaches, the proposed method achieves lower and values with comparable scores. Due to the discrete and differentiable nature of MBV, the proposed method can be used in other unsupervised end-to-end clustering or classification tasks, where other approaches may fail to generalize.
4 The zero resource speech challenge competition
Conclusions
We proposed to use multilabel-binary vectors to represent the content of human speech, as its discrete nature offers a strong extraction of speaker-independent representation. We show that these discrete units naturally possess the ability of disentangling speech content and style, which makes them extremely suitable for voice conversion tasks. Also, we show that these discrete units indeed produce better style disentanglement than ordinary settings, and finally we were able to improve voice conversion results through the addition of residual augmented signals.