Voice Conversion from Unaligned Corpora using Variational Autoencoding Wasserstein Generative Adversarial Networks

Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, Hsin-Min Wang

Introduction

The primary goal of voice conversion (VC) is to convert the speech from a source speaker to that of a target, without changing the linguistic or phonetic content. However, consider the case of converting one’s voice into that of another who speaks a different language. Traditional VC techniques would have trouble dealing with such cases because most of them require parallel training data in which many pairs of speakers uttered the same texts. In this paper, we are devoted to bridging the gap between parallel and non-parallel VC systems.

We pursue a unified generative model for speech that naturally accommodates VC (Sec. 2). In this framework, we do not have to align any frames or to cluster phones or frames explicitly. The idea is skeletonized in the probabilistic graphical model (PGM) in Fig. 1. In this model, our attention is directed away from seeking alignment. Rather, what we are concerned about are 1) finding a good inference model (Sec. 2.1) for the latent variable z\bm{{z}} and 2) building a good synthesizer whose outputs match the distribution of the real speech of the target (Sec. 2.2 through 2.3). In this paper, we present a specific implementation in which a variational autoencoder (VAE ) assumes the inference task and a Wasserstein generative adversarial network (W-GAN ) undertakes speech synthesis. Our contribution is two-fold:

We introduce the W-GAN to non-parallel voice conversion and elucidate the reason why it fits this task (Sec. 2).

We demonstrate the ability of W-GAN to synthesize more realistic spectra (Sec. 3).

Non-parallel voice conversion via deep generative models

Given spectral frames Xs={xs,n}n=1NsX_{s}=\{\bm{{x}}_{s,n}\}^{N_{s}}_{n=1} from the source speaker and those Xt={xt,n′}n′=1NtX_{t}=\{\bm{{x}}_{t,n^{\prime}}\}^{N_{t}}_{n^{\prime}=1} from the target, assume that the real data distributions of the source and the target respectively admit a density ps∗p^{*}_{s} and pt∗p^{*}_{t}. Let ff be a voice conversion function that induces a conditional distribution pt∣sp_{t\mid s}. The goal of VC is to estimate ff so that pt∣sp_{t\mid s} best approximates the real data distribution pt∗p^{*}_{t}:

We can decompose the VC function ff into two stages according to the PGM in Fig. 1(b). In the first stage, a speaker-independent encoder Eϕ\mathcal{E}_{\phi} infers a latent content zn\bm{{z}}_{n}. In the second stage, a speaker dependent decoder Gθ\mathcal{G}_{\theta} mixes zn\bm{{z}}_{n} with a speaker-specific variable y\bm{{y}} to reconstruct the input. The problem of VC is then reformulated as:

In short, this model explains the observation x\bm{{x}} using two latent variables y\bm{{y}} and z\bm{{z}}. We will drop the frame indices whenever readability is unharmed. We refer to y\bm{{y}} as the speaker representation vector because it is determined solely by the speaker identity. We refer to z\bm{{z}} as the phonetic content vector because with a fixed y\bm{{y}}, we can generate that speaker’s voice by varying z\bm{{z}}. Note that the term phonetic content is only valid in the context of our experimental settings where the speech is natural, noise-free, and non-emotional.

This encoder-decoder architecture facilitates VC from unaligned or non-parallel corpora. The function of the encoder is similar to a phone recognizer whereas the decoder operates as a synthesizer. The architecture enables voice conversion for the following reasons. The speaker representation y\bm{{y}} can be obtained from training. The encoder can infer the phonetic content z\bm{{z}}. The synthesizer can reconstruct any spectral frame x\bm{{x}} with the corresponding y\bm{{y}} and z\bm{{z}}. Combining these elements, we can build a non-parallel VC system via optimizing the encoder, the decoder (synthesizer), and the speaker representation. With the encoder, frame-wise alignment is no longer needed; frames that belong to the same phoneme class now hinge on a similar z\bm{{z}}. With this conditional synthesizer, VC becomes as simple as replacing the speaker representation y\bm{{y}}. (as illustrated in Fig. 1(c)).

We delineate our proposed method incrementally in three subsections: a conditional variational autoencoder (C-VAE) in Sec. 2.1, a generative adversarial nets (GAN) applied to improve the over-simplified C-VAE model in Sec. 2.2, and a Wasserstein GAN (W-GAN) that explicitly considers VC in the training objectives in Sec. 2.3.

Recent works have proven the viability of speech modeling with VAEs . A C-VAE that realizes the PGM in Fig. 1(a) maximizes a variational lower bound of the log-likelihood:

In order to train the C-VAE, we have to simplify the model in several aspects. First, we choose pθ(x∣z,y)p_{\theta}(\bm{{x}}|\bm{{z}},\bm{{y}}) to be a normal distribution whose covariance is an identity matrix. Second, we choose pθ(z)p_{\theta}(\bm{{z}}) to be a standard normal distribution. Third, the expectation over z\bm{{z}} is approximated by sampling methods. With these simplifications, we can avoid intractability and focus on modeling the statistics of the Gaussian distributions.

For the phonetic content z\bm{{z}}, we have:

where ϕ1∪ϕ2=ϕ{\phi_{1}}\cup{\phi_{2}}=\phi are the parameters of the encoder, and Eϕ1\mathcal{E}_{\phi_{1}} and Eϕ2\mathcal{E}_{\phi_{2}} are the inference models of mean and variance. For the reconstructed or converted spectral frames, we have:

Training this C-VAE means maximizing (3). For every input (xn,yn)(\bm{{x}}_{n},\bm{{y}}_{n}), we can sample the latent variable zn\bm{{z}}_{n} using the re-parameterization trick described in . With zn\bm{{z}}_{n} and yn\bm{{y}}_{n}, the model can reconstruct the input, and by replacing yn\bm{{y}}_{n}, it can convert voice. This means that we are building virtually multiple models in one. Conceptually, the speaker switch lies in the speaker representation y\bm{{y}} because the synthesis is conditioned on y\bm{{y}}.

2 Improving speech models with GANs

Despite the effectiveness of C-VAE, the simplification induces inaccuracy in the synthesis model. This defect originates from the fallible assumption that the observed data is normally distributed and uncorrelated across dimensions. This assumption gave us a defective learning objective, leading to muffled converted voices. Therefore, we are motivated to resort to models that side-step this defect.

where Dψ∗\mathcal{D}_{\psi}^{*} denotes the optimal discriminator, which is the density ratio in the second equality in (8). We can view this as a density ratio estimation problem without explicit specification of distributions.

Presumably, GANs produce sharper spectra because they optimize a loss function between two distributions in a more direct fashion. We can combine the objectives of VAE and GAN by assigning VAE’s decoder as GAN’s generator to form a VAE-GAN . However, the VAE-GAN does not consider VC explicitly. Therefore, we propose our final model: variational autoencoding Wasserstein GAN (VAW-GAN).

3 Direct consideration of voice conversion with W-GAN

The Wasserstein-1 distance is defined as follows:

where Π(p,q)\Pi(p,q) denotes the set of all joint distributions γ(x,x^)\gamma(\bm{{x}},\hat{\bm{{x}}}) whose marginals are respectively pt∗p^{*}_{t} and pt∣sp_{t\mid s}. According to the definition, the Wasserstein distance is calculated from the optimal transport, or the best frame alignment. Note that (9) is thus suitable for parallel VC.

On the other hand, the Kantorovich-Rubinstein duality of (9) allows us to explicitly approach non-parallel VC:

where the supremum is over all 1-Lipschitz functions D:X→R\mathcal{D}:\mathcal{X}\to R. If we have a parameterized family of functions Dψ∈Ψ\mathcal{D}_{\psi\in\Psi} that are all K-Lipschitz for some K, we could consider solving the problem:

Alignment is not required in this formulation because of the respective expectations. What we need now is a batch of real frames from the target speaker, another batch of synthetic frames converted from the source into the target speaker, and a good discriminator Dψ\mathcal{D}_{\psi}.

3.2 VAW-GAN

Incorporating the W-GAN loss (12) with (3) yields our final objective:

where α\alpha is a coefficient which emphasizes the W-GAN loss. This objective is shared across all three components: the encoder, the synthesizer, and the discriminator. The synthesizer minimizes this loss whereas the discriminator maximizes it; consequently the two components have to be optimized in alternating order. For clarity, we summarize the training procedures in Alg. 1. Note that we actually use an update schedule for Dψ\mathcal{D}_{\psi} instead of training it to real optimality.

Experiments

The proposed VC system was evaluated on the Voice Conversion Challenge 2016 dataset . The dataset was a parallel speech corpus; however, frame alignment was not performed in the following experiments.

We conducted experiments on a subset of 3 speakers. In the inter-gender experiment, we chose SF1 as the source and TM3 as the target. In the intra-gender experiment, we chose TF2 as the target. We used the first 150 utterances (around 10 minutes) per speaker for training, the succeeding 12 for validation, and 25 (out of 54) utterances in the official testing set for subjective evaluations.

2 The feature set

We used the STRAIGHT toolkit to extract speech parameters, including the STRAIGHT spectra (SP for short), aperiodicity (AP), and pitch contours (F0). The rest of the experimental settings were the same as in , except that we rescaled log energy-normalized SP (denoted by logSPenlog{SP}_{en}) to the range of $$ dimension-wise. Note that our system performed frame-by-frame conversion without post-filtering and that we utilized neither contextual nor dynamic features in our experiments.

3 Configurations and hyper-parameters

The baseline system was the C-VAE system (denoted simply as VAE) because its performance had been proven to be on par with another simple parallel baseline. In our proposed system, the encoder, the synthesizer, and the discriminator were convolutional neural networks. The phonetic space was 64-dimensional and assumed to have a standard normal distribution. The speaker representation were one-hot coded, and their embeddings were optimized as part of the generator parameters Due to space limitations, the rest of the specification of hyper-parameters and audio samples can be found on-line: https://github.com/JeremyCCHsu/vc-vawgan.

4 The training and conversion procedures

We first set α\alpha to 0 to exclude W-GAN, and trained the VAE till convergence to get the baseline model. Then, we proceeded on training the whole VAW-GAN via setting α\alpha to 50.

5 Subjective evaluations

Five-point mean opinion score (MOS) tests were conducted in a pairwise manner. Each of the 10 listeners graded the pairs of outputs from the VAW-GAN and the VAE. Inter-gender and intra-gender VC were evaluated respectively.

The MOS results on naturalness shown in Fig. 2 demonstrate that VAW-GAN significantly outperforms the VAE baseline (p-value ≪\ll 0.01 in paired t-tests). The results are in accordance with the converted spectra shown in Fig. 3, where the output spectra from VAW-GAN express richer variability across the frequency axis, hence reflecting clearer voices and enhanced intelligibility.

We did not report objective evaluations such as mean mel-cepstral coefficients because we found inconsistent results with the subjective evaluations. However, similar inconsistency is common in the VC literature because it is highly likely that those evaluations are inconsistent with human auditory systems . The performance of speaker similarity was also unreported because we found that it remained about the same as that of (System B in ).

Discussions

As we can see in Fig. 3, the spectral envelopes of the synthetic speech from VAW-GAN are more structured, with more observable peaks and troughs. Spectral structures are key to the speech intelligibility, indirectly contributing to the elevated MOS. In addition, the more detailed spectral shapes in the high-frequency region reflect clearer (non-muffled) voice of the synthetic speech.

2 W-GAN as a variance modeling alternative

The Wasserstein objective in (13) is minimized when the distribution of the converted spectrum pt∣sp_{t\mid s} is closest to the true data distribution pt∗p^{*}_{t}. Unlike VAE that assumes a Gaussian distribution on the observation, W-GAN models the observation implicitly through a series of stochastic procedures, without prescribing any density forms. In Fig. 4, we can observe that the output spectra of the VAW-GAN system have larger variance compared to those of the VAE system. The global variance (GV) of the VAW-GAN output may not be as good as that of the data but the higher values indicate that VAW-GAN does not centralize predicted values at the mean too severely. Since speech has a highly diverse distribution, it requires more sophisticated analysis on this phenomenon.

3 Imperfect speaker modeling in VAW-GAN

The reason that the speaker similarity of the converted voice is not improved reminds us of the fact that both VAE and VAW-GAN optimize the same PGM, thus the same speaker model. Therefore, modeling speaker with one global variable might be insufficient. As modeling speaker with a frame-wise variable may conflict with the phonetic vector z\bm{{z}}, we may have to resort to other PGMs. We will investigate this problem in the future.

Related work

To handle non-parallel VC, many researchers resort to frame-based, segment-based, or cluster-based alignment schemes. One of the most intuitive ways is to apply an automatic speech recognition (ASR) module to the utterances, and proceed with explicit alignment or model adaptation . The ASR module provides every frame with a phonetic label (usually the phonemic states). It is particularly suitable for text-to-speech (TTS) systems because they can readily utilize these labeled frames . A shortcoming with these approaches is that they require an extra mapping to realize cross-lingual VC. To this end, the INCA-based algorithms were proposed to iteratively seek frame-wise correspondence using converted surrogate frames. Another attempt is to separately build frame clusters for the source and the target, and then set up a mapping between them .

Recent advances include , in which the authors exploited i-vectors to represent speakers. Their work differed from ours in that they adopted explicit alignment during training. In , the authors represented the phonetic space with senone probabilities outputted from an ASR module, and then generated voice by means of a TTS module. Despite differences in realization, our models do share some similarity ideally.

Conclusions

We have presented a voice conversion framework that is able to directly incorporate a non-parallel VC criterion into the objective function. The proposed VAW-GAN framework improves the outputs with more realistic spectral shapes. Experimental results demonstrate significantly improved performance over the baseline system.

Acknowledgements

This work was supported in part by the Ministry of Science and Technology of Taiwan under Grant: MOST 105-2221-E-001-012-MY3.

References