Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks

Takuhiro Kaneko, Hirokazu Kameoka

Introduction

Voice conversion (VC) is a technique to modify non/para-linguistic information of speech while preserving linguistic information. This technique can be applied to various tasks such as speaker-identity modification for text-to-speech (TTS) systems , speaking assistance , speech enhancement , and pronunciation conversion .

Voice conversion can be formulated as a regression problem of estimating a mapping function from source to target speech. One successful approach involves statistical methods using a Gaussian mixture model (GMM) . Neural network (NN)-based methods, such as a restricted Boltzmann machine (RBM) , feed forward NN , recurrent NN (RNN) , and convolutional NN (CNN) , and exemplar-based methods, such as non-negative matrix factorization (NMF) , have also recently been proposed.

Many VC methods including those mentioned above typically use temporally aligned parallel data of source and target speech as training data. If perfectly aligned parallel data are available, obtaining the mapping function becomes relatively simple; however, collecting such data can be a painstaking process in real application scenarios. Even though we could collect such data, we need to perform automatic time alignment, which may occasionally fail. This can be problematic since misalignment involved in parallel data can cause speech-quality degradation; thus, careful pre-screening and manual correction may be required .

These facts motivated us to consider a VC problem that is free from parallel data. In this paper, we propose a parallel-data-free VC method, which is particularly noteworthy in that it (1) does not require any extra data, such as transcripts and reference speech, and extra modules, such as an automatic speech-recognition (ASR) module, (2) is not prone to over-smoothing, which is known to be one of the main factors leading to speech-quality degradation, and (3) captures a spectrotemporal structure without any alignment procedure.

To satisfy these requirements, our method, called CycleGAN-VC, uses a cycle-consistent adversarial network (CycleGAN) (i.e., DiscoGAN or DualGAN ) with gated CNNs and an identity-mapping loss . The CycleGAN was originally proposed for unpaired image-to-image translation. With this model, forward and inverse mappings are simultaneously learned using an adversarial loss and cycle-consistency loss . This makes it possible to find an optimal pseudo pair from unpaired data. Furthermore, the adversarial loss does not require explicit density estimation and results in reducing the over-smoothing effect . To use a CycleGAN for parallel-data-free VC, we configure a network using gated CNNs and train it with an identity-mapping loss. This allows the mapping function to capture sequential and hierarchical structures while preserving linguistic information.

We evaluated our method on a parallel-data-free VC task using the Voice Conversion Challenge 2016 (VCC 2016) dataset . An objective evaluation showed that the converted feature sequence was reasonably good in terms of global variance (GV) and modulation spectra (MS) . A subjective evaluation showed that the speech quality was comparable to that obtained with a GMM-based method trained using parallel and twice the amount of data. This is noteworthy since our method had a disadvantage in the training condition.

This paper is organized as follows. In Section 2, we describe related work. In Section 3, we review the CycleGAN and explain our proposed method (CycleGAN-VC). In Section 4, we report on the experimental results. In Section 5, we provide a discussion and conclude the paper.

Related work

Recently, several approaches for parallel-data-free VC have been proposed. One approach involves using an ASR module to find a pair of corresponding frames . This may work well if ASR performs robustly and accurately enough, but it requires a large amount of transcripts to train the ASR module. It would also be inherently difficult to capture nonverbal information. This may become a limitation to be applied in general situations. Other approaches involve methods using an adaptation technique or incorporating a pre-constructed speaker space . These methods do not require parallel data between source and target speakers but require parallel data among reference speakers. A few attempts have recently been made to develop methods that are completely free from parallel data and extra modules. With these methods, it is assumed that source and target speech lie in the same low-dimensional embeddings. This would not only limit applications but also cause difficulty in modeling complex structures, e.g., detailed spectrotemporal structure. In contrast, we learn a mapping function directly without embedding. We expect that this would make it possible to apply our method to various applications where complex structure modeling needs to be considered.

Parallel-data-free VC using CycleGAN

Our goal is to learn a mapping from source x∈Xx\in X to target y∈Yy\in Y without relying on parallel data. We solve this problem based on a CycleGAN . In this subsection, we briefly review the concept of CycleGAN and in the next subsection, we explain our proposed method for parallel-data-free VC.

With CycleGAN, a mapping GX→YG_{X\rightarrow Y} is learned using two losses, namely an adversarial loss and cycle-consistency loss . We illustrate the training procedure in Fig. 1.

Adversarial loss: An adversarial loss measures how distinguishable converted data GX→Y(x)G_{X\rightarrow Y}(x) are from target data yy. Hence, the closer the distribution of converted data PGX→Y(x)P_{G_{X\rightarrow Y}}(x) becomes to that of target data PData(y)P_{\rm Data}(y), the smaller this loss becomes. This objective is written as

The generator GX→YG_{X\rightarrow Y} attempts to generate data indistinguishable from target data yy by the discriminator DYD_{Y} by minimizing this loss, whereas DYD_{Y} attempts not to be deceived by GX→YG_{X\rightarrow Y} by maximizing this loss.

Cycle-consistency loss: Optimizing only the adversarial loss would not necessarily guarantee that the contextual information of xx and GX→Y(x)G_{X\rightarrow Y}(x) will be consistent. This is because the adversarial loss only tells us whether GX→Y(x)G_{X\rightarrow Y}(x) follows the target-data distribution and does not help preserve the contextual information of xx. The idea of CycleGAN is to introduce two additional terms. One is an adversarial loss Ladv(GY→X,DX){\cal L}_{adv}(G_{Y\rightarrow X},D_{X}) for an inverse mapping GY→XG_{Y\rightarrow X} and the other is a cycle-consistency loss, given as

These additional terms encourage GX→YG_{X\rightarrow Y} and GY→XG_{Y\rightarrow X} to find (x,y)(x,y) pairs with the same contextual information.

The full objective is written with trade-off parameter λcyc\lambda_{cyc}:

2 CycleGAN for parallel-data-free VC: CycleGAN-VC

To use a CycleGAN for parallel-data-free VC, we mainly made two modifications to the CycleGAN architecture: gated CNN and identity-mapping loss .

Gated CNN: One of the characteristics of speech is that it has sequential and hierarchical structures, e.g., voiced/unvoiced segments and phonemes/morphemes. An effective way to represent such structures would be to use an RNN, but it is computationally demanding due to the difficulty of parallel implementations. Instead, we configure a CycleGAN using gated CNNs , which not only allows parallelization over sequential data but also achieves state-of-the-art in language modeling and speech modeling . In a gated CNN, gated linear units (GLUs) are used as an activation function. A GLU is a data-driven activation function, and the (l+1)(l+1)-th layer output Hl+1\bm{H}_{l+1} is calculated using the ll-th layer output Hl\bm{H}_{l} and model parameters Wl\bm{W}_{l}, Vl\bm{V}_{l}, bl\bm{b}_{l}, and cl\bm{c}_{l},

where ⊗\otimes is the element-wise product and σ\sigma is the sigmoid function. This gated mechanism allows the information to be selectively propagated depending on the previous layer states.

Identity-mapping loss: A cycle-consistency loss provides constraints on a structure; however, it would not suffice to guarantee that the mappings always preserve linguistic information. To encourage linguistic-information preservation without relying on extra modules, we incorporate an identity-mapping loss ,

which encourages the generator to find the mapping that preserves composition between the input and output. In practice, weighted loss λidLid\lambda_{id}{\cal L}_{id} with trade-off parameter λid\lambda_{id} is added to Eq. 3.1. Note that the original study on CycleGANs showed the effectiveness of this loss for color preservation.

Experiments

We conducted experiments to evaluate our method on a parallel-data-free VC task. We used the VCC 2016 dataset , which was recorded by professional US English speakers, including five females and five males. Following a previous study , we used a subset of speakers for evaluation. A pair of female (SF1) and male (SM1) speakers were selected as sources and another pair (TF2 and TM3) were selected as targets. The audio files for each speaker were manually segmented into 216 short parallel sentences (about 13 minutes). Among them, 162 and 54 sentences were provided as training and evaluation sets, respectively. To evaluate our method under a parallel-data-free condition, we divided the training set into two subsets without overlap. For the first half, 81 sentences were used for the source and the other 81 sentences were used for the target. The speech data were downsampled to 16 kHz, and 24 Mel-cepstral coefficients (MCEPs), logarithmic fundamental frequency (log⁡F0\log F_{0}), and aperiodicities (APs) were then extracted every 5 ms using the WORLD analysis system . Among these features, we learned a mapping in the MCEP domain using our method. The F0F_{0} was converted using logarithm Gaussian normalized transformation . Aperiodicities were directly used without modification because a previous study showed that converting APs does not significantly affect speech quality.

Implementation details: We designed a network based on the recent success in image modeling and speech modeling . The network architecture is illustrated in Fig. 2. We designed the generator using a one-dimensional (1D) CNN to capture the relationship among the overall features while preserving the temporal structure. Inspired by a previous study for neural style transfer and super-resolution, we used the network that included downsampling, residual , and upsampling layers, as well as incorporating instance normalization . We used pixel shuffler for upsampling, which is effective for high-resolution image generation . We designed the discriminator using a 2D CNN to focus on a 2D spectral texture .

Training details: As a pre-process, we normalized the source and target MCEPs per dimension. To stabilize training, we used a least squares GAN , which replaces the negative log likelihood objective in Ladv{\cal L}_{adv} by a least squares loss. We set λcyc=10\lambda_{cyc}=10. We used Lid{\cal L}_{id} only for the first 10410^{4} iterations with λid=5\lambda_{id}=5 to guide the learning process. To increase the randomness of each batch, we did not use a sequence directly and cropped a fixed-length segment (128 frames) randomly from a randomly selected audio file. We trained the network using the Adam optimizer with a batch size of 11. We set the initial learning rates to 0.00020.0002 for the generator and 0.00010.0001 for the discriminator. We kept the same learning rate for the first 2×1052\times 10^{5} iterations and linearly decay over the next 2×1052\times 10^{5} iterations. We set the momentum term β1\beta_{1} to 0.50.5.

2 Objective evaluation

In these experiments, we focused on the conversion of MCEPs; therefore, we evaluated the quality of converted MCEPs. We compared our method (CycleGAN-VC) with a GMM-based method (GMM-VC) because it is still comparable to a DNN-based method in a relatively small dataset . Since this method requires parallel data, all the training data (162 sentences) for both source and target were used. This means that the amount of training data was twice as ours. As an ablation study, we examined our method without any GLUs. Instead of GLUs, we used typical GAN activation functions, i.e., rectified linear units (ReLUs) for the generator and leaky ReLUs for the discriminator. In the pre-experiment, we also examined our method without an identity-mapping loss. This revealed that the lack of this loss tends to cause significant degradation, e.g., collapse of the linguistic structure; thus, we did not examine this further.

Mel-cepstral distortion is a well-used measure to evaluate the quality of synthesized MCEPs, but recent studies indicate the limitation of this measure: it tends to prefer over-smoothing because it internally assumes Gaussian distribution. Therefore, as alternatives, we used two structural indicators highly correlated with subjective evaluation: GV and MS . We show the comparison of GV in Fig. 3. We list the comparison of root mean squared error (RMSE) between target and converted logarithmic MS in Table 1. We also show the comparison of MS per modulation frequency in Fig. 4. These results indicate that the MCEP sequences obtained with our method (CycleGAN-VC w/ GLU) are closest to the target in terms of GV and MS. We expect this is because (1) the adversarial loss does not require explicit density estimation; thus, avoids over-smoothing, and (2) the GLU is a data-driven activation function; therefore, it can represent sequential and hierarchical structures better than the ReLU and leaky ReLU. We show sample MCEP trajectories in Fig. 5. The trajectories of CycleGAN-VC w/ GLU have a similar global structure to those of GMM-VC w/ GV while preserving similar complexity to the source.

3 Subjective evaluation

We conducted listening tests to evaluate the performance of converted speechWe provide the converted speech samples at http://www.kecl. ntt.co.jp/people/kaneko.takuhiro/projects/ cyclegan-vc. By referring to the VCC 2016 , we evaluated the naturalness and speaker similarity of the converted samples. We compared our method with the baseline of the VCC 2016We used data at http://dx.doi.org/10.7488/ds/1575, which is a GMM-based method using parallel and twice the amount of data. To measure naturalness, we conducted a mean opinion score (MOS) test. As a reference, we used original and synthesized-and-analyzed (upper bound of our method) speeches of target speakers. Twenty sentences longer than 2 s and shorter than 5 s were randomly selected from the evaluation sets. To measure speaker similarity, we used the same/different paradigm . Ten sample pairs were randomly selected from the evaluation sets. There were nine participants who were well-educated English speakers. By referring to the study by , we evaluated on two subsets: intra-gender VC (SF1–TF2) and inter-gender VC (SF1–TM3).

We show the MOS for naturalness in Fig. 6. The results indicate that the proposed method significantly outperformed the baseline. We show the similarity to a source speaker and to a target speaker in Fig. 7. The results indicate that our method was slightly inferior to the baseline in SF1–TM3 VC but superior in SF1–TF2 VC. Overall, our method is comparable to the baseline. This is noteworthy since our method is trained under disadvantageous conditions with half the amount of and non-parallel data.

Discussion and conclusions

We proposed a parallel-data-free VC method called CycleGAN-VC, which uses a CycleGAN with gated CNNs and an identity-mapping loss. This method can learn a sequence-based mapping function without any extra data, modules, and time alignment procedure. An objective evaluation showed that the MCEP sequences obtained with our method are close to the target in terms of GV and MS. A subjective evaluation showed that the quality of converted speech was comparable to that obtained with the GMM-based method under advantageous conditions with parallel and twice the amount of data. However, there is still a margin between original and converted speeches. To fill the margin, we plan to apply our method to other features, such as STFT spectrograms , and other speech-synthesis frameworks, such as vocoder-free VC . Furthermore, our proposed method is a general framework, and possible future work includes applying the method to other VC applications .

Acknowledgements: This work was supported by JSPS KAKENHI Grant Number 17H01763.

References