A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation
Ehsan Hosseini-Asl, Yingbo Zhou, Caiming Xiong, Richard Socher
Introduction
Neural-based acoustic models have shown promising improvements in building automatic speech recognition (ASR) systems . However, it tends to perform poorly when evaluated on out-of-domain data, because of mismatch between the training and testing distribution (Table 1).
Domain mismatch is mainly due to variation in non-linguistic features, such as different speaker identity, unseen environmental noise, large accent variations, etc. Therefore, training a robust ASR system is highly dependent on factorizing linguistic features (text) from non-related variations, or adapting the inter-domain variations of source and target.
Voice conversion (VC) has been widely used to adapt the non-linguistic variations, such as statistical methods , and Neural-based models . However, traditional VC methods require parallel data of source and target that is difficult to obtain in practice. In addition, the requirement of parallel data prevent these methods from using more abundant unsupervised data. Therefore, an unsupervised domain adaptation is desirable for building a robust ASR system.
In this paper, we propose a new generative model based on CycleGAN for unsupervised non-parallel domain adaptation. Since differences in magnitude of frequency is the main variation across domains for spectrogram representations, it is imperative that CycleGAN correctly catch the spectro-temporal variations between different frequency bands across domains during training. This will allow the generator to learn the mapping function which can convert spectrogram from source to target domain. In this paper, we show that the original CycleGAN model is failing to learn the correct mapping function between domains, and the generator collapses into learning an identity mapping function, which results in generating a noisy and unnatural-sounding audio.
To accommodate generative adversarial network for training on non-parallel spectrogram domains, the generator should be back-propagated with multiple gradient signals (from different discriminators), that each represents the variations between source and target domains at a specific frequency band. To achieve this goal, we propose to use multiple and independent discriminators for each domain, similar to generative multi adversarial network (GMAN) . We show that the proposed Multi-Discriminator CycleGAN, without pretraining the discriminators, outperforms CycleGAN with pretrained discriminator, for spectrogram adaptation. Furthermore, we show that the multi discriminator architecture can overcome the checkerboard artifacts problem caused by deconvolution layer in generator , and generates natural clean audio. To evaluate the performance of the proposed model, gender-based domains are selected as domain adaptations.
Generative Adversarial Network (GAN) is a family of non-parametric density estimation models which learn to model the data generating distribution using adversarial training. Conditional GANs (CGAN) was first proposed for supervised (parallel) domain adaptation, where the goal is to convert source distribution to match the target. CGAN has been used in various data domains, especially image domains, both for parallel and non-parallel domain adaptation .
Recently, CGAN is used for speech enhancement on parallel datasets . Speech denoising is achieved by conditioning the generator on noisy speech to learn the de-noised version . Donahue et al. proposed a GAN model on audio (WaveGAN) and spectrogram (SpecGAN), which is actually trained CGAN on parallel domains. Kaneko et al. proposed a cycle-consistent adversarial network (CycleGAN) with gated convolutional neural network (CNN) as the generator part, where the model is trained on Mel-cepstral coefficients (MCEPs) features. Hsu et al. proposed a combination of variational inference network, using variational autoencoder (VAE) , and adversarial network, using Wasserstein GAN (WGAN) . In , the goal is to disentangle the linguistic from nuisance latent variables via VAE using spectra (SP for short), aperiodicity (AP), and pitch contours (F0) features, followed by adversarial training to learn the target distribution from the inferred linguistic latent distribution. A recurrent VAE is also proposed to capture the temporal relationships in the disentangled representation of sequence data, using Mel-scale filter bank (FBank).
Contributions of the proposed generative model are, (1) It is a robust GAN model developed for non-parallel unsupervised domains, compared to parallel-based SpecGAN and WaveGAN , (2) The choice of multiple discriminator is adjustable to the spectro-temporal structure of the intended domains, compared to domain-specific model design of , (3) Proposed GAN model training is robust and invariant to the choice of adversarial objective, i. e. binary cross-entropy or least square (LS-GAN ), while the CycleGAN in is only stable using least square loss, with additional using of identity mapping loss in generator, (4) Source and target domains in are sampled from same speakers, both including male and female, only uttering different sentences, while our approach is more natural as source and targets distribution is strongly diverged due to different speaker, gender, and uttered sentences. (5) Compared to FHVAE , our models improves ASR performance on TIMIT female set by PER (Table 3), when only trained on male.
Proposed Model
In this section, the proposed model is explained. We first describe the generative model based on adversarial network. Generative Model based on adversarial training (GAN) has been proposed by Goodfellow et al. to model the data generating distribution. Training GAN is based on minimizing Jensen-Shannon divergence between data generating distribution and model data distribution . Learning is through minimization of the adversarial loss between generator network , which learns a mapping function , and discriminator network . The generator is learning to model the data distribution by generating indistinguishable samples from , using a source noise signal to minimize (1), whereas discriminator is learning to discriminate between real data and generated by maximizing the adversarial loss,
For domain adaptation between parallel domains and , Conditional GAN (CGAN) is proposed, using a generator that directly learns the mapping function , by minimizing parallel conditional adversarial loss ,
where is discriminating between pair of real parallel data and generated pair . To apply CGAN for adaptation between non-parallel domains and , a conditional GAN using cycle consistent adversarial loss (CycleGAN) has been proposed . In CycleGAN , there are two conditional generators, i. e., and , each trained in adversarial setting with and , respectively. In other words, there are two pairs of Non-parallel conditional adversarial loss and , where,
Therefore, CycleGAN learns unsupervised mapping functions between and domains by combining (LABEL:eq:cgan-unparallel) and (4), to maximize the adversarial loss , where,
2 Multi-Discriminator CycleGAN (MD-CycleGAN)
In this section, we propose a multiple discriminator generative model based on cycle consistency loss (5). The model is based on generative multi adversarial network (GMAN) . In this paper, and represents spectrogram feature datasets of different speech domains. Spectrogram feature represents the frequency variation of audio data through time dimension. In order to allow CycleGAN to learn the mapping function of spectrogram between different speech domains, the generators should be able to learn the variations in each frequency band for each aligned time window, across domains.
In order to learn the frequency-dependent mapping functions that catch the variation per each frequency bands, we define multiple frequency-dependent discriminators , where represents the frequency band of domain with frequency bands, and represents -th frequency band od domain , respectively. The frequency band definition in each domain can share a portion of frequency spectrum, or be exclusive, based on the domain spectrogram distribution. We are also using the non-saturating version of GAN, NS-GAN, where the generator is learned through maximizing the probability of predicting generated samples as drawn from data generating distribution . Accordingly, the adversarial loss for each pair of generator and discriminator in (LABEL:eq:cgan-unparallel) and (5) is
The Multi-Discriminator CycleGAN (MD-CycleGAN) is training by maximizing , where,
A natural extension to the proposed MD-CycleGAN is to use multiple frequency-dependent generators jointly with discriminators as well. This can follow in two configurations. In one-one setting, each generator is trained on a specific frequency band with the corresponding discriminator, i. e., set of . Additionally, in one-many setting, each frequency-dependent generator is trained with all frequency-dependent discriminators, i. e., set of trained in adversarial setting.
Experiment
We used TIMIT and Wall Street Journal (WSJ) corporas to evaluate the performance of proposed model on domain adaptation. TIMIT dataset contains broadband kHz recordings of phonetically-balanced read speech of utterances ( hours). Male/Female ratio of speakers across train/validation/test sets are approximately % to %. WSJ contains hours of standard si284/dev93/eval92 for train/validation/test sets, with equally distributed genders.
The spectrogram representation of audio is used for training the CycleGAN and ASR models, which is computed with a ms window and ms step size. Each spectrogram is normalized to have zero mean and unit variance. To implement MD-CycleGAN, three non-overlapping frequency bands are defined, i. e. with $(\text{C, F, T, SF, ST})44410241024$ and final output layer.
In this section, ASR model is employed to evaluate the performance of proposed model, where domains are different genders. First, gender generators , are trained on gender-separated train set. These generators are then evaluated for traintest and testtrain adaptation using ASR model. In former, ASR model is retrained on the adapted train set, while in latter, a more applicable case, ASR model is fixed and evaluated on the new adapted test sets.
Results on adapting TIMIT train set are shown in Table 2 and 3. As ablation study to CycleGAN-VC , performance is significantly improved with three discriminator compared to single one. Compared to FHVAE , phoneme error rate is improved by in Table 3. To evaluate the generalization of the generators, we used them on WSJ dataset without retraining. As shown in Table 4, ASR performance is significantly improved by reducing the gap to the corresponding male and female baselines. For a fair comparison, ASR performance trained on WSJ train set is WER. It is worth mentioning that relatively lower performance on TIMIT is due to smaller size of dataset.
1.2 Test→→\rightarrowTrain Adaptation
Test set adaptation of TIMIT and WSJ are shown in Table 5 and 6. It is clear that using the proposed model, ASR performance is significantly improved by adapting testtrain , compared to original CycleGAN. Qualitative assessment of ASR predictions are shown in Tables 1 and Appendix A.
2 Qualitative Evaluation
In this section, the quality of generated spectrogram for malefemale adaptation is assessed. The characteristic difference between male and female spectrograms is the variation rate of frequency for a fixed time window. As shown in Figure 1, top row depicts the original spectrograms, where male is characterized by smooth frequency variation, opposed to peaky and high-rate variations of female. Well-trained generators should catch these inter-domain variations. As ablation study, we are also showing the generated spectrogram by CycleGAN (One-D CycleGAN), in middle row, comparing with Three-D CycleGAN in bottom row. One-D CycleGAN learns to convert the spectrogram only using a pretrained discriminator. It is noticeable that the converted spectrogram in One-D CycleGAN fails to match the target domain characteristics, at some frequency bands, and simply copied the source spectrogram. However, with no pretraining of Three-D CycleGAN, it learns a better mapping function, by either suitably smoothing the spectrogram (femalemale), or generating peaky variations (malefemale). The checkerboard artifacts is a common problem in deconvolution-based generators. This problem is visible in One-D CycleGAN, with discontinuous artifacts through time and frequency dimensions, which results in a noisy and unnatural-sounding audio. This problem is mitigated in Three-D CycleGAN, by learning the target domain characteristics using multiple independent discriminators.
Conclusion and Future Directions
In this paper, a new cyclic consistent generative adversarial network based on multiple discriminators is proposed (MD-CycleGAN) for unsupervised non-parallel speech domain adaptation. Based on the frequency variation of spectrogram between domains, the multiple discriminators enabled MD-CycleGAN to learn an appropriate mapping functions that catch the frequency variations between domains. The performance of MD-CycleGAN is measured by ASR prediction, when train and test set are sampled from different genders. It was shown that MD-CycleGAN can improve the ASR performance on unseen domains. As future extension, this model will be evaluated on datasets adaptation, e. g. TIMITWSJ, and accent, e. g. AmericanIndian adaptations.
References
Appendix A ASR Prediction Assessment
In this section, qualitative assessment of ASR predictions are shown, when test set is sampled from the opposite gender, and when MD-CycleGAN model is used to adapt the test set distribution to the train set. Tables 7, 8 show the selected ASR predictions when adapting femalemale, whereas Tables 9, 10 show adaptation for malefemale. Note that MD-CycleGAN is trained on TIMIT, and used for adaptation on WSJ dataset.